Phenotype-driven parallel embedding for microbiome multi-omic data integration
preprint
OA: closed
AI-generated summary
PAPRICA is a parallel autoencoder framework that embeds microbiome multi-omic data into separate latent spaces while aligning them and a phenotype space to enable cross-omic inference and improve phenotype prediction.
One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works
Abstract
The human microbiome is widely recognized as a key determinant of health and disease, yet most reported links between observed microbial features and clinical outcomes remain descriptive and lack an integrated system-level perspective. Multi-omic studies of the microbiome, which jointly profile and analyze multiple molecular aspects of the microbiome via metagenomics, metabolomics, proteomics, and transcriptomics assays, offers a more comprehensive view of this system, with the potential to uncover how microbial communities and functions influence host physiology. However, integration of such multi-omic data remains challenging due to high dimensionality, major differences in data properties across omics, and the need to utilize and preserve omic-specific information. Embedding omic data in low-dimensional spaces offer a promising avenue to capture complex patterns, reduce noise, and improve downstream analysis, yet most embedding-based microbiome studies to date exhibited limited predictive power or relied on a single joint embedding of all omics thus failing to preserve omic-species properties. To address this, we introduce PAPRICA (Phenotype-Aware Parallel Representation for Integrative omiC Analysis), an encoder-decoder framework for microbiome multi-omic integration that embeds each omic into its own latent space while jointly modeling their relationships. The model consists of parallel autoencoders trained with a loss function that promotes three objectives: (1) accurate reconstruction of each omic, (2) alignment of samples across omics such that proximity in one latent space reflects proximity in the others, and (3) alignment with a phenotype space to capture variation associated with continuous outcomes, such as fecal calprotectin levels in IBD. The resulting models support cross-omic inference and phenotype prediction from the learned latent representations, and enables integration without collapsing data into a single space. This modeling approach thus preserves omic-specific signals while capturing phenotype-associated variation. We compared PAPRICA to four alternative models that represent successive advances in multi-omic integration architectures. We found that across two complementary tasks, predicting one omic profile from another and predicting a continuous phenotype from an input omic profile, our parallel autoencoder approach, and particularly the PAPRICA model, demonstrated better performance across three multi-omic datasets (the Franzosa IBD cohort, Lifelines DEEP and the Dog Aging Project Precision Cohort). Combined, these findings suggest that our embedding strategy effectively captures and balances omic-specific structure, cross-omic relationships, and phenotype-relevant signals across diverse datasets, offering a flexible, scalable framework for embedding complex multi-omic microbiome data and advancing our ability to gain new insights into host-microbiome interactions.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00