Abstract
Background Microbiome data, like other high-throughput data, suffer from technical heterogeneity stemming from differential experimental designs and processing. In addition to measured artifacts such as batch effects, there is heterogeneity due to unknown or unmeasured factors, which lead to spurious conclusions if unaccounted for. With the advent of large-scale multi-center microbiome studies and the increasing availability of public datasets, this issue becomes more pronounced. Current approaches for addressing unmeasured heterogeneity in high-throughput data were developed for microarray and/or RNA sequencing data. They cannot accommodate the unique characteristics of microbiome data such as sparsity and over-dispersion.
Results
Here, we introduce Quantile Thresholding (QuanT), a novel non-parametric approach for identifying unmeasured heterogeneity tailored to microbiome data. QuanT applies quantile regression across multiple quantile levels to threshold the microbiome abundance data and uncovers latent heterogeneity using thresholded binary residual matrices. We validated QuanT using both synthetic and real microbiome datasets, demonstrating its superiority in capturing and mitigating heterogeneity and improving the accuracy of downstream analyses, such as prediction analysis, differential abundance tests, and community-level diversity evaluations.
Conclusions
We present QuanT, a novel tool for comprehensive identification of unmeasured heterogeneity in microbiome data. QuanT’s distinct non-parametric method markedly enhances downstream analyses, serving as a valuable tool for data integration and comprehensive analysis in microbiome research.
Competing Interest Statement
The authors have declared no competing interest.
Footnotes
jiuyaolu{at}wharton.upenn.edu
gsatten{at}emory.edu
ktmeyer{at}email.unc.edu
lenore_launer{at}nih.gov
wol4002{at}med.cornell.edu
Updated the algorithm and the results.
List of abbreviations
- QuanT
- quantile thresholding
- SVA
- surrogate variable analysis
- RUV
- remove unwanted variation
- SVD
- singular value decomposition
- iSV
- iterative singular vectors
- QSV
- quantile surrogate variables
- VOI
- variable of interest
- CARDIA
- Coronary Artery Risk Development in Young Adults
- CVD
- cardiovascular disease
- HBP
- high blood pressure
- TG
- true group
- NC
- no correction
- ROC-AUC
- area under the receiver operating characteristic curve
- PCoA
- principal coordinate analysis
- HIVRC
- HIV re-analysis consortium
- CRC
- colorectal cancer
- IBDMDB
- Inflammatory Bowel Disease Multi-omics Database
- MOMS-PI
- Multi-Omic Microbiome Study: Pregnancy Initiative
- DM
- Dirichlet-Multinomial
- MIDASim
- Microbiome Data Simulator
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.