Abstract
ABSTRACT Microbiome sequencing datasets are sparse, high-dimensional, compositional, and hierarchically structured. Predictive modelling from these data typically relies on ad hoc choices of feature representation, obscuring their impact on performance and biological interpretation. A standardized, compute-efficient framework is needed to jointly optimize microbial feature representation and model algorithms with transparent model evaluation. Here, we present ritme , an open-source software package implementing Combined Algorithm Selection and Hyperparameter Optimization tailored to microbial sequencing data. ritme systematically explores feature engineering methods — taxonomic aggregation, sparsity-aware selection, compositional transforms, and metadata enrichment — alongside diverse model classes using state-of-the-art optimizers and model trackers. Applied to three real-world use cases, ritme outperforms original study pipelines by 7-29% on the primary task metric — and surpasses a generic AutoML baseline across all three use cases — while using substantially fewer features, supporting model interpretability and downstream biological inspection. It further provides users with insights into how feature and model choices drive predictive performance. Together, these results establish ritme as a standardized framework for identifying optimal feature-model combinations from high-throughput sequencing data. ritme is an open-source Python package available at https://github.com/adamovanja/ritme . IMPORTANCE Predictive modelling from microbiome sequencing data is challenged by the sparse, high-dimensional, compositional, and hierarchical nature of these data, which existing automated machine learning tools do not address. Choices about how to summarize and transform these data are therefore made on a per-study basis, affecting both predictions and biological interpretation. We present ritme , the first open-source framework that jointly optimizes microbiome-specific feature representations — including taxonomic aggregation, sparsity-aware selection, compositional transforms, and metadata enrichment — together with predictive model class and hyperparameters. The search builds on state-of-the-art optimizers and scales across compute clusters. Across three real-world datasets, ritme outperformed the original studies and a generic AutoML baseline while exposing how feature and model choices shape predictive performance. By replacing ad hoc decisions with systematic optimization, ritme delivers more accurate, more parsimonious, and reproducible predictive models that can serve as a starting point for downstream biological investigation.
Full text
1,526 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
Microbiome sequencing datasets are sparse, high-dimensional, compositional, and hierarchically structured. Predictive modelling from these data typically relies on ad hoc choices of feature representation, obscuring their impact on performance and biological interpretation. A standardized, compute-efficient framework is needed to jointly optimize microbial feature representation and model algorithms with transparent model evaluation. Here, we present ritme, an open-source software package implementing Combined Algorithm Selection and Hyperparameter Optimization tailored to microbial sequencing data. ritme systematically explores feature engineering methods — taxonomic aggregation, sparsity-aware selection, compositional transforms, and metadata enrichment — alongside diverse model classes using state-of-the-art optimizers and model trackers. Applied to three real-world use cases, ritme outperforms original study pipelines and generic AutoML baselines. It further provides users with insights into how feature and model choices drive predictive performance. Together, these results establish ritme as a standardized framework for identifying optimal feature-model combinations from high-throughput sequencing data. ritme is an open-source Python package available at https://github.com/adamovanja/ritme.
Competing Interest Statement
The authors have declared no competing interest.
Footnotes
Figure 6 rasterised to enable quicker display in PDF; aded GitHub URL in abstract, added keywords and key points
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.