Target-driven optimization of feature representation and model selection for microbiome sequencing data with ritme

preprint OA: closed
Full text JSON View at publisher

Abstract

ABSTRACT Microbiome sequencing datasets are sparse, high-dimensional, compositional, and hierarchically structured. Predictive modelling from these data typically relies on ad hoc choices of feature representation, obscuring their impact on performance and biological interpretation. A standardized, compute-efficient framework is needed to jointly optimize microbial feature representation and model algorithms with transparent model evaluation. Here, we present ritme , an open-source software package implementing Combined Algorithm Selection and Hyperparameter Optimization tailored to microbial sequencing data. ritme systematically explores feature engineering methods — taxonomic aggregation, sparsity-aware selection, compositional transforms, and metadata enrichment — alongside diverse model classes using state-of-the-art optimizers and model trackers. Applied to three real-world use cases, ritme outperforms original study pipelines by 7-29% on the primary task metric — and surpasses a generic AutoML baseline across all three use cases — while using substantially fewer features, supporting model interpretability and downstream biological inspection. It further provides users with insights into how feature and model choices drive predictive performance. Together, these results establish ritme as a standardized framework for identifying optimal feature-model combinations from high-throughput sequencing data. ritme is an open-source Python package available at https://github.com/adamovanja/ritme . IMPORTANCE Predictive modelling from microbiome sequencing data is challenged by the sparse, high-dimensional, compositional, and hierarchical nature of these data, which existing automated machine learning tools do not address. Choices about how to summarize and transform these data are therefore made on a per-study basis, affecting both predictions and biological interpretation. We present ritme , the first open-source framework that jointly optimizes microbiome-specific feature representations — including taxonomic aggregation, sparsity-aware selection, compositional transforms, and metadata enrichment — together with predictive model class and hyperparameters. The search builds on state-of-the-art optimizers and scales across compute clusters. Across three real-world datasets, ritme outperformed the original studies and a generic AutoML baseline while exposing how feature and model choices shape predictive performance. By replacing ad hoc decisions with systematic optimization, ritme delivers more accurate, more parsimonious, and reproducible predictive models that can serve as a starting point for downstream biological investigation.
Full text 1,526 characters · extracted from oa-doi-fallback · click to expand
Abstract Microbiome sequencing datasets are sparse, high-dimensional, compositional, and hierarchically structured. Predictive modelling from these data typically relies on ad hoc choices of feature representation, obscuring their impact on performance and biological interpretation. A standardized, compute-efficient framework is needed to jointly optimize microbial feature representation and model algorithms with transparent model evaluation. Here, we present ritme, an open-source software package implementing Combined Algorithm Selection and Hyperparameter Optimization tailored to microbial sequencing data. ritme systematically explores feature engineering methods — taxonomic aggregation, sparsity-aware selection, compositional transforms, and metadata enrichment — alongside diverse model classes using state-of-the-art optimizers and model trackers. Applied to three real-world use cases, ritme outperforms original study pipelines and generic AutoML baselines. It further provides users with insights into how feature and model choices drive predictive performance. Together, these results establish ritme as a standardized framework for identifying optimal feature-model combinations from high-throughput sequencing data. ritme is an open-source Python package available at https://github.com/adamovanja/ritme. Competing Interest Statement The authors have declared no competing interest. Footnotes Figure 6 rasterised to enable quicker display in PDF; aded GitHub URL in abstract, added keywords and key points

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00