StrataBionn: a neural network supervised classification method for microbial communities

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

The classification of microbial communities into discrete states or “community state types” (CSTs) is fundamental to understanding host-microbiome interactions and their clinical implications. Traditional methods, such as the nearest-neighbor approaches, often struggle with the inherent noise, high dimensionality, and non-linear signatures of taxonomic profiles. We present a novel supervised framework for microbial community classification, leveraging an Artificial Neural Network (ANN) architecture implemented in a new tool we named StrataBionn. We rigorously evaluated our approach using large-scale vaginal microbiome datasets, directly benchmarking performance against VALENCIA and a Random Forest (RF) classifier. To demonstrate the versatility of our models, we further extended the framework to oral microbiome classification, assessing its stability across diverse anatomical sites. Our supervised models consistently outperformed the nearest-neighbor approach across all evaluated datasets. In the vaginal microbiome, our method achieved an 11.6% to 13.3% increase in performance across all primary metrics, including precision, recall, accuracy, and F1-score. Furthermore, we demonstrate that this performance advantage is maintained in the oral microbiome, highlighting the generalizability of our neural network and ensemble strategies to various microbial ecosystems without the need for niche-specific algorithmic adjustments. By capturing complex feature dependencies that distance-based methods overlook, our approach provides a more robust and accurate census of microbial community structures. StrataBionn’s ability to learn classification schemes for any microbiome with high accuracy and explainability, through the use of provided utilities to visualize feature-space classification boundaries and perform perturbation analysis on trained classifiers, makes it ideal for broad application in microecology research. This framework offers a scalable, high-performance alternative for microbiome researchers, facilitating more precise clinical stratification and biological insights across hosts body sites.
Full text 90,960 characters · extracted from oa-pdf · 11 sections · click to expand

Keywords

microbiome, microbiome classification, artificial neural network 16 17 18 19 20 21 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 2

Abstract

22 The classification of microbial communities into discrete states or "community state types" 23 (CSTs) is fundamental to understanding host-microbiome interactions and their clinical 24 implications. Traditional methods, such as the nearest-neighbor approaches, often struggle with 25 the inherent noise, high dimensionality, and non-linear signatures of taxonomic profiles. We 26 present a novel supervised framework for microbial community classification, leveraging an 27 Artificial Neural Network (ANN) architecture implemented in a new tool we named StrataBionn. 28 We rigorously evaluated our approach using large-scale vaginal microbiome datasets, directly 29 benchmarking performance against VALENCIA and a Random Forest (RF) classifier. To 30 demonstrate the versatility of our models, we further extended the framework to oral microbiome 31 classification, assessing its stability across diverse anatomical sites. Our supervised models 32 consistently outperformed the nearest-neighbor approach across all evaluated datasets. In the 33 vaginal microbiome, our method achieved an 11.6% to 13.3% increase in performance across all 34 primary metrics, including precision, recall, accuracy, and F1-score. Furthermore, we 35 demonstrate that this performance advantage is maintained in the oral microbiome, highlighting 36 the generalizability of our neural network and ensemble strategies to various microbial 37 ecosystems without the need for niche-specific algorithmic adjustments. By capturing complex 38 feature dependencies that distance-based methods overlook, our approach provides a more robust 39 and accurate census of microbial community structures. StrataBionn’s ability to learn 40 classification schemes for any microbiome with high accuracy and explainability, through the 41 use of provided utilities to visualize feature-space classification boundaries and perform 42 perturbation analysis on trained classifiers, makes it ideal for broad application in microecology 43 research. This framework offers a scalable, high-performance alternative for microbiome 44 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 3 researchers, facilitating more precise clinical stratification and biological insights across hosts 45 body sites. 46

Introduction

47 The study of microbial communities has fundamentally reshaped our understanding of human 48 health, disease, and biological fitness [1–14]. It has progressed from checking for the presence of 49 a specific taxa to tracing the shifts of community-level microbiome structures, defined by 50 relative species abundance and functional capabilities, which more accurately describes 51 phenotypic changes and disease risk in hosts [3,15–21]. Due to this paradigm shift, the rate at 52 which microbiome data is generated has outpaced the ability of analysis tools to derive meaning 53 from it. This has intensified the need for accurate, reproducible classification methods that can 54 categorize microbiome communities to identify new biomarkers across diverse populations. 55 The rapid maturation of high-throughput sequencing has shown that microbial communities 56 exhibit recurring patterns in composition across individuals, offering unprecedented 57 opportunities to map these microbial landscapes [22–24]. However, this "big data" era presents 58 significant analytical challenges. Extracting meaningful biological insights requires classification 59 frameworks that are computationally efficient and capable of high-level composition analysis 60 required in microbial ecology. Early research established the utility of composition-based 61 community level classifications, leading to the development of standardized classification 62 systems. Notable examples of such classifications include gut "enterotypes" [25,26] and vaginal 63 "community state types" (CSTs) [27–29], which are consistent assemblages of microbial species, 64 foundational in guiding modern comparative microbiome research and clinical diagnostics. 65 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 4 Despite their utility, the methodologies used to define these categories remain a bottleneck. 66 Historically, hierarchical clustering (HC) was the gold standard, yet its limitations in scalability 67 and reproducibility for rapidly growing datasets are well-documented [30–36]. To address this, 68 researchers have turned to supervised approaches like the nearest centroid classifier (NCC), 69 implemented in robust and easy-to-use tools such as VALENCIA, which are deterministic and 70 scale well [37]. However, NCC methods assume well-separated data and equal covariances; 71 consequently, they often struggle with the "fuzzy", non-linear edges and overlapping 72 distributions typical of complex biological datasets, such as those found in microbiome studies 73 [38,39]. Furthermore, NCC models fail to account for interactions between predictor variables—74 a critical oversight in microbiome data, where inter-species interactions define the community 75 structure [40]. Finally, NCC performance decays significantly when class variances are unequal 76 or when new data points fall outside the centroid-defined convex hulls [41,42]. 77 Here, we present StrataBionn, a novel neural network-based classification algorithm designed 78 to handle the non-linearities and high dimensionality of microbiome data. Unlike static 79 classifiers, StrataBionn uses a neural network architecture that can be trained, saved, and adapted 80 to a wide variety of microbial datasets. By integrating automated training methods and data 81 preprocessing, the model achieves superior generalization and consistency. While StrataBionn 82 offers a more sophisticated parameterization than current NCC tools, we provide comprehensive 83 guidelines to streamline the fine-tuning process for diverse research applications. 84 We benchmarked StrataBionn on microbial communities from two distinct human niches: the 85 vaginal and oral microbiomes. The vaginal microbiome is characterized by well-defined CSTs, 86 where four types (CST-I, II, III, and V) are dominated by specific Lactobacillus species, and one 87 (CST-IV) is defined by a diverse, anaerobic composition [43]. Using this established framework, 88 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 5 we show that StrataBionn achieves higher precision and recall than existing methods. 89 Furthermore, we applied our method to the oral microbiome, where the challenge lies in 90 distinguishing the "oral core" from the subtle deviations associated with periodontal and 91 systemic diseases. Finally, we show its usefulness in the assignment of samples from a novel oral 92 sample set generated for this study. 93 In summary, our results indicate that StrataBionn is a highly adaptable tool that is capable of 94 high-performance classification even with smaller reference datasets. While showcased here 95 using vaginal and oral communities, this method is intended for broad application across any 96 microbiome study, including comparisons between healthy and diseased cohorts. We anticipate 97 that StrataBionn will facilitate the characterization of large-scale metagenomic datasets, 98 ultimately deepening our understanding of the microbial drivers of host health. 99

Methods

100 Data Sources and Collection 101 Publicly Sourced Datasets: To train and evaluate StrataBionn, we utilized several previously 102 published microbiota composition datasets. Vaginal microbiome profiles from France et al. [37] 103 were used to train and validate the model against established Community State Type (CST) 104 classifications, which we then validated against a dataset from Hickey et al. [44] to demonstrate 105 performance on an independent dataset. Oral microbiome composition datasets were employed 106 to evaluate the framework’s utility as a generalized classifier. As the oral microbiome research 107 community has not established a common reference for the community types found in the oral 108 cavity, we generated a de novo naive classification with a large dataset recently generated by 109 Manghi et al. [45]. We then used the naive classification as “true” labels for the training and 110 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 6 validation of our method. We show that this approach produces similar results, in terms of 111 performance metrics, to those obtained in a more curated dataset like the vaginal microbiome. 112 We show how the classifier can be used to classify samples on a sample set generated for this 113 work and a dataset from published studies [46]. A freeze of the public datasets used for the study 114 are available in the github repository of the tool. 115 In-House Oral Microbiome Collection and metagenomic sequencing: To test the 116 classification consistency in the oral microbiome, plaque samples were collected from adult 117 refugees (≥18 years) at the Rwamwanja UN refugee settlement, Uganda. Ethics and Recruitment: 118 Collection was approved by the Institutional Review Board at Washington State University (IRB 119 #15196-002). A total of 54 participants were recruited using a random number sequence 120 generated in R to ensure unbiased selection. Informed consent was obtained from all participants 121 and a translator was present when necessary. Sampling and Storage: Dental plaque was collected 122 during routine cleanings, preserved in RNAlater, and stored at -80°C after transport to Pullman, 123 WA, under CDC Import Permit #2016-03-212. Library Preparation and Sequencing: Microbial 124 DNA was extracted using the MoBio Powerlizer™ DNA Isolation Kit (Mo Bio Laboratories, 125 Carlsbad, CA, USA), following the manufacturer’s protocol. For this, plaque samples preserved 126 in RNAlater were spun down, RNA later discarded and macerated manually with sterile pestles 127 with lysis solution. After DNA extraction, samples were sheared using a Covaris M220 to a 550 128 bp insert size. Libraries were prepared with the NEBNext Ultra DNA Library Prep kit (New 129 England Biolabs Inc.) and sequenced on a single lane of an Illumina HiSeq 2500 at the 130 Washington State University Genomics Core. 131 Bioinformatics and Quality Control of novel oral metagenomic data: Initial quality 132 assessment was performed using FastQC [47]. Out of 54 samples, 39 produced enough DNA 133 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 7 after extraction and/or passed FastQC’s quality control test, and were used in analysis. Sequence 134 trimming was performed with TrimGalore and Cutadapt and we hard trimmed the first 11 bp on 135 the 5’ end and soft trim for Phred scores < 25 [48,49]. Host Filtering: reads were aligned to the 136 human reference genome (GRCh38) using STAR [50], and human reads were removed and 137 deleted permanently. Taxonomic Classification: non-human reads were classified using 138 KrakenUniq [51], and to minimize false positives, we required ≥100 unique kmers and a 139 duplicity ratio ≤ 3. Count matrices were generated at the genus and species levels. Raw sequence 140 data is available at SRA-NCBI through accession PRJNA1445365. 141 Microbiome classifier: the StrataBionn Framework 142 We developed StrataBionn, stratification of biological data using neural networks, a supervised 143 classification tool designed to assign CST, or other community level labels, to microbiome data. 144 It accepts as an input a standardized data format (CSV) similar to that used by VALENCIA, 145 which includes columns for sample ID, read counts, labels, and a column for each taxa for which 146 count data was collected. The method incorporates a stratified data partitioning strategy to ensure 147 robust model training and evaluation, particularly in scenarios with varying data availability. 148 Stratified Data Partitioning: To maintain the inherent class balance of the original labeled data 149 within each subset and mitigate potential biases, a stratified splitting approach was employed. 150 This ensured that the amount of each CST in each subset used during training is consistent with 151 the distribution found in the original, unpartitioned dataset, to avoid sampling biases during 152 model training. Datasets were first shuffled and then partitioned proportionally by CST label. 153 Specifically, for each CST, the designated proportions of samples were allocated to the training, 154 testing, and validation sets. Following the split, the proportional distribution of CST labels within 155 each subset was recalculated and compared to the original dataset distribution. Consistency was 156 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 8 automatically assumed if the proportional representation of each CST in the resulting subsets 157 was within a predefined tolerance of 0.1% of the original distribution. We defined tolerance 158 using the equation ( !!"# !!"$%& − 1) ∗ 100, where 𝑝"#$ is the proportion of a given CST in the subset 159 of interest (training, test, validation), and 𝑝"#!%& is the proportion of that same CST in the 160 superset which contains all the training, testing, and validation samples. This rigorous 161 stratification ensures that each data subset provides a representative snapshot of the overall class 162 distribution, facilitating unbiased model training and evaluation. Two data partitioning schemes 163 were implemented and evaluated: (i) a 80/10/10 split for training, testing, and validation sets, 164 respectively, simulating conditions with ample data; and (ii) a 60/20/20 split for training, testing 165 and validation of the same sets, mimicking data-limited scenarios. The processed and partitioned 166 datasets were then exported as comma-separated value (CSV) files for subsequent model 167 training. These files were formatted in the same manner as those accepted by VALENCIA, 168 containing columns for the sample id, total number of reads, sample label (when training), and 169 bacteria species. This format was chosen to increase compatibility between StrataBionn and 170 other existing tools (i.e. VALENCIA). 171 Data Preprocessing within StrataBionn: Prior to model training, StrataBionn performs several 172 crucial preprocessing steps. To account for variations in sequencing depth or sample "quality," 173 the raw feature counts were normalized by the total counts per sample. This transformation 174 ensures that samples with differing sequencing depths contribute equally to the model training 175 process. Subsequently, any null or zero values, which could lead to computational errors during 176 training, were imputed with a small, non-interfering numerical value. While the StrataBionn 177 framework includes the option for additional data normalization techniques, initial evaluations 178 indicated that the total count normalization alone yielded the most effective model performance 179 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 9 for this specific application, and thus, further normalization steps were not performed in the final 180 model training and evaluation. 181 Classification Algorithms: Two distinct supervised learning algorithms were implemented and 182 evaluated within the StrataBionn framework: a Random Forest Classifier (RFC) and an Artificial 183 Neural Network (ANN). These methods were selected for their classification consistency and 184 evaluated along with the VALENCIA classifier which served as an accuracy baseline. 185 ● Random Forest Classifier: For this model, we used the RandomForestClassifier module 186 from the python scikit-learn library [52]. The RFC was parameterized to optimize its 187 performance. The number of estimators (n_estimators) and the number of features to 188 consider when looking for the best split (max_features) were systematically tuned. 189 Numbers of estimators ranging from 100 to 10,000 were tested, and while training time 190 increased significantly with more estimators, performance plateaued around 95.1% 191 (10,000 estimators), taking around 7 minutes. 192 193 ● Artificial Neural Network: The ANN model consisted of an input layer, a single hidden 194 layer, a Leaky Rectified Linear Unit (Leaky ReLU) activation function following the 195 hidden layer, and a dropout layer for regularization. The size of the hidden layer was 196 parameterized, with a default setting of 2 3 ∗ 𝑛'(!#) + 𝑛*#)!#) , where 𝑛'(!#) represents the 197 number of input neurons and 𝑛*#)!#) the number of output neurons. The equation for the 198 full system is as follows: 𝑊2(𝐿𝑅𝑒𝐿𝑈(𝑊1𝑥 + 𝑏1) ∘ + 1,!) + 𝑏1, where 𝑥 is the input data 199 vector, 𝑊1, 𝑊2 are learnable weight matrices, 𝑏1, 𝑏2 are learnable bias matrices, 𝑝 is the 200 probability that a neuron will dropout, and 𝑚 is a vector that masks neurons to be 201 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 10 dropped via multiplication with 0 or 1. The optimization algorithm was selectable 202 between Adaptive Moment Estimation (ADAM) and Stochastic Gradient Descent (SGD). 203 The SGD optimizer offered an optional class weighting scheme to address potential 204 imbalances in CST prevalence, allowing for increased emphasis on accurately classifying 205 less frequent classes. The learning rate and a learning rate reduction factor were also 206 tunable hyperparameters. A probability of assignment is reported for all CSTs for each 207 sample, and the CST with the highest probability is assigned as the label for each sample. 208 The model was constructed using the pytorch library and written python [53]. Training 209 time for the vaginal microbiome test set with default hyperparameters took around 27 210 seconds. 211 212 ● VALENCIA Baseline: The previously established VALENCIA method was included as 213 a baseline for performance comparison. It uses a nearest centroid classification model, 214 which employs pre-provided centroids constructed to classify vaginal microbiome 215 samples. 216 217 We also provide a naive assessment of the uncertainty in the classification by extracting the 218 logits from the softmax assignment and standardizing those values to generate probabilities of 219 samples belonging to the assigned class. For this we exponentiated the raw logits and divide 220 them by the sum of all exponentiated values according to the following formula, as implemented 221 in the pytorch library [53]: 222 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 11 223 Evaluation and Validation Strategy 224 Model Evaluation: All trained models, including the baseline VALENCIA classifier, were 225 evaluated using the same vaginal microbiome dataset. Critically, each proposed model was 226 evaluated on a distinct test set that was not used during training or validation, ensuring an 227 unbiased assessment of generalization performance. The VALENCIA baseline was evaluated 228 using its provided centroids, which were constructed based on the entire dataset. Model 229 performance was quantified using classification accuracy, defined as the proportion of correctly 230 classified samples. 231 Evaluation as General Microbiome Classifier: Performance of the ANN on oral microbiome 232 datasets was measured to evaluate its usefulness as a generalized microbiome community 233 classification method. We first performed a naive classification of the large oral microbiome 234 dataset using K-Means clustering to generate labels that would be employed for model training 235 and evaluation. Oral data was reformatted and treated identically to the vaginal microbiome data, 236 using the same processing methods previously described. Evaluation of the ANN on oral 237 microbiome data was performed using the same classification accuracy measurement used for 238 the vaginal microbiome. The Random Forest Classifier was not evaluated due to its much longer 239 training time to achieve a lower accuracy compared to the ANN, and VALENCIA was not 240 evaluated due to its specificity to the vaginal microbiome. We analyzed the performance of the 241

Method

with the oral microbiome. 242 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 12 Application of Trained Classifiers Across Studies: A vaginal microbiome test set containing 243 population data on different bacteria species [44] was classified using a model trained to identify 244 the CSTs from the VALENCIA training set. StrataBionn was first trained on the VALENCIA 245 training set while limited to considering only the set of bacteria species common to both the 246 training dataset and the target dataset. The trained classifier was then used to apply CST labels to 247 all samples in the target dataset, ignoring bacteria species which are not part of the common set. 248 These CST classifications were then manually examined and determined to be reasonable based 249 on their bacterial makeup using the findings of France et al [37]. 250

Results

251 Training, Validation and Test datasets 252 For the vaginal Microbiome, we generated stratified random sets from the France et al [37], for 253 training (80%: 10,580; 60%: 7,935), validation (80%: 1,334; 60%: 2,655), and testing (80%: 254 1,317; 60%: 2,641). For the oral microbiome, we generated stratified random sets from the 255 Manghi et al. [45] study for training (80%: 6,248; 60%: 4,686), validation (80%: 784; 60%: 256 1,565), and testing (80%: 780; 60%: 1,561). 257 We obtained read-counts supporting the presence of bacteria species in 39 newly generated oral 258 microbiomes for this study. The number of reads and estimated relative abundances for species 259 can be found in the github repository. 260 Preprocessing of Vaginal Microbiome Datasets 261 During the development and testing of StrataBionn, two training data sets with 80% and 60% 262 training allocations were included to test scenarios when training data is abundant or limited. 263 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 13 Stratified datasets from France et al. [37] containing curated CST classifications were used for 264 their high quality labels and large number of samples (13,231). The PACMAP [54] 265 representations of the truth labels for the high training data availability scenario are shown in 266 Figure 1, demonstrating similar CST composition among subsets. Consistent with our 267 expectations, more closely related CST communities are clustered closely in the PACMAP 268 representation (Figure 1). For example, CST-IA and CST-IB, which both feature a high 269 abundance of Lactobacillus crispatus, have a similar projection in the map and tend to occur 270 more closely than less related state types, such as CST-IA and CST-V. This observation supports 271 our use of PACMAP projections of the data for the purpose of representation. 272 The worst CST distribution tolerances across all CSTs are 0.005%, 7.2%, and 7.2% for the 80% 273 training (Figure 1A), 10% testing (Figure 1B), and 10% validation (Figure 1C) sets respectively. 274 For the 60% training, 20% testing, 20% validation sets, the tolerances are 0.05%, 13.4%, and 275 13.4% respectively. PACMAP representation of the 60%, 20%, 20% configuration is presented 276 in the supplementary materials (Supplementary Figure 1). 277 Evaluation of Classification Models 278 Each classification method was evaluated on both the high and low training data availability 279 labeled datasets using accuracy (Figure 2A), recall (Figure 2B), F1 score (Figure 2C), and 280 precision (Figure 2D) as metrics. These metrics were selected to measure the number of samples 281 correctly identified (accuracy), the rate of false negatives (recall), the rate of false positives 282 (precision), and balanced measure of both recall and accuracy (F1). Evaluation was performed 283 on the validation dataset after training a classifier on the corresponding training dataset for all 284 classifiers except for VALENCIA, where the original author provided centroids were used for 285 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 14 classification [37]. While the ANN and RF classifiers trained on a subset of the Vaginal 286 microbiome data, VALENCIA’s centroids were generated using the full microbiome dataset as 287 described in the author’s methodology [37]. For the ANN classifier, probabilities of assignment 288 were above 87% for 95% of samples (80% training: 87.8%; 60% training: 87.6%), and around 289 95% for 80% of classified samples (80% training: 95.3%; 60% training: 94.9%). In all cases the 290 neural network classifier and random forest classifier outperformed VALENCIA when 291 classifying both datasets. The ANN achieved a 11.6-13.3% increase compared to VALENCIA 292 across all metrics (Figure 2A-D), performing slightly better with higher data availability, while 293 the Random Forest approach achieved a 9-10.8% increase across all metrics compared to 294 VALENCIA (Figure 2A-D). The Random Forest classifier performed slightly better in lower 295 data availability scenarios, but never outperformed the ANN classifier by any measured metric. 296 297 To understand the conditions which led to our models' increased performance, we generated 298 confusion matrices for each classification method by comparing the validation dataset truth 299 values to StrataBionn-assigned CST labels. In VALENCIA, misidentification of CST-IA, CST-300 IB, CST-IIIA, CST-IIIB, and CST-IVA disproportionately contributes to its lower performance 301 (Figure 3A). In the case of the ANN classifier and random forest classifier implemented in 302 StrataBionn, communities CST-IA and CST-IIIA disproportionally contribute to the 303 misassignment of the community labels; while CST-IIIB and CST-IVB marginally contribute to 304 the misassignment (Figures 3B-C). To better demonstrate the cases where StrataBionn 305 outperforms VALENCIA in assignments, we estimated the differences in the number of 306 classified communities for each cell in the confusion matrix (Figure 3D). Cases where the 307 predicted label matched the true label were observed to be higher for the neural network 308 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 15 classifier compared to VALENCIA for most CSTs, especially those which are less common 309 (blue diagonals in Figure 3D). Hotspots denoted in red on the off-diagonal of the comparative 310 confusion matrix show that the neural network classifier is better able to differentiate between 311 some CSTs with similar composition distributions than VALENCIA (Figure 3D). This is likely 312 due to the ability to classify along non-linear boundaries which is not possible using a nearest 313 centroid classifier based approach. Due to the greater performance and higher ability to 314 differentiate between similar CSTs, we chose to proceed with a neural network classifier as the 315 basis for StrataBionn. 316 Evaluation of Novel Dataset Classifications 317 To test the ability of StrataBionn to apply accurate classifications to novel datasets, we generated 318 classifications for an unclassified vaginal microbiome dataset containing 458 samples from 319 Hickey et al. [44]. This dataset required no preprocessing other than reformatting and column 320 renaming. The classification model we used for this task was trained on the labeled France et al. 321 [37] dataset, and configured to consider only the bacteria species common to both datasets. We 322 then applied this classifier to the full Hickey et al. [44] dataset to get CST assignments for each 323 sample. We generated two PACMAPs using both the classified Hickey et al. [44] dataset and the 324 truth labels from the France et al. [37] dataset, shown in figure 4. We observed that the novel 325 classifications were assigned labels that matched their surrounding samples with similar 326 compositions. Assigned samples cluster around data points with the same classification from the 327 original France et al. [37] dataset. 328 Classification of Oral Microbiome Datasets 329 An oral microbiome bacterial community composition dataset was used to validate StrataBionn’s 330 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 16 ability as a general purpose classification tool on non-vaginal microbiome datasets. We used 331 7,812 samples collected by Manghi et al. [45] with pre-existing classifications removed due to 332 poor correlation with the data (see supplementary figure 2). As oral microbiomes have not been 333 characterized to the same level of detail of classification as the vaginal microbiomes, we 334 employed a K-Means clustering algorithm to determine the optimal number of clusters for this 335 dataset. We found that using 3 clusters provided the optimal classification by overserving the 336 silhouette and elbow plots provided in supplementary figures 3 and 4. After clustering labels 337 were applied, the data was then partitioned into high and low data training data availability 338 datasets using the stratified partitioning process previously described for the vaginal microbiome. 339 PACMAP visualizations of these subsets are provided in Figure 5. The worst CST distribution of 340 tolerance values for these datasets, as defined by the percent occurrences in a set above (positive) 341 or below (negative) the occurrences found in the reference set, were 0.008%, 0.0074%, and 342 0.0074% for the 80% training (Figure 5A), 10% testing (Figure 5B), and 10% validation (Figure 343 5C) sets respectively. For the 60% training, 20% testing, 20% validation sets, the tolerances were 344 0.019%, 0.35%, and 0.35% respectively. We observed that samples which were assigned to the 345 same cluster exhibit spatial locality, suggesting that our assignments capture the composition of 346 the samples, making them ideal for evaluating the performance of our model (Figure 5). 347 Labels assigned by StrataBionn were evaluated based on F1-score, recall, precision, and 348 accuracy metrics for high and low training data availability scenarios. StrataBionn achieves 98.8-349 98.9% across all metrics with low training data availability, and 99% across all metrics in high 350 training data availability scenarios (Figure 6A). We observed that this high level of accuracy was 351 consistent across all label groups, with a slightly higher level of accuracy for less common CSTs 352 given their size (Figure 6B). The least common CST, CST-2, was correctly identified in all 184 353 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 17 samples which truly belonged to that group, while only being mis-identified in two samples 354 belonging to CST-1. CST-1, the most common subtype, is the one most commonly mis-assigned 355 to a different cluster (7/369, 1.9%), likely due to the complicated boundaries between it and the 356 other CSTs observed in figure 6B. StrataBionn is able to correctly identify these 98.1% of the 357 time, indicating a strong ability to classify along complicated, non-linear boundaries. 358 The in-house generated oral microbiome dataset containing 39 samples, and 47 additional 359 samples from Baker et al [46]. were used to evaluate the ability of StrataBionn to apply 360 classifications to novel non-vaginal microbiome datasets, and required no additional processing 361 other than reformatting and column renaming. This new dataset was assigned labels using a 362 classifier trained on the 80% training data availability training subset for the oral microbiome 363 described above using only the bacteria species common to both datasets. All samples from 364 Baker et al. [46] were classified as CST-1 by StrataBionn, matching the label of the samples with 365 known CSTs with which they were compositionally most similar to (Fig. 7). In the in-house and 366 Baker et al. [46] datasets, we identified 7 belonging to CST-0, and 79 samples in CST-1. No 367 samples in these sets were identified as belonging to CST-2. 368

Discussion

369 There is increasing evidence of the implication of community-level microbial shifts with the 370 increased likelihood of disease, and human health more broadly [19,55–66]. This puts an 371 increased importance on our ability to accurately infer the microbial communities and 372 discriminate those that are commonly associated with healthy outcomes from those commonly 373 associated with disease. The complexity of microbial community data, characterized by high 374 dimensionality, compositional sparsity, and technical noise, requires robust, interpretable, and 375 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 18 transferable classification frameworks. The classification process is generally abstracted away 376 and hidden from the user, concealing useful information regarding how assignments are 377 determined during the classification process which hinders the interpretability of those 378 classifications. Additionally, current classification methods are fundamentally constrained to a 379 single dataset with a fixed set of bacteria species, preventing meaningful classifications derived 380 from one study from being applied generally to additional studies for the same microbiome. To 381 address these issues, we developed StrataBionn, an Artificial Neural Network (ANN)-based tool 382 designed to overcome the limitations of extant classifiers by providing superior accuracy and 383 cross-study applicability. By integrating transparent visualization and perturbation analysis, 384 StrataBionn transitions microbiome "state-typing" from a "black-box" abstraction to an 385 interpretable diagnostic process, where researchers can investigate the compositional causality of 386 the inferred community labels in a community. Although StrataBionn has been developed for 387 microbial communities, it could likely be adapted to study any ecological community data for 388 which abundant samples and training are available due to its community-independent 389 composition based classification architecture. 390 Comparative Performance and Architecture Selection 391 Our results demonstrate that StrataBionn consistently outperforms both the k-nearest neighbor-392 based approach of VALENCIA and a traditional Random Forest (RF) classifier implemented in 393 our package. While RF models showed competitive performance when compared to k-nearest 394 neighbor approaches, the ANN architecture was selected for the basis of StrataBionn due to its 395 superior computational efficiency and accuracy at scale. Notably, StrataBionn achieved a 12-396 13% increase in F1-score and precision over VALENCIA when benchmarked against the 397 vaginal microbiome, a dataset and type of community that VALENCIA was specifically 398 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 19 designed to resolve. The performance gain was most pronounced in the classification of sub-399 CSTs (e.g., CSTs IV-A and IV-C0), which often present overlapping composition profiles that 400 confound less sensitive models. Compared to the Random Forest classifier, StrataBionn achieves 401 an approximately 2% increase across all the same metrics, with the difference being slightly 402 more pronounced with more training data. The performance metric differences between the 60% 403 training dataset (7,935 training samples) and the 80% training dataset (10,580 training samples) 404 were negligible in our testing, but the 80% training dataset did achieve slightly higher 405 performance. More specifically, we can attribute this increase in correct assignments to a higher 406 precision for all sub-CSTs of CST I and III, and higher precision for CSTs IV-A, IV-C0, and V 407 (Fig. 3). This is important because often CST-IVs are associated with increased likelihood of 408 disease, and thus our method can dramatically reduce false positives in disease susceptibility 409 screenings by correctly attributing samples to their true, non-CST-IV type, without degraded 410 recall ability. What is less known is how the observed differences among CST-IV subgroups are 411 differentially associated with disease, making the distinction of those communities relevant for 412 any future study aimed at discovering the compositional and causal contributions of these 413 community differences to disease. 414 The marginal performance increase observed when expanding the training set from 60% to 80% 415 suggests that StrataBionn reaches an early accuracy plateau, indicating high data efficiency. This 416 is particularly important given that many datasets are going to be sample limited, and in those 417 circumstances, the strategy for the partition of the samples for training, testing and validation 418 might differ. While France et al. [37] argued that VALENCIA’s assignments may better reflect 419 biological states than the hierarchical clustering (HC) labels, our objective was to validate the 420 learning fidelity of the model. By using HC labels as a ground-truth proxy, we demonstrated that 421 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 20 StrataBionn possesses a superior capacity to internalize and replicate complex, user-defined 422 classification logic. The ANN learns classification assignments based on their compositions 423 during an iterative training process, where training is continued until accuracy improvements 424 reach a maximum. The result is a high accuracy, deterministic trained model which can apply 425 consistent classifications to novel data. As the training process learns distributions based on user-426 provided labels, it is not inherently tied to any one microbiome, which has been a major 427

Limitation

of previous models. Additionally, by allowing the user to specify a set of bacteria 428 species to consider during this process, disregarding other species, StrataBionn can produce 429 additional trained models which can be applied across separate datasets for meta-analysis. 430 In the PACMAP visualization of the data, it can be seen that some samples assigned to CSTs 431 appear isolated from the general clusters. This visualization can help researchers identify 432 samples with low probability of assignment. In Supplementary Table 1, we provide the 433 probabilities of assignment for a subset of samples from the vaginal microbiome that seem 434 isolated from their clusters. In these samples, the general trend is that they share a low 435 probability of assignment to their clusters. Because our method provides an empirical probability 436 of assignment, researchers can perform a post-assignment filter to declare ambiguities. 437 Interpretability and Compositional Causality 438 A persistent critique of deep learning in biology is the lack of model transparency. We took a 439 novel approach to address this concern by including a classification boundary visualization tool 440 and a perturbation analysis utility, which increases transparency in the classification process. 441 More specifically, StrataBionn addresses the issue of interpretability via two distinct 442 mechanisms: 443 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 21 1. Feature-Space Visualization: By projecting classification boundaries onto two-444 dimensional bacterial abundance coordinate planes, researchers can visually verify the 445 "decision logic" of the model. By allowing the user to select two bacteria for the x and y 446 axis and displaying the classification assignment in a scatter plot, insight as to what 447 bacteria correspond to which CST can be attained. For example, we can inspect which 448 bacteria are correlated with the different sub-CSTs by searching for a bacteria species 449 where all samples appear to vary widely in abundance, appearing distributed across an 450 axis. Then, we can find another species varying widely in abundance and observe the 451 clustering of certain CSTs, displayed in different colors, to determine the features and 452 thresholds which lead to their assignment. Examples of these plots and their 453 interpretations are provided in supplementary figures 5-7. 454 2. Stratified Perturbation Analysis: Stratabionn’s perturbation analysis utility quantifies 455 the sensitivity of the model to specific taxa. By permuting the counts of individual 456 species and measuring the resulting decay in F1-score, StrataBionn identifies which 457 bacteria are computationally "essential" for a specific CST assignment. StrataBionn 458 provides the impact of individual species perturbations for F1, precision, recall, and 459 accuracy metrics, providing additional insight into the impact of single species into the 460 increased sensitivity or specificity of the assignment to communities (supplementary 461 figure 8). 462 We propose that users implement the classification algorithm and use the tools provided by 463 StrataBionn to further investigate the compositional causality of specific community states. We 464 anticipate that StrataBionn will enable researchers to identify emerging properties that cause 465 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 22 microbial communities to differ across host populations in datasets of increasing sample size and 466 global distribution. 467 Generalizability and Metadata Integration 468 Unlike previous models fundamentally constrained to fixed taxa sets or specific datasets, 469 StrataBionn’s architecture is microbiome-agnostic. The ability to specify subsets of bacteria for 470 model training facilitates meta-analyses across disparate studies where sequencing depths or 471 taxonomic resolutions may vary. This flexibility is critical as the field moves toward large-scale 472 longitudinal cohorts involving the gut-brain axis, cardiovascular health, and susceptibility to 473 sexually transmitted infections (STIs). 474 The growing body of evidence confirming the mechanistic connections between the microbiome 475 and human health highlights a clear and increasing demand for highly accurate analytical tools. 476 These tools must also be adaptable to the study of the diverse microbial communities associated 477 with various hosts. Accurate characterization of the microbiome is essential to gain insights into 478 the gut-brain axis in neurodegenerative and psychological disorders [19,55], understanding the 479 role of various microbiotas in cardiovascular health [56,63,67], and determining susceptibility to 480 respiratory, renal, and sexually transmitted infections [57–64]. Within the gut microbiome, a 481 more granular understanding of the correlations between Community State Types (CSTs) and 482 pathology could enable targeted therapeutic interventions. These include fecal microbiota 483 transplants (FMT) designed to engraft disease-resistant profiles or the use of pre- and probiotics 484 to modulate the microbiome toward a protective state [11,19,65,66]. Indeed, CST-level analyses 485 have already proven more successful than lower-resolution studies in establishing the 486 correlations necessary to identify and treat conditions such as Celiac disease [65,66]. However, 487 because the performance of a state-typing model directly impacts the diagnostic accuracy of 488 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 23 microbiome-disease associations, algorithms must provide both maximal confidence and 489 biological explainability. StrataBionn addresses these requirements by delivering high-speed, 490 consistent, and interpretable labeling, allowing researchers to interrogate the classification 491 process without compromising predictive accuracy. 492 Technical Considerations and Limitations 493 StrataBionn was pre-configured with default hyperparameter values which allow the network to 494 accurately fit most classifications. Should the user desire to do so, they are enabled to manually 495 tune the hyperparameters to maximize the classification quality. We distribute the tool in a 496 GitHub repository that can be easily accessed by users and contains a provided conda 497 environment to quickly get started. In most cases, the user needs only to activate the environment 498 and provide labeled training and testing data to create a working classifier. 499 Despite the advantages of ANN architectures, the risk of overfitting remains a primary concern. 500 Neural network based classifiers inherit a concern of overfitting, where instead of learning 501 patterns in data, the model learns to memorize the samples in a dataset. We mitigated this by 502 implementing an "early stopping" protocol, where training is terminated once loss for an 503 independent test set plateaus, or decreases for several epochs. In addition to validation based 504 early stopping, the network depth was purposefully constrained to mitigate the risk of overfitting 505 to, or "memorizing" noise inherent to small-to-medium-sized datasets. If overfitting occurs, this 506 parameter is tunable and users can modify them accordingly. One of the benefits of our ANN, as 507 implemented in StrataBionn, is its ability to correctly assign community types in the presence of 508 non-linear boundaries. This is more clearly shown in the case of the oral microbiome (Figure 509 6B). These types of boundary classifications,which are likely frequent in microbiome data, are 510 problematic with simpler approaches like the nearest centroid classification algorithms. 511 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 24 Classifications from a pre-existing study with known biologically-relevant and composition-512 based CST labels [29] were used as a reference when validating StrataBionn’s performance on 513 the vaginal microbiome. As there are no generally accepted reference classifications for the oral 514 microbiome, we employed a K-means clustering model with 3 clusters (see supplementary 515 figures 3 and 4) to create naive reference classifications to validate our model on oral microbiota 516 data. As StrataBionn is designed to learn sample composition patterns from training data with 517 curated labels, the biological meaning behind the generated labels for the oral microbiome are 518 unimportant to the validation of the tool. 519 Studies have shown that the composition of microbiomes are inherently dynamic, changing due 520 to environmental and internal factors [68–70]. StrataBionn is limited by its inability to consider 521 temporal data, which could enhance the labelling ability of microbiome classification methods. 522 Future work could incorporate this data to gain a greater insight into the microbiome of an 523 individual, and how these state-change events correlate to disease acquisition and susceptibility. 524 Future iterations could incorporate Recurrent Neural Networks (RNNs), Long Short-Term 525 Memory (LSTM) networks, or Transformers to process temporal longitudinal data. Such an 526 evolution would allow for the prediction of state-change transitions, potentially identifying early 527 warning signs of dysbiosis before clinical symptoms manifest. As this information is collected 528 from the compilations of data points taken at discrete times, the high accuracy of StrataBionn 529 allows for the collection and aggregation of this data to further investigate the temporal evolution 530 of microbiome state-type in individuals. 531 As the field of microbiome research rapidly progresses, demanding faster and more accurate 532 community state typing, we attempt to contribute a solution. With StrataBionn, we offer a robust 533 and scalable classification tool designed to accommodate diverse studies, including the analysis 534 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 25 of ecological community data. Our goal is to enable future research to thoroughly investigate the 535 impact of microfauna composition on human and possibly ecosystem health. StrataBionn 536 achieves this by accurately classifying samples and providing tools to explore the decision 537 boundaries that govern the classification process. 538 Availability of data and materials 539 Information on how to obtain the vaginal microbiome datasets from France et al. [37] and 540 Hickey et al. [44], and the oral microbiome datasets from Manghi et al. [45] and Baker et al. [46] 541 can be found in their respective papers (we used the relative abundances reported by each one of 542 these studies). Nevertheless, a freeze of the tables with the relative abundances from these 543 studies is available in our github in the data folder. StrataBionn is available at 544 https://github.com/KelleyCornejoLabs/Microbiome_Classification. The raw FastQC for the in-545 house generated oral microbiome dataset is available at NCBI-SRA under BioProject 546 PRJNA1445365. 547 548 Figure 1. PACMAPs showing classifications from France et al. [37] for A) the 80% training set, 549 B) the 10% testing set, and C) the 10% validation set. 550 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 26 551 Figure 2. Precision metrics (A) Accuracy, (B) Recall, (C) F1-score, and (D) Precision for 552 Valencia, StrataBionn (60% training), StrataBionn (80% training), Random Forest (60% 553 training), Random Forest (80% training). 554 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 27 555 Figure 3. Confusion matrices for: (A) VALENCIA (classifying the 80% training set), (B) 556 StrataBionn (80% training), (C) Random Forest (80% training), and (D) Differential between 557 StrataBionn minus VALENCIA scores. In all figures the counts on the diagonal correspond to 558 successful assignments of communities by the methods to their true labels and the off-diagonals 559 correspond to misassignments. In figure 3D the positive values correspond to a larger number of 560 assignments to that category by StrataBionn, while negative numbers correspond to a lower 561 number of assignments by StrataBionn. Almost all elements on the diagonal of Figure D are 562 positive and all in the off-diagonal are negative, indicating that StrataBionn outperforms 563 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 28 VALENCIA in assigning communities to their true label, while reducing the misassignments. 564 The only exception is one fewer true assignment to CST IV-B compared to VALENCIA. 565 566 Figure 4. PACMAPs of reference dataset classifications and Hickey et al. [44] dataset 567 classifications generated by StrataBionn. In figure (A), both datasets are displayed with equal 568 opacity. In figure (B), classified Hickey et al. [44] data is displayed at full opacity while 569

Reference

data is displayed at 5% opacity. 570 571 Figure 5. PACMAPs showing Manghi et al. [45] dataset with KMeans clustering labels (for 572 three clusters: CST0, CST1, CST2), divided into A) the 80% training set, B) the 10% testing set, 573 and C) the 10% validation set. 574 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 29 575 Figure 6. (A) Accuracy, F1 score, Precision, and Recall achieved by StrataBionn on the 60% 576 validation and 80% validation clustered oral datasets. (B) Confusion matrix of StrataBionn 577 classifications on the 80% classified oral data validation set. Counts on the diagonal correspond 578 to successful assignments of communities by the methods to their true labels and the off-579 diagonals correspond to misassignments. 580 581 Figure 7. PACMAPs of Manghi et al. [45] oral microbiome reference dataset classifications, 582 Baker et al. [46] oral microbiome dataset classification, and in-house oral microbiome dataset 583 classifications generated by StrataBionn. In figure (A), both datasets are displayed with equal 584 opacity. In figure (B), classified in-house generated data is displayed at full opacity while 585 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 30

Reference

data is displayed at 5% opacity. In figure (C), Baker et al. [46] oral microbiome data is 586 displayed at full opacity while reference data is displayed at 5% opacity. 587

References

588 1. Björk JR, Bolte LA, Maltez Thomas A, Lee KA, Rossi N, Wind TT, et al. Longitudinal gut 589 microbiome changes in immune checkpoint blockade-treated advanced melanoma. Nat Med. 590 Nature Publishing Group; 2024;30:785–96. https://doi.org/10.1038/s41591-024-02803-3 591 2. Clauss M, Gérard P, Mosca A, Leclerc M. Interplay Between Exercise and Gut Microbiome in 592 the Context of Human Health and Performance. Front Nutr [Internet]. Frontiers; 2021 [cited 593 2026 Jan 7];8. https://doi.org/10.3389/fnut.2021.637010 594 3. Fan Y, Pedersen O. Gut microbiota in human metabolic health and disease. Nat Rev 595 Microbiol. Nature Publishing Group; 2021;19:55–71. https://doi.org/10.1038/s41579-020-0433-9 596 4. Gunjur A, Shao Y, Rozday T, Klein O, Mu A, Haak BW, et al. A gut microbial signature for 597 combination immune checkpoint blockade across cancer types. Nat Med. Nature Publishing 598 Group; 2024;30:797–809. https://doi.org/10.1038/s41591-024-02823-z 599 5. Jewell MD, van Moorsel SJ, Bell G. Presence of microbiome decreases fitness and modifies 600 phenotype in the aquatic plant Lemna minor. AoB PLANTS. 2023;15:plad026. 601 https://doi.org/10.1093/aobpla/plad026 602 6. Kandalai S, Li H, Zhang N, Peng H, Zheng Q. The human microbiome and cancer: a 603 diagnostic and therapeutic perspective. Cancer Biol Ther. 24:2240084. 604 https://doi.org/10.1080/15384047.2023.2240084 605 7. Liu B-N, Liu X-T, Liang Z-H, Wang J-H. Gut microbiota in obesity. World J Gastroenterol. 606 2021;27:3837–50. https://doi.org/10.3748/wjg.v27.i25.3837 607 8. Gould AL, Zhang V, Lamberti L, Jones EW, Obadia B, Korasidis N, et al. Microbiome 608 interactions shape host fitness. Proc Natl Acad Sci [Internet]. Proceedings of the National 609 Academy of Sciences; 2018 [cited 2026 Jan 7]; https://doi.org/10.1073/pnas.1809349115 610 9. Ogunrinola GA, Oyewale JO, Oshamika OO, Olasehinde GI. The Human Microbiome and Its 611 Impacts on Health. Int J Microbiol. 2020;2020:8045646. https://doi.org/10.1155/2020/8045646 612 10. Roelands J, Kuppen PJK, Ahmed EI, Mall R, Masoodi T, Singh P, et al. An integrated tumor, 613 immune and microbiome atlas of colon cancer. Nat Med. Nature Publishing Group; 614 2023;29:1273–86. https://doi.org/10.1038/s41591-023-02324-5 615 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 31 11. Sasidharan Pillai S, Gagnon CA, Foster C, Ashraf AP. Exploring the Gut Microbiota: Key 616 Insights Into Its Role in Obesity, Metabolic Syndrome, and Type 2 Diabetes. J Clin Endocrinol 617 Metab. 2024;109:2709–19. https://doi.org/10.1210/clinem/dgae499 618 12. Varghese S, Rao S, Khattak A, Zamir F, Chaari A. Physical Exercise and the Gut 619 Microbiome: A Bidirectional Relationship Influencing Health and Performance. Nutrients. 620 2024;16:3663. https://doi.org/10.3390/nu16213663 621 13. Xun W, Shao J, Shen Q, Zhang R. Rhizosphere microbiome: Functional compensatory 622 assembly for plant fitness. Comput Struct Biotechnol J. 2021;19:5487–93. 623 https://doi.org/10.1016/j.csbj.2021.09.035 624 14. Zhan M, Wang L, Xie C, Fu X, Zhang S, Wang A, et al. Succession of Gut Microbial 625 Structure in Twin Giant Pandas During the Dietary Change Stage and Its Role in Polysaccharide 626 Metabolism. Front Microbiol [Internet]. Frontiers; 2020 [cited 2026 Jan 7];11. 627 https://doi.org/10.3389/fmicb.2020.551038 628 15. Durack J, Lynch SV. The gut microbiome: Relationships with disease and opportunities for 629 therapy. J Exp Med. 2019;216:20–40. https://doi.org/10.1084/jem.20180448 630 16. Ghosh TS, Das M, Jeffery IB, O’Toole PW. Adjusting for age improves identification of gut 631 microbiome alterations in multiple diseases. Turnbaugh P, Garrett WS, Lozupone CA, 632 Turnbaugh P, editors. eLife. eLife Sciences Publications, Ltd; 2020;9:e50240. 633 https://doi.org/10.7554/eLife.50240 634 17. Grosso F, Zanetti D, Sanna S. Causal relationships between gut microbiome and hundreds of 635 age-related traits: evidence of a replicable effect on ApoM protein levels. Aging. 2025;17:1966–636 87. https://doi.org/10.18632/aging.206293 637 18. Gupta VK, Janda GS, Pump HK, Lele N, Cruz I, Cohen I, et al. Alterations in Gut 638 Microbiome-Host Relationships After Immune Perturbation in Patients With Multiple Sclerosis. 639 Neurol Neuroimmunol Neuroinflammation. Wolters Kluwer; 2025;12:e200355. 640 https://doi.org/10.1212/NXI.0000000000200355 641 19. Hou K, Wu Z-X, Chen X-Y, Wang J-Q, Zhang D, Xiao C, et al. Microbiota in health and 642 diseases. Signal Transduct Target Ther. Nature Publishing Group; 2022;7:135. 643 https://doi.org/10.1038/s41392-022-00974-4 644 20. Lewis FMT, Bernstein KT, Aral SO. Vaginal Microbiome and Its Relationship to Behavior, 645 Sexual Health, and Sexually Transmitted Diseases. Obstet Gynecol. 2017;129:643–54. 646 https://doi.org/10.1097/AOG.0000000000001932 647 21. Madhogaria B, Bhowmik P, Kundu A. Correlation between human gut microbiome and 648 diseases. Infect Med. 2022;1:180–91. https://doi.org/10.1016/j.imj.2022.08.004 649 22. Chen H, Jiang W. Application of high-throughput sequencing in understanding human oral 650 microbiome related with health and disease. Front Microbiol [Internet]. Frontiers; 2014 [cited 651 2026 Feb 24];5. https://doi.org/10.3389/fmicb.2014.00508 652 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 32 23. Di Bella JM, Bao Y, Gloor GB, Burton JP, Reid G. High throughput sequencing methods and 653 analysis for microbiome research. J Microbiol Methods. 2013;95:401–14. 654 https://doi.org/10.1016/j.mimet.2013.08.011 655 24. Compositional analysis: a valid approach to analyze microbiome high-throughput sequencing 656 data [Internet]. [cited 2026 Feb 24]. https://cdnsciencepub.com/doi/full/10.1139/cjm-2015-0821. 657 Accessed 24 Feb 2026 658 25. Arumugam M, Raes J, Pelletier E, Le Paslier D, Yamada T, Mende DR, et al. Enterotypes of 659 the human gut microbiome. Nature. Nature Publishing Group; 2011;473:174–80. 660 https://doi.org/10.1038/nature09944 661 26. Costea PI, Hildebrand F, Arumugam M, Bäckhed F, Blaser MJ, Bushman FD, et al. 662 Enterotypes in the landscape of gut microbial community composition. Nat Microbiol. 2018;3:8–663 16. https://doi.org/10.1038/s41564-017-0072-8 664 27. Callahan BJ, DiGiulio DB, Goltsman DSA, Sun CL, Costello EK, Jeganathan P, et al. 665 Replication and refinement of a vaginal microbial signature of preterm birth in two racially 666 distinct cohorts of US women. Proc Natl Acad Sci. Proceedings of the National Academy of 667 Sciences; 2017;114:9966–71. https://doi.org/10.1073/pnas.1705899114 668 28. DiGiulio DB, Callahan BJ, McMurdie PJ, Costello EK, Lyell DJ, Robaczewska A, et al. 669 Temporal and spatial variation of the human microbiota during pregnancy. Proc Natl Acad Sci. 670 Proceedings of the National Academy of Sciences; 2015;112:11060–5. 671 https://doi.org/10.1073/pnas.1502875112 672 29. Ravel J, Gajer P, Abdo Z, Schneider GM, Koenig SSK, McCulle SL, et al. Vaginal 673 microbiome of reproductive-age women. Proc Natl Acad Sci U S A. 2011;108 Suppl 1:4680–7. 674 https://doi.org/10.1073/pnas.1002611107 675 30. Ezugwu AE, Ikotun AM, Oyelade OO, Abualigah L, Agushaka JO, Eke CI, et al. A 676 comprehensive survey of clustering algorithms: State-of-the-art machine learning applications, 677 taxonomy, challenges, and future research prospects. Eng Appl Artif Intell. 2022;110:104743. 678 https://doi.org/10.1016/j.engappai.2022.104743 679 31. Fang Y, Subedi S. Clustering microbiome data using mixtures of logistic normal multinomial 680 models. Sci Rep. Nature Publishing Group; 2023;13:14758. https://doi.org/10.1038/s41598-023-681 41318-8 682 32. Gere A. Recommendations for validating hierarchical clustering in consumer sensory 683 projects. Curr Res Food Sci. 2023;6:100522. https://doi.org/10.1016/j.crfs.2023.100522 684 33. Gloor GB, Macklaim JM, Pawlowsky-Glahn V, Egozcue JJ. Microbiome Datasets Are 685 Compositional: And This Is Not Optional. Front Microbiol [Internet]. Frontiers; 2017 [cited 2026 686 Jan 7];8. https://doi.org/10.3389/fmicb.2017.02224 687 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 33 34. Liu Z, Yin X, Zhou Y, Li G, Chen K. Dissecting Microbial Community Structure and 688 Heterogeneity via Multivariate Covariate-Adjusted Clustering [Internet]. arXiv; 2025 [cited 2026 689 Jan 7]. https://doi.org/10.48550/arXiv.2508.11036 690 35. QIANG J, DING W, KUIJJER M, QUACKENBUSH J, CHEN P. Clustering Sparse Data 691 With Feature Correlation With Application to Discover Subtypes in Cancer. IEEE Access Pract 692 Innov Open Solut. 2020;8:67775–89. https://doi.org/10.1109/access.2020.2982569 693 36. Sinha R, Abu-Ali G, Vogtmann E, Fodor AA, Ren B, Amir A, et al. Assessment of variation 694 in microbial community amplicon sequencing by the Microbiome Quality Control (MBQC) 695 project consortium. Nat Biotechnol. 2017;35:1077–86. https://doi.org/10.1038/nbt.3981 696 37. France MT, Ma B, Gajer P, Brown S, Humphrys MS, Holm JB, et al. VALENCIA: a nearest 697 centroid classification method for vaginal microbial communities based on composition. 698 Microbiome. 2020;8:166. https://doi.org/10.1186/s40168-020-00934-6 699 38. Ronan T, Qi Z, Naegle KM. Avoiding common pitfalls when clustering biological data. Sci 700 Signal. 2016;9:re6. https://doi.org/10.1126/scisignal.aad1932 701 39. Levner I. Feature selection and nearest centroid classification for protein mass spectrometry. 702 BMC Bioinformatics. 2005;6:68. https://doi.org/10.1186/1471-2105-6-68 703 40. Ballabio D, Todeschini R. Multivariate Classification for Qualitative Analysis. Infrared 704 Spectrosc Food Qual Anal Control. Academic Press; 2009. p. 83–100. 705 https://doi.org/10.1016/B978-0-12-374136-3.00004-3 706 41. Blagus R, Lusa L. Class prediction for high-dimensional class-imbalanced data. BMC 707 Bioinformatics. 2010;11:523. https://doi.org/10.1186/1471-2105-11-523 708 42. Sánchez Reyna AG, Mendoza-Gonzalez R, Luna-García H, Celaya Padilla JM, Morgan 709 Benita JA, Espino-Salinas CH, et al. Synthetic data analysis for early detection of Alzheimer 710 progression through machine learning algorithms. PeerJ Comput Sci. 2024;10:e2437. 711 https://doi.org/10.7717/peerj-cs.2437 712 43. Ravel J, Gajer P, Abdo Z, Schneider GM, Koenig SSK, McCulle SL, et al. Vaginal 713 microbiome of reproductive-age women. Proc Natl Acad Sci. Proceedings of the National 714 Academy of Sciences; 2011;108:4680–7. https://doi.org/10.1073/pnas.1002611107 715 44. Hickey RJ, Zhou X, Settles ML, Erb J, Malone K, Hansmann MA, et al. Vaginal Microbiota 716 of Adolescent Girls Prior to the Onset of Menarche Resemble Those of Reproductive-Age 717 Women. mBio. American Society for Microbiology; 2015;6:10.1128/mbio.00097-15. 718 https://doi.org/10.1128/mbio.00097-15 719 45. Manghi P, Filosi M, Zolfo M, Casten LG, Garcia-Valiente A, Mattevi S, et al. Large-scale 720 metagenomic analysis of oral microbiomes reveals markers for autism spectrum disorders. Nat 721 Commun. Nature Publishing Group; 2024;15:9743. https://doi.org/10.1038/s41467-024-53934-7 722 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 34 46. Baker JL, Morton JT, Dinis M, Alvarez R, Tran NC, Knight R, et al. Deep metagenomics 723 examines the oral microbiome during dental caries, revealing novel taxa and co-occurrences with 724 host molecules. Genome Res. Cold Spring Harbor Lab; 2021;31:64–74. 725 https://doi.org/10.1101/gr.265645.120 726 47. FastQC [Internet]. 2015. https://qubeshub.org/resources/fastqc 727 48. Martin M. Cutadapt removes adapter sequences from high-throughput sequencing reads. 728 EMBnet.journal. 2011;17:10–2. https://doi.org/10.14806/ej.17.1.200 729 49. Babraham Bioinformatics - Trim Galore! [Internet]. [cited 2026 Feb 24]. 730 https://www.bioinformatics.babraham.ac.uk/projects/trim_galore/. Accessed 24 Feb 2026 731 50. Dobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, et al. STAR: ultrafast 732 universal RNA-seq aligner. Bioinformatics. 2013;29:15–21. 733 https://doi.org/10.1093/bioinformatics/bts635 734 51. Breitwieser FP, Baker DN, Salzberg SL. KrakenUniq: confident and fast metagenomics 735 classification using unique k-mer counts. Genome Biol. 2018;19:198. 736 https://doi.org/10.1186/s13059-018-1568-0 737 52. Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: 738 Machine Learning in Python. J Mach Learn Res. 2011;12:2825–30. 739 53. Ansel J, Yang E, He H, Gimelshein N, Jain A, Voznesensky M, et al. PyTorch 2: Faster 740 Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation. 741 29th ACM Int Conf Archit Support Program Lang Oper Syst Vol 2 ASPLOS 24 [Internet]. 742 ACM; 2024. https://doi.org/10.1145/3620665.3640366 743 54. Wang Y, Huang H, Rudin C, Shaposhnik Y. Understanding How Dimension Reduction 744 Tools Work: An Empirical Approach to Deciphering t-SNE, UMAP, TriMAP, and PaCMAP for 745 Data Visualization [Internet]. arXiv; 2021 [cited 2026 Feb 25]. 746 https://doi.org/10.48550/arXiv.2012.04456 747 55. The Microbiota-Gut-Brain Axis | Physiological Reviews | American Physiological Society 748 [Internet]. [cited 2026 Feb 24]. 749 https://journals.physiology.org/doi/full/10.1152/physrev.00018.2018?rfr_dat=cr_pu. Accessed 750 24 Feb 2026 751 56. Sanchez-Rodriguez E, Egea-Zorrilla A, Plaza-Díaz J, Aragón-Vela J, Muñoz-Quezada S, 752 Tercedor-Sánchez L, et al. The Gut Microbiota and Its Implication in the Development of 753 Atherosclerosis and Related Cardiovascular Diseases. Nutrients. Multidisciplinary Digital 754 Publishing Institute; 2020;12:605. https://doi.org/10.3390/nu12030605 755 57. Gosmann C, Anahtar MN, Handley SA, Farcasanu M, Abu-Ali G, Bowman BA, et al. 756 Lactobacillus-Deficient Cervicovaginal Bacterial Communities Are Associated with Increased 757 HIV Acquisition in Young South African Women. Immunity. 2017;46:29–37. 758 https://doi.org/10.1016/j.immuni.2016.12.013 759 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 35 58. Brotman RM, Bradford LL, Conrad M, Gajer P, Ault K, Peralta L, et al. Association between 760 Trichomonas vaginalis and vaginal bacterial community composition among reproductive-age 761 women. Sex Transm Dis. 2012;39:807–12. https://doi.org/10.1097/OLQ.0b013e3182631c79 762 59. van Houdt R, Ma B, Bruisten SM, Speksnijder AGCL, Ravel J, de Vries HJC. Lactobacillus 763 iners-dominated vaginal microbiota is associated with increased susceptibility to Chlamydia 764 trachomatis infection in Dutch women: a case-control study. Sex Transm Infect. 2018;94:117–765 23. https://doi.org/10.1136/sextrans-2017-053133 766 60. Hilty M, Burke C, Pedro H, Cardenas P, Bush A, Bossley C, et al. Disordered Microbial 767 Communities in Asthmatic Airways. PLOS ONE. Public Library of Science; 2010;5:e8578. 768 https://doi.org/10.1371/journal.pone.0008578 769 61. Millares L, Ferrari R, Gallego M, Garcia-Nuñez M, Pérez-Brocal V, Espasa M, et al. 770 Bronchial microbiome of severe COPD patients colonised by Pseudomonas aeruginosa. Eur J 771 Clin Microbiol Infect Dis. 2014;33:1101–11. https://doi.org/10.1007/s10096-013-2044-0 772 62. Huang YJ, Nelson CE, Brodie EL, DeSantis TZ, Baek MS, Liu J, et al. Airway microbiota 773 and bronchial hyperresponsiveness in patients with suboptimally controlled asthma. J Allergy 774 Clin Immunol. Elsevier; 2011;127:372-381.e3. https://doi.org/10.1016/j.jaci.2010.10.048 775 63. Barros AF, Borges NA, Ferreira DC, Carmo FL, Rosado AS, Fouque D, et al. Is there 776 Interaction Between Gut Microbial Profile and Cardiovascular Risk in Chronic Kidney Disease 777 Patients? Future Microbiol. Taylor & Francis; 2015;10:517–26. 778 https://doi.org/10.2217/fmb.14.140 779 64. Li L, Zhang Y-L, Liu X-Y, Meng X, Zhao R-Q, Ou L-L, et al. Periodontitis Exacerbates and 780 Promotes the Progression of Chronic Kidney Disease Through Oral Flora, Cytokines, and 781 Oxidative Stress. Front Microbiol. 2021;12:656372. https://doi.org/10.3389/fmicb.2021.656372 782 65. Krishnareddy S. The Microbiome in Celiac Disease. Gastroenterol Clin. Elsevier; 783 2019;48:115–26. https://doi.org/10.1016/j.gtc.2018.09.008 784 66. Valitutti F, Cucchiara S, Fasano A. Celiac Disease and the Microbiome. Nutrients. 785 Multidisciplinary Digital Publishing Institute; 2019;11:2403. 786 https://doi.org/10.3390/nu11102403 787 67. Prins FM, Collij V, Groot HE, Björk JR, Swarte JC, Andreu-Sánchez S, et al. The gut 788 microbiome across the cardiovascular risk spectrum. Eur J Prev Cardiol. 2024;31:935–44. 789 https://doi.org/10.1093/eurjpc/zwad377 790 68. Vandeputte D, De Commer L, Tito RY, Kathagen G, Sabino J, Vermeire S, et al. Temporal 791 variability in quantitative human gut microbiome profiles and implications for clinical research. 792 Nat Commun. Nature Publishing Group; 2021;12:6740. https://doi.org/10.1038/s41467-021-793 27098-7 794 69. Gajer P, Brotman RM, Bai G, Sakamoto J, Schütte UME, Zhong X, et al. Temporal 795 Dynamics of the Human Vaginal Microbiota. Sci Transl Med. American Association for the 796 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint 36 Advancement of Science; 2012;4:132ra52-132ra52. 797 https://doi.org/10.1126/scitranslmed.3003605 798 70. Gerber GK. The dynamic microbiome. FEBS Lett. 2014;588:4131–9. 799 https://doi.org/10.1016/j.febslet.2014.02.037 800 .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint Supplementary Materials: StrataBionn: a neural network supervised classification method for microbial communities Authors: Alex Symons1,2, Ashley Huynh3, Omar E. Cornejo2# Affiliations 1. Department of Computer Science and Engineering, University of California Santa Cruz, Santa Cruz, CA. 95064 2. Department of Ecology and Evolutionary Biology, University of California Santa Cruz, CA. 95064. 3. School of Biological Sciences, Washington State University, Pullman, WA. 99163 # corresponding author: Omar Cornejo, e-mail: [email protected] .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint Supplementary Table 1. This table shows the probabilities of assignment to each Sub-CST for all samples which appeared isolated compared to surrounding samples. Isolated in this context was defined as samples whose closest neighbor’s euclidean distance was above a threshold, after undergoing a PACMAP transformation. 13 such samples were discovered, and their assignment confidence values were gathered in the table above. We can see that samples found to be distant from other samples after a PACMAP transform, which groups samples of a similar composition, tend to receive a relatively low reported confidence value from StrataBionn. Supplementary Figure 1. PACMAPs showing classifications from France et al. for A) the 60% training set, B) the 20% testing set, and C) the 20% vaginal microbiome validation set. .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint Supplementary Figure 2. PACMAP containing 7,812 samples from Manghi et al. with Autism Spectrum Disorder (ASD) and control labels. While samples labeled ASD occur more commonly near the top of the PACMAP, ASD and control labeled samples co-occur throughout the entire figure. The lack of distinct clusters implies this classification scheme is not rooted in sample composition. .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint Supplementary Figure 3. A) Graph displaying silhouette scores for k clusters for naive classification of oral microbiome data using K-means clustering. We observe the silhouette score is highest using 3 clusters. B) Elbow plot showing within-cluster sum of squares (inertia) against number of clusters. Here we can identify that the inertia begins to slow when the cluster count reaches 3, which supports our decision to use three oral CST labels. .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint Supplementary Figure 4. Silhouette analysis plots showing the silhouette score for each sample in the oral microbiome dataset. Silhouette scores were calculated using naive K-means clustering-derived labels, with two (A), three (B), and four (C) clusters being tested. Using three clusters yields the highest average silhouette score, indicating the highest cluster separation. Supplementary Figure 5. This figure shows the decision boundaries learned by StrataBionn to classify the vaginal microbiome, focusing on the bacteria species Gardnerella vaginalis and Lactobacillus iners. We can clearly see in this graph that samples dominated by G. vaginalis (>60% composition) are always assigned to CST IV-B. Interestingly while there are many CST IV-B samples with a high proportion of G. vaginalis, few of them also contain as much L. iners as can be found in CST IV-B samples containing less G. vaginallis, although it would be possible. .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint Supplementary Figure 6. This figure shows the decision boundaries learned by StrataBionn to classify the vaginal microbiome, focusing on the bacteria species Lactobacillus crispatus and Lactobacillus iners. We can observe that CSTs I-A and I-B have a non-linear boundary when a sample is composed of ~70% L. crispatus. While CST I-A generally clusters in the corner where both bacteria species are low in relative abundance, this plot shows that the majority of the composition of these samples is L. crispatus. CST I-B has a complex but clear decision boundary with other CSTs. It is commonly assigned when a sample contains >~25% L. crispatus, but can occur with lower proportions of L. crispatus near ~50% L. iners. Sub-CSTs III-A and III-B seem to have a fuzzy decision boundary when a sample contains ~75% L. iners. Samples containing >~30% L. crispatus are generally assigned to neither III-A nor III-B, and are exclusively assigned to sub-CSTs of CST I. We also observe the compositional similarities between Sub- CSTs of CST I and CST III, as both share long boundaries where a single, clear line is difficult to draw using these axes. .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint Supplementary Figure 7. This figure shows the decision boundaries learned by StrataBionn to classify the vaginal microbiome, focusing on the bacteria species “Candidatus Lachnocurva vaginae” (formerly BVAB1) and Lactobacillus iners. In this plot we can observe that while most CSTs have relatively low abundances of “Ca. Lachnocurva vaginae”, Sub-CSTs IV-A, IV-B, and III-B are exceptions. Such samples with high levels of L. iners are generally classified as Sub- CST III-A, and samples with >~30% “Ca. Lachnocurva vaginae” are always classified as Sub- CST IV-A. .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint Supplementary Figure 8. Cleveland plots showing the impact of different species perturbation on the F1 score assignment to each CST type. We propose the use of this visualization in combination with the Feature space visualization from supplementary Figures 5-7 to identify defining features (species) that are relevant for the assignment of community labels. .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint Supplementary Figure 9. Heatmap for top 50 species in the oral dataset sequenced for this study. The values in the cells correspond to z normalized values of relative abundance and the annotation in the top corresponds to the subCST type that each sample was assigned to (CST0, CST1) and the probabilities of assignment to each CST type. In Supplementary Table 2. It can be seen that the uncertainty in the assignment, estimated as more evenness (Shannon-information index) in the values of probability of assignment is not significantly correlated with sequencing effort. Shannon diversity index in the probability of assignment was estimated as A normalized version of Shannon was estimated as .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint Comparisons were made between both the standard Shannon diversity index and the normalized version of it. Supplementary Table 2. Correlation values and probability .CC-BY 4.0 International licenseavailable under a was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made The copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: oa-pdf ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-08-12T06:43:03.944938+00:00
License: CC-BY-4.0