{"paper_id":"64280b9c-48d3-4ec7-84b5-fcb4761ea403","body_text":"1 \nStrataBionn: a neural network supervised 1 \nclassification method for microbial communities 2 \nAuthors: Alex Symons1,2, Ashley Huynh3, Omar E. Cornejo2# 3 \nAffiliations 4 \n 5 \n1. Department of Computer Science and Engineering, University of California Santa Cruz, Santa 6 \nCruz, CA. 95064 7 \n2. Department of Ecology and Evolutionary Biology, University of California Santa Cruz, CA. 8 \n95064. 9 \n3. School of Biological Sciences, Washington State University, Pullman, WA. 99163 10 \n 11 \n 12 \n# corresponding author: Omar Cornejo, e-mail: omcornej@ucsc.edu  13 \n 14 \n 15 \nKeywords: microbiome, microbiome classification, artificial neural network 16 \n 17 \n 18 \n 19 \n 20 \n 21 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n2 \nAbstract 22 \nThe classification of microbial communities into discrete states or \"community state types\" 23 \n(CSTs) is fundamental to understanding host-microbiome interactions and their clinical 24 \nimplications. Traditional methods, such as the nearest-neighbor approaches, often struggle with 25 \nthe inherent noise, high dimensionality, and non-linear signatures of taxonomic profiles. We 26 \npresent a novel supervised framework for microbial community classification, leveraging an 27 \nArtificial Neural Network (ANN) architecture implemented in a new tool we named StrataBionn. 28 \nWe rigorously evaluated our approach using large-scale vaginal microbiome datasets, directly 29 \nbenchmarking performance against VALENCIA and a Random Forest (RF) classifier. To 30 \ndemonstrate the versatility of our models, we further extended the framework to oral microbiome 31 \nclassification, assessing its stability across diverse anatomical sites. Our supervised models 32 \nconsistently outperformed the nearest-neighbor approach across all evaluated datasets. In the 33 \nvaginal microbiome, our method achieved an 11.6% to 13.3% increase in performance across all 34 \nprimary metrics, including precision, recall, accuracy, and F1-score. Furthermore, we 35 \ndemonstrate that this performance advantage is maintained in the oral microbiome, highlighting 36 \nthe generalizability of our neural network and ensemble strategies to various microbial 37 \necosystems without the need for niche-specific algorithmic adjustments. By capturing complex 38 \nfeature dependencies that distance-based methods overlook, our approach provides a more robust 39 \nand accurate census of microbial community structures. StrataBionn’s ability to learn 40 \nclassification schemes for any microbiome with high accuracy and explainability, through the 41 \nuse of provided utilities to visualize feature-space classification boundaries and perform 42 \nperturbation analysis on trained classifiers, makes it ideal for broad application in microecology 43 \nresearch. This framework offers a scalable, high-performance alternative for microbiome 44 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n3 \nresearchers, facilitating more precise clinical stratification and biological insights across hosts 45 \nbody sites. 46 \nIntroduction 47 \nThe study of microbial communities has fundamentally reshaped our understanding of human 48 \nhealth, disease, and biological fitness [1–14]. It has progressed from checking for the presence of 49 \na specific taxa to tracing the shifts of community-level microbiome structures, defined by 50 \nrelative species abundance and functional capabilities, which more accurately describes 51 \nphenotypic changes and disease risk in hosts [3,15–21]. Due to this paradigm shift, the rate at 52 \nwhich microbiome data is generated has outpaced the ability of analysis tools to derive meaning 53 \nfrom it. This has intensified the need for accurate, reproducible classification methods that can 54 \ncategorize microbiome communities to identify new biomarkers across diverse populations. 55 \nThe rapid maturation of high-throughput sequencing has shown that microbial communities 56 \nexhibit recurring patterns in composition across individuals, offering unprecedented 57 \nopportunities to map these microbial landscapes [22–24]. However, this \"big data\" era presents 58 \nsignificant analytical challenges. Extracting meaningful biological insights requires classification 59 \nframeworks that are computationally efficient and capable of high-level composition analysis 60 \nrequired in microbial ecology. Early research established the utility of composition-based 61 \ncommunity level classifications, leading to the development of standardized classification 62 \nsystems. Notable examples of such classifications include gut \"enterotypes\" [25,26] and vaginal 63 \n\"community state types\" (CSTs) [27–29], which are consistent assemblages of microbial species, 64 \nfoundational in guiding modern comparative microbiome research and clinical diagnostics. 65 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n4 \nDespite their utility, the methodologies used to define these categories remain a bottleneck. 66 \nHistorically, hierarchical clustering (HC) was the gold standard, yet its limitations in scalability 67 \nand reproducibility for rapidly growing datasets are well-documented [30–36]. To address this, 68 \nresearchers have turned to supervised approaches like the nearest centroid classifier (NCC), 69 \nimplemented in robust and easy-to-use tools such as VALENCIA, which are deterministic and 70 \nscale well [37]. However, NCC methods assume well-separated data and equal covariances; 71 \nconsequently, they often struggle with the \"fuzzy\", non-linear edges and overlapping 72 \ndistributions typical of complex biological datasets, such as those found in microbiome studies 73 \n[38,39]. Furthermore, NCC models fail to account for interactions between predictor variables—74 \na critical oversight in microbiome data, where inter-species interactions define the community 75 \nstructure [40]. Finally, NCC performance decays significantly when class variances are unequal 76 \nor when new data points fall outside the centroid-defined convex hulls [41,42]. 77 \nHere, we present StrataBionn, a novel neural network-based classification algorithm designed 78 \nto handle the non-linearities and high dimensionality of microbiome data. Unlike static 79 \nclassifiers, StrataBionn uses a neural network architecture that can be trained, saved, and adapted 80 \nto a wide variety of microbial datasets. By integrating automated training methods and data 81 \npreprocessing, the model achieves superior generalization and consistency. While StrataBionn 82 \noffers a more sophisticated parameterization than current NCC tools, we provide comprehensive 83 \nguidelines to streamline the fine-tuning process for diverse research applications. 84 \nWe benchmarked StrataBionn on microbial communities from two distinct human niches: the 85 \nvaginal and oral microbiomes. The vaginal microbiome is characterized by well-defined CSTs, 86 \nwhere four types (CST-I, II, III, and V) are dominated by specific Lactobacillus species, and one 87 \n(CST-IV) is defined by a diverse, anaerobic composition [43]. Using this established framework, 88 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n5 \nwe show that StrataBionn achieves higher precision and recall than existing methods. 89 \nFurthermore, we applied our method to the oral microbiome, where the challenge lies in 90 \ndistinguishing the \"oral core\" from the subtle deviations associated with periodontal and 91 \nsystemic diseases. Finally, we show its usefulness in the assignment of samples from a novel oral 92 \nsample set generated for this study. 93 \nIn summary, our results indicate that StrataBionn is a highly adaptable tool that is capable of 94 \nhigh-performance classification even with smaller reference datasets. While showcased here 95 \nusing vaginal and oral communities, this method is intended for broad application across any 96 \nmicrobiome study, including comparisons between healthy and diseased cohorts. We anticipate 97 \nthat StrataBionn will facilitate the characterization of large-scale metagenomic datasets, 98 \nultimately deepening our understanding of the microbial drivers of host health. 99 \nMethods 100 \nData Sources and Collection 101 \nPublicly Sourced Datasets: To train and evaluate StrataBionn, we utilized several previously 102 \npublished microbiota composition datasets. Vaginal microbiome profiles from France et al. [37] 103 \nwere used to train and validate the model against established Community State Type (CST) 104 \nclassifications, which we then validated against a dataset from Hickey et al. [44] to demonstrate 105 \nperformance on an independent dataset. Oral microbiome composition datasets were employed 106 \nto evaluate the framework’s utility as a generalized classifier. As the oral microbiome research 107 \ncommunity has not established a common reference for the community types found in the oral 108 \ncavity, we generated a de novo naive classification with a large dataset recently generated by 109 \nManghi et al.  [45]. We then used the naive classification as “true” labels for the training and 110 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n6 \nvalidation of our method. We show that this approach produces similar results, in terms of 111 \nperformance metrics, to those obtained in a more curated dataset like the vaginal microbiome. 112 \nWe show how the classifier can be used to classify samples on a sample set generated for this 113 \nwork and a dataset from published studies [46]. A freeze of the public datasets used for the study 114 \nare available in the github repository of the tool. 115 \nIn-House Oral Microbiome Collection and metagenomic sequencing: To test the 116 \nclassification consistency in the oral microbiome, plaque samples were collected from adult 117 \nrefugees (≥18 years) at the Rwamwanja UN refugee settlement, Uganda. Ethics and Recruitment: 118 \nCollection was approved by the Institutional Review Board at Washington State University (IRB 119 \n#15196-002). A total of 54 participants were recruited using a random number sequence 120 \ngenerated in R to ensure unbiased selection. Informed consent was obtained from all participants 121 \nand a translator was present when necessary. Sampling and Storage: Dental plaque was collected 122 \nduring routine cleanings, preserved in RNAlater, and stored at -80°C after transport to Pullman, 123 \nWA, under CDC Import Permit #2016-03-212. Library Preparation and Sequencing: Microbial 124 \nDNA was extracted using the MoBio Powerlizer™ DNA Isolation Kit (Mo Bio Laboratories, 125 \nCarlsbad, CA, USA), following the manufacturer’s protocol. For this, plaque samples preserved 126 \nin RNAlater were spun down, RNA later discarded and macerated manually with sterile pestles 127 \nwith lysis solution. After DNA extraction, samples were sheared using a Covaris M220 to a 550 128 \nbp insert size. Libraries were prepared with the NEBNext Ultra DNA Library Prep kit (New 129 \nEngland Biolabs Inc.) and sequenced on a single lane of an Illumina HiSeq 2500 at the 130 \nWashington State University Genomics Core. 131 \nBioinformatics and Quality Control of novel oral metagenomic data: Initial quality 132 \nassessment was performed using FastQC [47]. Out of 54 samples, 39 produced enough DNA 133 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n7 \nafter extraction and/or passed FastQC’s quality control test, and were used in analysis. Sequence 134 \ntrimming was performed with TrimGalore and Cutadapt  and we hard trimmed the first 11 bp on 135 \nthe 5’ end and soft trim for Phred scores < 25 [48,49]. Host Filtering: reads were aligned to the 136 \nhuman reference genome (GRCh38) using STAR [50], and human reads were removed and 137 \ndeleted permanently. Taxonomic Classification: non-human reads were classified using 138 \nKrakenUniq [51], and to minimize false positives, we required ≥100 unique kmers and a 139 \nduplicity ratio ≤ 3. Count matrices were generated at the genus and species levels. Raw sequence 140 \ndata is available at SRA-NCBI through accession PRJNA1445365. 141 \nMicrobiome classifier: the StrataBionn Framework 142 \nWe developed StrataBionn, stratification of biological data using neural networks, a supervised 143 \nclassification tool designed to assign CST, or other community level labels, to microbiome data. 144 \nIt accepts as an input a standardized data format (CSV) similar to that used by VALENCIA, 145 \nwhich includes columns for sample ID, read counts, labels, and a column for each taxa for which 146 \ncount data was collected. The method incorporates a stratified data partitioning strategy to ensure 147 \nrobust model training and evaluation, particularly in scenarios with varying data availability.  148 \nStratified Data Partitioning: To maintain the inherent class balance of the original labeled data 149 \nwithin each subset and mitigate potential biases, a stratified splitting approach was employed. 150 \nThis ensured that the amount of each CST in each subset used during training is consistent with 151 \nthe distribution found in the original, unpartitioned dataset, to avoid sampling biases during 152 \nmodel training. Datasets were first shuffled and then partitioned proportionally by CST label. 153 \nSpecifically, for each CST, the designated proportions of samples were allocated to the training, 154 \ntesting, and validation sets. Following the split, the proportional distribution of CST labels within 155 \neach subset was recalculated and compared to the original dataset distribution. Consistency was 156 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n8 \nautomatically assumed if the proportional representation of each CST in the resulting subsets 157 \nwas within a predefined tolerance of 0.1% of the original distribution. We defined tolerance 158 \nusing the equation (\t\n!!\"#\n!!\"$%&\n− 1) ∗ 100, where  𝑝\"#$ is the proportion of a given CST in the subset 159 \nof interest (training, test, validation), and  𝑝\"#!%&  is the proportion of that same CST in the 160 \nsuperset which contains all the training, testing, and validation samples. This rigorous 161 \nstratification ensures that each data subset provides a representative snapshot of the overall class 162 \ndistribution, facilitating unbiased model training and evaluation. Two data partitioning schemes 163 \nwere implemented and evaluated: (i) a 80/10/10 split for training, testing, and validation sets, 164 \nrespectively, simulating conditions with ample data; and (ii) a 60/20/20 split for training, testing 165 \nand validation of the same sets, mimicking data-limited scenarios. The processed and partitioned 166 \ndatasets were then exported as comma-separated value (CSV) files for subsequent model 167 \ntraining. These files were formatted in the same manner as those accepted by VALENCIA, 168 \ncontaining columns for the sample id, total number of reads, sample label (when training), and 169 \nbacteria species. This format was chosen to increase compatibility between StrataBionn and 170 \nother existing tools (i.e. VALENCIA).  171 \nData Preprocessing within StrataBionn: Prior to model training, StrataBionn performs several 172 \ncrucial preprocessing steps. To account for variations in sequencing depth or sample \"quality,\" 173 \nthe raw feature counts were normalized by the total counts per sample. This transformation 174 \nensures that samples with differing sequencing depths contribute equally to the model training 175 \nprocess. Subsequently, any null or zero values, which could lead to computational errors during 176 \ntraining, were imputed with a small, non-interfering numerical value. While the StrataBionn 177 \nframework includes the option for additional data normalization techniques, initial evaluations 178 \nindicated that the total count normalization alone yielded the most effective model performance 179 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n9 \nfor this specific application, and thus, further normalization steps were not performed in the final 180 \nmodel training and evaluation. 181 \nClassification Algorithms: Two distinct supervised learning algorithms were implemented and 182 \nevaluated within the StrataBionn framework: a Random Forest Classifier (RFC) and an Artificial 183 \nNeural Network (ANN). These methods were selected for their classification consistency and 184 \nevaluated along with the VALENCIA classifier which served as an accuracy baseline. 185 \n● Random Forest Classifier: For this model, we used the RandomForestClassifier module 186 \nfrom the python scikit-learn library [52]. The RFC was parameterized to optimize its 187 \nperformance. The number of estimators (n_estimators) and the number of features to 188 \nconsider when looking for the best split (max_features) were systematically tuned. 189 \nNumbers of estimators ranging from 100 to 10,000 were tested, and while training time 190 \nincreased significantly with more estimators, performance plateaued around 95.1% 191 \n(10,000 estimators), taking around 7 minutes. 192 \n 193 \n● Artificial Neural Network: The ANN model consisted of an input layer, a single hidden 194 \nlayer, a Leaky Rectified Linear Unit (Leaky ReLU) activation function following the 195 \nhidden layer, and a dropout layer for regularization. The size of the hidden layer was 196 \nparameterized, with a default setting of \n2\n3 ∗ 𝑛'(!#) + 𝑛*#)!#)  , where 𝑛'(!#) represents the 197 \nnumber of input neurons and 𝑛*#)!#)  the number of output neurons. The equation for the 198 \nfull system is as follows: 𝑊2(𝐿𝑅𝑒𝐿𝑈(𝑊1𝑥 + 𝑏1) \t ∘ \t\n+\n1,!) + 𝑏1, where 𝑥 is the input data 199 \nvector, 𝑊1, 𝑊2 are learnable weight matrices, 𝑏1, 𝑏2 are learnable bias matrices, 𝑝 is the 200 \nprobability that a neuron will dropout, and 𝑚 is a vector that masks neurons to be 201 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n10 \ndropped via multiplication with 0 or 1. The optimization algorithm was selectable 202 \nbetween Adaptive Moment Estimation (ADAM) and Stochastic Gradient Descent (SGD). 203 \nThe SGD optimizer offered an optional class weighting scheme to address potential 204 \nimbalances in CST prevalence, allowing for increased emphasis on accurately classifying 205 \nless frequent classes. The learning rate and a learning rate reduction factor were also 206 \ntunable hyperparameters. A probability of assignment is reported for all CSTs for each 207 \nsample, and the CST with the highest probability is assigned as the label for each sample. 208 \nThe model was constructed using the pytorch library and written python [53]. Training 209 \ntime for the vaginal microbiome test set with default hyperparameters took around 27 210 \nseconds. 211 \n 212 \n● VALENCIA Baseline: The previously established VALENCIA method was included as 213 \na baseline for performance comparison. It uses a nearest centroid classification model, 214 \nwhich employs pre-provided centroids constructed to classify vaginal microbiome 215 \nsamples. 216 \n 217 \nWe also provide a naive assessment of the uncertainty in the classification by extracting the 218 \nlogits from the softmax assignment and standardizing those values to generate probabilities of 219 \nsamples belonging to the assigned class. For this we exponentiated the raw logits and divide 220 \nthem by the sum of all exponentiated values according to the following formula, as implemented 221 \nin the pytorch library [53]: 222 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n11 \n 223 \nEvaluation and Validation Strategy 224 \nModel Evaluation: All trained models, including the baseline VALENCIA classifier, were 225 \nevaluated using the same vaginal microbiome dataset. Critically, each proposed model was 226 \nevaluated on a distinct test set that was not used during training or validation, ensuring an 227 \nunbiased assessment of generalization performance. The VALENCIA baseline was evaluated 228 \nusing its provided centroids, which were constructed based on the entire dataset. Model 229 \nperformance was quantified using classification accuracy, defined as the proportion of correctly 230 \nclassified samples.  231 \nEvaluation as General Microbiome Classifier: Performance of the ANN on oral microbiome 232 \ndatasets was measured to evaluate its usefulness as a generalized microbiome community 233 \nclassification method. We first performed a naive classification of the large oral microbiome 234 \ndataset using K-Means clustering to generate labels that would be employed for model training 235 \nand evaluation. Oral data was reformatted and treated identically to the vaginal microbiome data, 236 \nusing the same processing methods previously described. Evaluation of the ANN on oral 237 \nmicrobiome data was performed using the same classification accuracy measurement used for 238 \nthe vaginal microbiome. The Random Forest Classifier was not evaluated due to its much longer 239 \ntraining time to achieve a lower accuracy compared to the ANN, and VALENCIA was not 240 \nevaluated due to its specificity to the vaginal microbiome. We analyzed the performance of the 241 \nmethod with the oral microbiome. 242 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n12 \nApplication of Trained Classifiers Across Studies: A vaginal microbiome test set containing 243 \npopulation data on different bacteria species [44] was classified using a model trained to identify 244 \nthe CSTs from the VALENCIA training set. StrataBionn was first trained on the VALENCIA 245 \ntraining set while limited to considering only the set of bacteria species common to both the 246 \ntraining dataset and the target dataset. The trained classifier was then used to apply CST labels to 247 \nall samples in the target dataset, ignoring bacteria species which are not part of the common set. 248 \nThese CST classifications were then manually examined and determined to be reasonable based 249 \non their bacterial makeup using the findings of France et al [37]. 250 \nResults 251 \nTraining, Validation and Test datasets 252 \nFor the vaginal Microbiome, we generated stratified random sets from the France et al [37], for 253 \ntraining (80%: 10,580; 60%: 7,935), validation (80%: 1,334; 60%: 2,655), and testing (80%: 254 \n1,317; 60%: 2,641). For the oral microbiome, we generated stratified random sets from the 255 \nManghi et al. [45] study for training (80%: 6,248; 60%: 4,686), validation (80%: 784; 60%: 256 \n1,565), and testing (80%: 780; 60%: 1,561). 257 \nWe obtained read-counts supporting the presence of bacteria species in 39 newly generated oral 258 \nmicrobiomes for this study. The number of reads and estimated relative abundances for species 259 \ncan be found in the github repository. 260 \nPreprocessing of Vaginal Microbiome Datasets 261 \nDuring the development and testing of StrataBionn, two training data sets with 80% and 60% 262 \ntraining allocations were included to test scenarios when training data is abundant or limited. 263 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n13 \nStratified datasets from France et al. [37] containing curated CST classifications were used for 264 \ntheir high quality labels and large number of samples (13,231). The PACMAP [54] 265 \nrepresentations of the truth labels for the high training data availability scenario are shown in 266 \nFigure 1, demonstrating similar CST composition among subsets. Consistent with our 267 \nexpectations, more closely related CST communities are clustered closely in the PACMAP 268 \nrepresentation (Figure 1). For example, CST-IA and CST-IB, which both feature a high 269 \nabundance of Lactobacillus crispatus, have a similar projection in the map and tend to occur 270 \nmore closely than less related state types, such as CST-IA and CST-V. This observation supports 271 \nour use of PACMAP projections of the data for the purpose of representation. 272 \nThe worst CST distribution tolerances across all CSTs are 0.005%, 7.2%, and 7.2% for the 80% 273 \ntraining (Figure 1A), 10% testing (Figure 1B), and 10% validation (Figure 1C) sets respectively. 274 \nFor the 60% training, 20% testing, 20% validation sets, the tolerances are 0.05%, 13.4%, and 275 \n13.4% respectively. PACMAP representation of the 60%, 20%, 20% configuration is presented 276 \nin the supplementary materials (Supplementary Figure 1). 277 \nEvaluation of Classification Models 278 \nEach classification method was evaluated on both the high and low training data availability 279 \nlabeled datasets using accuracy (Figure 2A), recall (Figure 2B), F1 score (Figure 2C), and 280 \nprecision (Figure 2D) as metrics. These metrics were selected to measure the number of samples 281 \ncorrectly identified (accuracy), the rate of false negatives (recall), the rate of false positives 282 \n(precision), and balanced measure of both recall and accuracy (F1). Evaluation was performed 283 \non the validation dataset after training a classifier on the corresponding training dataset for all 284 \nclassifiers except for VALENCIA, where the original author provided centroids were used for 285 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n14 \nclassification [37]. While the ANN and RF classifiers trained on a subset of the Vaginal 286 \nmicrobiome data, VALENCIA’s centroids were generated using the full microbiome dataset as 287 \ndescribed in the author’s methodology [37]. For the ANN classifier, probabilities of assignment 288 \nwere above 87% for 95% of samples (80% training: 87.8%; 60% training: 87.6%), and around 289 \n95% for 80% of classified samples (80% training: 95.3%; 60% training: 94.9%). In all cases the 290 \nneural network classifier and random forest classifier outperformed VALENCIA when 291 \nclassifying both datasets. The ANN achieved a 11.6-13.3% increase compared to VALENCIA 292 \nacross all metrics (Figure 2A-D), performing slightly better with higher data availability, while 293 \nthe Random Forest approach achieved a 9-10.8% increase across all metrics compared to 294 \nVALENCIA (Figure 2A-D). The Random Forest classifier performed slightly better in lower 295 \ndata availability scenarios, but never outperformed the ANN classifier by any measured metric. 296 \n 297 \nTo understand the conditions which led to our models' increased performance, we generated 298 \nconfusion matrices for each classification method by comparing the validation dataset truth 299 \nvalues to StrataBionn-assigned CST labels. In VALENCIA, misidentification of CST-IA, CST-300 \nIB, CST-IIIA, CST-IIIB, and CST-IVA disproportionately contributes to its lower performance 301 \n(Figure 3A). In the case of the ANN classifier and random forest classifier implemented in 302 \nStrataBionn, communities CST-IA and CST-IIIA disproportionally contribute to the 303 \nmisassignment of the community labels; while CST-IIIB and CST-IVB marginally contribute to 304 \nthe misassignment (Figures 3B-C).  To better demonstrate the cases where StrataBionn 305 \noutperforms VALENCIA in assignments, we estimated the differences in the number of 306 \nclassified communities for each cell in the confusion matrix (Figure 3D). Cases where the 307 \npredicted label matched the true label were observed to be higher for the neural network 308 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n15 \nclassifier compared to VALENCIA for most CSTs, especially those which are less common 309 \n(blue diagonals in Figure 3D). Hotspots denoted in red on the off-diagonal of the comparative 310 \nconfusion matrix show that the neural network classifier is better able to differentiate between 311 \nsome CSTs with similar composition distributions than VALENCIA (Figure 3D). This is likely 312 \ndue to the ability to classify along non-linear boundaries which is not possible using a nearest 313 \ncentroid classifier based approach. Due to the greater performance and higher ability to 314 \ndifferentiate between similar CSTs, we chose to proceed with a neural network classifier as the 315 \nbasis for StrataBionn. 316 \nEvaluation of Novel Dataset Classifications 317 \nTo test the ability of StrataBionn to apply accurate classifications to novel datasets, we generated 318 \nclassifications for an unclassified vaginal microbiome dataset containing 458 samples from 319 \nHickey et al. [44]. This dataset required no preprocessing other than reformatting and column 320 \nrenaming. The classification model we used for this task was trained on the labeled France et al. 321 \n[37] dataset, and configured to consider only the bacteria species common to both datasets. We 322 \nthen applied this classifier to the full Hickey et al. [44] dataset to get CST assignments for each 323 \nsample. We generated two PACMAPs using both the classified Hickey et al. [44] dataset and the 324 \ntruth labels from the France et al. [37] dataset, shown in figure 4. We observed that the novel 325 \nclassifications were assigned labels that matched their surrounding samples with similar 326 \ncompositions. Assigned samples cluster around data points with the same classification from the 327 \noriginal France et al. [37] dataset. 328 \nClassification of Oral Microbiome Datasets 329 \nAn oral microbiome bacterial community composition dataset was used to validate StrataBionn’s 330 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n16 \nability as a general purpose classification tool on non-vaginal microbiome datasets. We used 331 \n7,812 samples collected by Manghi et al. [45] with pre-existing classifications removed due to 332 \npoor correlation with the data (see supplementary figure 2). As oral microbiomes have not been 333 \ncharacterized to the same level of detail of classification as the vaginal microbiomes, we 334 \nemployed a K-Means clustering algorithm to determine the optimal number of clusters for this 335 \ndataset. We found that using 3 clusters provided the optimal classification by overserving the 336 \nsilhouette and elbow plots provided in supplementary figures 3 and 4. After clustering labels 337 \nwere applied, the data was then partitioned into high and low data training data availability 338 \ndatasets using the stratified partitioning process previously described for the vaginal microbiome. 339 \nPACMAP visualizations of these subsets are provided in Figure 5. The worst CST distribution of 340 \ntolerance values for these datasets, as defined by the percent occurrences in a set above (positive) 341 \nor below (negative) the occurrences found in the reference set, were 0.008%, 0.0074%, and 342 \n0.0074% for the 80% training (Figure 5A), 10% testing (Figure 5B), and 10% validation (Figure 343 \n5C) sets respectively. For the 60% training, 20% testing, 20% validation sets, the tolerances were 344 \n0.019%, 0.35%, and 0.35% respectively. We observed that samples which were assigned to the 345 \nsame cluster exhibit spatial locality, suggesting that our assignments capture the composition of 346 \nthe samples, making them ideal for evaluating the performance of our model (Figure 5).  347 \nLabels assigned by StrataBionn were evaluated based on F1-score, recall, precision, and 348 \naccuracy metrics for high and low training data availability scenarios. StrataBionn achieves 98.8-349 \n98.9% across all metrics with low training data availability, and 99% across all metrics in high 350 \ntraining data availability scenarios (Figure 6A). We observed that this high level of accuracy was 351 \nconsistent across all label groups, with a slightly higher level of accuracy for less common CSTs 352 \ngiven their size (Figure 6B). The least common CST, CST-2, was correctly identified in all 184 353 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n17 \nsamples which truly belonged to that group, while only being mis-identified in two samples 354 \nbelonging to CST-1. CST-1, the most common subtype, is the one most commonly mis-assigned 355 \nto a different cluster (7/369, 1.9%), likely due to the complicated boundaries between it and the 356 \nother CSTs observed in figure 6B. StrataBionn is able to correctly identify these 98.1% of the 357 \ntime, indicating a strong ability to classify along complicated, non-linear boundaries. 358 \nThe in-house generated oral microbiome dataset containing 39 samples, and 47 additional 359 \nsamples from Baker et al [46]. were used to evaluate the ability of StrataBionn to apply 360 \nclassifications to novel non-vaginal microbiome datasets, and required no additional processing 361 \nother than reformatting and column renaming. This new dataset was assigned labels using a 362 \nclassifier trained on the 80% training data availability training subset for the oral microbiome 363 \ndescribed above using only the bacteria species common to both datasets. All samples from 364 \nBaker et al. [46] were classified as CST-1 by StrataBionn, matching the label of the samples with 365 \nknown CSTs with which they were compositionally most similar to (Fig. 7). In the in-house and 366 \nBaker et al. [46] datasets, we identified 7 belonging to CST-0, and 79 samples in CST-1. No 367 \nsamples in these sets were identified as belonging to CST-2.  368 \nDiscussion 369 \nThere is increasing evidence of the implication of community-level microbial shifts with the 370 \nincreased likelihood of disease, and human health more broadly [19,55–66]. This puts an 371 \nincreased importance on our ability to accurately infer the microbial communities and 372 \ndiscriminate those that are commonly associated with healthy outcomes from those commonly 373 \nassociated with disease. The complexity of microbial community data, characterized by high 374 \ndimensionality, compositional sparsity, and technical noise, requires robust, interpretable, and 375 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n18 \ntransferable classification frameworks. The classification process is generally abstracted away 376 \nand hidden from the user, concealing useful information regarding how assignments are 377 \ndetermined during the classification process which hinders the interpretability of those 378 \nclassifications. Additionally, current classification methods are fundamentally constrained to a 379 \nsingle dataset with a fixed set of bacteria species, preventing meaningful classifications derived 380 \nfrom one study from being applied generally to additional studies for the same microbiome. To 381 \naddress these issues, we developed StrataBionn, an Artificial Neural Network (ANN)-based tool 382 \ndesigned to overcome the limitations of extant classifiers by providing superior accuracy and 383 \ncross-study applicability. By integrating transparent visualization and perturbation analysis, 384 \nStrataBionn transitions microbiome \"state-typing\" from a \"black-box\" abstraction to an 385 \ninterpretable diagnostic process, where researchers can investigate the compositional causality of 386 \nthe inferred community labels in a community. Although StrataBionn has been developed for 387 \nmicrobial communities, it could likely be adapted to study any ecological community data for 388 \nwhich abundant samples and training are available due to its community-independent 389 \ncomposition based classification architecture. 390 \nComparative Performance and Architecture Selection 391 \nOur results demonstrate that StrataBionn consistently outperforms both the k-nearest neighbor-392 \nbased approach of VALENCIA and a traditional Random Forest (RF) classifier implemented in 393 \nour package. While RF models showed competitive performance when compared to k-nearest 394 \nneighbor approaches, the ANN architecture was selected for the basis of StrataBionn due to its 395 \nsuperior computational efficiency and accuracy at scale. Notably, StrataBionn achieved a 12-396 \n13% increase in F1-score and precision over VALENCIA when benchmarked against the 397 \nvaginal microbiome, a dataset and type of community that VALENCIA was specifically 398 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n19 \ndesigned to resolve. The performance gain was most pronounced in the classification of sub-399 \nCSTs (e.g., CSTs IV-A and IV-C0), which often present overlapping composition profiles that 400 \nconfound less sensitive models. Compared to the Random Forest classifier, StrataBionn achieves 401 \nan approximately 2% increase across all the same metrics, with the difference being slightly 402 \nmore pronounced with more training data. The performance metric differences between the 60% 403 \ntraining dataset (7,935 training samples) and the 80% training dataset (10,580 training samples) 404 \nwere negligible in our testing, but the 80% training dataset did achieve slightly higher 405 \nperformance. More specifically, we can attribute this increase in correct assignments to a higher 406 \nprecision for all sub-CSTs of CST I and III, and higher precision for CSTs IV-A, IV-C0, and V 407 \n(Fig. 3). This is important because often CST-IVs are associated with increased likelihood of 408 \ndisease, and thus our method can dramatically reduce false positives in disease susceptibility 409 \nscreenings by correctly attributing samples to their true, non-CST-IV type, without degraded 410 \nrecall ability. What is less known is how the observed differences among CST-IV subgroups are 411 \ndifferentially associated with disease, making the distinction of those communities relevant for 412 \nany future study aimed at discovering the compositional and causal contributions of these 413 \ncommunity differences to disease. 414 \nThe marginal performance increase observed when expanding the training set from 60% to 80% 415 \nsuggests that StrataBionn reaches an early accuracy plateau, indicating high data efficiency. This 416 \nis particularly important given that many datasets are going to be sample limited, and in those 417 \ncircumstances, the strategy for the partition of the samples for training, testing and validation 418 \nmight differ. While France et al. [37] argued that VALENCIA’s assignments may better reflect 419 \nbiological states than the hierarchical clustering (HC) labels, our objective was to validate the 420 \nlearning fidelity of the model. By using HC labels as a ground-truth proxy, we demonstrated that 421 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n20 \nStrataBionn possesses a superior capacity to internalize and replicate complex, user-defined 422 \nclassification logic. The ANN learns classification assignments based on their compositions 423 \nduring an iterative training process, where training is continued until accuracy improvements 424 \nreach a maximum. The result is a high accuracy, deterministic trained model which can apply 425 \nconsistent classifications to novel data. As the training process learns distributions based on user-426 \nprovided labels, it is not inherently tied to any one microbiome, which has been a major 427 \nlimitation of previous models. Additionally, by allowing the user to specify a set of bacteria 428 \nspecies to consider during this process, disregarding other species, StrataBionn can produce 429 \nadditional trained models which can be applied across separate datasets for meta-analysis. 430 \nIn the PACMAP visualization of the data, it can be seen that some samples assigned to CSTs 431 \nappear isolated from the general clusters. This visualization can help researchers identify 432 \nsamples with low probability of assignment. In Supplementary Table 1, we provide the 433 \nprobabilities of assignment for a subset of samples from the vaginal microbiome that seem 434 \nisolated from their clusters. In these samples, the general trend is that they share a low 435 \nprobability of assignment to their clusters. Because our method provides an empirical probability 436 \nof assignment, researchers can perform a post-assignment filter to declare ambiguities. 437 \nInterpretability and Compositional Causality 438 \nA persistent critique of deep learning in biology is the lack of model transparency. We took a 439 \nnovel approach to address this concern by including a classification boundary visualization tool 440 \nand a perturbation analysis utility, which increases transparency in the classification process. 441 \nMore specifically, StrataBionn addresses the issue of interpretability via two distinct 442 \nmechanisms: 443 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n21 \n1. Feature-Space Visualization: By projecting classification boundaries onto two-444 \ndimensional bacterial abundance coordinate planes, researchers can visually verify the 445 \n\"decision logic\" of the model. By allowing the user to select two bacteria for the x and y 446 \naxis and displaying the classification assignment in a scatter plot, insight as to what 447 \nbacteria correspond to which CST can be attained. For example, we can inspect which 448 \nbacteria are correlated with the different sub-CSTs by searching for a bacteria species 449 \nwhere all samples appear to vary widely in abundance, appearing distributed across an 450 \naxis. Then, we can find another species varying widely in abundance and observe the 451 \nclustering of certain CSTs, displayed in different colors, to determine the features and 452 \nthresholds which lead to their assignment. Examples of these plots and their 453 \ninterpretations are provided in supplementary figures 5-7. 454 \n2. Stratified Perturbation Analysis: Stratabionn’s perturbation analysis utility quantifies 455 \nthe sensitivity of the model to specific taxa. By permuting the counts of individual 456 \nspecies and measuring the resulting decay in F1-score, StrataBionn identifies which 457 \nbacteria are computationally \"essential\" for a specific CST assignment. StrataBionn 458 \nprovides the impact of individual species perturbations for F1, precision, recall, and 459 \naccuracy metrics, providing additional insight into the impact of single species into the 460 \nincreased sensitivity or specificity of the assignment to communities (supplementary 461 \nfigure 8). 462 \nWe propose that users implement the classification algorithm and use the tools provided by 463 \nStrataBionn to further investigate the compositional causality of specific community states. We 464 \nanticipate that StrataBionn will enable researchers to identify emerging properties that cause 465 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n22 \nmicrobial communities to differ across host populations in datasets of increasing sample size and 466 \nglobal distribution. 467 \nGeneralizability and Metadata Integration 468 \nUnlike previous models fundamentally constrained to fixed taxa sets or specific datasets, 469 \nStrataBionn’s architecture is microbiome-agnostic. The ability to specify subsets of bacteria for 470 \nmodel training facilitates meta-analyses across disparate studies where sequencing depths or 471 \ntaxonomic resolutions may vary. This flexibility is critical as the field moves toward large-scale 472 \nlongitudinal cohorts involving the gut-brain axis, cardiovascular health, and susceptibility to 473 \nsexually transmitted infections (STIs). 474 \nThe growing body of evidence confirming the mechanistic connections between the microbiome 475 \nand human health highlights a clear and increasing demand for highly accurate analytical tools. 476 \nThese tools must also be adaptable to the study of the diverse microbial communities associated 477 \nwith various hosts. Accurate characterization of the microbiome is essential to gain insights into 478 \nthe gut-brain axis in neurodegenerative and psychological disorders [19,55], understanding the 479 \nrole of various microbiotas in cardiovascular health [56,63,67], and determining susceptibility to 480 \nrespiratory, renal, and sexually transmitted infections [57–64]. Within the gut microbiome, a 481 \nmore granular understanding of the correlations between Community State Types (CSTs) and 482 \npathology could enable targeted therapeutic interventions. These include fecal microbiota 483 \ntransplants (FMT) designed to engraft disease-resistant profiles or the use of pre- and probiotics 484 \nto modulate the microbiome toward a protective state [11,19,65,66]. Indeed, CST-level analyses 485 \nhave already proven more successful than lower-resolution studies in establishing the 486 \ncorrelations necessary to identify and treat conditions such as Celiac disease [65,66]. However, 487 \nbecause the performance of a state-typing model directly impacts the diagnostic accuracy of 488 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n23 \nmicrobiome-disease associations, algorithms must provide both maximal confidence and 489 \nbiological explainability. StrataBionn addresses these requirements by delivering high-speed, 490 \nconsistent, and interpretable labeling, allowing researchers to interrogate the classification 491 \nprocess without compromising predictive accuracy. 492 \nTechnical Considerations and Limitations 493 \nStrataBionn was pre-configured with default hyperparameter values which allow the network to 494 \naccurately fit most classifications. Should the user desire to do so, they are enabled to manually 495 \ntune the hyperparameters to maximize the classification quality. We distribute the tool in a 496 \nGitHub repository that can be easily accessed by users and contains a provided conda 497 \nenvironment to quickly get started. In most cases, the user needs only to activate the environment 498 \nand provide labeled training and testing data to create a working classifier. 499 \nDespite the advantages of ANN architectures, the risk of overfitting remains a primary concern. 500 \nNeural network based classifiers inherit a concern of overfitting, where instead of learning 501 \npatterns in data, the model learns to memorize the samples in a dataset. We mitigated this by 502 \nimplementing an \"early stopping\" protocol, where training is terminated once loss for an 503 \nindependent test set plateaus, or decreases for several epochs. In addition to validation based 504 \nearly stopping, the network depth was purposefully constrained to mitigate the risk of overfitting 505 \nto, or \"memorizing\" noise inherent to small-to-medium-sized datasets. If overfitting occurs, this 506 \nparameter is tunable and users can modify them accordingly. One of the benefits of our ANN, as 507 \nimplemented in StrataBionn, is its ability to correctly assign community types in the presence of 508 \nnon-linear boundaries. This is more clearly shown in the case of the oral microbiome (Figure 509 \n6B). These types of boundary classifications,which are likely frequent in microbiome data, are 510 \nproblematic with simpler approaches like the nearest centroid classification algorithms. 511 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n24 \nClassifications from a pre-existing study with known biologically-relevant and composition-512 \nbased CST labels [29] were used as a reference when validating StrataBionn’s performance on 513 \nthe vaginal microbiome. As there are no generally accepted reference classifications for the oral 514 \nmicrobiome, we employed a K-means clustering model with 3 clusters (see supplementary 515 \nfigures 3 and 4) to create naive reference classifications to validate our model on oral microbiota 516 \ndata. As StrataBionn is designed to learn sample composition patterns from training data with 517 \ncurated labels, the biological meaning behind the generated labels for the oral microbiome are 518 \nunimportant to the validation of the tool. 519 \nStudies have shown that the composition of microbiomes are inherently dynamic, changing due 520 \nto environmental and internal factors [68–70]. StrataBionn is limited by its inability to consider 521 \ntemporal data, which could enhance the labelling ability of microbiome classification methods. 522 \nFuture work could incorporate this data to gain a greater insight into the microbiome of an 523 \nindividual, and how these state-change events correlate to disease acquisition and susceptibility. 524 \nFuture iterations could incorporate Recurrent Neural Networks (RNNs), Long Short-Term 525 \nMemory (LSTM) networks, or Transformers to process temporal longitudinal data. Such an 526 \nevolution would allow for the prediction of state-change transitions, potentially identifying early 527 \nwarning signs of dysbiosis before clinical symptoms manifest. As this information is collected 528 \nfrom the compilations of data points taken at discrete times, the high accuracy of StrataBionn 529 \nallows for the collection and aggregation of this data to further investigate the temporal evolution 530 \nof microbiome state-type in individuals. 531 \nAs the field of microbiome research rapidly progresses, demanding faster and more accurate 532 \ncommunity state typing, we attempt to contribute a solution. With StrataBionn, we offer a robust 533 \nand scalable classification tool designed to accommodate diverse studies, including the analysis 534 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n25 \nof ecological community data. Our goal is to enable future research to thoroughly investigate the 535 \nimpact of microfauna composition on human and possibly ecosystem health. StrataBionn 536 \nachieves this by accurately classifying samples and providing tools to explore the decision 537 \nboundaries that govern the classification process. 538 \nAvailability of data and materials 539 \nInformation on how to obtain the vaginal microbiome datasets from France et al. [37] and 540 \nHickey et al. [44], and the oral microbiome datasets from Manghi et al. [45] and Baker et al. [46] 541 \ncan be found in their respective papers (we used the relative abundances reported by each one of 542 \nthese studies). Nevertheless, a freeze of the tables with the relative abundances from these 543 \nstudies is available in our github in the data folder. StrataBionn is available at 544 \nhttps://github.com/KelleyCornejoLabs/Microbiome_Classification. The raw FastQC for the in-545 \nhouse generated oral microbiome dataset is available at NCBI-SRA under BioProject 546 \nPRJNA1445365. 547 \n 548 \nFigure 1. PACMAPs showing classifications from France et al. [37] for A) the 80% training set, 549 \nB) the 10% testing set, and C) the 10% validation set. 550 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n26 \n 551 \nFigure 2. Precision metrics (A) Accuracy, (B) Recall, (C) F1-score, and (D) Precision for 552 \nValencia, StrataBionn (60% training), StrataBionn (80% training), Random Forest (60% 553 \ntraining), Random Forest (80% training). 554 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n27 \n 555 \nFigure 3. Confusion matrices for: (A) VALENCIA (classifying the 80% training set), (B) 556 \nStrataBionn (80% training), (C) Random Forest (80% training), and (D) Differential between 557 \nStrataBionn minus VALENCIA scores. In all figures the counts on the diagonal correspond to 558 \nsuccessful assignments of communities by the methods to their true labels and the off-diagonals 559 \ncorrespond to misassignments. In figure 3D the positive values correspond to a larger number of 560 \nassignments to that category by StrataBionn, while negative numbers correspond to a lower 561 \nnumber of assignments by StrataBionn. Almost all elements on the diagonal of Figure D are 562 \npositive and all in the off-diagonal are negative, indicating that StrataBionn outperforms 563 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n28 \nVALENCIA in assigning communities to their true label, while reducing the misassignments. 564 \nThe only exception is one fewer true assignment to CST IV-B compared to VALENCIA. 565 \n 566 \nFigure 4. PACMAPs of reference dataset classifications and Hickey et al. [44] dataset 567 \nclassifications generated by StrataBionn. In figure (A), both datasets are displayed with equal 568 \nopacity. In figure (B), classified Hickey et al. [44] data is displayed at full opacity while 569 \nreference data is displayed at 5% opacity. 570 \n 571 \nFigure 5. PACMAPs showing Manghi et al. [45] dataset with KMeans clustering labels (for 572 \nthree clusters: CST0, CST1, CST2), divided into A) the 80% training set, B) the 10% testing set, 573 \nand C) the 10% validation set. 574 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n29 \n 575 \nFigure 6. (A) Accuracy, F1 score, Precision, and Recall achieved by StrataBionn on the 60% 576 \nvalidation and 80% validation clustered oral datasets. (B) Confusion matrix of StrataBionn 577 \nclassifications on the 80% classified oral data validation set. Counts on the diagonal correspond 578 \nto successful assignments of communities by the methods to their true labels and the off-579 \ndiagonals correspond to misassignments. 580 \n 581 \nFigure 7. PACMAPs of Manghi et al. [45] oral microbiome reference dataset classifications, 582 \nBaker et al. [46] oral microbiome dataset classification, and in-house oral microbiome dataset 583 \nclassifications generated by StrataBionn. In figure (A), both datasets are displayed with equal 584 \nopacity. In figure (B), classified in-house generated data is displayed at full opacity while 585 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n30 \nreference data is displayed at 5% opacity. In figure (C), Baker et al. [46] oral microbiome data is 586 \ndisplayed at full opacity while reference data is displayed at 5% opacity. 587 \nReferences 588 \n1. Björk JR, Bolte LA, Maltez Thomas A, Lee KA, Rossi N, Wind TT, et al. Longitudinal gut 589 \nmicrobiome changes in immune checkpoint blockade-treated advanced melanoma. Nat Med. 590 \nNature Publishing Group; 2024;30:785–96. https://doi.org/10.1038/s41591-024-02803-3  591 \n2. Clauss M, Gérard P, Mosca A, Leclerc M. Interplay Between Exercise and Gut Microbiome in 592 \nthe Context of Human Health and Performance. Front Nutr [Internet]. Frontiers; 2021 [cited 593 \n2026 Jan 7];8. https://doi.org/10.3389/fnut.2021.637010  594 \n3. Fan Y, Pedersen O. Gut microbiota in human metabolic health and disease. Nat Rev 595 \nMicrobiol. Nature Publishing Group; 2021;19:55–71. https://doi.org/10.1038/s41579-020-0433-9  596 \n4. Gunjur A, Shao Y, Rozday T, Klein O, Mu A, Haak BW, et al. A gut microbial signature for 597 \ncombination immune checkpoint blockade across cancer types. Nat Med. Nature Publishing 598 \nGroup; 2024;30:797–809. https://doi.org/10.1038/s41591-024-02823-z  599 \n5. Jewell MD, van Moorsel SJ, Bell G. Presence of microbiome decreases fitness and modifies 600 \nphenotype in the aquatic plant Lemna minor. AoB PLANTS. 2023;15:plad026. 601 \nhttps://doi.org/10.1093/aobpla/plad026  602 \n6. Kandalai S, Li H, Zhang N, Peng H, Zheng Q. The human microbiome and cancer: a 603 \ndiagnostic and therapeutic perspective. Cancer Biol Ther. 24:2240084. 604 \nhttps://doi.org/10.1080/15384047.2023.2240084  605 \n7. Liu B-N, Liu X-T, Liang Z-H, Wang J-H. Gut microbiota in obesity. World J Gastroenterol. 606 \n2021;27:3837–50. https://doi.org/10.3748/wjg.v27.i25.3837  607 \n8. Gould AL, Zhang V, Lamberti L, Jones EW, Obadia B, Korasidis N, et al. Microbiome 608 \ninteractions shape host fitness. Proc Natl Acad Sci [Internet]. Proceedings of the National 609 \nAcademy of Sciences; 2018 [cited 2026 Jan 7]; https://doi.org/10.1073/pnas.1809349115  610 \n9. Ogunrinola GA, Oyewale JO, Oshamika OO, Olasehinde GI. The Human Microbiome and Its 611 \nImpacts on Health. Int J Microbiol. 2020;2020:8045646. https://doi.org/10.1155/2020/8045646  612 \n10. Roelands J, Kuppen PJK, Ahmed EI, Mall R, Masoodi T, Singh P, et al. An integrated tumor, 613 \nimmune and microbiome atlas of colon cancer. Nat Med. Nature Publishing Group; 614 \n2023;29:1273–86. https://doi.org/10.1038/s41591-023-02324-5  615 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n31 \n11. Sasidharan Pillai S, Gagnon CA, Foster C, Ashraf AP. Exploring the Gut Microbiota: Key 616 \nInsights Into Its Role in Obesity, Metabolic Syndrome, and Type 2 Diabetes. J Clin Endocrinol 617 \nMetab. 2024;109:2709–19. https://doi.org/10.1210/clinem/dgae499  618 \n12. Varghese S, Rao S, Khattak A, Zamir F, Chaari A. Physical Exercise and the Gut 619 \nMicrobiome: A Bidirectional Relationship Influencing Health and Performance. Nutrients. 620 \n2024;16:3663. https://doi.org/10.3390/nu16213663  621 \n13. Xun W, Shao J, Shen Q, Zhang R. Rhizosphere microbiome: Functional compensatory 622 \nassembly for plant fitness. Comput Struct Biotechnol J. 2021;19:5487–93. 623 \nhttps://doi.org/10.1016/j.csbj.2021.09.035  624 \n14. Zhan M, Wang L, Xie C, Fu X, Zhang S, Wang A, et al. Succession of Gut Microbial 625 \nStructure in Twin Giant Pandas During the Dietary Change Stage and Its Role in Polysaccharide 626 \nMetabolism. Front Microbiol [Internet]. Frontiers; 2020 [cited 2026 Jan 7];11. 627 \nhttps://doi.org/10.3389/fmicb.2020.551038  628 \n15. Durack J, Lynch SV. The gut microbiome: Relationships with disease and opportunities for 629 \ntherapy. J Exp Med. 2019;216:20–40. https://doi.org/10.1084/jem.20180448  630 \n16. Ghosh TS, Das M, Jeffery IB, O’Toole PW. Adjusting for age improves identification of gut 631 \nmicrobiome alterations in multiple diseases. Turnbaugh P, Garrett WS, Lozupone CA, 632 \nTurnbaugh P, editors. eLife. eLife Sciences Publications, Ltd; 2020;9:e50240. 633 \nhttps://doi.org/10.7554/eLife.50240  634 \n17. Grosso F, Zanetti D, Sanna S. Causal relationships between gut microbiome and hundreds of 635 \nage-related traits: evidence of a replicable effect on ApoM protein levels. Aging. 2025;17:1966–636 \n87. https://doi.org/10.18632/aging.206293  637 \n18. Gupta VK, Janda GS, Pump HK, Lele N, Cruz I, Cohen I, et al. Alterations in Gut 638 \nMicrobiome-Host Relationships After Immune Perturbation in Patients With Multiple Sclerosis. 639 \nNeurol Neuroimmunol Neuroinflammation. Wolters Kluwer; 2025;12:e200355. 640 \nhttps://doi.org/10.1212/NXI.0000000000200355  641 \n19. Hou K, Wu Z-X, Chen X-Y, Wang J-Q, Zhang D, Xiao C, et al. Microbiota in health and 642 \ndiseases. Signal Transduct Target Ther. Nature Publishing Group; 2022;7:135. 643 \nhttps://doi.org/10.1038/s41392-022-00974-4  644 \n20. Lewis FMT, Bernstein KT, Aral SO. Vaginal Microbiome and Its Relationship to Behavior, 645 \nSexual Health, and Sexually Transmitted Diseases. Obstet Gynecol. 2017;129:643–54. 646 \nhttps://doi.org/10.1097/AOG.0000000000001932  647 \n21. Madhogaria B, Bhowmik P, Kundu A. Correlation between human gut microbiome and 648 \ndiseases. Infect Med. 2022;1:180–91. https://doi.org/10.1016/j.imj.2022.08.004  649 \n22. Chen H, Jiang W. Application of high-throughput sequencing in understanding human oral 650 \nmicrobiome related with health and disease. Front Microbiol [Internet]. Frontiers; 2014 [cited 651 \n2026 Feb 24];5. https://doi.org/10.3389/fmicb.2014.00508  652 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n32 \n23. Di Bella JM, Bao Y, Gloor GB, Burton JP, Reid G. High throughput sequencing methods and 653 \nanalysis for microbiome research. J Microbiol Methods. 2013;95:401–14. 654 \nhttps://doi.org/10.1016/j.mimet.2013.08.011  655 \n24. Compositional analysis: a valid approach to analyze microbiome high-throughput sequencing 656 \ndata [Internet]. [cited 2026 Feb 24]. https://cdnsciencepub.com/doi/full/10.1139/cjm-2015-0821. 657 \nAccessed 24 Feb 2026  658 \n25. Arumugam M, Raes J, Pelletier E, Le Paslier D, Yamada T, Mende DR, et al. Enterotypes of 659 \nthe human gut microbiome. Nature. Nature Publishing Group; 2011;473:174–80. 660 \nhttps://doi.org/10.1038/nature09944  661 \n26. Costea PI, Hildebrand F, Arumugam M, Bäckhed F, Blaser MJ, Bushman FD, et al. 662 \nEnterotypes in the landscape of gut microbial community composition. Nat Microbiol. 2018;3:8–663 \n16. https://doi.org/10.1038/s41564-017-0072-8  664 \n27. Callahan BJ, DiGiulio DB, Goltsman DSA, Sun CL, Costello EK, Jeganathan P, et al. 665 \nReplication and refinement of a vaginal microbial signature of preterm birth in two racially 666 \ndistinct cohorts of US women. Proc Natl Acad Sci. Proceedings of the National Academy of 667 \nSciences; 2017;114:9966–71. https://doi.org/10.1073/pnas.1705899114  668 \n28. DiGiulio DB, Callahan BJ, McMurdie PJ, Costello EK, Lyell DJ, Robaczewska A, et al. 669 \nTemporal and spatial variation of the human microbiota during pregnancy. Proc Natl Acad Sci. 670 \nProceedings of the National Academy of Sciences; 2015;112:11060–5. 671 \nhttps://doi.org/10.1073/pnas.1502875112  672 \n29. Ravel J, Gajer P, Abdo Z, Schneider GM, Koenig SSK, McCulle SL, et al. Vaginal 673 \nmicrobiome of reproductive-age women. Proc Natl Acad Sci U S A. 2011;108 Suppl 1:4680–7. 674 \nhttps://doi.org/10.1073/pnas.1002611107  675 \n30. Ezugwu AE, Ikotun AM, Oyelade OO, Abualigah L, Agushaka JO, Eke CI, et al. A 676 \ncomprehensive survey of clustering algorithms: State-of-the-art machine learning applications, 677 \ntaxonomy, challenges, and future research prospects. Eng Appl Artif Intell. 2022;110:104743. 678 \nhttps://doi.org/10.1016/j.engappai.2022.104743  679 \n31. Fang Y, Subedi S. Clustering microbiome data using mixtures of logistic normal multinomial 680 \nmodels. Sci Rep. Nature Publishing Group; 2023;13:14758. https://doi.org/10.1038/s41598-023-681 \n41318-8  682 \n32. Gere A. Recommendations for validating hierarchical clustering in consumer sensory 683 \nprojects. Curr Res Food Sci. 2023;6:100522. https://doi.org/10.1016/j.crfs.2023.100522  684 \n33. Gloor GB, Macklaim JM, Pawlowsky-Glahn V, Egozcue JJ. Microbiome Datasets Are 685 \nCompositional: And This Is Not Optional. Front Microbiol [Internet]. Frontiers; 2017 [cited 2026 686 \nJan 7];8. https://doi.org/10.3389/fmicb.2017.02224  687 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n33 \n34. Liu Z, Yin X, Zhou Y, Li G, Chen K. Dissecting Microbial Community Structure and 688 \nHeterogeneity via Multivariate Covariate-Adjusted Clustering [Internet]. arXiv; 2025 [cited 2026 689 \nJan 7]. https://doi.org/10.48550/arXiv.2508.11036  690 \n35. QIANG J, DING W, KUIJJER M, QUACKENBUSH J, CHEN P. Clustering Sparse Data 691 \nWith Feature Correlation With Application to Discover Subtypes in Cancer. IEEE Access Pract 692 \nInnov Open Solut. 2020;8:67775–89. https://doi.org/10.1109/access.2020.2982569  693 \n36. Sinha R, Abu-Ali G, Vogtmann E, Fodor AA, Ren B, Amir A, et al. Assessment of variation 694 \nin microbial community amplicon sequencing by the Microbiome Quality Control (MBQC) 695 \nproject consortium. Nat Biotechnol. 2017;35:1077–86. https://doi.org/10.1038/nbt.3981  696 \n37. France MT, Ma B, Gajer P, Brown S, Humphrys MS, Holm JB, et al. VALENCIA: a nearest 697 \ncentroid classification method for vaginal microbial communities based on composition. 698 \nMicrobiome. 2020;8:166. https://doi.org/10.1186/s40168-020-00934-6  699 \n38. Ronan T, Qi Z, Naegle KM. Avoiding common pitfalls when clustering biological data. Sci 700 \nSignal. 2016;9:re6. https://doi.org/10.1126/scisignal.aad1932  701 \n39. Levner I. Feature selection and nearest centroid classification for protein mass spectrometry. 702 \nBMC Bioinformatics. 2005;6:68. https://doi.org/10.1186/1471-2105-6-68  703 \n40. Ballabio D, Todeschini R. Multivariate Classification for Qualitative Analysis. Infrared 704 \nSpectrosc Food Qual Anal Control. Academic Press; 2009. p. 83–100. 705 \nhttps://doi.org/10.1016/B978-0-12-374136-3.00004-3  706 \n41. Blagus R, Lusa L. Class prediction for high-dimensional class-imbalanced data. BMC 707 \nBioinformatics. 2010;11:523. https://doi.org/10.1186/1471-2105-11-523  708 \n42. Sánchez Reyna AG, Mendoza-Gonzalez R, Luna-García H, Celaya Padilla JM, Morgan 709 \nBenita JA, Espino-Salinas CH, et al. Synthetic data analysis for early detection of Alzheimer 710 \nprogression through machine learning algorithms. PeerJ Comput Sci. 2024;10:e2437. 711 \nhttps://doi.org/10.7717/peerj-cs.2437  712 \n43. Ravel J, Gajer P, Abdo Z, Schneider GM, Koenig SSK, McCulle SL, et al. Vaginal 713 \nmicrobiome of reproductive-age women. Proc Natl Acad Sci. Proceedings of the National 714 \nAcademy of Sciences; 2011;108:4680–7. https://doi.org/10.1073/pnas.1002611107  715 \n44. Hickey RJ, Zhou X, Settles ML, Erb J, Malone K, Hansmann MA, et al. Vaginal Microbiota 716 \nof Adolescent Girls Prior to the Onset of Menarche Resemble Those of Reproductive-Age 717 \nWomen. mBio. American Society for Microbiology; 2015;6:10.1128/mbio.00097-15. 718 \nhttps://doi.org/10.1128/mbio.00097-15  719 \n45. Manghi P, Filosi M, Zolfo M, Casten LG, Garcia-Valiente A, Mattevi S, et al. Large-scale 720 \nmetagenomic analysis of oral microbiomes reveals markers for autism spectrum disorders. Nat 721 \nCommun. Nature Publishing Group; 2024;15:9743. https://doi.org/10.1038/s41467-024-53934-7  722 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n34 \n46. Baker JL, Morton JT, Dinis M, Alvarez R, Tran NC, Knight R, et al. Deep metagenomics 723 \nexamines the oral microbiome during dental caries, revealing novel taxa and co-occurrences with 724 \nhost molecules. Genome Res. Cold Spring Harbor Lab; 2021;31:64–74. 725 \nhttps://doi.org/10.1101/gr.265645.120  726 \n47. FastQC [Internet]. 2015. https://qubeshub.org/resources/fastqc  727 \n48. Martin M. Cutadapt removes adapter sequences from high-throughput sequencing reads. 728 \nEMBnet.journal. 2011;17:10–2. https://doi.org/10.14806/ej.17.1.200  729 \n49. Babraham Bioinformatics - Trim Galore! [Internet]. [cited 2026 Feb 24]. 730 \nhttps://www.bioinformatics.babraham.ac.uk/projects/trim_galore/. Accessed 24 Feb 2026  731 \n50. Dobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, et al. STAR: ultrafast 732 \nuniversal RNA-seq aligner. Bioinformatics. 2013;29:15–21. 733 \nhttps://doi.org/10.1093/bioinformatics/bts635  734 \n51. Breitwieser FP, Baker DN, Salzberg SL. KrakenUniq: confident and fast metagenomics 735 \nclassification using unique k-mer counts. Genome Biol. 2018;19:198. 736 \nhttps://doi.org/10.1186/s13059-018-1568-0  737 \n52. Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: 738 \nMachine Learning in Python. J Mach Learn Res. 2011;12:2825–30.  739 \n53. Ansel J, Yang E, He H, Gimelshein N, Jain A, Voznesensky M, et al. PyTorch 2: Faster 740 \nMachine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation. 741 \n29th ACM Int Conf Archit Support Program Lang Oper Syst Vol 2 ASPLOS 24 [Internet]. 742 \nACM; 2024. https://doi.org/10.1145/3620665.3640366  743 \n54. Wang Y, Huang H, Rudin C, Shaposhnik Y. Understanding How Dimension Reduction 744 \nTools Work: An Empirical Approach to Deciphering t-SNE, UMAP, TriMAP, and PaCMAP for 745 \nData Visualization [Internet]. arXiv; 2021 [cited 2026 Feb 25]. 746 \nhttps://doi.org/10.48550/arXiv.2012.04456  747 \n55. The Microbiota-Gut-Brain Axis | Physiological Reviews | American Physiological Society 748 \n[Internet]. [cited 2026 Feb 24]. 749 \nhttps://journals.physiology.org/doi/full/10.1152/physrev.00018.2018?rfr_dat=cr_pu. Accessed 750 \n24 Feb 2026  751 \n56. Sanchez-Rodriguez E, Egea-Zorrilla A, Plaza-Díaz J, Aragón-Vela J, Muñoz-Quezada S, 752 \nTercedor-Sánchez L, et al. The Gut Microbiota and Its Implication in the Development of 753 \nAtherosclerosis and Related Cardiovascular Diseases. Nutrients. Multidisciplinary Digital 754 \nPublishing Institute; 2020;12:605. https://doi.org/10.3390/nu12030605  755 \n57. Gosmann C, Anahtar MN, Handley SA, Farcasanu M, Abu-Ali G, Bowman BA, et al. 756 \nLactobacillus-Deficient Cervicovaginal Bacterial Communities Are Associated with Increased 757 \nHIV Acquisition in Young South African Women. Immunity. 2017;46:29–37. 758 \nhttps://doi.org/10.1016/j.immuni.2016.12.013  759 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n35 \n58. Brotman RM, Bradford LL, Conrad M, Gajer P, Ault K, Peralta L, et al. Association between 760 \nTrichomonas vaginalis and vaginal bacterial community composition among reproductive-age 761 \nwomen. Sex Transm Dis. 2012;39:807–12. https://doi.org/10.1097/OLQ.0b013e3182631c79  762 \n59. van Houdt R, Ma B, Bruisten SM, Speksnijder AGCL, Ravel J, de Vries HJC. Lactobacillus 763 \niners-dominated vaginal microbiota is associated with increased susceptibility to Chlamydia 764 \ntrachomatis infection in Dutch women: a case-control study. Sex Transm Infect. 2018;94:117–765 \n23. https://doi.org/10.1136/sextrans-2017-053133  766 \n60. Hilty M, Burke C, Pedro H, Cardenas P, Bush A, Bossley C, et al. Disordered Microbial 767 \nCommunities in Asthmatic Airways. PLOS ONE. Public Library of Science; 2010;5:e8578. 768 \nhttps://doi.org/10.1371/journal.pone.0008578  769 \n61. Millares L, Ferrari R, Gallego M, Garcia-Nuñez M, Pérez-Brocal V, Espasa M, et al. 770 \nBronchial microbiome of severe COPD patients colonised by Pseudomonas aeruginosa. Eur J 771 \nClin Microbiol Infect Dis. 2014;33:1101–11. https://doi.org/10.1007/s10096-013-2044-0  772 \n62. Huang YJ, Nelson CE, Brodie EL, DeSantis TZ, Baek MS, Liu J, et al. Airway microbiota 773 \nand bronchial hyperresponsiveness in patients with suboptimally controlled asthma. J Allergy 774 \nClin Immunol. Elsevier; 2011;127:372-381.e3. https://doi.org/10.1016/j.jaci.2010.10.048  775 \n63. Barros AF, Borges NA, Ferreira DC, Carmo FL, Rosado AS, Fouque D, et al. Is there 776 \nInteraction Between Gut Microbial Profile and Cardiovascular Risk in Chronic Kidney Disease 777 \nPatients? Future Microbiol. Taylor & Francis; 2015;10:517–26. 778 \nhttps://doi.org/10.2217/fmb.14.140  779 \n64. Li L, Zhang Y-L, Liu X-Y, Meng X, Zhao R-Q, Ou L-L, et al. Periodontitis Exacerbates and 780 \nPromotes the Progression of Chronic Kidney Disease Through Oral Flora, Cytokines, and 781 \nOxidative Stress. Front Microbiol. 2021;12:656372. https://doi.org/10.3389/fmicb.2021.656372  782 \n65. Krishnareddy S. The Microbiome in Celiac Disease. Gastroenterol Clin. Elsevier; 783 \n2019;48:115–26. https://doi.org/10.1016/j.gtc.2018.09.008  784 \n66. Valitutti F, Cucchiara S, Fasano A. Celiac Disease and the Microbiome. Nutrients. 785 \nMultidisciplinary Digital Publishing Institute; 2019;11:2403. 786 \nhttps://doi.org/10.3390/nu11102403  787 \n67. Prins FM, Collij V, Groot HE, Björk JR, Swarte JC, Andreu-Sánchez S, et al. The gut 788 \nmicrobiome across the cardiovascular risk spectrum. Eur J Prev Cardiol. 2024;31:935–44. 789 \nhttps://doi.org/10.1093/eurjpc/zwad377  790 \n68. Vandeputte D, De Commer L, Tito RY, Kathagen G, Sabino J, Vermeire S, et al. Temporal 791 \nvariability in quantitative human gut microbiome profiles and implications for clinical research. 792 \nNat Commun. Nature Publishing Group; 2021;12:6740. https://doi.org/10.1038/s41467-021-793 \n27098-7  794 \n69. Gajer P, Brotman RM, Bai G, Sakamoto J, Schütte UME, Zhong X, et al. Temporal 795 \nDynamics of the Human Vaginal Microbiota. Sci Transl Med. American Association for the 796 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n36 \nAdvancement of Science; 2012;4:132ra52-132ra52. 797 \nhttps://doi.org/10.1126/scitranslmed.3003605  798 \n70. Gerber GK. The dynamic microbiome. FEBS Lett. 2014;588:4131–9. 799 \nhttps://doi.org/10.1016/j.febslet.2014.02.037 800 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n \nSupplementary Materials: StrataBionn: a neural \nnetwork supervised classification method for \nmicrobial communities \nAuthors: Alex Symons1,2, Ashley Huynh3, Omar E. Cornejo2# \nAffiliations \n \n1. Department of Computer Science and Engineering, University of California Santa Cruz, Santa \nCruz, CA. 95064 \n2. Department of Ecology and Evolutionary Biology, University of California Santa Cruz, CA. \n95064. \n3. School of Biological Sciences, Washington State University, Pullman, WA. 99163 \n \n \n# corresponding author: Omar Cornejo, e-mail: omcornej@ucsc.edu  \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n \nSupplementary Table 1. This table shows the probabilities of assignment to each Sub-CST for all \nsamples which appeared isolated compared to surrounding samples. Isolated in this context was \ndefined as samples whose closest neighbor’s euclidean distance was above a threshold, after \nundergoing a PACMAP transformation. 13 such samples were discovered, and their assignment \nconfidence values were gathered in the table above. We can see that samples found to be distant \nfrom other samples after a PACMAP transform, which groups samples of a similar composition, \ntend to receive a relatively low reported confidence value from StrataBionn. \n \n \n \n \n \nSupplementary Figure 1. PACMAPs showing classifications from France et al. for A) the 60% \ntraining set, B) the 20% testing set, and C) the 20% vaginal microbiome validation set. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n \n \nSupplementary Figure 2. PACMAP containing 7,812 samples from Manghi et al. with Autism \nSpectrum Disorder (ASD) and control labels. While samples labeled ASD occur more commonly \nnear the top of the PACMAP, ASD and control labeled samples co-occur throughout the entire \nfigure. The lack of distinct clusters implies this classification scheme is not rooted in sample \ncomposition.  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n \n \n \nSupplementary Figure 3. A) Graph displaying silhouette scores for k clusters for naive \nclassification of oral microbiome data using K-means clustering. We observe the silhouette score \nis highest using 3 clusters. B) Elbow plot showing within-cluster sum of squares (inertia) against \nnumber of clusters. Here we can identify that the inertia begins to slow when the cluster count \nreaches 3, which supports our decision to use three oral CST labels. \n \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n \nSupplementary Figure 4. Silhouette analysis plots showing the silhouette score  for each sample \nin the oral microbiome dataset. Silhouette scores were calculated using naive K-means \nclustering-derived labels, with two (A), three (B), and four (C) clusters being tested. Using three \nclusters yields the highest average silhouette score, indicating the highest cluster separation. \n \n \nSupplementary Figure 5. This figure shows the decision boundaries learned by StrataBionn to \nclassify the vaginal microbiome, focusing on the bacteria species Gardnerella vaginalis and \nLactobacillus iners. We can clearly see in this graph that samples dominated by G. vaginalis \n(>60% composition) are always assigned to CST IV-B. Interestingly while there are many CST \nIV-B samples with a high proportion of G. vaginalis, few of them also contain as much L. iners \nas can be found in CST IV-B samples containing less G. vaginallis, although it would be \npossible. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n \n \nSupplementary Figure 6. This figure shows the decision boundaries learned by StrataBionn to \nclassify the vaginal microbiome, focusing on the bacteria species Lactobacillus crispatus and \nLactobacillus iners. We can observe that CSTs I-A and I-B have a non-linear boundary when a \nsample is composed of ~70% L. crispatus. While CST I-A generally clusters in the corner where \nboth bacteria species are low in relative abundance, this plot shows that the majority of the \ncomposition of these samples is L. crispatus. CST I-B has a complex but clear decision boundary \nwith other CSTs. It is commonly assigned when a sample contains >~25% L. crispatus, but can \noccur with lower proportions of L. crispatus near ~50% L. iners. Sub-CSTs III-A and III-B seem \nto have a fuzzy decision boundary when a sample contains ~75% L. iners. Samples containing \n>~30% L. crispatus are generally assigned to neither III-A nor III-B, and are exclusively \nassigned to sub-CSTs of CST I. We also observe the compositional similarities between Sub-\nCSTs of CST I and CST III, as both share long boundaries where a single, clear line is difficult \nto draw using these axes.  \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n \n \nSupplementary Figure 7. This figure shows the decision boundaries learned by StrataBionn to \nclassify the vaginal microbiome, focusing on the bacteria species “Candidatus Lachnocurva \nvaginae” (formerly BVAB1) and Lactobacillus iners. In this plot we can observe that while most \nCSTs have relatively low abundances of “Ca. Lachnocurva vaginae”, Sub-CSTs IV-A, IV-B, and \nIII-B are exceptions. Such samples with high levels of L. iners are generally classified as Sub-\nCST III-A, and samples with >~30% “Ca. Lachnocurva vaginae” are always classified as Sub-\nCST IV-A. \n \n \n \n \n \n \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n \n \nSupplementary Figure 8. Cleveland plots showing the impact of different species perturbation on \nthe F1 score assignment to each CST type. We propose the use of this visualization in \ncombination with the Feature space visualization from supplementary Figures 5-7 to identify \ndefining features (species) that are relevant for the assignment of community labels. \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n \n \nSupplementary Figure 9. Heatmap for top 50 species in the oral dataset sequenced for this study. \nThe values in the cells correspond to z normalized values of relative abundance and the \nannotation in the top corresponds to the subCST type that each sample was assigned to (CST0, \nCST1) and the probabilities of assignment to each CST type. \n \nIn Supplementary Table 2. It can be seen that the uncertainty in the assignment, estimated as \nmore evenness (Shannon-information index) in the values of probability of assignment is not \nsignificantly correlated with sequencing effort. Shannon diversity index in the probability of \nassignment was estimated as \n \n  \n \nA normalized version of Shannon was estimated as \n \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint \n\n \n \nComparisons were made between both the standard Shannon diversity index and the normalized \nversion of it. \n \nSupplementary Table 2. Correlation values and probability  \n \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted April 2, 2026. ; https://doi.org/10.64898/2026.03.31.715659doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}