Towards a Generative Paradigm for Large-scale Microbiome Analysis by Generative Language Model

preprint OA: closed
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-15

MGM 2.0, a generative language model, treats microbiome samples as sentences to extract nuanced relationships, predict colonization, generate disease-specific profiles, and optimize fecal microbiota transplantation donor selection.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-15 · read from full text

The study introduces MGM 2.0, an NLP-based generative framework that converts genus-level microbiome abundance tables into “sentence-like” representations (samples as sentences, taxa as words) and applies self-supervised and generative language-model techniques. Using large-scale microbiome data, the model showed robust predictive ability for identifying exogenous species colonization (AUROC = 0.86) and generated prompt-conditioned, disease-specific microbial profiles that were assessed using a “Microbiome Turing Test.” It was also applied to fecal microbiota transplantation by framing donor selection and post-transplant community composition prediction as a sequence-to-sequence task, reporting an average increase in C2R of 0.52 and the identification of potential “super donors.” The paper’s explicit limitation is not clearly stated in the provided text excerpt, but the caveat visible is that its generative and downstream performance are demonstrated within the framework’s chosen representation (genus-level normalization and ranked tokenization). The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Microbiome analysis has traditionally relied on taxonomic abundance tables, which, while effective, often constrain the exploration of deeper contextual relationships. In this study, we present MGM 2.0, a novel framework that applies advanced natural language processing (NLP) techniques to microbiome research. By reimagining microbiome samples as sentences and microbial species as words, MGM 2.0 enabled the extraction of nuanced patterns and relationships. The model demonstrated robust predictive performance in identifying exogenous species colonization (AUROC = 0.86). Additionally, through prompt-guided microbiome data generation, MGM 2.0 produced realistic microbial profiles conditioned on disease labels. The framework further revolutionized donor selection in fecal microbiota transplantation (FMT) by framing it as a sequence-to-sequence prediction task, enabling the prediction of post-transplantation community compositions and the identification of super donors for personalized treatments (average increase in C2R = 0.52). This innovative integration of NLP and microbiome science provides a versatile toolkit for predictive modeling, data generation, and personalized medicine. Highlights Introduced MGM 2.0, a generative language model utilizing sentence-like representation for microbiome analysis and generation. Demonstrated that sentence-like representation preserves sample distinctions, enabling accurate microbial sample classification tasks, such as colonization prediction. Generated realistic, disease-specific microbiome profiles using a prompt-guided approach, validated by a novel “Microbiome Turing Test.” Applied MGM 2.0 to fecal microbiota transplantation (FMT) donor selection, accurately predicting post-transplant community compositions and identifying potential “super donors” for personalized treatment strategies.
Full text 52,032 characters · extracted from preprint-html · click to expand
Towards a Generative Paradigm for Large-scale Microbiome Analysis by Generative Language Model | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results Towards a Generative Paradigm for Large-scale Microbiome Analysis by Generative Language Model View ORCID Profile Haohong Zhang , Zixin Kang , Yuli Zhang , Ronghua Yang , View ORCID Profile Kang Ning doi: https://doi.org/10.1101/2025.01.15.633278 Haohong Zhang 1 Key Laboratory of Molecular Biophysics of the Ministry of Education, Hubei Key Laboratory of Bioinformatics and Molecular-imaging, Center of AI Biology, Department of Bioinformatics and Systems Biology, College of Life Science and Technology, Huazhong University of Science and Technology , Wuhan 430074, Hubei, China Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Haohong Zhang Zixin Kang 1 Key Laboratory of Molecular Biophysics of the Ministry of Education, Hubei Key Laboratory of Bioinformatics and Molecular-imaging, Center of AI Biology, Department of Bioinformatics and Systems Biology, College of Life Science and Technology, Huazhong University of Science and Technology , Wuhan 430074, Hubei, China Find this author on Google Scholar Find this author on PubMed Search for this author on this site Yuli Zhang 1 Key Laboratory of Molecular Biophysics of the Ministry of Education, Hubei Key Laboratory of Bioinformatics and Molecular-imaging, Center of AI Biology, Department of Bioinformatics and Systems Biology, College of Life Science and Technology, Huazhong University of Science and Technology , Wuhan 430074, Hubei, China Find this author on Google Scholar Find this author on PubMed Search for this author on this site Ronghua Yang 2 Dovetree synbio Ltd. , 215 Qingnian street, Shenyang, China Find this author on Google Scholar Find this author on PubMed Search for this author on this site Kang Ning 1 Key Laboratory of Molecular Biophysics of the Ministry of Education, Hubei Key Laboratory of Bioinformatics and Molecular-imaging, Center of AI Biology, Department of Bioinformatics and Systems Biology, College of Life Science and Technology, Huazhong University of Science and Technology , Wuhan 430074, Hubei, China Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Kang Ning For correspondence: ningkang{at}hust.edu.cn Abstract Full Text Info/History Metrics Supplementary material Preview PDF Abstract Microbiome analysis has traditionally relied on taxonomic abundance tables, which, while effective, often constrain the exploration of deeper contextual relationships. In this study, we present MGM 2.0, a novel framework that applies advanced natural language processing (NLP) techniques to microbiome research. By reimagining microbiome samples as sentences and microbial species as words, MGM 2.0 enabled the extraction of nuanced patterns and relationships. The model demonstrated robust predictive performance in identifying exogenous species colonization (AUROC = 0.86). Additionally, through prompt-guided microbiome data generation, MGM 2.0 produced realistic microbial profiles conditioned on disease labels. The framework further revolutionized donor selection in fecal microbiota transplantation (FMT) by framing it as a sequence-to-sequence prediction task, enabling the prediction of post-transplantation community compositions and the identification of super donors for personalized treatments (average increase in C2R = 0.52). This innovative integration of NLP and microbiome science provides a versatile toolkit for predictive modeling, data generation, and personalized medicine. Highlights Introduced MGM 2.0, a generative language model utilizing sentence-like representation for microbiome analysis and generation. Demonstrated that sentence-like representation preserves sample distinctions, enabling accurate microbial sample classification tasks, such as colonization prediction. Generated realistic, disease-specific microbiome profiles using a prompt-guided approach, validated by a novel “Microbiome Turing Test.” Applied MGM 2.0 to fecal microbiota transplantation (FMT) donor selection, accurately predicting post-transplant community compositions and identifying potential “super donors” for personalized treatment strategies. Introduction Microbiome analysis has been instrumental in elucidating the composition, functions, and interactions of microbial communities across diverse environments [ 1 , 2 ]. Traditionally, taxonomic profiling abundance tables have served as the primary data representation, summarizing the relative proportions of microbial taxa within a sample [ 3 , 4 ]. While this approach effectively captures quantitative relationships, it often neglects the intricate contextual dependencies between microbial taxa and their communities, which restricts the potential of microbiome data to drive predictive and functional insights due to its sparsity, overdispersion and high-dimensionality [ 5 - 7 ]. Recently, several methods implemented phylogenetic information as a restriction to enhance microbial analysis [ 8 - 10 ]. However, despite these methods improve the performance of prediction task, the high-level task like data generati did not benefit from the integration of information from such approach. Recent advancements in natural language processing (NLP) offer a promising avenue to overcome these limitations [ 11 ]. By treating tabular data through the lens of sentence-like representations, analogous to text in NLP, we can leverage sophisticated computational frameworks that excel at modeling context and relationships [ 12 - 14 ]. In our previous study, we introduced the Microbial General Model (MGM), a microbiome foundation model that transforms species abundance tables into sentence-like representations, treating microbiome samples as “sentences” and microbial taxa as “words” [ 15 ]. MGM was pretrained on a large-scale corpus of 260,000 microbiome samples, enabling it to encode rich contextual information and capture complex inter-species interactions. This foundational framework provided a significant leap in microbiome research by leveraging state-of-the-art natural language processing (NLP) techniques such as self-supervised learning and attention mechanisms. Building on this foundation, we developed MGM 2.0, which enhances the generative capabilities of microbiome modeling. This upgraded framework introduces two transformative innovations: (1) generating realistic microbiome abundance profiles from prompts, facilitating controlled in silico experiments and simulation, and (2) predicting post-transplant microbiome compositions by modeling donor and recipient communities in fecal microbiota transplantation (FMT). By incorporating advanced NLP techniques, including prompt-based learning and sequence-to-sequence generation, MGM 2.0 transcends static data representations to enable dynamic, context-aware applications. These advancements hold immense potential for personalized medicine, ecological forecasting, and synthetic biology, highlighting the versatility and transformative impact of NLP-driven approaches in microbiome research. Results NLP techniques enabled deep understanding of microbial community MGM 2.0 leverages natural language processing (NLP) techniques to redefine microbial community analysis. By employing sentence-like data representations, it transforms microbial abundance data into discrete input formats through genus-level normalization, tokenization, and the ranking of taxa by relative abundance ( Fig. 1a ). Traditional machine learning algorithms have demonstrated strong performance in modeling microbial abundance tables, including linear methods such as LASSO [ 16 ], tree-based methods like Random Forests (RF) [ 17 ], and deep learning approaches like fully connected neural networks [ 18 ]. While these models primarily rely on supervised learning to capture inter-sample distinctions in labeled datasets [ 19 , 20 ] , MGM 2.0 employs self-supervised strategies inspired by NLP to focus on individual samples and the intricate structures within microbial communities ( Fig. 1b ) [ 21 , 22 ]. Download figure Open in new tab Figure 1. Comparation between the taxonomic profiling abundance table and sentence-like representation of microbial community. a. Tabular-to-sequence transformation process. b . Common methods and pretraining strategy to handle the tabular or sequential data. c . Specific methods on common microbiome problems and whether it benefits from pretraining. The integration of NLP methodologies enables common microbiome tasks, such as sample classification and biomarker discovery, to align with analogous tasks in NLP, including [ 23 ] and attention weight analysis [ 24 ]. Since these tasks focus on inter-sample differences, both supervised and self-supervised pretraining can enhance model performance. For generative tasks, MGM 2.0 demonstrates a notable advantage over traditional methods such as statistical models and Generative Adversarial Networks (GANs) [ 25 ]. These conventional approaches are constrained by their focus on the compositional characteristics of individual samples and derive limited benefits from supervised pretraining. In contrast, MGM 2.0 utilizes prompt-guided generation, leveraging pretrained models that capture microbial community structures without biases introduced by labels [ 21 ]. This capability is particularly powerful for generating realistic microbiome profiles conditioned on specific prompts, enabling in silico experimentation and hypothesis testing. Additionally, NLP techniques facilitate the prediction of community-level interventions through sequence-to-sequence modeling, addressing challenges that traditional abundance table-based approaches struggle to overcome. By incorporating NLP-driven methodologies, MGM 2.0 sets a new standard for microbial community modeling, offering transformative capabilities for predictive and generative microbiome research ( Fig. 1c , Supplementary Table 1 ). View this table: View inline View popup Table 1. Microbiome Turing Test of Different Methods. This table summarizes the performance metrics of three different methods: MGM, MB-GAN, and Random, along with a baseline for comparison. Each metric is reported with its mean value and 95% confidence interval where applicable. Sentence-like representation effectively captures microbial variations across samples To assess whether our sentence-like representation method retains the differential information between samples, we conducted an evaluation using a dataset from an exogenous species colonization experiment [ 26 ]. This dataset comprised 24 individuals, each subjected to antibiotic intervention and subsequently exposed to Enterococcus faecium ( E. faecium ) as the exogenous species. We aimed to evaluate whether sentence-like representation could effectively capture differential information across microbial communities by comparing the performance of two models: a RF model trained on the original species abundance table, and a MGM sentence classification model, which was applied to both classification and regression tasks. In the classification task, where the goal was to predict whether E. faecium would successfully colonize a given sample, the MGM model achieved comparable performance to the RF model, with an accuracy of 0.86 versus 0.83, respectively, based on five-fold cross-validation ( Fig 2a, b ). This result suggests that the sentence-like representation employed by the MGM model effectively preserves the essential differential information between samples, enabling it to capture variations in microbial communities despite the transformation from raw abundance data to sequence-based representations. Download figure Open in new tab Figure 2. The predicted colonization outcomes of E. faecium . a . ROC curve of MGM sentence classification models in binary classification (permissive vs. resistance) of the colonization outcomes of E. faecium . b . ROC curve of random forest models in binary classification (permissive vs. resistance) of the colonization outcomes of E. faecium . c . Correlation of MGM sentence classification models in regression of the colonization outcomes of E. faecium . d . Correlation of RF models in regression of the colonization outcomes of E. faecium . The regression task, aimed at predicting the post-intervention abundance of E. faecium , produced even more striking results. Here, the MGM model significantly outperformed the RF model, achieving an R 2 value of 0.22 compared to 0.03 for RF ( Fig 2c, d ). This outcome indicates that, despite the sentence-like transformation losing some of the original abundance information, the MGM model was still able to capture sufficient underlying structure to predict E. faecium abundance post-intervention. These findings demonstrate that sentence-like representation holds promise as an effective approach for colonization prediction tasks. The MGM model successfully retained key differential information, allowing it to perform robustly even when direct abundance values were not preserved. Prompt-guided microbiome generation and evaluation via a Microbiome Turing Test To showcase the generative capabilities of MGM 2.0, we developed a prompt-guided pipeline to synthesize realistic microbiome abundance profiles. Given that our sentence-like encoding omits original relative abundance information, we implemented a reconstructor network to convert generated rank sequences back into abundance tables. The generation process involved augmenting the model’s vocabulary with label tokens appended after the beginning-of-sequence token (’’) for each sample, followed by fine-tuning using next-token prediction. For example, the sentence-like representation of a sample labeled ‘CRC’ included a ‘CRC’ token following the ‘’ token. During generation, each label token served as a prompt, generating microbiome sequences with a length consistent with the original samples. These sequences were then transformed into abundance tables using the reconstructor ( Fig. 3a , Methods ). Download figure Open in new tab Figure 3. Evaluation of the prompt-guided MGM generative model. a . MGM generative model pipeline. b . Perplexity of different datasets. Perplexity in language models measures the degree to which the model considers a sentence to be realistic. c . Beta diversity analysis of the reconstructed model, based on Bray-Curtis distance. d . ROC-AUC curves of the classifier for gradient-generated data. e . UMAP dimensionality reduction plots of embeddings for three data types (train, test, generated). The ARI indices represent clustering based on data type and disease type, with clustering labels obtained through a Gaussian Mixture Model. We utilized 6,004 gut microbiota samples representing 17 diseases from the GMrepo database [ 27 ] to train and evaluate our generative model. The dataset was split into equal training and testing sets using hierarchical sampling. Initial evaluation focused on perplexity (PPL), a measure of how well the model predicts a sequence. Generated samples exhibited the lowest perplexity compared to training, external test, and randomly generated samples ( Fig. 3b ), indicating high realism. Beta diversity analysis, based on Bray-Curtis distance, further confirmed that the distribution of generated samples closely mirrored that of real data ( Fig. 3c ). To assess the biological relevance of the generated data, we trained a sentence classification model for the 17 diseases. This classifier achieved a ROC-AUC of 0.98 on the test data and 0.95 on the generated data, demonstrating that the generated samples retained disease-specific information. We also generated varying numbers of samples per disease (50, 100, 200, and 500) and found consistent performance of the disease classifier across these gradient-generated datasets ( Fig. 3d ), confirming the robustness of our generative model. UMAP dimensionality reduction of embeddings further revealed that generated samples clustered more strongly by disease type (ARI=0.085) than by data type (train, test, generated; ARI=5.4E-4), highlighting the model’s ability to capture disease-specific characteristics ( Fig. 3e ). To comprehensively evaluate the generative capabilities of MGM 2.0, we developed a novel assessment framework: the Microbiome Turing Test. This framework consists of two phases: evaluating the statistical properties of the generated data and assessing its biological significance. In the first phase, we examined the statistical properties of the generated microbiome data. Alpha diversity was calculated using Shannon’s index to assess species richness and evenness within samples. The Mean Absolute Error (MAE) between the relative abundances of real and generated samples quantified the average deviation, while Cosine Similarity measured the resemblance in community composition by evaluating the cosine of the angle between abundance vectors. Sparsity was analyzed by comparing the proportion of zero entries in abundance matrices, reflecting the similarity of sparsity patterns in real and generated data. Finally, Spearman Correlation coefficients between taxa abundances were computed to capture monotonic relationships within microbial communities. The second phase focused on evaluating the biological relevance of the generated data. A random forest classifier trained on real data achieved high ROC-AUC scores when applied to generated data, confirming the retention of disease-specific microbial patterns. Additionally, a classifier trained on the generated data exhibited significant overlap in the top 500 important microbes with classifiers trained on real data, further validating the preservation of key biomarkers. Relative network analysis was conducted to explore functional groupings and interactions within the microbiome. Co-occurrence networks were constructed, with nodes representing microbial taxa and edges denoting significant associations. Metrics such as network density (ratio of observed to possible edges), average degree (typical connectivity per node), and modularity (extent of distinct community structures) were used to compare the generated and real data. To benchmark MGM 2.0, we compared its performance against MB-GAN, a GAN-based microbiome data generation approach [ 28 ]. Across multiple metrics of the Microbiome Turing Test, MGM 2.0 demonstrated superior performance ( Table 1 ). Unlike MB-GAN, which requires retraining for each disease, MGM 2.0 leverages a prompt-guided approach, enabling efficient and flexible data generation conditioned on specific prompts. This advantage significantly enhances its utility for targeted microbiome research and in silico experimentation. Question answering strategy enabled predicting donor fitness in fecal microbiota transplantation To investigate the MGM 2.0’s generative capabilities for microbiome applications, we addressed donor selection in fecal microbiota transplantation (FMT) by predicting community-level perturbations. We framed FMT as a question-answering task, analogous to those in natural language processing. Specifically, recipient and donor sequence representations were concatenated to form the input query, and the model generated a predicted post-transplantation community composition using a sequence-to-sequence (seq2seq) approach. For example, given an FMT triad of recipient 1, donor 1, and post-transplantation sample 1, the input query consisted of the concatenated sentence-like representations of recipient 1 and donor 1, separated by a ‘sep’ (separate) token, with the sentence-like representation of post-transplantation sample 1 serving as the target sequence ( Fig. 4a , Supplementary Table 1 ). We fine-tuned the model on a dataset of 228 FMT experimental groups [ 29 ] and validated it using two external datasets: Khanna et al. (15 FMT pairs from inflammatory bowel disease (IBD) patients) [ 30 ], and Goyal et al. (38 FMT pairs from Clostridioides difficile infection (CDI) patients) [ 31 ]. Performance was evaluated using the ROUGE-1 metric, achieving consistent scores exceeding 0.6 across training, validation, and testing datasets, with particularly strong performance observed on the Khanna et al. dataset ( Fig. 4b ). Download figure Open in new tab Figure 4. Pipeline and evaluation of the community-level intervention prediction. a . The FMT experiment could be abstracted as a question-answering task. b . Rouge-1 evaluation on both inner and external dataset. c . C2R distribution of in silico reassignment of donors to different recipients. The large red points indicate the original donor-recipient pairs. d . SHAP summary plot highlighting genus-level contributors to C2R values. e . Waterfall plot explaining the SHAP values for the donor from FMT14. To assess our model’s utility for donor selection, we performed an additional validation using the Khanna et al. dataset. We systematically reassigned donors to different recipients within this dataset to generate alternative transplantation scenarios. The model predicted the resulting community composition for each new recipient-donor pairing. Community engraftment success was quantified using the Community-to-Recipient (C2R) metric. While the originally paired donors exhibited high C2R values, our analysis revealed that they were not always the optimal choice. Notably, the donor from FMT14 consistently yielded the highest C2R values across multiple recipient groups, suggesting potential as a “super donor” ( Fig. 4c , Supplementary Fig. 1 , Average increased C2R = 0.52). Conversely, the donor from FMT11 consistently resulted in the lowest C2R values when transplanted to other recipients, although it performed well in the original pairing ( Fig. 4c , Supplementary Fig 1 , Average increased C2R = -2.20). To investigate the factors contributing to donor effectiveness, we developed an LDA model to classify donors as having positive or negative impacts on C2R outcomes and employed SHapley Additive exPlanations (SHAP) [ 32 ] analysis to identify genus-level drivers influencing FMT success ( Fig. 4d, e ; Supplementary Fig. 2a-e ). This analysis revealed that donors with higher relative abundances of genera such as Desulfovibrio and Bifidobacterium were strongly associated with elevated C2R values, indicating enhanced community engraftment potential. Conversely, genera such as Bilophila and Holdmania were negatively associated with C2R values, suggesting they may hinder successful transplantation ( Fig. 4d ). Notably, previous studies have suggested that utilizing donor feces rich in Bifidobacterium can stimulate the recipient’s microbiota to recover from decreased diversity to levels comparable to the donor’s microbiota [ 33 ]. A detailed examination of the donor from FMT14, visualized using a waterfall plot, showed that Bifidobacterium and Faecalibacterium had a significant positive impact on the model output (+0.08 and +0.07, respectively) and were present at relative abundances of 0.033 and 0.202, respectively ( Fig. 4e ). These findings suggest that the enrichment of these genera may be critical to this donor as a “super donor.” These results reinforce the importance of specific microbial taxa in FMT success and underscore the potential of MGM 2.0 to provide interpretable and actionable insights into microbiome-based therapeutic strategies. By leveraging a question-answering framework inspired by natural language processing, MGM 2.0 offers a scalable, accurate, and interpretable method for advancing microbiome-based therapeutic interventions. These findings highlight its promise in optimizing donor selection and improving outcomes in FMT. Discussion This study introduces MGM 2.0, a generative language model for microbiome analysis and generation. By transforming microbiome data into sentence-like representations, MGM 2.0 enables the application of advanced NLP techniques, facilitating deeper contextual understanding and enhanced predictive capabilities. Our findings demonstrate the efficacy of MGM 2.0 across diverse tasks. The sentence-like representation preserves crucial sample-specific information, achieving classification accuracy comparable to traditional methods while significantly enhancing regression performance. This showcases the model’s ability to capture and utilize nuanced patterns and dependencies within microbial communities, underscoring the value of adapting NLP methodologies to microbiome research. A key innovation of MGM 2.0 lies in its robust generative capabilities for microbiome data generation. Through prompt-guided generation, the model produces realistic, biologically meaningful profiles conditioned on specific contexts, such as disease states. Fine-tuning a pre-trained model, combined with a reconstructor network, enabled us to generate profiles that closely mimic real-world distributions. The Microbiome Turing Test—a comprehensive evaluation framework incorporating biological significance and data distribution metrics—validated the superiority of MGM 2.0 over methods like MB-GAN. This capability is particularly valuable for generating synthetic datasets, which can augment limited real-world data, facilitate machine learning model development, and explore hypothetical microbiome states. Moreover, MGM 2.0’s ability to generate condition-specific data opens new avenues for targeted in silico experiments and simulations. MGM 2.0 also addresses critical challenges in personalized medicine, particularly in donor selection for FMT. By framing this problem as a question-answering task, the model effectively integrates recipient and donor information to predict post-transplant outcomes. The high ROUGE-1 scores achieved across datasets highlight the robustness of this approach, while in silico experiments demonstrate its potential for identifying “super donors” and optimizing transplant success. This innovation supports more precise donor selection, improving patient outcomes and minimizing risks associated with FMT. Despite its strengths, one limitation of MGM 2.0 is that its effectiveness has been verified primarily in human microbiome problems, leaving its applicability to environmental microbiomes unexplored. While the contextual information derived from the sentence-like representation has proven sufficient for tasks involving human-associated microbiomes, future studies should evaluate the model’s performance in diverse environmental settings. Addressing this gap would help establish its broader utility across various ecological and industrial applications. Enhancing the Microbiome Turing Test with expanded metrics may also provide an even more rigorous evaluation of generative model performance. In conclusion, MGM 2.0 represents a significant step forward in microbiome analysis harnessing the power of generative language models to transform how microbiome data is studied and applied. By bridging microbiome research and NLP, this sentence-like paradigm introduces new opportunities for data generation, predictive modeling, and precision medicine. The innovative methodologies and applications introduced by MGM 2.0 lay the groundwork for deeper insights, more robust predictions, and groundbreaking advancements in microbiome science. Methods Datasets and Preprocessing Colonization dataset Sequence data were retrieved from the European Nucleotide Archive (ENA) under study accession number PRJEB60398. Reads were preprocessed using fastp [ 34 ] according to the original study’s protocol: (1) Reads with more than 50% of bases below quality score 19 were removed. (2) Reads containing more than 5% N bases were removed. (3) Paired-end reads were discarded if read failed to meet the above criteria. Microbial community composition was then generated using metaphlan4 [ 35 ]. GMrepo Abundance tables were downloaded from the GMrepo homepage ( https://gmrepo.humangut.info/home ). Diseases with fewer than 100 samples, along with samples labeled as “Infant,” “Premature,” and “Pregnant,” were excluded from the analysis. FMT datasets Abundance tables for the training set were obtained from the original study’s Zotero repository ( https://doi.org/10.5281/zenodo.6611040 ). Sequence data for the two external datasets were retrieved from the Sequence Read Archive (SRA) under accession numbers PRJEB19232 and PRJNA380944. Microbial community composition for these datasets was generated using Qiime2 [ 3 ]. Data Encoding In this study, we employed a sentence-like approach to represent microbial community samples, transforming them into sequence-based representations for subsequent analysis. Each microbial community sample, denoted as { x 1 , x 2 , x 3 , …, x n } consists of a set of microbial taxa, where x i represents the relative abundance of a specific genus in the sample. The vocabulary size, denoted as X , refers to the total number of unique genera observed across all samples in the dataset. Next, the relative abundances x i for each genus are normalized based on the mean and standard deviation of the MicroCorpus-260K dataset [ 15 ]. Let μ j and σ j represent the mean and standard deviation of the relative abundances of genus j across all samples in the dataset. The normalized relative abundance for genus i in a sample is calculated as: Following normalization, the genera in each sample are ranked based on their normalized relative abundance. The genus with the highest normalized abundance is assigned rank 1, and ranks decrease in descending order of abundance. The rank r i of genus ii is defined as: Finally, the ranks are tokenized into discrete input representations. Each rank r i is mapped to a unique token t i from a predefined vocabulary V ={ t 1 , t 2 , …, t x }, where X is the total number of unique genera across the entire dataset. This tokenization process transforms the ordered ranks into a sequence format, such that the microbial community sample is represented as a sequence of tokens: In this sequence, each token corresponds to the rank of genus i , capturing the relative structure of the microbial community while reducing reliance on the absolute abundance values. Model Architecture We constructed MGM 2.0 model using eight layers of transformer blocks, with each block consisting of a self-attention layer and a feed forward neural network layer. Key hyperparameters were as follows: activation function, Gaussian Error Linear Unit (GELU); attention heads per layer, eight; embedding size: 256; feed forward size: 1024. The modeling framework was implemented in PyTorch, leveraging the Huggingface Transformers library for model configuration and training [ 36 ]. The self-attention mechanism employed in each transformer layer follows the scaled dot-product attention formula: Where Q (queries), K (keys), and V (values) are the linear projections of the input, and d k is the dimensionality of the keys. This formulation allows the model to capture contextual relationships between different tokens in the sequence, enabling more effective representation learning for downstream tasks. Downstream Fine-tuning For downstream tasks, the pre-trained MGM model is fine-tuned by replacing the language modeling head with a task-specific head. All downstream tasks in this study focused on microbial community classification employed a sequence classification head, which utilized the final token (‘eos’ (end of sentence) in this study) for classification. Fine-tuning was executed using Huggingface’s Trainer API. Key hyperparameters included, learning rate, 1e-3, batch size, 50, warmup steps, 1000, weight decay, 0.001, validation split, 10% of the data. Prompt-guided Generation Process To adapt the model for prompt-guided generation, we expanded the model’s vocabulary by introducing a new token immediately following the ’’ (beginning-of-sequence) token, designated as the label token. Let the extended vocabulary be denoted as V ′ ={ t 0 , t 1, t 2, … , t X , t label }, where t 0 corresponds to the ’ ’ token, t 1, t 2, …, t X are the tokens representing the ranks of genera, and t label is the new token added to indicate the sample identity. Each sample in the microbiome dataset is then associated with a unique label token , where i corresponds to the specific microbiome sample. For each microbiome sample S i ={ x 1 , x 2 , x 3 ,…, x n }, we assign it a label token to represent the sample identity, so that the model can generate a sequence specific to that sample. This vocabulary expansion ensures that the model can not only generate microbiome sequences but also tailor the sequences to reflect the characteristics of a specific sample. Formally, the input to the model is now a sequence of tokens , where denotes the label token for the sample i and t r 1 , t r 2 ,…, t rn are the tokens representing the ranks of microbial genera in the sample. Once this vocabulary expansion is made, the model is fine-tuned using next-token prediction. The training objective is to predict the next token in the sequence given the previous tokens. Let the current sequence of tokens be t 1 , t 2 ,…, t k , and the model’s goal is to predict the next token t k +1 . The loss function used for training is the negative log-likelihood of predicting the orrect token t k +1 , given the sequence of previously generated tokens: Here, θ represents the model parameters, and P(t k +1 t 1 , t 2 ,…, t k ;)is the probability predicted by the model for the next token t k +1 conditioned on the previous tokens. This is learned by minimizing the cross-entropy loss over the entire sequence. During the generation phase, the model is provided with the label token t k +1 as the prompt, which signals the model to generate a sequence corresponding to the sample i . The model then generates a sequence of tokens , where each t rk represents the rank of the k th genus in the sample Reconstructor Architecture Once the model generates the sentence-like sequence, it is necessary to convert these sequences back into relative abundance values. To achieve this, we utilized the reconstructor network that was trained in parallel with the main model. The reconstructor’s task is to take the generated rank sequence and map it back to the original relative abundance scale. This step ensures that the final generated microbiome sequences are not only in the correct rank order but also reflect the original quantitative properties of the microbiome samples. The reconstructor network was trained using a dataset of paired rank-encoded and abundance-encoded samples, learning to approximate the mapping from sentence-like representation back to relative abundance values. By using this approach, we preserve the diversity and abundance distributions of the microbiome, enabling the generation of realistic, contextually accurate microbiome data. The reconstructor is a deep-learning model used to reconstruct microbial abundance. The ranked corpus after tokenizing was firstly encoded to a vector X ϵ [0,2] N ,where X i =0 if species i is absent from this sample and X i = PE ( i ) + 1 if it is present. Here, PE ( i )is the position embedding from the Transformer with d model = 1. The microbial composition of this sample is represented by a vector y ϵ [0,1) N , where Y i is the relative abundance of species i in this sample. This deep-learning model (ReconstructorNet) was trained to learn the map from x to y . This model is a three-layer neural network with layers size N × 2 N × N × N , using ReLU activation in the first two layers and Softmax in the final layer, optimized using Adam. The last two layers also employ residual connections to improve training stability and performance. Pytorch-Lightning was used to build and train the model. Key hyperparameters included: Learning rate: 2e-4, Batch size: 64, Validation split: 20% of the data, Early stopping based on validation loss with 10 patience. Microbiome Turing Test In the Microbiome Turing Test framework, we computed five first-order metrics and five second-order metrics. For the first-order metrics, cosine similarity, mean absolute error (MAE), Spearman correlation coefficient, alpha diversity and sparsity were calculated by comparing the generated data with the real data for each disease individually, and then averaging the results across 17 diseases to obtain the evaluation metric. Specifically, cosine similarity was computed using the cosine_similarity function from the sklearn.metrics.pairwise module, while MAE was calculated by taking the mean absolute differences between the real and generated data for each sample. Spearman correlation coefficients were determined using the spearmanr function from the scipy.stats module. Alpha diversity was assessed using Shannon entropy and Simpson index, calculated with the alpha_diversity function from the skbio.diversity module. Sparsity was evaluated using entropy, and the distributions of these entropy values were compared using the Wilcoxon signed-rank test, implemented in the scipy.stats module. The second-order metrics focused on the biological significance of the data. For ROC-AUC, a Random Forest classifier trained on the real data was used to test the generated data from each model. The overlap of biomarkers was assessed by identifying the top 500 most important features in Random Forest classifiers trained on each dataset and then calculating the number of overlapping biomarkers between the real and generated data. Finally, we constructed Pearson correlation networks from the real and generated data using the networkx library. We evaluated network properties included network density, average degree, and modularity to determine how well the generated data captured the relationships between microbial communities. These processes were implemented using libraries such as numpy, pandas, torch, scipy, sklearn, and networkx. FMT Trial’s Representation In this setup, we used a sequence-to-sequence (seq2seq) approach to model the transformation from the donor-recipient pair to the post-transplantation microbiome composition. Specifically, the input query consisted of the concatenated sentence-like representations of the recipient and donor microbiomes, with the two sets of representations separated by a special ‘sep’ token. The sentence-like representation of the post-transplantation microbiome sample was treated as the target sequence. Mathematically, the input for an FMT triad involving recipient R , donor D , and post-transplantation sample is P represented as: Where t r 1 , t r 2 ,…, t rn and t d 1 , t d 2 ,…, t dm represent the ranked tokens corresponding to the genera in the recipient and donor microbiomes. The ′ sep ′ token separates the recipient and donor representations. The target sequence is the sentence-like representation of the post-transplantation microbiome: Code Availability The code for MGM model is available at https://github.com/HUST-NingKang-Lab/MGM . Author contributions KN conceived and proposed the idea. HZ and ZK designed and developed the framework. HZ, ZK, and YZ conducted the experiments, analyzed the data, and created the visualizations. RY provided valuable computing resources. HZ, ZK, YZ, RY, and KN contributed to editing and proofreading the manuscript. All authors reviewed and approved the final version of the manuscript. Competing interest The authors declare that they have no competing interests. Ethics approval and consent to participate Not applicable. Acknowledgments This work was partially supported by the National Key R&D Program of China (Grant No. 2023YFA1800900 and 2018YFC0910502), the National Natural Science Foundation of China (Grant Nos. 32071465, 31871334, 81827901). Numerical computations were performed on the Hefei Advanced Computing Center. References 1. ↵ Andersen , R. , S.J. Chapman , and R.R.E. Artz , Microbial communities in natural and disturbed peatlands: A review . Soil Biology and Biochemistry , 2013 . 57 : p. 979 – 994 . OpenUrl CrossRef 2. ↵ Integrative , H.M.P.R.N.C. , The Integrative Human Microbiome Project . Nature , 2019 . 569 ( 7758 ): p. 641 – 648 . OpenUrl CrossRef PubMed 3. ↵ Caporaso , J.G. , et al. , QIIME allows analysis of high-throughput community sequencing data . Nat Methods , 2010 . 7 ( 5 ): p. 335 – 6 . OpenUrl CrossRef PubMed Web of Science 4. ↵ Kuczynski , J. , et al. , Experimental and analytical tools for studying the human microbiome . Nat Rev Genet , 2011 . 13 ( 1 ): p. 47 – 58 . OpenUrl CrossRef PubMed 5. ↵ Abegaz , F. , et al. , A strategy for differential abundance analysis of sparse microbiome data with group-wise structured zeros . Sci Rep , 2024 . 14 ( 1 ): p. 12433 . OpenUrl CrossRef PubMed 6. Le Cao , K.A. , et al. , MixMC: A Multivariate Statistical Framework to Gain Insight into Microbial Communities . PLoS One , 2016 . 11 ( 8 ): p. e0160169 . OpenUrl CrossRef PubMed 7. ↵ Mallick , H. , et al. , Multivariable association discovery in population-scale meta-omics studies . PLoS Comput Biol , 2021 . 17 ( 11 ): p. e1009442 . OpenUrl CrossRef PubMed 8. ↵ Wang , Y. , et al. , A novel deep learning method for predictive modeling of microbiome data . Brief Bioinform , 2021 . 22 ( 3 ). 9. Wang , B. , et al. , DeepPhylo: Phylogeny-Aware Microbial Embeddings Enhanced Predictive Accuracy in Human Microbiome Data Analysis . Adv Sci (Weinh) , 2024 . 11 ( 45 ): p. e2404277 . OpenUrl CrossRef 10. ↵ Reiman , D. , et al. , PopPhy-CNN: A Phylogenetic Tree Embedded Architecture for Convolutional Neural Networks to Predict Host Phenotype From Metagenomic Data . IEEE J Biomed Health Inform , 2020 . 24 ( 10 ): p. 2993 – 3001 . OpenUrl CrossRef PubMed 11. ↵ Ma , Q. , et al. , Harnessing the deep learning power of foundation models in single-cell omics . Nat Rev Mol Cell Biol , 2024 . 25 ( 8 ): p. 593 – 594 . OpenUrl PubMed 12. ↵ Yang , F. , et al. , scBERT as a large-scale pretrained deep language model for cell type annotation of single-cell RNA-seq data . Nature Machine Intelligence , 2022 . 4 ( 10 ): p. 852 – 866 . OpenUrl CrossRef 13. Theodoris , C.V. , et al. , Transfer learning enables predictions in network biology . Nature , 2023 . 618 ( 7965 ): p. 616 – 624 . OpenUrl CrossRef PubMed 14. ↵ Cui , H. , et al. , scGPT: toward building a foundation model for single-cell multi-omics using generative AI . Nat Methods , 2024 . 15. ↵ Zhang , H. , et al. , MGM as a large-scale pretrained foundation model for microbiome analyses in diverse contexts . bioRxiv , 2025 : p. 2024.12.30.630825. 16. ↵ Tibshirani , R. , Regression Shrinkage and Selection Via the Lasso . Journal of the Royal Statistical Society: Series B (Methodological) , 1996 . 58 ( 1 ): p. 267 – 288 . OpenUrl CrossRef Web of Science 17. ↵ Tin Kam , H. , Random decision forests , in Proceedings of 3rd International Conference on Document Analysis and Recognition . 1995 . p. 278 - 282 vol .1. 18. ↵ LeCun , Y. , Y. Bengio , and G. Hinton , Deep learning . Nature , 2015 . 521 ( 7553 ): p. 436 – 44 . OpenUrl CrossRef PubMed 19. ↵ Tan , C. , et al. , A Survey on Deep Transfer Learning , in Artificial Neural Networks and Machine Learning -- ICANN 2018 . 2018 . p. 270 --279. 20. ↵ Segev , N. , et al. , Learn on Source, Refine on Target: A Model Transfer Learning Framework with Random Forests . IEEE Transactions on Pattern Analysis and Machine Intelligence , 2017 . 39 ( 9 ): p. 1811 – 1824 . OpenUrl CrossRef 21. ↵ Radford , A. , et al. Language Models are Unsupervised Multitask Learners . 2019 . 22. ↵ Devlin , J. , et al. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , Volume 1 ( Long and Short Papers ). 2019 . 23. ↵ Li , Q. , et al. , A Survey on Text Classification: From Traditional to Deep Learning . ACM Trans. Intell. Syst. Technol ., 2022 . 13 ( 2 ): p. Article 31. 24. ↵ Vaswani , A. , et al. , Attention is all you need , in Proceedings of the 31st International Conference on Neural Information Processing Systems . 2017 , Curran Associates Inc .: Long Beach, California, USA . p. 6000 – 6010 . 25. ↵ Goodfellow , I. , et al. , Generative adversarial networks . Commun. ACM , 2020 . 63 ( 11 ): p. 139 – 144 . OpenUrl CrossRef 26. ↵ Wu , L. , et al. , Data-driven prediction of colonization outcomes for complex microbial communities . Nat Commun , 2024 . 15 ( 1 ): p. 2406 . OpenUrl CrossRef PubMed 27. ↵ Wu , S. , et al. , GMrepo: a database of curated and consistently annotated human gut metagenomes . Nucleic Acids Res , 2020 . 48 ( D1 ): p. D545 – D553 . OpenUrl CrossRef PubMed 28. ↵ Rong , R. , et al. , MB-GAN: Microbiome Simulation via Generative Adversarial Network . Gigascience , 2021 . 10 ( 2 ). 29. ↵ Schmidt , T.S.B. , et al. , Drivers and determinants of strain dynamics following fecal microbiota transplantation . Nat Med , 2022 . 28 ( 9 ): p. 1902 – 1912 . OpenUrl CrossRef PubMed 30. ↵ Goyal , A. , et al. , Safety, Clinical Response, and Microbiome Findings Following Fecal Microbiota Transplant in Children With Inflammatory Bowel Disease . Inflamm Bowel Dis , 2018 . 24 ( 2 ): p. 410 – 421 . OpenUrl CrossRef PubMed 31. ↵ Khanna , S. , et al. , Changes in microbial ecology after fecal microbiota transplantation for recurrent C . difficile infection affected by underlying inflammatory bowel disease. Microbiome , 2017 . 5 ( 1 ): p. 55 . OpenUrl PubMed 32. ↵ Lundberg , S.M. and S.-I. Lee , A unified approach to interpreting model predictions , in Proceedings of the 31st International Conference on Neural Information Processing Systems . 2017 , Curran Associates Inc .: Long Beach, California, USA . p. 4768 – 4777 . 33. ↵ Mizuno , S. , et al. , Bifidobacterium-Rich Fecal Donor May Be a Positive Predictor for Successful Fecal Microbiota Transplantation in Patients with Irritable Bowel Syndrome . Digestion , 2017 . 96 ( 1 ): p. 29 – 38 . OpenUrl CrossRef PubMed 34. ↵ Chen , S. , et al. , fastp: an ultra-fast all-in-one FASTQ preprocessor . Bioinformatics , 2018 . 34 ( 17 ): p. i884 – i890 . OpenUrl CrossRef PubMed 35. ↵ Blanco-Miguez , A. , et al. , Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4 . Nat Biotechnol , 2023 . 41 ( 11 ): p. 1633 – 1644 . OpenUrl CrossRef PubMed 36. ↵ Wolf , T. , Huggingface’s transformers: State-of-the-art natural language processing . arXiv preprint arXiv: 1910.03771 , 2019 . View the discussion thread. Back to top Previous Next Posted January 19, 2025. Download PDF Supplementary Material Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Towards a Generative Paradigm for Large-scale Microbiome Analysis by Generative Language Model Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Towards a Generative Paradigm for Large-scale Microbiome Analysis by Generative Language Model Haohong Zhang , Zixin Kang , Yuli Zhang , Ronghua Yang , Kang Ning bioRxiv 2025.01.15.633278; doi: https://doi.org/10.1101/2025.01.15.633278 Share This Article: Copy Citation Tools Towards a Generative Paradigm for Large-scale Microbiome Analysis by Generative Language Model Haohong Zhang , Zixin Kang , Yuli Zhang , Ronghua Yang , Kang Ning bioRxiv 2025.01.15.633278; doi: https://doi.org/10.1101/2025.01.15.633278 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Bioinformatics Subject Areas All Articles Animal Behavior and Cognition (7642) Biochemistry (17715) Bioengineering (13907) Bioinformatics (42003) Biophysics (21470) Cancer Biology (18624) Cell Biology (25533) Clinical Trials (138) Developmental Biology (13390) Ecology (19935) Epidemiology (2067) Evolutionary Biology (24356) Genetics (15617) Genomics (22529) Immunology (17753) Microbiology (40432) Molecular Biology (17200) Neuroscience (88681) Paleontology (667) Pathology (2840) Pharmacology and Toxicology (4828) Physiology (7653) Plant Biology (15161) Scientific Communication and Education (2046) Synthetic Biology (4304) Systems Biology (9826) Zoology (2271)

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00