Integrative Analysis of Multi-Omics Data and Gut Microbiota Composition Reveals Prognostic Subtypes and Predicts Immunotherapy Response in Colorectal Cancer Using Machine Learning

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Colorectal cancer (CRC) exhibits substantial heterogeneity in molecular subtypes and clinical outcomes. We performed an integrative analysis of multi-omics data from 274 CRC patients to investigate the impact of gut microbiota composition on prognosis, identify novel subtypes, and develop a machine learning-based prognostic model. Our microbiome analysis revealed significant differences between CRC and normal tissues. Multi-omics clustering identified two major CRC subtypes, CS1 and CS2, with distinct molecular characteristics and survival outcomes. We developed the Multi-Omics Integrative Clustering and Machine Learning Score (MCMLS) model, which demonstrated strong prognostic value in predicting patient survival and outperformed existing models. The MCMLS low-score group exhibited higher immune cell infiltration, increased metabolic pathway activity, and potentially better immunotherapy response. In contrast, the MCMLS high-score group showed higher mutation burden, fibroblast infiltration, and enrichment of cell adhesion and migration pathways. Bacterial analysis revealed differentially abundant bacteria associated with prognosis. Importantly, MCMLS consistently predicted immunotherapy response across six independent datasets. Our findings highlight the complex interplay between the gut microbiome, tumor microenvironment, and immune landscape in CRC, providing valuable insights for improving patient stratification and personalized treatment strategies.
Full text 148,039 characters · extracted from preprint-html · click to expand
Integrative Analysis of Multi-Omics Data and Gut Microbiota Composition Reveals Prognostic Subtypes and Predicts Immunotherapy Response in Colorectal Cancer Using Machine Learning | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Integrative Analysis of Multi-Omics Data and Gut Microbiota Composition Reveals Prognostic Subtypes and Predicts Immunotherapy Response in Colorectal Cancer Using Machine Learning Jun Wang, Cong Yuan, Bo Tang, Juan Liu, Ke Pu This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5669193/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 12 Jul, 2025 Read the published version in Scientific Reports → Version 1 posted 4 You are reading this latest preprint version Abstract Colorectal cancer (CRC) exhibits substantial heterogeneity in molecular subtypes and clinical outcomes. We performed an integrative analysis of multi-omics data from 274 CRC patients to investigate the impact of gut microbiota composition on prognosis, identify novel subtypes, and develop a machine learning-based prognostic model. Our microbiome analysis revealed significant differences between CRC and normal tissues. Multi-omics clustering identified two major CRC subtypes, CS1 and CS2, with distinct molecular characteristics and survival outcomes. We developed the Multi-Omics Integrative Clustering and Machine Learning Score (MCMLS) model, which demonstrated strong prognostic value in predicting patient survival and outperformed existing models. The MCMLS low-score group exhibited higher immune cell infiltration, increased metabolic pathway activity, and potentially better immunotherapy response. In contrast, the MCMLS high-score group showed higher mutation burden, fibroblast infiltration, and enrichment of cell adhesion and migration pathways. Bacterial analysis revealed differentially abundant bacteria associated with prognosis. Importantly, MCMLS consistently predicted immunotherapy response across six independent datasets. Our findings highlight the complex interplay between the gut microbiome, tumor microenvironment, and immune landscape in CRC, providing valuable insights for improving patient stratification and personalized treatment strategies. Biological sciences/Cancer Biological sciences/Immunology Biological sciences/Microbiology Biological sciences/Systems biology Health sciences/Gastroenterology Colorectal cancer Multi-omics data integration Gut microbiota composition Machine learning-based prognostic model Immunotherapy response prediction Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Introduction Colorectal cancer (CRC) ranks as the third most commonly diagnosed cancer and the second leading cause of cancer-related mortality worldwide, with approximately 1.9 million new cases and 935,000 deaths reported in 2023 [ 1 , 2 ] . Despite advances in screening strategies and treatment modalities, patient outcomes remain suboptimal, with 5-year survival rates of approximately 65% [ 3 , 4 ] . The development of effective personalized treatment approaches is hindered by the heterogeneous nature of CRC, which is characterized by diverse molecular subtypes with varying clinical outcomes [ 5 , 6 ] . Recent research has established the tumor microenvironment (TME) as a critical determinant in CRC progression, therapeutic response, and patient prognosis [ 7 , 8 ] . The TME represents a complex ecosystem comprising various cellular components, including immune cells, fibroblasts, and endothelial cells, as well as non-cellular elements such as the extracellular matrix [ 9 ] . Additionally, the gut microbiome has emerged as a key contributor to CRC pathogenesis and treatment response. Multiple studies have demonstrated that specific microbial signatures are associated with CRC development, progression, and clinical outcomes. For example, Mouradov et al. identified three distinct oncomicrobial community subtypes (OCSs) with unique clinico-molecular features and prognostic implications, where OCS1 (characterized by Fusobacterium and oral pathogens) was associated with right-sided, high-grade, MSI-high tumors, while OCS2 (dominated by Firmicutes/Bacteroidetes) conferred better survival outcomes in microsatellite-stable tumors [ 10 ] . Similarly, Ibeanu et al. found that several bacterial species, including Fusobacterium nucleatum, promote tumorigenesis through inflammation induction and immune system modulation, making them promising candidates for targeted interventions [ 11 ] . These findings highlight the importance of integrating microbiome analysis into CRC research to enhance our understanding of disease mechanisms and improve patient stratification. The advent of high-throughput sequencing technologies has revolutionized our understanding of CRC's molecular landscape, enabling comprehensive multi-omics profiling [ 12 , 13 ] . Integrative analysis of these multi-dimensional datasets has revealed distinct molecular subtypes with important clinical implications. The consensus molecular subtypes (CMS) classification, which categorizes CRC into four subtypes based on gene expression patterns, has emerged as a robust framework for understanding tumor biology and guiding treatment decisions [ 5 , 14 , 15 ] . However, recent studies have shown that intra-tumor heterogeneity compromises the clinical utility of transcriptomic classifications, necessitating more sophisticated approaches for patient stratification [ 16 – 19 ] . However, the incorporation of gut microbiome data into these integrative analyses has been limited, despite the growing evidence supporting the pivotal role of the gut microbiome in CRC [ 20 ] . For instance, Li et al. demonstrated that proteomic characterization of CRC response to chemoradiation and targeted therapies could identify four distinct subtypes with unique biological and therapeutic characteristics, providing valuable insights for personalized treatment [ 21 – 23 ] . Machine learning (ML) approaches have emerged as powerful tools for analyzing complex biological data and developing predictive models in cancer research [ 24 , 25 ] . ML algorithms can effectively handle high-dimensional datasets, capture non-linear relationships, and identify subtle patterns that may be overlooked by traditional statistical methods [ 9 , 26 ] . In the context of CRC, ML has been applied to various tasks, such as predicting patient survival, classifying molecular subtypes, and identifying biomarkers for treatment response [ 27 – 29 ] . The integration of multi-omics data using ML approaches has shown promise in improving the accuracy and robustness of prognostic and predictive models. For example, Wang et al. developed a model based on epigenetically regulated gene expression profiles that identified four molecular subtypes with distinct clinical and molecular features, providing a better understanding of CRC heterogeneity [ 30 , 31 ] . Despite these advancements, the incorporation of gut microbiome data into multi-omics integrative analyses has been limited, despite the growing evidence supporting the pivotal role of the gut microbiome in CRC [ 32 – 34 ] . Recent studies have demonstrated that specific gut microbial signatures are associated with CRC prognosis and response to immunotherapy [ 35 ] . For instance, Zhang et al. developed a computational model that integrated spatial multi-omics data to predict immunotherapy response in cancer, highlighting the potential of multi-omics integration for biomarker discovery [ 36 ] . Furthermore, Carson et al. found that microbiome-CRC associations may differ by racial group, emphasizing the need for large, diverse population-based studies to determine the generalizability of these associations [ 37 ] . In this study, we aimed to conduct a comprehensive integrative analysis of multi-omics data and gut microbiota composition in CRC to identify novel prognostic subtypes and develop a robust ML-based model for predicting patient survival and immunotherapy response. We hypothesized that the integration of multi-omics data, including transcriptomics, epigenomics, and genomics, along with gut microbiome profiles, would provide a more holistic understanding of CRC biology and enable the development of accurate prognostic and predictive models. To achieve this, we collected multi-omics data and gut microbiome profiles from a large cohort of 274 CRC patients and performed an integrative analysis using state-of-the-art bioinformatics and ML approaches. We employed multiple unsupervised clustering algorithms to identify distinct molecular subtypes based on the integrated multi-omics data and evaluated their associations with clinical outcomes and immune features. Furthermore, we developed a novel ML-based prognostic model, the Multi-Omics Integrative Clustering and Machine Learning Score (MCMLS), which incorporates multi-omics data and gut microbiome profiles to predict patient survival and immunotherapy response. Our findings provide valuable insights into the complex interplay between the TME, gut microbiome, and molecular features of CRC, and highlight the potential of integrating multi-omics data and ML approaches for improving patient stratification and guiding personalized treatment strategies. Materials and methods Data Acquisition and Processing The tumor microbiome data for TCGA-COAD was obtained from the MicrobiomeX database [ 38 ] , which included 497 samples (456 tumor samples and 41 adjacent normal samples). The mRNA, lncRNA, miRNA, DNA CpG methylation sites, mutated genes, and clinical data were downloaded from The Cancer Genome Atlas (TCGA) database. A total of 274 patients had complete data across all platforms with available survival information. Clinical characteristics are provided in Supplementary Table 1. Multi-omics Integrative Clustering Analysis Multi-omics integrative clustering was performed using the MOVICS package in R (version 4.3.0) [ 39 ] . For feature selection, we applied criteria established by Chu et al [ 25 ] : for mRNA, lncRNA, and miRNA expression data, the top 3000, 3000, and 5000 features were selected based on median absolute deviation (MAD) and filtered using Cox regression with p-value cutoffs of 0.01, 0.001, and 0.001, respectively. For DNA methylation data, the top 3000 features were selected based on MAD and filtered with a p-value cutoff of 0.05. Mutation features were selected using a frequency cutoff of 0.15, and bacterial expression features were selected based on the top 15 highest standard deviations. The optimal number of clusters (2–8) was determined using cluster prediction analysis, and ten different clustering algorithms were applied. Silhouette analysis was performed to evaluate clustering quality. DNA methylation beta values were converted to M values for improved signal detection. Survival differences between identified subtypes were assessed using Kaplan-Meier analysis with log-rank tests, with p < 0.05 considered statistically significant. Immune Landscape Analysis and Subtype Validation The immune landscape of identified subtypes was characterized through gene set variation analysis (GSVA) [ 40 ] using immune-related pathways. Transcriptional regulatory networks were reconstructed using the RTN package, focusing on mucin regulation and chromatin remodeling. The methylation-based tumor-infiltrating lymphocyte (MeTIL) score was calculated, and stromal and immune scores were estimated using the ESTIMATE algorithm. We assessed immune checkpoint gene expression across the subtypes and estimated immune cell composition using the CIBERSORT deconvolution algorithm [ 41 ] . For subtype validation, four independent colorectal cancer datasets (GSE17536 [ 42 ] , GSE17537 [ 43 ] , GSE72970 [ 44 ] , and GSE161158 [ 45 ] ) were downloaded from the Gene Expression Omnibus (GEO) database. Batch effects were removed using the sva package, creating a meta-validation dataset. Nearest Template Prediction (NTP) and Prediction Analysis for Microarrays (PAM) methods were used to assign subtype labels. Agreement between subtyping methods was assessed using Cohen's kappa coefficient, with values > 0.6 considered substantial agreement. Multi-Omics Integrative Clustering and Machine Learning Score (MCMLS) Model Development For prognostic model construction, we employed a machine learning approach using the TCGA dataset as the training cohort and the meta-validation dataset as the validation cohort. Data preprocessing included extracting common genes between datasets and performing standardization. Following the methodology of Chu et al [ 26 ] , 101 different machine learning models were trained using various feature selection methods, implemented in the R environment with the caret package. Model performance was evaluated using the concordance index (C-index), with higher values indicating better discrimination between high-risk and low-risk patients. The optimal plsRcox algorithm was identified based on the highest average C-index across validation cohorts. For survival analysis, patients were divided into high-risk and low-risk groups based on the median MCMLS score, with p < 0.05 considered statistically significant. Comparison of Prognostic Signatures To compare the performance of our developed prognostic signature with existing signatures, we conducted a comprehensive analysis using the TCGA dataset as the training cohort and a meta-validation dataset as the validation cohort. The gene expression data and clinical survival information were obtained for both cohorts. We collected publicly available prognostic signatures from the supplementary materials of a previous study and manually curated additional signatures from the literature [ 26 ] . For each signature, we trained a Cox proportional hazards model and calculated C-indices for both training and validation cohorts. Time-dependent area under the receiver operating characteristic curve (AUC) analysis was conducted to compare the MCMLS model with clinical risk factors (tumor stage, T stage, N stage, M stage, and gender). AUC values > 0.7 were considered to indicate good discriminatory ability. Immune Infiltration Analysis and Molecular Characterization of MCMLS Subtypes The immune infiltration landscape was analyzed using seven deconvolution algorithms: CIBERSORT, EPIC, MCPcounter, xCell, ESTIMATE, TIMER, and quanTIseq [ 46 ] . Wilcoxon rank-sum tests were used to identify significant differences in immune cell infiltration between MCMLS subtypes, with p-values adjusted for multiple testing. DNA methylation and copy number variation data were integrated with gene expression data to characterize molecular features of the subtypes.. Functional Enrichment Analysis and Metabolic Pathway Activity Assessment Gene Set Enrichment Analysis (GSEA) was performed using the "h.all.v2023.2.Hs.symbols.gmt" and "KEGG.database.symbols.gmt" gene sets [ 47 – 50 ] . Differential expression analysis between MCMLS subtypes was conducted using the limma package. GSVA was performed using the "REACTOME_metabolism.gmt" gene set to analyze metabolic pathway activities [ 40 ] . Significantly enriched pathways were identified using a threshold of p < 0.05. Mutational Landscape Analysis of MCMLS Subtypes Somatic mutation data from the TCGA cohort was analyzed using the maftools package. Mutational profiles were compared between MCMLS subtypes using mafCompare analysis. Differentially mutated genes were identified using a significance threshold of p < 0.05. Tumor mutational burden (TMB) was calculated for each sample and compared between subtypes using the Wilcoxon rank-sum test. Somatic interaction analysis was performed to investigate co-occurrence and mutual exclusivity of gene mutations. Bacterial Differential Analysis and Correlations Bacteria with significant differences between MCMLS subtypes were identified using Linear discriminant analysis Effect Size (LEfSe) with thresholds of p 2. Survival analysis for differentially abundant bacteria was conducted using Cox proportional hazards models. Samples were grouped into "Low" and "High" abundance groups using the optimal cutoff determined by the surv_cutpoint function or the median value. Correlations between selected bacteria and immune cell infiltration or metabolic pathway activity were assessed using Pearson correlation analysis. Predicting Immunotherapy Response Using the MCMLS Model To assess the predictive power of the MCMLS model for immunotherapy response, we applied the model to multiple independent immunotherapy datasets, including Melanoma-GSE91061 [ 51 ] , Melanoma-GSE100797 [ 52 ] , Melanoma-phs000452 [ 53 ] , Melanoma-PRJEB23709 [ 54 ] , Urothelial Cancer-GSE176307 [ 55 ] , and Urothelial Cancer-IMvigor210 [ 56 ] . For each dataset, we performed a comprehensive analysis pipeline. First, we preprocessed the data by reading the gene expression data and clinical information from the corresponding CSV files and standardizing the gene expression data to ensure comparability across samples. Next, we calculated the risk score for each sample using the calRS function based on the MCMLS model genes. This function fitted a Cox proportional hazards model using the selected genes and calculated the risk score as the linear combination of the gene expression values weighted by their respective coefficients. The samples were then divided into high-risk and low-risk groups based on the optimal cutoff point determined by the surv_cutpoint function. We then conducted survival analysis by generating Kaplan-Meier survival curves to compare the overall survival between the high-risk and low-risk groups, using the log-rank test to assess the statistical significance of the survival differences. To investigate the association between the MCMLS risk groups and immunotherapy response, we harmonized the response categories into multiple (CR, PR, SD, PD) categories, depending on the available information in each dataset. We created violin plots with embedded box plots to visualize the distribution of risk scores across the response categories and used the Kruskal-Wallis test to compare the risk scores between the response groups. Additionally, we generated stacked bar plots to show the percentage of patients in each response category within the high-risk and low-risk groups, using the Fisher's exact test to assess the association between the risk groups and response categorie. Finally, we created a heatmap-like representation to visualize the risk scores and corresponding response categories for each sample, ordering the samples by decreasing risk score and color-coding the response categories. We used the Kruskal-Wallis test (for multiple response categories) to compare the risk scores across the response categories. Results Microbiome Composition and Diversity in Normal Intestinal Tissue versus Colorectal Cancer Comparative analysis of microbiome composition between normal intestinal tissue and colorectal cancer (CRC) samples revealed significant differences across taxonomic levels. The α-diversity, measured by Shannon index, was consistently lower in tumor samples than normal samples at all taxonomic levels (P < 0.001; Figs. 1 A,D,G,J,M,P). At the phylum level, Proteobacteria predominated in tumor samples while Firmicutes was more abundant in normal tissues (Fig. 1 B-C). At the class level, Gammaproteobacteria showed higher representation in tumor samples compared to Betaproteobacteria in normal samples (Fig. 1 E-F). The order-level analysis demonstrated elevated Enterobacterales in tumor samples versus Caudovirales in normal tissues (Fig. 1 H-I). Family-level examination revealed Enterobacteriaceae enrichment in tumors while Myoviridae dominated in normal samples (Fig. 1 K-L). At the genus level, Escherichia was more prevalent in tumor samples compared to Klebsiella in normal tissues (Fig. 1 N-O). Species-level analysis identified Escherichia coli enrichment in tumors whereas Klebsiella pneumoniae was more abundant in normal samples (Fig. 1 Q-R). No significant differences in species-level Shannon index were observed when stratifying CRC samples by gender, age, or cancer stage (Figures S1 A-I). Multi-omics Subtyping of Colorectal Cancer Patients Based on TCGA Database Multi-omics clustering analysis of CRC patients using TCGA data identified two distinct molecular subtypes. Cluster prediction and gap statistics analysis (Fig. 2 A) supported the application of ten different clustering methods (Fig. 2 B), consistently revealing two major subtypes designated as CS1 (n = 127) and CS2 (n = 147). Silhouette scores validated the internal similarity within each subtype (Fig. 2 C). A comprehensive subtype heatmap was generated integrating mRNA, lncRNA, miRNA, DNA methylation, mutated genes, and bacterial abundance data, highlighting ten representative features for each data type that differentiated the two subtypes (Fig. 2 D). Survival analysis demonstrated significantly longer overall survival for CS1 patients compared to CS2 patients (Fig. 2 E), confirming the clinical relevance of this multi-omics classification approach. Multi-omics Subtyping Landscape and Validation in Colorectal Cancer Further characterization of the molecular differences between CS1 and CS2 subtypes was performed using gene set enrichment analysis (GSEA). The P53 signaling pathway showed higher activity in CS1, while PI3K-AKT-MTOR signaling was more active in CS2 (Fig. 3 A). Immune landscape analysis revealed higher expression of immune checkpoint genes in CS1 compared to CS2 (Fig. 3 B). Assessment of transcriptional regulatory genes and chromatin remodeling factors demonstrated differential expression patterns, with PARB, PARA, and EGFR elevated in CS2, while HIF1A, ESR2, and FOXM1 were higher in CS1 (Fig. 3 C). The subtyping approach was validated in the META-COAD cohort using the nearest template prediction algorithm, which successfully classified patients into CS1 and CS2 subtypes (Fig. 3 D). Consistent with TCGA findings, META-COAD patients in CS1 demonstrated significantly longer overall survival than those in CS2 (Fig. 3 E). Principal component analysis confirmed successful integration of four independent CRC datasets (GSE17356, GSE17537, GSE72970, and GSE161158) after batch effect removal (Figures S2A-B). Development and Prognostic Value of the Multi-Omics Integrative Clustering and Machine Learning Score(MCMLS) We developed a prognostic model for CRC by evaluating 103 machine learning algorithm combinations (Fig. 4 A). The plsRcox algorithm achieved the highest C-index (0.661), demonstrating superior performance in predicting patient survival. This algorithm identified 18 hub genes with corresponding regression coefficients (Fig. 4 B) that formed the basis of our Multi-Omics integrative Clustering and Machine Learning Score (MCMLS). Patients stratified by median MCMLS score showed significant survival differences in both TCGA-COAD (P < 0.0001; Fig. 4 C) and META-COAD cohorts (P < 0.0001; Fig. 4 D). Univariate Cox regression analysis confirmed the significant association of the 18 hub genes with patient survival in both cohorts (Figs. 4 E-F), validating their inclusion in the MCMLS model. Clinical Utility of MCMLS in Colorectal Cancer The clinical relevance of MCMLS was assessed by examining its relationship with established clinicopathological factors. MCMLS was significantly associated with age, T stage, N stage, and overall stage, but not with gender or M stage (Figs. 5 A-F). Comparative analysis with 50 previously published CRC prognostic models demonstrated that MCMLS outperformed all other models in both TCGA-COAD and META-COAD cohorts (Fig. 5 G-H). MCMLS exhibited higher predictive accuracy than traditional clinical factors including stage, T stage, N stage, M stage, and gender in the TCGA-COAD cohort (Fig. 5 I), highlighting its potential as a complementary prognostic tool for clinical decision-making. Immune Features and Metabolic Pathways Associated with MCMLS Subgroups Immune cell infiltration analysis using seven algorithms revealed that the MCMLS low-score group had higher infiltration of CD8 + T cells (TIMER algorithm) and NK cells (xCell algorithm), while the MCMLS high-score group exhibited increased fibroblast infiltration (EPIC, MCPcounter, and xCell algorithms) (Fig. 6 A). Expression analysis of 74 immune-related genes showed that the MCMLS high-score group had elevated expression of immunotherapy-inhibiting molecules (Fig. 6 B). GSEA using KEGG gene sets demonstrated that the MCMLS low-score group was enriched in oxidative phosphorylation pathways, while the high-score group showed enrichment in cell adhesion and migration pathways (Fig. 7 A). Analysis with cancer gene sets confirmed these findings and additionally showed fatty acid metabolism enrichment in the low-score group and epithelial-mesenchymal transition enrichment in the high-score group (Fig. 7 B). Gene set variation analysis of metabolic pathways revealed generally higher metabolic activity in the low-score group, particularly in lipid metabolism, fatty acid metabolism, and the TCA cycle, while the high-score group showed elevated activity in PI metabolism (Fig. 7 C). Mutation Profiles of MCMLS Subgroups Mutation analysis demonstrated that the MCMLS high-score group had significantly higher mutation burden compared to the low-score group (P = 0.0041; Fig. 8 A). The most frequently mutated genes across all samples were APC, TP53, and KRAS (Fig. 8 B). Univariate Cox analysis identified TP53, TTN, and MUC16 as significantly more frequently mutated in the high-score group (P < 0.05; Fig. 8 C). Co-expression analysis revealed that in the low-score group, TP53 mutations were mutually exclusive with mutations in RYR1, MUC5B, CSMD3, OBSCN, and TTN (P < 0.01; Fig. 8 D), while in the high-score group, APC mutations were mutually exclusive with OBSCN mutations (P < 0.01; Fig. 8 E). The mutation frequency of TP53 was significantly higher in the high-score group (70.07%) compared to the low-score group (51.82%) (Fig. 8 F). Bacterial Composition and Prognostic Significance in MCMLS Subgroups LEfSe analysis identified 53 differentially abundant bacteria between MCMLS subgroups (P 2; Fig. 9 A). Univariate Cox analysis revealed that two bacteria, s_Streptomyces_lividans and s_Pseudomonas_sp._CIP.10, were associated with worse prognosis when their abundance was higher (Fig. 9 B). These bacteria were positively correlated with fibroblasts and negatively correlated with NK cells (Fig. 9 C) and most metabolic pathways (Fig. 9 D). Their higher abundance in the MCMLS high-score group aligned with the observed lower metabolic activity in this group. MCMLS Predicts Immunotherapy Response Across Multiple Datasets To validate the predictive capacity of MCMLS for immunotherapy response, we analyzed six independent immunotherapy-treated patient cohorts across different cancer types and immunotherapy regimens. The IMvigor210 cohort comprised 119 patients with locally advanced or metastatic urothelial carcinoma treated with atezolizumab (anti-PD-L1, 1200 mg intravenously every 21 days) as first-line therapy (PMID: 27939400). The GSE176307 dataset included 103 patients with metastatic urothelial cancer treated with various immune checkpoint inhibitors (ICB) at a single academic center. The PRJEB23709 cohort consisted of metastatic melanoma patients treated with either anti-PD-1 monotherapy or combined anti-PD-1 and anti-CTLA-4 therapy. The PHS00452 dataset included melanoma patients treated with immunotherapy as part of the Melanoma Genome Sequencing Project. The GSE91061 cohort comprised 65 melanoma patients (109 samples) treated with nivolumab (anti-PD-1) therapy (PMID: 29033130), while GSE100797 included 25 melanoma patients undergoing adoptive T cell therapy after previous immunotherapy treatment (PMID: 29170503). Across all datasets, the MCMLS high-score group consistently showed significantly worse prognosis compared to the low-score group (P < 0.001; Figs. 10 A,E,I,M,Q,U). MCMLS scores were significantly associated with immunotherapy response in all datasets (P < 0.05; Figs. 10 B,F,J,N,R,V), with a gradual increase in scores observed from complete response to progressive disease (P < 0.05; Figs. 10 C,G,K,O,S,W). The proportion of responders was significantly higher in the MCMLS low-score group across all datasets (Figs. 10 D,H,L,P,T,X), demonstrating the robust predictive power of MCMLS for immunotherapy response across different cancer types and therapeutic regimens.. Discussion The present study employed a comprehensive multi-omics approach to investigate the molecular heterogeneity, prognostic factors, and immunotherapy response in colorectal cancer (CRC). By integrating data from the gut microbiome, transcriptomics, epigenomics, and genomics, we aimed to gain a deeper understanding of the complex biological processes underlying CRC development and progression. Our findings highlight the importance of considering multiple omics data types in the characterization of CRC subtypes and the development of personalized treatment strategies. We began by comparing the microbiome composition and diversity between normal intestinal tissue and CRC samples at different taxonomic levels. The results revealed significant differences in bacterial abundance and diversity, with tumor samples exhibiting lower α-diversity and distinct dominant bacterial taxa compared to normal samples. These findings align with recent studies identifying distinct oncomicrobial community subtypes in CRC with unique clinico-molecular features [ 10 , 21 , 57 ] . The enrichment of certain bacterial species in tumor samples, such as Escherichia coli and members of Proteobacteria, and the depletion of beneficial bacteria like Klebsiella pneumoniae, reflect the complex interactions between the gut microbiome and the host in CRC pathogenesis. These distinct bacterial profiles may contribute to the creation of a pro-inflammatory microenvironment that promotes tumor growth and progression, as demonstrated by recent microbiome research in CRC [ 11 , 58 , 59 ] . Our multi-omics subtyping analysis identified two major CRC subtypes (CS1 and CS2) with distinct molecular characteristics and clinical outcomes. Notably, patients in the CS1 subtype exhibited significantly longer overall survival compared to those in the CS2 subtype. These findings underscore the power of integrating multiple omics data types in capturing the complex molecular landscape of CRC and identifying clinically relevant subtypes. The molecular differences between the CS1 and CS2 subtypes provide valuable insights into the underlying biological mechanisms driving the observed differences in prognosis [ 60 , 61 ] . For example, the CS1 subtype was characterized by higher activity of the P53 signaling pathway and increased expression of immune checkpoint genes, while the CS2 subtype showed greater activity in the PI3K-AKT-MTOR pathway. These distinct molecular features align with recent research demonstrating that CRC can be classified into different subtypes with unique biological characteristics and therapeutic vulnerabilities [ 23 ] . A novel contribution of our study is the development of the Multi-Omics Integrative Clustering and Machine Learning Score (MCMLS), which represents a significant advancement over existing prognostic models. The MCMLS model demonstrated robust prognostic performance in both the TCGA-COAD and META-COAD cohorts, outperforming 50 published CRC prognostic models and traditional clinical factors. The superior performance of MCMLS can be attributed to its comprehensive integration of multi-omics data, capturing complex interactions that may not be fully represented by individual omics platforms or clinical variables. This approach mirrors recent successful strategies in other cancer types, but our application to CRC with the inclusion of microbiome data represents a novel advancement in the field [ 25 , 62 ] Another key innovation of our study is the identification of distinct immune escape phenotypes associated with different MCMLS subgroups. The MCMLS low-score group exhibited higher infiltration of CD8 + T cells (TIMER algorithm) and NK cells (xCell algorithm), indicating a more favorable immune microenvironment. In contrast, the MCMLS high-score group showed increased fibroblast infiltration and expression of immunotherapy-inhibiting molecules. These findings extend current understanding of immune escape mechanisms in CRC by revealing specific cellular patterns associated with different molecular subtypes [ 8 ] . This characterization of immune escape phenotypes provides potential targets for therapeutic intervention and may guide immunotherapy strategies for different patient subgroups. Our metabolic pathway analysis revealed that the MCMLS low-score group was enriched in oxidative phosphorylation and fatty acid metabolism pathways, while the high-score group showed enrichment in cell adhesion, migration, and epithelial-mesenchymal transition pathways. These distinct metabolic profiles suggest different energy utilization strategies between the subtypes, which may contribute to their different clinical behaviors. Recent research has highlighted the importance of metabolic reprogramming in tumor immune evasion [ 63 ] , and our findings contribute to this emerging field by linking specific metabolic signatures to immune escape phenotypes and patient outcomes in CRC. The investigation of gene mutation differences between MCMLS subgroups revealed that the MCMLS high-score group had a significantly higher mutation burden compared to the low-score group, with specific genes such as TP53, TTN, and MUC16 exhibiting higher mutation frequencies. These findings align with recent comprehensive genomic studies of CRC [ 6 ] but uniquely link these mutational profiles to immune features and prognosis. The co-expression analysis further revealed novel mutually exclusive mutation patterns, with TP53 mutations being mutually exclusive with mutations in RYR1, MUC5B, CSMD3, OBSCN, and TTN in the MCMLS-Low group, while APC mutations were mutually exclusive with OBSCN mutations in the MCMLS-High group. These patterns may reflect different evolutionary trajectories of CRC and offer new insights into disease mechanisms. Perhaps the most innovative aspect of our study is the demonstrated ability of MCMLS to predict immunotherapy response across six independent datasets. Patients in the MCMLS low-score group consistently showed better responses to immunotherapy compared to those in the high-score group, with significant associations between MCMLS scores and clinical responses (complete response, partial response, stable disease, or progressive disease). This finding is particularly important given the current limitations of immunotherapy in CRC, where benefits are largely restricted to microsatellite instability-high tumors [ 35 ] . Our model's ability to identify potential responders beyond this subset could significantly expand the application of immunotherapy in CRC, addressing a critical unmet need. Finally, our analysis of the bacterial differences between MCMLS subgroups identified 53 differentially abundant bacteria, with two species (s_Streptomyces_lividans and s_Pseudomonas_sp._CIP.10) significantly associated with worse prognosis. Their correlation with fibroblast infiltration and negative association with NK cells and metabolic pathways suggests a potential role in shaping the tumor microenvironment. This represents one of the first demonstrations of specific bacterial signatures associated with both molecular subtypes and clinical outcomes in CRC, extending recent work on the diagnostic and prognostic implications of the gut microbiome [ 34 ] . In conclusion, our study presents several novel findings that advance the field of CRC research. First, we developed a robust multi-omics-based prognostic model (MCMLS) that outperforms existing methods and captures complex biological interactions. Second, we identified distinct immune escape phenotypes associated with different molecular subtypes, providing new insights into resistance mechanisms. Third, we demonstrated that MCMLS can accurately predict immunotherapy response across multiple independent datasets, potentially expanding the application of immunotherapy in CRC. Fourth, we uncovered specific bacterial signatures associated with prognosis and immune features, highlighting the importance of the microbiome in CRC biology. These innovations collectively contribute to a more comprehensive understanding of CRC heterogeneity and lay the foundation for more effective personalized treatment strategies. Despite these advancements, our study has several limitations that should be addressed in future research. Prospective clinical validation is necessary to confirm the clinical utility of the MCMLS model. The functional roles of the identified molecular features require further experimental validation. Integration of additional omics data types, such as proteomics and metabolomics, may provide an even more comprehensive understanding of CRC biology. Nevertheless, our study represents a significant step toward precision medicine in CRC, demonstrating the power of multi-omics data integration and machine learning approaches in advancing clinical management. Conclusion In summary, our comprehensive multi-omics analysis of CRC provides novel insights into the molecular heterogeneity and biological mechanisms underlying this complex disease. The developed MCMLS model represents a promising tool for precision medicine in CRC, enabling the stratification of patients based on their molecular profiles and the prediction of prognosis and immunotherapy response. The identified molecular features and pathways offer potential targets for future therapeutic interventions and mechanistic studies. Our findings underscore the importance of integrating multi-omics data and machine learning approaches in cancer research and highlight the potential of precision medicine in improving the clinical management and outcomes of CRC patients. Abbreviations CRC, Colorectal cancer TME, Tumor Microenvironment ICIs, Immune Checkpoint Inhibitors ML, Machine learning MCMLS, Multi-Omics Integrative Clustering and Machine Learning Score MAD, Median Absolute Deviation TRNs, Transcriptional Regulatory Networks TFs, transcription factors MeTIL, Methylation-based Tumor-Infiltrating Lymphocyte NTP, Nearest Template Prediction PAM, Prediction Analysis for Microarrays C-index, Concordance Index AUROC, Area Under the Receiver Operating Characteristic curve CNV, Copy Number Variation GSEA, Gene Set Enrichment Analysis GSVA, Gene Set Variation Analysis MAF, Mutation Annotation Format TILs, Tumor-Infiltrating Lymphocytes PCA, Principal component analysis CR, Complete Response PR, Partial Response SD, Stable Disease PD, Progressive Disease Declarations Acknowledgements Not applicable. Data Availability Statement The multi-omics data used in this study are publicly available from The Cancer Genome Atlas (TCGA-COAD) database. The microbiome data were obtained from the MicrobiomeX database. The validation datasets (GSE17536, GSE17537, GSE72970, and GSE161158) are available from the Gene Expression Omnibus (GEO) database. The immunotherapy datasets used for validation of the MCMLS model (IMvigor210, GSE176307, PRJEB23709, PHS00452, GSE91061, and GSE100797) are also publicly available from their respective repositories. The processed data, analysis code, and the MCMLS model implementation are available from the corresponding authors upon reasonable request. Restrictions apply to the availability of some patient-specific data, which were used under license for the current study, and so are not publicly available. Any additional data that support the findings of this study are included in the supplementary information files or are available from the corresponding authors upon reasonable request. Competing Interests The authors have no relevant financial or non-financial interests to disclose. Funding: This work was supported by the Intramural Fund of North Sichuan Medical College (CBY21-QD31), Nanchong City Talent Development Fund (CBY23-NCR06), and Fund of Bureau of Science&Technology Nanchong City (23JCYJPT0056) CRediT authorship contribution statement KP and JL contributed to the design and conceptualization of the study. JW, CY and BT wrote the manuscript. KP, JW, CY and BT were involved in the data processing, generation, analysis, interpretation, and figure and table production for this project. JL and KP guided the writing of the manuscript, revised the manuscript and provided financial support. All authors have read and agreed to the published version of the manuscript. References Morgan E, Arnold M, Gini A, et al. Global burden of colorectal cancer in 2020 and 2040: incidence and mortality estimates from GLOBOCAN[J]. Gut, 2023, 72(2): 338-344. Chung D C, Gray D M, 2nd, Singh H, et al. A Cell-free DNA Blood-Based Test for Colorectal Cancer Screening[J]. N Engl J Med, 2024, 390(11): 973-983. Xie Y H, Chen Y X, Fang J Y. Comprehensive review of targeted therapy for colorectal cancer[J]. Signal Transduct Target Ther, 2020, 5(1): 22. Li L, Jiang D, Liu H, et al. Comprehensive Proteogenomic Profiling Reveals the Molecular Characteristics of Colorectal Cancer at Distinct Stages of Progression[J]. Cancer Res, 2024, 84(17): 2888-2910. Shin A E, Giancotti F G, Rustgi A K. Metastatic colorectal cancer: mechanisms and emerging therapeutics[J]. Trends Pharmacol Sci, 2023, 44(4): 222-236. Nunes L, Li F, Wu M, et al. Prognostic genome and transcriptome signatures in colorectal cancers[J]. Nature, 2024, 633(8028): 137-146. Bejarano L, Jordao M J C, Joyce J A. Therapeutic Targeting of the Tumor Microenvironment[J]. Cancer Discov, 2021, 11(4): 933-959. Arashi K, Nishiyama T, Hosoya M, et al. Bilateral Deafness Due to Relapsing Polychondritis with Semicircular Canal Calcification Treated With Cochlear Implantation: A Case Report[J]. Ear Nose Throat J, 2023: 1455613231215173. Li Y, Wu X, Fang D, Luo Y. Informing immunotherapy with multi-omics driven machine learning[J]. NPJ Digit Med, 2024, 7(1): 67. Mouradov D, Greenfield P, Li S, et al. Oncomicrobial Community Profiling Identifies Clinicomolecular and Prognostic Subtypes of Colorectal Cancer[J]. Gastroenterology, 2023, 165(1): 104-120. Ibeanu G C, Rowaiye A B, Okoli J C, Eze D U. Microbiome Differences in Colorectal Cancer Patients and Healthy Individuals: Implications for Vaccine Antigen Discovery[J]. Immunotargets Ther, 2024, 13: 749-774. Espinosa E, Bautista R, Larrosa R, Plata O. Advancements in long-read genome sequencing technologies and algorithms[J]. Genomics, 2024, 116(3): 110842. Hu M, Zhu J, Peng G, et al. IMOVNN: incomplete multi-omics data integration variational neural networks for gut microbiome disease prediction and biomarker identification[J]. Brief Bioinform, 2023, 24(6). Paczkowska M, Barenboim J, Sintupisut N, et al. Integrative pathway enrichment analysis of multivariate omics data[J]. Nat Commun, 2020, 11(1): 735. Langerud J, Eilertsen I A, Moosavi S H, et al. Multiregional transcriptomics identifies congruent consensus subtypes with prognostic value beyond tumor heterogeneity of colorectal cancer[J]. Nat Commun, 2024, 15(1): 4342. Xu X, Gong C, Wang Y, et al. Multi-omics analysis to identify driving factors in colorectal cancer[J]. Epigenomics, 2020, 12(18): 1633-1650. Ma Y, Li J, Zhao X, et al. Multi-omics cluster defines the subtypes of CRC with distinct prognosis and tumor microenvironment[J]. Eur J Med Res, 2024, 29(1): 207. Wang F, Li Z, Xu T, et al. A comprehensive multi-omics analysis identifies a robust scoring system for cancer-associated fibroblasts and intervention targets in colorectal cancer[J]. J Cancer Res Clin Oncol, 2024, 150(3): 124. Stahler A, Hoppe B, Na I K, et al. Consensus Molecular Subtypes as Biomarkers of Fluorouracil and Folinic Acid Maintenance Therapy With or Without Panitumumab in RAS Wild-Type Metastatic Colorectal Cancer (PanaMa, AIO KRK 0212)[J]. J Clin Oncol, 2023, 41(16): 2975-2987. Dedecker L, Coppedge B, Avelar-Barragan J, et al. Microbiome distinctions between the CRC carcinogenic pathways[J]. Gut Microbes, 2021, 13(1): 1854641. Kong C, Liang L, Liu G, et al. Integrated metagenomic and metabolomic analysis reveals distinct gut-microbiome-derived phenotypes in early-onset colorectal cancer[J]. Gut, 2023, 72(6): 1129-1142. Liu Z, Zhang X, Zhang H, et al. Multi-Omics Analysis Reveals Intratumor Microbes as Immunomodulators in Colorectal Cancer[J]. Microbiol Spectr, 2023, 11(2): e0503822. Li Y, Wang B, Ma F, et al. Proteomic characterization of the colorectal cancer response to chemoradiation and targeted therapies reveals potential therapeutic strategies[J]. Cell Rep Med, 2023, 4(12): 101311. Deng F, Huang J, Yuan X, et al. Performance and efficiency of machine learning algorithms for analyzing rectangular biomedical data[J]. Lab Invest, 2021, 101(4): 430-441. Chu G, Ji X, Wang Y, Niu H. Integrated multiomics analysis and machine learning refine molecular subtypes and prognosis for muscle-invasive urothelial cancer[J]. Mol Ther Nucleic Acids, 2023, 33: 110-126. Liu Z, Liu L, Weng S, et al. Machine learning-based integration develops an immune-derived lncRNA signature for improving outcomes in colorectal cancer[J]. Nat Commun, 2022, 13(1): 816. Wang R, Dai W, Gong J, et al. Development of a novel combined nomogram model integrating deep learning-pathomics, radiomics and immunoscore to predict postoperative outcome of colorectal cancer lung metastasis patients[J]. J Hematol Oncol, 2022, 15(1): 11. Reel P S, Reel S, Van Kralingen J C, et al. Machine learning for classification of hypertension subtypes using multi-omics: A multi-centre, retrospective, data-driven study[J]. EBioMedicine, 2022, 84: 104276. Preto A J, Chanana S, Ence D, et al. Multi-omics data integration identifies novel biomarkers and patient subgroups in inflammatory bowel disease[J]. J Crohns Colitis, 2025, 19(1). Foersch S, Glasner C, Woerl A C, et al. Multistain deep learning for prediction of prognosis and therapy response in colorectal cancer[J]. Nat Med, 2023, 29(2): 430-439. Wang X, Liu J, Wang D, et al. Epigenetically regulated gene expression profiles reveal four molecular subtypes with prognostic and therapeutic implications in colorectal cancer[J]. Brief Bioinform, 2021, 22(4). Wei W, Li Y, Huang T. Using Machine Learning Methods to Study Colorectal Cancer Tumor Micro-Environment and Its Biomarkers[J]. Int J Mol Sci, 2023, 24(13). Liu Z, Guo C, Dang Q, et al. Integrative analysis from multi-center studies identities a consensus machine learning-derived lncRNA signature for stage II/III colorectal cancer[J]. EBioMedicine, 2022, 75: 103750. Chen G, Ren Q, Zhong Z, et al. Exploring the gut microbiome's role in colorectal cancer: diagnostic and prognostic implications[J]. Front Immunol, 2024, 15: 1431747. Du W, Frankel T L, Green M, Zou W. IFNgamma signaling integrity in colorectal cancer immunity and immunotherapy[J]. Cell Mol Immunol, 2022, 19(1): 23-32. Zhang S, Deshpande A, Verma B K, et al. Integration of Clinical Trial Spatial Multiomics Analysis and Virtual Clinical Trials Enables Immunotherapy Response Prediction and Biomarker Discovery[J]. Cancer Res, 2024, 84(16): 2734-2748. Carson T L, Byrd D A, Smith K S, et al. A case-control study of the association between the gut microbiota and colorectal cancer: exploring the roles of diet, stress, and race[J]. Gut Pathog, 2024, 16(1): 13. Sheng D, Jin C, Yue K, et al. Pan-cancer atlas of tumor-resident microbiome, immunity and prognosis[J]. Cancer Lett, 2024, 598: 217077. Lu X, Meng J, Zhou Y, et al. MOVICS: an R package for multi-omics integration and visualization in cancer subtyping[J]. Bioinformatics, 2021, 36(22-23): 5539-5541. Hanzelmann S, Castelo R, Guinney J. GSVA: gene set variation analysis for microarray and RNA-seq data[J]. BMC Bioinformatics, 2013, 14: 7. Newman A M, Liu C L, Green M R, et al. Robust enumeration of cell subsets from tissue expression profiles[J]. Nat Methods, 2015, 12(5): 453-7. Smith J J, Deane N G, Wu F, et al. Experimentally derived metastasis gene expression profile predicts recurrence and death in patients with colon cancer[J]. Gastroenterology, 2010, 138(3): 958-68. Freeman T J, Smith J J, Chen X, et al. Smad4-mediated signaling inhibits intestinal neoplasia by inhibiting expression of beta-catenin[J]. Gastroenterology, 2012, 142(3): 562-571 e2. Del Rio M, Mollevi C, Bibeau F, et al. Molecular subtypes of metastatic colorectal cancer are associated with patient response to irinotecan-based therapies[J]. Eur J Cancer, 2017, 76: 68-75. Szeglin B C, Wu C, Marco M R, et al. A SMAD4-modulated gene profile predicts disease-free survival in stage II and III colorectal cancer[J]. Cancer Rep (Hoboken), 2022, 5(1): e1423. Sturm G, Finotello F, Petitprez F, et al. Comprehensive evaluation of transcriptome-based cell-type quantification methods for immuno-oncology[J]. Bioinformatics, 2019, 35(14): i436-i445. Subramanian A, Tamayo P, Mootha V K, et al. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles[J]. Proc Natl Acad Sci U S A, 2005, 102(43): 15545-50. Kanehisa M, Furumichi M, Sato Y, et al. KEGG: biological systems database as a model of the real world[J]. Nucleic Acids Res, 2025, 53(D1): D672-D677. Kanehisa M. Toward understanding the origin and evolution of cellular organisms[J]. Protein Sci, 2019, 28(11): 1947-1951. Kanehisa M, Goto S. KEGG: kyoto encyclopedia of genes and genomes[J]. Nucleic Acids Res, 2000, 28(1): 27-30. Riaz N, Havel J J, Makarov V, et al. Tumor and Microenvironment Evolution during Immunotherapy with Nivolumab[J]. Cell, 2017, 171(4): 934-949.e16. Lauss M, Donia M, Harbst K, et al. Mutational and putative neoantigen load predict clinical benefit of adoptive T cell therapy in melanoma[J]. Nat Commun, 2017, 8(1): 1738. Van Allen E M, Miao D, Schilling B, et al. Genomic correlates of response to CTLA-4 blockade in metastatic melanoma[J]. Science, 2015, 350(6257): 207-211. Gide T N, Quek C, Menzies A M, et al. Distinct Immune Cell Populations Define Response to Anti-PD-1 Monotherapy and Anti-PD-1/Anti-CTLA-4 Combined Therapy[J]. Cancer Cell, 2019, 35(2): 238-255.e6. Rose T L, Weir W H, Mayhew G M, et al. Fibroblast growth factor receptor 3 alterations and response to immune checkpoint inhibition in metastatic urothelial cancer: a real world experience[J]. Br J Cancer, 2021, 125(9): 1251-1260. Mariathasan S, Turley S J, Nickles D, et al. TGFbeta attenuates tumour response to PD-L1 blockade by contributing to exclusion of T cells[J]. Nature, 2018, 554(7693): 544-548. Song M, Chan A T, Sun J. Influence of the Gut Microbiome, Diet, and Environment on Risk of Colorectal Cancer[J]. Gastroenterology, 2020, 158(2): 322-340. Brennan C A, Garrett W S. Fusobacterium nucleatum - symbiont, opportunist and oncobacterium[J]. Nat Rev Microbiol, 2019, 17(3): 156-166. Huang Y, Cao J, Zhu M, et al. Nontoxigenic Bacteroides fragilis: A double-edged sword[J]. Microbiol Res, 2024, 286: 127796. Lopez-Siles M, Duncan S H, Garcia-Gil L J, Martinez-Medina M. Faecalibacterium prausnitzii: from microbiology to diagnostics and prognostics[J]. ISME J, 2017, 11(4): 841-852. Sanders M E, Merenstein D J, Reid G, et al. Probiotics and prebiotics in intestinal health and disease: from biology to the clinic[J]. Nat Rev Gastroenterol Hepatol, 2019, 16(10): 605-616. Wong S H, Yu J. Gut microbiota in colorectal cancer: mechanisms of action and clinical applications[J]. Nat Rev Gastroenterol Hepatol, 2019, 16(11): 690-704. Nicolini A, Ferrari P. Involvement of tumor immune microenvironment metabolic reprogramming in colorectal cancer progression, immune escape, and response to immunotherapy[J]. Front Immunol, 2024, 15: 1353787. Additional Declarations No competing interests reported. Supplementary Files SupplementaryMaterial.docx SupplementaryTable1.xlsx SupplementaryTable2.xlsx SupplementaryTable3.xlsx SupplementaryTable4.xlsx SupplementaryTable5.xlsx SupplementaryTable6.xlsx SupplementaryTable7.xlsx SupplementaryTable8.xlsx Cite Share Download PDF Status: Published Journal Publication published 12 Jul, 2025 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Accepted 24 Jun, 2025 Reviewers invited by journal 21 Apr, 2025 Submission checks completed at journal 10 Apr, 2025 First submitted to journal 29 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5669193","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":445938413,"identity":"afe0c4b9-143c-4321-b1ec-8a4f35ec0283","order_by":0,"name":"Jun Wang","email":"","orcid":"","institution":"Affiliated Hospital of North Sichuan Medical College","correspondingAuthor":false,"prefix":"","firstName":"Jun","middleName":"","lastName":"Wang","suffix":""},{"id":445938415,"identity":"26baaf5e-12ca-4c95-b049-cf26a7fcf4c8","order_by":1,"name":"Cong Yuan","email":"","orcid":"","institution":"Affiliated Hospital of North Sichuan Medical College","correspondingAuthor":false,"prefix":"","firstName":"Cong","middleName":"","lastName":"Yuan","suffix":""},{"id":445938417,"identity":"ec1a5fe7-a7b4-40bd-aca6-4d74d559661a","order_by":2,"name":"Bo Tang","email":"","orcid":"","institution":"University of Electronic Science and Technology of China","correspondingAuthor":false,"prefix":"","firstName":"Bo","middleName":"","lastName":"Tang","suffix":""},{"id":445938419,"identity":"de71d234-43aa-4fc5-8266-8e8a22e20133","order_by":3,"name":"Juan Liu","email":"","orcid":"","institution":"Affiliated Hospital of North Sichuan Medical College","correspondingAuthor":false,"prefix":"","firstName":"Juan","middleName":"","lastName":"Liu","suffix":""},{"id":445938421,"identity":"c6c18ebc-4727-4242-a28f-df1da808cd6c","order_by":4,"name":"Ke Pu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAw0lEQVRIiWNgGAWjYFACNsYHCRVQNg+RWpgNPpwhUQub5Mw2UrSYS6QlSPPOu2Ov236A8cHbNgZ5c0JaLGekHTDm3fYscduZBGbDuW0MhjsbCGgxuJHekMy77XCC2Q0GNmneNoYEgwNEaDnMO+ewPVAL+28itaQdbJzZcJhxG9AWZqK0WPY8S2b4cOww0C+JzZJzzkkYbiCkxZw9zfxHQg3QYccPH/zwpsxGnrDDEEzGBiAhQUA9qpZRMApGwSgYBTgAALXRQfu1/J7XAAAAAElFTkSuQmCC","orcid":"","institution":"Affiliated Hospital of North Sichuan Medical College","correspondingAuthor":true,"prefix":"","firstName":"Ke","middleName":"","lastName":"Pu","suffix":""}],"badges":[],"createdAt":"2024-12-18 11:53:29","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5669193/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5669193/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-025-08915-1","type":"published","date":"2025-07-12T15:57:57+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":81512977,"identity":"3931d5cf-9dbe-4a75-973a-a6de724b12e6","added_by":"auto","created_at":"2025-04-28 06:34:41","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":25485750,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparison of microbiome composition and diversity between normal intestinal tissues and colorectal cancer samples. (A\u003c/strong\u003e) α diversity at the bacterial phylum level. (\u003cstrong\u003eB\u003c/strong\u003e) Average relative abundance bar plot at the bacterial phylum level, comparing normal intestinal tissues and colorectal cancer samples. (\u003cstrong\u003eC\u003c/strong\u003e) Relative abundance bar plot of all samples at the bacterial phylum level, comparing normal intestinal tissues and colorectal cancer samples. (\u003cstrong\u003eD\u003c/strong\u003e) α diversity at the bacterial class level. (\u003cstrong\u003eE\u003c/strong\u003e) Average relative abundance bar plot at the bacterial class level, comparing normal intestinal tissues and colorectal cancer samples. (\u003cstrong\u003eF\u003c/strong\u003e) Relative abundance bar plot of all samples at the bacterial class level, comparing normal intestinal tissues and colorectal cancer samples. (\u003cstrong\u003eG\u003c/strong\u003e) α diversity at the bacterial order level. (\u003cstrong\u003eH\u003c/strong\u003e) Average relative abundance bar plot at the bacterial order level, comparing normal intestinal tissues and colorectal cancer samples. (\u003cstrong\u003eI\u003c/strong\u003e) Relative abundance bar plot of all samples at the bacterial order level, comparing normal intestinal tissues and colorectal cancer samples. (\u003cstrong\u003eJ\u003c/strong\u003e) α diversity at the bacterial family level. (\u003cstrong\u003eK\u003c/strong\u003e) Average relative abundance bar plot at the bacterial family level, comparing normal intestinal tissues and colorectal cancer samples. (\u003cstrong\u003eL\u003c/strong\u003e) Relative abundance bar plot of all samples at the bacterial family level, comparing normal intestinal tissues and colorectal cancer samples. (\u003cstrong\u003eM\u003c/strong\u003e) α diversity at the bacterial genus level. (\u003cstrong\u003eN\u003c/strong\u003e) Average relative abundance bar plot at the bacterial genus level, comparing normal intestinal tissues and colorectal cancer samples. (\u003cstrong\u003eO\u003c/strong\u003e) Relative abundance bar plot of all samples at the bacterial genus level, comparing normal intestinal tissues and colorectal cancer samples. (\u003cstrong\u003eP\u003c/strong\u003e) α diversity at the bacterial species level. (\u003cstrong\u003eQ\u003c/strong\u003e) Average relative abundance bar plot at the bacterial species level, comparing normal intestinal tissues and colorectal cancer samples. (\u003cstrong\u003eR\u003c/strong\u003e) Relative abundance bar plot of all samples at the bacterial species level, comparing normal intestinal tissues and colorectal cancer samples.\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/bee8beb38e55320296d1d82b.png"},{"id":81512353,"identity":"791e1cd2-5c3b-47ac-896b-39bf1cf00941","added_by":"auto","created_at":"2025-04-28 06:26:39","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":26314316,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eMulti-omics subtyping of colorectal cancer patients based on TCGA database. \u003c/strong\u003e(\u003cstrong\u003eA\u003c/strong\u003e) Cluster prediction index and gap statistic analysis plot for multi-omics clustering. (\u003cstrong\u003eB\u003c/strong\u003e) Clustering of colorectal cancer patients using 10 multi-omics clustering methods. (\u003cstrong\u003eC\u003c/strong\u003e) Evaluation of sample similarity within each group by calculating Silhouette scores. (\u003cstrong\u003eD\u003c/strong\u003e) Consensus set subtype comprehensive heatmap, including subtype heatmaps of mRNA, lncRNA, miRNA, DNA CpG methylation sites, mutated genes, and bacterial abundance. The heatmap displays 10 feature molecules of mRNA, lncRNA, miRNA, DNA CpG methylation sites, mutated genes, and bacterial abundance subtypes. (\u003cstrong\u003eE\u003c/strong\u003e) Survival curves of the two subtypes.\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/4d94b8fc5bdc9865b6def4d7.png"},{"id":81512355,"identity":"32a5f4f8-adea-4e39-b56d-1c0d2383dba2","added_by":"auto","created_at":"2025-04-28 06:26:40","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":86508316,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eMulti-omics subtyping atlas of colorectal cancer and validation. \u003c/strong\u003e(\u003cstrong\u003eA\u003c/strong\u003e) Enrichment of cancer-related gene sets in different subtypes. (\u003cstrong\u003eB\u003c/strong\u003e) Immune landscape in the TCGA-COAD cohort. The top annotation of the heatmap shows the immune enrichment score, stromal enrichment score, and DNA methylation of tumor-infiltrating lymphocytes. The upper panel shows the expression of typical immune checkpoint genes, and the lower panel shows the enrichment levels of 24 TME-related immune cells. (\u003cstrong\u003eC\u003c/strong\u003e) Regulatory activity profiles of 23 transcriptional regulatory genes (upper panel) and potential regulatory factors associated with chromatin remodeling (lower panel) in the two subtypes. (\u003cstrong\u003eD\u003c/strong\u003e) Validation of TCGA-COAD cohort subtyping in the nearest template of the META-COAD cohort. (\u003cstrong\u003eE\u003c/strong\u003e) Survival analysis of subtyping in the META-COAD cohort.\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/42e2cbe637cae3804682dcac.png"},{"id":81512390,"identity":"6eb46dfe-3c6c-48e2-9600-49b8ae4a3faa","added_by":"auto","created_at":"2025-04-28 06:26:42","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":24031722,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGeneration and prognostic value of MCMLS (Multi-Omics Integrative Clustering and Machine Learning Score). \u003c/strong\u003e(\u003cstrong\u003eA\u003c/strong\u003e) A combination of 103 machine learning algorithms was generated through an integrated computational framework. The C-index of each model was calculated through the TCGA-COAD and META-COAD cohorts and ranked according to the average C-index of the validation set. (\u003cstrong\u003eB\u003c/strong\u003e) The central genes and regression coefficients selected by the optimal algorithm. (\u003cstrong\u003eC-D\u003c/strong\u003e) Survival analysis of patients with high MCMLS and low MCMLS in the TCGA-COAD and META-COAD cohorts. (\u003cstrong\u003eE\u003c/strong\u003e) Univariate Cox regression analysis results of 18 hub genes in the TCGA-COAD cohort. (\u003cstrong\u003eF\u003c/strong\u003e) Univariate Cox regression analysis results of 18 hub genes in the META-COAD cohort.\u003c/p\u003e","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/3c26a775d5a2081cf02bb9c1.png"},{"id":81512367,"identity":"bd9b8e86-dec6-41f0-9c07-c8f7df261f9f","added_by":"auto","created_at":"2025-04-28 06:26:41","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":6975708,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eClinical practice value of MCMLS. \u003c/strong\u003e(\u003cstrong\u003eA-F\u003c/strong\u003e) Distribution of gender, age, T stage, M stage, N stage, and Stage stage between different MCMLS groups. (\u003cstrong\u003eG\u003c/strong\u003e)-(\u003cstrong\u003eH\u003c/strong\u003e) Comparison of the MCMLS model with 50 other published colorectal cancer models in the TCGA-COAD and META-COAD cohorts. (\u003cstrong\u003eI\u003c/strong\u003e) Comparison of the prognostic prediction ability of MCMLS with other clinical information in the TCGA-COAD cohort.\u003c/p\u003e","description":"","filename":"Figure5.png","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/4c175c1d3f3449fc463677d5.png"},{"id":81512370,"identity":"7120d3d7-898f-4f39-92c0-346053dcd989","added_by":"auto","created_at":"2025-04-28 06:26:41","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":13728231,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eImmune escape phenotype of different MCMLS subgroups. \u003c/strong\u003e(\u003cstrong\u003eA\u003c/strong\u003e) Heatmap of immune cell infiltration between different groups based on 7 immune infiltration algorithms. (\u003cstrong\u003eB\u003c/strong\u003e) From left to right: mRNA expression (median normalized expression level); expression versus methylation (correlation between gene expression and DNA methylation beta values); amplification frequency (difference between the proportion of samples with IM amplification in a specific subtype and the proportion of amplification in all samples); and deletion frequency of 74 IM genes divided by immune subtype (as amplification).\u003c/p\u003e","description":"","filename":"Figure6.png","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/9c9d4710de477a7cf9cdfb3d.png"},{"id":81512408,"identity":"2c77a735-d166-4d33-bda3-6ac3751f6b35","added_by":"auto","created_at":"2025-04-28 06:26:43","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":22423672,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePotential pathways of different MCMLS subgroups. \u003c/strong\u003e(\u003cstrong\u003eA\u003c/strong\u003e) GSEA enrichment analysis of KEGG gene sets in different subgroups. (\u003cstrong\u003eB\u003c/strong\u003e) GSEA enrichment analysis of cancer gene sets in different subgroups. (\u003cstrong\u003eC\u003c/strong\u003e) GSVA enrichment analysis of metabolic gene sets in different subgroups.\u003c/p\u003e","description":"","filename":"Figure7.png","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/53e5face8bd0661e6dee3dde.png"},{"id":81512990,"identity":"19202a8c-e88d-48ae-9dfd-7021125de99c","added_by":"auto","created_at":"2025-04-28 06:34:43","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":9585437,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGene mutation analysis between different MCMLS subgroups. \u003c/strong\u003e(\u003cstrong\u003eA\u003c/strong\u003e) Box plot of mutation load differences between different MCMLS groups. (\u003cstrong\u003eB\u003c/strong\u003e) Waterfall plot of the top 20 most mutated genes between different MCMLS groups. (\u003cstrong\u003eC\u003c/strong\u003e) Univariate Cox analysis forest plot of the top 20 most mutated genes between different MCMLS groups to clarify which genes belong to which MCMLS groups. (\u003cstrong\u003eD\u003c/strong\u003e) Co-expression correlation plot of the top 20 mutated genes in the MCMLS-Low group. (\u003cstrong\u003eE\u003c/strong\u003e) Co-expression correlation plot of the top 20 mutated genes in the MCMLS-High group.\u003cstrong\u003e \u003c/strong\u003e(\u003cstrong\u003eF\u003c/strong\u003e) Schematic diagram of TP53 gene mutation structure between different MCMLS groups.\u003c/p\u003e","description":"","filename":"Figure8.png","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/782398c2ac0879244a1555c9.png"},{"id":81512400,"identity":"63a8784b-f121-4576-a821-5abcf3cffe30","added_by":"auto","created_at":"2025-04-28 06:26:43","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":28033868,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eBacterial difference analysis between different MCMLS subgroups. \u003c/strong\u003e(\u003cstrong\u003eA\u003c/strong\u003e) Bar plot of bacterial differences between different MCMLS subgroups. (\u003cstrong\u003eB\u003c/strong\u003e) Univariate Cox forest plot of bacteria between different MCMLS subgroups. (\u003cstrong\u003eC\u003c/strong\u003e) Bubble plot of the correlation between prognostic differential bacteria and immune cells. (\u003cstrong\u003eD\u003c/strong\u003e) Bubble plot of the correlation between prognostic differential bacteria and metabolic pathways.\u003c/p\u003e","description":"","filename":"Figure9.png","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/8ff9cdeb65b18aa09e51fc4c.png"},{"id":81512362,"identity":"21b46da4-ea06-481c-b034-5e6f57e47425","added_by":"auto","created_at":"2025-04-28 06:26:41","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":14274340,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eMCMLS model can accurately predict immunotherapy response. \u003c/strong\u003e(\u003cstrong\u003eA-D\u003c/strong\u003e) Relationship between MCMLS score and immunotherapy in the IMvigor210 dataset (survival curve, immunotherapy response bar plot, score and immune response difference box plot, MCMLS grouping and immunotherapy response distribution percentage bar plot). (\u003cstrong\u003eE-H\u003c/strong\u003e) Relationship between MCMLS score and immunotherapy in the GSE176307 dataset (survival curve, immunotherapy response bar plot, score and immune response difference box plot, MCMLS grouping and immunotherapy response distribution percentage bar plot). (\u003cstrong\u003eI-L\u003c/strong\u003e) Relationship between MCMLS score and immunotherapy in the PRJEB23709 dataset (survival curve, immunotherapy response bar plot, score and immune response difference box plot, MCMLS grouping and immunotherapy response distribution percentage bar plot). (\u003cstrong\u003eM-P\u003c/strong\u003e) Relationship between MCMLS score and immunotherapy in the PHS00452 dataset (survival curve, immunotherapy response bar plot, score and immune response difference box plot, MCMLS grouping and immunotherapy response distribution percentage bar plot). (\u003cstrong\u003eQ-T\u003c/strong\u003e) Relationship between MCMLS score and immunotherapy in the GSE91061 dataset (survival curve, immunotherapy response bar plot, score and immune response difference box plot, MCMLS grouping and immunotherapy response distribution percentage bar plot). (\u003cstrong\u003eU-X\u003c/strong\u003e) Relationship between MCMLS score and immunotherapy in the GSE100797 dataset (survival curve, immunotherapy response bar plot, score and immune response difference box plot, MCMLS grouping and immunotherapy response distribution percentage bar plot).\u003c/p\u003e","description":"","filename":"Figure10.png","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/f1848ae20cf488de0c6373a7.png"},{"id":81512972,"identity":"c8b3edb4-3006-4200-9e17-7c408951cf2a","added_by":"auto","created_at":"2025-04-28 06:34:41","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":1554628,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/3630f58533a116c38542eb0c.docx"},{"id":81512354,"identity":"7fb41a70-4570-4208-b273-6c4af7ec1805","added_by":"auto","created_at":"2025-04-28 06:26:40","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":19976,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/7b8db40bfc55e278efc06320.xlsx"},{"id":81512357,"identity":"b59edf3f-32a1-45c2-93bd-aa029f32d25e","added_by":"auto","created_at":"2025-04-28 06:26:40","extension":"xlsx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":14098,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/8d046d118345bc437d1b44ce.xlsx"},{"id":81512985,"identity":"efafec11-0489-4892-a328-430e23fc651e","added_by":"auto","created_at":"2025-04-28 06:34:42","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":31486,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable3.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/ea0db92c6904671a59efc274.xlsx"},{"id":81512375,"identity":"6f20154d-42aa-43ce-b463-5ef25cd29953","added_by":"auto","created_at":"2025-04-28 06:26:41","extension":"xlsx","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":19202,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable4.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/53a13561b2171b67fbd051ea.xlsx"},{"id":81513792,"identity":"7341e8d5-fbd6-46d2-8645-d0d7326b4bb0","added_by":"auto","created_at":"2025-04-28 06:42:41","extension":"xlsx","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":16108,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable5.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/09dfc103b3f3934c012ff79b.xlsx"},{"id":81512358,"identity":"82e3a014-fff0-454e-9ad2-ba14133a2b56","added_by":"auto","created_at":"2025-04-28 06:26:40","extension":"xlsx","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":19970,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable6.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/b990a4afbdae7831666c3191.xlsx"},{"id":81512983,"identity":"afa1f8ae-aae3-435f-843b-bd08cdc59298","added_by":"auto","created_at":"2025-04-28 06:34:42","extension":"xlsx","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":16079,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable7.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/7f06a9d6d149514109f131fb.xlsx"},{"id":81512393,"identity":"372c5513-575b-4fe7-9aae-3db8f7e0d43e","added_by":"auto","created_at":"2025-04-28 06:26:42","extension":"xlsx","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":12608,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTable8.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5669193/v1/3cce0204a08f22ee58eefbd6.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Integrative Analysis of Multi-Omics Data and Gut Microbiota Composition Reveals Prognostic Subtypes and Predicts Immunotherapy Response in Colorectal Cancer Using Machine Learning","fulltext":[{"header":"Introduction","content":"\u003cp\u003eColorectal cancer (CRC) ranks as the third most commonly diagnosed cancer and the second leading cause of cancer-related mortality worldwide, with approximately 1.9\u0026nbsp;million new cases and 935,000 deaths reported in 2023\u003csup\u003e[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]\u003c/sup\u003e. Despite advances in screening strategies and treatment modalities, patient outcomes remain suboptimal, with 5-year survival rates of approximately 65%\u003csup\u003e[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]\u003c/sup\u003e. The development of effective personalized treatment approaches is hindered by the heterogeneous nature of CRC, which is characterized by diverse molecular subtypes with varying clinical outcomes\u003csup\u003e[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]\u003c/sup\u003e. Recent research has established the tumor microenvironment (TME) as a critical determinant in CRC progression, therapeutic response, and patient prognosis\u003csup\u003e[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]\u003c/sup\u003e. The TME represents a complex ecosystem comprising various cellular components, including immune cells, fibroblasts, and endothelial cells, as well as non-cellular elements such as the extracellular matrix \u003csup\u003e[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]\u003c/sup\u003e. Additionally, the gut microbiome has emerged as a key contributor to CRC pathogenesis and treatment response. Multiple studies have demonstrated that specific microbial signatures are associated with CRC development, progression, and clinical outcomes. For example, Mouradov et al. identified three distinct oncomicrobial community subtypes (OCSs) with unique clinico-molecular features and prognostic implications, where OCS1 (characterized by Fusobacterium and oral pathogens) was associated with right-sided, high-grade, MSI-high tumors, while OCS2 (dominated by Firmicutes/Bacteroidetes) conferred better survival outcomes in microsatellite-stable tumors\u003csup\u003e[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]\u003c/sup\u003e. Similarly, Ibeanu et al. found that several bacterial species, including Fusobacterium nucleatum, promote tumorigenesis through inflammation induction and immune system modulation, making them promising candidates for targeted interventions\u003csup\u003e[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]\u003c/sup\u003e. These findings highlight the importance of integrating microbiome analysis into CRC research to enhance our understanding of disease mechanisms and improve patient stratification.\u003c/p\u003e \u003cp\u003eThe advent of high-throughput sequencing technologies has revolutionized our understanding of CRC's molecular landscape, enabling comprehensive multi-omics profiling\u003csup\u003e[\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]\u003c/sup\u003e. Integrative analysis of these multi-dimensional datasets has revealed distinct molecular subtypes with important clinical implications. The consensus molecular subtypes (CMS) classification, which categorizes CRC into four subtypes based on gene expression patterns, has emerged as a robust framework for understanding tumor biology and guiding treatment decisions\u003csup\u003e[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]\u003c/sup\u003e. However, recent studies have shown that intra-tumor heterogeneity compromises the clinical utility of transcriptomic classifications, necessitating more sophisticated approaches for patient stratification\u003csup\u003e[\u003cspan additionalcitationids=\"CR17 CR18\" citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]\u003c/sup\u003e. However, the incorporation of gut microbiome data into these integrative analyses has been limited, despite the growing evidence supporting the pivotal role of the gut microbiome in CRC \u003csup\u003e[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]\u003c/sup\u003e. For instance, Li et al. demonstrated that proteomic characterization of CRC response to chemoradiation and targeted therapies could identify four distinct subtypes with unique biological and therapeutic characteristics, providing valuable insights for personalized treatment\u003csup\u003e[\u003cspan additionalcitationids=\"CR22\" citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eMachine learning (ML) approaches have emerged as powerful tools for analyzing complex biological data and developing predictive models in cancer research\u003csup\u003e[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]\u003c/sup\u003e. ML algorithms can effectively handle high-dimensional datasets, capture non-linear relationships, and identify subtle patterns that may be overlooked by traditional statistical methods\u003csup\u003e[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]\u003c/sup\u003e. In the context of CRC, ML has been applied to various tasks, such as predicting patient survival, classifying molecular subtypes, and identifying biomarkers for treatment response\u003csup\u003e[\u003cspan additionalcitationids=\"CR28\" citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]\u003c/sup\u003e. The integration of multi-omics data using ML approaches has shown promise in improving the accuracy and robustness of prognostic and predictive models. For example, Wang et al. developed a model based on epigenetically regulated gene expression profiles that identified four molecular subtypes with distinct clinical and molecular features, providing a better understanding of CRC heterogeneity\u003csup\u003e[\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]\u003c/sup\u003e. Despite these advancements, the incorporation of gut microbiome data into multi-omics integrative analyses has been limited, despite the growing evidence supporting the pivotal role of the gut microbiome in CRC\u003csup\u003e[\u003cspan additionalcitationids=\"CR33\" citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]\u003c/sup\u003e. Recent studies have demonstrated that specific gut microbial signatures are associated with CRC prognosis and response to immunotherapy\u003csup\u003e[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]\u003c/sup\u003e. For instance, Zhang et al. developed a computational model that integrated spatial multi-omics data to predict immunotherapy response in cancer, highlighting the potential of multi-omics integration for biomarker discovery\u003csup\u003e[\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]\u003c/sup\u003e. Furthermore, Carson et al. found that microbiome-CRC associations may differ by racial group, emphasizing the need for large, diverse population-based studies to determine the generalizability of these associations\u003csup\u003e[\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eIn this study, we aimed to conduct a comprehensive integrative analysis of multi-omics data and gut microbiota composition in CRC to identify novel prognostic subtypes and develop a robust ML-based model for predicting patient survival and immunotherapy response. We hypothesized that the integration of multi-omics data, including transcriptomics, epigenomics, and genomics, along with gut microbiome profiles, would provide a more holistic understanding of CRC biology and enable the development of accurate prognostic and predictive models. To achieve this, we collected multi-omics data and gut microbiome profiles from a large cohort of 274 CRC patients and performed an integrative analysis using state-of-the-art bioinformatics and ML approaches. We employed multiple unsupervised clustering algorithms to identify distinct molecular subtypes based on the integrated multi-omics data and evaluated their associations with clinical outcomes and immune features. Furthermore, we developed a novel ML-based prognostic model, the Multi-Omics Integrative Clustering and Machine Learning Score (MCMLS), which incorporates multi-omics data and gut microbiome profiles to predict patient survival and immunotherapy response. Our findings provide valuable insights into the complex interplay between the TME, gut microbiome, and molecular features of CRC, and highlight the potential of integrating multi-omics data and ML approaches for improving patient stratification and guiding personalized treatment strategies.\u003c/p\u003e"},{"header":"Materials and methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eData Acquisition and Processing\u003c/h2\u003e \u003cp\u003eThe tumor microbiome data for TCGA-COAD was obtained from the MicrobiomeX database\u003csup\u003e[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]\u003c/sup\u003e, which included 497 samples (456 tumor samples and 41 adjacent normal samples). The mRNA, lncRNA, miRNA, DNA CpG methylation sites, mutated genes, and clinical data were downloaded from The Cancer Genome Atlas (TCGA) database. A total of 274 patients had complete data across all platforms with available survival information. Clinical characteristics are provided in Supplementary Table\u0026nbsp;1.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eMulti-omics Integrative Clustering Analysis\u003c/h3\u003e\n\u003cp\u003eMulti-omics integrative clustering was performed using the MOVICS package in R (version 4.3.0)\u003csup\u003e[\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]\u003c/sup\u003e. For feature selection, we applied criteria established by Chu et al\u003csup\u003e[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]\u003c/sup\u003e: for mRNA, lncRNA, and miRNA expression data, the top 3000, 3000, and 5000 features were selected based on median absolute deviation (MAD) and filtered using Cox regression with p-value cutoffs of 0.01, 0.001, and 0.001, respectively. For DNA methylation data, the top 3000 features were selected based on MAD and filtered with a p-value cutoff of 0.05. Mutation features were selected using a frequency cutoff of 0.15, and bacterial expression features were selected based on the top 15 highest standard deviations. The optimal number of clusters (2\u0026ndash;8) was determined using cluster prediction analysis, and ten different clustering algorithms were applied. Silhouette analysis was performed to evaluate clustering quality. DNA methylation beta values were converted to M values for improved signal detection. Survival differences between identified subtypes were assessed using Kaplan-Meier analysis with log-rank tests, with p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 considered statistically significant.\u003c/p\u003e\n\u003ch3\u003eImmune Landscape Analysis and Subtype Validation\u003c/h3\u003e\n\u003cp\u003eThe immune landscape of identified subtypes was characterized through gene set variation analysis (GSVA)\u003csup\u003e[\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]\u003c/sup\u003e using immune-related pathways. Transcriptional regulatory networks were reconstructed using the RTN package, focusing on mucin regulation and chromatin remodeling. The methylation-based tumor-infiltrating lymphocyte (MeTIL) score was calculated, and stromal and immune scores were estimated using the ESTIMATE algorithm. We assessed immune checkpoint gene expression across the subtypes and estimated immune cell composition using the CIBERSORT deconvolution algorithm\u003csup\u003e[\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eFor subtype validation, four independent colorectal cancer datasets (GSE17536\u003csup\u003e[\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]\u003c/sup\u003e, GSE17537\u003csup\u003e[\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]\u003c/sup\u003e, GSE72970\u003csup\u003e[\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]\u003c/sup\u003e, and GSE161158\u003csup\u003e[\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]\u003c/sup\u003e) were downloaded from the Gene Expression Omnibus (GEO) database. Batch effects were removed using the sva package, creating a meta-validation dataset. Nearest Template Prediction (NTP) and Prediction Analysis for Microarrays (PAM) methods were used to assign subtype labels. Agreement between subtyping methods was assessed using Cohen's kappa coefficient, with values\u0026thinsp;\u0026gt;\u0026thinsp;0.6 considered substantial agreement.\u003c/p\u003e\n\u003ch3\u003eMulti-Omics Integrative Clustering and Machine Learning Score (MCMLS) Model Development\u003c/h3\u003e\n\u003cp\u003eFor prognostic model construction, we employed a machine learning approach using the TCGA dataset as the training cohort and the meta-validation dataset as the validation cohort. Data preprocessing included extracting common genes between datasets and performing standardization. Following the methodology of Chu et al\u003csup\u003e[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]\u003c/sup\u003e, 101 different machine learning models were trained using various feature selection methods, implemented in the R environment with the caret package. Model performance was evaluated using the concordance index (C-index), with higher values indicating better discrimination between high-risk and low-risk patients. The optimal plsRcox algorithm was identified based on the highest average C-index across validation cohorts. For survival analysis, patients were divided into high-risk and low-risk groups based on the median MCMLS score, with p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 considered statistically significant.\u003c/p\u003e\n\u003ch3\u003eComparison of Prognostic Signatures\u003c/h3\u003e\n\u003cp\u003eTo compare the performance of our developed prognostic signature with existing signatures, we conducted a comprehensive analysis using the TCGA dataset as the training cohort and a meta-validation dataset as the validation cohort. The gene expression data and clinical survival information were obtained for both cohorts. We collected publicly available prognostic signatures from the supplementary materials of a previous study and manually curated additional signatures from the literature\u003csup\u003e[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]\u003c/sup\u003e. For each signature, we trained a Cox proportional hazards model and calculated C-indices for both training and validation cohorts. Time-dependent area under the receiver operating characteristic curve (AUC) analysis was conducted to compare the MCMLS model with clinical risk factors (tumor stage, T stage, N stage, M stage, and gender). AUC values\u0026thinsp;\u0026gt;\u0026thinsp;0.7 were considered to indicate good discriminatory ability.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eImmune Infiltration Analysis and Molecular Characterization of MCMLS Subtypes\u003c/h2\u003e \u003cp\u003eThe immune infiltration landscape was analyzed using seven deconvolution algorithms: CIBERSORT, EPIC, MCPcounter, xCell, ESTIMATE, TIMER, and quanTIseq\u003csup\u003e[\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e]\u003c/sup\u003e. Wilcoxon rank-sum tests were used to identify significant differences in immune cell infiltration between MCMLS subtypes, with p-values adjusted for multiple testing. DNA methylation and copy number variation data were integrated with gene expression data to characterize molecular features of the subtypes..\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eFunctional Enrichment Analysis and Metabolic Pathway Activity Assessment\u003c/h3\u003e\n\u003cp\u003eGene Set Enrichment Analysis (GSEA) was performed using the \"h.all.v2023.2.Hs.symbols.gmt\" and \"KEGG.database.symbols.gmt\" gene sets\u003csup\u003e[\u003cspan additionalcitationids=\"CR48 CR49\" citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e]\u003c/sup\u003e. Differential expression analysis between MCMLS subtypes was conducted using the limma package. GSVA was performed using the \"REACTOME_metabolism.gmt\" gene set to analyze metabolic pathway activities\u003csup\u003e[\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]\u003c/sup\u003e. Significantly enriched pathways were identified using a threshold of p\u0026thinsp;\u0026lt;\u0026thinsp;0.05.\u003c/p\u003e\n\u003ch3\u003eMutational Landscape Analysis of MCMLS Subtypes\u003c/h3\u003e\n\u003cp\u003eSomatic mutation data from the TCGA cohort was analyzed using the maftools package. Mutational profiles were compared between MCMLS subtypes using mafCompare analysis. Differentially mutated genes were identified using a significance threshold of p\u0026thinsp;\u0026lt;\u0026thinsp;0.05. Tumor mutational burden (TMB) was calculated for each sample and compared between subtypes using the Wilcoxon rank-sum test. Somatic interaction analysis was performed to investigate co-occurrence and mutual exclusivity of gene mutations.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eBacterial Differential Analysis and Correlations\u003c/h2\u003e \u003cp\u003eBacteria with significant differences between MCMLS subtypes were identified using Linear discriminant analysis Effect Size (LEfSe) with thresholds of p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 and LDA\u0026thinsp;\u0026gt;\u0026thinsp;2. Survival analysis for differentially abundant bacteria was conducted using Cox proportional hazards models. Samples were grouped into \"Low\" and \"High\" abundance groups using the optimal cutoff determined by the surv_cutpoint function or the median value. Correlations between selected bacteria and immune cell infiltration or metabolic pathway activity were assessed using Pearson correlation analysis.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003ePredicting Immunotherapy Response Using the MCMLS Model\u003c/h2\u003e \u003cp\u003eTo assess the predictive power of the MCMLS model for immunotherapy response, we applied the model to multiple independent immunotherapy datasets, including Melanoma-GSE91061\u003csup\u003e[\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e]\u003c/sup\u003e, Melanoma-GSE100797\u003csup\u003e[\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e]\u003c/sup\u003e, Melanoma-phs000452\u003csup\u003e[\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e]\u003c/sup\u003e, Melanoma-PRJEB23709\u003csup\u003e[\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e]\u003c/sup\u003e, Urothelial Cancer-GSE176307\u003csup\u003e[\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e]\u003c/sup\u003e, and Urothelial Cancer-IMvigor210\u003csup\u003e[\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e]\u003c/sup\u003e. For each dataset, we performed a comprehensive analysis pipeline. First, we preprocessed the data by reading the gene expression data and clinical information from the corresponding CSV files and standardizing the gene expression data to ensure comparability across samples. Next, we calculated the risk score for each sample using the calRS function based on the MCMLS model genes. This function fitted a Cox proportional hazards model using the selected genes and calculated the risk score as the linear combination of the gene expression values weighted by their respective coefficients. The samples were then divided into high-risk and low-risk groups based on the optimal cutoff point determined by the surv_cutpoint function. We then conducted survival analysis by generating Kaplan-Meier survival curves to compare the overall survival between the high-risk and low-risk groups, using the log-rank test to assess the statistical significance of the survival differences. To investigate the association between the MCMLS risk groups and immunotherapy response, we harmonized the response categories into multiple (CR, PR, SD, PD) categories, depending on the available information in each dataset. We created violin plots with embedded box plots to visualize the distribution of risk scores across the response categories and used the Kruskal-Wallis test to compare the risk scores between the response groups. Additionally, we generated stacked bar plots to show the percentage of patients in each response category within the high-risk and low-risk groups, using the Fisher's exact test to assess the association between the risk groups and response categorie. Finally, we created a heatmap-like representation to visualize the risk scores and corresponding response categories for each sample, ordering the samples by decreasing risk score and color-coding the response categories. We used the Kruskal-Wallis test (for multiple response categories) to compare the risk scores across the response categories.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eMicrobiome Composition and Diversity in Normal Intestinal Tissue versus Colorectal Cancer\u003c/h2\u003e \u003cp\u003eComparative analysis of microbiome composition between normal intestinal tissue and colorectal cancer (CRC) samples revealed significant differences across taxonomic levels. The α-diversity, measured by Shannon index, was consistently lower in tumor samples than normal samples at all taxonomic levels (P\u0026thinsp;\u0026lt;\u0026thinsp;0.001; Figs.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA,D,G,J,M,P). At the phylum level, Proteobacteria predominated in tumor samples while Firmicutes was more abundant in normal tissues (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB-C). At the class level, Gammaproteobacteria showed higher representation in tumor samples compared to Betaproteobacteria in normal samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eE-F). The order-level analysis demonstrated elevated Enterobacterales in tumor samples versus Caudovirales in normal tissues (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eH-I). Family-level examination revealed Enterobacteriaceae enrichment in tumors while Myoviridae dominated in normal samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eK-L). At the genus level, Escherichia was more prevalent in tumor samples compared to Klebsiella in normal tissues (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eN-O). Species-level analysis identified Escherichia coli enrichment in tumors whereas Klebsiella pneumoniae was more abundant in normal samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eQ-R). No significant differences in species-level Shannon index were observed when stratifying CRC samples by gender, age, or cancer stage (Figures \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eA-I).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eMulti-omics Subtyping of Colorectal Cancer Patients Based on TCGA Database\u003c/h2\u003e \u003cp\u003eMulti-omics clustering analysis of CRC patients using TCGA data identified two distinct molecular subtypes. Cluster prediction and gap statistics analysis (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA) supported the application of ten different clustering methods (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB), consistently revealing two major subtypes designated as CS1 (n\u0026thinsp;=\u0026thinsp;127) and CS2 (n\u0026thinsp;=\u0026thinsp;147). Silhouette scores validated the internal similarity within each subtype (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC). A comprehensive subtype heatmap was generated integrating mRNA, lncRNA, miRNA, DNA methylation, mutated genes, and bacterial abundance data, highlighting ten representative features for each data type that differentiated the two subtypes (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD). Survival analysis demonstrated significantly longer overall survival for CS1 patients compared to CS2 patients (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eE), confirming the clinical relevance of this multi-omics classification approach.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eMulti-omics Subtyping Landscape and Validation in Colorectal Cancer\u003c/h2\u003e \u003cp\u003eFurther characterization of the molecular differences between CS1 and CS2 subtypes was performed using gene set enrichment analysis (GSEA). The P53 signaling pathway showed higher activity in CS1, while PI3K-AKT-MTOR signaling was more active in CS2 (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA). Immune landscape analysis revealed higher expression of immune checkpoint genes in CS1 compared to CS2 (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB). Assessment of transcriptional regulatory genes and chromatin remodeling factors demonstrated differential expression patterns, with PARB, PARA, and EGFR elevated in CS2, while HIF1A, ESR2, and FOXM1 were higher in CS1 (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC). The subtyping approach was validated in the META-COAD cohort using the nearest template prediction algorithm, which successfully classified patients into CS1 and CS2 subtypes (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD). Consistent with TCGA findings, META-COAD patients in CS1 demonstrated significantly longer overall survival than those in CS2 (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eE). Principal component analysis confirmed successful integration of four independent CRC datasets (GSE17356, GSE17537, GSE72970, and GSE161158) after batch effect removal (Figures S2A-B).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eDevelopment and Prognostic Value of the Multi-Omics Integrative Clustering and Machine Learning Score(MCMLS)\u003c/h2\u003e \u003cp\u003eWe developed a prognostic model for CRC by evaluating 103 machine learning algorithm combinations (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA). The plsRcox algorithm achieved the highest C-index (0.661), demonstrating superior performance in predicting patient survival. This algorithm identified 18 hub genes with corresponding regression coefficients (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB) that formed the basis of our Multi-Omics integrative Clustering and Machine Learning Score (MCMLS). Patients stratified by median MCMLS score showed significant survival differences in both TCGA-COAD (P\u0026thinsp;\u0026lt;\u0026thinsp;0.0001; Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC) and META-COAD cohorts (P\u0026thinsp;\u0026lt;\u0026thinsp;0.0001; Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eD). Univariate Cox regression analysis confirmed the significant association of the 18 hub genes with patient survival in both cohorts (Figs.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eE-F), validating their inclusion in the MCMLS model.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eClinical Utility of MCMLS in Colorectal Cancer\u003c/h2\u003e \u003cp\u003eThe clinical relevance of MCMLS was assessed by examining its relationship with established clinicopathological factors. MCMLS was significantly associated with age, T stage, N stage, and overall stage, but not with gender or M stage (Figs.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA-F). Comparative analysis with 50 previously published CRC prognostic models demonstrated that MCMLS outperformed all other models in both TCGA-COAD and META-COAD cohorts (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eG-H). MCMLS exhibited higher predictive accuracy than traditional clinical factors including stage, T stage, N stage, M stage, and gender in the TCGA-COAD cohort (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eI), highlighting its potential as a complementary prognostic tool for clinical decision-making.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eImmune Features and Metabolic Pathways Associated with MCMLS Subgroups\u003c/h2\u003e \u003cp\u003eImmune cell infiltration analysis using seven algorithms revealed that the MCMLS low-score group had higher infiltration of CD8\u0026thinsp;+\u0026thinsp;T cells (TIMER algorithm) and NK cells (xCell algorithm), while the MCMLS high-score group exhibited increased fibroblast infiltration (EPIC, MCPcounter, and xCell algorithms) (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eA). Expression analysis of 74 immune-related genes showed that the MCMLS high-score group had elevated expression of immunotherapy-inhibiting molecules (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eB).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eGSEA using KEGG gene sets demonstrated that the MCMLS low-score group was enriched in oxidative phosphorylation pathways, while the high-score group showed enrichment in cell adhesion and migration pathways (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eA). Analysis with cancer gene sets confirmed these findings and additionally showed fatty acid metabolism enrichment in the low-score group and epithelial-mesenchymal transition enrichment in the high-score group (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eB). Gene set variation analysis of metabolic pathways revealed generally higher metabolic activity in the low-score group, particularly in lipid metabolism, fatty acid metabolism, and the TCA cycle, while the high-score group showed elevated activity in PI metabolism (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eC).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eMutation Profiles of MCMLS Subgroups\u003c/h2\u003e \u003cp\u003eMutation analysis demonstrated that the MCMLS high-score group had significantly higher mutation burden compared to the low-score group (P\u0026thinsp;=\u0026thinsp;0.0041; Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eA). The most frequently mutated genes across all samples were APC, TP53, and KRAS (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eB). Univariate Cox analysis identified TP53, TTN, and MUC16 as significantly more frequently mutated in the high-score group (P\u0026thinsp;\u0026lt;\u0026thinsp;0.05; Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eC). Co-expression analysis revealed that in the low-score group, TP53 mutations were mutually exclusive with mutations in RYR1, MUC5B, CSMD3, OBSCN, and TTN (P\u0026thinsp;\u0026lt;\u0026thinsp;0.01; Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eD), while in the high-score group, APC mutations were mutually exclusive with OBSCN mutations (P\u0026thinsp;\u0026lt;\u0026thinsp;0.01; Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eE). The mutation frequency of TP53 was significantly higher in the high-score group (70.07%) compared to the low-score group (51.82%) (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eF).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eBacterial Composition and Prognostic Significance in MCMLS Subgroups\u003c/h2\u003e \u003cp\u003eLEfSe analysis identified 53 differentially abundant bacteria between MCMLS subgroups (P\u0026thinsp;\u0026lt;\u0026thinsp;0.05, LDA\u0026thinsp;\u0026gt;\u0026thinsp;2; Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003eA). Univariate Cox analysis revealed that two bacteria, s_Streptomyces_lividans and s_Pseudomonas_sp._CIP.10, were associated with worse prognosis when their abundance was higher (Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003eB). These bacteria were positively correlated with fibroblasts and negatively correlated with NK cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003eC) and most metabolic pathways (Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003eD). Their higher abundance in the MCMLS high-score group aligned with the observed lower metabolic activity in this group.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003eMCMLS Predicts Immunotherapy Response Across Multiple Datasets\u003c/h2\u003e \u003cp\u003eTo validate the predictive capacity of MCMLS for immunotherapy response, we analyzed six independent immunotherapy-treated patient cohorts across different cancer types and immunotherapy regimens. The IMvigor210 cohort comprised 119 patients with locally advanced or metastatic urothelial carcinoma treated with atezolizumab (anti-PD-L1, 1200 mg intravenously every 21 days) as first-line therapy (PMID: 27939400). The GSE176307 dataset included 103 patients with metastatic urothelial cancer treated with various immune checkpoint inhibitors (ICB) at a single academic center. The PRJEB23709 cohort consisted of metastatic melanoma patients treated with either anti-PD-1 monotherapy or combined anti-PD-1 and anti-CTLA-4 therapy. The PHS00452 dataset included melanoma patients treated with immunotherapy as part of the Melanoma Genome Sequencing Project. The GSE91061 cohort comprised 65 melanoma patients (109 samples) treated with nivolumab (anti-PD-1) therapy (PMID: 29033130), while GSE100797 included 25 melanoma patients undergoing adoptive T cell therapy after previous immunotherapy treatment (PMID: 29170503). Across all datasets, the MCMLS high-score group consistently showed significantly worse prognosis compared to the low-score group (P\u0026thinsp;\u0026lt;\u0026thinsp;0.001; Figs.\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003eA,E,I,M,Q,U). MCMLS scores were significantly associated with immunotherapy response in all datasets (P\u0026thinsp;\u0026lt;\u0026thinsp;0.05; Figs.\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003eB,F,J,N,R,V), with a gradual increase in scores observed from complete response to progressive disease (P\u0026thinsp;\u0026lt;\u0026thinsp;0.05; Figs.\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003eC,G,K,O,S,W). The proportion of responders was significantly higher in the MCMLS low-score group across all datasets (Figs.\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003eD,H,L,P,T,X), demonstrating the robust predictive power of MCMLS for immunotherapy response across different cancer types and therapeutic regimens..\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe present study employed a comprehensive multi-omics approach to investigate the molecular heterogeneity, prognostic factors, and immunotherapy response in colorectal cancer (CRC). By integrating data from the gut microbiome, transcriptomics, epigenomics, and genomics, we aimed to gain a deeper understanding of the complex biological processes underlying CRC development and progression. Our findings highlight the importance of considering multiple omics data types in the characterization of CRC subtypes and the development of personalized treatment strategies.\u003c/p\u003e \u003cp\u003eWe began by comparing the microbiome composition and diversity between normal intestinal tissue and CRC samples at different taxonomic levels. The results revealed significant differences in bacterial abundance and diversity, with tumor samples exhibiting lower α-diversity and distinct dominant bacterial taxa compared to normal samples. These findings align with recent studies identifying distinct oncomicrobial community subtypes in CRC with unique clinico-molecular features\u003csup\u003e[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e]\u003c/sup\u003e. The enrichment of certain bacterial species in tumor samples, such as Escherichia coli and members of Proteobacteria, and the depletion of beneficial bacteria like Klebsiella pneumoniae, reflect the complex interactions between the gut microbiome and the host in CRC pathogenesis. These distinct bacterial profiles may contribute to the creation of a pro-inflammatory microenvironment that promotes tumor growth and progression, as demonstrated by recent microbiome research in CRC\u003csup\u003e[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e, \u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e]\u003c/sup\u003e. Our multi-omics subtyping analysis identified two major CRC subtypes (CS1 and CS2) with distinct molecular characteristics and clinical outcomes. Notably, patients in the CS1 subtype exhibited significantly longer overall survival compared to those in the CS2 subtype. These findings underscore the power of integrating multiple omics data types in capturing the complex molecular landscape of CRC and identifying clinically relevant subtypes. The molecular differences between the CS1 and CS2 subtypes provide valuable insights into the underlying biological mechanisms driving the observed differences in prognosis\u003csup\u003e[\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e, \u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e]\u003c/sup\u003e. For example, the CS1 subtype was characterized by higher activity of the P53 signaling pathway and increased expression of immune checkpoint genes, while the CS2 subtype showed greater activity in the PI3K-AKT-MTOR pathway. These distinct molecular features align with recent research demonstrating that CRC can be classified into different subtypes with unique biological characteristics and therapeutic vulnerabilities\u003csup\u003e[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eA novel contribution of our study is the development of the Multi-Omics Integrative Clustering and Machine Learning Score (MCMLS), which represents a significant advancement over existing prognostic models. The MCMLS model demonstrated robust prognostic performance in both the TCGA-COAD and META-COAD cohorts, outperforming 50 published CRC prognostic models and traditional clinical factors. The superior performance of MCMLS can be attributed to its comprehensive integration of multi-omics data, capturing complex interactions that may not be fully represented by individual omics platforms or clinical variables. This approach mirrors recent successful strategies in other cancer types, but our application to CRC with the inclusion of microbiome data represents a novel advancement in the field\u003csup\u003e[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e]\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eAnother key innovation of our study is the identification of distinct immune escape phenotypes associated with different MCMLS subgroups. The MCMLS low-score group exhibited higher infiltration of CD8\u0026thinsp;+\u0026thinsp;T cells (TIMER algorithm) and NK cells (xCell algorithm), indicating a more favorable immune microenvironment. In contrast, the MCMLS high-score group showed increased fibroblast infiltration and expression of immunotherapy-inhibiting molecules. These findings extend current understanding of immune escape mechanisms in CRC by revealing specific cellular patterns associated with different molecular subtypes\u003csup\u003e[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]\u003c/sup\u003e. This characterization of immune escape phenotypes provides potential targets for therapeutic intervention and may guide immunotherapy strategies for different patient subgroups.\u003c/p\u003e \u003cp\u003eOur metabolic pathway analysis revealed that the MCMLS low-score group was enriched in oxidative phosphorylation and fatty acid metabolism pathways, while the high-score group showed enrichment in cell adhesion, migration, and epithelial-mesenchymal transition pathways. These distinct metabolic profiles suggest different energy utilization strategies between the subtypes, which may contribute to their different clinical behaviors. Recent research has highlighted the importance of metabolic reprogramming in tumor immune evasion\u003csup\u003e[\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e]\u003c/sup\u003e, and our findings contribute to this emerging field by linking specific metabolic signatures to immune escape phenotypes and patient outcomes in CRC.\u003c/p\u003e \u003cp\u003eThe investigation of gene mutation differences between MCMLS subgroups revealed that the MCMLS high-score group had a significantly higher mutation burden compared to the low-score group, with specific genes such as TP53, TTN, and MUC16 exhibiting higher mutation frequencies. These findings align with recent comprehensive genomic studies of CRC\u003csup\u003e[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]\u003c/sup\u003ebut uniquely link these mutational profiles to immune features and prognosis. The co-expression analysis further revealed novel mutually exclusive mutation patterns, with TP53 mutations being mutually exclusive with mutations in RYR1, MUC5B, CSMD3, OBSCN, and TTN in the MCMLS-Low group, while APC mutations were mutually exclusive with OBSCN mutations in the MCMLS-High group. These patterns may reflect different evolutionary trajectories of CRC and offer new insights into disease mechanisms.\u003c/p\u003e \u003cp\u003ePerhaps the most innovative aspect of our study is the demonstrated ability of MCMLS to predict immunotherapy response across six independent datasets. Patients in the MCMLS low-score group consistently showed better responses to immunotherapy compared to those in the high-score group, with significant associations between MCMLS scores and clinical responses (complete response, partial response, stable disease, or progressive disease). This finding is particularly important given the current limitations of immunotherapy in CRC, where benefits are largely restricted to microsatellite instability-high tumors\u003csup\u003e[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]\u003c/sup\u003e. Our model's ability to identify potential responders beyond this subset could significantly expand the application of immunotherapy in CRC, addressing a critical unmet need.\u003c/p\u003e \u003cp\u003eFinally, our analysis of the bacterial differences between MCMLS subgroups identified 53 differentially abundant bacteria, with two species (s_Streptomyces_lividans and s_Pseudomonas_sp._CIP.10) significantly associated with worse prognosis. Their correlation with fibroblast infiltration and negative association with NK cells and metabolic pathways suggests a potential role in shaping the tumor microenvironment. This represents one of the first demonstrations of specific bacterial signatures associated with both molecular subtypes and clinical outcomes in CRC, extending recent work on the diagnostic and prognostic implications of the gut microbiome\u003csup\u003e[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eIn conclusion, our study presents several novel findings that advance the field of CRC research. First, we developed a robust multi-omics-based prognostic model (MCMLS) that outperforms existing methods and captures complex biological interactions. Second, we identified distinct immune escape phenotypes associated with different molecular subtypes, providing new insights into resistance mechanisms. Third, we demonstrated that MCMLS can accurately predict immunotherapy response across multiple independent datasets, potentially expanding the application of immunotherapy in CRC. Fourth, we uncovered specific bacterial signatures associated with prognosis and immune features, highlighting the importance of the microbiome in CRC biology. These innovations collectively contribute to a more comprehensive understanding of CRC heterogeneity and lay the foundation for more effective personalized treatment strategies.\u003c/p\u003e \u003cp\u003eDespite these advancements, our study has several limitations that should be addressed in future research. Prospective clinical validation is necessary to confirm the clinical utility of the MCMLS model. The functional roles of the identified molecular features require further experimental validation. Integration of additional omics data types, such as proteomics and metabolomics, may provide an even more comprehensive understanding of CRC biology. Nevertheless, our study represents a significant step toward precision medicine in CRC, demonstrating the power of multi-omics data integration and machine learning approaches in advancing clinical management.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn summary, our comprehensive multi-omics analysis of CRC provides novel insights into the molecular heterogeneity and biological mechanisms underlying this complex disease. The developed MCMLS model represents a promising tool for precision medicine in CRC, enabling the stratification of patients based on their molecular profiles and the prediction of prognosis and immunotherapy response. The identified molecular features and pathways offer potential targets for future therapeutic interventions and mechanistic studies. Our findings underscore the importance of integrating multi-omics data and machine learning approaches in cancer research and highlight the potential of precision medicine in improving the clinical management and outcomes of CRC patients.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eCRC, Colorectal cancer\u003c/p\u003e\n\u003cp\u003eTME, Tumor Microenvironment\u003c/p\u003e\n\u003cp\u003eICIs, Immune Checkpoint Inhibitors\u003c/p\u003e\n\u003cp\u003eML, Machine learning\u003c/p\u003e\n\u003cp\u003eMCMLS, Multi-Omics Integrative Clustering and Machine Learning Score\u003c/p\u003e\n\u003cp\u003eMAD, Median Absolute Deviation\u003c/p\u003e\n\u003cp\u003eTRNs, Transcriptional Regulatory Networks\u003c/p\u003e\n\u003cp\u003eTFs, transcription factors\u003c/p\u003e\n\u003cp\u003eMeTIL, Methylation-based Tumor-Infiltrating Lymphocyte\u003c/p\u003e\n\u003cp\u003eNTP, Nearest Template Prediction\u003c/p\u003e\n\u003cp\u003ePAM, Prediction Analysis for Microarrays\u003c/p\u003e\n\u003cp\u003eC-index, Concordance Index\u003c/p\u003e\n\u003cp\u003eAUROC, Area Under the Receiver Operating Characteristic curve\u003c/p\u003e\n\u003cp\u003eCNV, Copy Number Variation\u003c/p\u003e\n\u003cp\u003eGSEA, Gene Set Enrichment Analysis\u003c/p\u003e\n\u003cp\u003eGSVA, Gene Set Variation Analysis\u003c/p\u003e\n\u003cp\u003eMAF, Mutation Annotation Format\u003c/p\u003e\n\u003cp\u003eTILs, Tumor-Infiltrating Lymphocytes\u003c/p\u003e\n\u003cp\u003ePCA, Principal component analysis\u003c/p\u003e\n\u003cp\u003eCR, Complete Response\u003c/p\u003e\n\u003cp\u003ePR, Partial Response\u003c/p\u003e\n\u003cp\u003eSD, Stable Disease\u003c/p\u003e\n\u003cp\u003ePD, Progressive Disease\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe multi-omics data used in this study are publicly available from The Cancer Genome Atlas (TCGA-COAD) database. The microbiome data were obtained from the MicrobiomeX database. The validation datasets (GSE17536, GSE17537, GSE72970, and GSE161158) are available from the Gene Expression Omnibus (GEO) database. The immunotherapy datasets used for validation of the MCMLS model (IMvigor210, GSE176307, PRJEB23709, PHS00452, GSE91061, and GSE100797) are also publicly available from their respective repositories. The processed data, analysis code, and the MCMLS model implementation are available from the corresponding authors upon reasonable request. Restrictions apply to the availability of some patient-specific data, which were used under license for the current study, and so are not publicly available. Any additional data that support the findings of this study are included in the supplementary information files or are available from the corresponding authors upon reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors have no relevant financial or non-financial interests to disclose.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u003c/strong\u003e This work was supported by the Intramural Fund of North Sichuan Medical College (CBY21-QD31), Nanchong City Talent Development Fund (CBY23-NCR06), and Fund of Bureau of Science\u0026amp;Technology Nanchong City (23JCYJPT0056)\u003c/p\u003e\n\u003cp\u003eCRediT authorship contribution statement\u003c/p\u003e\n\u003cp\u003eKP and JL contributed to the design and conceptualization of the study. JW, CY and BT wrote the manuscript. KP, JW, CY and BT were involved in the data processing, generation, analysis, interpretation, and figure and table production for this project. JL and KP guided the writing of the manuscript, revised the manuscript and provided financial support. All authors have read and agreed to the published version of the manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eMorgan E, Arnold M, Gini A, et al. Global burden of colorectal cancer in 2020 and 2040: incidence and mortality estimates from GLOBOCAN[J]. Gut, 2023, 72(2): 338-344.\u003c/li\u003e\n\u003cli\u003eChung D C, Gray D M, 2nd, Singh H, et al. A Cell-free DNA Blood-Based Test for Colorectal Cancer Screening[J]. N Engl J Med, 2024, 390(11): 973-983.\u003c/li\u003e\n\u003cli\u003eXie Y H, Chen Y X, Fang J Y. Comprehensive review of targeted therapy for colorectal cancer[J]. Signal Transduct Target Ther, 2020, 5(1): 22.\u003c/li\u003e\n\u003cli\u003eLi L, Jiang D, Liu H, et al. Comprehensive Proteogenomic Profiling Reveals the Molecular Characteristics of Colorectal Cancer at Distinct Stages of Progression[J]. Cancer Res, 2024, 84(17): 2888-2910.\u003c/li\u003e\n\u003cli\u003eShin A E, Giancotti F G, Rustgi A K. Metastatic colorectal cancer: mechanisms and emerging therapeutics[J]. Trends Pharmacol Sci, 2023, 44(4): 222-236.\u003c/li\u003e\n\u003cli\u003eNunes L, Li F, Wu M, et al. Prognostic genome and transcriptome signatures in colorectal cancers[J]. Nature, 2024, 633(8028): 137-146.\u003c/li\u003e\n\u003cli\u003eBejarano L, Jordao M J C, Joyce J A. Therapeutic Targeting of the Tumor Microenvironment[J]. Cancer Discov, 2021, 11(4): 933-959.\u003c/li\u003e\n\u003cli\u003eArashi K, Nishiyama T, Hosoya M, et al. Bilateral Deafness Due to Relapsing Polychondritis with Semicircular Canal Calcification Treated With Cochlear Implantation: A Case Report[J]. Ear Nose Throat J, 2023: 1455613231215173.\u003c/li\u003e\n\u003cli\u003eLi Y, Wu X, Fang D, Luo Y. Informing immunotherapy with multi-omics driven machine learning[J]. NPJ Digit Med, 2024, 7(1): 67.\u003c/li\u003e\n\u003cli\u003eMouradov D, Greenfield P, Li S, et al. Oncomicrobial Community Profiling Identifies Clinicomolecular and Prognostic Subtypes of Colorectal Cancer[J]. Gastroenterology, 2023, 165(1): 104-120.\u003c/li\u003e\n\u003cli\u003eIbeanu G C, Rowaiye A B, Okoli J C, Eze D U. Microbiome Differences in Colorectal Cancer Patients and Healthy Individuals: Implications for Vaccine Antigen Discovery[J]. Immunotargets Ther, 2024, 13: 749-774.\u003c/li\u003e\n\u003cli\u003eEspinosa E, Bautista R, Larrosa R, Plata O. Advancements in long-read genome sequencing technologies and algorithms[J]. Genomics, 2024, 116(3): 110842.\u003c/li\u003e\n\u003cli\u003eHu M, Zhu J, Peng G, et al. IMOVNN: incomplete multi-omics data integration variational neural networks for gut microbiome disease prediction and biomarker identification[J]. Brief Bioinform, 2023, 24(6).\u003c/li\u003e\n\u003cli\u003ePaczkowska M, Barenboim J, Sintupisut N, et al. Integrative pathway enrichment analysis of multivariate omics data[J]. Nat Commun, 2020, 11(1): 735.\u003c/li\u003e\n\u003cli\u003eLangerud J, Eilertsen I A, Moosavi S H, et al. Multiregional transcriptomics identifies congruent consensus subtypes with prognostic value beyond tumor heterogeneity of colorectal cancer[J]. Nat Commun, 2024, 15(1): 4342.\u003c/li\u003e\n\u003cli\u003eXu X, Gong C, Wang Y, et al. Multi-omics analysis to identify driving factors in colorectal cancer[J]. Epigenomics, 2020, 12(18): 1633-1650.\u003c/li\u003e\n\u003cli\u003eMa Y, Li J, Zhao X, et al. Multi-omics cluster defines the subtypes of CRC with distinct prognosis and tumor microenvironment[J]. Eur J Med Res, 2024, 29(1): 207.\u003c/li\u003e\n\u003cli\u003eWang F, Li Z, Xu T, et al. A comprehensive multi-omics analysis identifies a robust scoring system for cancer-associated fibroblasts and intervention targets in colorectal cancer[J]. J Cancer Res Clin Oncol, 2024, 150(3): 124.\u003c/li\u003e\n\u003cli\u003eStahler A, Hoppe B, Na I K, et al. Consensus Molecular Subtypes as Biomarkers of Fluorouracil and Folinic Acid Maintenance Therapy With or Without Panitumumab in RAS Wild-Type Metastatic Colorectal Cancer (PanaMa, AIO KRK 0212)[J]. J Clin Oncol, 2023, 41(16): 2975-2987.\u003c/li\u003e\n\u003cli\u003eDedecker L, Coppedge B, Avelar-Barragan J, et al. Microbiome distinctions between the CRC carcinogenic pathways[J]. Gut Microbes, 2021, 13(1): 1854641.\u003c/li\u003e\n\u003cli\u003eKong C, Liang L, Liu G, et al. Integrated metagenomic and metabolomic analysis reveals distinct gut-microbiome-derived phenotypes in early-onset colorectal cancer[J]. Gut, 2023, 72(6): 1129-1142.\u003c/li\u003e\n\u003cli\u003eLiu Z, Zhang X, Zhang H, et al. Multi-Omics Analysis Reveals Intratumor Microbes as Immunomodulators in Colorectal Cancer[J]. Microbiol Spectr, 2023, 11(2): e0503822.\u003c/li\u003e\n\u003cli\u003eLi Y, Wang B, Ma F, et al. Proteomic characterization of the colorectal cancer response to chemoradiation and targeted therapies reveals potential therapeutic strategies[J]. Cell Rep Med, 2023, 4(12): 101311.\u003c/li\u003e\n\u003cli\u003eDeng F, Huang J, Yuan X, et al. Performance and efficiency of machine learning algorithms for analyzing rectangular biomedical data[J]. Lab Invest, 2021, 101(4): 430-441.\u003c/li\u003e\n\u003cli\u003eChu G, Ji X, Wang Y, Niu H. Integrated multiomics analysis and machine learning refine molecular subtypes and prognosis for muscle-invasive urothelial cancer[J]. Mol Ther Nucleic Acids, 2023, 33: 110-126.\u003c/li\u003e\n\u003cli\u003eLiu Z, Liu L, Weng S, et al. Machine learning-based integration develops an immune-derived lncRNA signature for improving outcomes in colorectal cancer[J]. Nat Commun, 2022, 13(1): 816.\u003c/li\u003e\n\u003cli\u003eWang R, Dai W, Gong J, et al. Development of a novel combined nomogram model integrating deep learning-pathomics, radiomics and immunoscore to predict postoperative outcome of colorectal cancer lung metastasis patients[J]. J Hematol Oncol, 2022, 15(1): 11.\u003c/li\u003e\n\u003cli\u003eReel P S, Reel S, Van Kralingen J C, et al. Machine learning for classification of hypertension subtypes using multi-omics: A multi-centre, retrospective, data-driven study[J]. EBioMedicine, 2022, 84: 104276.\u003c/li\u003e\n\u003cli\u003ePreto A J, Chanana S, Ence D, et al. Multi-omics data integration identifies novel biomarkers and patient subgroups in inflammatory bowel disease[J]. J Crohns Colitis, 2025, 19(1).\u003c/li\u003e\n\u003cli\u003eFoersch S, Glasner C, Woerl A C, et al. Multistain deep learning for prediction of prognosis and therapy response in colorectal cancer[J]. Nat Med, 2023, 29(2): 430-439.\u003c/li\u003e\n\u003cli\u003eWang X, Liu J, Wang D, et al. Epigenetically regulated gene expression profiles reveal four molecular subtypes with prognostic and therapeutic implications in colorectal cancer[J]. Brief Bioinform, 2021, 22(4).\u003c/li\u003e\n\u003cli\u003eWei W, Li Y, Huang T. Using Machine Learning Methods to Study Colorectal Cancer Tumor Micro-Environment and Its Biomarkers[J]. Int J Mol Sci, 2023, 24(13).\u003c/li\u003e\n\u003cli\u003eLiu Z, Guo C, Dang Q, et al. Integrative analysis from multi-center studies identities a consensus machine learning-derived lncRNA signature for stage II/III colorectal cancer[J]. EBioMedicine, 2022, 75: 103750.\u003c/li\u003e\n\u003cli\u003eChen G, Ren Q, Zhong Z, et al. Exploring the gut microbiome\u0026apos;s role in colorectal cancer: diagnostic and prognostic implications[J]. Front Immunol, 2024, 15: 1431747.\u003c/li\u003e\n\u003cli\u003eDu W, Frankel T L, Green M, Zou W. IFNgamma signaling integrity in colorectal cancer immunity and immunotherapy[J]. Cell Mol Immunol, 2022, 19(1): 23-32.\u003c/li\u003e\n\u003cli\u003eZhang S, Deshpande A, Verma B K, et al. Integration of Clinical Trial Spatial Multiomics Analysis and Virtual Clinical Trials Enables Immunotherapy Response Prediction and Biomarker Discovery[J]. Cancer Res, 2024, 84(16): 2734-2748.\u003c/li\u003e\n\u003cli\u003eCarson T L, Byrd D A, Smith K S, et al. A case-control study of the association between the gut microbiota and colorectal cancer: exploring the roles of diet, stress, and race[J]. Gut Pathog, 2024, 16(1): 13.\u003c/li\u003e\n\u003cli\u003eSheng D, Jin C, Yue K, et al. Pan-cancer atlas of tumor-resident microbiome, immunity and prognosis[J]. Cancer Lett, 2024, 598: 217077.\u003c/li\u003e\n\u003cli\u003eLu X, Meng J, Zhou Y, et al. MOVICS: an R package for multi-omics integration and visualization in cancer subtyping[J]. Bioinformatics, 2021, 36(22-23): 5539-5541.\u003c/li\u003e\n\u003cli\u003eHanzelmann S, Castelo R, Guinney J. GSVA: gene set variation analysis for microarray and RNA-seq data[J]. BMC Bioinformatics, 2013, 14: 7.\u003c/li\u003e\n\u003cli\u003eNewman A M, Liu C L, Green M R, et al. Robust enumeration of cell subsets from tissue expression profiles[J]. Nat Methods, 2015, 12(5): 453-7.\u003c/li\u003e\n\u003cli\u003eSmith J J, Deane N G, Wu F, et al. Experimentally derived metastasis gene expression profile predicts recurrence and death in patients with colon cancer[J]. Gastroenterology, 2010, 138(3): 958-68.\u003c/li\u003e\n\u003cli\u003eFreeman T J, Smith J J, Chen X, et al. Smad4-mediated signaling inhibits intestinal neoplasia by inhibiting expression of beta-catenin[J]. Gastroenterology, 2012, 142(3): 562-571 e2.\u003c/li\u003e\n\u003cli\u003eDel Rio M, Mollevi C, Bibeau F, et al. Molecular subtypes of metastatic colorectal cancer are associated with patient response to irinotecan-based therapies[J]. Eur J Cancer, 2017, 76: 68-75.\u003c/li\u003e\n\u003cli\u003eSzeglin B C, Wu C, Marco M R, et al. A SMAD4-modulated gene profile predicts disease-free survival in stage II and III colorectal cancer[J]. Cancer Rep (Hoboken), 2022, 5(1): e1423.\u003c/li\u003e\n\u003cli\u003eSturm G, Finotello F, Petitprez F, et al. Comprehensive evaluation of transcriptome-based cell-type quantification methods for immuno-oncology[J]. Bioinformatics, 2019, 35(14): i436-i445.\u003c/li\u003e\n\u003cli\u003eSubramanian A, Tamayo P, Mootha V K, et al. Gene set enrichment analysis: a knowledge-based approach for interpreting genome-wide expression profiles[J]. Proc Natl Acad Sci U S A, 2005, 102(43): 15545-50.\u003c/li\u003e\n\u003cli\u003eKanehisa M, Furumichi M, Sato Y, et al. KEGG: biological systems database as a model of the real world[J]. Nucleic Acids Res, 2025, 53(D1): D672-D677.\u003c/li\u003e\n\u003cli\u003eKanehisa M. Toward understanding the origin and evolution of cellular organisms[J]. Protein Sci, 2019, 28(11): 1947-1951.\u003c/li\u003e\n\u003cli\u003eKanehisa M, Goto S. KEGG: kyoto encyclopedia of genes and genomes[J]. Nucleic Acids Res, 2000, 28(1): 27-30.\u003c/li\u003e\n\u003cli\u003eRiaz N, Havel J J, Makarov V, et al. Tumor and Microenvironment Evolution during Immunotherapy with Nivolumab[J]. Cell, 2017, 171(4): 934-949.e16.\u003c/li\u003e\n\u003cli\u003eLauss M, Donia M, Harbst K, et al. Mutational and putative neoantigen load predict clinical benefit of adoptive T cell therapy in melanoma[J]. Nat Commun, 2017, 8(1): 1738.\u003c/li\u003e\n\u003cli\u003eVan Allen E M, Miao D, Schilling B, et al. Genomic correlates of response to CTLA-4 blockade in metastatic melanoma[J]. Science, 2015, 350(6257): 207-211.\u003c/li\u003e\n\u003cli\u003eGide T N, Quek C, Menzies A M, et al. Distinct Immune Cell Populations Define Response to Anti-PD-1 Monotherapy and Anti-PD-1/Anti-CTLA-4 Combined Therapy[J]. Cancer Cell, 2019, 35(2): 238-255.e6.\u003c/li\u003e\n\u003cli\u003eRose T L, Weir W H, Mayhew G M, et al. Fibroblast growth factor receptor 3 alterations and response to immune checkpoint inhibition in metastatic urothelial cancer: a real world experience[J]. Br J Cancer, 2021, 125(9): 1251-1260.\u003c/li\u003e\n\u003cli\u003eMariathasan S, Turley S J, Nickles D, et al. TGFbeta attenuates tumour response to PD-L1 blockade by contributing to exclusion of T cells[J]. Nature, 2018, 554(7693): 544-548.\u003c/li\u003e\n\u003cli\u003eSong M, Chan A T, Sun J. Influence of the Gut Microbiome, Diet, and Environment on Risk of Colorectal Cancer[J]. Gastroenterology, 2020, 158(2): 322-340.\u003c/li\u003e\n\u003cli\u003eBrennan C A, Garrett W S. Fusobacterium nucleatum - symbiont, opportunist and oncobacterium[J]. Nat Rev Microbiol, 2019, 17(3): 156-166.\u003c/li\u003e\n\u003cli\u003eHuang Y, Cao J, Zhu M, et al. Nontoxigenic Bacteroides fragilis: A double-edged sword[J]. Microbiol Res, 2024, 286: 127796.\u003c/li\u003e\n\u003cli\u003eLopez-Siles M, Duncan S H, Garcia-Gil L J, Martinez-Medina M. Faecalibacterium prausnitzii: from microbiology to diagnostics and prognostics[J]. ISME J, 2017, 11(4): 841-852.\u003c/li\u003e\n\u003cli\u003eSanders M E, Merenstein D J, Reid G, et al. Probiotics and prebiotics in intestinal health and disease: from biology to the clinic[J]. Nat Rev Gastroenterol Hepatol, 2019, 16(10): 605-616.\u003c/li\u003e\n\u003cli\u003eWong S H, Yu J. Gut microbiota in colorectal cancer: mechanisms of action and clinical applications[J]. Nat Rev Gastroenterol Hepatol, 2019, 16(11): 690-704.\u003c/li\u003e\n\u003cli\u003eNicolini A, Ferrari P. Involvement of tumor immune microenvironment metabolic reprogramming in colorectal cancer progression, immune escape, and response to immunotherapy[J]. Front Immunol, 2024, 15: 1353787.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":false,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Colorectal cancer, Multi-omics data integration, Gut microbiota composition, Machine learning-based prognostic model, Immunotherapy response prediction","lastPublishedDoi":"10.21203/rs.3.rs-5669193/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5669193/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eColorectal cancer (CRC) exhibits substantial heterogeneity in molecular subtypes and clinical outcomes. We performed an integrative analysis of multi-omics data from 274 CRC patients to investigate the impact of gut microbiota composition on prognosis, identify novel subtypes, and develop a machine learning-based prognostic model. Our microbiome analysis revealed significant differences between CRC and normal tissues. Multi-omics clustering identified two major CRC subtypes, CS1 and CS2, with distinct molecular characteristics and survival outcomes. We developed the Multi-Omics Integrative Clustering and Machine Learning Score (MCMLS) model, which demonstrated strong prognostic value in predicting patient survival and outperformed existing models. The MCMLS low-score group exhibited higher immune cell infiltration, increased metabolic pathway activity, and potentially better immunotherapy response. In contrast, the MCMLS high-score group showed higher mutation burden, fibroblast infiltration, and enrichment of cell adhesion and migration pathways. Bacterial analysis revealed differentially abundant bacteria associated with prognosis. Importantly, MCMLS consistently predicted immunotherapy response across six independent datasets. Our findings highlight the complex interplay between the gut microbiome, tumor microenvironment, and immune landscape in CRC, providing valuable insights for improving patient stratification and personalized treatment strategies.\u003c/p\u003e","manuscriptTitle":"Integrative Analysis of Multi-Omics Data and Gut Microbiota Composition Reveals Prognostic Subtypes and Predicts Immunotherapy Response in Colorectal Cancer Using Machine Learning","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-28 06:26:32","doi":"10.21203/rs.3.rs-5669193/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Accepted","date":"2025-06-24T08:12:09+00:00","index":"","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-04-21T21:11:49+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-04-10T08:11:10+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2025-03-29T09:41:33+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"981bf0fe-24d9-420d-bcd5-04680ed41820","owner":[],"postedDate":"April 28th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":47468339,"name":"Biological sciences/Cancer"},{"id":47468340,"name":"Biological sciences/Immunology"},{"id":47468341,"name":"Biological sciences/Microbiology"},{"id":47468342,"name":"Biological sciences/Systems biology"},{"id":47468343,"name":"Health sciences/Gastroenterology"}],"tags":[],"updatedAt":"2025-07-14T16:04:13+00:00","versionOfRecord":{"articleIdentity":"rs-5669193","link":"https://doi.org/10.1038/s41598-025-08915-1","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2025-07-12 15:57:57","publishedOnDateReadable":"July 12th, 2025"},"versionCreatedAt":"2025-04-28 06:26:32","video":"","vorDoi":"10.1038/s41598-025-08915-1","vorDoiUrl":"https://doi.org/10.1038/s41598-025-08915-1","workflowStages":[]},"version":"v1","identity":"rs-5669193","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5669193","identity":"rs-5669193","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00