Benchmarking multi-omics integrative clustering methods for subtype identification in colorectal cancer

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Background: and objectives Colorectal cancer (CRC) represents a heterogeneous malignancy that has concerned global burden of incidence and mortality. The traditional tumor-node-metastasis staging system has exhibited certain limitations. With the advancement of omics technologies, researchers are directing their focus on developing a more precise multi-omics molecular classification. Therefore, the utilization of unsupervised multi-omics integrative clustering methods in CRC, advocating for the establishment of a comprehensive benchmark with practical guidelines. In this study, we obtained CRC multi-omics data, encompassing DNA methylation, gene expression, and protein expression from the TCGA database. We then generated interrelated CRC multi-omics data with various structures based on realistic multi-omics correlations, and performed a comprehensive evaluation of eight representative methods categorized as early integration, intermediate integration, and late integration using complementary benchmarks for subtype classification accuracy. Lastly, we employed these methods to integrate real-world CRC multi-omics data, survival and differential analysis were used to highlight differences among newly identified multi-omics subtypes. Results Through in-depth comparisons, we observed that similarity network fusion (SNF) exhibited exceptional performance in integrating multi-omics data derived from simulations. Additionally, SNF effectively distinguished CRC patients into five subgroups with the highest classification accuracy. Moreover, we found significant survival differences and molecular distinctions among SNF subtypes. Conclusions The findings consistently demonstrate that SNF outperforms other methods in CRC multi-omics integrative clustering. The significant survival differences and molecular distinctions among SNF subtypes provide novel insights into the multi-omics perspective on CRC heterogeneity with potential clinical treatment. The code and its implementation are available in GitHub https://github.com/zsbvb/Comparison-of-Multiomics-Integration-Methods-for-CRC.
Full text 143,935 characters · extracted from preprint-html · click to expand
Benchmarking multi-omics integrative clustering methods for subtype identification in colorectal cancer | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Benchmarking multi-omics integrative clustering methods for subtype identification in colorectal cancer Shuai Zhang, Jiali Lv, Zhe Fan, Bingbing Gu, Bingbing Fan, Chunxia Li, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4106569/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background and objectives Colorectal cancer (CRC) represents a heterogeneous malignancy that has concerned global burden of incidence and mortality. The traditional tumor-node-metastasis staging system has exhibited certain limitations. With the advancement of omics technologies, researchers are directing their focus on developing a more precise multi-omics molecular classification. Therefore, the utilization of unsupervised multi-omics integrative clustering methods in CRC, advocating for the establishment of a comprehensive benchmark with practical guidelines. In this study, we obtained CRC multi-omics data, encompassing DNA methylation, gene expression, and protein expression from the TCGA database. We then generated interrelated CRC multi-omics data with various structures based on realistic multi-omics correlations, and performed a comprehensive evaluation of eight representative methods categorized as early integration, intermediate integration, and late integration using complementary benchmarks for subtype classification accuracy. Lastly, we employed these methods to integrate real-world CRC multi-omics data, survival and differential analysis were used to highlight differences among newly identified multi-omics subtypes. Results Through in-depth comparisons, we observed that similarity network fusion (SNF) exhibited exceptional performance in integrating multi-omics data derived from simulations. Additionally, SNF effectively distinguished CRC patients into five subgroups with the highest classification accuracy. Moreover, we found significant survival differences and molecular distinctions among SNF subtypes. Conclusions The findings consistently demonstrate that SNF outperforms other methods in CRC multi-omics integrative clustering. The significant survival differences and molecular distinctions among SNF subtypes provide novel insights into the multi-omics perspective on CRC heterogeneity with potential clinical treatment. The code and its implementation are available in GitHub https://github.com/zsbvb/Comparison-of-Multiomics-Integration-Methods-for-CRC . Multi-omics Colorectal cancer Data integration Subtype identification Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Colorectal cancer (CRC) is a malignancy that exhibits a highly global incidence and mortality rate [ 1 ]. Staging and classification play a crucial role in developing treatment plans and assessing prognosis after a qualitative diagnosis of CRC. Despite there has been a substantial amount of significant research on primary CRC subtypes, the current clinical guidelines for classifying CRC patients solely depend on tumor-node-metastasis (TNM) staging [ 2 , 3 ]. Nonetheless, it is common to observe CRC patients at the same TNM stage exhibiting heterogeneous responses and prognoses to identical treatment [ 4 , 5 ]. Therefore, there is a significant emphasis on exploring more comprehensive and effective CRC classification system in clinical research. The rapid advancements in high-throughput sequencing, mass-spectrometric technology, and computational science have produced a vast amount of omics data derived from biological molecules, including epigenomics, genomics, transcriptomics, proteomics, and metabolomics. These molecular events exhibit hierarchical relationships and interactions, culminating in a complex molecular network, and it is challenging to capture the entirety of this complexity using only single-level omics data [ 6 , 7 ]. Multi-omics integration analysis combines diverse omics data to systematically investigate biological samples and uncover interactions among various components within biological systems [ 8 ]. Consequently, it has resulted in the transformation of TNM staging into the integration of microscopic factors derived from muti-omics approaches, promoting the improvement and enhancement of the prognostic and predictive evaluation system in the field of clinical oncology [ 9 – 15 ]. Various integrative approaches, classified in early integration, intermediate integration, and late integration based on integrative strategies [ 16 , 17 ], enable the joint clustering analysis of multiple omics data. Recent studies have discussed some of these multi-omics integrative clustering methods. Pierre-Jean and colleagues conducted a comparative analysis of 13 unsupervised multi-omics integrative methods with respect to clustering and variable selection using eight simulated benchmark datasets [ 18 ]. Duan and colleagues comprehensively combined classification accuracy, clinical significance, robustness, and computational efficiency to assess ten representative multi-omics integrative clustering methods established on benchmark datasets of nine cancer types [ 19 ]. Tini and associates evaluated five unsupervised multi-omics integrative clustering methods and explored their classification accuracy in the context of various experimental designs, feature selection, parameter training, noise using simulated datasets and three real datasets [ 20 ]. Indeed, various cancer types exhibit a multitude of heterogeneous biological variables and complex multi-omics data structures, conducting comprehensive evaluations simultaneously may not accurately depict the unique characteristics of each cancer [ 19 ]. Moreover, expect for the attributes of limited sample sizes, high dimensionality, and substantial noise, multi-omics data following the central dogma exhibits significant correlations across various omics levels [ 21 ]. For instance, hypermethylation of CpG islands in the upstream region typically leads to lower gene expression in downstream, and higher gene expression is frequently associated with higher protein expression in downstream [ 22 ]. High correlation represents a crucial feature that must not be overlooked in multi-omics integration analysis. Therefore, the application of unsupervised multi-omics integrative clustering methods in CRC requires the establishment of a comprehensive benchmark with practical guidelines. In our study, we generated interrelated CRC multi-omics data with extensive simulated scenarios and evaluated eight representative multi-omics integrative clustering methods (LRAcluster [ 23 ], SNF [ 24 ], CIMLR [ 25 ], Mocluster [ 26 ], intNMF [ 27 ], MCIA [ 28 ], iClusterPlus [ 29 ] and PINSPlus [ 30 ]) using complementary benchmarks for CRC subtype classification accuracy. Building upon the results of the systematic simulated comparison, we further investigated the classification accuracy of these methods for CRC subtyping using realistic multi-omics data and demonstrated that similarity network fusion (SNF) consistently outperforms other tools. The novel SNF subtypes exhibited significant survival differences and molecular distinctions, providing valuable insights into the underlying biology of CRC heterogeneity with potential clinical treatment (Fig. 1 ). Methods We conducted a systematic multi-omics integrative evaluation using eight widely adopted and representative methods in various integration strategies. These methods were outlined in Table 1 and introduced in Supplementary Methods . It is important to note that within these categorizations, a specific method may be classified into multiple integration strategies. Table 1 Description of selected multi-omics integrative clustering methods. Integration strategy Type Method Full title Early integration LRAcluster Low-rank approximation based multi-omics data clustering Intermediate integration Similarity-based SNF Similarity network fusion CIMLR Cancer integration via multikernel learning Dimension reduction Mocluster Identifying joint patterns across multiple omics data intNMF Integrative non-negative matrix factorization MCIA Multiple co-inertia analysis Statistical modeling iClusterPlus Integrative clusterings in cancer genomic data Late integration PINSPlus Perturbation clustering for data integration and subtyping Generation of simulated datasets The generation process of simulated CRC multi-omics data was illustrated in Fig. 2 A. Initially, we identified a series of regulatory pathways involving 182 methylated sites, 81 genes, 93 proteins based on TCGA dataset by querying biological databases. Subsequently, following the simulation design by Chalise and colleagues [ 31 ], we calculated the Pearson correlation \(\rho\) ( Table. S1 and Table. S2 ) and the variance-covariance matrix for regulatory pairs to generate simulated data. We assumed that partial variables within each omics exhibit actual differences between subgroups, while other variables are irrelevant to subgroups. Here, \(n\) represents the sample size in each omics dataset, \(p\) denotes to the proportion of differentially expressed variables, and \(\sigma\) signifies the standard deviation of irrelevant variables (noise). We postulated that the number of ground-truth subgroups among \(n\) patients is denoted as \(k\) , and the mean level of differentially variable differs in subgroups. Where \({ \mu }_{ }\) represents the mean level, \({\delta }_{ }\) denotes the presetting shift among subgroups. DNA methylation Methylation \(\beta\) -values, ranging from 0 to 1 and approximating a beta distribution, were generated based on 182 methylated sites, with 0 representing unmethylated, and 1 indicating methylated completely. Initially, we performed the logit transformation on methylation \(\beta\) -values, producing methylation \(M\) -values spanning from − \(\infty\) to \(\infty\) , which follow a Gaussian distribution. Subsequently, we generated differentially methylated sites (DMSs) using the binomial distribution with probability \(p\) , and calculated the effect of the methylated site as follows: $${\text{ Effect }}_{d}={\mu }_{d}+DMS\times {{\delta }_{DMS}}_{ }$$ Where \({ \mu }_{d}\) represents the mean level of the methylation \(M\) -value, \(DMS\) is a binary variable where 1 signifies a differentially methylated site and 0 means an irrelevant methylated site, \({\delta }_{DMS}\) denotes the presetting shift among subgroups. Then, we simulated the multivariate Gaussian distribution with a sample size of \(n\) based on \({\text{Effect }}_{d}\) and the variance-covariance matrix, we introduced noise of irrelevant variables by adjusting the standard deviation \(\sigma\) . Finally, we applied the reverse logit transformation to convert methylation \(M\) -values back to methylation \(\beta\) -values. Gene expression We generated simulated gene expression data using the expression levels of 81 genes and the correlations between methylated sites and genes, and determined differentially expressed genes (DEGs) using a binomial distribution with probability \(p\) , as described in DMSs. The effect on gene expression was calculated as follows. If multiple methylated sites regulate a gene, the expression levels of those methylated sites were averaged to calculate a unified correlation coefficient. $${\text{ Effect }}_{g}=\left({\rho }_{d}\times {\mu }_{d}+\sqrt{1-{{\rho }_{d}}^{2}}\times {\mu }_{g}\right)+DEG\times {{\delta }_{DEG}}_{ }$$ Here, \({\rho }_{d}\) represents the Pearson correlation of methylation-gene regulatory pairs, and \({\mu }_{g}\) is the mean value of gene expression. A multivariate Gaussian distribution of gene expression, with a sample size of \(n\) , was produced based on \({\text{Effect }}_{g}\) and the variance-covariance matrix. We introduced noise of irrelevant variables by adjusting the standard deviation \(\sigma\) . Protein expression Increased gene expression may result in elevated expression of downstream proteins. We generated simulated protein expression data based on the expression levels of 93 proteins and their correlations with genes. In accordance with DEGs, we determined the effect of differentially expressed proteins (DEPs) as follows: $${\text{ Effect }}_{p}=\left({\rho }_{g}\times {\mu }_{g}+\sqrt{1-{{\rho }_{g}}^{2}}\times {\mu }_{p}\right)+DEP\times {{\delta }_{DEP}}_{ }$$ Where \({\rho }_{g}\) represents the Pearson correlation of gene-protein regulatory pairs, \({\mu }_{p}\) is the mean value of protein expression. We generated a multivariate Gaussian distribution of protein expression with a sample size of \(n\) based on \({\text{Effect }}_{p}\) and the variance-covariance matrix, and introduced noise of irrelevant variables by adjusting the standard deviation \(\sigma\) . Setting of simulated scenarios We investigated the impact of the following simulation parameters on classification accuracy of multi-omics integrative clustering methods: (1) sample sizes, (2) numbers of subgroups, (3) shifts among subgroups, (4) proportions of differentially expressed variables, (5) noises, (6) ratios of sample size among subgroups, and (7) combinations of omics data. Each simulated scenario repeated 50 times, and we calculated the average results from these 50 simulations. Table 2 provided an overview of the parameter setting in the simulation experiment. The basic simulated scenario was \(n\) = 150, \(k\) = 3, \(\delta\) = 2, \(p\) = 10%, \(\sigma\) = 1, and \(c\) = (0.2, 0.3, 0.5). We evaluated the consistency between integrative clusters and ground-truth subgroups using the adjusted rand index (ARI), normalized mutual information (NMI), and F1-score that described in detail in Supplementary Methods . Table 2 Overview of parameter setting in simulated experiment. Simulation parameter Symbol Values Sample size \(n\) 50, 100, 150, 200, 250, 300, 350, 400, 450, 500 Number of subgroups \(k\) 2, 3, 4, 5, 6 Shift among subgroups \(\delta\) 0, 0.5, 1, 1.5, 2, 2.5, 3 Proportion of differentially expressed variables \(p\) 1%, 5%, 10%, 15%, 20%, 25% Noise \(\sigma\) 0.5, 1, 2, 3 Ratio of sample size among subgroups \(c\) balanced, moderately unbalanced, and extremely unbalanced Combination of omics data four combinations of DNA methylation, gene expression, and protein expression Processing and analysis of real-world colorectal cancer data We used the real CRC multi-omics data from TCGA to verify the classification accuracy of those methods. Initially, we selected a cohort of 331 CRC patients with epigenomics, transcriptomics, proteomics, and clinical data. After mitigating the low-variation variables based on median absolute deviation (MAD), we applied standardization to all features to remove scaling-related biases. Subsequently, we characterized the consensus molecular subtypes (CMS) as the ground-truth subgroups, CMS is a well-established transcriptomic-based classification of colorectal cancer comprising four subtypes (CMS1/microsatellite instability immune, CMS2/canonical, CMS3/metabolic, and CMS4/mesenchymal) [ 32 ]. We classified patients into CMS subtypes using random forest method, and evaluated the classification accuracy of each method with optimal clustering number based on ARI, NMI, F1-score. Consistent with the results observed in the simulated study, the real CRC multi-omics data shows that the classification accuracy of SNF remained superior compared to other methods. In addition, regarding clustering number determination, we assessed the classification accuracy across various number of SNF subgroups. The SNF subgroup exhibiting the highest classification accuracy was preserved for further analysis. Lastly, we conducted a comparative analysis of SNF subgroup in patient survival, molecular expression and somatic mutation. Associations between subtypes and overall survival were then calculated by Kaplan–Meier analysis using a log-rank test. Cox proportional hazard regression analysis was performed to estimated hazard ratios and corresponding P values between specific groups of interest. Somatic mutation calls from the TCGA database, processed by Varscan2 [ 33 ], were utilized to investigate genomic alteration among SNF subgroups. The downloaded mutation annotation format (MAF) file comprised 285 CRC patients and 16,981 genes (Hugo Symbol). We also examined the differences on transcriptomics features between pairwise SNF subgroups, and applied the Benjamini–Hochberg false discovery rate (FDR) for multiple testing correction. DEGs were confirmed according to the negative binomial distribution model implemented in edgeR [ 34 ] (FDR 1.5). Results Generation of simulated colorectal cancer datasets Simulated CRC multi-omics data, encompassing DNA methylation, gene expression, and protein expression were generated. Noteworthy, ensuring high quality of simulated data is of paramount importance. Here, we compared the probability density distributions of both the original data and simulated data under the condition of no subgroup ( \(\delta\) = 0). As depicted in Fig. S1 , the consistency in distributional properties and inherent patterns of single-omics features between the original and simulated data suggests that the simulated data effectively substitutes for the original data in our study. To further demonstrate the influence of the critical parameter \(\delta\) on the simulated data, we produced simulated datasets with varying shifts among subgroups ( \(\delta\) = 0, 2, and 2.5) while maintaining \(k\) = 3. As exhibited in Fig. S2 and Fig. S3 , the heatmaps and principal component analysis (PCA) plots for each omics dataset highlight that the subgroup separation becomes more pronounced as \(\delta\) increases. Ensemble multi-omics integrative methods on simulated datasets Each multi-omics integrative clustering method determines the optimal number of clusters according to own criteria. Our study evaluated the consistency between the optimal clustering number of methods and the number of ground-truth subgroups ( \(k\) = 3) based on the basic simulated scenario. Figure 2 B illustrates that the range of optimal clustering numbers for these methods, varying from 2 to 6. Except for iClusterPlus, all methods had a median optimal clustering number that matched the ground-truth value ( \(k\) = 3). iClusterPlus exhibited the least favorable performance, with a range from 3 to 5 and a median of 5. LRAcluster and CIMLR consistently outperformed other methods in each simulation, yielding an optimal clustering number of 3. SNF and intNMF showed a consistent distribution of optimal clustering numbers, ranging from 2 to 3. In contrast, PINSPlus, MCIA, and MoCluster displayed a relatively dispersed distribution, varying from 2 to 6. The evaluation results revealed that the optimal clustering number varies across different methods. To minimize the effect of varying optimal clustering numbers on subtypes classification accuracy, we adjusted the clustering numbers of each method to match the number of ground-truth subgroups during the subsequent simulated evaluation. Afterwards, we investigated the impact of the various simulation parameters on classification accuracy and evaluated the consistency between integrative clusters and ground-truth subgroups using the ARI, NMI, and F1-score. To investigate the impact of sample size on the classification accuracy of these methods, we generated simulated data by adjusting the sample size (Table 2 ) within the basic simulated scenario. From Fig. 3 A, we observed that CIMLR exhibits lower classification accuracy for small sample sizes ( \(n\) < 100). As the sample size increases, the classification accuracy of CIMLR substantially improves and eventually reaches stability. In contrast, the classification accuracy of other methods displayed a slight fluctuation, typically within a 0.1 range. For LRAcluster, SNF, MoCluster, and MCIA, the ARI, NMI, and F1-scores consistently exceeded 0.95 and remained stable across various sample sizes. Conversely, iClusterPlus displayed a decreasing trend in classification accuracy as the sample size increased, although all three evaluated indicators were still close to 0.95. Notably, intNMF and PINSPlus underperformed across different sample sizes, achieving ARI and NMI scores below 0.7, with an F1-score near 0.75. Detailed simulation results are available in Table. S3 . We next performed multi-omics integration analysis by varying the number of subgroups (ranging from 2 to 6) under basic simulated scenario. It became evident that most methods demonstrated preferable classification accuracy when dealing with 2 or 3 subgroups. As the number of subgroups increased, the classification accuracy of the methods generally decreased (Fig. 3 B). LRAcluster and SNF exhibited superior classification accuracy compared to other methods, with all three evaluation indicators approaching 1 and displaying stability across various simulated scenarios. iClusterPlus showed slight fluctuations in ARI and NMI, hovering around 0.95, and achieved an F1-score exceeding 0.95. While the classification accuracy of MCIA and CIMLR notably declined. Specifically, MoCluster was greatly influenced by the number of subgroups, resulting in a decrease of 0.2 to 0.3 in classification accuracy. Detailed simulation results can be found in Table. S4 . Each Cluster and associated features were created by applying a fixed shift to their mean values. We varied shift among subgroups through a range of values (Table 2 ) under basic simulated scenario and observed as the shift increased, all methods exhibited enhanced classification accuracy. SNF performed exceptionally well overall but struggled to accurately identify subgroups only when a weak shift was present among subgroups ( \(\delta\) < 1). Mocluster, LRAcluster, MCIA, and iClusterPlus demonstrated similar patterns, with substantial increases in classification accuracy before \(\delta\) = 1.5, followed by stability. In contrast, CIMLR exhibited substantial improvements until \(\delta\) = 2. On the other hand, PINSPlus and intNMF achieved ARI and NMI less than 0.7, with an F1-score approaching 0.75 when \(\delta\) = 2 (Fig. 3 C). Further details can be found in Table. S5 . We determined differentially expressed variables by adjusting the probability in the binomial distribution, and traversed \(p\) across a range of values (Table 2 ) under basic simulated scenario. Figure 3 D illustrates that as \(p\) increased, there was a noticeable enhancement in classification accuracy for all methods. Most methods experienced substantial improvements before \(p\) = 10% and plateaued thereafter. Notably, SNF, Mocluster, LRAcluster, and MCIA displayed poor performance in the extreme scenario of \(p\) = 1%, while in other scenarios, their performance was close to 1 and remained stable. In contrast, CIMLR and iClusterPlus achieved satisfactory classification accuracy at \(p\) = 10%. PINSPlus and intNMF required a higher proportion of differentially expressed variables ( \(p\) = 25%) to achieve comparable performance levels. Detailed simulation results can be found in Table. S6 . Subsequently, we adjusted the standard deviation of irrelevant variables to generate the noise level \(\sigma\) , and varied \(\sigma\) by times of 0.5, 1, 2, and 3 compared to the original standard deviation under basic simulated scenario. Overall, as noise increased, most methods exhibited a decrease in classification accuracy, particularly when \(\sigma\) reached 2 to 3 (Fig. 3 E). SNF and iClusterPlus consistently maintained high classification accuracy. CIMLR, Mocluster, MCIA, and LRAcluster displayed a noticeable decline in classification accuracy after \(\sigma\) reached 2, with LRAcluster experiencing the most pronounced drop. For, intNMF and PINSPlus, the decline began when \(\sigma\) reached 1. Comprehensive simulation results can be found in Table. S7 . Here, we explored the influence of varying ratios of sample size among subgroups (balanced, moderately imbalanced, and extremely imbalanced) on these methods under basic simulated scenario. The balanced, moderately imbalanced, and extremely imbalanced ratios of the three subgroups were \(c\) = (1/3, 1/3, 1/3), \(c\) = (0.2, 0.3, 0.5), and \(c\) = (0.1, 0.3, 0.6) respectively. In the scenario of four subgroups, the balanced, moderately imbalanced, and extremely imbalanced ratios were \(c\) = (0.25, 0.25, 0.25, 0.25), \(c\) = (0.1, 0.2, 0.3, 0.4), and \(c\) = (0.1, 0.1, 0.2, 0.6) respectively. We observed that LRAcluster, SNF, and MCIA consistently achieved classification accuracy exceeding 0.95 across various scenarios of \(k\) and \(c\) . For \(k\) = 3, iClusterPlus, CIMLR, and Mocluster displayed satisfactory classification accuracy. However, in cases of moderately and extremely imbalanced with \(k\) = 4, these three algorithms exhibited a downward trend to varying degrees. Similarly, intNMF and PINSPlus delivered comparatively less satisfactory performance (Fig. 3 F). Detailed simulation results are presented in Table. S8 . To explore the effectiveness of various omics combinations, we simulated four combinations of multi-omics data scenarios: M + G + P, M + G, G + P, and M + P (M, G, and P represent DNA methylation, gene expression, and protein expression, respectively) under the basic simulated scenario. As shown in Fig. 3 G, there is no combination was universally effective for all eight methods. Generally, the combination involving three omics data demonstrated superior classification accuracy compared to other combinations, with except for MCIA and PINSPlus. Notably, the combination of M + G showcased outstanding classification accuracy for these two algorithms. SNF and CIMLR consistently delivered stable performance across all combinations. The ARI, NMI, and F1-score of SNF approached 1 in all combinations, and the medians of the three evaluation indicators for CIMLR exceeded 0.85. Particularly, the classification accuracy of LRAcluster and iClusterPlus decreased largely in the M + P and G + P combinations, respectively. Application in realistic colorectal cancer datasets of multi-omics integrative methods In our practical study, we included a cohort of 331 CRC patients with epigenomics, transcriptomics, proteomics, and clinical data from the TCGA database ( Fig. S4 ). Comprehensive basic information regarding the enrolled CRC patients was presented in Table. S9 . The initial number of features for epigenomics, transcriptomics, and proteomics were 22,601, 55,026, and 198, respectively. Subsequently, we selected the top 1000 features with the highest MAD values in epigenomics and transcriptomics data and the leading 100 features in proteomics for further analysis. Consistent with our findings in simulated experiments, the analysis of real CRC multi-omics data demonstrated that the classification accuracy of SNF consistently surpassed methodologies when employing the optimal number of subgroups, yielding ARI of 0.227, NMI of 0.254, and F1-score of 0.576 ( Fig. S5 ). Figure 4 A illustrated the classification performance of SNF with clustering numbers ranging from 2 to 6 on the real dataset. SNF effectively segregated CRC patients into 5 distinct subgroups, exhibiting the highest classification accuracy: an ARI of 0.326, NMI of 0.337, and F1-score of 0.630. Subsequently, we adopted the 5-category SNF clustering as the primary CRC subgroups for further analysis. The t-SNE plots for dimensional reduction revealed that a discernible separation trend among SNF subgroups within each omics dataset ( Fig. S6 ). SNF subgroups had distinct overall survival (OS), which SNF-V was associated with poor survival (Fig. 4 B), and SNF I-IV were significantly associated with improved OS compared to SNF-V in Cox proportional hazard regression analysis. CMS subgroups also exhibited significant differences in OS, however, contrary to expectations, CMS1 and CMS2 subgroups were not significantly associated with improved OS compared to CMS4 (Fig. 4 C). In the Cox multivariate analysis, SNF remained significantly associated with improved OS (with age and pathological stage), whereas CMS subtypes did not ( Table. S10 ). The relationships between SNF subgroups and CMS subgroups were illustrated in Fig. 4 D. Specifically, the majority of SNF-I (84.3%) belonged to the CMS4 subtype. SNF-II predominantly (76.4%) included CMS2 subtype patients, while SNF-III primarily comprised 21 CMS2 patients (29.6%) and 42 CMS3 patients (59.2%). Moreover, SNF-IV and SNF-V mainly consisted of CMS2 (74.2%) and CMS1 (75.0%) subgroups, respectively. In contrast, the SNF subgroups for other clustering numbers did not significantly stratify patients with survival differences ( Fig. S7 ), and their relationships with CMS subgroups were less discernible compared to the 5-category SNF clustering. Additionally, we conducted a Wilcoxon rank sum test to compare acknowledged CRC biomarkers among SNF subgroups (Fig. 4 E-H). SNF subgroups also showed significant differences in term of TMB, MSI score, PD-L1 expression and TP53 expression. Notably, the SNF-V subgroup displayed the highest levels of TMB, MSI score and PD-L1 expression, providing additional justification for considering immune checkpoint blockade (ICB) as a potential therapeutic intervention. The panoramic landscapes of the 15 most frequently somatic altered genes for the SNF subgroups were represented in Fig. 4 I. Overall, the frequencies of specific mutated genes exhibited notable variations among SNF subgroups. APC , TP53 , TTN , and KRAS occupied the top four positions in SNF-I, SNF-II, SNF-III, and SNF-IV subgroups. However, in the SNF-V subgroup, the mutated frequencies of APC (33%), TP53 (38%), and KRAS (27%) were lower compared to other subtypes, while TTN (67%) remained displayed frequently mutated rates. Specifically, the SNF-V patients significantly enriched in somatic mutations of SYNE1 (52%), MUC16 (54%), FAT4 (52%), DNAH5 (31%), OBSCN (48%), ZFHX4 (42%), DNAH11 (38%), CSMD1 (35%), RYR2 (35%), and FLG (27%) in contrast to other SNF subgroups ( P < 0.01, Fisher exact test), serving as potential predictive biomarkers for ICB activity. In addition, we performed differential expression testing to illustrate significant differences in transcriptomics features among SNF subgroups. The upset plot visualized the number of differentially expressed genes between any pairwise SNF subgroups ( Fig. S8 ). We identified 323 differentially expressed genes in total, with the most substantial difference (37 genes) observed between the SNF-II and SNF-V subgroups. Moreover, 24 genes displayed significant expression difference in at least half of the 10 pairwise SNF subgroup comparisons, while 3 genes ( CTSE , REG1A , REG4 ) showed significant differences in 7 pairwise SNF subgroups. These findings suggest that these molecular alterations among SNF subgroups likely contribute to the carcinogenesis and progression of colorectal cancer through diverse mechanisms. Discussion Multi-omics integrative clustering methods have been developed to enhance the precision and personalization of treatments for CRC patients by identifying disease subtypes. Our study generated CRC multi-omics data with realistic correlations through simulated experiments, and evaluated the optimal number of clusters and classification accuracy of eight representative multi-omics integrative clustering methods. The simulated results revealed that CIMLR and MCIA provided the most accurate estimates for the number of clusters, and SNF displayed outstanding classification accuracy and robustness in various simulation scenarios. In summary, the similarity-based methods consider the heterogeneity of CRC multi-omics data and individual similarities by computing sample similarity matrices for each omics dataset, and had been widely used in multi-omics integration analysis. In contrast, other methods have certain limitations in CRC multi-omics integration. For instance, LRAcluster is sensitive to noise, iClusterPlus excels in classification accuracy and robustness but suffers from computational inefficiency. In addition, it is important to clarify that, while PINSPlus has unsatisfactory classification accuracy, which could be attributed to re-analyzing clustering results, it is a user-friendly method for automatically determining the optimal number of clusters without requiring complex parameter adjustments. Therefore, researchers can consider suitable methods for certain purposes or circumstances. Furthermore, the classification accuracy of these methods is sensitive to the number of subgroups, shifts among subgroups, proportions of differentially expressed variables, noise levels, and combinations of omics data. Researchers should take these factors into account when employing multi-omics integrative clustering methods. In agreement with our results, Pierre-Jean and colleagues also concluded SNF as the suitable method for individual classification [ 18 ]. Duan and associates recommended SNF to conduct general cancer subtyping tasks [ 19 ], our study also found the superiority of SNF of multi-omics integrative analysis in CRC classification accuracy. The possible reason is that SNF not only captures the shared structure from diverse omics data but also extracts the specific structure, enhancing its ability to preserve more information within the sample similarity network. Regarding the determination of the optimal number of clusters, we affirmed that LRAcluster performs excellently with highly correlated multi-omics data. In contrast, our observations deviate from the findings of Chauvel and colleagues, who reported that MoCluster appropriately identifies the number of clusters [ 35 ], we noted that the optimal clustering numbers determined by MoCluster under highly correlated multi-omics data quite differ from the number of ground-truth subgroups. One potential factor of interference might be that the inter-omics correlations mask MoCluster's capacity to discern subgroups. In the comprehensive analysis of realistic CRC multi-omics data, consistent with our simulation findings, SNF exhibited superior classification accuracy compared to other methods. SNF categorized CRC patients into five multi-omics molecular subgroups, demonstrating a distinct correspondence between these subgroups and CMS subtypes. Furthermore, there were various types of disparities among the SNF subgroups. SNF-V was associated with significant poor survival. Regarding somatic mutations, the SNF-V subgroup particularly exhibited notable variations in the frequencies of mutated genes. For instance, SYNE1 , MUC16 , and FAT4 had significantly higher mutation rates in SNF-V compared to other subgroups. These genes have previously been linked to the prognosis of CRC patients [ 36 – 38 ]. In the differentially expressed analysis, we identified differentially expressed markers among the SNF subgroups. In addition to well-established cancer-related biomarkers such as CXCL9 , MUC2 , and MUC4 . Furthermore, CTSE , REG1A , and REG4 exhibited differential expression between numerous pairwise SNF subgroups, may play roles in the progression of colorectal cancer through various mechanisms. Certainly, prior studies conducted by Astrosini et al. revealed REG1A as a molecular marker with prognostic value, associated specifically with peritoneal carcinomatosis in CRC [ 39 ]. Additionally, Rafa et al. highlighted REG4 as a multifunctional secreted protein that exerts effects on CRC cells through autocrine and paracrine mechanisms, potentially holding significance in the development and progression of CRC [ 40 ].Collectively, the disparities observed among SNF subtypes offer a novel multi-omics perspective on CRC heterogeneity with potential clinical treatment. Certain limitations in our work must be acknowledged. Firstly, our research exclusively focused on the epigenome-transcriptome-proteome due to their evident regulatory interplay. Additionally, in our simulated experiment, the number of CRC multi-omics features remained constant, precluding an investigation into the impact of varying feature numbers on classification accuracy, and during the assessment of classification accuracy, we fixed the number of clusters in the multi-omics integrative clustering analysis to match the actual number of ground-truth subgroups, without considering the optimal cluster numbers of these methods. Furthermore, we configured the model parameters of multi-omics integrative clustering methods based on the recommendations of each method, without accounting for the potential influence of these parameters. Lastly, due to limited access to expansive multi-omics CRC datasets, our practical analysis was solely based on TCGA, and the effectiveness of the SNF subgroups needs to be verified in subsequent studies. As the range of multi-omics data categories expands, there has been a notable advancement of multi-omics integration methods, especially those based on deep learning techniques. Future benchmarking frameworks for integration should encompass diverse multi-omics data types, comprising discrete data and single-cell data. Additionally, given that numerous methods can analyze multiple omics from the same samples, there is potential to develop methods capable of integrating incomplete data. Conclusions Accurately classifying CRC patients into molecular subtypes is essential for understanding tumor heterogeneity. The progression in computer science and omics technologies has facilitated the development of multi-omics integrative clustering algorithms. Therefore, the application of unsupervised multi-omics integration clustering methods in CRC necessitates the establishment of a comprehensive benchmark with practical guidance. Through simulation studies and case analyses, we demonstrate that the superior classification accuracy, robustness, and efficacy of SNF render it an exemplary method for identifying subtypes indicative of tumor heterogeneity that can be further extended to clinical practice. Abbreviations CRC Colorectal Cancer SNF Similarity Network Fusion TNM Tumor-Node-Metastasis DMSs Differentially Methylated Sites DEGs Differentially Expressed Genes DEPs Differentially Expressed Proteins ARI Adjusted Rand Index NMI Normalized Mutual Information MAD Median Absolute Deviation CMS Consensus Molecular Subtypes MAF Mutation Annotation Format FDR False Discovery Rate PCA Principal Component Analysis OS Overall Survival ICB Immune Checkpoint Blockade Declarations Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Availability of data and materials The simulated data are generated using our R function and can be found in Github repository https://github.com/zsbvb/Comparison-of-Multiomics-Integration-Methods-for-CRC. The epigenomics, proteomics, and clinical data for TCGA colorectal cancer are available at https://www.cbioportal.org. The transcriptomics and MAF file are downloaded from https://portal.gdc.cancer.gov. Competing interests The authors do not have any conflict of interest. Funding This study was supported by National Natural Science Foundation of China (82222064, 81973147 to TZ, 82304247 to CW), the Young Scholars Program of Shandong University (21320082164070 to CW), National Key Research and Development Program (2022YFC2010100 to TZ), and Shandong University Distinguished Young Scholars (to TZ). Authors' contributions SZ, JL, CW and TZ generated the hypothesis, compiled the simulated datasets and performed the evaluation. SZ wrote the original manuscript. FZ and BG contributed in result interpretation. BF and CL contributed to analytic strategy, supervised the study. All authors read and approved the final manuscript. Acknowledgements We would like to thank reviewers for their feedback on our manuscript. We would also like to thank the authors of all multi-omics integration methods used in this study for making their code and data freely available. Last, we thank the joint effort of many investigators and staff members in TCGA, and most importantly, the colorectal patients who have participated in this work. References Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin. 2021;71(3):209–49. Cervantes A, Adam R, Rosello S, Arnold D, Normanno N, Taieb J, Seligmann J, De Baere T, Osterlund P, Yoshino T, et al. Metastatic colorectal cancer: ESMO Clinical Practice Guideline for diagnosis, treatment and follow-up. Ann Oncol. 2023;34(1):10–32. Dekker E, Tanis PJ, Vleugels JLA, Kasi PM, Wallace MB. Colorectal cancer. Lancet. 2019;394(10207):1467–80. Nagtegaal ID, Quirke P, Schmoll HJ. Has the new TNM classification for colorectal cancer improved care? Nat Rev Clin Oncol. 2011;9(2):119–23. Nitsche U, Maak M, Schuster T, Kunzli B, Langer R, Slotta-Huspenina J, Janssen KP, Friess H, Rosenberg R. Prediction of prognosis is not improved by the seventh and latest edition of the TNM classification for colorectal cancer in a single-center collective. Ann Surg. 2011;254(5):793–800. discussion 800 – 791. Ruan W, Yuan X, Eltzschig HK. Circadian rhythm as a therapeutic target. Nat Rev Drug Discovery. 2021;20(4):287–307. Hu FB. Metabolic profiling of diabetes: from black-box epidemiology to systems epidemiology. Clin Chem. 2011;57(9):1224–6. Karczewski KJ, Snyder MP. Integrative omics for health and disease. Nat Rev Genet. 2018;19(5):299–310. Dumbill E. A Revolution That Will Transform How We Live, Work, and Think: An Interview with the Authors of Big Data. Big Data. 2013;1(2):73–7. Kristensen VN, Lingjaerde OC, Russnes HG, Vollan HK, Frigessi A, Borresen-Dale AL. Principles and methods of integrative genomic analyses in cancer. Nat Rev Cancer. 2014;14(5):299–313. Gunther OP, Chen V, Freue GC, Balshaw RF, Tebbutt SJ, Hollander Z, Takhar M, McMaster WR, McManus BM, Keown PA, et al. A computational pipeline for the development of multi-marker bio-signature panels and ensemble classifiers. BMC Bioinformatics. 2012;13:326. Singh A, Shannon CP, Gautier B, Rohart F, Vacher M, Tebbutt SJ, Le Cao KA. DIABLO: an integrative approach for identifying key molecular drivers from multi-omics assays. Bioinf (Oxford England). 2019;35(17):3055–62. van de Wiel MA, Lien TG, Verlaat W, van Wieringen WN, Wilting SM. Better prediction by use of co-data: adaptive group-regularized ridge regression. Stat Med. 2016;35(3):368–81. Wang T, Shao W, Huang Z, Tang H, Zhang J, Ding Z, Huang K. MOGONET integrates multi-omics data using graph convolutional networks allowing patient classification and biomarker identification. Nat Commun. 2021;12(1):3445. Roelands J, Kuppen PJK, Ahmed EI, Mall R, Masoodi T, Singh P, Monaco G, Raynaud C, de Miranda N, Ferraro L, et al. An integrated tumor, immune and microbiome atlas of colon cancer. Nat Med. 2023;29(5):1273–86. Rappoport N, Shamir R. Multi-omic and multi-view clustering algorithms: review and cancer benchmark. Nucleic Acids Res. 2018;46(20):10546–62. Zitnik M, Nguyen F, Wang B, Leskovec J, Goldenberg A, Hoffman MM. Machine Learning for Integrating Data in Biology and Medicine: Principles, Practice, and Opportunities. Inf Fusion. 2019;50:71–91. Pierre-Jean M, Deleuze JF, Le Floch E, Mauger F. Clustering and variable selection evaluation of 13 unsupervised methods for multi-omics data integration. Brief Bioinform. 2020;21(6):2011–30. Duan R, Gao L, Gao Y, Hu Y, Xu H, Huang M, Song K, Wang H, Dong Y, Jiang C, et al. Evaluation and comparison of multi-omics data integration methods for cancer subtyping. PLoS Comput Biol. 2021;17(8):e1009224. Tini G, Marchetti L, Priami C, Scott-Boyer MP. Multi-omics integration-a comparison of unsupervised clustering methodologies. Brief Bioinform. 2019;20(4):1269–79. Crick F. Central dogma of molecular biology. Nature. 1970;227(5258):561–3. Gry M, Rimini R, Stromberg S, Asplund A, Ponten F, Uhlen M, Nilsson P. Correlations between RNA and protein expression profiles in 23 human cell lines. BMC Genomics. 2009;10:365. Wu D, Wang D, Zhang MQ, Gu J. Fast dimension reduction and integrative clustering of multi-omics data using low-rank approximation: application to cancer molecular classification. BMC Genomics. 2015;16:1022. Wang B, Mezlini AM, Demir F, Fiume M, Tu Z, Brudno M, Haibe-Kains B, Goldenberg A. Similarity network fusion for aggregating data types on a genomic scale. Nat Methods. 2014;11(3):333–7. Ramazzotti D, Lal A, Wang B, Batzoglou S, Sidow A. Multi-omic tumor data reveal diversity of molecular mechanisms that correlate with survival. Nat Commun. 2018;9(1):4453. Meng C, Helm D, Frejno M, Kuster B. moCluster: Identifying Joint Patterns Across Multiple Omics Data Sets. J Proteome Res. 2016;15(3):755–65. Chalise P, Fridley BL. Integrative clustering of multi-level 'omic data based on non-negative matrix factorization algorithm. PLoS ONE. 2017;12(5):e0176278. Meng C, Kuster B, Culhane AC, Gholami AM. A multivariate approach to the integration of multi-omics datasets. BMC Bioinformatics. 2014;15:162. Mo Q, Wang S, Seshan VE, Olshen AB, Schultz N, Sander C, Powers RS, Ladanyi M, Shen R. Pattern discovery and cancer gene identification in integrated cancer genomic data. Proc Natl Acad Sci U S A. 2013;110(11):4245–50. Nguyen T, Tagett R, Diaz D, Draghici S. A novel approach for data integration and disease subtyping. Genome Res. 2017;27(12):2025–39. Chalise P, Raghavan R, Fridley BL. InterSIM: Simulation tool for multiple integrative 'omic datasets'. Comput Methods Programs Biomed. 2016;128:69–74. Guinney J, Dienstmann R, Wang X, de Reynies A, Schlicker A, Soneson C, Marisa L, Roepman P, Nyamundanda G, Angelino P, et al. The consensus molecular subtypes of colorectal cancer. Nat Med. 2015;21(11):1350–6. Koboldt DC, Zhang Q, Larson DE, Shen D, McLellan MD, Lin L, Miller CA, Mardis ER, Ding L, Wilson RK. VarScan 2: somatic mutation and copy number alteration discovery in cancer by exome sequencing. Genome Res. 2012;22(3):568–76. Robinson MD, McCarthy DJ, Smyth GK. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinf (Oxford England). 2010;26(1):139–40. Chauvel C, Novoloaca A, Veyre P, Reynier F, Becker J. Evaluation of integrative clustering methods for the analysis of multi-omics data. Brief Bioinform. 2020;21(2):541–52. Ge W, Hu H, Cai W, Xu J, Hu W, Weng X, Qin X, Huang Y, Han W, Hu Y, et al. High-risk Stage III colon cancer patients identified by a novel five-gene mutational signature are characterized by upregulation of IL-23A and gut bacterial translocation of the tumor microenvironment. Int J Cancer. 2020;146(7):2027–35. Li C, Xu J, Wang X, Zhang C, Yu Z, Liu J, Tai Z, Luo Z, Yi X, Zhong Z. Whole exome and transcriptome sequencing reveal clonal evolution and exhibit immune-related features in metastatic colorectal tumors. Cell Death Discov. 2021;7(1):222. Yu J, Wu WK, Li X, He J, Li XX, Ng SS, Yu C, Gao Z, Yang J, Li M, et al. Novel recurrently mutated genes and a prognostic mutation signature in colorectal cancer. Gut. 2015;64(4):636–45. Astrosini C, Roeefzaad C, Dai YY, Dieckgraefe BK, Jons T, Kemmner W. REG1A expression is a prognostic marker in colorectal cancer and associated with peritoneal carcinomatosis. Int J Cancer. 2008;123(2):409–13. Rafa L, Dessein AF, Devisme L, Buob D, Truant S, Porchet N, Huet G, Buisine MP, Lesuffleur T. REG4 acts as a mitogenic, motility and pro-invasive factor for colon cancer cells. Int J Oncol. 2010;36(3):689–98. Additional Declarations No competing interests reported. Supplementary Files SupplementaryMethods.docx SupplementaryResults.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4106569","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":281503820,"identity":"08981574-68b1-4d6e-83a8-3a47a112205d","order_by":0,"name":"Shuai Zhang","email":"","orcid":"","institution":"Department of Biostatistics, School of Public Health, Cheeloo College of Medicine, Shandong University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Shuai","middleName":"","lastName":"Zhang","suffix":""},{"id":281503821,"identity":"71b66aad-bc0c-4636-8973-e2945bad5784","order_by":1,"name":"Jiali Lv","email":"","orcid":"","institution":"Department of Biostatistics, School of Public Health, Cheeloo College of Medicine, Shandong University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jiali","middleName":"","lastName":"Lv","suffix":""},{"id":281503822,"identity":"681b554a-6080-4c75-b51f-f4e42853cc28","order_by":2,"name":"Zhe Fan","email":"","orcid":"","institution":"Department of Biostatistics, School of Public Health, Cheeloo College of Medicine, Shandong University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Zhe","middleName":"","lastName":"Fan","suffix":""},{"id":281503823,"identity":"dd4b83b2-dd02-46c3-b597-106f5e36bd87","order_by":3,"name":"Bingbing Gu","email":"","orcid":"","institution":"Department of Biostatistics, School of Public Health, Cheeloo College of Medicine, Shandong University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Bingbing","middleName":"","lastName":"Gu","suffix":""},{"id":281503824,"identity":"6d83336e-3fd5-4d57-a28e-a4ea2e41f8ed","order_by":4,"name":"Bingbing Fan","email":"","orcid":"","institution":"Department of Biostatistics, School of Public Health, Cheeloo College of Medicine, Shandong University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Bingbing","middleName":"","lastName":"Fan","suffix":""},{"id":281503825,"identity":"616aaadb-46de-44fc-b274-b381b6ac80fd","order_by":5,"name":"Chunxia Li","email":"","orcid":"","institution":"Department of Biostatistics, School of Public Health, Cheeloo College of Medicine, Shandong University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Chunxia","middleName":"","lastName":"Li","suffix":""},{"id":281503826,"identity":"5e6e728d-95fa-43d2-984e-eb023a53b159","order_by":6,"name":"Cheng Wang","email":"","orcid":"","institution":"Department of Biostatistics, School of Public Health, Cheeloo College of Medicine, Shandong University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Cheng","middleName":"","lastName":"Wang","suffix":""},{"id":281503827,"identity":"be95a33b-2918-47fe-832d-44be2a63b68d","order_by":7,"name":"Tao Zhang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA10lEQVRIiWNgGAWjYJACZhDBz8BgAGPjBzwwZZINJGsxOECsFnv2w4c/F9Tcsdt8/vA2CYYK68QG9rMH8NvCk5ZgPOPYs+RtN9LKJBjOpCc28OQlEHBYjkEyD9vhZLMbPGYSjG2HExskeAzwa+F/Y3CY59/hZOP+M0At/4jRIpFj2MzbdtjOgCEHqKWBGC03niUz8/YdTpC4kVZskXAs3biNJwe/Fvb+5MOfeb4dtufvP7zxxocaa9l+9jP4tcBAYgOITABiNqLUA4E9sQpHwSgYBaNgBAIAo4xBMyDKB/4AAAAASUVORK5CYII=","orcid":"","institution":"Department of Biostatistics, School of Public Health, Cheeloo College of Medicine, Shandong University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Tao","middleName":"","lastName":"Zhang","suffix":""}],"badges":[],"createdAt":"2024-03-15 09:23:15","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4106569/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4106569/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":53200530,"identity":"a73cbbae-5773-4e36-bf5a-08166698f6d9","added_by":"auto","created_at":"2024-03-21 19:27:44","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":681128,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eWorkflow overviewof simulated study and realistic data analysis based on multi-omics integration methods. \u003c/strong\u003eWorkflow overview of our study that divided into two subparts. Firstly, we identified a series of multi-omics regulatory pathways, calculated the Pearson correlation and variance-covariance matrix of regulatory pairs based on TCGA colorectal cancer database, and simulated multi-omics data with realistic correlations to evaluating the performance of the eight multi-omics integration methods in optimal number of clusters and classification accuracy. Additionally, we verify the performance of multi-omics integration methods in classification accuracy based on realistic TCGA colorectal cancer multi-omics data, and employed the best-performing SNF algorithm to conduct survival analysis and differential analysis. Abbreviations: TCGA, the cancer genome atlas; SNF, similarity network fusion.\u003c/p\u003e","description":"","filename":"Figure1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4106569/v1/bc3a7ce809f6af0e7dd8b96a.jpg"},{"id":53200527,"identity":"8c1861c6-5193-4926-a644-692b39671137","added_by":"auto","created_at":"2024-03-21 19:27:44","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":200769,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDistribution of optimal clustering numbers for multi-omics integration methods in simulated data. (A)\u003c/strong\u003eWorkflow of simulated data generation that follow the regulatory pathways from generating DMP using the binomial distribution, to generating DEP according to DEG. \u003cstrong\u003e(B)\u003c/strong\u003e Boxplot of the optimal clustering number of each method on the basic simulated scenario repeated 50 times. The black dashed line represented the ground-truth subgroups (k = 3). The boxplot compactly displays the distribution of a continuous variable. Data are visualized as five summary statistics (the median, two hinges and two whiskers), and all outlying points individually. Abbreviations: DMS, differentially methylated site; DEG, differentially expressed gene; DEP, differentially expressed protein.\u003c/p\u003e","description":"","filename":"Figure2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4106569/v1/cea1437b15d5b23a285c7a6b.jpg"},{"id":53200531,"identity":"f1d438fb-862b-41b1-97ed-99bb527c46cc","added_by":"auto","created_at":"2024-03-21 19:27:44","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":478356,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDistribution of classification accuracy for multi-omics integration methods in simulated data. \u003c/strong\u003eLine charts of the corresponding ARI, NMI, and F1-score between the clusters identified by each method and the ground-truth subgroups on the various simulated scenarios.\u003cstrong\u003e (A)\u003c/strong\u003e Sample size \u003cstrong\u003e(B)\u003c/strong\u003e Number of subgroups \u003cstrong\u003e(C)\u003c/strong\u003eShift among subgroups \u003cstrong\u003e(D)\u003c/strong\u003e Proportion of differentially expressed variables \u003cstrong\u003e(E)\u003c/strong\u003eNoise \u003cstrong\u003e(F)\u003c/strong\u003e Ratio of sample size among subgroups. \u003cstrong\u003e(G) \u003c/strong\u003eBox plot\u003cstrong\u003e \u003c/strong\u003eof the corresponding ARI, NMI, and F1-score for the four combinations of omics data. The x-coordinate of each line chart is the value of concerned simulated parameter when other simulated parameters are fixed. Each simulated scenario repeated 50 times, and we averaged the results of 50 simulations. Abbreviations: ARI, adjusted rand index; NMI, normalized mutual information; M, DNA methylation; G, gene expression; P, protein expression.\u003c/p\u003e","description":"","filename":"Figure3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4106569/v1/4dda738e9e54189a4f4ec5c8.jpg"},{"id":53200529,"identity":"1165bfba-9723-4cf8-93c1-496d1bbdf4e7","added_by":"auto","created_at":"2024-03-21 19:27:44","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":632183,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eIntegrative subtypes of colorectal cancer identified by multi-omics integration methods. (A) \u003c/strong\u003eBar charts of the corresponding ARI, NMI, and F1-score for the SNF clusters on preprocessed colorectal cancer multi-omics data. We take the CMS subgroups as the ground-truth labels to assess the classification accuracy across various number of SNF clusters ranging from 2 to 6. \u003cstrong\u003e(B)\u003c/strong\u003e Kaplan–Meier survival curve of SNF clusters for overall survival. \u003cstrong\u003e(C)\u003c/strong\u003e Kaplan–Meier survival curve of CMS subgroups for overall survival. HRs and 95% CIs are calculated by Cox proportional hazard regression. Overall \u003cem\u003eP\u003c/em\u003e value is calculated by log-rank test. Vertical lines indicate censor points. \u003cem\u003eP\u003c/em\u003e values are two-sided.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(D) \u003c/strong\u003eCircos plot of the relations between SNF and CMS subgroups. Size of each element is proportional to number of samples in each respective category. The corresponding number of patients are presented in Table. S9. \u003cstrong\u003e(E-H)\u003c/strong\u003e Box plots represent the distributions of TMB, MSI score, \u003cem\u003ePD-L1\u003c/em\u003e expression and \u003cem\u003eTP53\u003c/em\u003eexpression among SNF clusters. Overall \u003cem\u003eP\u003c/em\u003e value is calculated by Wilcoxon rank sum test.\u003cstrong\u003e (I)\u003c/strong\u003e Oncoprint of the somatic mutation distribution of the top 15 most frequently mutated genes in each SNF cluster. Each column represents a patient that ordered by SNF subgroups. Genes are ordered by mutational frequency. The dark green grids represent somatic mutation. 323 differentially expressed genes (FDR \u0026lt; 0.05 and log2 FC \u0026gt; 1.5 in each pairwise SNF clusters) are shown in heatmap, and we performed hierarchical clustering on the rows. Abbreviations: SNF, similarity network fusion; CMS, consensus molecular subtypes; ARI, adjusted rand index; NMI, normalized mutual information; FDR, false discovery rate; FC, fold change; TMB, tumor mutational burden; MSI, microsatellite instability; ****, \u003cem\u003eP\u003c/em\u003e\u0026lt; 0.0001; **, \u003cem\u003eP \u003c/em\u003e\u0026lt; 0.01.\u003c/p\u003e","description":"","filename":"Figure4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4106569/v1/30d3886a280083f6e1670a46.jpg"},{"id":53201445,"identity":"0934d31a-303a-4984-bcbb-5a841d370c87","added_by":"auto","created_at":"2024-03-21 19:43:46","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1186346,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4106569/v1/ed74dbf7-b804-45b6-85e6-a43e5213b811.pdf"},{"id":53201285,"identity":"429676ac-2d36-455f-af24-9c079b2129a9","added_by":"auto","created_at":"2024-03-21 19:35:44","extension":"docx","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":44643,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMethods.docx","url":"https://assets-eu.researchsquare.com/files/rs-4106569/v1/110116c78b0c01807ab9f618.docx"},{"id":53200532,"identity":"c70594cf-c091-4643-8c8b-6a09b48f37fb","added_by":"auto","created_at":"2024-03-21 19:27:45","extension":"docx","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":7856395,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryResults.docx","url":"https://assets-eu.researchsquare.com/files/rs-4106569/v1/663900cfea67b2f60631e93e.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Benchmarking multi-omics integrative clustering methods for subtype identification in colorectal cancer","fulltext":[{"header":"Introduction","content":"\u003cp\u003eColorectal cancer (CRC) is a malignancy that exhibits a highly global incidence and mortality rate [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Staging and classification play a crucial role in developing treatment plans and assessing prognosis after a qualitative diagnosis of CRC. Despite there has been a substantial amount of significant research on primary CRC subtypes, the current clinical guidelines for classifying CRC patients solely depend on tumor-node-metastasis (TNM) staging [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Nonetheless, it is common to observe CRC patients at the same TNM stage exhibiting heterogeneous responses and prognoses to identical treatment [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Therefore, there is a significant emphasis on exploring more comprehensive and effective CRC classification system in clinical research.\u003c/p\u003e \u003cp\u003eThe rapid advancements in high-throughput sequencing, mass-spectrometric technology, and computational science have produced a vast amount of omics data derived from biological molecules, including epigenomics, genomics, transcriptomics, proteomics, and metabolomics. These molecular events exhibit hierarchical relationships and interactions, culminating in a complex molecular network, and it is challenging to capture the entirety of this complexity using only single-level omics data [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Multi-omics integration analysis combines diverse omics data to systematically investigate biological samples and uncover interactions among various components within biological systems [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Consequently, it has resulted in the transformation of TNM staging into the integration of microscopic factors derived from muti-omics approaches, promoting the improvement and enhancement of the prognostic and predictive evaluation system in the field of clinical oncology [\u003cspan additionalcitationids=\"CR10 CR11 CR12 CR13 CR14\" citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eVarious integrative approaches, classified in early integration, intermediate integration, and late integration based on integrative strategies [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], enable the joint clustering analysis of multiple omics data. Recent studies have discussed some of these multi-omics integrative clustering methods. Pierre-Jean and colleagues conducted a comparative analysis of 13 unsupervised multi-omics integrative methods with respect to clustering and variable selection using eight simulated benchmark datasets [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Duan and colleagues comprehensively combined classification accuracy, clinical significance, robustness, and computational efficiency to assess ten representative multi-omics integrative clustering methods established on benchmark datasets of nine cancer types [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Tini and associates evaluated five unsupervised multi-omics integrative clustering methods and explored their classification accuracy in the context of various experimental designs, feature selection, parameter training, noise using simulated datasets and three real datasets [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIndeed, various cancer types exhibit a multitude of heterogeneous biological variables and complex multi-omics data structures, conducting comprehensive evaluations simultaneously may not accurately depict the unique characteristics of each cancer [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Moreover, expect for the attributes of limited sample sizes, high dimensionality, and substantial noise, multi-omics data following the central dogma exhibits significant correlations across various omics levels [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. For instance, hypermethylation of CpG islands in the upstream region typically leads to lower gene expression in downstream, and higher gene expression is frequently associated with higher protein expression in downstream [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. High correlation represents a crucial feature that must not be overlooked in multi-omics integration analysis. Therefore, the application of unsupervised multi-omics integrative clustering methods in CRC requires the establishment of a comprehensive benchmark with practical guidelines.\u003c/p\u003e \u003cp\u003eIn our study, we generated interrelated CRC multi-omics data with extensive simulated scenarios and evaluated eight representative multi-omics integrative clustering methods (LRAcluster [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e], SNF [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e], CIMLR [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e], Mocluster [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e], intNMF [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e], MCIA [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], iClusterPlus [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] and PINSPlus [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]) using complementary benchmarks for CRC subtype classification accuracy. Building upon the results of the systematic simulated comparison, we further investigated the classification accuracy of these methods for CRC subtyping using realistic multi-omics data and demonstrated that similarity network fusion (SNF) consistently outperforms other tools. The novel SNF subtypes exhibited significant survival differences and molecular distinctions, providing valuable insights into the underlying biology of CRC heterogeneity with potential clinical treatment (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eWe conducted a systematic multi-omics integrative evaluation using eight widely adopted and representative methods in various integration strategies. These methods were outlined in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and introduced in \u003cb\u003eSupplementary Methods\u003c/b\u003e. It is important to note that within these categorizations, a specific method may be classified into multiple integration strategies.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDescription of selected multi-omics integrative clustering methods.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIntegration strategy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eType\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFull title\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEarly integration\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLRAcluster\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLow-rank approximation based multi-omics data clustering\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"5\" rowspan=\"6\"\u003e \u003cp\u003eIntermediate integration\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSimilarity-based\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSNF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSimilarity network fusion\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCIMLR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCancer integration via multikernel learning\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eDimension reduction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMocluster\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eIdentifying joint patterns across multiple omics data\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eintNMF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eIntegrative non-negative matrix factorization\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMCIA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMultiple co-inertia analysis\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eStatistical modeling\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eiClusterPlus\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eIntegrative clusterings in cancer genomic data\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLate integration\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePINSPlus\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePerturbation clustering for data integration and subtyping\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eGeneration of simulated datasets\u003c/h2\u003e \u003cp\u003eThe generation process of simulated CRC multi-omics data was illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA. Initially, we identified a series of regulatory pathways involving 182 methylated sites, 81 genes, 93 proteins based on TCGA dataset by querying biological databases. Subsequently, following the simulation design by Chalise and colleagues [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e], we calculated the Pearson correlation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\rho\\)\u003c/span\u003e\u003c/span\u003e (\u003cb\u003eTable. S1\u003c/b\u003e and \u003cb\u003eTable. S2\u003c/b\u003e) and the variance-covariance matrix for regulatory pairs to generate simulated data.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe assumed that partial variables within each omics exhibit actual differences between subgroups, while other variables are irrelevant to subgroups. Here, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(n\\)\u003c/span\u003e\u003c/span\u003e represents the sample size in each omics dataset, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\\)\u003c/span\u003e\u003c/span\u003e denotes to the proportion of differentially expressed variables, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma\\)\u003c/span\u003e\u003c/span\u003e signifies the standard deviation of irrelevant variables (noise). We postulated that the number of ground-truth subgroups among \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(n\\)\u003c/span\u003e\u003c/span\u003e patients is denoted as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(k\\)\u003c/span\u003e\u003c/span\u003e, and the mean level of differentially variable differs in subgroups. Where\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({ \\mu }_{ }\\)\u003c/span\u003e\u003c/span\u003erepresents the mean level, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\delta }_{ }\\)\u003c/span\u003e\u003c/span\u003e denotes the presetting shift among subgroups.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eDNA methylation\u003c/h2\u003e \u003cp\u003eMethylation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\beta\\)\u003c/span\u003e\u003c/span\u003e-values, ranging from 0 to 1 and approximating a beta distribution, were generated based on 182 methylated sites, with 0 representing unmethylated, and 1 indicating methylated completely.\u003c/p\u003e \u003cp\u003eInitially, we performed the logit transformation on methylation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\beta\\)\u003c/span\u003e\u003c/span\u003e-values, producing methylation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(M\\)\u003c/span\u003e\u003c/span\u003e-values spanning from \u0026minus;\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\infty\\)\u003c/span\u003e\u003c/span\u003e to \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\infty\\)\u003c/span\u003e\u003c/span\u003e, which follow a Gaussian distribution. Subsequently, we generated differentially methylated sites (DMSs) using the binomial distribution with probability \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\\)\u003c/span\u003e\u003c/span\u003e, and calculated the effect of the methylated site as follows:\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$${\\text{ Effect }}_{d}={\\mu }_{d}+DMS\\times {{\\delta }_{DMS}}_{ }$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eWhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({ \\mu }_{d}\\)\u003c/span\u003e\u003c/span\u003erepresents the mean level of the methylation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(M\\)\u003c/span\u003e\u003c/span\u003e-value, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(DMS\\)\u003c/span\u003e\u003c/span\u003e is a binary variable where 1 signifies a differentially methylated site and 0 means an irrelevant methylated site, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\delta }_{DMS}\\)\u003c/span\u003e\u003c/span\u003e denotes the presetting shift among subgroups. Then, we simulated the multivariate Gaussian distribution with a sample size of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(n\\)\u003c/span\u003e\u003c/span\u003e based on \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\text{Effect }}_{d}\\)\u003c/span\u003e\u003c/span\u003e and the variance-covariance matrix, we introduced noise of irrelevant variables by adjusting the standard deviation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma\\)\u003c/span\u003e\u003c/span\u003e. Finally, we applied the reverse logit transformation to convert methylation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(M\\)\u003c/span\u003e\u003c/span\u003e-values back to methylation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\beta\\)\u003c/span\u003e\u003c/span\u003e-values.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eGene expression\u003c/h2\u003e \u003cp\u003eWe generated simulated gene expression data using the expression levels of 81 genes and the correlations between methylated sites and genes, and determined differentially expressed genes (DEGs) using a binomial distribution with probability \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\\)\u003c/span\u003e\u003c/span\u003e, as described in DMSs. The effect on gene expression was calculated as follows. If multiple methylated sites regulate a gene, the expression levels of those methylated sites were averaged to calculate a unified correlation coefficient.\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$${\\text{ Effect }}_{g}=\\left({\\rho }_{d}\\times {\\mu }_{d}+\\sqrt{1-{{\\rho }_{d}}^{2}}\\times {\\mu }_{g}\\right)+DEG\\times {{\\delta }_{DEG}}_{ }$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eHere, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\rho }_{d}\\)\u003c/span\u003e\u003c/span\u003e represents the Pearson correlation of methylation-gene regulatory pairs, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mu }_{g}\\)\u003c/span\u003e\u003c/span\u003e is the mean value of gene expression. A multivariate Gaussian distribution of gene expression, with a sample size of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(n\\)\u003c/span\u003e\u003c/span\u003e, was produced based on \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\text{Effect }}_{g}\\)\u003c/span\u003e\u003c/span\u003e and the variance-covariance matrix. We introduced noise of irrelevant variables by adjusting the standard deviation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eProtein expression\u003c/h2\u003e \u003cp\u003eIncreased gene expression may result in elevated expression of downstream proteins. We generated simulated protein expression data based on the expression levels of 93 proteins and their correlations with genes. In accordance with DEGs, we determined the effect of differentially expressed proteins (DEPs) as follows:\u003cdiv id=\"Equc\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$${\\text{ Effect }}_{p}=\\left({\\rho }_{g}\\times {\\mu }_{g}+\\sqrt{1-{{\\rho }_{g}}^{2}}\\times {\\mu }_{p}\\right)+DEP\\times {{\\delta }_{DEP}}_{ }$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eWhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\rho }_{g}\\)\u003c/span\u003e\u003c/span\u003e represents the Pearson correlation of gene-protein regulatory pairs, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mu }_{p}\\)\u003c/span\u003e\u003c/span\u003e is the mean value of protein expression. We generated a multivariate Gaussian distribution of protein expression with a sample size of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(n\\)\u003c/span\u003e\u003c/span\u003e based on \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\text{Effect }}_{p}\\)\u003c/span\u003e\u003c/span\u003e and the variance-covariance matrix, and introduced noise of irrelevant variables by adjusting the standard deviation \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eSetting of simulated scenarios\u003c/h2\u003e \u003cp\u003eWe investigated the impact of the following simulation parameters on classification accuracy of multi-omics integrative clustering methods: (1) sample sizes, (2) numbers of subgroups, (3) shifts among subgroups, (4) proportions of differentially expressed variables, (5) noises, (6) ratios of sample size among subgroups, and (7) combinations of omics data. Each simulated scenario repeated 50 times, and we calculated the average results from these 50 simulations. Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e provided an overview of the parameter setting in the simulation experiment. The basic simulated scenario was \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(n\\)\u003c/span\u003e\u003c/span\u003e = 150, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(k\\)\u003c/span\u003e\u003c/span\u003e = 3, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\delta\\)\u003c/span\u003e\u003c/span\u003e = 2, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\\)\u003c/span\u003e\u003c/span\u003e = 10%, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma\\)\u003c/span\u003e\u003c/span\u003e = 1, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(c\\)\u003c/span\u003e\u003c/span\u003e = (0.2, 0.3, 0.5). We evaluated the consistency between integrative clusters and ground-truth subgroups using the adjusted rand index (ARI), normalized mutual information (NMI), and F1-score that described in detail in \u003cb\u003eSupplementary Methods\u003c/b\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eOverview of parameter setting in simulated experiment.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSimulation parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSymbol\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eValues\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSample size\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(n\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e50, 100, 150, 200, 250, 300, 350, 400, 450, 500\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNumber of subgroups\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(k\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2, 3, 4, 5, 6\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eShift among subgroups\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\delta\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0, 0.5, 1, 1.5, 2, 2.5, 3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProportion of differentially expressed variables\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1%, 5%, 10%, 15%, 20%, 25%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNoise\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.5, 1, 2, 3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRatio of sample size among subgroups\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(c\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ebalanced, moderately unbalanced, and extremely unbalanced\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCombination of omics data\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003efour combinations of DNA methylation, gene expression, and protein expression\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eProcessing and analysis of real-world colorectal cancer data\u003c/h2\u003e \u003cp\u003eWe used the real CRC multi-omics data from TCGA to verify the classification accuracy of those methods. Initially, we selected a cohort of 331 CRC patients with epigenomics, transcriptomics, proteomics, and clinical data. After mitigating the low-variation variables based on median absolute deviation (MAD), we applied standardization to all features to remove scaling-related biases.\u003c/p\u003e \u003cp\u003eSubsequently, we characterized the consensus molecular subtypes (CMS) as the ground-truth subgroups, CMS is a well-established transcriptomic-based classification of colorectal cancer comprising four subtypes (CMS1/microsatellite instability immune, CMS2/canonical, CMS3/metabolic, and CMS4/mesenchymal) [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. We classified patients into CMS subtypes using random forest method, and evaluated the classification accuracy of each method with optimal clustering number based on ARI, NMI, F1-score. Consistent with the results observed in the simulated study, the real CRC multi-omics data shows that the classification accuracy of SNF remained superior compared to other methods. In addition, regarding clustering number determination, we assessed the classification accuracy across various number of SNF subgroups. The SNF subgroup exhibiting the highest classification accuracy was preserved for further analysis.\u003c/p\u003e \u003cp\u003eLastly, we conducted a comparative analysis of SNF subgroup in patient survival, molecular expression and somatic mutation. Associations between subtypes and overall survival were then calculated by Kaplan\u0026ndash;Meier analysis using a log-rank test. Cox proportional hazard regression analysis was performed to estimated hazard ratios and corresponding \u003cem\u003eP\u003c/em\u003e values between specific groups of interest. Somatic mutation calls from the TCGA database, processed by Varscan2 [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e], were utilized to investigate genomic alteration among SNF subgroups. The downloaded mutation annotation format (MAF) file comprised 285 CRC patients and 16,981 genes (Hugo Symbol). We also examined the differences on transcriptomics features between pairwise SNF subgroups, and applied the Benjamini\u0026ndash;Hochberg false discovery rate (FDR) for multiple testing correction. DEGs were confirmed according to the negative binomial distribution model implemented in edgeR [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e] (FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05 and log2 Fold Change\u0026thinsp;\u0026gt;\u0026thinsp;1.5).\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eGeneration of simulated colorectal cancer datasets\u003c/h2\u003e \u003cp\u003eSimulated CRC multi-omics data, encompassing DNA methylation, gene expression, and protein expression were generated. Noteworthy, ensuring high quality of simulated data is of paramount importance. Here, we compared the probability density distributions of both the original data and simulated data under the condition of no subgroup (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\delta\\)\u003c/span\u003e\u003c/span\u003e = 0). As depicted in \u003cb\u003eFig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e, the consistency in distributional properties and inherent patterns of single-omics features between the original and simulated data suggests that the simulated data effectively substitutes for the original data in our study. To further demonstrate the influence of the critical parameter \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\delta\\)\u003c/span\u003e\u003c/span\u003e on the simulated data, we produced simulated datasets with varying shifts among subgroups (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\delta\\)\u003c/span\u003e\u003c/span\u003e = 0, 2, and 2.5) while maintaining \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(k\\)\u003c/span\u003e\u003c/span\u003e = 3. As exhibited in \u003cb\u003eFig. \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e\u003c/b\u003e and \u003cb\u003eFig. S3\u003c/b\u003e, the heatmaps and principal component analysis (PCA) plots for each omics dataset highlight that the subgroup separation becomes more pronounced as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\delta\\)\u003c/span\u003e\u003c/span\u003e increases.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eEnsemble multi-omics integrative methods on simulated datasets\u003c/h2\u003e \u003cp\u003eEach multi-omics integrative clustering method determines the optimal number of clusters according to own criteria. Our study evaluated the consistency between the optimal clustering number of methods and the number of ground-truth subgroups (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(k\\)\u003c/span\u003e\u003c/span\u003e = 3) based on the basic simulated scenario. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB illustrates that the range of optimal clustering numbers for these methods, varying from 2 to 6. Except for iClusterPlus, all methods had a median optimal clustering number that matched the ground-truth value (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(k\\)\u003c/span\u003e\u003c/span\u003e = 3). iClusterPlus exhibited the least favorable performance, with a range from 3 to 5 and a median of 5. LRAcluster and CIMLR consistently outperformed other methods in each simulation, yielding an optimal clustering number of 3. SNF and intNMF showed a consistent distribution of optimal clustering numbers, ranging from 2 to 3. In contrast, PINSPlus, MCIA, and MoCluster displayed a relatively dispersed distribution, varying from 2 to 6.\u003c/p\u003e \u003cp\u003eThe evaluation results revealed that the optimal clustering number varies across different methods. To minimize the effect of varying optimal clustering numbers on subtypes classification accuracy, we adjusted the clustering numbers of each method to match the number of ground-truth subgroups during the subsequent simulated evaluation. Afterwards, we investigated the impact of the various simulation parameters on classification accuracy and evaluated the consistency between integrative clusters and ground-truth subgroups using the ARI, NMI, and F1-score.\u003c/p\u003e \u003cp\u003eTo investigate the impact of sample size on the classification accuracy of these methods, we generated simulated data by adjusting the sample size (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e) within the basic simulated scenario. From Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA, we observed that CIMLR exhibits lower classification accuracy for small sample sizes (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(n\\)\u003c/span\u003e\u003c/span\u003e \u0026lt; 100). As the sample size increases, the classification accuracy of CIMLR substantially improves and eventually reaches stability. In contrast, the classification accuracy of other methods displayed a slight fluctuation, typically within a 0.1 range. For LRAcluster, SNF, MoCluster, and MCIA, the ARI, NMI, and F1-scores consistently exceeded 0.95 and remained stable across various sample sizes. Conversely, iClusterPlus displayed a decreasing trend in classification accuracy as the sample size increased, although all three evaluated indicators were still close to 0.95. Notably, intNMF and PINSPlus underperformed across different sample sizes, achieving ARI and NMI scores below 0.7, with an F1-score near 0.75. Detailed simulation results are available in \u003cb\u003eTable. S3\u003c/b\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe next performed multi-omics integration analysis by varying the number of subgroups (ranging from 2 to 6) under basic simulated scenario. It became evident that most methods demonstrated preferable classification accuracy when dealing with 2 or 3 subgroups. As the number of subgroups increased, the classification accuracy of the methods generally decreased (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB). LRAcluster and SNF exhibited superior classification accuracy compared to other methods, with all three evaluation indicators approaching 1 and displaying stability across various simulated scenarios. iClusterPlus showed slight fluctuations in ARI and NMI, hovering around 0.95, and achieved an F1-score exceeding 0.95. While the classification accuracy of MCIA and CIMLR notably declined. Specifically, MoCluster was greatly influenced by the number of subgroups, resulting in a decrease of 0.2 to 0.3 in classification accuracy. Detailed simulation results can be found in \u003cb\u003eTable. S4\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eEach Cluster and associated features were created by applying a fixed shift to their mean values. We varied shift among subgroups through a range of values (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e) under basic simulated scenario and observed as the shift increased, all methods exhibited enhanced classification accuracy. SNF performed exceptionally well overall but struggled to accurately identify subgroups only when a weak shift was present among subgroups (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\delta\\)\u003c/span\u003e\u003c/span\u003e \u0026lt; 1). Mocluster, LRAcluster, MCIA, and iClusterPlus demonstrated similar patterns, with substantial increases in classification accuracy before \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\delta\\)\u003c/span\u003e\u003c/span\u003e = 1.5, followed by stability. In contrast, CIMLR exhibited substantial improvements until \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\delta\\)\u003c/span\u003e\u003c/span\u003e = 2. On the other hand, PINSPlus and intNMF achieved ARI and NMI less than 0.7, with an F1-score approaching 0.75 when \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\delta\\)\u003c/span\u003e\u003c/span\u003e = 2 (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC). Further details can be found in \u003cb\u003eTable. S5\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eWe determined differentially expressed variables by adjusting the probability in the binomial distribution, and traversed \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\\)\u003c/span\u003e\u003c/span\u003e across a range of values (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e) under basic simulated scenario. Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD illustrates that as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\\)\u003c/span\u003e\u003c/span\u003e increased, there was a noticeable enhancement in classification accuracy for all methods. Most methods experienced substantial improvements before \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\\)\u003c/span\u003e\u003c/span\u003e = 10% and plateaued thereafter. Notably, SNF, Mocluster, LRAcluster, and MCIA displayed poor performance in the extreme scenario of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\\)\u003c/span\u003e\u003c/span\u003e= 1%, while in other scenarios, their performance was close to 1 and remained stable. In contrast, CIMLR and iClusterPlus achieved satisfactory classification accuracy at \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\\)\u003c/span\u003e\u003c/span\u003e = 10%. PINSPlus and intNMF required a higher proportion of differentially expressed variables (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(p\\)\u003c/span\u003e\u003c/span\u003e = 25%) to achieve comparable performance levels. Detailed simulation results can be found in \u003cb\u003eTable. S6\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eSubsequently, we adjusted the standard deviation of irrelevant variables to generate the noise level \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma\\)\u003c/span\u003e\u003c/span\u003e, and varied \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma\\)\u003c/span\u003e\u003c/span\u003e by times of 0.5, 1, 2, and 3 compared to the original standard deviation under basic simulated scenario. Overall, as noise increased, most methods exhibited a decrease in classification accuracy, particularly when \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma\\)\u003c/span\u003e\u003c/span\u003e reached 2 to 3 (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eE). SNF and iClusterPlus consistently maintained high classification accuracy. CIMLR, Mocluster, MCIA, and LRAcluster displayed a noticeable decline in classification accuracy after \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma\\)\u003c/span\u003e\u003c/span\u003e reached 2, with LRAcluster experiencing the most pronounced drop. For, intNMF and PINSPlus, the decline began when \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\sigma\\)\u003c/span\u003e\u003c/span\u003e reached 1. Comprehensive simulation results can be found in \u003cb\u003eTable. S7\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eHere, we explored the influence of varying ratios of sample size among subgroups (balanced, moderately imbalanced, and extremely imbalanced) on these methods under basic simulated scenario. The balanced, moderately imbalanced, and extremely imbalanced ratios of the three subgroups were \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(c\\)\u003c/span\u003e\u003c/span\u003e = (1/3, 1/3, 1/3), \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(c\\)\u003c/span\u003e\u003c/span\u003e = (0.2, 0.3, 0.5), and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(c\\)\u003c/span\u003e\u003c/span\u003e = (0.1, 0.3, 0.6) respectively. In the scenario of four subgroups, the balanced, moderately imbalanced, and extremely imbalanced ratios were \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(c\\)\u003c/span\u003e\u003c/span\u003e = (0.25, 0.25, 0.25, 0.25), \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(c\\)\u003c/span\u003e\u003c/span\u003e = (0.1, 0.2, 0.3, 0.4), and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(c\\)\u003c/span\u003e\u003c/span\u003e = (0.1, 0.1, 0.2, 0.6) respectively. We observed that LRAcluster, SNF, and MCIA consistently achieved classification accuracy exceeding 0.95 across various scenarios of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(k\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(c\\)\u003c/span\u003e\u003c/span\u003e. For \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(k\\)\u003c/span\u003e\u003c/span\u003e = 3, iClusterPlus, CIMLR, and Mocluster displayed satisfactory classification accuracy. However, in cases of moderately and extremely imbalanced with \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(k\\)\u003c/span\u003e\u003c/span\u003e = 4, these three algorithms exhibited a downward trend to varying degrees. Similarly, intNMF and PINSPlus delivered comparatively less satisfactory performance (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eF). Detailed simulation results are presented in \u003cb\u003eTable. S8\u003c/b\u003e.\u003c/p\u003e \u003cp\u003eTo explore the effectiveness of various omics combinations, we simulated four combinations of multi-omics data scenarios: M\u0026thinsp;+\u0026thinsp;G\u0026thinsp;+\u0026thinsp;P, M\u0026thinsp;+\u0026thinsp;G, G\u0026thinsp;+\u0026thinsp;P, and M\u0026thinsp;+\u0026thinsp;P (M, G, and P represent DNA methylation, gene expression, and protein expression, respectively) under the basic simulated scenario. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eG, there is no combination was universally effective for all eight methods. Generally, the combination involving three omics data demonstrated superior classification accuracy compared to other combinations, with except for MCIA and PINSPlus. Notably, the combination of M\u0026thinsp;+\u0026thinsp;G showcased outstanding classification accuracy for these two algorithms. SNF and CIMLR consistently delivered stable performance across all combinations. The ARI, NMI, and F1-score of SNF approached 1 in all combinations, and the medians of the three evaluation indicators for CIMLR exceeded 0.85. Particularly, the classification accuracy of LRAcluster and iClusterPlus decreased largely in the M\u0026thinsp;+\u0026thinsp;P and G\u0026thinsp;+\u0026thinsp;P combinations, respectively.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eApplication in realistic colorectal cancer datasets of multi-omics integrative methods\u003c/h2\u003e \u003cp\u003eIn our practical study, we included a cohort of 331 CRC patients with epigenomics, transcriptomics, proteomics, and clinical data from the TCGA database (\u003cb\u003eFig. S4\u003c/b\u003e). Comprehensive basic information regarding the enrolled CRC patients was presented in \u003cb\u003eTable. S9\u003c/b\u003e. The initial number of features for epigenomics, transcriptomics, and proteomics were 22,601, 55,026, and 198, respectively. Subsequently, we selected the top 1000 features with the highest MAD values in epigenomics and transcriptomics data and the leading 100 features in proteomics for further analysis.\u003c/p\u003e \u003cp\u003eConsistent with our findings in simulated experiments, the analysis of real CRC multi-omics data demonstrated that the classification accuracy of SNF consistently surpassed methodologies when employing the optimal number of subgroups, yielding ARI of 0.227, NMI of 0.254, and F1-score of 0.576 (\u003cb\u003eFig. S5\u003c/b\u003e). Figure\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA illustrated the classification performance of SNF with clustering numbers ranging from 2 to 6 on the real dataset. SNF effectively segregated CRC patients into 5 distinct subgroups, exhibiting the highest classification accuracy: an ARI of 0.326, NMI of 0.337, and F1-score of 0.630. Subsequently, we adopted the 5-category SNF clustering as the primary CRC subgroups for further analysis.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe t-SNE plots for dimensional reduction revealed that a discernible separation trend among SNF subgroups within each omics dataset (\u003cb\u003eFig. S6\u003c/b\u003e). SNF subgroups had distinct overall survival (OS), which SNF-V was associated with poor survival (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB), and SNF I-IV were significantly associated with improved OS compared to SNF-V in Cox proportional hazard regression analysis. CMS subgroups also exhibited significant differences in OS, however, contrary to expectations, CMS1 and CMS2 subgroups were not significantly associated with improved OS compared to CMS4 (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC). In the Cox multivariate analysis, SNF remained significantly associated with improved OS (with age and pathological stage), whereas CMS subtypes did not (\u003cb\u003eTable. S10\u003c/b\u003e). The relationships between SNF subgroups and CMS subgroups were illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eD. Specifically, the majority of SNF-I (84.3%) belonged to the CMS4 subtype. SNF-II predominantly (76.4%) included CMS2 subtype patients, while SNF-III primarily comprised 21 CMS2 patients (29.6%) and 42 CMS3 patients (59.2%). Moreover, SNF-IV and SNF-V mainly consisted of CMS2 (74.2%) and CMS1 (75.0%) subgroups, respectively. In contrast, the SNF subgroups for other clustering numbers did not significantly stratify patients with survival differences (\u003cb\u003eFig. S7\u003c/b\u003e), and their relationships with CMS subgroups were less discernible compared to the 5-category SNF clustering. Additionally, we conducted a Wilcoxon rank sum test to compare acknowledged CRC biomarkers among SNF subgroups (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eE-H). SNF subgroups also showed significant differences in term of TMB, MSI score, \u003cem\u003ePD-L1\u003c/em\u003e expression and \u003cem\u003eTP53\u003c/em\u003e expression. Notably, the SNF-V subgroup displayed the highest levels of TMB, MSI score and \u003cem\u003ePD-L1\u003c/em\u003e expression, providing additional justification for considering immune checkpoint blockade (ICB) as a potential therapeutic intervention.\u003c/p\u003e \u003cp\u003eThe panoramic landscapes of the 15 most frequently somatic altered genes for the SNF subgroups were represented in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eI. Overall, the frequencies of specific mutated genes exhibited notable variations among SNF subgroups. \u003cem\u003eAPC\u003c/em\u003e, \u003cem\u003eTP53\u003c/em\u003e, \u003cem\u003eTTN\u003c/em\u003e, and \u003cem\u003eKRAS\u003c/em\u003e occupied the top four positions in SNF-I, SNF-II, SNF-III, and SNF-IV subgroups. However, in the SNF-V subgroup, the mutated frequencies of \u003cem\u003eAPC\u003c/em\u003e (33%), \u003cem\u003eTP53\u003c/em\u003e (38%), and \u003cem\u003eKRAS\u003c/em\u003e (27%) were lower compared to other subtypes, while \u003cem\u003eTTN\u003c/em\u003e (67%) remained displayed frequently mutated rates. Specifically, the SNF-V patients significantly enriched in somatic mutations of \u003cem\u003eSYNE1\u003c/em\u003e (52%), \u003cem\u003eMUC16\u003c/em\u003e (54%), \u003cem\u003eFAT4\u003c/em\u003e (52%), \u003cem\u003eDNAH5\u003c/em\u003e (31%), \u003cem\u003eOBSCN\u003c/em\u003e (48%), \u003cem\u003eZFHX4\u003c/em\u003e (42%), \u003cem\u003eDNAH11\u003c/em\u003e (38%), \u003cem\u003eCSMD1\u003c/em\u003e (35%), \u003cem\u003eRYR2\u003c/em\u003e (35%), and \u003cem\u003eFLG\u003c/em\u003e (27%) in contrast to other SNF subgroups (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.01, Fisher exact test), serving as potential predictive biomarkers for ICB activity. In addition, we performed differential expression testing to illustrate significant differences in transcriptomics features among SNF subgroups. The upset plot visualized the number of differentially expressed genes between any pairwise SNF subgroups (\u003cb\u003eFig. S8\u003c/b\u003e). We identified 323 differentially expressed genes in total, with the most substantial difference (37 genes) observed between the SNF-II and SNF-V subgroups. Moreover, 24 genes displayed significant expression difference in at least half of the 10 pairwise SNF subgroup comparisons, while 3 genes (\u003cem\u003eCTSE\u003c/em\u003e, \u003cem\u003eREG1A\u003c/em\u003e, \u003cem\u003eREG4\u003c/em\u003e) showed significant differences in 7 pairwise SNF subgroups. These findings suggest that these molecular alterations among SNF subgroups likely contribute to the carcinogenesis and progression of colorectal cancer through diverse mechanisms.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eMulti-omics integrative clustering methods have been developed to enhance the precision and personalization of treatments for CRC patients by identifying disease subtypes. Our study generated CRC multi-omics data with realistic correlations through simulated experiments, and evaluated the optimal number of clusters and classification accuracy of eight representative multi-omics integrative clustering methods.\u003c/p\u003e \u003cp\u003eThe simulated results revealed that CIMLR and MCIA provided the most accurate estimates for the number of clusters, and SNF displayed outstanding classification accuracy and robustness in various simulation scenarios. In summary, the similarity-based methods consider the heterogeneity of CRC multi-omics data and individual similarities by computing sample similarity matrices for each omics dataset, and had been widely used in multi-omics integration analysis. In contrast, other methods have certain limitations in CRC multi-omics integration. For instance, LRAcluster is sensitive to noise, iClusterPlus excels in classification accuracy and robustness but suffers from computational inefficiency. In addition, it is important to clarify that, while PINSPlus has unsatisfactory classification accuracy, which could be attributed to re-analyzing clustering results, it is a user-friendly method for automatically determining the optimal number of clusters without requiring complex parameter adjustments. Therefore, researchers can consider suitable methods for certain purposes or circumstances. Furthermore, the classification accuracy of these methods is sensitive to the number of subgroups, shifts among subgroups, proportions of differentially expressed variables, noise levels, and combinations of omics data. Researchers should take these factors into account when employing multi-omics integrative clustering methods.\u003c/p\u003e \u003cp\u003eIn agreement with our results, Pierre-Jean and colleagues also concluded SNF as the suitable method for individual classification [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Duan and associates recommended SNF to conduct general cancer subtyping tasks [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], our study also found the superiority of SNF of multi-omics integrative analysis in CRC classification accuracy. The possible reason is that SNF not only captures the shared structure from diverse omics data but also extracts the specific structure, enhancing its ability to preserve more information within the sample similarity network. Regarding the determination of the optimal number of clusters, we affirmed that LRAcluster performs excellently with highly correlated multi-omics data. In contrast, our observations deviate from the findings of Chauvel and colleagues, who reported that MoCluster appropriately identifies the number of clusters [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e], we noted that the optimal clustering numbers determined by MoCluster under highly correlated multi-omics data quite differ from the number of ground-truth subgroups. One potential factor of interference might be that the inter-omics correlations mask MoCluster's capacity to discern subgroups.\u003c/p\u003e \u003cp\u003eIn the comprehensive analysis of realistic CRC multi-omics data, consistent with our simulation findings, SNF exhibited superior classification accuracy compared to other methods. SNF categorized CRC patients into five multi-omics molecular subgroups, demonstrating a distinct correspondence between these subgroups and CMS subtypes. Furthermore, there were various types of disparities among the SNF subgroups. SNF-V was associated with significant poor survival. Regarding somatic mutations, the SNF-V subgroup particularly exhibited notable variations in the frequencies of mutated genes. For instance, \u003cem\u003eSYNE1\u003c/em\u003e, \u003cem\u003eMUC16\u003c/em\u003e, and \u003cem\u003eFAT4\u003c/em\u003e had significantly higher mutation rates in SNF-V compared to other subgroups. These genes have previously been linked to the prognosis of CRC patients [\u003cspan additionalcitationids=\"CR37\" citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. In the differentially expressed analysis, we identified differentially expressed markers among the SNF subgroups. In addition to well-established cancer-related biomarkers such as \u003cem\u003eCXCL9\u003c/em\u003e, \u003cem\u003eMUC2\u003c/em\u003e, and \u003cem\u003eMUC4\u003c/em\u003e. Furthermore, \u003cem\u003eCTSE\u003c/em\u003e, \u003cem\u003eREG1A\u003c/em\u003e, and \u003cem\u003eREG4\u003c/em\u003e exhibited differential expression between numerous pairwise SNF subgroups, may play roles in the progression of colorectal cancer through various mechanisms. Certainly, prior studies conducted by Astrosini et al. revealed REG1A as a molecular marker with prognostic value, associated specifically with peritoneal carcinomatosis in CRC [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. Additionally, Rafa et al. highlighted REG4 as a multifunctional secreted protein that exerts effects on CRC cells through autocrine and paracrine mechanisms, potentially holding significance in the development and progression of CRC [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e].Collectively, the disparities observed among SNF subtypes offer a novel multi-omics perspective on CRC heterogeneity with potential clinical treatment.\u003c/p\u003e \u003cp\u003eCertain limitations in our work must be acknowledged. Firstly, our research exclusively focused on the epigenome-transcriptome-proteome due to their evident regulatory interplay. Additionally, in our simulated experiment, the number of CRC multi-omics features remained constant, precluding an investigation into the impact of varying feature numbers on classification accuracy, and during the assessment of classification accuracy, we fixed the number of clusters in the multi-omics integrative clustering analysis to match the actual number of ground-truth subgroups, without considering the optimal cluster numbers of these methods. Furthermore, we configured the model parameters of multi-omics integrative clustering methods based on the recommendations of each method, without accounting for the potential influence of these parameters. Lastly, due to limited access to expansive multi-omics CRC datasets, our practical analysis was solely based on TCGA, and the effectiveness of the SNF subgroups needs to be verified in subsequent studies.\u003c/p\u003e \u003cp\u003eAs the range of multi-omics data categories expands, there has been a notable advancement of multi-omics integration methods, especially those based on deep learning techniques. Future benchmarking frameworks for integration should encompass diverse multi-omics data types, comprising discrete data and single-cell data. Additionally, given that numerous methods can analyze multiple omics from the same samples, there is potential to develop methods capable of integrating incomplete data.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eAccurately classifying CRC patients into molecular subtypes is essential for understanding tumor heterogeneity. The progression in computer science and omics technologies has facilitated the development of multi-omics integrative clustering algorithms. Therefore, the application of unsupervised multi-omics integration clustering methods in CRC necessitates the establishment of a comprehensive benchmark with practical guidance. Through simulation studies and case analyses, we demonstrate that the superior classification accuracy, robustness, and efficacy of SNF render it an exemplary method for identifying subtypes indicative of tumor heterogeneity that can be further extended to clinical practice.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCRC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eColorectal Cancer\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSNF\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eSimilarity Network Fusion\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTNM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eTumor-Node-Metastasis\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eDMSs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eDifferentially Methylated Sites\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eDEGs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eDifferentially Expressed Genes\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eDEPs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eDifferentially Expressed Proteins\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eARI\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eAdjusted Rand Index\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eNMI\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eNormalized Mutual Information\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMAD\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMedian Absolute Deviation\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCMS\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eConsensus Molecular Subtypes\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMAF\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMutation Annotation Format\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eFDR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eFalse Discovery Rate\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003ePCA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ePrincipal Component Analysis\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eOS\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eOverall Survival\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eICB\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eImmune Checkpoint Blockade\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe simulated data are generated using our R function and can be found in Github repository https://github.com/zsbvb/Comparison-of-Multiomics-Integration-Methods-for-CRC. The epigenomics, proteomics, and clinical data for TCGA colorectal cancer are available at https://www.cbioportal.org. The transcriptomics and MAF file are downloaded from https://portal.gdc.cancer.gov.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors do not have any conflict of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was supported by National Natural Science Foundation of China (82222064,\u0026nbsp;81973147 to TZ, 82304247 to CW), the Young Scholars Program of\u0026nbsp;Shandong University (21320082164070 to CW), National Key Research and Development Program\u0026nbsp;(2022YFC2010100 to TZ), and Shandong University Distinguished Young Scholars (to TZ).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSZ, JL, CW and TZ generated the hypothesis, compiled the simulated datasets and performed the evaluation. SZ wrote the original manuscript. FZ and BG contributed in result interpretation. BF and CL contributed to analytic strategy, supervised the study. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe would like to thank reviewers for their feedback on our manuscript. We would also like to thank the authors of all multi-omics integration methods used in this study for making their code and data freely available. Last, we thank the joint effort of many investigators and staff members in TCGA, and most importantly, the colorectal patients who have participated in this work.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eSung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin. 2021;71(3):209\u0026ndash;49.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCervantes A, Adam R, Rosello S, Arnold D, Normanno N, Taieb J, Seligmann J, De Baere T, Osterlund P, Yoshino T, et al. Metastatic colorectal cancer: ESMO Clinical Practice Guideline for diagnosis, treatment and follow-up. Ann Oncol. 2023;34(1):10\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDekker E, Tanis PJ, Vleugels JLA, Kasi PM, Wallace MB. Colorectal cancer. Lancet. 2019;394(10207):1467\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNagtegaal ID, Quirke P, Schmoll HJ. Has the new TNM classification for colorectal cancer improved care? Nat Rev Clin Oncol. 2011;9(2):119\u0026ndash;23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNitsche U, Maak M, Schuster T, Kunzli B, Langer R, Slotta-Huspenina J, Janssen KP, Friess H, Rosenberg R. Prediction of prognosis is not improved by the seventh and latest edition of the TNM classification for colorectal cancer in a single-center collective. Ann Surg. 2011;254(5):793\u0026ndash;800. discussion 800\u0026thinsp;\u0026ndash;\u0026thinsp;791.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRuan W, Yuan X, Eltzschig HK. Circadian rhythm as a therapeutic target. Nat Rev Drug Discovery. 2021;20(4):287\u0026ndash;307.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu FB. Metabolic profiling of diabetes: from black-box epidemiology to systems epidemiology. Clin Chem. 2011;57(9):1224\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKarczewski KJ, Snyder MP. Integrative omics for health and disease. Nat Rev Genet. 2018;19(5):299\u0026ndash;310.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDumbill E. A Revolution That Will Transform How We Live, Work, and Think: An Interview with the Authors of Big Data. Big Data. 2013;1(2):73\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKristensen VN, Lingjaerde OC, Russnes HG, Vollan HK, Frigessi A, Borresen-Dale AL. Principles and methods of integrative genomic analyses in cancer. Nat Rev Cancer. 2014;14(5):299\u0026ndash;313.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGunther OP, Chen V, Freue GC, Balshaw RF, Tebbutt SJ, Hollander Z, Takhar M, McMaster WR, McManus BM, Keown PA, et al. A computational pipeline for the development of multi-marker bio-signature panels and ensemble classifiers. BMC Bioinformatics. 2012;13:326.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSingh A, Shannon CP, Gautier B, Rohart F, Vacher M, Tebbutt SJ, Le Cao KA. DIABLO: an integrative approach for identifying key molecular drivers from multi-omics assays. Bioinf (Oxford England). 2019;35(17):3055\u0026ndash;62.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan de Wiel MA, Lien TG, Verlaat W, van Wieringen WN, Wilting SM. Better prediction by use of co-data: adaptive group-regularized ridge regression. Stat Med. 2016;35(3):368\u0026ndash;81.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang T, Shao W, Huang Z, Tang H, Zhang J, Ding Z, Huang K. MOGONET integrates multi-omics data using graph convolutional networks allowing patient classification and biomarker identification. Nat Commun. 2021;12(1):3445.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRoelands J, Kuppen PJK, Ahmed EI, Mall R, Masoodi T, Singh P, Monaco G, Raynaud C, de Miranda N, Ferraro L, et al. An integrated tumor, immune and microbiome atlas of colon cancer. Nat Med. 2023;29(5):1273\u0026ndash;86.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRappoport N, Shamir R. Multi-omic and multi-view clustering algorithms: review and cancer benchmark. Nucleic Acids Res. 2018;46(20):10546\u0026ndash;62.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZitnik M, Nguyen F, Wang B, Leskovec J, Goldenberg A, Hoffman MM. Machine Learning for Integrating Data in Biology and Medicine: Principles, Practice, and Opportunities. Inf Fusion. 2019;50:71\u0026ndash;91.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePierre-Jean M, Deleuze JF, Le Floch E, Mauger F. Clustering and variable selection evaluation of 13 unsupervised methods for multi-omics data integration. Brief Bioinform. 2020;21(6):2011\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDuan R, Gao L, Gao Y, Hu Y, Xu H, Huang M, Song K, Wang H, Dong Y, Jiang C, et al. Evaluation and comparison of multi-omics data integration methods for cancer subtyping. PLoS Comput Biol. 2021;17(8):e1009224.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTini G, Marchetti L, Priami C, Scott-Boyer MP. Multi-omics integration-a comparison of unsupervised clustering methodologies. Brief Bioinform. 2019;20(4):1269\u0026ndash;79.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCrick F. Central dogma of molecular biology. Nature. 1970;227(5258):561\u0026ndash;3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGry M, Rimini R, Stromberg S, Asplund A, Ponten F, Uhlen M, Nilsson P. Correlations between RNA and protein expression profiles in 23 human cell lines. BMC Genomics. 2009;10:365.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu D, Wang D, Zhang MQ, Gu J. Fast dimension reduction and integrative clustering of multi-omics data using low-rank approximation: application to cancer molecular classification. BMC Genomics. 2015;16:1022.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang B, Mezlini AM, Demir F, Fiume M, Tu Z, Brudno M, Haibe-Kains B, Goldenberg A. Similarity network fusion for aggregating data types on a genomic scale. Nat Methods. 2014;11(3):333\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRamazzotti D, Lal A, Wang B, Batzoglou S, Sidow A. Multi-omic tumor data reveal diversity of molecular mechanisms that correlate with survival. Nat Commun. 2018;9(1):4453.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeng C, Helm D, Frejno M, Kuster B. moCluster: Identifying Joint Patterns Across Multiple Omics Data Sets. J Proteome Res. 2016;15(3):755\u0026ndash;65.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChalise P, Fridley BL. Integrative clustering of multi-level 'omic data based on non-negative matrix factorization algorithm. PLoS ONE. 2017;12(5):e0176278.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeng C, Kuster B, Culhane AC, Gholami AM. A multivariate approach to the integration of multi-omics datasets. BMC Bioinformatics. 2014;15:162.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMo Q, Wang S, Seshan VE, Olshen AB, Schultz N, Sander C, Powers RS, Ladanyi M, Shen R. Pattern discovery and cancer gene identification in integrated cancer genomic data. Proc Natl Acad Sci U S A. 2013;110(11):4245\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNguyen T, Tagett R, Diaz D, Draghici S. A novel approach for data integration and disease subtyping. Genome Res. 2017;27(12):2025\u0026ndash;39.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChalise P, Raghavan R, Fridley BL. InterSIM: Simulation tool for multiple integrative 'omic datasets'. Comput Methods Programs Biomed. 2016;128:69\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuinney J, Dienstmann R, Wang X, de Reynies A, Schlicker A, Soneson C, Marisa L, Roepman P, Nyamundanda G, Angelino P, et al. The consensus molecular subtypes of colorectal cancer. Nat Med. 2015;21(11):1350\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKoboldt DC, Zhang Q, Larson DE, Shen D, McLellan MD, Lin L, Miller CA, Mardis ER, Ding L, Wilson RK. VarScan 2: somatic mutation and copy number alteration discovery in cancer by exome sequencing. Genome Res. 2012;22(3):568\u0026ndash;76.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRobinson MD, McCarthy DJ, Smyth GK. edgeR: a Bioconductor package for differential expression analysis of digital gene expression data. Bioinf (Oxford England). 2010;26(1):139\u0026ndash;40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChauvel C, Novoloaca A, Veyre P, Reynier F, Becker J. Evaluation of integrative clustering methods for the analysis of multi-omics data. Brief Bioinform. 2020;21(2):541\u0026ndash;52.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGe W, Hu H, Cai W, Xu J, Hu W, Weng X, Qin X, Huang Y, Han W, Hu Y, et al. High-risk Stage III colon cancer patients identified by a novel five-gene mutational signature are characterized by upregulation of IL-23A and gut bacterial translocation of the tumor microenvironment. Int J Cancer. 2020;146(7):2027\u0026ndash;35.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi C, Xu J, Wang X, Zhang C, Yu Z, Liu J, Tai Z, Luo Z, Yi X, Zhong Z. Whole exome and transcriptome sequencing reveal clonal evolution and exhibit immune-related features in metastatic colorectal tumors. Cell Death Discov. 2021;7(1):222.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu J, Wu WK, Li X, He J, Li XX, Ng SS, Yu C, Gao Z, Yang J, Li M, et al. Novel recurrently mutated genes and a prognostic mutation signature in colorectal cancer. Gut. 2015;64(4):636\u0026ndash;45.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAstrosini C, Roeefzaad C, Dai YY, Dieckgraefe BK, Jons T, Kemmner W. REG1A expression is a prognostic marker in colorectal cancer and associated with peritoneal carcinomatosis. Int J Cancer. 2008;123(2):409\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRafa L, Dessein AF, Devisme L, Buob D, Truant S, Porchet N, Huet G, Buisine MP, Lesuffleur T. REG4 acts as a mitogenic, motility and pro-invasive factor for colon cancer cells. Int J Oncol. 2010;36(3):689\u0026ndash;98.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Multi-omics, Colorectal cancer, Data integration, Subtype identification","lastPublishedDoi":"10.21203/rs.3.rs-4106569/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4106569/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground and objectives\u003c/h2\u003e \u003cp\u003eColorectal cancer (CRC) represents a heterogeneous malignancy that has concerned global burden of incidence and mortality. The traditional tumor-node-metastasis staging system has exhibited certain limitations. With the advancement of omics technologies, researchers are directing their focus on developing a more precise multi-omics molecular classification. Therefore, the utilization of unsupervised multi-omics integrative clustering methods in CRC, advocating for the establishment of a comprehensive benchmark with practical guidelines. In this study, we obtained CRC multi-omics data, encompassing DNA methylation, gene expression, and protein expression from the TCGA database. We then generated interrelated CRC multi-omics data with various structures based on realistic multi-omics correlations, and performed a comprehensive evaluation of eight representative methods categorized as early integration, intermediate integration, and late integration using complementary benchmarks for subtype classification accuracy. Lastly, we employed these methods to integrate real-world CRC multi-omics data, survival and differential analysis were used to highlight differences among newly identified multi-omics subtypes.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eThrough in-depth comparisons, we observed that similarity network fusion (SNF) exhibited exceptional performance in integrating multi-omics data derived from simulations. Additionally, SNF effectively distinguished CRC patients into five subgroups with the highest classification accuracy. Moreover, we found significant survival differences and molecular distinctions among SNF subtypes.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eThe findings consistently demonstrate that SNF outperforms other methods in CRC multi-omics integrative clustering. The significant survival differences and molecular distinctions among SNF subtypes provide novel insights into the multi-omics perspective on CRC heterogeneity with potential clinical treatment. The code and its implementation are available in GitHub \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/zsbvb/Comparison-of-Multiomics-Integration-Methods-for-CRC\u003c/span\u003e\u003cspan address=\"https://github.com/zsbvb/Comparison-of-Multiomics-Integration-Methods-for-CRC\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e","manuscriptTitle":"Benchmarking multi-omics integrative clustering methods for subtype identification in colorectal cancer","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-03-21 19:27:39","doi":"10.21203/rs.3.rs-4106569/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ed6199c7-71a5-4d43-bb57-e6e4f67c779f","owner":[],"postedDate":"March 21st, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-03-21T19:27:41+00:00","versionOfRecord":[],"versionCreatedAt":"2024-03-21 19:27:39","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4106569","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4106569","identity":"rs-4106569","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-28T02:00:01.590549+00:00
License: CC-BY-4.0