Consensus molecular subtyping in colon cancer: from the design of in-house assay to prognostic nomogram | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Consensus molecular subtyping in colon cancer: from the design of in-house assay to prognostic nomogram Mariapia Caputo, Leonarda Maurmo, Debora Traversa, Tommaso Maria Marvulli, and 10 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9178611/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 9 You are reading this latest preprint version Abstract Background. Consensus Molecular Subtypes (CMS) provide a biologically robust classification of colorectal cancer (CRC) with prognostic relevance; however, their clinical implementation remains limited by reliance on genome-wide transcriptomics and platform-dependent classifiers. We aimed to develop and validate a compact, FFPE-compatible in-house CMS assay and to integrate CMS information into a clinically applicable prognostic nomogram. Methods. An in silico training cohort was generated from TCGA-COAD and TCGA-READ datasets, including transcriptomic profiles, somatic mutations, and microsatellite instability (MSI) status. Feature selection was performed using a minimum Redundancy Maximum Relevance (mRMR) wrapper approach combined with Random Forest and Support Vector Machine classifiers to identify a minimal CMS-informative feature set. Based on selected features, a custom amplicon-based DNA/RNA NGS assay was designed. The assay was validated in an independent cohort of 100 CRC patients with bulk RNAseq–based CMS classification. Classification performance was assessed against CMSclassifier and CMScaller. A multivariable Cox proportional hazards model incorporating CMS, clinicopathological variables, MSI, and driver mutations was used to develop a prognostic nomogram for 3- and 5-year overall survival. Results. The optimal model achieved a global accuracy of 0.76 using 70 features, comprising 7 coding genes, 32 lncRNAs, 30 genomic alterations, and MSI status. The in-house CMS assay demonstrated high balanced accuracy for CMS1 and CMS2, with lower performance for CMS3 and CMS4. Integration of CMSassay classification into the prognostic model showed that CMS classification and tumor stage contributed the largest effects on survival risk. The resulting nomogram provided individualized 3- and 5-year overall survival estimates with clear risk stratification. Conclusions. We developed a clinically feasible, FFPE-compatible CMS assay that preserves the prognostic relevance of transcriptome-based CMS classification while reducing technical and computational complexity. When integrated into a multivariable nomogram, CMS provides substantial prognostic value beyond standard molecular and clinicopathological factors. This approach facilitates the translation of CMS into routine clinical workflows and supports its use as a quantitative prognostic tool in colon cancer. CMS colon cancer in house panel transcriptomics nomogram Figures Figure 1 Figure 2 Figure 3 Figure 4 Background Colorectal cancer (CRC) represents the third most commonly diagnosed cancer and the second leading cause of cancer-related deaths [ 1 , 2 ], according to the global statistics. CRC is a highly heterogeneous disease, characterized by complex interactions between genetic, epigenetic and environmental factors. This heterogeneity highlights significant challenges for diagnosis, prognosis and treatment [ 3 ]. One of the innovative solutions to address these unmet clinical needs, consists in the development of the molecular classification systems to stratify CRC into biologically different subtypes with clinical relevance. Among these, the Consensus Molecular Subtypes (CMS) classification has emerged as the most robust and widely accepted framework. The CMS classification, developed by Guinney et al. [ 4 ] in 2015, categorizes CRC into four distinct subtypes based on transcriptomic data: CMS1 (Microsatellite Instability Immune, MSI), CMS2 (Canonical), CMS3 (Metabolic), and CMS4 (Mesenchymal) [ 4 ]. Each CMS subtype is associated with unique molecular, histological, and clinical characteristics. CMS1 is characterized by MSI [ 5 ], strong immune infiltration, and a favorable prognosis but poor response to conventional chemotherapy. CMS2, the most prevalent subtype, exhibits chromosomal instability, epithelial differentiation, and activation of the WNT and MYC signaling pathways [ 6 ], displaying a better response to chemotherapy and improved survival. CMS3 shows dysregulated metabolic pathways, particularly in glucose metabolism, and tends to occur in tumors with KRAS mutations [ 7 ]. CMS4, associated with mesenchymal activation, stromal invasion, and angiogenesis, carries the worst prognosis due to its aggressive nature and resistance to standard treatments [ 6 , 8 , 9 ]. The CMS framework has provided valuable insights into the biological bases of CRC and has facilitated the development of precision medicine strategies. By identifying molecular subtypes, clinicians and researchers can better understand tumor behavior, predict treatment responses, and design targeted therapies. Several challenges still remain in the clinical implementation of CMS classification, including the integration of transcriptomic profiling into routine practice and the need for subtype-specific therapeutic options. In these years, various algorithms have been developed to improve the CMS classification in clinical and research settings, including the CMScaller [ 10 ], CMSclassifier [ 11 ], frameworks. Each of these algorithms offers unique strengths and limitations. The CMScaller is a computational tool that uses gene expression profiles to assign CRC tumors to one of the four CMS categories [ 10 ]. It relies on a predefined set of gene expression signatures and applies statistical methods to ensure accurate subtype assignment. The CMScaller is known for its high reproducibility and robustness, making it a preferred choice in transcriptomic studies. However, its reliance on gene expression data necessitates high-quality RNA samples, which can be a limitation in certain clinical scenarios. The CMSclassifier , similar to the CMScaller , employs a supervised machine learning approach to classify CRC tumors [ 11 ]. It incorporates additional biological parameters, such as mutation status and pathway activation scores, to enhance the accuracy of subtype assignments. This algorithm is particularly advantageous in settings where multi-omic data are available, as it provides a more integrative understanding of tumor biology. However, the CMSclassifier requires more computational resources and data integration expertise, potentially limiting its accessibility for routine clinical use. While the CMScaller and CMSclassifier share a common goal of implementing the CMS framework, their methodological differences highlight the importance of algorithm selection based on the research or clinical context. Together, these tools could provide a multifaceted approach to understanding CRC heterogeneity and advancing precision oncology. Materials and Methods TCGA-COAD and TCGA-READ Data retrieval RNA expression data and corresponding clinical information from patients with colorectal cancer were downloaded from the Genomic Data Commons (GDC) Data Portal [ 12 ]. The datasets were obtained from The Cancer Genome Atlas (TCGA) project, specifically from the TCGA-COAD (colon adenocarcinoma) and TCGA-READ (rectum adenocarcinoma) cohorts. TCGAbiolinks R package [ 13 ] has been used to retrieve clinical data, HT-Seq count transcriptome data and miRNA Expression Quantification data. Ensemble ID has been converted into HUGO gene symbols by BiomaRt R package [ 14 ]. LncRNA expression data has been extrapolated by LNCpedia v5.2 conversion table [ 15 ]. TCGAbiolinks R package has been also used to retrieve alterations data, thus obtaining MAF-files. MSI status has been obtained by GDCquery (data.category=’Other’ and clinical.info=’msi’) function. CMSclassifier was used for the classification of samples in CMSs. It was implemented by Guinney et al.[ 4 ],and it is based on the 273 genes included in the original random forest (RF) CMS classifier. Pre-processing: sample filtering and data adjusting For the analysis of count data from RNA-seq for the detection of differentially expressed genes, has been used DESeq2 R package [ 16 , 17 ]. The count data, previously downloaded, are presented as a table (the rows correspond to the genes and the columns to the patients) which reports, for each sample, the number of sequence fragments that have been assigned to each gene. Firstly, a Summarized Experiment (SE) has been obtained by GDCprepare (summarizedExperiment=TRUE) function. This SE was used to create a DESeqDataSet class. While it is not necessary to pre-filter low count genes before running the DESeq2 functions, there are two reasons which make pre-filtering useful: by removing rows in which there are very few reads, the memory size of DESeqDataSet object is reduced, and the speed of the transformation and testing functions within DESeq2 is increased. In our study, has been performed a minimal pre-filtering to keep only rows that have at least 10 reads total. After that, it has been used the varianceStabilizingTranformation function which removes the dependence of the variance on the mean and then transforms the count data, obtaining a matrix of values which have constant variance along the range of mean values[ 16 ]. To reduce the number of miRNAs and lncRNAs, before data normalization, further filtering was applied: all rows with reads count < 10 in more than 10% of patients were eliminated. For data mutations we selected the clinically relevant mutations by OncoKB , a precision oncology knowledge base [ 18 ]. After that, have been selected only 'Nonsense' and 'Missense' mutations grouped by genes and has been created a matrix which reports, for each sample, the presence or absence of the mutation for each gene. Feature Selection and classification Wrapper features selection method To reduce the computational burden of classification and obtain a compact feature set that could be more easily applied in clinical practice, a feature selection step was performed. We adopted a wrapper-based feature selection approach, which evaluates subsets of features using a machine learning algorithm combined with a search strategy that explores the space of possible feature combinations. Each subset is assessed according to the predictive performance achieved by the selected algorithm. In general, wrapper methods follow these steps: (1) a subset of features is selected from the full set using a search strategy; (2) a machine learning model is trained on the selected subset; (3) model performance is evaluated using a predefined metric; and (4) the procedure is repeated iteratively across different feature subsets and corresponding trained models. In our study, we used minimum Redundancy Maximum Relevance (mRMR) [ 19 ] for feature selection together with two classifiers, Support-Vector-Machine (SVM) [ 20 ] and Random Forest (RF) [ 21 ], both of which are well suited for multi-label classification. Our specific aim was to assign patients to the four CMS classes. mRMR Feature Selection Method Feature selection was performed using the mRMR.classic fuction, implemented in the mRMRe R package [ 19 ], which enables efficient identification of relevant and non-redundant features. Let y denote the output variable and X= {x1,.., xn} the set of n input features. The method ranks features in X by maximizing their mutual information (MI) with y (maximum relevance) while minimizing the average MI with features already selected (minimum redundancy). Mutual information between two datasets is estimated on the basis of the correlation between variables Pearson’s or Spearman’s correlation coefficients are used for continuous variables, Cramer’s V for discrete variables, and Somers’ Dxy index for the association between continuous variables and survival data. Support Vector Machine and Random Forest SVMs are supervised learning methods commonly used for classification and regression tasks [ 20 ]. By means of kernel functions, they allow efficient classification of non-linear data. In this framework, input samples are represented as points in an n-dimensional space, where each dimension corresponds to one feature. The algorithm then iteratively identifies a function defining a hyperplane that best separates the regions occupied by the different output classes. In our study, SVM classification was performed using the svm() function from the e1071 R package. The RF algorithm, instead, combines multiple randomized decision trees and aggregates their predictions by majority voting [ 21 ]. This approach has shown strong performance, particularly in settings where the number of variables greatly exceeds the number of observations. Its general workflow includes drawing random samples from the original dataset, building a decision tree for each sample, generating predictions from each tree, and selecting the class receiving the highest number of votes as the final output. In our study, RF classification was carried out using the randomForest function from the randomForest R package, based on Breiman’s original implementation [ 21 ]. In-house CMSassay: DNA/RNA custom panel design and bioinformatic analyses The in-house CMSassay included two panels, both of them amplicon-based and designed through Ion Ampliseq Designer web tool (ThermoFisher). RNA panel required a great effort because it was built through multiple rounds of tiling/pooling in order to increase coverage rate. No SNPs were allowed under primers for the bulk of the design. Quality Control (QC) to avoid cross-binding of primers was performed for each panel. Two lncRNAs were excluded from the final design due technical issues in the identification of primers able to correctly address them. Alignment and variant calling was performed through Torrent Server plugins, Torrent Variant Caller in somatic configuration. From Ion Torrent Server we retrieved VCF files for each sample, which were annotated by OncoKB™ annotator [ 22 ]. RNA panel was analyzed with Ampliseq RNA plugin to obtain read count matrices, that were exported and normalized through DESeq2 R package. DNA/RNA extraction and sequencing Nucleic acid extraction was performed using the MagMAX™ FFPE DNA/RNA Ultra Kit (ThermoFisher Scientific) on the KingFisher™ Duo Prime System (ThermoFisher Scientific) following the manufacturer’s instructions. Briefly, 6 µm-thick formalin-fixed paraffin-embedded (FFPE) tissue sections were deparaffinized and subjected to proteinase K digestion at 56°C overnight to ensure effective nucleic acid release. Nucleic acids were captured using paramagnetic beads and washed with buffers provided in the kit to remove contaminants. The extracted DNA and RNA were eluted in 50 µL of the elution buffer provided and quantified using Qubit™ Fluorometric Quantification (ThermoFisher Scientific) to assess concentration. Extracted nucleic acids were stored at -20°C and − 80°C, respectively, until further processing. Library preparation was performed using the Ion AmpliSeq™ Library Kit Plus (ThermoFisher Scientific) with a custom gene panel targeting regions of interest. Briefly, the extracted material was used for multiplex PCR amplification of target regions, followed by partial digestion of primer sequences using FuPa reagent. The amplicons were then ligated with Ion-compatible barcoded adapters using DNA Ligase and purified with magnetic bead-based selection. The final libraries were quantified using the Ion Library TaqMan® Quantitation Kit and diluted to a final concentration of 100 pM for template preparation. The prepared library pool was loaded onto the Ion Chef™ System (ThermoFisher Scientific) for automated loading onto the Ion 530™ Chip. Sequencing was performed on the Ion S5™ System (ThermoFisher Scientific) following standard protocols for sequencing run setup and data acquisition. Microsatellite Instability (MSI) evaluation MSI status was assessed using the Idylla™ MSI Test (Biocartis), an automated, real-time PCR–based assay designed for the qualitative detection of MSI in solid tumors. FFPE tumor tissue sections were used as input material according to the manufacturer’s instructions. Bulk RNAseq of the validation cohort A local cohort including 100 patients was enrolled and bulk RNAseq was performed to perform CMS subtyping. The study was approved by the Institutional Ethics Committee “Gabriella Serio” of the IRCCS Istituto Tumori “Giovanni Paolo II”, and all patients provided written informed consent. In light of the retrospective design, a Data Protection Impact Assessment was prepared by the Institutional Data Protection Officer in accordance with Article 36 of the EU Genaral Data Protection Regulation (Regulation (EU) 2016/679). This document was subsequently reviewed and approved by the Institutional Ethics Committee “Gabriella Serio” (Prot. n. 780/CE). The study was conducted in accordance with the Declaration of Helsinki. In detail, Next Generation Sequencing (NGS) experiments were performed by Genomix4life S.R.L. (Baronissi, Salerno, Italy) as described in [ 23 ]. Bioinformatic analysis and CMS classification Once the quality of all RNAseq data was confirmed to be adequate, the reference genome (hg38, downloaded from UCSC ( https://genome.ucsc.edu/ ) , was indexed using STAR (v.2.7.11b), and alignment was performed. Following alignment, BAM files were quantified using RSEM[ 24 ], generating two types of quantification for each sample: one at the gene level and one at the isoform level. Raw read count gene level matrix was used to run CMSclassifier R package [ 25 ] CMSclassifier accepted as input log2_scaled Gene Expression Profiles data values. Thus, we prepared two matrices, one log2 transformed and the other normalized using the Variance Stabilizing Transformation (VST) method, both obtained through DESeq2 R package. We use RandomForest methods in CMSclassify function. In addition, CMScaller classification was also performed using the CMScaller package[ 26 ]. CMScaller accepts as input only RNA-seq data and it can be run with counts or RSEM values by setting the parameter RNAseq=TRUE ; for VST normalized counts, the parameter was set to RNAseq=FALSE . Statistical analyses and Development of the prognostic nomogram To calculate metrics to evaluate validation performance of the in-house CMSassay, sensitivity, specificity, accuracy was calculated in R. Heatmaps with discordant cases were generated through pheatmap R package. Survival model construction Overall survival (OS) was defined as the time from diagnosis to death from any cause or last follow-up, expressed in months. Survival status was coded as an event for deceased patients and as censored for patients alive at last follow-up. A multivariable Cox proportional hazards regression model [ 27 ] was fitted using the cph function from the rms R package. The model included the following covariates selected a priori based on clinical relevance and biological plausibility: in-house CMSassay classification, sex, tumor sidedness, pathological stage, MSI status, and BRAF, KRAS, and PIK3CA mutational status. All categorical variables were modeled as factors. Model fitting retained the design matrix and response ( x = TRUE, y = TRUE ) and enabled survival probability estimation ( surv = TRUE ). The proportional hazards assumption was assessed graphically and found to be acceptable for all included covariates. Nomogram development Based on the fitted Cox model, a prognostic nomogram was constructed using the nomogram function [ 28 ] of the rms R package to provide individualized predictions of 3-year and 5-year overall survival. For each predictor, the nomogram assigns a number of points proportional to its regression coefficient. The sum of points across all variables yields a total prognostic score, which is mapped to the model’s linear predictor (LP) and corresponding survival probabilities at predefined time horizons (36 and 60 months). Survival probabilities at 3 and 5 years were derived from the Cox model using the baseline survival function estimated by the Survival() function in rms R package. To visualize the relationship between prognostic risk and outcome, survival probabilities were plotted as continuous functions of the LP. Additionally, a transformation was applied to convert the LP into total nomogram points, allowing a dual-axis graphical representation linking LP values, total point scores, and absolute survival probabilities. This visualization enables direct translation of the cumulative nomogram score into individualized 3- and 5-year OS estimates. All analyses and graphical outputs were performed with ggplot2 R package [ 29 ]. Results Best model identification To be able to design a custom NGS panel usable in clinical routine for CMS identification, the bioinformatic workflow was depicted in Fig. 1 a. At the very first instance, an in silico cohort was built up starting from TCGA-COAD and TCGA-READ cohorts. Patients to be included needed to have transcriptomic and mutational data available and MSI status information. Thus, the final matrix included 232 samples and 3804 features which correspond to 273 genes, 3491 lncRNAs, 39 alterations and MSI status (Table 1 ). mRMR technique was used to perform feature selection through two ML models, SVM and RF. Table 1 Number of original features and final selected features used for the design of the CMS assay Data Category Number features original dataset Final features number Gene Expression 273 7 Lnc RNAs 3491 32 MSI status 1 1 Alterations 39 30 As mentioned above, we used a wrapper feature selection method: we selected features from the overall dataset to form a small feature set. The variables were sorted by their scores based on mRMR algorithm; the first N variables were selected as our features. The best N was searched by step 10, from 10 to 200. For each subset of N features the classification result was observed. The best global accuracy achieved was 0.76 obtained when N = 70 with RF classifier (Fig. 1 b). The final model included 39 expression-related features (7 protein-coding genes and 32 lncRNAs), 30 DNA-related features, together with MSI status (Table 2 ). In Fig. 1 c, precision, recall, F1 and accuracy per CMS class are shown, highlighting a poorer performance for CMS2/3 than for CMS1/4. Table 2 List of final lncRNA, gene expression and mutation features selected for the in-house CMSassay panel. LncRNA Gene expression Alterations lnc-COX7C-7 DPYSL3 MTOR lnc-MBL2-2 SLIT2 NRAS LINC01980 RECK FANCL LINC00397 MRVI1 BARD1 LINC00355 IGF1 IDH1 lnc-SHISA9-3 SYNM RAF1 lnc-EVX1-14 UCHL1 PIK3CA lnc-KCNC2-2 TSC1 LINC02178 PTCH1 POU6F2-AS2 FGFR2 lnc-SNCA-5 PTEN LINC00393 HRAS lnc-CDH8-10 CHEK1 C21orf62-AS1 ATM lnc-TSPY2-6 KRAS LINC01036 FLT3 SERTAD4-AS1 BRCA2 lnc-TACR3-1 RAD51B lnc-PWWP2B-4 TSC2 LINC02405 CDK12 lnc-FAT1-1 RAD51C RPARP-AS1 NF1 lnc-HOXC8-1 BRIP1 lnc-CIT-1 BRCA1 lnc-TMEM132B-2 ERCC2 LINC00639 SMARCB1 LINC01967 CHEK2 SREBF2-AS1 KDM6A FAM66C BRAF SNHG1 MAP2K1 lnc-ALX1 lnc-ALDH1A3 Design of the CMS in-house assay The selected features by the approach mRMR + RF were used to design a custom amplicon-based NGS assay. In detail, the assay is a combination of two custom panels, one for the evaluation of genomic alterations and the other for expression, both coding genes and lncRNAs. As described in Methods section, MSI status was evaluated through an IVD assay. The DNA panel included the full coding sequence of 30 genes (Table 2 ), with a final coverage of 99.21% with 110690 bases covered over 111562 bases. The design included 1508 amplicons with an insert size ranging 125-175bp. In Fig. 2 , it is possible to observe the distribution of amplicon insert size which is tight and unimodal and with no long tails beyond ~ 175 bp. This design was considered optimal for FFPE samples to obtain a uniform amplification and a balanced coverage across targets. Indeed, we observed almost 3800*10 3 mapped reads, 98% of reads were on target and almost a mean depth of 2300. Regarding RNA-based panel, one/amplicon per target was designed taking into account the use of FFPE samples. We were able to obtain 100% of targets detected with almost 70% of valid reads and 2500*10 3 of mapped ones. Validation in a local independent cohort To validate the in-house CMSassay, an internal retrospective cohort including 100 patients has been enrolled. For the internal cohort, bulk RNA-seq was performed to be able to run CMSclassifier . As displayed in Fig. 3 a-b, the CMS distribution is very similar in our cohort and in TCGA cohort, which was used to identify the model and the design the in-house CMSassay. The general features of the local cohort are listed in Table 3 . The performances of in-house CMSassay were CMS-related. Indeed, we observed very good performances for CMS1 and CMS2 reaching a balanced accuracy of 0.87 and 0.82 and a specificity of 0.88 and 0.64, respectively (Table 4 ). Table 3 Features of internal validation cohort. Characteristic N = 100 1 Sex F 48 (48%) M 52 (52%) Age 74 (65, 79) Sideness Left 47 (47%) Right 53 (53%) Pathological Stage Stage I 8 (8.0%) Stage II 23 (23%) Stage III 38 (38%) Stage IV 31 (31%) CMS CMS1 17 (17%) CMS2 37 (37%) CMS3 14 (14%) CMS4 32 (32%) 1 n (%); Median (Q1, Q3) To deeply dive into classification performances in terms of discordant cases per CMS, the raw read count martix of the internal cohort was normalized using VST function in DESeq2 R package and log2 transformation. These two matrices were used to run CMScaller algorithm and CMSclassifier , the latter being run on the VST-normalized matrix. We observed that, across all classification methods, the largest number of discordant cases involved misclassification from CMS2 to CMS3 and from CMS2 to CMS4. Notably, CMScaller , in both configurations, showed poor performance in distinguishing CMS3 from CMS1 and CMS2. Moreover, CMSclassifier applied to VST-normalized data misclassified four CMS4 cases as CMS3 (Fig. 3 b). Overall, these findings suggest the presence of method- and normalization-dependent internal biases in CMS classification. Table 4 Sensitivity, specificity and balanced accuracy for each CMS subtypes obtained with the in-house CMSassay sensitivity CMS1 CMS2 CMS3 CMS4 0,857143 1 0,145833 0,340909 specificity 0,88172 0,636364 0,865385 0,696429 balanced_accuracy 0,869432 0,818182 0,505609 0,518669 Biological insights of in-house CMS assay and prognostic nomogram Biologically speaking CMS classification has intrinsically both a predictive and a prognostic function. Indeed, CMS1 immune subtype includes patients which could benefit from immunecheckpoint inhibitors and CMS4 patients have a very poor prognosis. Thus, we explore the distribution of in-house CMSassay classification with respect to routinely evaluated diagnostic (MS status, BRAF, KRAS and PIK3CA mutational status) and known prognostic features (sideness and stage). CMS1 is enriched of patients with II stage, right-sided tumors, BRAF and PIK3CA alterations, besides the predictable high number of MSI. CMS3 and CMS4 patients were found to be enriched of KRAS and PIK3CA alterations (Fig. 4 a). The chance to include in-house CMSassay into the diagnostic routine led us to the built-up of a nomogram. The nomogram (Fig. 4 b) incorporates in-house CMSassay classification, sex, tumor sidedness, pathological stage, MSI status, and BRAF, KRAS, and PIK3CA mutational status. Each covariate contributes a variable-specific number of points proportional to its effect size in the Cox proportional hazards model. The cumulative score, obtained by summing individual points across all predictors, corresponds to a linear predictor value, which is subsequently translated into predicted 3- and 5-year overall survival probabilities. Among all variables, tumor stage and CMS classification contributed the largest range of points, indicating a dominant impact on survival risk. Molecular alterations (MSI and oncogenic mutations) and clinicopathological features provided additional, incremental prognostic information. Figure 4 c illustrates the continuous relationship between the linear predictor derived from the nomogram and the corresponding predicted survival probabilities at 3 years (solid line) and 5 years (dashed line). Increasing LP values—reflecting higher total point scores—were associated with a monotonic decrease in survival probability, with a steeper decline observed for 5-year survival compared to 3-year survival. The upper x-axis translates LP values into total nomogram points, enabling direct visual conversion from the cumulative nomogram score to absolute survival probabilities. Patients with low total point scores exhibited high predicted survival at both time horizons, whereas high-risk patients showed markedly reduced long-term survival, approaching zero at extreme LP values. Discussion The CMS taxonomy has consolidated transcriptomic heterogeneity of colorectal cancer into four biologically and clinically meaningful groups. Indeed, CMS framework has profoundly shaped the biological interpretation of colorectal cancer heterogeneity, providing a reproducible transcriptomic taxonomy with well-defined associations to prognosis, tumor microenvironment, and key molecular alterations [ 4 , 30 , 31 ]. In the present study, we demonstrate that CMS classification retains robust prognostic relevance when implemented through a compact, FFPE-compatible in-house assay and integrated into a multivariable survival model alongside established clinicopathological and molecular factors. In line with previous reports, the in‑house CMSassay recapitulated these patterns, with CMS1 enriched for right‑sided, MSI‑high, BRAF/PIK3CA‑mutant, advanced-stage tumors and CMS4 associated with adverse clinicopathological features and poorer survival, supporting the biological plausibility of the panel and its potential integration into routine FFPE-based workflows [ 31 ]. Systematic reviews and meta-analyses confirm CMS as a prognostic biomarker across stages, with CMS4 linked to worst outcomes and CMS2 to best, independent of standard clinicopathological factors [ 32 ]. Emerging evidence also suggests predictive utility, such as CMS1 enrichment in immunotherapy responders and CMS2 benefit from anti‑EGFR agents, though prospective trials validating CMS‑guided therapy allocation remain scarce. By enabling CMS assessment from targeted FFPE panels, the assay bridges this gap toward precision oncology implementation [ 33 ]. Current implementations of CMS in clinical research mainly rely on CMSclassifier [ 4 ] and CMScaller [ 10 ], both of which were derived from, or benchmarked on, large meta-cohorts profiled predominantly with microarray or high‑quality bulk RNA‑seq and assume well‑normalized, genome‑wide gene expression matrices. In our study, direct comparison of CMSclassifier and CMScaller on the same validation cohort revealed substantial discordance, particularly at the CMS2–CMS3 and CMS2–CMS4 boundaries. Moreover, CMS assignments proved sensitive to data preprocessing choices, with normalization strategy (VST versus log2 transformation) influencing subtype calls. These findings are consistent with previous reports showing reduced classification stability for CMS3 and CMS4 and increased uncertainty near subtype borders[ 34 ]. Multiple independent cohorts and meta‑analyses have established CMS as an independent prognostic factor in both early‑stage and metastatic colorectal cancer, with CMS1 and CMS4 generally associated with worse long‑term outcomes compared with CMS2 and, variably, CMS3 [ 32 ]. The present study confirms this prognostic impact in a real‑world FFPE cohort, where CMS classification derived from the in‑house assay contributed one of the largest point ranges in the multivariable Cox‑based nomogram, second only to pathological stage, and provided incremental risk stratification beyond MSI, BRAF, KRAS and PIK3CA status. By integrating molecular (CMS, MSI, driver mutations) and clinicopathological variables (sex, sidedness, stage), the nomogram translates CMS information into individualized 3‑ and 5‑year overall survival estimates, which can be readily interpreted at the bedside. Thus, this approach moves from subtype labels to quantitative survival probabilities that may support patient counseling, follow‑up intensity, and trial stratification. Despite its robust prognostic value, the direct impact of CMS on systemic therapy selection remains modest compared with single‑marker biomarkers such as RAS, BRAF, MSI and HER2 [ 35 ]. Retrospective analyses suggest differential chemosensitivity and variable benefit from targeted agents across CMS groups, but the benefit of intensified regimens such as FOLFOXIRI plus bevacizumab appears largely independent of CMS, and prospective trials specifically powered to allocate therapy based on CMS are still lacking [ 33 ]. In this context, the in‑house CMSassay should currently be viewed primarily as a high‑resolution prognostic and biological stratification tool, rather than a stand‑alone predictive test guiding drug choice. This study has several limitations. First, discordance observed between CMSclassifier and CMScaller highlights the inherent instability of current CMS tools when applied across platforms and normalization strategies. This limitation reflects broader challenges in CMS standardization rather than assay-specific shortcomings, but it underscores the need for assay-calibrated classifiers. Finally, the prognostic nomogram was internally derived and requires external validation in independent, multi-center cohorts before clinical adoption. Prospective studies will be necessary to determine whether CMS-integrated prognostic tools can meaningfully influence clinical decision-making and patient outcomes. Conclusion By enabling CMS classification from routine FFPE material using a compact DNA/RNA panel combined with MSI testing, the present assay addresses a key translational barrier to CMS implementation. While this approach does not overcome the current limitations of CMS-guided therapy selection, it provides a clinically feasible platform to incorporate CMS into prognostic algorithms and real-world datasets. Future advances are likely to emerge from the integration of CMS with additional layers of tumor biology, including immune composition, stromal architecture, and functional drug sensitivity profiling. Such integrative strategies may ultimately convert the currently limited therapeutic output of CMS into actionable subtype-specific treatment algorithms. Declarations Ethics approval and consent to participate The Institutional Ethics Committee “Gabriella Serio” of the IRCCS Istituto Tumori “Giovanni Paolo II” approved the study Given the retrospective nature of the study, a Data Protection Impact Assessment document was redacted by the institutional Data Protection Officer in compliance with Article 36 of the EU General Data Protection Regulation (Regulation (EU) 2016/679). The document was evaluated and approved by the Institutional Ethics Committee “Gabriella Serio” (Prot n. 780/CE). The study is compliant with the Declaration of Helsinki. The authors affiliated to the IRCCS Istituto Tumori “Giovanni Paolo II”, Bari, are responsible for the views expressed in this article, which do not necessarily represent the Institute. Clinical trial number: not applicable. Consent for publication Not applicable Availability of data and materials Bulk RNAseq raw data were deposited in SRA repository under the accession number PRJNA1304982 (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1304982). Competing interests The authors report no disclosures relating to this work. Funding “Tecnopolo per la Medicina di Precisione”, CUP: B84I18000540002. Authors' contributions MC performed bioinformatic analyses and feature selection; LM: data management; DT, BP, GM performed all sequencing experiments; TMM validation analyses; EM, FAZ, AR, FC histological revision and FFPE sample retrieval; RDL, GB enrollment; ST, SDS, conceived the study and supervised the project. Acknowledgements All the patients. References Siegel R, Desantis C, Jemal A. Colorectal cancer statistics, 2014. CA Cancer J Clin. 2014;64:104–17. Sung H, Ferlay J, Siegel RL, et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. Cancer J Clin. 2021;71:209–49. Dang Q, Zuo L, Hu X, et al. Molecular subtypes of colorectal cancer in the era of precision oncotherapy: Current inspirations and future challenges. Cancer Med. 2024;13:e70041. Guinney J, Dienstmann R, Wang X, et al. The consensus molecular subtypes of colorectal cancer. Nat Med. 2015;21:1350–6. Eso Y, Shimizu T, Takeda H, et al. Microsatellite instability and immune checkpoint inhibitors: toward precision medicine against gastrointestinal and hepatobiliary cancers. J Gastroenterol. 2020;55:15–26. Ding X, Huang H, Fang Z, et al. From Subtypes to Solutions: Integrating CMS Classification with Precision Therapeutics in Colorectal Cancer. Curr Treat Options Oncol. 2024;25:1580–93. Valenzuela G, Canepa J, Simonetti C, et al. Consensus molecular subtypes of colorectal cancer in clinical practice: A translational approach. World J Clin Oncol. 2021;12:1000–8. Feliu J, Gámez-Pozo A, Martínez-Pérez D, et al. Functional proteomics of colon cancer Consensus Molecular Subtypes. Br J Cancer. 2024;130:1670–8. Chong W, Zhu X, Ren H, et al. Integrated multi-omics characterization of KRAS mutant colorectal cancer. Theranostics. 2022;12:5138–54. Eide PW, Bruun J, Lothe RA, et al. CMScaller: an R package for consensus molecular subtyping of colorectal cancer pre-clinical models. Sci Rep. 2017;7:16618. Quinn GP, Sessler T, Ahmaderaghi B, et al. classifieR a flexible interactive cloud-application for functional annotation of cancer transcriptomes. BMC Bioinformatics. 2022;23:114. GDC Data Portal Homepage. https://portal.gdc.cancer.gov/ (accessed 21 January 2026). Colaprico A, Silva TC, Olsen C, et al. TCGAbiolinks: an R/Bioconductor package for integrative analysis of TCGA data. Nucleic Acids Res. 2016;44:e71. Durinck S, Moreau Y, Kasprzyk A, et al. BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis. Bioinformatics. 2005;21:3439–40. Volders P-J, Anckaert J, Verheggen K, et al. LNCipedia 5: towards a reference set of human long non-coding RNAs. Nucleic Acids Res. 2019;47:D135–9. Anders S, Huber W. Differential expression analysis for sequence count data. Genome Biol. 2010;11:R106. Love MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15:550. Chakravarty D, Gao J, Phillips S et al. OncoKB: A Precision Oncology Knowledge Base. JCO Precis Oncol 2017; 1: PO.17.00011. De Jay N, Papillon-Cavanagh S, Olsen C, et al. mRMRe: an R package for parallelized mRMR ensemble feature selection. Bioinformatics. 2013;29:2365–8. Mayoraz E, Alpaydin E. Support vector machines for multi-class classification. In: Mira J, Sánchez-Andrés JV, editors. Engineering Applications of Bio-Inspired Artificial Neural Networks. Berlin, Heidelberg: Springer; 1999. pp. 833–42. Breiman L. Random Forests. Mach Learn. 2001;45:5–32. OncoKB. ™ · GitHub, https://github.com/oncokb/oncokb-annotator (accessed 19 January 2026). Monaco D, Traversa D, Mattioli E, et al. Biological and prognostic relevance of A-to-I RNA editing across consensus molecular subtypes of colon cancer. Sci Rep. 2026;16:4018. Li B, Dewey CN. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics. 2011;12:323. Sage-Bionetworks/CMSclassifier. https://github.com/Sage-Bionetworks/CMSclassifier (2025, accessed 20 January 2026). peterawe. peterawe/CMScaller. https://github.com/peterawe/CMScaller (2025, accessed 20 January 2026). Cox DR. Regression Models and Life-Tables. J Roy Stat Soc: Ser B (Methodol). 1972;34:187–202. Zhang Z, Kattan MW. Drawing Nomograms with R: applications to categorical outcome and survival data. Ann Transl Med. 2017;5:211. Wilkinson L. WICKHAM Biometrics. 2011;67:678–9. ggplot2: Elegant Graphics for Data Analysis by H. Dienstmann R, Vermeulen L, Guinney J, et al. Consensus molecular subtypes and the evolution of precision medicine in colorectal cancer. Nat Rev Cancer. 2017;17:79–92. Sawayama H, Miyamoto Y, Ogawa K, et al. Investigation of colorectal cancer in accordance with consensus molecular subtype classification. Ann Gastroenterol Surg. 2020;4:528–39. Ten Hoorn S, de Back TR, Sommeijer DW, et al. Clinical Value of Consensus Molecular Subtypes in Colorectal Cancer: A Systematic Review and Meta-Analysis. J Natl Cancer Inst. 2022;114:503–16. Borelli B, Fontana E, Giordano M, et al. Prognostic and predictive impact of consensus molecular subtypes and CRCAssigner classifications in metastatic colorectal cancer: a translational analysis of the TRIBE2 study. ESMO Open. 2021;6:100073. Trinh A, Trumpi K, De Sousa E, Melo F, et al. Practical and Robust Identification of Molecular Subtypes in Colorectal Cancer by Immunohistochemistry. Clin Cancer Res. 2017;23:387–98. Martelli V, Pastorino A, Sobrero AF. Prognostic and predictive molecular biomarkers in advanced colorectal cancer. Pharmacol Ther. 2022;236:108239. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 15 May, 2026 Reviews received at journal 05 May, 2026 Reviewers agreed at journal 03 May, 2026 Reviewers agreed at journal 25 Apr, 2026 Reviewers invited by journal 21 Apr, 2026 Editor assigned by journal 20 Apr, 2026 Editor invited by journal 27 Mar, 2026 Submission checks completed at journal 27 Mar, 2026 First submitted to journal 27 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9178611","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":631592641,"identity":"e0acbf69-9620-4e7c-a26a-687ff2d6d8e1","order_by":0,"name":"Mariapia Caputo","email":"","orcid":"","institution":"IRCCS Istituto Tumori \"Giovanni Paolo II\"","correspondingAuthor":false,"prefix":"","firstName":"Mariapia","middleName":"","lastName":"Caputo","suffix":""},{"id":631592642,"identity":"f041fcf7-1100-4cf1-bc2e-bc08f8598f64","order_by":1,"name":"Leonarda Maurmo","email":"","orcid":"","institution":"LHA BAT","correspondingAuthor":false,"prefix":"","firstName":"Leonarda","middleName":"","lastName":"Maurmo","suffix":""},{"id":631592643,"identity":"29680521-cbc6-4184-8f99-78ad4274ab64","order_by":2,"name":"Debora Traversa","email":"","orcid":"","institution":"IRCCS Istituto Tumori \"Giovanni Paolo II\"","correspondingAuthor":false,"prefix":"","firstName":"Debora","middleName":"","lastName":"Traversa","suffix":""},{"id":631592644,"identity":"88f0ff44-8d04-4682-8db6-9185d81aadf3","order_by":3,"name":"Tommaso Maria Marvulli","email":"","orcid":"","institution":"IRCCS Istituto Tumori \"Giovanni Paolo II\"","correspondingAuthor":false,"prefix":"","firstName":"Tommaso","middleName":"Maria","lastName":"Marvulli","suffix":""},{"id":631592645,"identity":"4e8e121d-b5b4-45f0-80ed-c6954eb6dca6","order_by":4,"name":"Eliseo Mattioli","email":"","orcid":"","institution":"IRCCS Istituto Tumori \"Giovanni Paolo II\"","correspondingAuthor":false,"prefix":"","firstName":"Eliseo","middleName":"","lastName":"Mattioli","suffix":""},{"id":631592646,"identity":"9bf6cb5d-a791-4707-a077-ce694b261ff1","order_by":5,"name":"Francesco Alfredo Zito","email":"","orcid":"","institution":"IRCCS Istituto Tumori \"Giovanni Paolo II\"","correspondingAuthor":false,"prefix":"","firstName":"Francesco","middleName":"Alfredo","lastName":"Zito","suffix":""},{"id":631592647,"identity":"07262412-1840-4769-8a9c-fac5e3f0fc78","order_by":6,"name":"Raffaele De Luca","email":"","orcid":"","institution":"IRCCS Istituto Tumori \"Giovanni Paolo II\"","correspondingAuthor":false,"prefix":"","firstName":"Raffaele","middleName":"","lastName":"De Luca","suffix":""},{"id":631592648,"identity":"f4088f24-9c90-4ce0-b362-4369f4aa54f6","order_by":7,"name":"Graziana Barile","email":"","orcid":"","institution":"S.S.D. Chirurgia Generale ad Indirizzo Senologico, IRCCS Istituto Tumori “Giovanni Paolo II”","correspondingAuthor":false,"prefix":"","firstName":"Graziana","middleName":"","lastName":"Barile","suffix":""},{"id":631592649,"identity":"9d377276-4015-425d-a22f-61b82931d583","order_by":8,"name":"Angela Ricco","email":"","orcid":"","institution":"IRCCS Istituto Tumori \"Giovanni Paolo II\"","correspondingAuthor":false,"prefix":"","firstName":"Angela","middleName":"","lastName":"Ricco","suffix":""},{"id":631592651,"identity":"d6a20075-6438-4c1e-9897-abd7398af321","order_by":9,"name":"Fulvia Colonna","email":"","orcid":"","institution":"IRCCS Istituto Tumori \"Giovanni Paolo II\"","correspondingAuthor":false,"prefix":"","firstName":"Fulvia","middleName":"","lastName":"Colonna","suffix":""},{"id":631592653,"identity":"136ca163-28b4-46f0-a392-8231bd990d30","order_by":10,"name":"Brunella Pilato","email":"","orcid":"","institution":"IRCCS Istituto Tumori \"Giovanni Paolo II\"","correspondingAuthor":false,"prefix":"","firstName":"Brunella","middleName":"","lastName":"Pilato","suffix":""},{"id":631592654,"identity":"4d028a08-008c-4816-9f88-6378de4fcce2","order_by":11,"name":"Giuseppina Matera","email":"","orcid":"","institution":"IRCCS Istituto Tumori \"Giovanni Paolo II\"","correspondingAuthor":false,"prefix":"","firstName":"Giuseppina","middleName":"","lastName":"Matera","suffix":""},{"id":631592656,"identity":"d6e5d142-a321-4f9c-b77b-04be11364a80","order_by":12,"name":"Stefania Tommasi","email":"","orcid":"","institution":"IRCCS Istituto Tumori \"Giovanni Paolo II\"","correspondingAuthor":false,"prefix":"","firstName":"Stefania","middleName":"","lastName":"Tommasi","suffix":""},{"id":631592657,"identity":"222f2237-bb17-4fac-869c-85cd3a334582","order_by":13,"name":"Simona De Summa","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABFklEQVRIiWNgGAWjYHACAwjF3gAiLRgYmIEUI5DDBibxaeE5ACIlULQ04tAD1SKRANXCANWCoFCBfAPzxoc/27ZF8898fOzBjwoJOXl35ocPGHfY5fMxMLc/wGbFAbZiY96227kzbqelG/ackTA2PMxmbMB4JtmyDYfDDBh4zKQZgVoabueAGBKJG5t52CQY25gNcPlFvoHH/OdPoJb5N8+gaKnHqYXhAI8ZA8hhG27wQLTMZwZrOYxTi8FhtmJpnnO3czeeSUuTBPnFgBnol8Qzxw3YmBkbZ2BzWHvzxo8/ym7nzjt++JjEjwobOfn+ww8ffNxRbSDf3v7gAzaHMWOGIZBIwCqFC8hjc/8oGAWjYBSMaAAAtzxeqi1Na4cAAAAASUVORK5CYII=","orcid":"","institution":"IRCCS Istituto Tumori \"Giovanni Paolo II\"","correspondingAuthor":true,"prefix":"","firstName":"Simona","middleName":"","lastName":"De Summa","suffix":""}],"badges":[],"createdAt":"2026-03-20 11:53:30","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9178611/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9178611/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":108398318,"identity":"d412669f-ad60-4479-8725-c70412d7da2c","added_by":"auto","created_at":"2026-05-04 08:29:14","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":239446,"visible":true,"origin":"","legend":"\u003cp\u003e(a) Bioinformatic workflow for the development of the in-house CMSassay. (b) Classification accuracy as function of the number of selected features using the mRMR features selection method for SVM classifier (blue dots) and RF classifier (red dots). (c) Precision, recall, F1 score and accuracy obtained by selecting 70 features using the RF classifier for each CMS subtypes.\u003c/p\u003e","description":"","filename":"Figure1211.png","url":"https://assets-eu.researchsquare.com/files/rs-9178611/v1/3af668b22a30e0ebb348ea2b.png"},{"id":108804228,"identity":"f0beaace-a715-44dc-9327-67cd62f3c4b1","added_by":"auto","created_at":"2026-05-08 15:18:12","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":111930,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of amplicon insert sizes for the DNA panel.\u003c/p\u003e","description":"","filename":"Figure1212.png","url":"https://assets-eu.researchsquare.com/files/rs-9178611/v1/e26bb093fda3deed36751227.png"},{"id":108493257,"identity":"19191e58-2682-4197-b432-5626307ae73a","added_by":"auto","created_at":"2026-05-05 09:59:48","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":142004,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of CMS subtypes in (a) the TCGA cohort and in (b) the internal validation cohort. (c) Heatmap of discordant predictions obtained using log2-normalized and VST-normalized data classified with CMScaller, VST-normalized data classified with CMSclassifier, and with the in-house CMSassay\u003cem\u003e.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"Figure1213.png","url":"https://assets-eu.researchsquare.com/files/rs-9178611/v1/00ae4d6c800f21d642ef30f5.png"},{"id":108398321,"identity":"4d87fc75-f7e5-4f8f-bea2-1e3ba59cd83f","added_by":"auto","created_at":"2026-05-04 08:29:14","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":108590,"visible":true,"origin":"","legend":"\u003cp\u003e(a) Biological insights from the in-house CMSassay, showing the distribution of tumor sidedness, pathological stage, MSI status, and BRAF, KRAS and PIK3CA mutations across CMS subtypes. (b) Nomogram incorporating in-house CMSassay classification, sex, tumor sidedness, pathological stage, MSI status, and BRAF, KRAS, and PIK3CA mutational status. (c) Continuous relationship between the linear predictor derived from the nomogram and the corresponding predicted 3-year (solid line) and 5-year (dashed line) survival probabilities.\u003c/p\u003e","description":"","filename":"Figure1214.png","url":"https://assets-eu.researchsquare.com/files/rs-9178611/v1/1b78d18d599118f0ba276b22.png"},{"id":109081165,"identity":"a6aae3fc-63c7-44ac-ba32-d05a4c915f77","added_by":"auto","created_at":"2026-05-12 12:02:51","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1020868,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9178611/v1/131229fa-524d-4e69-9f07-8e1f5cf6efd1.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Consensus molecular subtyping in colon cancer: from the design of in-house assay to prognostic nomogram","fulltext":[{"header":"Background","content":"\u003cp\u003eColorectal cancer (CRC) represents the third most commonly diagnosed cancer and the second leading cause of cancer-related deaths [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e], according to the global statistics.\u003c/p\u003e \u003cp\u003eCRC is a highly heterogeneous disease, characterized by complex interactions between genetic, epigenetic and environmental factors. This heterogeneity highlights significant challenges for diagnosis, prognosis and treatment [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eOne of the innovative solutions to address these unmet clinical needs, consists in the development of the molecular classification systems to stratify CRC into biologically different subtypes with clinical relevance. Among these, the Consensus Molecular Subtypes (CMS) classification has emerged as the most robust and widely accepted framework.\u003c/p\u003e \u003cp\u003eThe CMS classification, developed by Guinney et al. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e] in 2015, categorizes CRC into four distinct subtypes based on transcriptomic data: CMS1 (Microsatellite Instability Immune, MSI), CMS2 (Canonical), CMS3 (Metabolic), and CMS4 (Mesenchymal) [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eEach CMS subtype is associated with unique molecular, histological, and clinical characteristics. CMS1 is characterized by MSI [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], strong immune infiltration, and a favorable prognosis but poor response to conventional chemotherapy. CMS2, the most prevalent subtype, exhibits chromosomal instability, epithelial differentiation, and activation of the WNT and MYC signaling pathways [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], displaying a better response to chemotherapy and improved survival. CMS3 shows dysregulated metabolic pathways, particularly in glucose metabolism, and tends to occur in tumors with KRAS mutations [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. CMS4, associated with mesenchymal activation, stromal invasion, and angiogenesis, carries the worst prognosis due to its aggressive nature and resistance to standard treatments [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe CMS framework has provided valuable insights into the biological bases of CRC and has facilitated the development of precision medicine strategies. By identifying molecular subtypes, clinicians and researchers can better understand tumor behavior, predict treatment responses, and design targeted therapies.\u003c/p\u003e \u003cp\u003eSeveral challenges still remain in the clinical implementation of CMS classification, including the integration of transcriptomic profiling into routine practice and the need for subtype-specific therapeutic options.\u003c/p\u003e \u003cp\u003eIn these years, various algorithms have been developed to improve the CMS classification in clinical and research settings, including the \u003cem\u003eCMScaller\u003c/em\u003e [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], \u003cem\u003eCMSclassifier\u003c/em\u003e [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e], frameworks. Each of these algorithms offers unique strengths and limitations.\u003c/p\u003e \u003cp\u003eThe \u003cem\u003eCMScaller\u003c/em\u003e is a computational tool that uses gene expression profiles to assign CRC tumors to one of the four CMS categories [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. It relies on a predefined set of gene expression signatures and applies statistical methods to ensure accurate subtype assignment. The \u003cem\u003eCMScaller\u003c/em\u003e is known for its high reproducibility and robustness, making it a preferred choice in transcriptomic studies. However, its reliance on gene expression data necessitates high-quality RNA samples, which can be a limitation in certain clinical scenarios.\u003c/p\u003e \u003cp\u003eThe \u003cem\u003eCMSclassifier\u003c/em\u003e, similar to the \u003cem\u003eCMScaller\u003c/em\u003e, employs a supervised machine learning approach to classify CRC tumors [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. It incorporates additional biological parameters, such as mutation status and pathway activation scores, to enhance the accuracy of subtype assignments. This algorithm is particularly advantageous in settings where multi-omic data are available, as it provides a more integrative understanding of tumor biology. However, the \u003cem\u003eCMSclassifier\u003c/em\u003e requires more computational resources and data integration expertise, potentially limiting its accessibility for routine clinical use.\u003c/p\u003e \u003cp\u003eWhile the \u003cem\u003eCMScaller\u003c/em\u003e and \u003cem\u003eCMSclassifier\u003c/em\u003e share a common goal of implementing the CMS framework, their methodological differences highlight the importance of algorithm selection based on the research or clinical context. Together, these tools could provide a multifaceted approach to understanding CRC heterogeneity and advancing precision oncology.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eTCGA-COAD and TCGA-READ Data retrieval\u003c/h2\u003e \u003cp\u003eRNA expression data and corresponding clinical information from patients with colorectal cancer were downloaded from the Genomic Data Commons (GDC) Data Portal [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. The datasets were obtained from The Cancer Genome Atlas (TCGA) project, specifically from the TCGA-COAD (colon adenocarcinoma) and TCGA-READ (rectum adenocarcinoma) cohorts. \u003cem\u003eTCGAbiolinks\u003c/em\u003e R package [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] has been used to retrieve clinical data, HT-Seq count transcriptome data and miRNA Expression Quantification data. Ensemble ID has been converted into HUGO gene symbols by \u003cem\u003eBiomaRt\u003c/em\u003e R package [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. LncRNA expression data has been extrapolated by LNCpedia v5.2 conversion table [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. \u003cem\u003eTCGAbiolinks\u003c/em\u003e R package has been also used to retrieve alterations data, thus obtaining MAF-files. MSI status has been obtained by \u003cem\u003eGDCquery\u003c/em\u003e (data.category=\u0026rsquo;Other\u0026rsquo; and clinical.info=\u0026rsquo;msi\u0026rsquo;) function. CMSclassifier was used for the classification of samples in CMSs. It was implemented by Guinney et al.[\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e],and it is based on the 273 genes included in the original random forest (RF) CMS classifier.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003ePre-processing: sample filtering and data adjusting\u003c/h3\u003e\n\u003cp\u003eFor the analysis of count data from RNA-seq for the detection of differentially expressed genes, has been used \u003cem\u003eDESeq2\u003c/em\u003e R package [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. The count data, previously downloaded, are presented as a table (the rows correspond to the genes and the columns to the patients) which reports, for each sample, the number of sequence fragments that have been assigned to each gene. Firstly, a Summarized Experiment (SE) has been obtained by \u003cem\u003eGDCprepare\u003c/em\u003e (summarizedExperiment=TRUE) function. This SE was used to create a \u003cem\u003eDESeqDataSet\u003c/em\u003e class. While it is not necessary to pre-filter low count genes before running the \u003cem\u003eDESeq2\u003c/em\u003e functions, there are two reasons which make pre-filtering useful: by removing rows in which there are very few reads, the memory size of \u003cem\u003eDESeqDataSet\u003c/em\u003e object is reduced, and the speed of the transformation and testing functions within \u003cem\u003eDESeq2\u003c/em\u003e is increased. In our study, has been performed a minimal pre-filtering to keep only rows that have at least 10 reads total. After that, it has been used the \u003cem\u003evarianceStabilizingTranformation\u003c/em\u003e function which removes the dependence of the variance on the mean and then transforms the count data, obtaining a matrix of values which have constant variance along the range of mean values[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTo reduce the number of miRNAs and lncRNAs, before data normalization, further filtering was applied: all rows with reads count\u0026thinsp;\u0026lt;\u0026thinsp;10 in more than 10% of patients were eliminated.\u003c/p\u003e \u003cp\u003eFor data mutations we selected the clinically relevant mutations by \u003cem\u003eOncoKB\u003c/em\u003e, a precision oncology knowledge base [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. After that, have been selected only \u003cem\u003e'Nonsense'\u003c/em\u003e and \u003cem\u003e'Missense'\u003c/em\u003e mutations grouped by genes and has been created a matrix which reports, for each sample, the presence or absence of the mutation for each gene.\u003c/p\u003e\n\u003ch3\u003eFeature Selection and classification\u003c/h3\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eWrapper features selection method\u003c/h2\u003e \u003cp\u003eTo reduce the computational burden of classification and obtain a compact feature set that could be more easily applied in clinical practice, a feature selection step was performed. We adopted a wrapper-based feature selection approach, which evaluates subsets of features using a machine learning algorithm combined with a search strategy that explores the space of possible feature combinations. Each subset is assessed according to the predictive performance achieved by the selected algorithm. In general, wrapper methods follow these steps: (1) a subset of features is selected from the full set using a search strategy; (2) a machine learning model is trained on the selected subset; (3) model performance is evaluated using a predefined metric; and (4) the procedure is repeated iteratively across different feature subsets and corresponding trained models. In our study, we used minimum Redundancy Maximum Relevance (mRMR) [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e] for feature selection together with two classifiers, Support-Vector-Machine (SVM) [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] and Random Forest (RF) [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e], both of which are well suited for multi-label classification. Our specific aim was to assign patients to the four CMS classes.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003emRMR Feature Selection Method\u003c/h3\u003e\n\u003cp\u003eFeature selection was performed using the \u003cem\u003emRMR.classic\u003c/em\u003e fuction, implemented in the \u003cem\u003emRMRe\u003c/em\u003e R package [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], which enables efficient identification of relevant and non-redundant features. Let \u003cem\u003ey\u003c/em\u003e denote the output variable and \u003cem\u003eX=\u003c/em\u003e {x1,.., xn} the set of \u003cem\u003en\u003c/em\u003e input features. The method ranks features in \u003cem\u003eX\u003c/em\u003e by maximizing their mutual information (MI) with \u003cem\u003ey\u003c/em\u003e (maximum relevance) while minimizing the average MI with features already selected (minimum redundancy). Mutual information between two datasets is estimated on the basis of the correlation between variables Pearson\u0026rsquo;s or Spearman\u0026rsquo;s correlation coefficients are used for continuous variables, Cramer\u0026rsquo;s V for discrete variables, and Somers\u0026rsquo; Dxy index for the association between continuous variables and survival data.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eSupport Vector Machine and Random Forest\u003c/h2\u003e \u003cp\u003eSVMs are supervised learning methods commonly used for classification and regression tasks [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. By means of kernel functions, they allow efficient classification of non-linear data. In this framework, input samples are represented as points in an n-dimensional space, where each dimension corresponds to one feature. The algorithm then iteratively identifies a function defining a hyperplane that best separates the regions occupied by the different output classes. In our study, SVM classification was performed using the \u003cem\u003esvm()\u003c/em\u003e function from the \u003cem\u003ee1071\u003c/em\u003e R package. The RF algorithm, instead, combines multiple randomized decision trees and aggregates their predictions by majority voting [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. This approach has shown strong performance, particularly in settings where the number of variables greatly exceeds the number of observations. Its general workflow includes drawing random samples from the original dataset, building a decision tree for each sample, generating predictions from each tree, and selecting the class receiving the highest number of votes as the final output. In our study, RF classification was carried out using the \u003cem\u003erandomForest\u003c/em\u003e function from the \u003cem\u003erandomForest\u003c/em\u003e R package, based on Breiman\u0026rsquo;s original implementation [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eIn-house CMSassay: DNA/RNA custom panel design and bioinformatic analyses\u003c/h3\u003e\n\u003cp\u003eThe in-house CMSassay included two panels, both of them amplicon-based and designed through Ion Ampliseq Designer web tool (ThermoFisher). RNA panel required a great effort because it was built through multiple rounds of tiling/pooling in order to increase coverage rate. No SNPs were allowed under primers for the bulk of the design. Quality Control (QC) to avoid cross-binding of primers was performed for each panel. Two lncRNAs were excluded from the final design due technical issues in the identification of primers able to correctly address them. Alignment and variant calling was performed through Torrent Server plugins, Torrent Variant Caller in somatic configuration. From Ion Torrent Server we retrieved VCF files for each sample, which were annotated by OncoKB\u0026trade; annotator [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eRNA panel was analyzed with Ampliseq RNA plugin to obtain read count matrices, that were exported and normalized through \u003cem\u003eDESeq2\u003c/em\u003e R package.\u003c/p\u003e\n\u003ch3\u003eDNA/RNA extraction and sequencing\u003c/h3\u003e\n\u003cp\u003eNucleic acid extraction was performed using the MagMAX\u0026trade; FFPE DNA/RNA Ultra Kit (ThermoFisher Scientific) on the KingFisher\u0026trade; Duo Prime System (ThermoFisher Scientific) following the manufacturer\u0026rsquo;s instructions. Briefly, 6 \u0026micro;m-thick formalin-fixed paraffin-embedded (FFPE) tissue sections were deparaffinized and subjected to proteinase K digestion at 56\u0026deg;C overnight to ensure effective nucleic acid release. Nucleic acids were captured using paramagnetic beads and washed with buffers provided in the kit to remove contaminants. The extracted DNA and RNA were eluted in 50 \u0026micro;L of the elution buffer provided and quantified using Qubit\u0026trade; Fluorometric Quantification (ThermoFisher Scientific) to assess concentration. Extracted nucleic acids were stored at -20\u0026deg;C and \u0026minus;\u0026thinsp;80\u0026deg;C, respectively, until further processing.\u003c/p\u003e \u003cp\u003eLibrary preparation was performed using the Ion AmpliSeq\u0026trade; Library Kit Plus (ThermoFisher Scientific) with a custom gene panel targeting regions of interest. Briefly, the extracted material was used for multiplex PCR amplification of target regions, followed by partial digestion of primer sequences using FuPa reagent. The amplicons were then ligated with Ion-compatible barcoded adapters using DNA Ligase and purified with magnetic bead-based selection. The final libraries were quantified using the Ion Library TaqMan\u0026reg; Quantitation Kit and diluted to a final concentration of 100 pM for template preparation.\u003c/p\u003e \u003cp\u003eThe prepared library pool was loaded onto the Ion Chef\u0026trade; System (ThermoFisher Scientific) for automated loading onto the Ion 530\u0026trade; Chip. Sequencing was performed on the Ion S5\u0026trade; System (ThermoFisher Scientific) following standard protocols for sequencing run setup and data acquisition.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eMicrosatellite Instability (MSI) evaluation\u003c/h2\u003e \u003cp\u003eMSI status was assessed using the Idylla\u0026trade; MSI Test (Biocartis), an automated, real-time PCR\u0026ndash;based assay designed for the qualitative detection of MSI in solid tumors. FFPE tumor tissue sections were used as input material according to the manufacturer\u0026rsquo;s instructions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eBulk RNAseq of the validation cohort\u003c/h2\u003e \u003cp\u003eA local cohort including 100 patients was enrolled and bulk RNAseq was performed to perform CMS subtyping.\u003c/p\u003e \u003cp\u003eThe study was approved by the Institutional Ethics Committee \u0026ldquo;Gabriella Serio\u0026rdquo; of the IRCCS Istituto Tumori \u0026ldquo;Giovanni Paolo II\u0026rdquo;, and all patients provided written informed consent. In light of the retrospective design, a Data Protection Impact Assessment was prepared by the Institutional Data Protection Officer in accordance with Article 36 of the EU Genaral Data Protection Regulation (Regulation (EU) 2016/679). This document was subsequently reviewed and approved by the Institutional Ethics Committee \u0026ldquo;Gabriella Serio\u0026rdquo; (Prot. n. 780/CE). The study was conducted in accordance with the Declaration of Helsinki.\u003c/p\u003e \u003cp\u003eIn detail, Next Generation Sequencing (NGS) experiments were performed by Genomix4life S.R.L. (Baronissi, Salerno, Italy) as described in [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eBioinformatic analysis and CMS classification\u003c/h2\u003e \u003cp\u003eOnce the quality of all RNAseq data was confirmed to be adequate, the reference genome (hg38, downloaded from UCSC (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://genome.ucsc.edu/\u003c/span\u003e\u003cspan address=\"https://genome.ucsc.edu/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cem\u003e)\u003c/em\u003e, was indexed using STAR (v.2.7.11b), and alignment was performed. Following alignment, BAM files were quantified using RSEM[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e], generating two types of quantification for each sample: one at the gene level and one at the isoform level.\u003c/p\u003e \u003cp\u003eRaw read count gene level matrix was used to run \u003cem\u003eCMSclassifier\u003c/em\u003e R package [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] \u003cem\u003eCMSclassifier\u003c/em\u003e accepted as input log2_scaled Gene Expression Profiles data values. Thus, we prepared two matrices, one log2 transformed and the other normalized using the Variance Stabilizing Transformation (VST) method, both obtained through \u003cem\u003eDESeq2\u003c/em\u003e R package. We use \u003cem\u003eRandomForest\u003c/em\u003e methods in \u003cem\u003eCMSclassify\u003c/em\u003e function. In addition, \u003cem\u003eCMScaller\u003c/em\u003e classification was also performed using the \u003cem\u003eCMScaller\u003c/em\u003e package[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. CMScaller accepts as input only RNA-seq data and it can be run with counts or RSEM values by setting the parameter \u003cem\u003eRNAseq=TRUE\u003c/em\u003e; for VST normalized counts, the parameter was set to \u003cem\u003eRNAseq=FALSE\u003c/em\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eStatistical analyses and Development of the prognostic nomogram\u003c/h2\u003e \u003cp\u003eTo calculate metrics to evaluate validation performance of the in-house CMSassay, sensitivity, specificity, accuracy was calculated in R. Heatmaps with discordant cases were generated through \u003cem\u003epheatmap\u003c/em\u003e R package.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eSurvival model construction\u003c/h2\u003e \u003cp\u003eOverall survival (OS) was defined as the time from diagnosis to death from any cause or last follow-up, expressed in months. Survival status was coded as an event for deceased patients and as censored for patients alive at last follow-up. A multivariable Cox proportional hazards regression model [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e] was fitted using the \u003cem\u003ecph\u003c/em\u003e function from the \u003cem\u003erms\u003c/em\u003e R package. The model included the following covariates selected a priori based on clinical relevance and biological plausibility: in-house CMSassay classification, sex, tumor sidedness, pathological stage, MSI status, and BRAF, KRAS, and PIK3CA mutational status. All categorical variables were modeled as factors. Model fitting retained the design matrix and response (\u003cem\u003ex\u0026thinsp;=\u0026thinsp;TRUE, y\u0026thinsp;=\u0026thinsp;TRUE\u003c/em\u003e) and enabled survival probability estimation (\u003cem\u003esurv\u0026thinsp;=\u0026thinsp;TRUE\u003c/em\u003e). The proportional hazards assumption was assessed graphically and found to be acceptable for all included covariates.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eNomogram development\u003c/h2\u003e \u003cp\u003eBased on the fitted Cox model, a prognostic nomogram was constructed using the \u003cem\u003enomogram\u003c/em\u003e function [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e] of the \u003cem\u003erms\u003c/em\u003e R package to provide individualized predictions of 3-year and 5-year overall survival. For each predictor, the nomogram assigns a number of points proportional to its regression coefficient. The sum of points across all variables yields a total prognostic score, which is mapped to the model\u0026rsquo;s linear predictor (LP) and corresponding survival probabilities at predefined time horizons (36 and 60 months). Survival probabilities at 3 and 5 years were derived from the Cox model using the baseline survival function estimated by the \u003cem\u003eSurvival()\u003c/em\u003e function in \u003cem\u003erms\u003c/em\u003e R package. To visualize the relationship between prognostic risk and outcome, survival probabilities were plotted as continuous functions of the LP. Additionally, a transformation was applied to convert the LP into total nomogram points, allowing a dual-axis graphical representation linking LP values, total point scores, and absolute survival probabilities. This visualization enables direct translation of the cumulative nomogram score into individualized 3- and 5-year OS estimates.\u003c/p\u003e \u003cp\u003eAll analyses and graphical outputs were performed with \u003cem\u003eggplot2\u003c/em\u003e R package [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eBest model identification\u003c/h2\u003e \u003cp\u003eTo be able to design a custom NGS panel usable in clinical routine for CMS identification, the bioinformatic workflow was depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea. At the very first instance, an \u003cem\u003ein silico\u003c/em\u003e cohort was built up starting from TCGA-COAD and TCGA-READ cohorts.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003ePatients to be included needed to have transcriptomic and mutational data available and MSI status information. Thus, the final matrix included 232 samples and 3804 features which correspond to 273 genes, 3491 lncRNAs, 39 alterations and MSI status (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003emRMR technique was used to perform feature selection through two ML models, SVM and RF.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eNumber of original features and final selected features used for the design of the CMS assay\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eData Category\u003c/span\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eNumber features original dataset\u003c/span\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eFinal features number\u003c/span\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eGene Expression\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003e273\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003e7\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eLnc RNAs\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003e3491\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003e32\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eMSI status\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003e1\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003e1\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cspan type=\"ItalicSmallCaps\" class=\"ItalicSmallCaps\" name=\"Emphasis\"\u003eAlterations\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cspan type=\"ItalicSmallCaps\" class=\"ItalicSmallCaps\" name=\"Emphasis\"\u003e39\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cspan type=\"ItalicSmallCaps\" class=\"ItalicSmallCaps\" name=\"Emphasis\"\u003e30\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAs mentioned above, we used a wrapper feature selection method: we selected features from the overall dataset to form a small feature set. The variables were sorted by their scores based on mRMR algorithm; the first \u003cem\u003eN\u003c/em\u003e variables were selected as our features. The best \u003cem\u003eN\u003c/em\u003e was searched by step 10, from 10 to 200. For each subset of \u003cem\u003eN\u003c/em\u003e features the classification result was observed. The best global accuracy achieved was 0.76 obtained when \u003cem\u003eN\u003c/em\u003e\u0026thinsp;=\u0026thinsp;70 with RF classifier (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb). The final model included 39 expression-related features (7 protein-coding genes and 32 lncRNAs), 30 DNA-related features, together with MSI status (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). In Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ec, precision, recall, F1 and accuracy per CMS class are shown, highlighting a poorer performance for CMS2/3 than for CMS1/4.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eList of final lncRNA, gene expression and mutation features selected for the in-house CMSassay panel.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLncRNA\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eGene expression\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAlterations\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elnc-COX7C-7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eDPYSL3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMTOR\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elnc-MBL2-2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eSLIT2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNRAS\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLINC01980\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eRECK\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFANCL\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLINC00397\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eMRVI1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBARD1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLINC00355\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eIGF1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eIDH1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elnc-SHISA9-3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eSYNM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRAF1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elnc-EVX1-14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eUCHL1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePIK3CA\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elnc-KCNC2-2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTSC1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLINC02178\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePTCH1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePOU6F2-AS2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFGFR2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elnc-SNCA-5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePTEN\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLINC00393\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHRAS\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elnc-CDH8-10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCHEK1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC21orf62-AS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eATM\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elnc-TSPY2-6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eKRAS\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLINC01036\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eFLT3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSERTAD4-AS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBRCA2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elnc-TACR3-1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRAD51B\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elnc-PWWP2B-4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eTSC2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLINC02405\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCDK12\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003elnc-FAT1-1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRAD51C\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eRPARP-AS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNF1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003elnc-HOXC8-1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBRIP1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003elnc-CIT-1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBRCA1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003elnc-TMEM132B-2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eERCC2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eLINC00639\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSMARCB1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eLINC01967\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCHEK2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eSREBF2-AS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eKDM6A\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eFAM66C\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBRAF\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eSNHG1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMAP2K1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003elnc-ALX1\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003elnc-ALDH1A3\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eDesign of the CMS in-house assay\u003c/h2\u003e \u003cp\u003eThe selected features by the approach mRMR\u0026thinsp;+\u0026thinsp;RF were used to design a custom amplicon-based NGS assay. In detail, the assay is a combination of two custom panels, one for the evaluation of genomic alterations and the other for expression, both coding genes and lncRNAs. As described in Methods section, MSI status was evaluated through an IVD assay.\u003c/p\u003e \u003cp\u003eThe DNA panel included the full coding sequence of 30 genes (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e), with a final coverage of 99.21% with 110690 bases covered over 111562 bases.\u003c/p\u003e \u003cp\u003eThe design included 1508 amplicons with an insert size ranging 125-175bp. In Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, it is possible to observe the distribution of amplicon insert size which is tight and unimodal and with no long tails beyond ~\u0026thinsp;175 bp. This design was considered optimal for FFPE samples to obtain a uniform amplification and a balanced coverage across targets. Indeed, we observed almost 3800*10\u003csup\u003e3\u003c/sup\u003e mapped reads, 98% of reads were on target and almost a mean depth of 2300.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eRegarding RNA-based panel, one/amplicon per target was designed taking into account the use of FFPE samples. We were able to obtain 100% of targets detected with almost 70% of valid reads and 2500*10\u003csup\u003e3\u003c/sup\u003e of mapped ones.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eValidation in a local independent cohort\u003c/h2\u003e \u003cp\u003eTo validate the in-house CMSassay, an internal retrospective cohort including 100 patients has been enrolled. For the internal cohort, bulk RNA-seq was performed to be able to run \u003cem\u003eCMSclassifier\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eAs displayed in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea-b, the CMS distribution is very similar in our cohort and in TCGA cohort, which was used to identify the model and the design the in-house CMSassay. The general features of the local cohort are listed in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe performances of in-house CMSassay were CMS-related. Indeed, we observed very good performances for CMS1 and CMS2 reaching a balanced accuracy of 0.87 and 0.82 and a specificity of 0.88 and 0.64, respectively (Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eFeatures of internal validation cohort.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCharacteristic\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eN\u0026thinsp;=\u0026thinsp;100\u003csup\u003e1\u003c/sup\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSex\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e48 (48%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e52 (52%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAge\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e74 (65, 79)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eSideness\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLeft\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e47 (47%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRight\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e53 (53%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ePathological Stage\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStage I\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8 (8.0%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStage II\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e23 (23%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStage III\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e38 (38%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStage IV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e31 (31%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eCMS\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCMS1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e17 (17%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCMS2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e37 (37%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCMS3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e14 (14%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCMS4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e32 (32%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"2\"\u003e\u003csup\u003e1\u003c/sup\u003en (%); Median (Q1, Q3)\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTo deeply dive into classification performances in terms of discordant cases per CMS, the raw read count martix of the internal cohort was normalized using \u003cem\u003eVST\u003c/em\u003e function in \u003cem\u003eDESeq2\u003c/em\u003e R package and log2 transformation. These two matrices were used to run \u003cem\u003eCMScaller\u003c/em\u003e algorithm and \u003cem\u003eCMSclassifier\u003c/em\u003e, the latter being run on the VST-normalized matrix.\u003c/p\u003e \u003cp\u003eWe observed that, across all classification methods, the largest number of discordant cases involved misclassification from CMS2 to CMS3 and from CMS2 to CMS4. Notably, \u003cem\u003eCMScaller\u003c/em\u003e, in both configurations, showed poor performance in distinguishing CMS3 from CMS1 and CMS2. Moreover, CMSclassifier applied to VST-normalized data misclassified four CMS4 cases as CMS3 (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eb). Overall, these findings suggest the presence of method- and normalization-dependent internal biases in CMS classification.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSensitivity, specificity and balanced accuracy for each CMS subtypes obtained with the in-house CMSassay\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cb\u003esensitivity\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCMS1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCMS2\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCMS3\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCMS4\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0,857143\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,145833\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,340909\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003especificity\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0,88172\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,636364\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,865385\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,696429\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ebalanced_accuracy\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0,869432\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0,818182\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0,505609\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0,518669\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eBiological insights of in-house CMS assay and prognostic nomogram\u003c/h2\u003e \u003cp\u003eBiologically speaking CMS classification has intrinsically both a predictive and a prognostic function. Indeed, CMS1 immune subtype includes patients which could benefit from immunecheckpoint inhibitors and CMS4 patients have a very poor prognosis. Thus, we explore the distribution of in-house CMSassay classification with respect to routinely evaluated diagnostic (MS status, BRAF, KRAS and PIK3CA mutational status) and known prognostic features (sideness and stage). CMS1 is enriched of patients with II stage, right-sided tumors, BRAF and PIK3CA alterations, besides the predictable high number of MSI. CMS3 and CMS4 patients were found to be enriched of KRAS and PIK3CA alterations (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ea).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe chance to include in-house CMSassay into the diagnostic routine led us to the built-up of a nomogram. The nomogram (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eb) incorporates in-house CMSassay classification, sex, tumor sidedness, pathological stage, MSI status, and BRAF, KRAS, and PIK3CA mutational status. Each covariate contributes a variable-specific number of points proportional to its effect size in the Cox proportional hazards model. The cumulative score, obtained by summing individual points across all predictors, corresponds to a linear predictor value, which is subsequently translated into predicted 3- and 5-year overall survival probabilities.\u003c/p\u003e \u003cp\u003eAmong all variables, tumor stage and CMS classification contributed the largest range of points, indicating a dominant impact on survival risk. Molecular alterations (MSI and oncogenic mutations) and clinicopathological features provided additional, incremental prognostic information.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ec illustrates the continuous relationship between the linear predictor derived from the nomogram and the corresponding predicted survival probabilities at 3 years (solid line) and 5 years (dashed line).\u003c/p\u003e \u003cp\u003eIncreasing LP values\u0026mdash;reflecting higher total point scores\u0026mdash;were associated with a monotonic decrease in survival probability, with a steeper decline observed for 5-year survival compared to 3-year survival. The upper x-axis translates LP values into total nomogram points, enabling direct visual conversion from the cumulative nomogram score to absolute survival probabilities. Patients with low total point scores exhibited high predicted survival at both time horizons, whereas high-risk patients showed markedly reduced long-term survival, approaching zero at extreme LP values.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe CMS taxonomy has consolidated transcriptomic heterogeneity of colorectal cancer into four biologically and clinically meaningful groups. Indeed, CMS framework has profoundly shaped the biological interpretation of colorectal cancer heterogeneity, providing a reproducible transcriptomic taxonomy with well-defined associations to prognosis, tumor microenvironment, and key molecular alterations [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn the present study, we demonstrate that CMS classification retains robust prognostic relevance when implemented through a compact, FFPE-compatible in-house assay and integrated into a multivariable survival model alongside established clinicopathological and molecular factors.\u003c/p\u003e \u003cp\u003eIn line with previous reports, the in‑house CMSassay recapitulated these patterns, with CMS1 enriched for right‑sided, MSI‑high, BRAF/PIK3CA‑mutant, advanced-stage tumors and CMS4 associated with adverse clinicopathological features and poorer survival, supporting the biological plausibility of the panel and its potential integration into routine FFPE-based workflows [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSystematic reviews and meta-analyses confirm CMS as a prognostic biomarker across stages, with CMS4 linked to worst outcomes and CMS2 to best, independent of standard clinicopathological factors [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eEmerging evidence also suggests predictive utility, such as CMS1 enrichment in immunotherapy responders and CMS2 benefit from anti‑EGFR agents, though prospective trials validating CMS‑guided therapy allocation remain scarce. By enabling CMS assessment from targeted FFPE panels, the assay bridges this gap toward precision oncology implementation [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eCurrent implementations of CMS in clinical research mainly rely on \u003cem\u003eCMSclassifier\u003c/em\u003e [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e] and \u003cem\u003eCMScaller\u003c/em\u003e [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], both of which were derived from, or benchmarked on, large meta-cohorts profiled predominantly with microarray or high‑quality bulk RNA‑seq and assume well‑normalized, genome‑wide gene expression matrices.\u003c/p\u003e \u003cp\u003eIn our study, direct comparison of \u003cem\u003eCMSclassifier\u003c/em\u003e and \u003cem\u003eCMScaller\u003c/em\u003e on the same validation cohort revealed substantial discordance, particularly at the CMS2\u0026ndash;CMS3 and CMS2\u0026ndash;CMS4 boundaries. Moreover, CMS assignments proved sensitive to data preprocessing choices, with normalization strategy (VST versus log2 transformation) influencing subtype calls. These findings are consistent with previous reports showing reduced classification stability for CMS3 and CMS4 and increased uncertainty near subtype borders[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eMultiple independent cohorts and meta‑analyses have established CMS as an independent prognostic factor in both early‑stage and metastatic colorectal cancer, with CMS1 and CMS4 generally associated with worse long‑term outcomes compared with CMS2 and, variably, CMS3 [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. The present study confirms this prognostic impact in a real‑world FFPE cohort, where CMS classification derived from the in‑house assay contributed one of the largest point ranges in the multivariable Cox‑based nomogram, second only to pathological stage, and provided incremental risk stratification beyond MSI, BRAF, KRAS and PIK3CA status.\u003c/p\u003e \u003cp\u003eBy integrating molecular (CMS, MSI, driver mutations) and clinicopathological variables (sex, sidedness, stage), the nomogram translates CMS information into individualized 3‑ and 5‑year overall survival estimates, which can be readily interpreted at the bedside. Thus, this approach moves from subtype labels to quantitative survival probabilities that may support patient counseling, follow‑up intensity, and trial stratification.\u003c/p\u003e \u003cp\u003eDespite its robust prognostic value, the direct impact of CMS on systemic therapy selection remains modest compared with single‑marker biomarkers such as RAS, BRAF, MSI and HER2 [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. Retrospective analyses suggest differential chemosensitivity and variable benefit from targeted agents across CMS groups, but the benefit of intensified regimens such as FOLFOXIRI plus bevacizumab appears largely independent of CMS, and prospective trials specifically powered to allocate therapy based on CMS are still lacking [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. In this context, the in‑house CMSassay should currently be viewed primarily as a high‑resolution prognostic and biological stratification tool, rather than a stand‑alone predictive test guiding drug choice.\u003c/p\u003e \u003cp\u003eThis study has several limitations. First, discordance observed between \u003cem\u003eCMSclassifier\u003c/em\u003e and \u003cem\u003eCMScaller\u003c/em\u003e highlights the inherent instability of current CMS tools when applied across platforms and normalization strategies. This limitation reflects broader challenges in CMS standardization rather than assay-specific shortcomings, but it underscores the need for assay-calibrated classifiers.\u003c/p\u003e \u003cp\u003eFinally, the prognostic nomogram was internally derived and requires external validation in independent, multi-center cohorts before clinical adoption. Prospective studies will be necessary to determine whether CMS-integrated prognostic tools can meaningfully influence clinical decision-making and patient outcomes.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eBy enabling CMS classification from routine FFPE material using a compact DNA/RNA panel combined with MSI testing, the present assay addresses a key translational barrier to CMS implementation. While this approach does not overcome the current limitations of CMS-guided therapy selection, it provides a clinically feasible platform to incorporate CMS into prognostic algorithms and real-world datasets.\u003c/p\u003e \u003cp\u003eFuture advances are likely to emerge from the integration of CMS with additional layers of tumor biology, including immune composition, stromal architecture, and functional drug sensitivity profiling. Such integrative strategies may ultimately convert the currently limited therapeutic output of CMS into actionable subtype-specific treatment algorithms.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Institutional Ethics Committee \u0026ldquo;Gabriella Serio\u0026rdquo; of the IRCCS Istituto Tumori \u0026ldquo;Giovanni Paolo II\u0026rdquo; approved the study Given the retrospective nature of the study, a Data Protection Impact Assessment document was redacted by the institutional Data Protection Officer in compliance with Article 36 of the EU General Data Protection Regulation (Regulation (EU) 2016/679). The document was evaluated and approved by the Institutional Ethics Committee \u0026ldquo;Gabriella Serio\u0026rdquo; (Prot n. 780/CE). The study is compliant with the Declaration of Helsinki.\u003c/p\u003e\n\u003cp\u003eThe authors affiliated to the IRCCS Istituto Tumori \u0026ldquo;Giovanni Paolo II\u0026rdquo;, Bari, are responsible for the views expressed in this article, which do not necessarily represent the Institute.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eClinical trial number: not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBulk RNAseq raw data were deposited in SRA repository under the accession number PRJNA1304982 (https://www.ncbi.nlm.nih.gov/bioproject/PRJNA1304982).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors report no disclosures relating to this work.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u0026ldquo;Tecnopolo per la Medicina di Precisione\u0026rdquo;, CUP: B84I18000540002.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMC performed bioinformatic analyses and feature selection; LM: data management; DT, BP, GM performed all sequencing experiments; TMM validation analyses; EM, FAZ, AR, FC histological revision and FFPE sample retrieval; RDL, GB enrollment; ST, SDS, conceived the study and supervised the project.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll the patients.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eSiegel R, Desantis C, Jemal A. Colorectal cancer statistics, 2014. CA Cancer J Clin. 2014;64:104\u0026ndash;17.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSung H, Ferlay J, Siegel RL, et al. Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. Cancer J Clin. 2021;71:209\u0026ndash;49.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDang Q, Zuo L, Hu X, et al. Molecular subtypes of colorectal cancer in the era of precision oncotherapy: Current inspirations and future challenges. Cancer Med. 2024;13:e70041.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuinney J, Dienstmann R, Wang X, et al. The consensus molecular subtypes of colorectal cancer. Nat Med. 2015;21:1350\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEso Y, Shimizu T, Takeda H, et al. Microsatellite instability and immune checkpoint inhibitors: toward precision medicine against gastrointestinal and hepatobiliary cancers. J Gastroenterol. 2020;55:15\u0026ndash;26.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDing X, Huang H, Fang Z, et al. From Subtypes to Solutions: Integrating CMS Classification with Precision Therapeutics in Colorectal Cancer. Curr Treat Options Oncol. 2024;25:1580\u0026ndash;93.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eValenzuela G, Canepa J, Simonetti C, et al. Consensus molecular subtypes of colorectal cancer in clinical practice: A translational approach. World J Clin Oncol. 2021;12:1000\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFeliu J, G\u0026aacute;mez-Pozo A, Mart\u0026iacute;nez-P\u0026eacute;rez D, et al. Functional proteomics of colon cancer Consensus Molecular Subtypes. Br J Cancer. 2024;130:1670\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChong W, Zhu X, Ren H, et al. Integrated multi-omics characterization of KRAS mutant colorectal cancer. Theranostics. 2022;12:5138\u0026ndash;54.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEide PW, Bruun J, Lothe RA, et al. CMScaller: an R package for consensus molecular subtyping of colorectal cancer pre-clinical models. Sci Rep. 2017;7:16618.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQuinn GP, Sessler T, Ahmaderaghi B, et al. classifieR a flexible interactive cloud-application for functional annotation of cancer transcriptomes. BMC Bioinformatics. 2022;23:114.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGDC Data Portal Homepage. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://portal.gdc.cancer.gov/\u003c/span\u003e\u003cspan address=\"https://portal.gdc.cancer.gov/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (accessed 21 January 2026).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eColaprico A, Silva TC, Olsen C, et al. TCGAbiolinks: an R/Bioconductor package for integrative analysis of TCGA data. Nucleic Acids Res. 2016;44:e71.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDurinck S, Moreau Y, Kasprzyk A, et al. BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis. Bioinformatics. 2005;21:3439\u0026ndash;40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVolders P-J, Anckaert J, Verheggen K, et al. LNCipedia 5: towards a reference set of human long non-coding RNAs. Nucleic Acids Res. 2019;47:D135\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAnders S, Huber W. Differential expression analysis for sequence count data. Genome Biol. 2010;11:R106.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLove MI, Huber W, Anders S. Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol. 2014;15:550.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChakravarty D, Gao J, Phillips S et al. OncoKB: A Precision Oncology Knowledge Base. JCO Precis Oncol 2017; 1: PO.17.00011.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDe Jay N, Papillon-Cavanagh S, Olsen C, et al. mRMRe: an R package for parallelized mRMR ensemble feature selection. Bioinformatics. 2013;29:2365\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMayoraz E, Alpaydin E. Support vector machines for multi-class classification. In: Mira J, S\u0026aacute;nchez-Andr\u0026eacute;s JV, editors. Engineering Applications of Bio-Inspired Artificial Neural Networks. Berlin, Heidelberg: Springer; 1999. pp. 833\u0026ndash;42.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBreiman L. Random Forests. Mach Learn. 2001;45:5\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOncoKB. \u003csup\u003e\u0026trade;\u003c/sup\u003e \u0026middot; GitHub, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/oncokb/oncokb-annotator\u003c/span\u003e\u003cspan address=\"https://github.com/oncokb/oncokb-annotator\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (accessed 19 January 2026).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMonaco D, Traversa D, Mattioli E, et al. Biological and prognostic relevance of A-to-I RNA editing across consensus molecular subtypes of colon cancer. Sci Rep. 2026;16:4018.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi B, Dewey CN. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics. 2011;12:323.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSage-Bionetworks/CMSclassifier. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/Sage-Bionetworks/CMSclassifier\u003c/span\u003e\u003cspan address=\"https://github.com/Sage-Bionetworks/CMSclassifier\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025, accessed 20 January 2026).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003epeterawe. peterawe/CMScaller. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/peterawe/CMScaller\u003c/span\u003e\u003cspan address=\"https://github.com/peterawe/CMScaller\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025, accessed 20 January 2026).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCox DR. Regression Models and Life-Tables. J Roy Stat Soc: Ser B (Methodol). 1972;34:187\u0026ndash;202.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Z, Kattan MW. Drawing Nomograms with R: applications to categorical outcome and survival data. Ann Transl Med. 2017;5:211.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWilkinson L. WICKHAM Biometrics. 2011;67:678\u0026ndash;9. ggplot2: Elegant Graphics for Data Analysis by H.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDienstmann R, Vermeulen L, Guinney J, et al. Consensus molecular subtypes and the evolution of precision medicine in colorectal cancer. Nat Rev Cancer. 2017;17:79\u0026ndash;92.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSawayama H, Miyamoto Y, Ogawa K, et al. Investigation of colorectal cancer in accordance with consensus molecular subtype classification. Ann Gastroenterol Surg. 2020;4:528\u0026ndash;39.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTen Hoorn S, de Back TR, Sommeijer DW, et al. Clinical Value of Consensus Molecular Subtypes in Colorectal Cancer: A Systematic Review and Meta-Analysis. J Natl Cancer Inst. 2022;114:503\u0026ndash;16.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBorelli B, Fontana E, Giordano M, et al. Prognostic and predictive impact of consensus molecular subtypes and CRCAssigner classifications in metastatic colorectal cancer: a translational analysis of the TRIBE2 study. ESMO Open. 2021;6:100073.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTrinh A, Trumpi K, De Sousa E, Melo F, et al. Practical and Robust Identification of Molecular Subtypes in Colorectal Cancer by Immunohistochemistry. Clin Cancer Res. 2017;23:387\u0026ndash;98.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMartelli V, Pastorino A, Sobrero AF. Prognostic and predictive molecular biomarkers in advanced colorectal cancer. Pharmacol Ther. 2022;236:108239.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-gastroenterology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmge","sideBox":"Learn more about [BMC Gastroenterology](http://bmcgastroenterol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmge/default.aspx","title":"BMC Gastroenterology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"CMS, colon cancer, in house panel, transcriptomics, nomogram","lastPublishedDoi":"10.21203/rs.3.rs-9178611/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9178611/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground.\u003c/h2\u003e \u003cp\u003eConsensus Molecular Subtypes (CMS) provide a biologically robust classification of colorectal cancer (CRC) with prognostic relevance; however, their clinical implementation remains limited by reliance on genome-wide transcriptomics and platform-dependent classifiers. We aimed to develop and validate a compact, FFPE-compatible in-house CMS assay and to integrate CMS information into a clinically applicable prognostic nomogram.\u003c/p\u003e\u003ch2\u003eMethods.\u003c/h2\u003e \u003cp\u003eAn in silico training cohort was generated from TCGA-COAD and TCGA-READ datasets, including transcriptomic profiles, somatic mutations, and microsatellite instability (MSI) status. Feature selection was performed using a minimum Redundancy Maximum Relevance (mRMR) wrapper approach combined with Random Forest and Support Vector Machine classifiers to identify a minimal CMS-informative feature set. Based on selected features, a custom amplicon-based DNA/RNA NGS assay was designed. The assay was validated in an independent cohort of 100 CRC patients with bulk RNAseq\u0026ndash;based CMS classification. Classification performance was assessed against CMSclassifier and CMScaller. A multivariable Cox proportional hazards model incorporating CMS, clinicopathological variables, MSI, and driver mutations was used to develop a prognostic nomogram for 3- and 5-year overall survival.\u003c/p\u003e\u003ch2\u003eResults.\u003c/h2\u003e \u003cp\u003eThe optimal model achieved a global accuracy of 0.76 using 70 features, comprising 7 coding genes, 32 lncRNAs, 30 genomic alterations, and MSI status. The in-house CMS assay demonstrated high balanced accuracy for CMS1 and CMS2, with lower performance for CMS3 and CMS4. Integration of CMSassay classification into the prognostic model showed that CMS classification and tumor stage contributed the largest effects on survival risk. The resulting nomogram provided individualized 3- and 5-year overall survival estimates with clear risk stratification.\u003c/p\u003e\u003ch2\u003eConclusions.\u003c/h2\u003e \u003cp\u003eWe developed a clinically feasible, FFPE-compatible CMS assay that preserves the prognostic relevance of transcriptome-based CMS classification while reducing technical and computational complexity. When integrated into a multivariable nomogram, CMS provides substantial prognostic value beyond standard molecular and clinicopathological factors. This approach facilitates the translation of CMS into routine clinical workflows and supports its use as a quantitative prognostic tool in colon cancer.\u003c/p\u003e","manuscriptTitle":"Consensus molecular subtyping in colon cancer: from the design of in-house assay to prognostic nomogram","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-04 08:29:09","doi":"10.21203/rs.3.rs-9178611/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2026-05-15T09:57:17+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-05T23:09:18+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"264299152576923589618513542292492787543","date":"2026-05-03T13:46:33+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"216872876738553231686923622537875784877","date":"2026-04-25T12:00:43+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-04-21T19:49:33+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-04-20T08:17:51+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-03-27T13:37:49+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-27T09:54:46+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Gastroenterology","date":"2026-03-27T09:45:33+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-gastroenterology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmge","sideBox":"Learn more about [BMC Gastroenterology](http://bmcgastroenterol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmge/default.aspx","title":"BMC Gastroenterology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"820b691f-60ce-4954-9128-1773798e1067","owner":[],"postedDate":"May 4th, 2026","published":true,"recentEditorialEvents":[{"type":"editorInvitedReview","content":"","date":"2026-05-15T09:57:17+00:00","index":70,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-05T23:09:18+00:00","index":67,"fulltext":""},{"type":"reviewerAgreed","content":"264299152576923589618513542292492787543","date":"2026-05-03T13:46:33+00:00","index":65,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-04T08:29:09+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-04 08:29:09","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9178611","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9178611","identity":"rs-9178611","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.