Single cell sequencing combined with transcriptome sequencing to build a prognostic model for colorectal cancer | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Single cell sequencing combined with transcriptome sequencing to build a prognostic model for colorectal cancer Dan Yang, Fengsheng Dai, Xianghua Zeng, Jianhua Wang, Lijuan Zhao, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6455265/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Introduction: Colorectal cancer (CRC) is a ubiquitous malignant tumor, ranking third in the global incidence and second in mortality. Despite advancements in screening techniques like colonoscopies, the lack of trustworthy biomarkers makes it difficult to diagnose CRC accurately. The objective of this study is to construct a prognostic risk model of CRC for accurate prediction of CRC results. Materials and Methods: We obtained GSE231559 data set from GEO database for single cell sequencing analysis, selected marker genes of endothelial cells with the highest contribution to CRC to construct a prognosis model, determined 16 model genes intimately associated with the survival of patients and risk scores that can predict the prognosis of CRC patients, and verified the performance of the model through external verification queues GSE17537 and GSE87211. Then, the molecular mechanism of the prognosis model was discussed from the aspects of immune infiltration, mutation map, immunotherapy and signal pathways involved in risk score. Results: The findings indicated that the immune system of high-risk patients had poor response ability to tumors, and the tumor microenvironment was more conducive to immune escape, leading to tumor progress. The immunotherapy effect of high-risk patients is worse, and as compared to low-risk patients, the OS is much lower. The enriched pathways are related to tumor growth, invasion, metastasis and the use of immune escape mechanism, which may provide powerful clues for the prognosis of high-risk patients or malignant tumors. Discussion: We identified a total of 16 model genes that are expected to serve as new CRC biomarkers. However, the biological functions and potential regulatory mechanisms of these genes should be confirmed through experiments in the future. Conclusion: Our findings demonstrate that the prognostic risk of CRC patients may be evaluated using this approach. Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Introduction Colorectal cancer (CRC) is a common global malignant tumor with a 5-year survival rate of less than 15% [ 1 , 2 ] . It comes from the colon's or the rectal mucosa's epithelial cells. [ 3 ] . The prognosis of CRC is closely related to factors such as tumor stage, pathological type, patient's age and overall health status [ 4 ] . According to recent studies, the prognosis of colorectal cancer can be significantly improved by early detection and targeted therapy [ 5 – 7 ] . Scientists are also working hard to create new diagnosis and treatment methods and explore new biomarkers in order to provide promising anti-CRC tools and methods, as well as better strategies for diagnosing and treating CRC. With the progress in sequencing technology in recent years, a growing number of new CRC biomarkers and targets have been identified [ 8 , 9 ] . As one of the emerging sequencing technologies, single-cell RNA sequencing (scRNA-seq) technology can perform population analysis of tumor cells at the single-cell resolution level, allowing researchers to have a deeper understanding of transcriptome expression profiles and also help improve Early diagnosis and precise treatment of CRC [ 10 – 12 ] . Chu et al. established a high-quality single cell map of human intestinal tissue through scRNA-seq, revealing tumor-specific cell subtypes and transcriptome changes [ 13 ] . Qian et al. integrated scRNA-seq and bulk RNA transcriptome sequencing data to reveal the heterogeneous immune landscape and key cell subsets in CRC, providing new insights and potential therapeutic targets for the immunotherapy of CRC [ 14 ] . Therefore, in order to comprehensively identify predictive biomarkers and new molecular targets for CRC gene therapy, comprehensive and in-depth analysis and utilization of single-cell transcriptome sequencing data analysis can accurately stratify and identify patients. Here, by combining a significant quantity of RNA-seq data with the publicly available scRNA-seq database, we created a novel predictive model based on genes linked to single-cell endothelial cells in colorectal cancer. We identified a total of 16 model genes (MTUS1, EMP2, NCOA7, EHD4, PTPRG, EGFL7, INSR, JAG2, PTTG1IP, CLU, CAV1, TMEM88, PODXL, IGFBP3, NPDC1, and PALMD) that are expected to serve as new CRC biomarkers. In summary, our findings demonstrate that the prognostic risk of CRC patients may be evaluated using this approach. Materials 1. Data download The TCGA, the biggest database of cancer genetic information, is available at https://portal.gdc.cancer.gov/ and contains information on SNPs, gene expression, and copy number variants, among other things. We obtained raw mRNA expression data for CRC, gathering a total of 701 samples, with 51 in the normal group and 650 in the tumor group. Download the single cell data file of GSE231559 from the NCBI GEO public database, including a total of 9 sample data with complete single cell expression profiles, including 3 normal samples and 6 tumor samples; the Series Matrix File data file and annotation file of GSE17537 It is GPL570 and includes a total of expression profile data of 55 samples; the Series Matrix File data file of GSE87211 and the annotation file is GPL13497, which includes a total of expression profile data of 196 samples. 2. Quality control Using the Seurat program, we first imported expression profiles and filtered cells according to the proportions of mitochondrial and ribosomal readings per cell, the number of genes expressed, and the total number of UMIs per cell. Three median absolute deviations (MAD) from the median are considered outliers. It is widely acknowledged that cells displaying excessively high total UMI counts and a doubled number of expressed genes are classified as doublets, while cells with increased mitochondrial and ribosomal read percentages are indicative of subpar quality, potentially approaching apoptosis or having degraded into cellular debris. Upon completing these preliminary procedures, we utilized Doublet Finder (V2.0.4) to eliminate the doublets from each sample individually, thereby concluding the cell quality control process. 3. Data standardization Utilize the Normalize Data function to standardize the dataset, employ Cell Cycle Scoring to compute the cell cycle score, identify Variable Features to detect hypervariable genes, perform PCA on the expression matrix and scale the data to standardize it and lessen the impact of the cell cycle, ribosomal genes, and mitochondrial genes on further analysis. By primarily consulting the Cell Marker database and relevant literature, augmented by automated annotation using Single R software, for cell annotation, we identify the cell types and associated marker genes found in the relevant tissue. 4. Determine the contribution of cell subpopulations to CRC We characterize the contribution of different cell subpopulations to disease by examining changes in cell number and gene expression. First, each subpopulation's distinctive genes were found. In summary, we conducted bulk differential gene expression analysis and described how the tumor group's genes differed from those of the normal group in this process. Then, we defined FC score, which measures changes in the number and expression levels of characteristic genes in biological processes. Finally, the average FC score of all distinctive genes in the cluster characterizes the contribution of various cell subpopulations to the illness. 5. Model construction and prognosis Genes linked to prognosis were found, and a prognosis-related model was further developed using lasso regression. A risk score equation for every sample was created by taking into account the expression levels of each individual gene. It was then weighted based on the predicted regression coefficient of the lasso regression analysis. Using the median risk score value as the cutoff, the samples were divided into two groups based on the risk score equation: low-risk and high-risk. 6. Immune characteristics and tumor microenvironment analysis The approach employed to assess immune infiltration in this investigation is GSV analysis, which computes an enrichment score that signifies the degree to which genes in a particular gene set are simultaneously up- or down-regulated in an individual sample. 28 immune cell types, including Treg, Th1, Th2, Th17, central memory T cells, and Tem, were assessed using the GSVA-R software program. Furthermore, the application of the ESTIMATE algorithm to determine the immune score and stromal score provides insights into the abundance of stromal and immune cell gene characteristics. 7. Immune infiltration The CIBERSORT methodology is a widely used method for determining the kinds of immune cells present in the microenvironment. This technique applies deconvolution analysis on the expression matrix of immune cell subtypes and is founded on the support vector regression approach. T cells, B cells, plasma cells, and myeloid cell subsets are among the 22 human immune cell phenotypes that may be identified using their 547 biomarkers. 8. GSEA analysis Using a predefined gene set, GSEA analysis ranks genes according to how differently they are expressed in two distinct types of samples. The preset gene set's enrichment at the top or bottom of this ranking list is then determined. In addition to exploring the molecular processes of the core genes from both sets of data, this study uses GSEA to analyze the differences in signaling pathways between the high expression cohort and the low expression cohort. The number of substitutions is established at 1000, with the substitution type designated as phenotype. 9. GSVA analysis GSVA is a non-parametric, unsupervised technique for evaluating transcriptome gene set enrichment. GSVA translates gene-level alterations into pathway-level adjustments by meticulously assessing the gene set of interest, after which it determines the biological function of the sample. This study will collect gene sets from the Molecular Signatures Database and meticulously score each gene set using the GSVA algorithm in order to assess potential biological function differences among samples. 10. Quasi-timing analysis Single-cell investigations allow the transcriptional control of very varied cell populations and the characterisation of complex physiological processes. As a result of these investigations, genes that identify certain cell subtypes, genes that indicate intermediate phases of biological processes, and genes that exist in transitional states between different cell fates have been identified. In numerous single-cell investigations, individual cells execute gene expression processes in a non-synchronous manner, with each cell representing a snapshot in time of the transcriptional process under examination. Monocle presents a methodology to sequence individual cells within pseudotime (pseudochronology), leveraging the asynchronous activities of individual cells to position them along trajectories aligned with biological processes such as cell differentiation. 11. Statistical analysis R (version 4.3.0) was used for all statistical analyses, and a significance level of p < 0.05 was applied. Results 1.Quality control We initially imported expression profiles using the Seurat program to filter cells by the number of UMIs per cell, the number of expressed genes, the percentage of mitochondrial readings per cell, and the percentage of ribosomal reads per cell. Where nFeature_RNA indicates the number of genes and nCount_RNA indicates the total number of UMIs for the cell, outliers are defined as three median absolute deviations (MAD) from the median. Less than 200 genes will be extracted from cells. Filtered violin and scatter plots are presented (Fig S1 A-B). 2. Data standardization Subsequently, employ the Normalize Data function to standardize the dataset, utilize Cell Cycle Scoring to compute the cell cycle score, and identify 2000 hypervariable genes (Fig S1 C) through the Find Variable Features method. To normalize the dataset and exclude cell cycle pairs, ribosomal genes, and mitochondrial genes for additional analysis, use Scale Data. To assess the impact, execute PCA for linear dimensionality reduction on the matrix, selecting 20 principal components (Fig S1 D-E), and implement Harmony to mitigate batch effects. In order to preserve batch variability within each cluster, this procedure iteratively groups comparable cells from various batches inside the PCA space (Fig S1 F). Finally, use unified manifold approximation and projection, or UMAP, to reduce nonlinear dimensionality., use Find Neighbors to identify neighboring cell points, and apply Find Clusters to categorize cells into distinct clusters (Fig. 1 A). 3. Contribution of cell annotation and cell subpopulations to CRC This research additionally classified each subtype, with all clusters designated into nine cellular categories: Epithelial cells, CD4 + T lymphocytes, Monocytes, CD8 + T lymphocytes, B lymphocytes, Fibroblasts, Plasma cells, Endothelial cells, and Dendritic cells (Fig. 1 B). The traditional markers for the nine cell types are displayed in a bubble plot (Fig. 1 C). Next, by taking into account variations in cell counts and gene expression, we were able to define the contribution of several cell subpopulations to CRC. We selected differential genes from the tumor group to the normal group to describe the changes in this process. Then FC score was defined, which is used to measure changes in the number and expression levels of characteristic genes in biological processes. Finally, it was found that Endothelial cells contribute the most to CRC (Fig. 1 D). 4. Construction of prognostic model We gathered clinical data from CRC samples within the TCGA database and employed marker genes of endothelial cells to identify prognostic genes in CRC via Cox univariate regression. The findings indicated that a total of 28 genes linked to prognosis were found (p value < 0.05) (Fig. 2 A). Based on the prognostic genes, unique genes in colorectal cancer were extracted using the lasso regression feature selection approach. We used lasso regression analysis to get each sample after randomly dividing the processed CRC dataset, which contained survival data samples from the TCGA database, into training and testing sets at a 4:1 ratio (Fig. 2 B-D). The corresponding optimal risk score value is used for subsequent analysis. RiskScore = MTUS1×(-0.174441201921624) + EMP2×(-0.107481846357423) + NCOA7×(-0.082921817145568) + EHD4×(-0.078686080062342) + PTPRG×(-0.0553784848516573) + EGFL7×0.00453034652452583 + INSR×0.0161439754146179 + JAG2×0.0306554904877961 + PTTG1IP×0.0421245179315935 + CLU×0.0518270906986344 + CAV1×0.052522780662374 + TMEM88×0.0620618376532079 + PODXL×0.0740276615814645 + IGFBP3×0.0911907751517928 + NPDC1×0.105466421387826 + PALMD×0.132677159075287.The samples were divided into high-risk and low-risk groups according to risk ratings, and Kaplan-Meier curves were employed for analysisThe OS of the high-risk group was significantly poorer than that of the low-risk group in both the training and test sets (Fig. 2 E-F). Furthermore, both the training and test sets' ROC curve findings show that the model performs well in verification (Fig. 2 G-F). 5. Robustness of external validation set Download the processed CRC samples accompanied by survival data from the GEO database (GSE17537, GSE87211), forecast the clinical classification of the CRC samples within the GEO database utilizing the model, and assess the survival disparity between the two groups through Kaplan-Meier analysis. The findings demonstrated that the OS of the high-risk group was noticeably lower than the low-risk group (Fig. 3 A-B). We used external datasets to perform ROC curve analysis on the model in order to verify its correctness. It was shown by the findings that the model was highly predictive in predicting the samples' prognosis (Fig. 3 C-D). 6.Immune characteristics and tumor microenvironment analysis This investigation further examined whether high- and low-risk cohorts exhibit distinct immune profiles. Initially, the immune response was quantified utilizing the ssGSEA algorithm (Fig. 4 A), followed by a differential analysis aimed at uncovering risk score-specific immune signatures. The findings indicated that the immune characteristics were distinctly categorized between the risk score groups, with immune features such as Tem, TFH, TReg, and Tgd in the high-risk score group being significantly elevated compared to those in the low-risk score group (Fig. 4 B). Additionally, compared to the low-risk score group, the high-risk score group showed higher activation-related scores for the HIPPO, WNT, and Angiogenesis pathways (Fig. 4 C). The scientists then calculated immunological and stromal scores using the ESTIMATE method, which showed differences between the two groups' results (Fig. 4 D-E). 7. Examination of the connection between immune infiltration and risk score The tumor microenvironment is mostly composed of tumor-associated fibroblasts, extracellular matrix, immune cells, unique physical, inflammatory mediators and chemical properties, different growth factors, and the cancer cells themselves. This microenvironment profoundly influences tumor diagnosis and survival, as well as clinical treatment responsiveness. We further explore the possible biological pathways via which risk score influences the development of CRC by investigating the relationship between risk score and tumor immune infiltration. In a variety of ways, we display the immune cell composition and correlations between high- and low-risk groups (Fig. 5 A-B). Samples from the high-risk group showed considerably lower levels of resting and activating CD 4+ memory T cells than the low-risk group, but M0 macrophages and regulatory T cells (Tregs) showed much higher levels (Fig. 5 C). We then looked into the relationship between immune cells and risk score. Among other things, the results showed that risk score had a substantial negative association with resting CD 4 + memory T cells and activated CD 4 + memory T cells, but a significant positive correlation with M2 and M0 macrophages (Fig. 5 D). 8. Variations in mutations between populations at high and low risk The CRC SNP-related data was processed, the TOP30 genes with the greatest mutation frequency were chosen for presentation, the differences in mutated genes between the two sample groups were compared, and the mutation landscape was displayed using the R package Complex Heatmap (Fig. 6 ). According to the results, samples from the high-risk group had a substantially lower percentage of APC and other gene mutations than samples from the low-risk group. Subsequently, we predicted both high- and low-risk groups' sensitivity to anti-tumor immunotherapy, and the findings showed that the high-risk group responded less well to immunotherapy (Fig. 7 A). 9. Pathway Enrichment Analysis Next, the molecular mechanisms via which risk scores influence tumor growth were investigated, along with the specific signaling pathways implicated in the high- and low-risk correlation models. GSEA findings indicate that the relevant pathways include KEGG AMINOACYL TRNA BIOSYNTHESIS, KEGG GLYCOSAMINOGLYCAN BIOSYNTHESIS CHONDROITIN SULFATE, KEGG CITRATE CYCLE TCA CYCLE, among others (Fig. 7 B). The two sample groups' unique pathways were mostly enriched in signaling pathways such MYOGENESIS, KRAS_SIGNALING_DN, and EPITHELIAL MESENCHYMAL TRANSITION, according to the GSVA data (Fig. 7 C). 10. Expression of model genes and immunometabolism pathways We examined how model genes were expressed in individual cells and demonstrated how they were expressed in B cells, fibroblasts, plasma cells, endothelial cells, dendritic cells, monocytes, CD4 + T cells, epithelial cells, and CD8 + T cells (Fig. 8 A-B). Finally, we quantitatively scored the immunometabolism-related pathway genes and used bubble charts to display the activity differences of the model genes in the immunometabolism-related pathways. The results showed that the 16 model genes had higher activity in the angiogenesis, epithelial mesenchymal transition, and TGF-β signaling pathways (Fig. 8 C). Finally, the developmental trajectory of model genes in Endothelial cells is displayed (Fig. 9 A-D). Discussion One of the most prevalent malignant tumors worldwide is colorectal cancer [ 2 ] . Its incidence and mortality vary from region to region, and the incidence in developed countries is usually higher than that in developing countries [ 3 ] . There are many risk factors affecting the occurrence of CRC, including hereditary diseases (such as familial adenomatosis and Lynch syndrome), bad eating habits, high-fat and low-fiber diet, obesity and age [ 3 , 5 ] . Early screening is particularly important, which can effectively reduce the mortality of CRC and promote early diagnosis [ 10 ] . In addition, individualized treatment, metastasis monitoring and nutritional and psychological support can effectively improve the therapeutic effect. In addition, regular imaging examination can help monitor tumor recurrence or metastasis, so as to adjust the treatment plan in time. Meanwhile, the development of a treatment and prognosis model for colorectal cancer is clinically significant. This model can not only improve the prediction ability of patients' disease progress and survival probability, but also optimize the treatment plan, provide personalized care for patients with different risk levels and rationally allocate medical resources [ 13 ] . In addition, the prognosis model also provides data support for basic and clinical research, promotes the research and development of new therapies, enhances patients' sense of education and participation, and improves their compliance and satisfaction with the treatment process. Here, we selected the marker genes of endothelial cells with the highest contribution to CRC from the single cell sequencing analysis by obtaining GSE231559 data set, and identified 16 model genes that are strongly associated with patient survival and the risk score that may be used to forecast the prognosis of patients with colorectal cancer. At the same time, the performance of the model was verified by external verification queues GSE17537 and GSE87211. Among 16 genes, it has been reported MTUS1, EGFL7, INSR, CLU, CAV1, TMEM88, PODXL, and IGFBP3 could use to be CRC biomarkers [ 15 – 23 ] . These studies once again confirm the importance of our model genes In recent years, the research on immune infiltration of CRC has made remarkable progress, revealing the important role of immune cells in tumor microenvironment. Many types of immune cells, including CD8 + T lymphocytes, B cells and macrophages, can be detected in colon cancer tissues, and their abundance and activity are closely related to the development, prognosis and response to treatment of the tumor [ 24 , 25 ] . For example, a high level of CD8 + T cell infiltration is usually associated with a good clinical prognosis, while tumor-associated macrophages are often associated with enhanced tumor invasion [ 25 ] . In addition, with the application of single cell RNA sequencing and other technologies, researchers can evaluate the characteristics of immune microenvironment more accurately, which provides new opportunities for individualized treatment [ 26 – 29 ] . Therefore, immune infiltration is not only an important factor in understanding the biology of CRC, but also has the potential as a prognostic marker and guiding treatment decision. According to our findings, the high-risk group's T cells' CD4 memory resting and activation are much lower than those of the low-risk group; The Macrophages M0 and Tregs were higher. Then we discussed the relationship between risk score and immune cells. The findings indicated that the risk score was positively correlated with MacrophagesM2 and MacrophagesM0, and inversely connected with T cells CD4 memory resting and TcellsCD4 memory activated. In a word, these results are very helpful to further investigate the mechanism of immune infiltration and optimize the immunotherapy strategy for CRC in the future. However, this study also has limitations. First, the biological functions and potential regulatory mechanisms of these genes will be confirmed through experiments in the future. Secondly, we will collect information on clinical samples in our hospital and once again highlight the accuracy and importance of this model, in order to provide reliable targets for early diagnosis of clinical CRC and provide reference for the formulation of personalized treatment plans for CRC. Declarations Funding This work was supported by the Chongqing Natural Science Foundation [grant number CSTB2022NSCQ-MSX1116] and the Health Commission of Chongqing Scientific Research Project [grant number 2023MSXM173]. Data availability The datasets used and/or analyzed during the current study are available from the corresponding author on reasonable request. Compliance Statement Clinical trial number Not applicable. Consent to Publish declaration Not applicable. Consent to Participate declaration Not applicable. Ethics declaration Not applicable. Author Contribution Dan Yang conceptualized the study, designed the methodology, and wrote the original draft.Fengsheng Dai performed formal data analysis, curated datasets, and contributed to manuscript revisions.Xianghua Zeng developed visualization tools, validated experimental results, and assisted in data interpretation.Jianhua Wang secured funding, supervised project administration, and critically reviewed the manuscript.Lijuan Zhao conducted investigations, organized resources, and participated in editing the final draft.Tingxiu Xiang contributed to statistical validation, reviewed the literature, and edited technical sections.Zhongjun Wu oversaw the research framework, coordinated interdisciplinary collaboration, and approved the final manuscript. References Patel SG, Karlitz JJ, Yen T, Lieu CH, Boland CR. The rising tide of early-onset colorectal cancer: a comprehensive review of epidemiology, clinical features, biology, risk factors, prevention, and early detection. Lancet Gastroenterol Hepatol. 2022;7(3):262–74. Yan H, Talty R, Johnson CH. Targeting ferroptosis to treat colorectal cancer. Trends Cell Biol. 2023;33(3):185–8. Abedizadeh R, Majidi F, Khorasani HR, Abedi H, Sabour D. Colorectal cancer: a comprehensive review of carcinogenesis, diagnosis, and novel strategies for classified treatments. Cancer Metastasis Rev. 2024;43(2):729–53. Williams H, Lee C, Garcia-Aguilar J. Nonoperative management of rectal cancer. Front Oncol. 2024;14:1477510. Andrei P, Battuello P, Grasso G, Rovera E, Tesio N, Bardelli A. Integrated approaches for precision oncology in colorectal cancer: The more you know, the better. Semin Cancer Biol. 2022;84:199–213. Hu H, Wang Q, Yu D et al. Berberine Derivative B68 Promotes Tumor Immune Clearance by Dual-Targeting BMI1 for Senescence Induction and CSN5 for PD-L1 Degradation. Adv Sci (Weinh). 2024: e2413122. Wang N, Fang JY. Fusobacterium nucleatum, a key pathogenic factor and microbial biomarker for colorectal cancer. Trends Microbiol. 2023;31(2):159–72. Lin Q, Wang Z, Wang J, et al. Innovative strategies to optimise colorectal cancer immunotherapy through molecular mechanism insights. Front Immunol. 2024;15:1509658. Fan X, Li B, Zhang F et al. FGF19-Activated Hepatic Stellate Cells Release ANGPTL4 that Promotes Colorectal Cancer Liver Metastasis. Adv Sci (Weinh). 2024: e2413525. Zhu AK, Li GY, Chen FC, et al. Integrated Analysis of Single-Cell and Bulk RNA-Sequencing Based on EcoTyper Machine Learning Framework Identifies Cell-State-Specific M2 Macrophage Markers Associated with Gastric Cancer Prognosis. Immunotargets Ther. 2024;13:721–34. Hu C, Huang X, Chen J, et al. Dissecting the cellular reprogramming and tumor microenvironment in left- and right-sided Colorectal Cancer by single cell RNA sequencing. Transl Res. 2024;276:22–37. Ahluwalia P, Mondal AK, Vashisht A, et al. Identification of a distinctive immunogenomic gene signature in stage-matched colorectal cancer. J Cancer Res Clin Oncol. 2024;151(1):9. Chu X, Li X, Zhang Y, et al. Integrative single-cell analysis of human colorectal cancer reveals patient stratification with distinct immune evasion mechanisms. Nat Cancer. 2024;5(9):1409–26. Zhang Q, Liu Y, Wang X, Zhang C, Hou M, Liu Y. Integration of single-cell RNA sequencing and bulk RNA transcriptome sequencing reveals a heterogeneous immune landscape and pivotal cell subpopulations associated with colorectal cancer prognosis. Front Immunol. 2023;14:1184167. Cheng LY, Huang MS, Zhong HG, et al. MTUS1 is a promising diagnostic and prognostic biomarker for colorectal cancer. World J Surg Oncol. 2022;20(1):257. Juan Z, Dake C, Tanaka K, Shuixiang H. EGFL7 as a novel therapeutic candidate regulates cell invasion and anoikis in colorectal cancer through PI3K/AKT signaling pathway. Int J Clin Oncol. 2021;26(6):1099–108. de Oliveira C, Martins S, Gonçalves PG, et al. Low EGFL7 expression is associated with high lymph node spread and invasion of lymphatic vessels in colorectal cancer. Sci Rep. 2023;13(1):19783. Landi D, Moreno V, Guino E et al. Polymorphisms affecting micro-RNA regulation and associated with the risk of dietary-related cancers: a review from the literature and new evidence for a functional role of rs17281995 (CD86) and rs1051690 (INSR), previously associated with colorectal cancer. Mutat Res. 2011. 717(1–2): 109 – 15. He W, Chan CM, Wong SC, et al. Jagged 2 silencing inhibits motility and invasiveness of colorectal cancer cell lines. Oncol Lett. 2016;12(6):5193–8. Artemaki PI, Sklirou AD, Kontos CK, et al. High clusterin (CLU) mRNA expression levels in tumors of colorectal cancer patients predict a poor prognostic outcome. Clin Biochem. 2020;75:62–9. Peng Y, Zhao J, Yin F, et al. A methylation-driven gene panel predicts survival in patients with colon cancer. FEBS Open Bio. 2021;11(9):2490–506. Lee H, Kong JS, Lee SS, Kim A. Radiation-Induced Overexpression of TGFβ and PODXL Contributes to Colorectal Cancer Cell Radioresistance through Enhanced Motility. Cells. 2021. 10(8): 2087. Perez-Carbonell L, Balaguer F, Toiyama Y, et al. IGFBP3 methylation is a novel diagnostic and predictive biomarker in colorectal cancer. PLoS ONE. 2014;9(8):e104285. Picard E, Verschoor CP, Ma GW, Pawelec G. Relationships Between Immune Landscapes, Genetic Subtypes and Responses to Immunotherapy in Colorectal Cancer. Front Immunol. 2020;11:369. Giannakis M, Mu XJ, Shukla SA, et al. Genomic Correlates of Immune-Cell Infiltrates in Colorectal Carcinoma. Cell Rep. 2016;15(4):857–65. Masuda K, Kornberg A, Miller J, et al. Multiplexed single-cell analysis reveals prognostic and nonprognostic T cell types in human colorectal cancer. JCI Insight. 2022;7(7):e154646. Li J, Wu C, Hu H, et al. Remodeling of the immune and stromal cell compartment by PD-1 blockade in mismatch repair-deficient colorectal cancer. Cancer Cell. 2023;41(6):1152–e11697. Wang R, Li J, Zhou X, et al. Single-cell genomic and transcriptomic landscapes of primary and metastatic colorectal cancer tumors. Genome Med. 2022;14(1):93. Mei Y, Xiao W, Hu H, et al. Single-cell analyses reveal suppressive tumor microenvironment of human colorectal cancer. Clin Transl Med. 2021;11(6):e422. Additional Declarations No competing interests reported. Supplementary Files supplyment.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6455265","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":472589678,"identity":"5009668e-d003-4ac4-90fe-bf0a04f4136a","order_by":0,"name":"Dan Yang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABA0lEQVRIiWNgGAWjYDCCAwwMEkhsmwQwK6GAeC1pCQxsIC0GRGoBgsMQLQx4tPAdP3vwxs8dh+XM23sPHvhRcz6PX7478cMDAwZ5frEDWLVInslLtuw9c9hY5sy5hIM9x24XS7bxbpYAOsxw5uwErFoMDuSYSfC2pSXOkMgxOMzAdjtxwzHeDSAtCQa3cWg5/8ZM8m9bWj1Ey79zIC2bf+DVciPHTJq3zSZBAqSFse0ASMs2vLZI3nhjbC3bZmM4g+eMwcHevuTEmW252ywSDCRw+oXvfI7hzbdtEvIS7D3GH358s0vsZz67+eaPCht5fmnsWnACCcJKRsEoGAWjYBTgBAAnLmN+2gMkvQAAAABJRU5ErkJggg==","orcid":"","institution":"The First Affiliated Hospital of Chongqing Medical University","correspondingAuthor":true,"prefix":"","firstName":"Dan","middleName":"","lastName":"Yang","suffix":""},{"id":472589679,"identity":"850d0805-6c64-4f95-8172-325af11229e7","order_by":1,"name":"Fengsheng Dai","email":"","orcid":"","institution":"Chongqing University Cancer Hospital","correspondingAuthor":false,"prefix":"","firstName":"Fengsheng","middleName":"","lastName":"Dai","suffix":""},{"id":472589680,"identity":"6d544741-af45-41a7-9031-58601699c8bc","order_by":2,"name":"Xianghua Zeng","email":"","orcid":"","institution":"Chongqing University Cancer Hospital","correspondingAuthor":false,"prefix":"","firstName":"Xianghua","middleName":"","lastName":"Zeng","suffix":""},{"id":472589681,"identity":"d167a17c-cfc3-4a32-a626-7d84bb8367bd","order_by":3,"name":"Jianhua Wang","email":"","orcid":"","institution":"the First Affiliated Hospital of Dali University","correspondingAuthor":false,"prefix":"","firstName":"Jianhua","middleName":"","lastName":"Wang","suffix":""},{"id":472589682,"identity":"4ae62e19-56f9-477f-a630-c700b6fcdeee","order_by":4,"name":"Lijuan Zhao","email":"","orcid":"","institution":"The Affiliated Yongchuan Hospital of Chongqing Medical University,Chongqing","correspondingAuthor":false,"prefix":"","firstName":"Lijuan","middleName":"","lastName":"Zhao","suffix":""},{"id":472589683,"identity":"e2fd7bea-f046-44ed-a3c3-c33a6cd2238e","order_by":5,"name":"Tingxiu Xiang","email":"","orcid":"","institution":"Chongqing University Cancer Hospital","correspondingAuthor":false,"prefix":"","firstName":"Tingxiu","middleName":"","lastName":"Xiang","suffix":""},{"id":472589684,"identity":"6918fdad-90ce-48b5-b108-d23fc193324b","order_by":6,"name":"Zhongjun Wu","email":"","orcid":"","institution":"The First Affiliated Hospital of Chongqing Medical University","correspondingAuthor":false,"prefix":"","firstName":"Zhongjun","middleName":"","lastName":"Wu","suffix":""}],"badges":[],"createdAt":"2025-04-15 13:23:28","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6455265/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6455265/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":85171189,"identity":"510e5846-8345-4c80-a504-0aead750fa1c","added_by":"auto","created_at":"2025-06-23 05:40:30","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1720336,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAnnotation of Cells.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) The UMAP visualization illustrates the outcomes of cellular clustering (based on RNA expression, with a resolution of 0.2). Each point symbolizes a cell, while distinct colors differentiate various clusters. (B) The annotation results for each type of cell are shown in the UMAP diagram. Cells are then divided into several cell kinds. The distribution of these different cell types is shown by a variety of colors and labels. (C) The bubble graphic shows the patterns of marker gene expression in different cell types. Genes are shown on the x-axis, while cell types are shown on the y-axis. Each bubble's color shows the average gene expression (blue denotes low expression, red denotes high expression), while its size indicates the proportion of cells in the associated cell type (% expressed). (D) Pie charts and bar graphs illustrate the contribution of various cell groups to colorectal cancer.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6455265/v1/987c4a045356363d339124e6.jpeg"},{"id":85171186,"identity":"b9dd31fb-ecf0-4690-8930-ee2877dc268c","added_by":"auto","created_at":"2025-06-23 05:40:30","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1130900,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eConstruction of CRC prognostic model.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) Prognostic gene single factor forest map. (B) LASSO regression coefficient path diagram. Lines with different colors represent different variables. With the increase of the penalty term λ, the coefficient of unimportant variables will soon become 0, while the more important variables, the less influence of the penalty term on the coefficient of variables, the more they can stay in the end. The characteristic variables with 16 non-zero coefficients are the variables selected by Lasso regression. (C) Cross-validation curve of LASSO regression. The number of variables needed for the relevant model is displayed by the abscissa at the top, which progressively drops from left to right; The logarithm of the penalty coefficient λ is the abscissa below; The ordinate represents mean squared error (MSE), which is the discrepancy between the computation model's actual value and its anticipated value. After dimensionality reduction, 16 variables with non-zero coefficient characteristics were selected for modeling. (D) The histogram shows the coefficients of model genes and the size of log2(HR). (E-F) Patients in the low-risk category have a better prognosis, and the training set tests the KM survival curves of the high-risk and low-risk groups. (G-H) The training set tests the ROC curve, and the classification ability of the model is good.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6455265/v1/aeb594e12a47df2652be65ad.jpeg"},{"id":85171937,"identity":"f2db270a-06b2-412f-bcea-3196089f9bbc","added_by":"auto","created_at":"2025-06-23 05:48:30","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":777087,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eRobustness of external verification.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A-B) Verify the KM survival curve of high and low risk groups. (C-D) ROC curve of verification set.\u003c/p\u003e","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6455265/v1/3b3a7aaa0cda7993e530231b.jpeg"},{"id":85171193,"identity":"0064db54-c066-4b6e-926b-d11f5a1a6cde","added_by":"auto","created_at":"2025-06-23 05:40:30","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":2892206,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAnalysis of immune characteristics and tumor microenvironment.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) Thermogram of correlation between clinical features of colorectal cancer and 28 immune features. (B) Risk score and differential expression of 28 immune-related cells. (C) Differential expression of risk score and pathway. (D-E) High and low risk scores differ in terms of immunity and matrix scores.\u003c/p\u003e","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6455265/v1/a6a0faed5c8bfac25d290d48.jpeg"},{"id":85171198,"identity":"23719e60-a31d-46bd-a4c6-2118851e007f","added_by":"auto","created_at":"2025-06-23 05:40:30","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":2430324,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eImmune infiltration.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) The proportions of 22 distinct immune cell subsets are compared between patients in high-risk and low-risk groups. (B) A correlation thermogram illustrates the relationships among the different immune cell populations, with red indicating positive correlations and blue representing negative correlations, as quantified by correlation coefficient values. (C) The immune cell makeup of the high-risk and low-risk cohorts is shown differently; the low-risk cohort is shown in blue, while the high-risk cohort is shown in pink. (D) The correlation between the immune cell composition and the risk score is illustrated, highlighting statistical significance (P value), while the size and color of the data points reflect both the strength of the correlation and the level of significance.\u003c/p\u003e","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6455265/v1/4d08f30c91dc39d0ea6c9970.jpeg"},{"id":85171190,"identity":"16275dfb-91ef-4d1d-b9b2-b7b8e58097cd","added_by":"auto","created_at":"2025-06-23 05:40:30","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":176329,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCorrelation mutation map of risk scoring model.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) CRC SNP-related data. A high-mutation-frequency TOP30 gene was chosen for display, and the variations in mutant genes between high- and low-risk groups were contrasted.\u003c/p\u003e","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-6455265/v1/c2220fa6589ae51b66e727b1.png"},{"id":85171192,"identity":"18b2a1d7-eb2a-4eb5-abce-5a8abcaea0d4","added_by":"auto","created_at":"2025-06-23 05:40:30","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":121713,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePrediction of immunotherapy response.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) The analysis illustrates the disparity in immunotherapy responses between two cohorts, indicating that individuals classified in the high-risk group exhibit a poorer reaction to immunotherapeutic interventions. (B) The signaling pathways linked to the risk score are displayed from the KEGG. (C) The signaling pathways linked to the risk score are depicted, with blue representing those associated with highly expressed genes and green denoting pathways related to lowly expressed genes, while the underlying gene set is derived from hallmark genes.\u003c/p\u003e","description":"","filename":"Onlinefloatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-6455265/v1/d83cde3a74a616a51b2f716c.png"},{"id":85171941,"identity":"a29e2ac3-8177-4eec-8896-b8e2dcb99211","added_by":"auto","created_at":"2025-06-23 05:48:31","extension":"jpeg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":695066,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGeneral situation of model gene expression.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) The UMAP visualization illustrates the distribution of expression levels for 16 selected model genes within single-cell transcriptomic data. The spatial positioning and intensity variation of distinct genes throughout the cellular population suggest their potential roles that may be specific to particular cell types. (B) The bubble plot represents the expression levels of model genes, with the horizontal axis denoting the genes and the vertical axis indicating various cell types; the size of each bubble corresponds to the gene expression ratio, while the color reflects the average expression level, with deep red indicating high expression. The findings indicate a pronounced expression of these model genes in endothelial cells. (C) The pathway enrichment analysis bubble diagram shows the substantial enrichment results of model genes across different functional pathways: the associated genes are displayed on the vertical axis, while the functional pathways are listed on the horizontal axis. The size of the bubbles reflects the significance level of the enrichment (FDR value), and the color represents the direction of the enrichment effect (with red indicating up-regulation and blue indicating down-regulation); additional classification labels further categorize the pathways into groups such as metabolism, immunity, signaling, and others.\u003c/p\u003e","description":"","filename":"floatimage8.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6455265/v1/2e5d77d4b6843fb16530772d.jpeg"},{"id":85171942,"identity":"c8ac125b-0630-472f-b917-7be31ae7b8b7","added_by":"auto","created_at":"2025-06-23 05:48:31","extension":"jpeg","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":2701074,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eQuasi-time series analysis.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A-B) Picture of cell differentiation trajectory constructed according to quasi-time, with dark blue indicating the initial stage of cell differentiation. (C) Branching thermogram showed different expression patterns of genes. The pre-branch of CellType is the data before the branch point, the Cellfate1 is the smaller side of the state after the branch point, and the Cellfate2 is the larger side of the state after the branch point, which is divided into 6 cluster by default according to the genetic changes. (D) displaying the gene expression changes of the model genes from the beginning to the end of the pseudo-time process.\u003c/p\u003e","description":"","filename":"floatimage9.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6455265/v1/0afecf36706d09c61de4b668.jpeg"},{"id":91616731,"identity":"2f1bf35f-630b-41df-a08b-8ca4df13a305","added_by":"auto","created_at":"2025-09-18 10:40:22","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":13833358,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6455265/v1/58c140e1-05b6-4b46-b131-c342d1410bdb.pdf"},{"id":85171938,"identity":"5e0be85f-f02a-4444-b386-4fb8596d01ab","added_by":"auto","created_at":"2025-06-23 05:48:30","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":719620,"visible":true,"origin":"","legend":"","description":"","filename":"supplyment.docx","url":"https://assets-eu.researchsquare.com/files/rs-6455265/v1/caccd3de4dbbf763474ccadd.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Single cell sequencing combined with transcriptome sequencing to build a prognostic model for colorectal cancer","fulltext":[{"header":"Introduction","content":"\u003cp\u003eColorectal cancer (CRC) is a common global malignant tumor with a 5-year survival rate of less than 15%\u003csup\u003e[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]\u003c/sup\u003e. It comes from the colon's or the rectal mucosa's epithelial cells.\u003csup\u003e[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]\u003c/sup\u003e. The prognosis of CRC is closely related to factors such as tumor stage, pathological type, patient's age and overall health status\u003csup\u003e[\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]\u003c/sup\u003e. According to recent studies, the prognosis of colorectal cancer can be significantly improved by early detection and targeted therapy\u003csup\u003e[\u003cspan additionalcitationids=\"CR6\" citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e–\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]\u003c/sup\u003e. Scientists are also working hard to create new diagnosis and treatment methods and explore new biomarkers in order to provide promising anti-CRC tools and methods, as well as better strategies for diagnosing and treating CRC.\u003c/p\u003e \u003cp\u003eWith the progress in sequencing technology in recent years, a growing number of new CRC biomarkers and targets have been identified \u003csup\u003e[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]\u003c/sup\u003e. As one of the emerging sequencing technologies, single-cell RNA sequencing (scRNA-seq) technology can perform population analysis of tumor cells at the single-cell resolution level, allowing researchers to have a deeper understanding of transcriptome expression profiles and also help improve Early diagnosis and precise treatment of CRC \u003csup\u003e[\u003cspan additionalcitationids=\"CR11\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e–\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]\u003c/sup\u003e. Chu et al. established a high-quality single cell map of human intestinal tissue through scRNA-seq, revealing tumor-specific cell subtypes and transcriptome changes\u003csup\u003e[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]\u003c/sup\u003e. Qian et al. integrated scRNA-seq and bulk RNA transcriptome sequencing data to reveal the heterogeneous immune landscape and key cell subsets in CRC, providing new insights and potential therapeutic targets for the immunotherapy of CRC\u003csup\u003e[\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]\u003c/sup\u003e. Therefore, in order to comprehensively identify predictive biomarkers and new molecular targets for CRC gene therapy, comprehensive and in-depth analysis and utilization of single-cell transcriptome sequencing data analysis can accurately stratify and identify patients.\u003c/p\u003e \u003cp\u003eHere, by combining a significant quantity of RNA-seq data with the publicly available scRNA-seq database, we created a novel predictive model based on genes linked to single-cell endothelial cells in colorectal cancer. We identified a total of 16 model genes (MTUS1, EMP2, NCOA7, EHD4, PTPRG, EGFL7, INSR, JAG2, PTTG1IP, CLU, CAV1, TMEM88, PODXL, IGFBP3, NPDC1, and PALMD) that are expected to serve as new CRC biomarkers. In summary, our findings demonstrate that the prognostic risk of CRC patients may be evaluated using this approach.\u003c/p\u003e "},{"header":"Materials","content":"\u003ch3\u003e1. Data download\u003c/h3\u003e\u003cp\u003eThe TCGA, the biggest database of cancer genetic information, is available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://portal.gdc.cancer.gov/\u003c/span\u003e\u003cspan address=\"https://portal.gdc.cancer.gov/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e and contains information on SNPs, gene expression, and copy number variants, among other things. We obtained raw mRNA expression data for CRC, gathering a total of 701 samples, with 51 in the normal group and 650 in the tumor group. Download the single cell data file of GSE231559 from the NCBI GEO public database, including a total of 9 sample data with complete single cell expression profiles, including 3 normal samples and 6 tumor samples; the Series Matrix File data file and annotation file of GSE17537 It is GPL570 and includes a total of expression profile data of 55 samples; the Series Matrix File data file of GSE87211 and the annotation file is GPL13497, which includes a total of expression profile data of 196 samples.\u003c/p\u003e\u003ch2\u003e2. Quality control\u003c/h2\u003e\u003cp\u003eUsing the Seurat program, we first imported expression profiles and filtered cells according to the proportions of mitochondrial and ribosomal readings per cell, the number of genes expressed, and the total number of UMIs per cell. Three median absolute deviations (MAD) from the median are considered outliers. It is widely acknowledged that cells displaying excessively high total UMI counts and a doubled number of expressed genes are classified as doublets, while cells with increased mitochondrial and ribosomal read percentages are indicative of subpar quality, potentially approaching apoptosis or having degraded into cellular debris. Upon completing these preliminary procedures, we utilized Doublet Finder (V2.0.4) to eliminate the doublets from each sample individually, thereby concluding the cell quality control process.\u003c/p\u003e\u003ch3\u003e3. Data standardization\u003c/h3\u003e\u003cp\u003eUtilize the Normalize Data function to standardize the dataset, employ Cell Cycle Scoring to compute the cell cycle score, identify Variable Features to detect hypervariable genes, perform PCA on the expression matrix and scale the data to standardize it and lessen the impact of the cell cycle, ribosomal genes, and mitochondrial genes on further analysis. By primarily consulting the Cell Marker database and relevant literature, augmented by automated annotation using Single R software, for cell annotation, we identify the cell types and associated marker genes found in the relevant tissue.\u003c/p\u003e\u003ch3\u003e4. Determine the contribution of cell subpopulations to CRC\u003c/h3\u003e\u003cp\u003eWe characterize the contribution of different cell subpopulations to disease by examining changes in cell number and gene expression. First, each subpopulation's distinctive genes were found. In summary, we conducted bulk differential gene expression analysis and described how the tumor group's genes differed from those of the normal group in this process. Then, we defined FC score, which measures changes in the number and expression levels of characteristic genes in biological processes. Finally, the average FC score of all distinctive genes in the cluster characterizes the contribution of various cell subpopulations to the illness.\u003c/p\u003e\u003ch3\u003e5. Model construction and prognosis\u003c/h3\u003e\u003cp\u003eGenes linked to prognosis were found, and a prognosis-related model was further developed using lasso regression. A risk score equation for every sample was created by taking into account the expression levels of each individual gene. It was then weighted based on the predicted regression coefficient of the lasso regression analysis. Using the median risk score value as the cutoff, the samples were divided into two groups based on the risk score equation: low-risk and high-risk.\u003c/p\u003e\u003ch3\u003e6. Immune characteristics and tumor microenvironment analysis\u003c/h3\u003e\u003cp\u003eThe approach employed to assess immune infiltration in this investigation is GSV analysis, which computes an enrichment score that signifies the degree to which genes in a particular gene set are simultaneously up- or down-regulated in an individual sample. 28 immune cell types, including Treg, Th1, Th2, Th17, central memory T cells, and Tem, were assessed using the GSVA-R software program. Furthermore, the application of the ESTIMATE algorithm to determine the immune score and stromal score provides insights into the abundance of stromal and immune cell gene characteristics.\u003c/p\u003e\u003ch2\u003e7. Immune infiltration\u003c/h2\u003e\u003cp\u003eThe CIBERSORT methodology is a widely used method for determining the kinds of immune cells present in the microenvironment. This technique applies deconvolution analysis on the expression matrix of immune cell subtypes and is founded on the support vector regression approach. T cells, B cells, plasma cells, and myeloid cell subsets are among the 22 human immune cell phenotypes that may be identified using their 547 biomarkers.\u003c/p\u003e\u003ch3\u003e8. GSEA analysis\u003c/h3\u003e\u003cp\u003eUsing a predefined gene set, GSEA analysis ranks genes according to how differently they are expressed in two distinct types of samples. The preset gene set's enrichment at the top or bottom of this ranking list is then determined. In addition to exploring the molecular processes of the core genes from both sets of data, this study uses GSEA to analyze the differences in signaling pathways between the high expression cohort and the low expression cohort. The number of substitutions is established at 1000, with the substitution type designated as phenotype.\u003c/p\u003e\u003ch3\u003e9. GSVA analysis\u003c/h3\u003e\u003cp\u003eGSVA is a non-parametric, unsupervised technique for evaluating transcriptome gene set enrichment. GSVA translates gene-level alterations into pathway-level adjustments by meticulously assessing the gene set of interest, after which it determines the biological function of the sample. This study will collect gene sets from the Molecular Signatures Database and meticulously score each gene set using the GSVA algorithm in order to assess potential biological function differences among samples.\u003c/p\u003e\u003ch2\u003e10. Quasi-timing analysis\u003c/h2\u003e\u003cp\u003eSingle-cell investigations allow the transcriptional control of very varied cell populations and the characterisation of complex physiological processes. As a result of these investigations, genes that identify certain cell subtypes, genes that indicate intermediate phases of biological processes, and genes that exist in transitional states between different cell fates have been identified. In numerous single-cell investigations, individual cells execute gene expression processes in a non-synchronous manner, with each cell representing a snapshot in time of the transcriptional process under examination. Monocle presents a methodology to sequence individual cells within pseudotime (pseudochronology), leveraging the asynchronous activities of individual cells to position them along trajectories aligned with biological processes such as cell differentiation.\u003c/p\u003e\u003ch2\u003e11. Statistical analysis\u003c/h2\u003e\u003cp\u003eR (version 4.3.0) was used for all statistical analyses, and a significance level of p \u0026lt; 0.05 was applied.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e1.Quality control\u003c/h2\u003e \u003cp\u003eWe initially imported expression profiles using the Seurat program to filter cells by the number of UMIs per cell, the number of expressed genes, the percentage of mitochondrial readings per cell, and the percentage of ribosomal reads per cell. Where nFeature_RNA indicates the number of genes and nCount_RNA indicates the total number of UMIs for the cell, outliers are defined as three median absolute deviations (MAD) from the median. Less than 200 genes will be extracted from cells. Filtered violin and scatter plots are presented (Fig \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eA-B).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e2. Data standardization\u003c/h2\u003e \u003cp\u003eSubsequently, employ the Normalize Data function to standardize the dataset, utilize Cell Cycle Scoring to compute the cell cycle score, and identify 2000 hypervariable genes (Fig \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eC) through the Find Variable Features method. To normalize the dataset and exclude cell cycle pairs, ribosomal genes, and mitochondrial genes for additional analysis, use Scale Data. To assess the impact, execute PCA for linear dimensionality reduction on the matrix, selecting 20 principal components (Fig \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eD-E), and implement Harmony to mitigate batch effects. In order to preserve batch variability within each cluster, this procedure iteratively groups comparable cells from various batches inside the PCA space (Fig \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eF). Finally, use unified manifold approximation and projection, or UMAP, to reduce nonlinear dimensionality., use Find Neighbors to identify neighboring cell points, and apply Find Clusters to categorize cells into distinct clusters (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e3. Contribution of cell annotation and cell subpopulations to CRC\u003c/h2\u003e \u003cp\u003eThis research additionally classified each subtype, with all clusters designated into nine cellular categories: Epithelial cells, CD4\u0026thinsp;+\u0026thinsp;T lymphocytes, Monocytes, CD8\u0026thinsp;+\u0026thinsp;T lymphocytes, B lymphocytes, Fibroblasts, Plasma cells, Endothelial cells, and Dendritic cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB). The traditional markers for the nine cell types are displayed in a bubble plot (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC). Next, by taking into account variations in cell counts and gene expression, we were able to define the contribution of several cell subpopulations to CRC. We selected differential genes from the tumor group to the normal group to describe the changes in this process. Then FC score was defined, which is used to measure changes in the number and expression levels of characteristic genes in biological processes. Finally, it was found that Endothelial cells contribute the most to CRC (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e4. Construction of prognostic model\u003c/h2\u003e \u003cp\u003eWe gathered clinical data from CRC samples within the TCGA database and employed marker genes of endothelial cells to identify prognostic genes in CRC via Cox univariate regression. The findings indicated that a total of 28 genes linked to prognosis were found (p value\u0026thinsp;\u0026lt;\u0026thinsp;0.05) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA). Based on the prognostic genes, unique genes in colorectal cancer were extracted using the lasso regression feature selection approach. We used lasso regression analysis to get each sample after randomly dividing the processed CRC dataset, which contained survival data samples from the TCGA database, into training and testing sets at a 4:1 ratio (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB-D). The corresponding optimal risk score value is used for subsequent analysis. RiskScore\u0026thinsp;=\u0026thinsp;MTUS1\u0026times;(-0.174441201921624)\u0026thinsp;+\u0026thinsp;EMP2\u0026times;(-0.107481846357423)\u0026thinsp;+\u0026thinsp;NCOA7\u0026times;(-0.082921817145568)\u0026thinsp;+\u0026thinsp;EHD4\u0026times;(-0.078686080062342)\u0026thinsp;+\u0026thinsp;PTPRG\u0026times;(-0.0553784848516573)\u0026thinsp;+\u0026thinsp;EGFL7\u0026times;0.00453034652452583\u0026thinsp;+\u0026thinsp;INSR\u0026times;0.0161439754146179\u0026thinsp;+\u0026thinsp;JAG2\u0026times;0.0306554904877961\u0026thinsp;+\u0026thinsp;PTTG1IP\u0026times;0.0421245179315935\u0026thinsp;+\u0026thinsp;CLU\u0026times;0.0518270906986344\u0026thinsp;+\u0026thinsp;CAV1\u0026times;0.052522780662374\u0026thinsp;+\u0026thinsp;TMEM88\u0026times;0.0620618376532079\u0026thinsp;+\u0026thinsp;PODXL\u0026times;0.0740276615814645\u0026thinsp;+\u0026thinsp;IGFBP3\u0026times;0.0911907751517928\u0026thinsp;+\u0026thinsp;NPDC1\u0026times;0.105466421387826\u0026thinsp;+\u0026thinsp;PALMD\u0026times;0.132677159075287.The samples were divided into high-risk and low-risk groups according to risk ratings, and Kaplan-Meier curves were employed for analysisThe OS of the high-risk group was significantly poorer than that of the low-risk group in both the training and test sets (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eE-F). Furthermore, both the training and test sets' ROC curve findings show that the model performs well in verification (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eG-F).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003e5. Robustness of external validation set\u003c/h2\u003e \u003cp\u003eDownload the processed CRC samples accompanied by survival data from the GEO database (GSE17537, GSE87211), forecast the clinical classification of the CRC samples within the GEO database utilizing the model, and assess the survival disparity between the two groups through Kaplan-Meier analysis. The findings demonstrated that the OS of the high-risk group was noticeably lower than the low-risk group (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA-B). We used external datasets to perform ROC curve analysis on the model in order to verify its correctness. It was shown by the findings that the model was highly predictive in predicting the samples' prognosis (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC-D).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003e6.Immune characteristics and tumor microenvironment analysis\u003c/h2\u003e \u003cp\u003eThis investigation further examined whether high- and low-risk cohorts exhibit distinct immune profiles. Initially, the immune response was quantified utilizing the ssGSEA algorithm (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA), followed by a differential analysis aimed at uncovering risk score-specific immune signatures. The findings indicated that the immune characteristics were distinctly categorized between the risk score groups, with immune features such as Tem, TFH, TReg, and Tgd in the high-risk score group being significantly elevated compared to those in the low-risk score group (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB). Additionally, compared to the low-risk score group, the high-risk score group showed higher activation-related scores for the HIPPO, WNT, and Angiogenesis pathways (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC). The scientists then calculated immunological and stromal scores using the ESTIMATE method, which showed differences between the two groups' results (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eD-E).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003e7. Examination of the connection between immune infiltration and risk score\u003c/h2\u003e \u003cp\u003eThe tumor microenvironment is mostly composed of tumor-associated fibroblasts, extracellular matrix, immune cells, unique physical, inflammatory mediators and chemical properties, different growth factors, and the cancer cells themselves. This microenvironment profoundly influences tumor diagnosis and survival, as well as clinical treatment responsiveness. We further explore the possible biological pathways via which risk score influences the development of CRC by investigating the relationship between risk score and tumor immune infiltration. In a variety of ways, we display the immune cell composition and correlations between high- and low-risk groups (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA-B). Samples from the high-risk group showed considerably lower levels of resting and activating CD\u003csup\u003e4+\u003c/sup\u003e memory T cells than the low-risk group, but M0 macrophages and regulatory T cells (Tregs) showed much higher levels (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eC). We then looked into the relationship between immune cells and risk score. Among other things, the results showed that risk score had a substantial negative association with resting CD 4\u003csup\u003e+\u003c/sup\u003e memory T cells and activated CD 4\u003csup\u003e+\u003c/sup\u003e memory T cells, but a significant positive correlation with M2 and M0 macrophages (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eD).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003e8. Variations in mutations between populations at high and low risk\u003c/h2\u003e \u003cp\u003eThe CRC SNP-related data was processed, the TOP30 genes with the greatest mutation frequency were chosen for presentation, the differences in mutated genes between the two sample groups were compared, and the mutation landscape was displayed using the R package Complex Heatmap (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e). According to the results, samples from the high-risk group had a substantially lower percentage of APC and other gene mutations than samples from the low-risk group. Subsequently, we predicted both high- and low-risk groups' sensitivity to anti-tumor immunotherapy, and the findings showed that the high-risk group responded less well to immunotherapy (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003e9. Pathway Enrichment Analysis\u003c/h2\u003e \u003cp\u003eNext, the molecular mechanisms via which risk scores influence tumor growth were investigated, along with the specific signaling pathways implicated in the high- and low-risk correlation models. GSEA findings indicate that the relevant pathways include KEGG AMINOACYL TRNA BIOSYNTHESIS, KEGG GLYCOSAMINOGLYCAN BIOSYNTHESIS CHONDROITIN SULFATE, KEGG CITRATE CYCLE TCA CYCLE, among others (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eB). The two sample groups' unique pathways were mostly enriched in signaling pathways such MYOGENESIS, KRAS_SIGNALING_DN, and EPITHELIAL MESENCHYMAL TRANSITION, according to the GSVA data (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eC).\u003c/p\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003e10. Expression of model genes and immunometabolism pathways\u003c/h2\u003e \u003cp\u003eWe examined how model genes were expressed in individual cells and demonstrated how they were expressed in B cells, fibroblasts, plasma cells, endothelial cells, dendritic cells, monocytes, CD4\u0026thinsp;+\u0026thinsp;T cells, epithelial cells, and CD8\u0026thinsp;+\u0026thinsp;T cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eA-B). Finally, we quantitatively scored the immunometabolism-related pathway genes and used bubble charts to display the activity differences of the model genes in the immunometabolism-related pathways. The results showed that the 16 model genes had higher activity in the angiogenesis, epithelial mesenchymal transition, and TGF-β signaling pathways (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003eC). Finally, the developmental trajectory of model genes in Endothelial cells is displayed (Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003eA-D).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eOne of the most prevalent malignant tumors worldwide is colorectal cancer\u003csup\u003e[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]\u003c/sup\u003e. Its incidence and mortality vary from region to region, and the incidence in developed countries is usually higher than that in developing countries\u003csup\u003e[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]\u003c/sup\u003e. There are many risk factors affecting the occurrence of CRC, including hereditary diseases (such as familial adenomatosis and Lynch syndrome), bad eating habits, high-fat and low-fiber diet, obesity and age\u003csup\u003e[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]\u003c/sup\u003e. Early screening is particularly important, which can effectively reduce the mortality of CRC and promote early diagnosis\u003csup\u003e[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]\u003c/sup\u003e. In addition, individualized treatment, metastasis monitoring and nutritional and psychological support can effectively improve the therapeutic effect. In addition, regular imaging examination can help monitor tumor recurrence or metastasis, so as to adjust the treatment plan in time. Meanwhile, the development of a treatment and prognosis model for colorectal cancer is clinically significant. This model can not only improve the prediction ability of patients' disease progress and survival probability, but also optimize the treatment plan, provide personalized care for patients with different risk levels and rationally allocate medical resources\u003csup\u003e[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]\u003c/sup\u003e. In addition, the prognosis model also provides data support for basic and clinical research, promotes the research and development of new therapies, enhances patients' sense of education and participation, and improves their compliance and satisfaction with the treatment process. Here, we selected the marker genes of endothelial cells with the highest contribution to CRC from the single cell sequencing analysis by obtaining GSE231559 data set, and identified 16 model genes that are strongly associated with patient survival and the risk score that may be used to forecast the prognosis of patients with colorectal cancer. At the same time, the performance of the model was verified by external verification queues GSE17537 and GSE87211. Among 16 genes, it has been reported MTUS1, EGFL7, INSR, CLU, CAV1, TMEM88, PODXL, and IGFBP3 could use to be CRC biomarkers\u003csup\u003e[\u003cspan additionalcitationids=\"CR16 CR17 CR18 CR19 CR20 CR21 CR22\" citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]\u003c/sup\u003e. These studies once again confirm the importance of our model genes\u003c/p\u003e \u003cp\u003eIn recent years, the research on immune infiltration of CRC has made remarkable progress, revealing the important role of immune cells in tumor microenvironment. Many types of immune cells, including CD8\u0026thinsp;+\u0026thinsp;T lymphocytes, B cells and macrophages, can be detected in colon cancer tissues, and their abundance and activity are closely related to the development, prognosis and response to treatment of the tumor\u003csup\u003e[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]\u003c/sup\u003e. For example, a high level of CD8\u0026thinsp;+\u0026thinsp;T cell infiltration is usually associated with a good clinical prognosis, while tumor-associated macrophages are often associated with enhanced tumor invasion\u003csup\u003e[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]\u003c/sup\u003e. In addition, with the application of single cell RNA sequencing and other technologies, researchers can evaluate the characteristics of immune microenvironment more accurately, which provides new opportunities for individualized treatment\u003csup\u003e[\u003cspan additionalcitationids=\"CR27 CR28\" citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]\u003c/sup\u003e. Therefore, immune infiltration is not only an important factor in understanding the biology of CRC, but also has the potential as a prognostic marker and guiding treatment decision. According to our findings, the high-risk group's T cells' CD4 memory resting and activation are much lower than those of the low-risk group; The Macrophages M0 and Tregs were higher. Then we discussed the relationship between risk score and immune cells. The findings indicated that the risk score was positively correlated with MacrophagesM2 and MacrophagesM0, and inversely connected with T cells CD4 memory resting and TcellsCD4 memory activated. In a word, these results are very helpful to further investigate the mechanism of immune infiltration and optimize the immunotherapy strategy for CRC in the future.\u003c/p\u003e \u003cp\u003eHowever, this study also has limitations. First, the biological functions and potential regulatory mechanisms of these genes will be confirmed through experiments in the future. Secondly, we will collect information on clinical samples in our hospital and once again highlight the accuracy and importance of this model, in order to provide reliable targets for early diagnosis of clinical CRC and provide reference for the formulation of personalized treatment plans for CRC.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by the Chongqing Natural Science Foundation [grant number CSTB2022NSCQ-MSX1116] and the Health Commission of Chongqing Scientific Research Project [grant number 2023MSXM173].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets used and/or analyzed during the current study are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompliance Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eClinical trial number\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eNot applicable.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to Publish declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eNot applicable.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to Participate declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eNot applicable.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eDan Yang conceptualized the study, designed the methodology, and wrote the original draft.Fengsheng Dai performed formal data analysis, curated datasets, and contributed to manuscript revisions.Xianghua Zeng developed visualization tools, validated experimental results, and assisted in data interpretation.Jianhua Wang secured funding, supervised project administration, and critically reviewed the manuscript.Lijuan Zhao conducted investigations, organized resources, and participated in editing the final draft.Tingxiu Xiang contributed to statistical validation, reviewed the literature, and edited technical sections.Zhongjun Wu oversaw the research framework, coordinated interdisciplinary collaboration, and approved the final manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003ePatel SG, Karlitz JJ, Yen T, Lieu CH, Boland CR. The rising tide of early-onset colorectal cancer: a comprehensive review of epidemiology, clinical features, biology, risk factors, prevention, and early detection. Lancet Gastroenterol Hepatol. 2022;7(3):262\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYan H, Talty R, Johnson CH. Targeting ferroptosis to treat colorectal cancer. Trends Cell Biol. 2023;33(3):185\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbedizadeh R, Majidi F, Khorasani HR, Abedi H, Sabour D. Colorectal cancer: a comprehensive review of carcinogenesis, diagnosis, and novel strategies for classified treatments. Cancer Metastasis Rev. 2024;43(2):729\u0026ndash;53.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWilliams H, Lee C, Garcia-Aguilar J. Nonoperative management of rectal cancer. Front Oncol. 2024;14:1477510.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAndrei P, Battuello P, Grasso G, Rovera E, Tesio N, Bardelli A. Integrated approaches for precision oncology in colorectal cancer: The more you know, the better. Semin Cancer Biol. 2022;84:199\u0026ndash;213.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu H, Wang Q, Yu D et al. Berberine Derivative B68 Promotes Tumor Immune Clearance by Dual-Targeting BMI1 for Senescence Induction and CSN5 for PD-L1 Degradation. Adv Sci (Weinh). 2024: e2413122.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang N, Fang JY. Fusobacterium nucleatum, a key pathogenic factor and microbial biomarker for colorectal cancer. Trends Microbiol. 2023;31(2):159\u0026ndash;72.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin Q, Wang Z, Wang J, et al. Innovative strategies to optimise colorectal cancer immunotherapy through molecular mechanism insights. Front Immunol. 2024;15:1509658.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFan X, Li B, Zhang F et al. FGF19-Activated Hepatic Stellate Cells Release ANGPTL4 that Promotes Colorectal Cancer Liver Metastasis. Adv Sci (Weinh). 2024: e2413525.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhu AK, Li GY, Chen FC, et al. Integrated Analysis of Single-Cell and Bulk RNA-Sequencing Based on EcoTyper Machine Learning Framework Identifies Cell-State-Specific M2 Macrophage Markers Associated with Gastric Cancer Prognosis. Immunotargets Ther. 2024;13:721\u0026ndash;34.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu C, Huang X, Chen J, et al. Dissecting the cellular reprogramming and tumor microenvironment in left- and right-sided Colorectal Cancer by single cell RNA sequencing. Transl Res. 2024;276:22\u0026ndash;37.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAhluwalia P, Mondal AK, Vashisht A, et al. Identification of a distinctive immunogenomic gene signature in stage-matched colorectal cancer. J Cancer Res Clin Oncol. 2024;151(1):9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChu X, Li X, Zhang Y, et al. Integrative single-cell analysis of human colorectal cancer reveals patient stratification with distinct immune evasion mechanisms. Nat Cancer. 2024;5(9):1409\u0026ndash;26.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang Q, Liu Y, Wang X, Zhang C, Hou M, Liu Y. Integration of single-cell RNA sequencing and bulk RNA transcriptome sequencing reveals a heterogeneous immune landscape and pivotal cell subpopulations associated with colorectal cancer prognosis. Front Immunol. 2023;14:1184167.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCheng LY, Huang MS, Zhong HG, et al. MTUS1 is a promising diagnostic and prognostic biomarker for colorectal cancer. World J Surg Oncol. 2022;20(1):257.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJuan Z, Dake C, Tanaka K, Shuixiang H. EGFL7 as a novel therapeutic candidate regulates cell invasion and anoikis in colorectal cancer through PI3K/AKT signaling pathway. Int J Clin Oncol. 2021;26(6):1099\u0026ndash;108.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ede Oliveira C, Martins S, Gon\u0026ccedil;alves PG, et al. Low EGFL7 expression is associated with high lymph node spread and invasion of lymphatic vessels in colorectal cancer. Sci Rep. 2023;13(1):19783.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLandi D, Moreno V, Guino E et al. Polymorphisms affecting micro-RNA regulation and associated with the risk of dietary-related cancers: a review from the literature and new evidence for a functional role of rs17281995 (CD86) and rs1051690 (INSR), previously associated with colorectal cancer. Mutat Res. 2011. 717(1\u0026ndash;2): 109\u0026thinsp;\u0026ndash;\u0026thinsp;15.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe W, Chan CM, Wong SC, et al. Jagged 2 silencing inhibits motility and invasiveness of colorectal cancer cell lines. Oncol Lett. 2016;12(6):5193\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArtemaki PI, Sklirou AD, Kontos CK, et al. High clusterin (CLU) mRNA expression levels in tumors of colorectal cancer patients predict a poor prognostic outcome. Clin Biochem. 2020;75:62\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeng Y, Zhao J, Yin F, et al. A methylation-driven gene panel predicts survival in patients with colon cancer. FEBS Open Bio. 2021;11(9):2490\u0026ndash;506.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee H, Kong JS, Lee SS, Kim A. Radiation-Induced Overexpression of TGFβ and PODXL Contributes to Colorectal Cancer Cell Radioresistance through Enhanced Motility. Cells. 2021. 10(8): 2087.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePerez-Carbonell L, Balaguer F, Toiyama Y, et al. IGFBP3 methylation is a novel diagnostic and predictive biomarker in colorectal cancer. PLoS ONE. 2014;9(8):e104285.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePicard E, Verschoor CP, Ma GW, Pawelec G. Relationships Between Immune Landscapes, Genetic Subtypes and Responses to Immunotherapy in Colorectal Cancer. Front Immunol. 2020;11:369.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGiannakis M, Mu XJ, Shukla SA, et al. Genomic Correlates of Immune-Cell Infiltrates in Colorectal Carcinoma. Cell Rep. 2016;15(4):857\u0026ndash;65.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMasuda K, Kornberg A, Miller J, et al. Multiplexed single-cell analysis reveals prognostic and nonprognostic T cell types in human colorectal cancer. JCI Insight. 2022;7(7):e154646.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi J, Wu C, Hu H, et al. Remodeling of the immune and stromal cell compartment by PD-1 blockade in mismatch repair-deficient colorectal cancer. Cancer Cell. 2023;41(6):1152\u0026ndash;e11697.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang R, Li J, Zhou X, et al. Single-cell genomic and transcriptomic landscapes of primary and metastatic colorectal cancer tumors. Genome Med. 2022;14(1):93.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMei Y, Xiao W, Hu H, et al. Single-cell analyses reveal suppressive tumor microenvironment of human colorectal cancer. Clin Transl Med. 2021;11(6):e422.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-6455265/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6455265/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eIntroduction: \u003c/strong\u003eColorectal cancer (CRC) is a ubiquitous malignant tumor, ranking third in the global incidence and second in mortality. Despite advancements in screening techniques like colonoscopies, the lack of trustworthy biomarkers makes it difficult to diagnose CRC accurately. The objective of this study is to construct a prognostic risk model of CRC for accurate prediction of CRC results.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMaterials and Methods: \u003c/strong\u003eWe obtained GSE231559 data set from GEO database for single cell sequencing analysis, selected marker genes of endothelial cells with the highest contribution to CRC to construct a prognosis model, determined 16 model genes intimately associated with the survival of patients and risk scores that can predict the prognosis of CRC patients, and verified the performance of the model through external verification queues GSE17537 and GSE87211. Then, the molecular mechanism of the prognosis model was discussed from the aspects of immune infiltration, mutation map, immunotherapy and signal pathways involved in risk score.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults: \u003c/strong\u003eThe findings indicated that the immune system of high-risk patients had poor response ability to tumors, and the tumor microenvironment was more conducive to immune escape, leading to tumor progress. The immunotherapy effect of high-risk patients is worse, and as compared to low-risk patients, the OS is much lower. The enriched pathways are related to tumor growth, invasion, metastasis and the use of immune escape mechanism, which may provide powerful clues for the prognosis of high-risk patients or malignant tumors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDiscussion: \u003c/strong\u003eWe identified a total of 16 model genes that are expected to serve as new CRC biomarkers. However, the biological functions and potential regulatory mechanisms of these genes should be confirmed through experiments in the future.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion:\u003c/strong\u003e Our findings demonstrate that the prognostic risk of CRC patients may be evaluated using this approach.\u003c/p\u003e","manuscriptTitle":"Single cell sequencing combined with transcriptome sequencing to build a prognostic model for colorectal cancer","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-23 05:40:25","doi":"10.21203/rs.3.rs-6455265/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"55669bc2-5cad-4e49-8ee5-339b17cc9dc2","owner":[],"postedDate":"June 23rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-09-18T10:39:52+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-23 05:40:25","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6455265","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6455265","identity":"rs-6455265","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.