Integrative Bioinformatics of Bulk and Single-Cell Transcriptomes Identifies Predictive Molecular Features of Ulcerative Colitis-Associated Colorectal Cancer | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Integrative Bioinformatics of Bulk and Single-Cell Transcriptomes Identifies Predictive Molecular Features of Ulcerative Colitis-Associated Colorectal Cancer Changming Huang, Xian Zhang, Liying Wang, Jingxuan Zhang, Yang Gao, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8417562/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Ulcerative colitis (UC) is a chronic inflammatory disease that increases the risk of colorectal cancer (CRC), yet the molecular features linking these conditions remain unclear. Here, we integrated bulk RNA sequencing with single-cell transcriptomic data to investigate shared transcriptional programs between UC and CRC. UC-associated differentially expressed genes (DEGs) were mapped onto UC and sporadic CRC single-cell datasets, allowing cell type-resolved analysis of conserved alterations. We found that a subset of UC DEGs displayed consistent dysregulation in sporadic CRC, with strong cell specificity. Myeloid cells-particularly macrophages, neutrophils, and dendritic cells-showed coordinated upregulation of inflammatory genes in both diseases, whereas epithelial cells exhibited convergent downregulation of protective gene modules. These patterns suggest that persistent immune activation and epithelial dysfunction represent shared pathological axes that may contribute to inflammation-driven tumorigenesis. Using bulk data from UC-associated CRC, machine-learning analysis identified four genes with predictive value for UC-to-CRC progression. Overall, this study provides a refined cellular framework for understanding UC-CRC molecular convergence and highlights candidate biomarkers for early detection. Ulcerative colitis Colorectal cancer scRNA sequencing Inflammation-driven tumorigenesis Machine Learning Tumor microenvironment Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Introduction Inflammatory bowel disease (IBD) is a prevalent global disorder in the 21st century [1] . Among IBD subtypes, ulcerative colitis (UC) is a chronic inflammatory condition of the colon characterized by recurrent mucosal inflammation and epithelial injury. Patients with long-standing UC have an increased risk of developing colorectal cancer (CRC), which represents a major cause of cancer-related morbidity and mortality worldwide [2, 3] . Although the elevated cancer risk in UC has been recognized clinically, the molecular mechanisms driving the transition from chronic inflammation to malignant transformation remain incompletely understood. Because UC is incurable, affected patients often have reduced life expectancy and are at heightened risk for CRC. Recent studies have reported numerous new findings regarding UC biomarkers and clinical manifestations [4] , as well as therapeutic targets and strategies for CRC [5, 6] . However, analyses exploring the relationship between UC and CRC, particularly sporadic CRC in patients without chronic colitis, remain limited; most studies have focused on UC-associated CRC (UC-CRC) [7, 8] . Although UC-CRC and sporadic CRC display distinct pathological features, whether these two conditions share molecular correlations or overlapping transcriptional programs is still unclear. Identifying common molecular signatures between UC and sporadic CRC is crucial for understanding mechanisms that may predispose UC patients to malignant transformation. Such shared programs or convergent cellular pathways could provide novel insights into the early events of inflammation-driven tumorigenesis and highlight potential biomarkers or therapeutic targets for predicting and preventing UC-associated CRC. Advances in high-throughput transcriptomic technologies have enabled the systematic identification of disease-associated molecular signatures in complex disorders [9] . Bulk RNA sequencing allows the detection of differentially expressed genes (DEGs) between diseased and normal tissues [10] , providing insights into the global molecular alterations associated with UC. However, bulk transcriptomic approaches are limited in their capacity to resolve cell type-specific expression patterns, which are essential for understanding the heterogeneous cellular composition of the inflamed colon. Single-cell RNA sequencing (scRNA-seq) has emerged as a powerful method to dissect cellular heterogeneity and to characterize cell type–specific transcriptional programs [11] . By integrating bulk and single-cell transcriptomic data, it is possible to capture both global and cell-resolved gene expression changes [12] , thereby facilitating the identification of key molecular features shared between UC and CRC. In this study, we first identified differentially expressed genes (DEGs) between UC and healthy colonic tissues using bulk RNA-seq data. We then evaluated the expression patterns of these DEGs in both UC and sporadic CRC single-cell RNA-seq datasets to investigate conserved cellular and molecular features. Finally, we aimed to pinpoint potential key genes that may drive the progression from UC to CRC, thereby highlighting candidate biomarkers for early detection and molecular targets for therapeutic intervention. By integrating multi-scale transcriptomic data, this study establishes a comprehensive framework for elucidating the molecular interplay between chronic inflammation and colorectal tumorigenesis. Result 1. Key gene sets underlying transcriptional alterations in UC patients To systematically evaluate the transcriptional characteristics of ulcerative colitis (UC) patients, a total of 670 samples from five independent datasets were included in this study. These datasets shared more than 15,000 overlapping genes (Fig. S1 A). After removing batch effects (Fig. S1 B–D), differential expression analysis revealed a substantial number of significantly dysregulated genes in rectal tissues of UC patients compared with normal controls (Fig. 1 A). Specifically, 193 genes were markedly upregulated and 144 genes were downregulated (P < 0.05, |log2FC| ≥ 1), indicating prominent transcriptional alterations in diseased tissues. To link these transcriptomic changes with disease phenotypes, we further identified disease-related genes using weighted gene co-expression network analysis (WGCNA) in combination with clinical variables (AGE, DISEASE, GENDER, and ANCESTRY), aiming to uncover potential biomarkers and underlying mechanisms. The optimal soft-thresholding power for network construction was determined based on algorithmic recommendations (Fig. 1 B), resulting in the identification of 12 co-expression modules with distinct expression patterns (Fig. 1 C). Module–trait correlation analysis showed that the yellow module exhibited the strongest positive correlation with UC (132 genes), while the blue module showed the strongest negative correlation (595 genes) (Fig. 1 D). Module Membership (MM) versus Gene Significance (GS) analysis indicated moderate overall correlations between module eigengenes and disease traits (r > 0.5) for both modules (Fig. 1 E). Therefore, the yellow and blue modules were selected for subsequent functional analyses. Complete gene lists for these two modules are provided in Table S1 . Among the genes significantly upregulated in UC, 60 overlapped with the disease-associated yellow module, whereas 69 downregulated genes overlapped with the negatively correlated blue module (Fig. 1 F). To visualize the expression patterns of these key genes, we generated a heatmap of 129 representative genes across all samples (Fig. 1 G). Functional enrichment analysis revealed that the 60 upregulated, disease-associated genes (up_GeneSignature) were predominantly involved in immune and inflammatory responses, including cytokine–receptor interactions and pathways such as IL-17, TNF, and NF-κB signaling, as well as chemokine, Toll-like receptor, and NOD-like receptor signaling pathways (Fig. 2 A,B). These pathways are closely linked to immune regulation and inflammatory injury, consistent with the known pathogenesis of UC. In contrast, the 69 downregulated, negatively correlated genes (down_GeneSignature) were mainly enriched in metabolic and transport-related pathways, including mineral absorption, nitrogen metabolism, PPAR signaling, and AMPK signaling (Fig. 2 C,D), suggesting that this gene set may play a critical role in maintaining metabolic homeostasis under physiological conditions. 2. Immune cell infiltration characteristics in UC patients To further characterize the immune microenvironment in ulcerative colitis (UC), we analyzed immune cell infiltration patterns based on the upregulated gene set (up_GeneSignature) and its associated enrichment pathways. The immune infiltration analysis revealed substantial heterogeneity in immune cell composition across samples and between groups, with T cells and macrophages showing notably higher infiltration levels in UC tissues (Fig. 3 A). Correlation analysis among immune cell subsets demonstrated complex intercellular interactions (Fig. 3 B). Specifically, resting mast cells, NK cells, and CD4⁺ memory T cells were negatively correlated with their activated counterparts. M0/M1 macrophages exhibited negative correlations with B cells, activated mast cells, and activated CD4⁺ memory T cells, but positive correlations with their resting states, suggesting that macrophage polarization may be closely linked to the activation states of T cells and mast cells. Moreover, M2 macrophages showed negative correlations with B cells and monocytes(Fig. 3 B). Myeloid cells, including macrophages, mast cells, and neutrophils, displayed marked infiltration differences between UC and control tissues (Fig. 3 C), underscoring their potential central roles in disease progression. Notably, expression levels of the upregulated 60-gene signature were strongly correlated with the infiltration of macrophages and dendritic cells (Fig. 3 D), implying their involvement in immune cell recruitment and activation. In contrast, the downregulated 69-gene set (down_GeneSignature) showed only weak associations with immune infiltration (Fig. 3 E). Collectively, these findings indicate that aberrant activation of macrophages, dendritic cells, and mast cells may contribute to persistent mucosal inflammation in UC, and that the upregulated gene signature likely represents a core immune regulatory network driving disease pathogenesis. 3. Comparative single-cell transcriptomic profiling of UC and CRC patients Single-cell transcriptomic data from UC and CRC patients were integrated to compare cellular composition and transcriptional patterns. After quality control (Fig. S2 A–D), identification of highly variable genes (Fig. S2 I, J), PCA-based dimensionality reduction (Fig. S2 E–H, K, L), and batch correction (Fig. 4 A–D), the datasets achieved good consistency and comparability. Clustering and cell-type annotation identified major cell populations, including T cells, B cells, myeloid cells, endothelial cells, and epithelial cells (Fig. 4 E, F). Feature and dot plots confirmed the accuracy of annotation through canonical markers such as CD3D (T cells), CD19 (B cells), CD68 (macrophages), EPCAM (epithelial cells), and PECAM1 (endothelial cells) (Fig. 4 G–J). Compared with normal tissues, UC samples displayed a marked increase in immune cell infiltration and a reduction in stromal components such as fibroblasts and epithelial cells (Fig. 4 K, M). NK cells, B cells, and myeloid cells were significantly enriched, consistent with inflammation-driven immune activation and tissue remodeling. In contrast, CRC tissues exhibited a distinct pattern: NK and B cell proportions decreased, whereas myeloid cells remained elevated and epithelial cells expanded markedly within tumor regions (Fig. 4 L, N). Despite opposite trends in immune composition, UC and CRC shared similar myeloid transcriptional signatures, suggesting conserved activation states under inflammatory and tumor microenvironments. Notably, the 60 upregulated genes from the UC signature (up_GeneSignature) showed the highest enrichment scores in myeloid cells across both diseases (Fig. 4 O, P), indicating that these cells are likely the principal effectors mediating immune microenvironment remodeling in UC and CRC. (A–D) Batch effect correction before and after integration for UC (A, B) and CRC (C, D) datasets. Each point represents a single cell, colored by sample group. (E, F) UMAP plots of UC (E) and CRC (F) datasets, showing annotated cell types and spatial distributions. (G, H) Feature plots of canonical marker genes in major cell types for UC (G) and CRC (H). (I, J) Dot plots showing marker gene expression across cell clusters used for annotation in UC (I) and CRC (J). (K, L) UMAP plots comparing cellular composition between normal and disease samples in UC (K) and CRC (L). (M, N) Bar plots showing proportion changes of major cell types between normal and disease groups in UC (M) and CRC (N). (O, P) Violin plots of up_GeneSignature scores across cell types in UC (O) and CRC (P). Scores were calculated using the Seurat AddModuleScore function, with the y-axis representing the module score for each cell and the x-axis representing cell types. 4. Comparative transcriptional features of myeloid cells in UC and CRC patients To further explore myeloid transcriptional features in UC and CRC, myeloid cells were re-clustered and subtyped into M1 macrophages, M2 macrophages, dendritic cells (DCs), neutrophils, and mast cells. In UC tissues, macrophages were the dominant myeloid population, with M1 macrophages being the most abundant (Fig. 5 A). In the UC group, the proportion of M1 macrophages showed no obvious change, whereas M2 macrophages were significantly upregulated; DCs and neutrophils exhibited an increasing trend, and mast cells decreased (Fig. 5 B). Dot plot analysis confirmed marker expression across clusters and indicated that M2 macrophages display gene expression features resembling tumor-associated macrophages (TAMs), suggesting possible roles in modulating the inflammatory microenvironment (Fig. 5 C). The 60-gene up_GeneSignature was highly expressed in macrophages and neutrophils (Fig. 5 D). Compared with controls, UC samples showed significant upregulation of this signature in M2 macrophages, mast cells, DCs, and neutrophils, while M1 macrophages showed a downward trend (Fig. 5 E). In CRC, myeloid compartments were also dominated by macrophages, with a notable increase in TAM-like macrophages compared with UC (Fig. 5 F). Subclassification in CRC showed increases across the three macrophage subsets and a decrease in neutrophils (Fig. 5 G). Marker profiles highlighted pronounced M2/TAM features in TAM macrophages (Fig. 5 H). The up_GeneSignature exhibited overall high expression in M1, M2 macrophages, and neutrophils, consistent with the expression pattern observed in UC samples(Fig. 5 I), and relative to controls, CRC samples exhibited significant increases of this signature in macrophages, neutrophils, and DCs; the M2, TAM phenotype was more prominent in CRC than in UC (Fig. 5 J). Single-cell differential expression comparing CRC myeloid cells to controls identified 16 significantly upregulated genes that overlap with the UC-derived up_GeneSignature (Fig. 5 K–L). Violin plots show the expression differences of these 16 genes across myeloid subsets and groups (Fig. 5 M). Overall, both UC and CRC myeloid cells show activation of M2-like programs and related inflammatory pathways, with a stronger TAM signature in CRC. The UC-derived up_GeneSignature remains enriched in myeloid cells across both diseases, suggesting shared immune-regulatory roles and potential predictive value. 5. Comparative transcriptional features of epithelial cells in UC and CRC patients We next examined the expression patterns of the down_GeneSignature identified from UC transcriptomes in single-cell samples of UC and CRC. The analysis revealed that this gene set was predominantly expressed in epithelial cells in both UC and CRC samples (Fig. 6 A, B). Specifically, compared with controls, epithelial cells from UC and CRC patients exhibited a significant overall downregulation of down_GeneSignature genes, whereas in immune cells these genes were largely significantly upregulated (Fig. 6 C, D), consistent with our earlier immune infiltration analysis showing low correlation between down_GeneSignature and immune cell abundance. Further analysis indicated that 21 down_GeneSignature genes were significantly downregulated in epithelial cells of CRC tumor samples (Fig. 6 E), with their distribution visualized in a volcano plot (Fig. 6 F). Taken together, the expression patterns across UC and CRC suggest that these 21 downregulated genes in CRC epithelial cells may have potential predictive value. 6. Machine learning–based identification and validation of key genes in UC-to-CRC progression To further identify and validate the key roles of up_GeneSignature and down_GeneSignature genes in the progression from UC to secondary CRC, we employed the GSE3629 dataset, which includes RNA-seq data from UC-related CRC tumor tissues and adjacent non-tumor tissues. This dataset allows us to assess whether these gene signatures maintain continuous involvement during the transition from UC to CRC. Using a 10-fold cross-validation strategy, the LASSO model was applied (Fig. 7 A, B), resulting in six genes with significant discriminative power: IL1B, FPR2, OSM, GBP1, MT1G, and MT1H. The directions of their regression coefficients indicated that IL1B, FPR2, OSM, and GBP1 acted as potential risk factors, whereas MT1G and MT1H might function as protective factors. Random forest analysis confirmed these six genes as important features (Fig. 7 C, D). ROC curves were plotted for each of the six candidate genes, and all exhibited AUC values above 0.9, indicating high diagnostic potential for distinguishing CRC tumor tissues from adjacent tissues (Fig. 7 E). A logistic regression model constructed using these six genes demonstrated good overall fit, and the nomogram (Fig. 7 F) provides a quantitative reference for evaluating the contribution of each gene to UC-derived CRC risk in clinical prediction. Survival analysis revealed that MT1G and MT1H, two down_GeneSignature genes, showed a trend toward worse relapse-free survival (RFS) in the low-expression group, although the log-rank P values did not reach statistical significance; the hazard ratios (HRs for the high-expression group) were < 1, consistent with their potential protective roles (Fig. 7 G, H). In contrast, among up_GeneSignature genes, FPR2 and OSM exhibited a trend toward reduced survival in the high-expression group, with HRs > 1 (Fig. 7 I, J). IL1B and GBP1 did not show a survival trend, which may reflect their predictive relevance being restricted to myeloid cells, as observed in Section 4, rather than at the bulk tissue level. Discussion Although UC-associated colorectal cancer (UC-CRC) and sporadic CRC exhibit distinct clinical and pathological characteristics, our study identified shared transcriptional features between these two conditions. Differentially expressed genes (DEGs) observed in UC patients relative to healthy controls were also significantly altered in CRC patients, particularly within immune and epithelial cell populations. This overlap in transcriptomic alterations suggests that certain molecular programs initiated during chronic colonic inflammation in UC may persist and contribute to the early events of colorectal tumorigenesis. Notably, among these DEGs, we identified potential predictive genes for UC-CRC, including MT1G, MT1H, and FPR2. These genes appear to be continuously dysregulated from the inflammatory stage in UC through the progression toward CRC and exhibit consistent changes in sporadic CRC as well. This observation indicates that a subset of genes may represent a molecular continuum, linking chronic inflammation with subsequent malignant transformation, and highlights the possibility that these genes may serve as early biomarkers or mechanistic mediators of inflammation-driven colorectal carcinogenesis. Furthermore, the consistent transcriptomic alterations in immune and epithelial cells imply that both cell-intrinsic changes and the immune microenvironment may play key roles in mediating the transition from chronic inflammation to tumorigenesis. Future studies exploring the functional roles of these shared DEGs and their pathways could provide mechanistic insights and potentially inform strategies for early detection or targeted intervention in patients at high risk for UC-CRC. By integrating single-cell and bulk transcriptomic data, we extended our analysis to the cellular level and observed that many genes significantly dysregulated in UC also exhibited pronounced alterations in CRC, albeit with strong cell type specificity. Interestingly, DEGs identified solely at the bulk transcriptome level in UC did not show substantial changes when assessed across the overall transcriptome of sporadic CRC, which may result in their underappreciation in previous studies. However, the single-cell perspective revealed that UC-associated upregulated genes were predominantly correlated with immune cells, particularly myeloid populations such as macrophages, neutrophils, and dendritic cells, indicating that these cell types are critical contexts in which these genes exert their effects. This finding suggests that the high expression of these genes in specific immune cell subsets may contribute to disease exacerbation. Conversely, UC-associated downregulated genes were significantly enriched in epithelial cells, suggesting a potential role in negatively regulating disease progression. Although the precise contribution of these genes to CRC pathogenesis requires further functional validation, their consistent cell-specific dysregulation in both UC and sporadic CRC highlights the importance of considering cellular resolution when investigating inflammation-driven tumorigenesis. Importantly, we identified four genes with potential predictive value for UC-CRC progression. Due to the lack of suitable single-cell data for UC-associated CRC, these candidates were evaluated using bulk UC-CRC datasets through machine learning and prognostic analyses. This approach underscores a key advantage of our study: genes that do not appear significantly altered at the bulk transcriptome level may still exert critical effects through cell type–specific expression changes, and such effects can be captured using integrated single-cell analyses. Our findings provide a framework for identifying candidate molecular mediators of UC-associated CRC progression and illustrate the value of combining bulk and single-cell transcriptomics to uncover subtle but biologically meaningful signals. We observed that FPR2 and OSM are significantly upregulated in immune cells and represent shared expression features between UC and sporadic CRC, serving as potential predictive genes for UC-associated CRC. FPR2 encodes a G protein-coupled receptor that mediates chemotaxis and modulates inflammatory responses. Previous studies have linked FPR2 upregulation to dysregulated immune responses and disease pathogenesis, including influenza infection [13] , pulmonary fibrosis [14] , and intestinal inflammation [15] . In the context of UC and CRC, persistent FPR2 upregulation in myeloid cells, such as macrophages, neutrophils, and dendritic cells, may sustain chronic inflammatory signaling, promoting an environment conducive to tumor initiation and progression. This observation highlights FPR2 as a potential biomarker and suggests that targeting FPR2-mediated pathways could modulate inflammation-driven carcinogenesis. Oncostatin M (OSM) is a pleiotropic cytokine involved in regulating diverse inflammatory processes, including tissue repair, liver regeneration, and bone remodeling [16, 17] . Since its discovery in 1986, OSM has been implicated as a key mediator in various inflammatory conditions and cancers, including arthritis [18] , inflammatory bowel disease, pulmonary and skeletal disorders [19, 20] , as well as pancreatic, renal, and cervical cancers [21–23] . In our study, the concurrent upregulation of FPR2 and OSM in myeloid cells suggests a coordinated inflammatory program that may persist in both UC and CRC, potentially sustaining a pro-tumorigenic microenvironment. This finding highlights a cell type-specific mechanism, where chronic activation of myeloid cells could contribute to tissue damage, immune dysregulation, and progression toward malignant transformation. Notably, these transcriptional changes were not apparent at the bulk tissue level, emphasizing the importance of single-cell resolution for detecting key drivers of disease. The significance of our findings lies in identifying shared molecular programs in inflammation-driven carcinogenesis that are conserved between UC and sporadic CRC. By pinpointing FPR2 and OSM as highly expressed in myeloid compartments, we highlight candidate biomarkers for early detection and potential targets for modulating chronic inflammation. Integrating these insights with predictive modeling for UC-associated CRC further illustrates the value of combining single-cell transcriptomics with bulk RNA-seq data to uncover cell type-specific alterations that may drive disease progression, which would be overlooked in conventional bulk analyses. In addition, we observed that MT1G and MT1H are significantly downregulated in colonic epithelial cells of both UC and sporadic CRC patients. Further analyses revealed a negative correlation between their expression and disease progression, suggesting that reduced levels may contribute to persistent inflammation, epithelial barrier dysfunction, and increased risk of malignant transformation. MT1G and MT1H, members of the metallothionein (MT) family, are small cysteine-rich proteins involved in metal homeostasis, protection against oxidative stress, and regulation of cellular proliferation and apoptosis [24–26] . Their downregulation in epithelial cells implies a potential loss of protective function, which may facilitate disease exacerbation and progression toward CRC. These findings indicate that MT1G and MT1H may exert protective or suppressive roles in intestinal epithelial cells, and their decreased expression could promote disease worsening and malignant transformation. Although the precise mechanisms remain to be fully elucidated, these results provide novel insights into the cell type-specific regulation underlying the transition from chronic inflammation to tumorigenesis, and they may serve as potential targets for therapeutic intervention or as prognostic biomarkers. Materials and Methods Bulk RNA-seq Data Analysis All bulk RNA-seq datasets were downloaded from GEO ( http://www.ncbi.nlm.nih.gov/geo ). The datasets used to screen key gene changes in UC included GSE66407 (41 UC, 16 controls), GSE158952 (26 UC), GSE193677 (309 UC, 220 controls), GSE235236 (24 UC, 7 controls), and GSE245890 (27 UC). Rectum tissue samples from UC patients (n = 427) and healthy controls (n = 243) were used. To validate the role of the key gene sets in UC-to-CRC disease progression, the dataset GSE3629 (10 adjacent normal, 6 tumor) was used, including in situ tumor and adjacent tissue of UC-associated CRC. For the five bulk RNA-seq datasets, probe sequences were mapped to gene symbols based on their respective sequencing platforms. For multiple probes corresponding to the same gene symbol, the probe with the highest measured intensity (maximum mean) was retained. The datasets were then merged and batch effects were removed. The effectiveness of batch correction was confirmed using boxplots and PCA. Data download was performed using the R package GEOquery, and batch effect removal was conducted using the R package sva. Boxplots were drawn using the boxplot function, and PCA plots were generated with draw_pca. For GSE3629, after probe-to-symbol mapping, the raw data were log-transformed and quantile normalized. Differential expression analysis for bulk RNA-seq was performed using the R package limma to compare UC versus control colon samples, as well as UC-associated CRC tumor versus adjacent normal tissue. Volcano plots and heatmaps of differential expression results were generated using ggplot2 and pheatmap, respectively. Weighted Gene Co-expression Network Analysis (WGCNA) A gene co-expression network was constructed based on the integrated UC-control dataset to identify gene modules highly correlated with UC. First, the gene expression matrix was filtered to retain the top 25% most variable genes. Outlier samples were identified and removed. The recommended soft-thresholding power (9) was selected to achieve scale-free topology. Modules were detected using the dynamic tree cut method, with a minimum module size of 30 and a merging threshold of 0.25. Visualization of the analysis results was performed using the WGCNA R package. Functional Enrichment Analysis Gene symbols were converted to ENTREZ IDs using the bitr function from the clusterProfiler package (v4.8.2) with the org.Hs.eg.db database (v3.17.0). KEGG pathway enrichment analysis was performed using the enrichKEGG function with the following parameters: organism = "hsa" (Homo sapiens), p-value cutoff = 0.05, p-value adjustment method = Benjamini-Hochberg, and q-value cutoff = 0.2. Significantly enriched pathways were visualized as bubble plots showing enrichment factor, gene count, and adjusted p-value. To illustrate the relationship between genes and pathways, chord diagrams were generated using the circlize package (v0.4.15). All functional enrichment analyses were conducted in R (v4.3.1) following the standardized protocols of the clusterProfiler suite. Immune Microenvironment Analysis The relative proportions of 22 immune cell subtypes in UC versus control colon samples were estimated using the CIBERSORT algorithm with the LM22 signature matrix. CIBERSORT was run with default parameters and 1,000 permutations to ensure robust deconvolution results. Relative immune cell abundances were visualized using stacked bar plots, and correlations among immune cell types were calculated and plotted using the Hmisc package. Group comparisons of immune cell proportions were performed using the Wilcoxon rank-sum test with p-values adjusted for multiple comparisons. Associations between immune cell proportions and gene expression in the yellow and blue WGCNA modules were assessed using Spearman rank correlation analysis. Single-Cell RNA Sequencing Data Analysis Single-cell RNA sequencing (scRNA-seq) datasets used in this study were obtained from the GEO database ( https://www.ncbi.nlm.nih.gov/geo/ ). The datasets include: UC scRNA-seq datasets — GSE214695 (3 UC rectum, 6 Normal), GSE231993 (4 UC, 4 Normal), GSE125527 (7 UC, 8 Normal); CRC scRNA-seq datasets — GSE161277 (4 CRC Tumor, 3 Normal), GSE261388 (3 CRC Tumor, 3 Normal), GSE231559 (6 CRC Tumor, 3 Normal). Data analysis was performed in R (v4.4.1), primarily using the Seurat package (v5.3.0) for quality control, normalization, dimensionality reduction, clustering, and visualization. Each sample underwent quality control with thresholds set according to sequencing characteristics, including number of detected genes (nFeature_RNA), UMI counts (nCount_RNA), mitochondrial gene percentage (percent.mt), and ribosomal gene percentage (percent.ribo), to remove low-quality cells and potential doublets. Filtered data were normalized using the LogNormalize method. Highly variable genes were identified using the FindVariableFeatures function for downstream dimensionality reduction. To integrate multiple samples and correct for batch effects, the Harmony algorithm was applied. Principal component analysis (PCA) was used for linear dimensionality reduction, and UMAP was applied for visualization of the lower-dimensional embedding. Cell type annotation was manually performed based on canonical marker genes, supported by literature and database references. Expression of marker genes and gene sets of interest was visualized using FeaturePlot, VlnPlot (violin plots), and DotPlot functions. Module scores for specific gene sets were calculated using the AddModuleScore function and visualized in violin plots to illustrate gene set expression across cell types. Boxplots were generated using ggplot2, displaying each cluster separately. Boxes represent the median and interquartile range, and group comparisons were conducted using the Wilcoxon rank-sum test. Machine Learning Analysis Data Preprocessing: The GSE3629 dataset, containing expression profiles of UC-associated CRC tumor tissues (n = 6) and adjacent normal tissues (n = 10), was downloaded from the GEO database. Raw data were imported and processed in R. Preprocessing steps included: removing rows with all zero values; log2(x + 1) transformation; for genes mapped to multiple probes, retaining the probe with the highest median expression; quantile normalization using the normalizeBetweenArrays() function from the limma package. The normalized data were used for downstream analysis. Machine Learning: LASSO (Least Absolute Shrinkage and Selection Operator) regression analysis was performed using the glmnet package. The model was specified as binomial (family = "binomial") with a penalty coefficient α = 1. Ten-fold cross-validation (nfolds = 10) was applied to determine the optimal regularization parameter λ. Features corresponding to the minimum mean squared error (lambda.min) and the 1-standard-error simplified model (lambda.1se) were extracted. Final key genes were selected based on the lambda.1se results, and their coefficient directions and magnitudes were visualized using ggplot2. A Random Forest model was further built using the randomForest package. Input variables were the expression levels of the key genes, and the grouping variable was sample type (UC_CRC_Tumor vs. UC_CRC_Rectal), with ntree = 500. Gene importance was evaluated using the importance() function, providing MeanDecreaseAccuracy and MeanDecreaseGini values, and a feature ranking plot was generated. Model performance was assessed using the confusion matrix and the error rate curve (err.rate). Receiver Operating Characteristic (ROC) curves for each key gene were plotted using the pROC package, and the Area Under the Curve (AUC) with 95% confidence intervals was calculated. Gene expression values were standardized using z-score normalization prior to analysis. Finally, a logistic regression model was constructed using the six key genes identified jointly by LASSO and Random Forest (IL1B, FPR2, OSM, GBP1, MT1G, MT1H), and a nomogram was built using the rms package to predict the risk of UC progression to CRC. Survival Analysis Survival analysis was performed using the GEPIA web server ( http://gepia2.cancer-pku.cn ). Disease-Free Survival (RFS) was selected as the outcome. Patients were stratified into high- and low-expression groups using the 70th percentile as the high cutoff and the 30th percentile as the low cutoff. Kaplan-Meier survival curves were generated to evaluate the association between gene expression and CRC patient prognosis. Final figures were assembled and visually enhanced using Adobe Illustrator. Statistical Analysis All statistical analyses were performed in R. Differential expression, enrichment, module-trait correlations, immune cell comparisons, and single-cell module score analyses were conducted using the methods described in the corresponding R scripts. Specific statistical tests, including Pearson or Spearman correlation, Wilcoxon rank-sum test, and LASSO/Random Forest modeling, are described in the figure legends. P-values < 0.05 were considered statistically significant unless otherwise stated. Declarations Acknowledgements We acknowledge the availability of public datasets that made this study possible. Data Availability Statement All datasets analyzed in this study are publicly available. The scripts used for data analysis are available from the corresponding author upon reasonable request. All bulk RNA-seq datasets were downloaded from the Gene Expression Omnibus (GEO)-a public functional genomics data repository supporting MIAME-compliant data submissions-under the accession numbers: GSE66407, GSE158952, GSE193677, GSE235236, GSE245890, and GSE3629. Single-cell RNA sequencing (scRNA-seq) datasets used in this study were obtained from the Gene Expression Omnibus (GEO) database, with the following accession numbers: GSE214695, GSE231993, GSE125527, GSE161277, GSE261388, and GSE231559. For Machine Learning Analysis: The dataset (accession number: GSE3629) was downloaded from the Gene Expression Omnibus (GEO) database. All aforementioned datasets are publicly accessible via the GEO repository’s permanent links: http://www.ncbi.nlm.nih.gov/geo and https://www.ncbi.nlm.nih.gov/geo/. All accession numbers and associated files have been fully released and are available for verification. Authors' contributions CM and XZ contributed to the study conception and design. LY performed data acquisition and preprocessing. JX, YG, and JY conducted the bioinformatic analyses and interpreted the results. PC, CL, and YD assisted with data analysis and visualization. CM drafted the manuscript. ZX supervised the project, revised the manuscript critically for important intellectual content, and confirmed the authenticity of all data. All authors read and approved the final manuscript. All authors read and approved the final version of the manuscript. Funding Declaration This research received no external funding. Ethics, Consent to Participate, and Consent to Publish declarations: not applicable. Competing interests The authors declare that they have no competing interests. Clinical Trial Numbe Clinical trial number: not applicable. References Ng, S.C., et al., Worldwide incidence and prevalence of inflammatory bowel disease in the 21st century: a systematic review of population-based studies. Lancet, 2017. 390 (10114): p. 2769-2778. Winther, K.V., et al., Long-term risk of cancer in ulcerative colitis: a population-based cohort study from Copenhagen County. Clin Gastroenterol Hepatol, 2004. 2 (12): p. 1088-95. Rutter, M.D., et al., Thirty-year analysis of a colonoscopic surveillance program for neoplasia in ulcerative colitis. Gastroenterology, 2006. 130 (4): p. 1030-8. Wangchuk, P., K. Yeshi, and A. Loukas, Ulcerative colitis: clinical biomarkers, therapeutic targets, and emerging treatments. Trends Pharmacol Sci, 2024. 45 (10): p. 892-903. Li, Q., et al., Signaling pathways involved in colorectal cancer: pathogenesis and targeted therapy. Signal Transduct Target Ther, 2024. 9 (1): p. 266. Singh, M., et al., Advancements in combining targeted therapy and immunotherapy for colorectal cancer. Trends Cancer, 2024. 10 (7): p. 598-609. Rogler, G., Chronic ulcerative colitis and colorectal cancer. Cancer Lett, 2014. 345 (2): p. 235-41. Li, Y., et al., Disease-related expression of the IL6/STAT3/SOCS3 signalling pathway in ulcerative colitis and ulcerative colitis-related carcinogenesis. Gut, 2010. 59 (2): p. 227-35. Judes, G., et al., High-throughput <> technologies: New tools for the study of triple-negative breast cancer. Cancer Lett, 2016. 382 (1): p. 77-85. D'Agostino, N., W. Li, and D. Wang, High-throughput transcriptomics. Sci Rep, 2022. 12 (1): p. 20313. Slovin, S., et al., Single-Cell RNA Sequencing Analysis: A Step-by-Step Overview. Methods Mol Biol, 2021. 2284 : p. 343-365. Chi, H., et al., T-cell exhaustion signatures characterize the immune landscape and predict HCC prognosis via integrating single-cell RNA-seq and bulk RNA-sequencing. Front Immunol, 2023. 14 : p. 1137025. Alessi, M.C., et al., FPR2: A Novel Promising Target for the Treatment of Influenza. Front Microbiol, 2017. 8 : p. 1719. Liu, X., et al., Serum amyloid A contributes to radiation-induced lung injury by activating macrophages through FPR2/Rac1/NF-kappaB pathway. Int J Biol Sci, 2024. 20 (12): p. 4941-4956. Wu, M.Y., et al., Enhancement of efferocytosis through biased FPR2 signaling attenuates intestinal inflammation. EMBO Mol Med, 2023. 15 (12): p. e17815. Zarling, J.M., et al., Oncostatin M: a growth regulator produced by differentiated histiocytic lymphoma cells. Proc Natl Acad Sci U S A, 1986. 83 (24): p. 9739-43. Lantieri, F. and T. Bachetti, OSM/OSMR and Interleukin 6 Family Cytokines in Physiological and Pathological Condition. Int J Mol Sci, 2022. 23 (19). Fearon, U., et al., Oncostatin M induces angiogenesis and cartilage degradation in rheumatoid arthritis synovial tissue and human cartilage cocultures. Arthritis Rheum, 2006. 54 (10): p. 3152-62. Hoagland, D.A., et al., Macrophage-derived oncostatin M repairs the lung epithelial barrier during inflammatory damage. Science, 2025. 389 (6756): p. 169-175. Domaniku-Waraich, A., et al., Oncostatin M signaling drives cancer-associated skeletal muscle wasting. Cell Rep Med, 2024. 5 (4): p. 101498. Lee, B.Y., et al., Heterocellular OSM-OSMR signalling reprograms fibroblasts to promote pancreatic cancer growth and metastasis. Nat Commun, 2021. 12 (1): p. 7336. Wei, S., et al., OSM May Serve as a Biomarker of Poor Prognosis in Clear Cell Renal Cell Carcinoma and Promote Tumor Cell Invasion and Migration. Int J Genomics, 2023. 2023 : p. 6665452. Noh, J., et al., Activation of OSM-STAT3 Epigenetically Regulates Tumor-Promoting Transcriptional Programs in Cervical Cancer. Cancers (Basel), 2022. 14 (24). Si, M. and J. Lang, The roles of metallothioneins in carcinogenesis. J Hematol Oncol, 2018. 11 (1): p. 107. Higuera, M., et al., Impact of zinc on hepatocellular carcinoma cell behavior and metallothionein expression: Insights from preclinical models. Biomed Pharmacother, 2025. 185 : p. 117918. Jeong, S.H., et al., MTF1 Is Essential for the Expression of MT1B, MT1F, MT1G, and MT1H Induced by PHMG, but Not CMIT, in the Human Pulmonary Alveolar Epithelial Cells. Toxics, 2021. 9 (9). Additional Declarations No competing interests reported. Supplementary Files TableS1.xlsx FIGS1.tif Supplementary Figure S1. Batch effect removal in bulk transcriptome data. (A) Edwards Venn diagram of five datasets. Each ellipse represents a dataset’s gene set; overlaps indicate shared genes. Red-numbered regions show the number of effective genes after integration. (B)Boxplots showing distribution of gene expression before and after batch effect removal and standardization. Median and interquartile range are indicated. (C, D) PCA plots before (left) and after (right) batch correction. Each point represents a sample; colors indicate sample groups (C) or sources (D). Changes in clustering patterns reflect batch effect correction. FigS2.tif Supplementary Figure S2. Integration and comparison workflow of UC and CRC single-cell datasets. (A–D) Distribution of UC (A, B) and CRC (C, D) single-cell data before and after quality control, illustrating filtering of low-quality cells. (E, F) PCA plots showing sample distribution along principal components (PCs) for UC (E) and CRC (F) datasets. X- and Y-axes represent PCs; colors indicate different sample groups. (G, H) Gene loadings in PCs 1–4 for UC (G) and CRC (H) datasets, highlighting genes with large contributions to each PC. (I, J) Scatter plots of mean expression vs. standardized variance for gene selection. Red dots indicate highly variable genes (HVGs) in UC (I) and CRC (J) datasets. (K, L) Heatmaps of top-loading genes in PCs 1–4 for UC (K) and CRC (L), showing expression differences across cells or samples. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8417562","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":573726418,"identity":"77e3d825-ada1-4e49-b001-5f46037acc8f","order_by":0,"name":"Changming Huang","email":"","orcid":"","institution":"First Affiliated Hospital of Soochow University","correspondingAuthor":false,"prefix":"","firstName":"Changming","middleName":"","lastName":"Huang","suffix":""},{"id":573726422,"identity":"d958ce5f-635c-48fd-9967-2f6ff4b3aec2","order_by":1,"name":"Xian Zhang","email":"","orcid":"","institution":"Jiangsu Province (Suqian) Hospital","correspondingAuthor":false,"prefix":"","firstName":"Xian","middleName":"","lastName":"Zhang","suffix":""},{"id":573726425,"identity":"21957d43-b1b8-4875-bb9d-323b7ec5e096","order_by":2,"name":"Liying Wang","email":"","orcid":"","institution":"Nanjing Medical University","correspondingAuthor":false,"prefix":"","firstName":"Liying","middleName":"","lastName":"Wang","suffix":""},{"id":573726428,"identity":"7471e717-cd3a-4a58-af17-fc45ba0e1e50","order_by":3,"name":"Jingxuan Zhang","email":"","orcid":"","institution":"Xuzhou Medical College","correspondingAuthor":false,"prefix":"","firstName":"Jingxuan","middleName":"","lastName":"Zhang","suffix":""},{"id":573726432,"identity":"7a47181d-2a31-49a5-9e7c-cf832b46e855","order_by":4,"name":"Yang Gao","email":"","orcid":"","institution":"Xuzhou Medical College","correspondingAuthor":false,"prefix":"","firstName":"Yang","middleName":"","lastName":"Gao","suffix":""},{"id":573726436,"identity":"21acee85-03d5-418e-aadb-36a441c61f24","order_by":5,"name":"Jinyue Wang","email":"","orcid":"","institution":"Xuzhou Medical College","correspondingAuthor":false,"prefix":"","firstName":"Jinyue","middleName":"","lastName":"Wang","suffix":""},{"id":573726440,"identity":"5a35857e-0b42-4682-be82-11fd7564e6ca","order_by":6,"name":"Pengcheng Ji","email":"","orcid":"","institution":"Jiangsu Province (Suqian) Hospital","correspondingAuthor":false,"prefix":"","firstName":"Pengcheng","middleName":"","lastName":"Ji","suffix":""},{"id":573726443,"identity":"26f4ede9-d4aa-4441-b910-a087b0cfd244","order_by":7,"name":"Chi Liang","email":"","orcid":"","institution":"Jiangsu Province (Suqian) Hospital","correspondingAuthor":false,"prefix":"","firstName":"Chi","middleName":"","lastName":"Liang","suffix":""},{"id":573726447,"identity":"ece8e0bf-3008-4a5f-8608-de9348be9ea7","order_by":8,"name":"Yuan Ding","email":"","orcid":"","institution":"Jiangsu Province (Suqian) Hospital","correspondingAuthor":false,"prefix":"","firstName":"Yuan","middleName":"","lastName":"Ding","suffix":""},{"id":573726451,"identity":"7ef1b097-fdcc-48df-8752-2ef3a2289fe6","order_by":9,"name":"Zixiang Zhang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7klEQVRIiWNgGAWjYDACZgaGAyDaAER8MLCxI00L44yCtGTibQNpYeb5cIixgZBK+XYewwM/d9Tam7OfPfzaxuAAMwP74aMb8GlhbGZLONh75jizZU9emnWOwR0+Bp60tBt4vcLMfOAAb9sxNoMDOWbGOQbPmBkkeMzwamFjZmw4+LftGI/B+TdmxhYGhxkbCGnhAdpymLetRsLgRo7xYwZitEgwsyUclm07YGBw440ZY49BWjIbIb/I958x/vi2rc7e4HyO8Ycff2zs+NkPH8OrBQoOg/0lASaJUA4CdSCC+QORqkfBKBgFo2CEAQCJ90nMCrUgPgAAAABJRU5ErkJggg==","orcid":"","institution":"First Affiliated Hospital of Soochow University","correspondingAuthor":true,"prefix":"","firstName":"Zixiang","middleName":"","lastName":"Zhang","suffix":""}],"badges":[],"createdAt":"2025-12-21 13:38:39","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8417562/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8417562/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":100371223,"identity":"c48e8fbc-7b97-4979-9121-845d2c84e4a8","added_by":"auto","created_at":"2026-01-16 08:09:40","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":4396199,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.docx","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/3bfb7359bff811cd80e224f5.docx"},{"id":100369376,"identity":"b9a1e603-3a7d-492e-bc53-c4f6fbb26aad","added_by":"auto","created_at":"2026-01-16 07:58:59","extension":"tif","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":12402440,"visible":true,"origin":"","legend":"","description":"","filename":"FIG1.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/7d07647aa57931a4fd2b45ac.tif"},{"id":100211843,"identity":"1a9e02e7-45ce-491f-8def-ef981bd3daf0","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"tif","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":16556948,"visible":true,"origin":"","legend":"","description":"","filename":"FIG3.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/6bbaa04c1775f737bdf362c3.tif"},{"id":100211840,"identity":"f7f46e94-a3e2-4b22-9d94-6eb2784cdfb8","added_by":"auto","created_at":"2026-01-14 07:50:18","extension":"tif","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":9531460,"visible":true,"origin":"","legend":"","description":"","filename":"FIG4.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/77c93c23d8480ef254fb208d.tif"},{"id":100211851,"identity":"c93557e4-5bdd-4fcd-9d65-c2a631b4ead6","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"tif","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6661344,"visible":true,"origin":"","legend":"","description":"","filename":"FIG5.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/42a4ef7664b5923eda8ecc2c.tif"},{"id":100370057,"identity":"bb1c31b2-918e-48f7-b003-9738ed0b3cbf","added_by":"auto","created_at":"2026-01-16 07:59:51","extension":"tif","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":7600496,"visible":true,"origin":"","legend":"","description":"","filename":"Fig2.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/33c4325178ea5eb6c71f487b.tif"},{"id":100211854,"identity":"1cef9baf-da78-42f2-a3c7-392723c95d2d","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"tif","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":2840336,"visible":true,"origin":"","legend":"","description":"","filename":"Fig6.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/05207ef6085158615e0c2ddc.tif"},{"id":100211873,"identity":"8070de5b-1eb9-45d0-ab59-ebce495582b7","added_by":"auto","created_at":"2026-01-14 07:50:20","extension":"tif","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5649508,"visible":true,"origin":"","legend":"","description":"","filename":"Fig7.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/f1c31e7b9f66e601ea6d2d13.tif"},{"id":100211837,"identity":"5af5123b-08c5-436e-83d1-6aaff68ad743","added_by":"auto","created_at":"2026-01-14 07:50:18","extension":"json","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":11293,"visible":true,"origin":"","legend":"","description":"","filename":"9338d1abb68d4c35b7570ed5df9be4d8.json","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/a2338d1320f7bb2674bddcfc.json"},{"id":100370804,"identity":"7753ea19-fde4-472e-b307-45973f1fa287","added_by":"auto","created_at":"2026-01-16 08:08:21","extension":"tif","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":10220124,"visible":true,"origin":"","legend":"","description":"","filename":"FIGS1.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/68a6d817f0e697105a364883.tif"},{"id":100211849,"identity":"462a4270-fe0b-46c6-af73-86c6c551d2c7","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"tif","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":26533116,"visible":true,"origin":"","legend":"","description":"","filename":"FigS2.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/65193493c6edd48c858f0dbf.tif"},{"id":100211845,"identity":"5564499f-847f-4c80-b68c-9b1d58e0da94","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"xlsx","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":20828,"visible":true,"origin":"","legend":"","description":"","filename":"TableS1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/792faeea8dc679a9ec140413.xlsx"},{"id":100211841,"identity":"c3606c94-e937-4a58-aaa0-998d0d284ac7","added_by":"auto","created_at":"2026-01-14 07:50:18","extension":"xml","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":81153,"visible":true,"origin":"","legend":"","description":"","filename":"9338d1abb68d4c35b7570ed5df9be4d81enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/e887d5a7c482329419865fa4.xml"},{"id":100211866,"identity":"500fb78b-6acb-4045-8929-00cba278b919","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"tif","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":12402440,"visible":true,"origin":"","legend":"","description":"","filename":"FIG1.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/414c63f323324c96fb5052fb.tif"},{"id":100370139,"identity":"93e2e9fe-9f4b-4ea6-8bea-a2dae8b990ce","added_by":"auto","created_at":"2026-01-16 08:00:07","extension":"tif","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":16556948,"visible":true,"origin":"","legend":"","description":"","filename":"FIG3.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/f72f68a7ba9d6ff4350f6932.tif"},{"id":100369884,"identity":"e2aeada8-7eb5-42de-836e-b6844721a325","added_by":"auto","created_at":"2026-01-16 07:59:36","extension":"tif","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":9531460,"visible":true,"origin":"","legend":"","description":"","filename":"FIG4.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/bc06a024be78fbee09ea812d.tif"},{"id":100369452,"identity":"d6ef32fa-0519-4c83-91c6-074a839416dd","added_by":"auto","created_at":"2026-01-16 07:59:03","extension":"tif","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6661344,"visible":true,"origin":"","legend":"","description":"","filename":"FIG5.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/9979e4b79af8fa1e965b5678.tif"},{"id":100369853,"identity":"f373b9b5-9a98-4621-a806-1b1f151a8073","added_by":"auto","created_at":"2026-01-16 07:59:33","extension":"tif","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":7600496,"visible":true,"origin":"","legend":"","description":"","filename":"Fig2.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/86020e631e9afeda4725fd50.tif"},{"id":100370267,"identity":"2ac36518-bbfe-4eb8-8297-86305eae9716","added_by":"auto","created_at":"2026-01-16 08:05:09","extension":"tif","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":2840336,"visible":true,"origin":"","legend":"","description":"","filename":"Fig6.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/bac8d956e560a13e7cf749f7.tif"},{"id":100369592,"identity":"02c180c1-768a-49c8-8b57-2e6b64558b10","added_by":"auto","created_at":"2026-01-16 07:59:09","extension":"tif","order_by":19,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5649508,"visible":true,"origin":"","legend":"","description":"","filename":"Fig7.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/cefa9e8e30e77d0dffa2cad6.tif"},{"id":100211848,"identity":"32b10c2e-a701-409e-9413-b3eb57001655","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"jpeg","order_by":20,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":791440,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/7f24496d7a6798d50b1fb25d.jpeg"},{"id":100370499,"identity":"aa348830-ac53-4c3d-9c2e-f65c1d8acbec","added_by":"auto","created_at":"2026-01-16 08:06:09","extension":"jpeg","order_by":21,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":553450,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/4a3da9dc426a9befc41ca639.jpeg"},{"id":100211850,"identity":"a186954f-515d-48ca-9d65-64fac211bf02","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"jpeg","order_by":22,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1300500,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/234e00981134f47c1e975e44.jpeg"},{"id":100211869,"identity":"d37882f9-c85d-4d99-8a33-f94cad35adae","added_by":"auto","created_at":"2026-01-14 07:50:20","extension":"jpeg","order_by":23,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":835088,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/ca2ae9c1881796d966a184d4.jpeg"},{"id":100369868,"identity":"e442bc7b-abbd-4338-9249-56271363d6e8","added_by":"auto","created_at":"2026-01-16 07:59:35","extension":"jpeg","order_by":24,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":568644,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/95215f0240a9ea3a8726993d.jpeg"},{"id":100211878,"identity":"736f6891-d080-4ce3-ab6b-64e87affe8cc","added_by":"auto","created_at":"2026-01-14 07:50:20","extension":"jpeg","order_by":25,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":219348,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage6.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/d47d96f4369e4bbcae211c1f.jpeg"},{"id":100211872,"identity":"fa76c896-3072-400f-8d18-27feae616734","added_by":"auto","created_at":"2026-01-14 07:50:20","extension":"jpeg","order_by":26,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":392778,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage7.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/5d796978ba6e640e76b552c5.jpeg"},{"id":100370479,"identity":"cdb21af7-0b17-4d84-93ca-a717184c40ab","added_by":"auto","created_at":"2026-01-16 08:06:02","extension":"png","order_by":27,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":3124449,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFIG1.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/6ad82ebaa83fb74ecbf2880b.png"},{"id":100211853,"identity":"88fb02e8-26fb-4aef-92f8-08ee654d7c44","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"png","order_by":28,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6044821,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFIG3.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/8daa6722b3d5b757243bf2da.png"},{"id":100211881,"identity":"b585397e-f333-4b5f-8f84-6c142b786ebd","added_by":"auto","created_at":"2026-01-14 07:50:20","extension":"png","order_by":29,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":3980486,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFIG4.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/40a33700c90e9a407efd85ad.png"},{"id":100211886,"identity":"00647775-2f5c-4174-89d9-ca8cee8a0854","added_by":"auto","created_at":"2026-01-14 07:50:20","extension":"png","order_by":30,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":2267056,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFIG5.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/74603349c554c8165fa564ec.png"},{"id":100211863,"identity":"c8b5d7ff-0c91-4d1a-8052-e8804fc50d76","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"png","order_by":31,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":3241503,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFig2.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/b95d933983dd1afc3d952502.png"},{"id":100211877,"identity":"ca3e5118-2b5b-4abe-8696-53d4951a65e7","added_by":"auto","created_at":"2026-01-14 07:50:20","extension":"png","order_by":32,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1003867,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFig6.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/75c9bc3c9c0678045b6ec385.png"},{"id":100211867,"identity":"6226ad37-473c-44e2-86fc-8f503ea63ea4","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"png","order_by":33,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":388565,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFig7.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/3f1d10bcfb60e7a82f019e5a.png"},{"id":100369895,"identity":"09cbac15-8605-49dd-95fe-32b94124301e","added_by":"auto","created_at":"2026-01-16 07:59:36","extension":"png","order_by":34,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":120619,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/bd947bbe82cd70c27c250a1b.png"},{"id":100211879,"identity":"1a76c660-4ef2-4c91-a2e1-8ecf7306b718","added_by":"auto","created_at":"2026-01-14 07:50:20","extension":"png","order_by":35,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":114856,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/d9ed1d7b200128decfe136dc.png"},{"id":100211884,"identity":"9cf82ac6-2710-4fe7-a908-35e6dde790dd","added_by":"auto","created_at":"2026-01-14 07:50:20","extension":"png","order_by":36,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":235852,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/5fde20b358e47704cbf7a65e.png"},{"id":100211875,"identity":"4bf0e912-99dd-40b2-a8e1-c152f64fd5bb","added_by":"auto","created_at":"2026-01-14 07:50:20","extension":"png","order_by":37,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":158881,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/ff53936bee9bc8717371c3e6.png"},{"id":100370343,"identity":"9f40baa3-a749-4b90-af73-940e36aa4c14","added_by":"auto","created_at":"2026-01-16 08:05:29","extension":"png","order_by":38,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":102571,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/008617a0c6ac441a4fd23b7f.png"},{"id":100211880,"identity":"f64be18c-13fa-4d65-9996-1e8fe6e5e365","added_by":"auto","created_at":"2026-01-14 07:50:20","extension":"png","order_by":39,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":44984,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/6061bbf9e51d58c3f9c727f6.png"},{"id":100370468,"identity":"b7b7ef8b-5cca-42fc-ae08-843c6a91b916","added_by":"auto","created_at":"2026-01-16 08:05:54","extension":"png","order_by":40,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":76487,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/220a37ff8ca500e3fd0f03de.png"},{"id":100371005,"identity":"a1eb3ea4-dc7a-4e44-bab0-b24b62647fd6","added_by":"auto","created_at":"2026-01-16 08:09:09","extension":"xml","order_by":41,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":80705,"visible":true,"origin":"","legend":"","description":"","filename":"9338d1abb68d4c35b7570ed5df9be4d81structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/27b576cea64184b91b98b7fa.xml"},{"id":100370486,"identity":"b69ef4f7-9efd-4dff-8fb5-608c3e745c1e","added_by":"auto","created_at":"2026-01-16 08:06:03","extension":"html","order_by":42,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":93739,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/97752781c68e0078b293da34.html"},{"id":100370054,"identity":"4ab66972-db63-4d1d-981b-79c9f3baa037","added_by":"auto","created_at":"2026-01-16 07:59:51","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":5381179,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTranscriptomic features of UC patients.\u003c/strong\u003e \u003cstrong\u003e(A)\u003c/strong\u003e Volcano plot of DEGs between UC patients and healthy controls; red: significantly upregulated (padj \u0026lt; 0.05, logFC ≥ 1), blue: significantly downregulated (padj \u0026lt; 0.05, logFC ≤ –1). \u003cstrong\u003e(B) \u003c/strong\u003eWGCNA soft-threshold selection. Left: scale-free topology fit (R²) versus power; red dashed line: recommended threshold (R² = 0.8). Right: mean connectivity at each power. \u003cstrong\u003e(C) \u003c/strong\u003eGene co-expression modules. Top: hierarchical clustering dendrogram; bottom: modules identified by dynamic tree cut (min. size = 30, color-coded). \u003cstrong\u003e(D)\u003c/strong\u003eHeatmap of module–trait correlations. Rows represent gene modules, and columns represent clinical traits. Pearson correlation coefficients are shown in each cell. The corresponding p-values were calculated using Pearson correlation significance tests, with *p \u0026lt; 0.05, **p \u0026lt; 0.01, ***p \u0026lt; 0.001 indicating increasing levels of statistical significance. Red indicates positive correlation, blue indicates negative correlation, and color intensity reflects correlation strength. \u003cstrong\u003e(E)\u003c/strong\u003eScatter plot of module membership (MM) versus gene significance (GS) for selected modules; MM: correlation with module eigengene; GS: correlation with disease trait.\u003cstrong\u003e (F)\u003c/strong\u003e Venn diagram showing overlap between upregulated DEGs and positively correlated module (yellow), and overlap between downregulated DEGs and negatively correlated module (blue). \u003cstrong\u003e(G)\u003c/strong\u003e Heatmap displaying expression levels of the selected key genes across samples. Color intensity represents relative expression levels.\u003c/p\u003e","description":"","filename":"FIG1.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/dd811f679b602c3fb9e23480.jpg"},{"id":100370278,"identity":"e14ac00e-e1f5-4a14-b108-7e5e1a9bfe44","added_by":"auto","created_at":"2026-01-16 08:05:14","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":4238240,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFunctional enrichment of up_GeneSignature and down_GeneSignature. (A, C)\u003c/strong\u003e Bubble plots showing KEGG pathway enrichment for up_GeneSignature \u003cstrong\u003e(A)\u003c/strong\u003eand down_GeneSignature \u003cstrong\u003e(C)\u003c/strong\u003e. Y-axis: pathway names; X-axis: enrichment factor (Rich Factor). Bubble size represents the number of enriched genes, and color indicates significance (-log10 adjusted p-value). Only the top 20 pathways are shown. \u003cstrong\u003e(B, D)\u003c/strong\u003e Chord diagrams linking key genes to enriched KEGG pathways for up_GeneSignature \u003cstrong\u003e(B)\u003c/strong\u003e and down_GeneSignature \u003cstrong\u003e(D)\u003c/strong\u003e. Left sectors: enriched pathways; right sectors: genes in those pathways. Connecting chords indicate significant enrichment, highlighting shared key genes across multiple pathways.\u003c/p\u003e","description":"","filename":"Fig2.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/9ec2aa47f4d7b41a50b20623.jpg"},{"id":100211834,"identity":"e5da02b9-5937-4a26-a8f2-ecdf533f66e9","added_by":"auto","created_at":"2026-01-14 07:50:18","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":8881076,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eImmune microenvironment analysis in UC patients. (A)\u003c/strong\u003e Relative proportions of 22 immune cell subtypes in each sample. Bars represent individual samples grouped as Control and UC. Colors indicate different immune cell types.\u003cstrong\u003e (B) \u003c/strong\u003eHeatmap showing pairwise Spearman correlations between immune cell types across samples. Red: positive correlation, blue: negative; color intensity reflects correlation strength.\u003cstrong\u003e(C) \u003c/strong\u003eBoxplots of immune cells with significantly different infiltration between UC and Control groups. Statistical significance assessed by Wilcoxon rank-sum test (*p ≤ 0.05, **p \u0026lt; 0.01, ***p \u0026lt; 0.001).\u003cstrong\u003e (D, E) \u003c/strong\u003eDiagonal heatmaps of Spearman correlations between key module genes and immune cell infiltration. Numbers in the lower half of each cell indicate correlation coefficients; numbers in parentheses in the upper half indicate p-values.\u003c/p\u003e","description":"","filename":"FIG3.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/8e3615c9a77353db3254b94d.jpg"},{"id":100370086,"identity":"389b06f9-d915-4e5e-82cd-863735351ad1","added_by":"auto","created_at":"2026-01-16 07:59:57","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":5701204,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSingle-cell transcriptomic features in UC and CRC patients.\u003c/strong\u003e\u003cbr\u003e\n \u003cstrong\u003e(A–D)\u003c/strong\u003e Batch effect correction before and after integration for UC (A, B) and CRC (C, D) datasets. Each point represents a single cell, colored by sample group. \u003cstrong\u003e(E, F)\u003c/strong\u003e UMAP plots of UC (E) and CRC (F) datasets, showing annotated cell types and spatial distributions. \u003cstrong\u003e(G, H)\u003c/strong\u003e Feature plots of canonical marker genes in major cell types for UC (G) and CRC (H). \u003cstrong\u003e(I, J)\u003c/strong\u003eDot plots showing marker gene expression across cell clusters used for annotation in UC (I) and CRC (J). \u003cstrong\u003e(K, L)\u003c/strong\u003e UMAP plots comparing cellular composition between normal and disease samples in UC (K) and CRC (L). \u003cstrong\u003e(M, N)\u003c/strong\u003eBar plots showing proportion changes of major cell types between normal and disease groups in UC (M) and CRC (N). \u003cstrong\u003e(O, P)\u003c/strong\u003e Violin plots of up_GeneSignature scores across cell types in UC (O) and CRC (P). Scores were calculated using the Seurat AddModuleScore function, with the y-axis representing the module score for each cell and the x-axis representing cell types.\u003c/p\u003e","description":"","filename":"FIG4.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/2b0cc12ae3e90942542b061d.jpg"},{"id":100370200,"identity":"9482d3b5-e4df-4386-abbb-41280d58f99d","added_by":"auto","created_at":"2026-01-16 08:00:28","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":4715394,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSingle-cell transcriptomic features of myeloid cells in UC and CRC patients. (A)\u003c/strong\u003e UMAP of myeloid cells in UC dataset after secondary annotation. Right panel shows UMAP colored by normal and disease groups. \u003cstrong\u003e(B)\u003c/strong\u003e Bar plot showing proportions of myeloid cell subtypes in UC dataset between normal and disease groups. \u003cstrong\u003e(C) \u003c/strong\u003eDot plot showing expression levels of marker genes used for annotating myeloid cell subtypes in UC dataset. Dot size represents fraction of cells expressing the gene; color intensity represents average expression. \u003cstrong\u003e(D) \u003c/strong\u003eViolin plot of up_GeneSignature scores across myeloid cell subtypes in UC dataset. Scores were calculated using Seurat AddModuleScore function; y-axis represents module score per cell, x-axis represents myeloid subtypes. \u003cstrong\u003e(E) \u003c/strong\u003eBox plot showing distribution of up_GeneSignature scores across myeloid subtypes and between normal/disease groups in UC dataset. \u003cstrong\u003e(F)\u003c/strong\u003e UMAP of myeloid cells in CRC dataset after secondary annotation, with right panel showing normal vs tumor group distributions. \u003cstrong\u003e(G)\u003c/strong\u003e Bar plot showing proportions of myeloid cell subtypes in CRC dataset between normal and tumor groups. \u003cstrong\u003e(H)\u003c/strong\u003e Dot plot showing expression of marker genes across myeloid subtypes in CRC dataset. \u003cstrong\u003e(I) \u003c/strong\u003eViolin plot showing up_GeneSignature scores across myeloid subtypes in CRC dataset. \u003cstrong\u003e(J)\u003c/strong\u003e Box plot showing up_GeneSignature scores across myeloid subtypes and between normal/tumor groups in CRC dataset. \u003cstrong\u003e(K)\u003c/strong\u003e Venn diagram showing overlap between UC up_GeneSignature genes and genes significantly upregulated in CRC myeloid cells (tumor vs normal). \u003cstrong\u003e(L) \u003c/strong\u003eVolcano plot of differentially expressed genes in CRC myeloid cells (tumor vs normal). Red: significantly upregulated genes; blue: significantly downregulated genes; gray: non-significant genes. \u003cstrong\u003e(M) \u003c/strong\u003eViolin plots showing expression levels of 16 overlapping genes across myeloid cell subtypes in CRC dataset. Y-axis represents normalized expression, and x-axis represents myeloid subtypes.\u003c/p\u003e","description":"","filename":"FIG5.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/e60582afe0e4760890fce569.jpg"},{"id":100211857,"identity":"5b98ff0f-817b-4674-bf6a-7ed21f468086","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":2038870,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSingle-cell expression features of down_GeneSignature in UC and CRC patients.\u003c/strong\u003e \u003cstrong\u003e(A, B)\u003c/strong\u003eViolin plots showing down_GeneSignature module scores across cell types in UC (A) and CRC (B) single-cell datasets. Module scores were calculated using Seurat AddModuleScore function to reflect overall expression of the gene set in each cell. X-axis represents cell types, Y-axis represents module scores. Boxplots overlaid on the violins display median and interquartile range. \u003cstrong\u003e(C, D)\u003c/strong\u003e Box plots showing expression differences of down_GeneSignature across cell types between normal and disease groups in UC (C) and CRC (D) datasets. X-axis represents sample groups; Y-axis represents expression levels (module scores). Statistical significance was assessed using Wilcoxon rank-sum test (wilcox.test). \u003cstrong\u003e(E)\u003c/strong\u003e Venn diagram showing overlap between down_GeneSignature genes in UC and genes significantly downregulated in CRC epithelial cells (tumor vs normal). \u003cstrong\u003e(F)\u003c/strong\u003e Volcano plot of differentially expressed genes in CRC epithelial cells (tumor vs normal). Red: significantly upregulated genes; blue: significantly downregulated genes; gray: non-significant genes.\u003c/p\u003e","description":"","filename":"Fig6.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/ae27f7a6012c0a1106fb3c60.jpg"},{"id":100370682,"identity":"b37346b6-9d72-42f0-8769-5de56a402a5a","added_by":"auto","created_at":"2026-01-16 08:07:23","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":4471807,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eMachine learning-based identification and validation of key genes in UC-to-CRC progression.\u003c/strong\u003e \u003cstrong\u003e(A)\u003c/strong\u003e LASSO regression coefficient path plot showing how gene coefficients change across different λ values. X-axis: log(λ); Y-axis: regression coefficients. Each line represents a gene. \u003cstrong\u003e(B)\u003c/strong\u003e Ten-fold cross-validation curve for LASSO. Red dashed line: λ with minimum cross-validated error (lambda.min); blue dashed line: 1-SE λ (lambda.1se), representing a parsimonious model. \u003cstrong\u003e(C)\u003c/strong\u003e Random forest model error rate curve showing how the out-of-bag error changes with increasing number of decision trees. \u003cstrong\u003e(D)\u003c/strong\u003e Gene importance ranking from random forest, including MeanDecreaseAccuracy and MeanDecreaseGini, indicating each gene’s contribution to classification accuracy and node purity. \u003cstrong\u003e(E)\u003c/strong\u003e ROC curves of six key genes (IL1B, FPR2, OSM, GBP1, MT1G, MT1H) distinguishing UC adjacent tissue from tumor tissue. AUC indicates predictive performance (closer to 1 = higher accuracy). \u003cstrong\u003e(F)\u003c/strong\u003e Nomogram based on six key genes to predict the probability of UC progression to CRC. Each gene contributes a weighted score to the total risk score, which corresponds to predicted probability of disease progression. \u003cstrong\u003e(G–J)\u003c/strong\u003e Kaplan-Meier survival curves for MT1G (G), MT1H (H), FPR2 (I), and OSM (J) in CRC patients, stratified by high vs. low expression. Log-rank P tests differences between groups (smaller P = stronger evidence). HR(high) compares high-expression to low-expression groups: HR \u0026gt; 1 indicates higher risk; HR \u0026lt; 1 indicates protective effect.\u003c/p\u003e","description":"","filename":"Fig7.tif.jpg","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/3746b5dc30e596d28ee8f1f3.jpg"},{"id":102906103,"identity":"0e60b30a-4f50-45b7-9118-31e2c9386dda","added_by":"auto","created_at":"2026-02-18 09:12:31","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":36572868,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/c0a04f8f-17fd-429c-94cb-eb270f731c50.pdf"},{"id":100370007,"identity":"cd337d5b-5504-4bac-94c1-837088f24110","added_by":"auto","created_at":"2026-01-16 07:59:45","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":20828,"visible":true,"origin":"","legend":"","description":"","filename":"TableS1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/536fd8a967780ac4b1f23f22.xlsx"},{"id":100211861,"identity":"4aab689f-218a-4f16-8ddc-b39b432f9215","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"tif","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":10220124,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupplementary Figure S1. Batch effect removal in bulk transcriptome data.\u003c/strong\u003e \u003cstrong\u003e(A)\u003c/strong\u003e Edwards Venn diagram of five datasets. Each ellipse represents a dataset’s gene set; overlaps indicate shared genes. Red-numbered regions show the number of effective genes after integration. \u003cstrong\u003e(B)\u003c/strong\u003eBoxplots showing distribution of gene expression before and after batch effect removal and standardization. Median and interquartile range are indicated. \u003cstrong\u003e(C, D)\u003c/strong\u003e PCA plots before (left) and after (right) batch correction. Each point represents a sample; colors indicate sample groups (C) or sources (D). Changes in clustering patterns reflect batch effect correction.\u003c/p\u003e","description":"","filename":"FIGS1.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/c0453125316d946aae45b9dd.tif"},{"id":100211855,"identity":"cab4133f-8c0e-45b5-b06e-17d256582c90","added_by":"auto","created_at":"2026-01-14 07:50:19","extension":"tif","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":26533116,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupplementary Figure S2. Integration and comparison workflow of UC and CRC single-cell datasets.\u003c/strong\u003e \u003cstrong\u003e(A–D)\u003c/strong\u003e Distribution of UC (A, B) and CRC (C, D) single-cell data before and after quality control, illustrating filtering of low-quality cells. \u003cstrong\u003e(E, F)\u003c/strong\u003e PCA plots showing sample distribution along principal components (PCs) for UC (E) and CRC (F) datasets. X- and Y-axes represent PCs; colors indicate different sample groups. \u003cstrong\u003e(G, H)\u003c/strong\u003e Gene loadings in PCs 1–4 for UC (G) and CRC (H) datasets, highlighting genes with large contributions to each PC. \u003cstrong\u003e(I, J)\u003c/strong\u003e Scatter plots of mean expression vs. standardized variance for gene selection. Red dots indicate highly variable genes (HVGs) in UC (I) and CRC (J) datasets. \u003cstrong\u003e(K, L)\u003c/strong\u003e Heatmaps of top-loading genes in PCs 1–4 for UC (K) and CRC (L), showing expression differences across cells or samples.\u003c/p\u003e","description":"","filename":"FigS2.tif","url":"https://assets-eu.researchsquare.com/files/rs-8417562/v1/3fe173a6180acb7877f671ed.tif"}],"financialInterests":"No competing interests reported.","formattedTitle":"Integrative Bioinformatics of Bulk and Single-Cell Transcriptomes Identifies Predictive Molecular Features of Ulcerative Colitis-Associated Colorectal Cancer","fulltext":[{"header":"Introduction","content":"\u003cp\u003eInflammatory bowel disease (IBD) is a prevalent global disorder in the 21st century\u003csup\u003e\u003cb\u003e[1]\u003c/b\u003e\u003c/sup\u003e. Among IBD subtypes, ulcerative colitis (UC) is a chronic inflammatory condition of the colon characterized by recurrent mucosal inflammation and epithelial injury. Patients with long-standing UC have an increased risk of developing colorectal cancer (CRC), which represents a major cause of cancer-related morbidity and mortality worldwide\u003csup\u003e\u003cb\u003e[2, 3]\u003c/b\u003e\u003c/sup\u003e. Although the elevated cancer risk in UC has been recognized clinically, the molecular mechanisms driving the transition from chronic inflammation to malignant transformation remain incompletely understood. Because UC is incurable, affected patients often have reduced life expectancy and are at heightened risk for CRC. Recent studies have reported numerous new findings regarding UC biomarkers and clinical manifestations\u003csup\u003e\u003cb\u003e[4]\u003c/b\u003e\u003c/sup\u003e, as well as therapeutic targets and strategies for CRC\u003csup\u003e\u003cb\u003e[5, 6]\u003c/b\u003e\u003c/sup\u003e. However, analyses exploring the relationship between UC and CRC, particularly sporadic CRC in patients without chronic colitis, remain limited; most studies have focused on UC-associated CRC (UC-CRC)\u003csup\u003e\u003cb\u003e[7, 8]\u003c/b\u003e\u003c/sup\u003e. Although UC-CRC and sporadic CRC display distinct pathological features, whether these two conditions share molecular correlations or overlapping transcriptional programs is still unclear. Identifying common molecular signatures between UC and sporadic CRC is crucial for understanding mechanisms that may predispose UC patients to malignant transformation. Such shared programs or convergent cellular pathways could provide novel insights into the early events of inflammation-driven tumorigenesis and highlight potential biomarkers or therapeutic targets for predicting and preventing UC-associated CRC.\u003c/p\u003e \u003cp\u003eAdvances in high-throughput transcriptomic technologies have enabled the systematic identification of disease-associated molecular signatures in complex disorders\u003csup\u003e\u003cb\u003e[9]\u003c/b\u003e\u003c/sup\u003e. Bulk RNA sequencing allows the detection of differentially expressed genes (DEGs) between diseased and normal tissues\u003csup\u003e\u003cb\u003e[10]\u003c/b\u003e\u003c/sup\u003e, providing insights into the global molecular alterations associated with UC. However, bulk transcriptomic approaches are limited in their capacity to resolve cell type-specific expression patterns, which are essential for understanding the heterogeneous cellular composition of the inflamed colon. Single-cell RNA sequencing (scRNA-seq) has emerged as a powerful method to dissect cellular heterogeneity and to characterize cell type–specific transcriptional programs\u003csup\u003e\u003cb\u003e[11]\u003c/b\u003e\u003c/sup\u003e. By integrating bulk and single-cell transcriptomic data, it is possible to capture both global and cell-resolved gene expression changes\u003csup\u003e\u003cb\u003e[12]\u003c/b\u003e\u003c/sup\u003e, thereby facilitating the identification of key molecular features shared between UC and CRC.\u003c/p\u003e \u003cp\u003eIn this study, we first identified differentially expressed genes (DEGs) between UC and healthy colonic tissues using bulk RNA-seq data. We then evaluated the expression patterns of these DEGs in both UC and sporadic CRC single-cell RNA-seq datasets to investigate conserved cellular and molecular features. Finally, we aimed to pinpoint potential key genes that may drive the progression from UC to CRC, thereby highlighting candidate biomarkers for early detection and molecular targets for therapeutic intervention. By integrating multi-scale transcriptomic data, this study establishes a comprehensive framework for elucidating the molecular interplay between chronic inflammation and colorectal tumorigenesis.\u003c/p\u003e"},{"header":"Result","content":"\u003cp\u003e \u003cb\u003e1. Key gene sets underlying transcriptional alterations in UC patients\u003c/b\u003e \u003c/p\u003e\u003cp\u003eTo systematically evaluate the transcriptional characteristics of ulcerative colitis (UC) patients, a total of 670 samples from five independent datasets were included in this study. These datasets shared more than 15,000 overlapping genes (Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eA). After removing batch effects (Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eB–D), differential expression analysis revealed a substantial number of significantly dysregulated genes in rectal tissues of UC patients compared with normal controls (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA). Specifically, 193 genes were markedly upregulated and 144 genes were downregulated (P \u0026lt; 0.05, |log2FC| ≥ 1), indicating prominent transcriptional alterations in diseased tissues. To link these transcriptomic changes with disease phenotypes, we further identified disease-related genes using weighted gene co-expression network analysis (WGCNA) in combination with clinical variables (AGE, DISEASE, GENDER, and ANCESTRY), aiming to uncover potential biomarkers and underlying mechanisms. The optimal soft-thresholding power for network construction was determined based on algorithmic recommendations (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB), resulting in the identification of 12 co-expression modules with distinct expression patterns (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC). Module–trait correlation analysis showed that the yellow module exhibited the strongest positive correlation with UC (132 genes), while the blue module showed the strongest negative correlation (595 genes) (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD). Module Membership (MM) versus Gene Significance (GS) analysis indicated moderate overall correlations between module eigengenes and disease traits (r \u0026gt; 0.5) for both modules (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eE). Therefore, the yellow and blue modules were selected for subsequent functional analyses. Complete gene lists for these two modules are provided in Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e. Among the genes significantly upregulated in UC, 60 overlapped with the disease-associated yellow module, whereas 69 downregulated genes overlapped with the negatively correlated blue module (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eF). To visualize the expression patterns of these key genes, we generated a heatmap of 129 representative genes across all samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eG).\u003c/p\u003e\u003cp\u003eFunctional enrichment analysis revealed that the 60 upregulated, disease-associated genes (up_GeneSignature) were predominantly involved in immune and inflammatory responses, including cytokine–receptor interactions and pathways such as IL-17, TNF, and NF-κB signaling, as well as chemokine, Toll-like receptor, and NOD-like receptor signaling pathways (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA,B). These pathways are closely linked to immune regulation and inflammatory injury, consistent with the known pathogenesis of UC. In contrast, the 69 downregulated, negatively correlated genes (down_GeneSignature) were mainly enriched in metabolic and transport-related pathways, including mineral absorption, nitrogen metabolism, PPAR signaling, and AMPK signaling (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC,D), suggesting that this gene set may play a critical role in maintaining metabolic homeostasis under physiological conditions.\u003c/p\u003e\u003cp\u003e \u003cb\u003e2. Immune cell infiltration characteristics in UC patients\u003c/b\u003e \u003c/p\u003e\u003cp\u003eTo further characterize the immune microenvironment in ulcerative colitis (UC), we analyzed immune cell infiltration patterns based on the upregulated gene set (up_GeneSignature) and its associated enrichment pathways. The immune infiltration analysis revealed substantial heterogeneity in immune cell composition across samples and between groups, with T cells and macrophages showing notably higher infiltration levels in UC tissues (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA). Correlation analysis among immune cell subsets demonstrated complex intercellular interactions (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB). Specifically, resting mast cells, NK cells, and CD4⁺ memory T cells were negatively correlated with their activated counterparts. M0/M1 macrophages exhibited negative correlations with B cells, activated mast cells, and activated CD4⁺ memory T cells, but positive correlations with their resting states, suggesting that macrophage polarization may be closely linked to the activation states of T cells and mast cells. Moreover, M2 macrophages showed negative correlations with B cells and monocytes(Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB). Myeloid cells, including macrophages, mast cells, and neutrophils, displayed marked infiltration differences between UC and control tissues (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC), underscoring their potential central roles in disease progression. Notably, expression levels of the upregulated 60-gene signature were strongly correlated with the infiltration of macrophages and dendritic cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD), implying their involvement in immune cell recruitment and activation. In contrast, the downregulated 69-gene set (down_GeneSignature) showed only weak associations with immune infiltration (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eE). Collectively, these findings indicate that aberrant activation of macrophages, dendritic cells, and mast cells may contribute to persistent mucosal inflammation in UC, and that the upregulated gene signature likely represents a core immune regulatory network driving disease pathogenesis.\u003c/p\u003e\u003cp\u003e \u003cb\u003e3. Comparative single-cell transcriptomic profiling of UC and CRC patients\u003c/b\u003e \u003c/p\u003e\u003cp\u003eSingle-cell transcriptomic data from UC and CRC patients were integrated to compare cellular composition and transcriptional patterns. After quality control (Fig. \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003eA–D), identification of highly variable genes (Fig. \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003eI, J), PCA-based dimensionality reduction (Fig. \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003eE–H, K, L), and batch correction (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA–D), the datasets achieved good consistency and comparability. Clustering and cell-type annotation identified major cell populations, including T cells, B cells, myeloid cells, endothelial cells, and epithelial cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eE, F). Feature and dot plots confirmed the accuracy of annotation through canonical markers such as CD3D (T cells), CD19 (B cells), CD68 (macrophages), EPCAM (epithelial cells), and PECAM1 (endothelial cells) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eG–J). Compared with normal tissues, UC samples displayed a marked increase in immune cell infiltration and a reduction in stromal components such as fibroblasts and epithelial cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eK, M). NK cells, B cells, and myeloid cells were significantly enriched, consistent with inflammation-driven immune activation and tissue remodeling. In contrast, CRC tissues exhibited a distinct pattern: NK and B cell proportions decreased, whereas myeloid cells remained elevated and epithelial cells expanded markedly within tumor regions (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eL, N). Despite opposite trends in immune composition, UC and CRC shared similar myeloid transcriptional signatures, suggesting conserved activation states under inflammatory and tumor microenvironments. Notably, the 60 upregulated genes from the UC signature (up_GeneSignature) showed the highest enrichment scores in myeloid cells across both diseases (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eO, P), indicating that these cells are likely the principal effectors mediating immune microenvironment remodeling in UC and CRC.\u003c/p\u003e\u003cp\u003e \u003cb\u003e(A–D)\u003c/b\u003e Batch effect correction before and after integration for UC (A, B) and CRC (C, D) datasets. Each point represents a single cell, colored by sample group. \u003cb\u003e(E, F)\u003c/b\u003e UMAP plots of UC (E) and CRC (F) datasets, showing annotated cell types and spatial distributions. \u003cb\u003e(G, H)\u003c/b\u003e Feature plots of canonical marker genes in major cell types for UC (G) and CRC (H). \u003cb\u003e(I, J)\u003c/b\u003e Dot plots showing marker gene expression across cell clusters used for annotation in UC (I) and CRC (J). \u003cb\u003e(K, L)\u003c/b\u003e UMAP plots comparing cellular composition between normal and disease samples in UC (K) and CRC (L). \u003cb\u003e(M, N)\u003c/b\u003e Bar plots showing proportion changes of major cell types between normal and disease groups in UC (M) and CRC (N). \u003cb\u003e(O, P)\u003c/b\u003e Violin plots of up_GeneSignature scores across cell types in UC (O) and CRC (P). Scores were calculated using the Seurat AddModuleScore function, with the y-axis representing the module score for each cell and the x-axis representing cell types.\u003c/p\u003e\u003cp\u003e \u003cb\u003e4. Comparative transcriptional features of myeloid cells in UC and CRC patients\u003c/b\u003e \u003c/p\u003e\u003cp\u003eTo further explore myeloid transcriptional features in UC and CRC, myeloid cells were re-clustered and subtyped into M1 macrophages, M2 macrophages, dendritic cells (DCs), neutrophils, and mast cells. In UC tissues, macrophages were the dominant myeloid population, with M1 macrophages being the most abundant (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA). In the UC group, the proportion of M1 macrophages showed no obvious change, whereas M2 macrophages were significantly upregulated; DCs and neutrophils exhibited an increasing trend, and mast cells decreased (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB). Dot plot analysis confirmed marker expression across clusters and indicated that M2 macrophages display gene expression features resembling tumor-associated macrophages (TAMs), suggesting possible roles in modulating the inflammatory microenvironment (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eC). The 60-gene up_GeneSignature was highly expressed in macrophages and neutrophils (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eD). Compared with controls, UC samples showed significant upregulation of this signature in M2 macrophages, mast cells, DCs, and neutrophils, while M1 macrophages showed a downward trend (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eE). In CRC, myeloid compartments were also dominated by macrophages, with a notable increase in TAM-like macrophages compared with UC (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eF). Subclassification in CRC showed increases across the three macrophage subsets and a decrease in neutrophils (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eG). Marker profiles highlighted pronounced M2/TAM features in TAM macrophages (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eH). The up_GeneSignature exhibited overall high expression in M1, M2 macrophages, and neutrophils, consistent with the expression pattern observed in UC samples(Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eI), and relative to controls, CRC samples exhibited significant increases of this signature in macrophages, neutrophils, and DCs; the M2, TAM phenotype was more prominent in CRC than in UC (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eJ). Single-cell differential expression comparing CRC myeloid cells to controls identified 16 significantly upregulated genes that overlap with the UC-derived up_GeneSignature (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eK–L). Violin plots show the expression differences of these 16 genes across myeloid subsets and groups (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eM). Overall, both UC and CRC myeloid cells show activation of M2-like programs and related inflammatory pathways, with a stronger TAM signature in CRC. The UC-derived up_GeneSignature remains enriched in myeloid cells across both diseases, suggesting shared immune-regulatory roles and potential predictive value.\u003c/p\u003e\u003cp\u003e \u003cb\u003e5. Comparative transcriptional features of epithelial cells in UC and CRC patients\u003c/b\u003e \u003c/p\u003e\u003cp\u003eWe next examined the expression patterns of the down_GeneSignature identified from UC transcriptomes in single-cell samples of UC and CRC. The analysis revealed that this gene set was predominantly expressed in epithelial cells in both UC and CRC samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eA, B). Specifically, compared with controls, epithelial cells from UC and CRC patients exhibited a significant overall downregulation of down_GeneSignature genes, whereas in immune cells these genes were largely significantly upregulated (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eC, D), consistent with our earlier immune infiltration analysis showing low correlation between down_GeneSignature and immune cell abundance. Further analysis indicated that 21 down_GeneSignature genes were significantly downregulated in epithelial cells of CRC tumor samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eE), with their distribution visualized in a volcano plot (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003eF). Taken together, the expression patterns across UC and CRC suggest that these 21 downregulated genes in CRC epithelial cells may have potential predictive value.\u003c/p\u003e\u003cp\u003e \u003cb\u003e6. Machine learning–based identification and validation of key genes in UC-to-CRC progression\u003c/b\u003e \u003c/p\u003e\u003cp\u003eTo further identify and validate the key roles of up_GeneSignature and down_GeneSignature genes in the progression from UC to secondary CRC, we employed the GSE3629 dataset, which includes RNA-seq data from UC-related CRC tumor tissues and adjacent non-tumor tissues. This dataset allows us to assess whether these gene signatures maintain continuous involvement during the transition from UC to CRC. Using a 10-fold cross-validation strategy, the LASSO model was applied (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eA, B), resulting in six genes with significant discriminative power: IL1B, FPR2, OSM, GBP1, MT1G, and MT1H. The directions of their regression coefficients indicated that IL1B, FPR2, OSM, and GBP1 acted as potential risk factors, whereas MT1G and MT1H might function as protective factors. Random forest analysis confirmed these six genes as important features (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eC, D). ROC curves were plotted for each of the six candidate genes, and all exhibited AUC values above 0.9, indicating high diagnostic potential for distinguishing CRC tumor tissues from adjacent tissues (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eE). A logistic regression model constructed using these six genes demonstrated good overall fit, and the nomogram (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eF) provides a quantitative reference for evaluating the contribution of each gene to UC-derived CRC risk in clinical prediction. Survival analysis revealed that MT1G and MT1H, two down_GeneSignature genes, showed a trend toward worse relapse-free survival (RFS) in the low-expression group, although the log-rank P values did not reach statistical significance; the hazard ratios (HRs for the high-expression group) were \u0026lt; 1, consistent with their potential protective roles (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eG, H). In contrast, among up_GeneSignature genes, FPR2 and OSM exhibited a trend toward reduced survival in the high-expression group, with HRs \u0026gt; 1 (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eI, J). IL1B and GBP1 did not show a survival trend, which may reflect their predictive relevance being restricted to myeloid cells, as observed in Section 4, rather than at the bulk tissue level.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eAlthough UC-associated colorectal cancer (UC-CRC) and sporadic CRC exhibit distinct clinical and pathological characteristics, our study identified shared transcriptional features between these two conditions. Differentially expressed genes (DEGs) observed in UC patients relative to healthy controls were also significantly altered in CRC patients, particularly within immune and epithelial cell populations. This overlap in transcriptomic alterations suggests that certain molecular programs initiated during chronic colonic inflammation in UC may persist and contribute to the early events of colorectal tumorigenesis. Notably, among these DEGs, we identified potential predictive genes for UC-CRC, including MT1G, MT1H, and FPR2. These genes appear to be continuously dysregulated from the inflammatory stage in UC through the progression toward CRC and exhibit consistent changes in sporadic CRC as well. This observation indicates that a subset of genes may represent a molecular continuum, linking chronic inflammation with subsequent malignant transformation, and highlights the possibility that these genes may serve as early biomarkers or mechanistic mediators of inflammation-driven colorectal carcinogenesis. Furthermore, the consistent transcriptomic alterations in immune and epithelial cells imply that both cell-intrinsic changes and the immune microenvironment may play key roles in mediating the transition from chronic inflammation to tumorigenesis. Future studies exploring the functional roles of these shared DEGs and their pathways could provide mechanistic insights and potentially inform strategies for early detection or targeted intervention in patients at high risk for UC-CRC.\u003c/p\u003e \u003cp\u003eBy integrating single-cell and bulk transcriptomic data, we extended our analysis to the cellular level and observed that many genes significantly dysregulated in UC also exhibited pronounced alterations in CRC, albeit with strong cell type specificity. Interestingly, DEGs identified solely at the bulk transcriptome level in UC did not show substantial changes when assessed across the overall transcriptome of sporadic CRC, which may result in their underappreciation in previous studies. However, the single-cell perspective revealed that UC-associated upregulated genes were predominantly correlated with immune cells, particularly myeloid populations such as macrophages, neutrophils, and dendritic cells, indicating that these cell types are critical contexts in which these genes exert their effects. This finding suggests that the high expression of these genes in specific immune cell subsets may contribute to disease exacerbation. Conversely, UC-associated downregulated genes were significantly enriched in epithelial cells, suggesting a potential role in negatively regulating disease progression. Although the precise contribution of these genes to CRC pathogenesis requires further functional validation, their consistent cell-specific dysregulation in both UC and sporadic CRC highlights the importance of considering cellular resolution when investigating inflammation-driven tumorigenesis.\u003c/p\u003e \u003cp\u003eImportantly, we identified four genes with potential predictive value for UC-CRC progression. Due to the lack of suitable single-cell data for UC-associated CRC, these candidates were evaluated using bulk UC-CRC datasets through machine learning and prognostic analyses. This approach underscores a key advantage of our study: genes that do not appear significantly altered at the bulk transcriptome level may still exert critical effects through cell type\u0026ndash;specific expression changes, and such effects can be captured using integrated single-cell analyses. Our findings provide a framework for identifying candidate molecular mediators of UC-associated CRC progression and illustrate the value of combining bulk and single-cell transcriptomics to uncover subtle but biologically meaningful signals.\u003c/p\u003e \u003cp\u003eWe observed that FPR2 and OSM are significantly upregulated in immune cells and represent shared expression features between UC and sporadic CRC, serving as potential predictive genes for UC-associated CRC. FPR2 encodes a G protein-coupled receptor that mediates chemotaxis and modulates inflammatory responses. Previous studies have linked FPR2 upregulation to dysregulated immune responses and disease pathogenesis, including influenza infection\u003csup\u003e\u003cb\u003e[13]\u003c/b\u003e\u003c/sup\u003e, pulmonary fibrosis\u003csup\u003e\u003cb\u003e[14]\u003c/b\u003e\u003c/sup\u003e, and intestinal inflammation\u003csup\u003e\u003cb\u003e[15]\u003c/b\u003e\u003c/sup\u003e. In the context of UC and CRC, persistent FPR2 upregulation in myeloid cells, such as macrophages, neutrophils, and dendritic cells, may sustain chronic inflammatory signaling, promoting an environment conducive to tumor initiation and progression. This observation highlights FPR2 as a potential biomarker and suggests that targeting FPR2-mediated pathways could modulate inflammation-driven carcinogenesis. Oncostatin M (OSM) is a pleiotropic cytokine involved in regulating diverse inflammatory processes, including tissue repair, liver regeneration, and bone remodeling\u003csup\u003e\u003cb\u003e[16, 17]\u003c/b\u003e\u003c/sup\u003e. Since its discovery in 1986, OSM has been implicated as a key mediator in various inflammatory conditions and cancers, including arthritis\u003csup\u003e\u003cb\u003e[18]\u003c/b\u003e\u003c/sup\u003e, inflammatory bowel disease, pulmonary and skeletal disorders\u003csup\u003e\u003cb\u003e[19, 20]\u003c/b\u003e\u003c/sup\u003e, as well as pancreatic, renal, and cervical cancers\u003csup\u003e\u003cb\u003e[21\u0026ndash;23]\u003c/b\u003e\u003c/sup\u003e. In our study, the concurrent upregulation of FPR2 and OSM in myeloid cells suggests a coordinated inflammatory program that may persist in both UC and CRC, potentially sustaining a pro-tumorigenic microenvironment. This finding highlights a cell type-specific mechanism, where chronic activation of myeloid cells could contribute to tissue damage, immune dysregulation, and progression toward malignant transformation. Notably, these transcriptional changes were not apparent at the bulk tissue level, emphasizing the importance of single-cell resolution for detecting key drivers of disease.\u003c/p\u003e \u003cp\u003eThe significance of our findings lies in identifying shared molecular programs in inflammation-driven carcinogenesis that are conserved between UC and sporadic CRC. By pinpointing FPR2 and OSM as highly expressed in myeloid compartments, we highlight candidate biomarkers for early detection and potential targets for modulating chronic inflammation. Integrating these insights with predictive modeling for UC-associated CRC further illustrates the value of combining single-cell transcriptomics with bulk RNA-seq data to uncover cell type-specific alterations that may drive disease progression, which would be overlooked in conventional bulk analyses.\u003c/p\u003e \u003cp\u003eIn addition, we observed that MT1G and MT1H are significantly downregulated in colonic epithelial cells of both UC and sporadic CRC patients. Further analyses revealed a negative correlation between their expression and disease progression, suggesting that reduced levels may contribute to persistent inflammation, epithelial barrier dysfunction, and increased risk of malignant transformation. MT1G and MT1H, members of the metallothionein (MT) family, are small cysteine-rich proteins involved in metal homeostasis, protection against oxidative stress, and regulation of cellular proliferation and apoptosis\u003csup\u003e\u003cb\u003e[24\u0026ndash;26]\u003c/b\u003e\u003c/sup\u003e. Their downregulation in epithelial cells implies a potential loss of protective function, which may facilitate disease exacerbation and progression toward CRC.\u003c/p\u003e \u003cp\u003eThese findings indicate that MT1G and MT1H may exert protective or suppressive roles in intestinal epithelial cells, and their decreased expression could promote disease worsening and malignant transformation. Although the precise mechanisms remain to be fully elucidated, these results provide novel insights into the cell type-specific regulation underlying the transition from chronic inflammation to tumorigenesis, and they may serve as potential targets for therapeutic intervention or as prognostic biomarkers.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eBulk RNA-seq Data Analysis\u003c/h2\u003e \u003cp\u003eAll bulk RNA-seq datasets were downloaded from GEO (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.ncbi.nlm.nih.gov/geo\u003c/span\u003e\u003cspan address=\"http://www.ncbi.nlm.nih.gov/geo\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The datasets used to screen key gene changes in UC included GSE66407 (41 UC, 16 controls), GSE158952 (26 UC), GSE193677 (309 UC, 220 controls), GSE235236 (24 UC, 7 controls), and GSE245890 (27 UC). Rectum tissue samples from UC patients (n\u0026thinsp;=\u0026thinsp;427) and healthy controls (n\u0026thinsp;=\u0026thinsp;243) were used. To validate the role of the key gene sets in UC-to-CRC disease progression, the dataset GSE3629 (10 adjacent normal, 6 tumor) was used, including in situ tumor and adjacent tissue of UC-associated CRC.\u003c/p\u003e \u003cp\u003eFor the five bulk RNA-seq datasets, probe sequences were mapped to gene symbols based on their respective sequencing platforms. For multiple probes corresponding to the same gene symbol, the probe with the highest measured intensity (maximum mean) was retained. The datasets were then merged and batch effects were removed. The effectiveness of batch correction was confirmed using boxplots and PCA. Data download was performed using the R package GEOquery, and batch effect removal was conducted using the R package sva. Boxplots were drawn using the boxplot function, and PCA plots were generated with draw_pca. For GSE3629, after probe-to-symbol mapping, the raw data were log-transformed and quantile normalized. Differential expression analysis for bulk RNA-seq was performed using the R package limma to compare UC versus control colon samples, as well as UC-associated CRC tumor versus adjacent normal tissue. Volcano plots and heatmaps of differential expression results were generated using ggplot2 and pheatmap, respectively.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eWeighted Gene Co-expression Network Analysis (WGCNA)\u003c/h3\u003e\n\u003cp\u003eA gene co-expression network was constructed based on the integrated UC-control dataset to identify gene modules highly correlated with UC. First, the gene expression matrix was filtered to retain the top 25% most variable genes. Outlier samples were identified and removed. The recommended soft-thresholding power (9) was selected to achieve scale-free topology. Modules were detected using the dynamic tree cut method, with a minimum module size of 30 and a merging threshold of 0.25. Visualization of the analysis results was performed using the WGCNA R package.\u003c/p\u003e\n\u003ch3\u003eFunctional Enrichment Analysis\u003c/h3\u003e\n\u003cp\u003eGene symbols were converted to ENTREZ IDs using the bitr function from the clusterProfiler package (v4.8.2) with the org.Hs.eg.db database (v3.17.0). KEGG pathway enrichment analysis was performed using the enrichKEGG function with the following parameters: organism = \"hsa\" (Homo sapiens), p-value cutoff\u0026thinsp;=\u0026thinsp;0.05, p-value adjustment method\u0026thinsp;=\u0026thinsp;Benjamini-Hochberg, and q-value cutoff\u0026thinsp;=\u0026thinsp;0.2. Significantly enriched pathways were visualized as bubble plots showing enrichment factor, gene count, and adjusted p-value. To illustrate the relationship between genes and pathways, chord diagrams were generated using the circlize package (v0.4.15). All functional enrichment analyses were conducted in R (v4.3.1) following the standardized protocols of the clusterProfiler suite.\u003c/p\u003e\n\u003ch3\u003eImmune Microenvironment Analysis\u003c/h3\u003e\n\u003cp\u003eThe relative proportions of 22 immune cell subtypes in UC versus control colon samples were estimated using the CIBERSORT algorithm with the LM22 signature matrix. CIBERSORT was run with default parameters and 1,000 permutations to ensure robust deconvolution results. Relative immune cell abundances were visualized using stacked bar plots, and correlations among immune cell types were calculated and plotted using the Hmisc package. Group comparisons of immune cell proportions were performed using the Wilcoxon rank-sum test with p-values adjusted for multiple comparisons. Associations between immune cell proportions and gene expression in the yellow and blue WGCNA modules were assessed using Spearman rank correlation analysis.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eSingle-Cell RNA Sequencing Data Analysis\u003c/h2\u003e \u003cp\u003eSingle-cell RNA sequencing (scRNA-seq) datasets used in this study were obtained from the GEO database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/geo/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/geo/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The datasets include: UC scRNA-seq datasets \u0026mdash; GSE214695 (3 UC rectum, 6 Normal), GSE231993 (4 UC, 4 Normal), GSE125527 (7 UC, 8 Normal); CRC scRNA-seq datasets \u0026mdash; GSE161277 (4 CRC Tumor, 3 Normal), GSE261388 (3 CRC Tumor, 3 Normal), GSE231559 (6 CRC Tumor, 3 Normal).\u003c/p\u003e \u003cp\u003eData analysis was performed in R (v4.4.1), primarily using the Seurat package (v5.3.0) for quality control, normalization, dimensionality reduction, clustering, and visualization. Each sample underwent quality control with thresholds set according to sequencing characteristics, including number of detected genes (nFeature_RNA), UMI counts (nCount_RNA), mitochondrial gene percentage (percent.mt), and ribosomal gene percentage (percent.ribo), to remove low-quality cells and potential doublets. Filtered data were normalized using the LogNormalize method. Highly variable genes were identified using the FindVariableFeatures function for downstream dimensionality reduction. To integrate multiple samples and correct for batch effects, the Harmony algorithm was applied. Principal component analysis (PCA) was used for linear dimensionality reduction, and UMAP was applied for visualization of the lower-dimensional embedding. Cell type annotation was manually performed based on canonical marker genes, supported by literature and database references. Expression of marker genes and gene sets of interest was visualized using FeaturePlot, VlnPlot (violin plots), and DotPlot functions. Module scores for specific gene sets were calculated using the AddModuleScore function and visualized in violin plots to illustrate gene set expression across cell types. Boxplots were generated using ggplot2, displaying each cluster separately. Boxes represent the median and interquartile range, and group comparisons were conducted using the Wilcoxon rank-sum test.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eMachine Learning Analysis\u003c/h3\u003e\n\u003cp\u003eData Preprocessing: The GSE3629 dataset, containing expression profiles of UC-associated CRC tumor tissues (n\u0026thinsp;=\u0026thinsp;6) and adjacent normal tissues (n\u0026thinsp;=\u0026thinsp;10), was downloaded from the GEO database. Raw data were imported and processed in R. Preprocessing steps included: removing rows with all zero values; log2(x\u0026thinsp;+\u0026thinsp;1) transformation; for genes mapped to multiple probes, retaining the probe with the highest median expression; quantile normalization using the normalizeBetweenArrays() function from the limma package. The normalized data were used for downstream analysis.\u003c/p\u003e \u003cp\u003eMachine Learning: LASSO (Least Absolute Shrinkage and Selection Operator) regression analysis was performed using the glmnet package. The model was specified as binomial (family = \"binomial\") with a penalty coefficient α\u0026thinsp;=\u0026thinsp;1. Ten-fold cross-validation (nfolds\u0026thinsp;=\u0026thinsp;10) was applied to determine the optimal regularization parameter λ. Features corresponding to the minimum mean squared error (lambda.min) and the 1-standard-error simplified model (lambda.1se) were extracted. Final key genes were selected based on the lambda.1se results, and their coefficient directions and magnitudes were visualized using ggplot2. A Random Forest model was further built using the randomForest package. Input variables were the expression levels of the key genes, and the grouping variable was sample type (UC_CRC_Tumor vs. UC_CRC_Rectal), with ntree\u0026thinsp;=\u0026thinsp;500. Gene importance was evaluated using the importance() function, providing MeanDecreaseAccuracy and MeanDecreaseGini values, and a feature ranking plot was generated. Model performance was assessed using the confusion matrix and the error rate curve (err.rate). Receiver Operating Characteristic (ROC) curves for each key gene were plotted using the pROC package, and the Area Under the Curve (AUC) with 95% confidence intervals was calculated. Gene expression values were standardized using z-score normalization prior to analysis. Finally, a logistic regression model was constructed using the six key genes identified jointly by LASSO and Random Forest (IL1B, FPR2, OSM, GBP1, MT1G, MT1H), and a nomogram was built using the rms package to predict the risk of UC progression to CRC.\u003c/p\u003e\n\u003ch3\u003eSurvival Analysis\u003c/h3\u003e\n\u003cp\u003eSurvival analysis was performed using the GEPIA web server (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://gepia2.cancer-pku.cn\u003c/span\u003e\u003cspan address=\"http://gepia2.cancer-pku.cn\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Disease-Free Survival (RFS) was selected as the outcome. Patients were stratified into high- and low-expression groups using the 70th percentile as the high cutoff and the 30th percentile as the low cutoff. Kaplan-Meier survival curves were generated to evaluate the association between gene expression and CRC patient prognosis. Final figures were assembled and visually enhanced using Adobe Illustrator.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eAll statistical analyses were performed in R. Differential expression, enrichment, module-trait correlations, immune cell comparisons, and single-cell module score analyses were conducted using the methods described in the corresponding R scripts. Specific statistical tests, including Pearson or Spearman correlation, Wilcoxon rank-sum test, and LASSO/Random Forest modeling, are described in the figure legends. P-values\u0026thinsp;\u0026lt;\u0026thinsp;0.05 were considered statistically significant unless otherwise stated.\u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe acknowledge the availability of public datasets that made this study possible.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll datasets analyzed in this study are publicly available. The scripts used for data analysis are available from the corresponding author upon reasonable request. All bulk RNA-seq datasets were downloaded from the Gene Expression Omnibus (GEO)-a public functional genomics data repository supporting MIAME-compliant data submissions-under the accession numbers: GSE66407, GSE158952, GSE193677, GSE235236, GSE245890, and GSE3629. Single-cell RNA sequencing (scRNA-seq) datasets used in this study were obtained from the Gene Expression Omnibus (GEO) database, with the following accession numbers: GSE214695, GSE231993, GSE125527, GSE161277, GSE261388, and GSE231559. For Machine Learning Analysis: The dataset (accession number: GSE3629) was downloaded from the Gene Expression Omnibus (GEO) database.\u003c/p\u003e\n\u003cp\u003eAll aforementioned datasets are publicly accessible via the GEO repository\u0026rsquo;s permanent links: http://www.ncbi.nlm.nih.gov/geo and https://www.ncbi.nlm.nih.gov/geo/. All accession numbers and associated files have been fully released and are available for verification.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCM and XZ contributed to the study conception and design. LY performed data acquisition and preprocessing. JX, YG, and JY conducted the bioinformatic analyses and interpreted the results. PC, CL, and YD assisted with data analysis and visualization. CM drafted the manuscript. ZX supervised the project, revised the manuscript critically for important intellectual content, and confirmed the authenticity of all data. All authors read and approved the final manuscript. All authors read and approved the final version of the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding Declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research received no external funding.\u003c/p\u003e\n\u003cp\u003eEthics, Consent to Participate, and Consent to Publish declarations: not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eClinical Trial Numbe\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eClinical trial number: not applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eNg, S.C., et al., \u003cem\u003eWorldwide incidence and prevalence of inflammatory bowel disease in the 21st century: a systematic review of population-based studies.\u003c/em\u003e Lancet, 2017. \u003cstrong\u003e390\u003c/strong\u003e(10114): p. 2769-2778.\u003c/li\u003e\n\u003cli\u003eWinther, K.V., et al., \u003cem\u003eLong-term risk of cancer in ulcerative colitis: a population-based cohort study from Copenhagen County.\u003c/em\u003e Clin Gastroenterol Hepatol, 2004. \u003cstrong\u003e2\u003c/strong\u003e(12): p. 1088-95.\u003c/li\u003e\n\u003cli\u003eRutter, M.D., et al., \u003cem\u003eThirty-year analysis of a colonoscopic surveillance program for neoplasia in ulcerative colitis.\u003c/em\u003e Gastroenterology, 2006. \u003cstrong\u003e130\u003c/strong\u003e(4): p. 1030-8.\u003c/li\u003e\n\u003cli\u003eWangchuk, P., K. Yeshi, and A. Loukas, \u003cem\u003eUlcerative colitis: clinical biomarkers, therapeutic targets, and emerging treatments.\u003c/em\u003e Trends Pharmacol Sci, 2024. \u003cstrong\u003e45\u003c/strong\u003e(10): p. 892-903.\u003c/li\u003e\n\u003cli\u003eLi, Q., et al., \u003cem\u003eSignaling pathways involved in colorectal cancer: pathogenesis and targeted therapy.\u003c/em\u003e Signal Transduct Target Ther, 2024. \u003cstrong\u003e9\u003c/strong\u003e(1): p. 266.\u003c/li\u003e\n\u003cli\u003eSingh, M., et al., \u003cem\u003eAdvancements in combining targeted therapy and immunotherapy for colorectal cancer.\u003c/em\u003e Trends Cancer, 2024. \u003cstrong\u003e10\u003c/strong\u003e(7): p. 598-609.\u003c/li\u003e\n\u003cli\u003eRogler, G., \u003cem\u003eChronic ulcerative colitis and colorectal cancer.\u003c/em\u003e Cancer Lett, 2014. \u003cstrong\u003e345\u003c/strong\u003e(2): p. 235-41.\u003c/li\u003e\n\u003cli\u003eLi, Y., et al., \u003cem\u003eDisease-related expression of the IL6/STAT3/SOCS3 signalling pathway in ulcerative colitis and ulcerative colitis-related carcinogenesis.\u003c/em\u003e Gut, 2010. \u003cstrong\u003e59\u003c/strong\u003e(2): p. 227-35.\u003c/li\u003e\n\u003cli\u003eJudes, G., et al., \u003cem\u003eHigh-throughput \u0026lt;\u0026lt;Omics\u0026gt;\u0026gt; technologies: New tools for the study of triple-negative breast cancer.\u003c/em\u003e Cancer Lett, 2016. \u003cstrong\u003e382\u003c/strong\u003e(1): p. 77-85.\u003c/li\u003e\n\u003cli\u003eD\u0026apos;Agostino, N., W. Li, and D. Wang, \u003cem\u003eHigh-throughput transcriptomics.\u003c/em\u003e Sci Rep, 2022. \u003cstrong\u003e12\u003c/strong\u003e(1): p. 20313.\u003c/li\u003e\n\u003cli\u003eSlovin, S., et al., \u003cem\u003eSingle-Cell RNA Sequencing Analysis: A Step-by-Step Overview.\u003c/em\u003e Methods Mol Biol, 2021. \u003cstrong\u003e2284\u003c/strong\u003e: p. 343-365.\u003c/li\u003e\n\u003cli\u003eChi, H., et al., \u003cem\u003eT-cell exhaustion signatures characterize the immune landscape and predict HCC prognosis via integrating single-cell RNA-seq and bulk RNA-sequencing.\u003c/em\u003e Front Immunol, 2023. \u003cstrong\u003e14\u003c/strong\u003e: p. 1137025.\u003c/li\u003e\n\u003cli\u003eAlessi, M.C., et al., \u003cem\u003eFPR2: A Novel Promising Target for the Treatment of Influenza.\u003c/em\u003e Front Microbiol, 2017. \u003cstrong\u003e8\u003c/strong\u003e: p. 1719.\u003c/li\u003e\n\u003cli\u003eLiu, X., et al., \u003cem\u003eSerum amyloid A contributes to radiation-induced lung injury by activating macrophages through FPR2/Rac1/NF-kappaB pathway.\u003c/em\u003e Int J Biol Sci, 2024. \u003cstrong\u003e20\u003c/strong\u003e(12): p. 4941-4956.\u003c/li\u003e\n\u003cli\u003eWu, M.Y., et al., \u003cem\u003eEnhancement of efferocytosis through biased FPR2 signaling attenuates intestinal inflammation.\u003c/em\u003e EMBO Mol Med, 2023. \u003cstrong\u003e15\u003c/strong\u003e(12): p. e17815.\u003c/li\u003e\n\u003cli\u003eZarling, J.M., et al., \u003cem\u003eOncostatin M: a growth regulator produced by differentiated histiocytic lymphoma cells.\u003c/em\u003e Proc Natl Acad Sci U S A, 1986. \u003cstrong\u003e83\u003c/strong\u003e(24): p. 9739-43.\u003c/li\u003e\n\u003cli\u003eLantieri, F. and T. Bachetti, \u003cem\u003eOSM/OSMR and Interleukin 6 Family Cytokines in Physiological and Pathological Condition.\u003c/em\u003e Int J Mol Sci, 2022. \u003cstrong\u003e23\u003c/strong\u003e(19).\u003c/li\u003e\n\u003cli\u003eFearon, U., et al., \u003cem\u003eOncostatin M induces angiogenesis and cartilage degradation in rheumatoid arthritis synovial tissue and human cartilage cocultures.\u003c/em\u003e Arthritis Rheum, 2006. \u003cstrong\u003e54\u003c/strong\u003e(10): p. 3152-62.\u003c/li\u003e\n\u003cli\u003eHoagland, D.A., et al., \u003cem\u003eMacrophage-derived oncostatin M repairs the lung epithelial barrier during inflammatory damage.\u003c/em\u003e Science, 2025. \u003cstrong\u003e389\u003c/strong\u003e(6756): p. 169-175.\u003c/li\u003e\n\u003cli\u003eDomaniku-Waraich, A., et al., \u003cem\u003eOncostatin M signaling drives cancer-associated skeletal muscle wasting.\u003c/em\u003e Cell Rep Med, 2024. \u003cstrong\u003e5\u003c/strong\u003e(4): p. 101498.\u003c/li\u003e\n\u003cli\u003eLee, B.Y., et al., \u003cem\u003eHeterocellular OSM-OSMR signalling reprograms fibroblasts to promote pancreatic cancer growth and metastasis.\u003c/em\u003e Nat Commun, 2021. \u003cstrong\u003e12\u003c/strong\u003e(1): p. 7336.\u003c/li\u003e\n\u003cli\u003eWei, S., et al., \u003cem\u003eOSM May Serve as a Biomarker of Poor Prognosis in Clear Cell Renal Cell Carcinoma and Promote Tumor Cell Invasion and Migration.\u003c/em\u003e Int J Genomics, 2023. \u003cstrong\u003e2023\u003c/strong\u003e: p. 6665452.\u003c/li\u003e\n\u003cli\u003eNoh, J., et al., \u003cem\u003eActivation of OSM-STAT3 Epigenetically Regulates Tumor-Promoting Transcriptional Programs in Cervical Cancer.\u003c/em\u003e Cancers (Basel), 2022. \u003cstrong\u003e14\u003c/strong\u003e(24).\u003c/li\u003e\n\u003cli\u003eSi, M. and J. Lang, \u003cem\u003eThe roles of metallothioneins in carcinogenesis.\u003c/em\u003e J Hematol Oncol, 2018. \u003cstrong\u003e11\u003c/strong\u003e(1): p. 107.\u003c/li\u003e\n\u003cli\u003eHiguera, M., et al., \u003cem\u003eImpact of zinc on hepatocellular carcinoma cell behavior and metallothionein expression: Insights from preclinical models.\u003c/em\u003e Biomed Pharmacother, 2025. \u003cstrong\u003e185\u003c/strong\u003e: p. 117918.\u003c/li\u003e\n\u003cli\u003eJeong, S.H., et al., \u003cem\u003eMTF1 Is Essential for the Expression of MT1B, MT1F, MT1G, and MT1H Induced by PHMG, but Not CMIT, in the Human Pulmonary Alveolar Epithelial Cells.\u003c/em\u003e Toxics, 2021. \u003cstrong\u003e9\u003c/strong\u003e(9).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Ulcerative colitis, Colorectal cancer, scRNA sequencing, Inflammation-driven tumorigenesis, Machine Learning, Tumor microenvironment","lastPublishedDoi":"10.21203/rs.3.rs-8417562/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8417562/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eUlcerative colitis (UC) is a chronic inflammatory disease that increases the risk of colorectal cancer (CRC), yet the molecular features linking these conditions remain unclear. Here, we integrated bulk RNA sequencing with single-cell transcriptomic data to investigate shared transcriptional programs between UC and CRC. UC-associated differentially expressed genes (DEGs) were mapped onto UC and sporadic CRC single-cell datasets, allowing cell type-resolved analysis of conserved alterations. We found that a subset of UC DEGs displayed consistent dysregulation in sporadic CRC, with strong cell specificity. Myeloid cells-particularly macrophages, neutrophils, and dendritic cells-showed coordinated upregulation of inflammatory genes in both diseases, whereas epithelial cells exhibited convergent downregulation of protective gene modules. These patterns suggest that persistent immune activation and epithelial dysfunction represent shared pathological axes that may contribute to inflammation-driven tumorigenesis. Using bulk data from UC-associated CRC, machine-learning analysis identified four genes with predictive value for UC-to-CRC progression. Overall, this study provides a refined cellular framework for understanding UC-CRC molecular convergence and highlights candidate biomarkers for early detection.\u003c/p\u003e","manuscriptTitle":"Integrative Bioinformatics of Bulk and Single-Cell Transcriptomes Identifies Predictive Molecular Features of Ulcerative Colitis-Associated Colorectal Cancer","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-01-14 07:50:08","doi":"10.21203/rs.3.rs-8417562/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"1ec1e62c-4687-4535-a474-57353c0133b5","owner":[],"postedDate":"January 14th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-02-18T09:11:02+00:00","versionOfRecord":[],"versionCreatedAt":"2026-01-14 07:50:08","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8417562","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8417562","identity":"rs-8417562","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.