Investigating combined hypoxia and stemness indices for prognostic transcripts in gastric cancer: Machine learning and network analysis approaches.

OA: gold
AI-generated deep summary by qwen3.7-flash, 2026-08-27 · read from full text

This study utilized machine learning and network analysis to develop a prognostic decision tree for gastric cancer by integrating hypoxia and stemness indices derived from TCGA-STAD RNA-seq data. The researchers calculated hypoxia scores using gene set variation analysis and stemness scores via the mRNA stemness index method, subsequently performing hierarchical clustering and weighted gene co-expression network analysis to identify survival-associated hub genes. Key findings indicated that combining these molecular features improved the prediction of patient outcomes compared to existing biomarkers, although the model relies entirely on retrospective computational analysis without experimental validation. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

IntroductionGastric cancer (GC) is among the deadliest malignancies globally, characterized by hypoxia-driven pathways that promote cancer progression, including stemness mechanisms facilitating invasion and metastasis. This study aimed to develop a prognostic decision tree using genes implicated in hypoxia and stemness pathways to predict outcomes in GC patients.Materials and methodsGC RNA-seq data from The Cancer Genome Atlas (TCGA) were analyzed to compute hypoxia and stemness scores using Gene Set Variation Analysis (GSVA) and the mRNA expression-based stemness index (mRNAsi). Hierarchical clustering identified clusters with distinct survival outcomes, and differentially expressed genes (DEGs) between clusters were identified. Weighted Gene Co-expression Network Analysis (WGCNA) identified modules and hub genes associated with clinical traits. Overlapping DEGs and hub genes underwent functional enrichment, protein-protein interaction (PPI) network analysis, and survival analysis. A prognostic decision tree was constructed using survival-associated shared genes.ResultsHierarchical clustering identified six clusters among 375 TCGA GC patients, with significant survival differences between cluster 1 (low hypoxia, high stemness) and cluster 4 (high hypoxia, high stemness). Validation in the GSE62254 dataset corroborated these findings. WGCNA revealed modules linked to clinical traits and survival, with functional enrichment highlighting pathways like cell adhesion and calcium signaling. The decision tree, based on genes such as AKAP6, GLRB, and RUNX1T1, achieved an AUC of 0.81 (training) and 0.67 (test), demonstrating the utility of combined scores in patient stratification.ConclusionThis study introduces a novel hypoxia-stemness-based prognostic decision tree for GC. The identified genes show promise as prognostic biomarkers, warranting further clinical validation.
Full text 45,395 characters · extracted from pmc-nxml · 9 sections · click to expand

Credit

Sharareh Mahmoudian-Hamedani: Writing – review & editing, Writing – original draft, Visualization, Software, Methodology, Investigation, Formal analysis, Data curation. Maryam Lotfi-Shahreza: Writing – review & editing, Validation, Methodology, Formal analysis. Parvaneh Nikpour: Writing – review & editing, Writing – original draft, Validation, Supervision, Resources, Project administration, Investigation, Funding acquisition, Conceptualization.

Funding

This work was partially supported by a grant from Isfahan University of Medical Sciences , Isfahan, Iran [grant number 340122 ]. The sponsor (Isfahan University of Medical Sciences) was not involved in the study design, data collection, analysis and interpretation, manuscript writing, or the decision to submit the manuscript for publication.

Results

To discover the correlation between combined hypoxia and stemness and the survival of GC patients, an unsupervised two-dimensional hierarchical clustering was performed based on the combined hypoxia and stemness scores of 375 GC patients whose tumoral samples’ transcriptomic data was retrieved from TCGA. The result of clustering is represented in Fig. 2 a which included: cluster 1 (66 samples, low hypoxia and high stemness), cluster 2 (45 samples, low hypoxia and stemness), cluster 3 (84 samples, rather high hypoxia and high stemness), cluster 4 (53 samples, high hypoxia and stemness), cluster 5 (80 samples, rather high hypoxia and rather low stemness) and cluster 6 (44 samples, low hypoxia and stemness). According to the survival analysis performed on these clusters, the highest survival difference was attributed to clusters 1 (higher survival rates) and 4 (lower survival rates), and log-rank test result trended towards significance ( P. value = 0.06, Fig. 2 b). Then, we determined whether we could identify the previously described hypoxia-stemness clusters in an external validation dataset, GSE62254 . We applied the same hierarchical-clustering algorithm to subgroup the GC patients according to the combined hypoxia and stemness scores ( Fig. 2 c) and survival rates were then calculated for each cluster. Our results showed that the difference in survival rate between the two clusters with low hypoxia and high stemness (cluster 4 with higher survival rates) and high hypoxia and stemness (cluster 6 with lower survival rates) was statistically significant ( P. value = 0.01, Fig. 2 d). These results confirmed the accuracy of the combined hypoxia and stemness scores in determining patient risk stratification and prognosis. After confirming the prognostic efficacy of the combined scores, we extracted 1446 differentially-expressed genes between clusters 1 (higher survival rates) and 4 (lower survival rates) of the TCGA-STAD samples. Fig. 2 Clustering and survival analysis of TCGA-STAD and GSE62254 datasets. TCGA-STAD and GSE62254 dataset samples were clustered based on hypoxia and stemness scores (a and c), then survival analysis were performed for the clusters and Kaplan-Meier plots were extracted (b and d). STAD: Stomach adenocarcinoma, TCGA: The Cancer Genome Atlas. Fig. 2 Clustering and survival analysis of TCGA-STAD and GSE62254 datasets. TCGA-STAD and GSE62254 dataset samples were clustered based on hypoxia and stemness scores (a and c), then survival analysis were performed for the clusters and Kaplan-Meier plots were extracted (b and d). STAD: Stomach adenocarcinoma, TCGA: The Cancer Genome Atlas. To define the prognosis-related key gene clusters and hub genes of GC, we performed WGCNA analysis using the normalized transcriptomic data of 375 GC samples retrieved from TCGA and the matched clinical data including pathologic stage, pathologic_T, pathologic_N, pathologic_M, days to death and overall survival. The soft threshold was set to 9, and the scale-free topology fitting index reached 0.89 ( Fig. 3 a and b). Thirty modules were then detected using the WGCNA calculation which were merged into 15, each with a unique color, by a cut-off of 0.4, while the minimum number of genes in each cluster was set to 30 ( Fig. 3 c). The genes contained in the dark gray module are genes that do not belong to any other modules. Pearson correlation analysis was employed to identify modules that were significantly linked to the prognosis-related clinical traits. There were found to be 15 modules, of which dark turquoise (R = 0.16, P . value = 0.04), dark red (R = 0.15, P . value = 0.003), and dark orange (R = −0.15, P . value = 0.005) were related to the pathologic stage. Blue (R = 0.15, P . value = 0.004), dark turquoise (R = 0.19, P . value = 2e-04), dark red (R = 0.17, P . value = 0.001), dark green (R = −0.14, P . value = 0.006), dark orange (R = −0.2, P . value = 1e-04) and royal blue (R = −0.16, P . value = 0.001) were related to pathologic_T, magenta (R = 0.12, P . value = 0.02) and dark orange (R = −0.16, P . value = 0.002) were related to pathologic_N and greenyellow (R = 0.16, P . value = 0.002) and purple (R = 0.11, P . value = 0.03) modules were related to overall survival ( Fig. 3 d). With module membership (MM)≥0.7 being set as a threshold to select hub genes, there were overall 1885 genes selected for further analysis. These genes were then intersected with the differentially expressed genes between clusters 1 and 4 of the TCGA-STAD samples. A total of 569 overlapping genes, referred to as the "shared genes" throughout the text, were identified and selected for further analysis ( Fig. 4 ). Fig. 3 The WGCNA results The plots represent scale-free fit index versus soft-thresholding power (a) and mean connectivity versus soft-thresholding power (b), respectively. The cluster dendrogram of gene expression data shows distinct modules with different colors representing each one (c). Module-trait relationships has been represented as a heatmap (d). ME: Module eigengene, WGCNA: Weighted gene co-expression network analysis. Fig. 3 Fig. 4 Venn diagram representing the shared genes between the hub genes and the DEGs. The Venn diagram shows the number of shared genes (#596) between the hub genes resulted from WGCNA (#1885) and differential expression analysis of cluster 1 vs. 4 of TCGA-STAD clustering results (#1446). DEGs: Differentially-expressed genes, STAD: Stomach adenocarcinoma, TCGA: The Cancer Genome Atlas, WGCNA: Weighted gene co-expression network analysis. Fig. 4 The WGCNA results The plots represent scale-free fit index versus soft-thresholding power (a) and mean connectivity versus soft-thresholding power (b), respectively. The cluster dendrogram of gene expression data shows distinct modules with different colors representing each one (c). Module-trait relationships has been represented as a heatmap (d). ME: Module eigengene, WGCNA: Weighted gene co-expression network analysis. Venn diagram representing the shared genes between the hub genes and the DEGs. The Venn diagram shows the number of shared genes (#596) between the hub genes resulted from WGCNA (#1885) and differential expression analysis of cluster 1 vs. 4 of TCGA-STAD clustering results (#1446). DEGs: Differentially-expressed genes, STAD: Stomach adenocarcinoma, TCGA: The Cancer Genome Atlas, WGCNA: Weighted gene co-expression network analysis. In order to gain further insight into the functionality of the “ shared genes ”, KEGG and GO analyses were performed. As the results show in Fig. 5 a, the most significant GO biological process (BP) terms associated with the shared genes included cell adhesion, negative regulation of cell proliferation and nervous system development, while GO cellular component (CC) analysis revealed that the related proteins to these genes were mostly localized in plasma membrane and extracellular spaces and GO molecular function (MF) analysis showed that the related proteins functions to be calcium ion binding, extracellular matrix structural constituent, heparin, integrin and collagen binding. According to KEGG functional enrichment analysis, the shared genes participated in focal adhesion, calcium and cAMP signaling pathways ( Fig. 5 b). Fig. 5 Dot plots of GO and KEGG functional enrichment analysis based on the shared genes. The horizontal and vertical axes represent gene ratio and terms, respectively. For each term, FDR<0.05 was considered significant. Top ten terms in each plot are presented. GO terms (a), KEGG terms (b). BP: biological process, CC: Cellular component, DEGs: differentially expressed genes, GO: Gene Ontology, FDR: false discovery rate, KEGG: Kyoto Encyclopedia of Genes and Genomes, MF: Molecular function, STAD: Stomach adenocarcinoma, TCGA: The Cancer Genome Atlas. Fig. 5 Dot plots of GO and KEGG functional enrichment analysis based on the shared genes. The horizontal and vertical axes represent gene ratio and terms, respectively. For each term, FDR<0.05 was considered significant. Top ten terms in each plot are presented. GO terms (a), KEGG terms (b). BP: biological process, CC: Cellular component, DEGs: differentially expressed genes, GO: Gene Ontology, FDR: false discovery rate, KEGG: Kyoto Encyclopedia of Genes and Genomes, MF: Molecular function, STAD: Stomach adenocarcinoma, TCGA: The Cancer Genome Atlas. Next, to investigate the interactions between the proteins associated with the “ shared genes ”, a PPI network was constructed which included 538 nodes and 587 edges ( Fig. 6 a). The highest node degree centrality was attributed to FN1 (Fibronectin 1), followed by COL1A1 (Collagen Type I Alpha 1 Chain), COL3A1 (collagen type III alpha 1 chain), MMP2 (matrix metallopeptidase 2), COL1A2 (collagen type I alpha 2 chain), DCN (decorin), COL5A1 (collagen type V alpha 1 chain), THBS1 (thrombospondin 1), LUM (lumican), COL6A2 (collagen type VI alpha 2 chain) ( Fig. 6 b). Fig. 6 PPI analysis of the shared genes The PPI network (a) and the bar plot of the top 10 nodes with the highest degree-centralities (b) are represented. PPI: Protein-protein interaction. Fig. 6 PPI analysis of the shared genes The PPI network (a) and the bar plot of the top 10 nodes with the highest degree-centralities (b) are represented. PPI: Protein-protein interaction. The association of the 569 “ shared genes ” with TCGA-STAD overall survival was investigated. Sixty-two genes were found to be significantly ( P . value < 0.05) associated with overall survival, which are referred to as the "survival-associated shared genes" throughout the text ( Suppl Table 1 ). Next, these 62 genes were submitted for Recursive Partitioning to construct a prognostic decision tree model ( Fig. 7 a). The 375 tumoral samples of the TCGA-STAD dataset were partitioned into training (0.75) and test (0.25) sets. The decision tree was constructed using eight genes, including AKAP6 , GLRB , LINC00578, LINC00968, MIR145 , NBEA , NEGR1 and RUNX1T1 . Each node in the decision tree signifies a decision point based on normalized gene expression levels, with each branch indicating a potential outcome. To predict an outcome for a new patient, the tree navigates through the branches according to the patient's gene expression levels until it reaches a leaf node, which represents the predicted outcome. In this decision tree, GLRB represents the root node and NBEA , LINC00968 and MIR145 together represent the leaf nodes. The efficiency of the model was evaluated using the ROC curves, according to which the AUC of the training and the test sets were 0.81 and 0.67, respectively ( Fig. 7 b and c). Other performance indices are presented in Table 1 . Fig. 7 The prognostic decision tree model for GC constructed using the shared genes. The prognostic decision tree (a) and ROC curves of the training (b) and test (c) sets are represented. GC: Gastric cancer, ROC: Receiver operating characteristic. Fig. 7 Table 1 Decision tree model performance metrices Table 1 AUC a Precision Recall F1 score Youden's index Maximum classification accuracy 0.68 0.73 0.76 0.75 0.26 0.32 a Area Under the Curve. The prognostic decision tree model for GC constructed using the shared genes. The prognostic decision tree (a) and ROC curves of the training (b) and test (c) sets are represented. GC: Gastric cancer, ROC: Receiver operating characteristic. Decision tree model performance metrices Area Under the Curve.

Materials

RNA-seq data of 375 STAD samples (tumoral) and related clinical data were retrieved from TCGA using the TCGAbiolinks R package [ 20 ]. The RNA-seq data in the form of STAR (Spliced Transcripts Alignment to a Reference) counts were normalized using the trimmed mean of M values (TMM) method and log 2 -transformed using the edgeR package [ 21 ]. As a validation set, the microarray gene expression data of 300GC patients ( GSE62254 [ 22 ]) were downloaded from Gene Expression Omnibus (GEO) database via the GEOquery package [ 23 ] and normalized using the limma [ 24 ] package. The related clinical data of GSE62254 were retrieved as well [ 22 ]. HALLMARK_HYPOXIA.v7.5.1, a gene set containing 200 hypoxia-related genes was retrieved from the Molecular Signatures Database (MSigDB, https://www.gsea-msigdb.org/gsea/msigdb ), which provides a vast collection of gene sets with annotations that can be utilized to detect patterns of different pathways using samples’ gene expression data. Hypoxia scores of TCGA-STAD and GSE62254 samples were subsequently estimated employing the GSVA R package by using the HALLMARK_HYPOXIA.v7.5.1 as the gene signature and ssGSEA as the method. In 2018, Malta et al. proposed a method to calculate the mRNA expression-based stemness index (mRNAsi) of cancer cells, using the one-class logistic regression (OCLR) machine-learning algorithm [ 25 , 26 ]. In the present study, mRNAsi was calculated by correlation analysis between stemness signatures' weight vector and mRNA expression in GC samples obtained from TCGA-STAD and GEO database ( GSE62254 ) was calculated using Spearman's method. To identify and confirm the association between hypoxia and stemness scores and prognosis of GC patients, first an unsupervised two-dimensional hierarchical clustering based on these two scores was carried out on the TCGA-STAD and GES62254 samples. The results of the hypoxia score and mRNAsi of GC samples were first standardized using the “scale” function in the R. To perform hierarchical clustering, the Manhattan method for distance measurement and Ward's minimum-variance method for linkage analysis were employed and the number of clusters was set to six (k = 6). Distance measurement and linkage methods are used to calculate the closeness of the samples and the distance of the clusters, respectively [ 27 ]. The same methods were applied to the results of hypoxia scores and mRNAsi of GC samples of the GSE62254 dataset (As the validation set). The clustering was performed and visualized using the hclust and ComplexHeatmap [ 28 ] packages, respectively. Survival analysis was performed on the clustering results of TCGA-STAD and GSE62254 validation set to unravel the association of hypoxia and stemness-based hierarchical clustering with prognosis and possible pattern of hypoxia and stemness in groups with high and low survival rates. To this aim, the survival package in R was employed. The overall survival (OS) of distinct clusters was evaluated using the Kaplan-Meier curve and log-rank test. Next, DEGs of the two TCGA-STAD derived clusters with the highest difference in survival rate were identified using false discovery rate (FDR) 1.5 by employing the edgeR package in R. WGCNA is a systems biology analysis method for categorizing the genes in the microarray/RNA-seq samples into distinct modules based on their correlation. It can also be utilized for associating modules with samples' traits in order to find the most pertinent modules for targeted therapy or introducing potential biomarkers [ 29 ]. To construct a weighted gene co-expression network, first TCGA-STAD samples were clustered using the average method. Next, an adjacency matrix was formed as a result of converting the similarity matrix using the β value as the soft threshold. To create a scale-free network, a topological overlap matrix (TOM) was then created and by applying DynamicTreeCut method, TOM was converted into a weighted gene co-expression network and each module was identified by a specific color. To correlate each module with different clinical traits, module eigengenes were calculated. Module eigengenes are described as the first principal components that summarize the overall gene expression level in individual modules. After selecting those modules that were correlated with clinical traits related to GC patients’ prognosis, the module membership (MM) ≥ 0.7 was used to identify hub genes which were associated with prognosis. The higher value of MM represents the more prognostic value it holds for the patient. To identify the shared genes between the previously mentioned DEGs and hub genes attained from WGCNA, InteractiVenn [ 30 ], an online tool for creating Venn diagrams was employed. The Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Ontology (GO) terms associated with the shared genes were then obtained from the Database for Annotation, Visualization and Integrated Discovery (DAVID) [ 31 ] and visualized using the ggplot2 package in R. To construct a PPI network, list of shared genes was applied to STRING database and the network was constructed with minimum interaction score of 0.7 (high confidence) [ 32 ], and the resulted network was exported to Cytoscape software [ 33 ]. To explore the association of the “shared genes” with overall survival, the survival package was employed in R. Statistical significance of the association was calculated using the log-rank test and P . value < 0.05 was considered as the threshold. A prognostic decision tree model was constructed using the expression of “survival-associated shared genes” and patients' prognosis (patients with alive status were labeled as good prognosis, and those with dead status were labeled as bad prognosis) as inputs to CART (Classification and Regression Tree) algorithm which is implemented in the rpart R package [ 34 ]. The CART algorithm creates a decision tree model by first splitting the dataset into training and testing subsets. It then recursively selects the optimal features to split the data until a stopping criterion is met, such as maximum depth or minimum samples per node. The algorithm takes the dataset as input and constructs a tree structure starting from the root node down to the leaf nodes [ 35 ]. Data was split into a training (75 %) and a testing (25 %) set and the model was fitted on the training set. The rpart.plot package in R [ 36 ] was used to visualize the decision tree. The performance of the model was then evaluated on the testing set by plotting the receiver operating characteristic (ROC) curve and calculating the area under the curve (AUC) using the pROC package in R. To further investigate the performance of the model, other indices including precision, recall, F1 score, Youden's index and maximum classification accuracy were calculated.

Conclusion

In conclusion, this study integrates bioinformatics approaches to elucidate the interplay between hypoxia and stemness in gastric cancer and its implications for patient prognosis. By analyzing TCGA data and an external validation dataset, the study shows that combining hypoxia and stemness scores can categorize GC patients into different prognostic groups. Specifically, clusters characterized by high hypoxia and stemness exhibit poorer survival outcomes compared to those with low hypoxia and high stemness. Differentially expressed genes identified between these clusters, along with hub genes derived from WGCNA, highlight pathways related to cell adhesion, proliferation, and nervous system development, emphasizing their role in GC progression. Furthermore, the construction of a prognostic decision tree model based on survival-associated shared genes provides a framework for predicting patient outcomes. This approach highlights the potential of hypoxia and stemness as biomarkers for GC prognosis and suggests further validation and therapeutic strategies for managing this cancer.

Discussion

GC is one of the leading cancer-related mortalities worldwide. One of the characteristics of solid tumors is the hypoxic microenvironment, which by itself can lead to high stemness in the tumor cells [ 37 ]. High stemness can result in de-differentiation of the tumor cells and increase the chances of invasion and metastasis [ 38 ]. The resulting aggressive phenotype is associated with poor prognosis, and that's the reason behind many studies' interest in finding novel hypoxia- and stemness-related prognostic biomarkers [ [39] , [40] , [41] , [42] ]. There are studies that have examined the combined effects of hypoxia and stemness on tumor progression [ 43 , 44 ] and in this regard, there is a lot to be discovered in the field of gastric cancer. In the present study, the RNA-seq data of gastric cancer were retrieved from TCGA database and hypoxia and stemness scores were calculated. Hierarchical clustering and survival analysis were then carried out resulting in two clusters (one with low hypoxia and high stemness (cluster 1) and another one with both high hypoxia and high stemness (cluster 4)) showing the most difference in survival rates. Of note, performing the same pipeline on an external validation set ( GSE62254 ) resulted in the same pattern. Hypoxia within tumors has been linked to the decreased disease-free survival (DFS) outcomes across a variety of cancer types [ [45] , [46] , [47] , [48] , [49] ]. In the same way, in the context of gastric cancer, several studies have reported the association between hypoxia and hypoxia-related gene expressions and poorer overall survival [ [50] , [51] , [52] ]. The hypoxic environment has the potential to alter the gene expression levels that regulate various metabolic and other physiological processes. Furthermore, the hypoxic signaling pathways interact with other cellular pathways to modify the malignant behavior of cancer cells including the proliferation, migration, invasion, and angiogenesis and significantly influences the treatment outcomes of cancer [ 53 , 54 ]. The current study utilized TCGA data to evaluate hypoxia scores for gastric cancer and assess their correlation with survival rates. Our findings align with prior studies, demonstrating a direct correlation between elevated hypoxia levels and diminished survival rates. The concept of “stemness” encompasses the coordinated molecular mechanisms that regulate and sustain the characteristics of stem cells [ 55 ]. Since its introduction by Malta et al. [ 25 ], the mRNAsi score, derived from mRNA expression profiles, has become a widely utilized metric in numerous studies to evaluate the stemness properties of tumor specimens. Its application reveals divergent relationships with survival outcomes across distinct cancer types. While some malignancies like breast and hepatocellular carcinoma demonstrate an inverse correlation between mRNAsi levels and overall survival rates [ 56 , 57 ], indicating that higher mRNAsi scores are associated with poorer prognoses, others like lung adenocarcinoma, gastric and bladder cancers exhibit a positive correlation, suggesting that elevated mRNAsi scores correspond to improved overall survival rates [ 43 , 58 , 59 ]. Of note, in the study on bladder cancer [ 59 ], when they got the results which were different from general understanding of the characteristics of CSCs [ 25 ], they assumed it maybe because of not considering the tumor purity factor. With considering this and re-calculating mRNAsi (referred as corrected mRNAsi), they observed the expected association pattern between mRNAsi and survival rate (higher mRNAsi scores showing worse overall survival rates). Moreover, in certain instances, no statistically significant associations between mRNAsi scores and survival outcomes are discerned [ 60 , 61 ]. These findings suggest that the relationship between mRNAsi scores and survival outcomes may be tissue-dependent, varying significantly across different cancer types. Of note, in the current study, we did not investigate the association between the mRNAsi index and survival outcomes in TCGA-STAD samples in isolation. Instead, we combined hypoxia and mRNAsi scores to assess their overall association with overall survival. Based on the findings of the present study, samples with high hypoxia and high stemness showed the worse survival rates compared to ones with low hypoxia and high stemness. Our analysis of the TCGA-STAD as well as an external GC dataset indicated that hypoxia is the most critical factor in determining the survival status of GC patients, whereas stemness appears to have a minimal impact on survival. That can be explained regarding the heterogenous nature of gastric cancer, causing the prognosis be linked to many factors such as tumor location [ 62 ] and type [ 63 ]. Furthermore, A study by Yoon et al. indicated that in contrast to many other studies, increased levels of gastric stem cell markers, such as LGR5 and VIL1, are observed in chromosomal instability (CIN)/intestinal subtypes and correlate with improved survival [ 64 ]. This finding highlights the complexity of molecular pathways that contribute to tumor progression. Future studies can focus on single-cell and spatial sequencing analysis to better understand the complex pathways underlying tumorigenesis and progression in distinct types of GC. Differential expression analysis between clusters 1 (higher survival rates) and 4 (lower survival rates) of the TCGA-STAD samples resulted in 1446 genes. Performing WGCNA on the normalized RNA-seq data of TCGA-STAD resulted in 1885 hub genes showing associations with selected clinical traits. GO-BP analysis on the shared genes between these two gene sets showed enrichment in cell adhesion and proliferation as well as nervous system development . Hypoxia, through HIF factors, has significant effects on cell adhesion-related molecules expression in the tumor microenvironment. For example, Zhang et al. showed that hypoxia can lead to up-regulation of epithelial cell adhesion molecule (EpCAM), which increases the expression of stemness markers such as NANOG, SOX2, and OCT4 and epithelial mesenchymal transition (EMT) markers such as N‐cadherin and vimentin, and thus contributes to increasing the stemness properties of breast cancer cells [ 65 ]. Another study that supports the idea of hypoxia affecting expression of cell adhesion molecules has been reported by Chen et al., where they showed that VEGF-A from T cells under hypoxic conditions mediates the regulation of VCAM-1, affecting immune cell adhesion and homing through the endothelial barrier. This suggests that hypoxia-induced VEGF-A expression in tumors could influence immune escape mechanisms by modulating cell adhesion [ 54 ]. Furthermore, the role of hypoxia in inducing cell adhesion molecules have been studies in other conditions, such as endometriosis [ 66 ], showing that it induces the expression of molecules like integrin α5, αV, β3, and β5, and ANTXR2, via TGF-β1/Smad and HIF-1α signaling, respectively, thereby enhancing cell adhesion and promoting the development of endometriosis. In addition to promoting stemness and EMT, hypoxia can increase gastric cancer cells proliferation [ 67 ]. In the hypoxic and low pH environment of tumors, glucose metabolism shifts to high glycolysis instead of oxidative phosphorylation, known as the Warburg effect, which is a hallmark of most tumor cells and is linked to rapid proliferation, drug resistance, and stemness [ 68 ]. Furthermore, hypoxia can contribute to gastric cancer cell proliferation through miRNA-regulation: hypoxia-inducible miR-224 acts as an oncogene by promoting gastric cancer cell growth, migration, and invasion through the downregulation of RASSF8 [ 69 ]. In addition, Chronic hypoxia in tumors, caused by an imbalance between cell proliferation and angiogenesis, leads to abnormal microvasculature and acute hypoxia, with re-oxygenation generating reactive oxygen species (ROS) that exacerbate the effects of hypoxia on tumors [ 70 ]. ROS in tumors transmit self-proliferative signals, reduce sensitivity to anti-proliferative signals, and promote invasion and metastasis [ 71 ]. Previous studies have shown that the nervous system plays a crucial role in promoting tumor metastasis by regulating metastatic cascades through the secretion of neural-related factors from nerve endings, such as neurotrophies, neurotransmitters, and neuropeptides [ 72 ]. Colorectal and gastric cancer stem cells can differentiate into neurons that support tumor growth, and targeting the neural differentiation capability of CSCs may enhance cancer treatment by preventing tumor neurogenesis and resistance to therapies [ 73 ]. Furthermore, the enteric nervous system in particular may probably be capable of providing a dissemination rout for gastric cancer cells metastasis, especially in diffuse-type gastric cancer [ 74 ]. In addition, dysregulation of serotonin receptors, dopamine receptors (DRs), and catechol- O -methyltransferase (COMT) in GC suggests an intermediate role of the brain-gastrointestinal axis in GC development [ 75 , 76 ], further confirming the influence of nervous system in GC metastasis. KEGG functional analysis also revealed that the shared genes were mostly involved in focal adhesion, calcium and cAMP signaling pathways, all of which are involved in cancer progression and invasion [ [77] , [78] , [79] ]. After performing survival analysis on the shared genes, a decision tree was constructed. Our prognostic decision tree achieved an AUC of 0.68, reflecting moderate discriminative power. The F1 score of 0.75, alongside precision and recall values of 0.73 and 0.76, respectively, indicates the model's ability to maintain a balance between identifying true positives and avoiding false positives. Other studies have also explored hypoxia and stemness indices to create prediction models for GC patients. For example, Chen et al. [ 80 ] in 2020, constructed a LASSO Cox regression-based prognostic risk score using mRNAsi index with an AUC of 0.68. Moreover, Wang et al. [ 81 ] in 2022 constructed a hypoxia-angiogenesis based prognostic model using random forest, demonstrating AUC values of 0.67, 0.67, and 0.78 for 1-, 3-, and 5-year predictions. Furthermore, in 2024, Li et al. [ 82 ] developed a prognostic nomogram model based on hypoxia- and mitochondrial dysfunction-related genes (HMDRGs) along with age and stage of the disease, with an AUC of 0.70. Our decision tree was constructed with eight genes including AKAP6 , GLRB , LINC00578, LINC00968, MIR145 , NBEA , NEGR1 and RUNX1T1 . AKAP6 is a member of the family of A-kinase anchoring proteins (AKAPs). AKAPs act as scaffolds that localize protein kinase A (PKA) to its target proteins, thereby modulating and enhancing the biological outcomes of cAMP signaling. cAMP–PKA signaling is a key pathway that influences cancer cell proliferation, motility, invasion and metabolism [ 79 , 83 ]. Several studies have revealed the link between AKAP6 polymorphisms and cancer. For instance, it has been demonstrated that AKAP6 variants can influence the risk and outcome of glioma [ 84 ]. Another study reported that AKAP6 mutations are increased in gastric cancer [ 85 ]. Further research is needed to elucidate the mechanisms by which AKAP6 dysregulation impacts cancer progression. GLRB encodes the beta subunit of the glycine receptor chloride channel (GlyR β), which has a significant role in inhibitory neurotransmission in central nervous system and also modulates neurotransmitter release presynaptically [ 86 ]. Studies have elucidated the importance of GLRB in cancers. For example, the widespread transcription of GlyR β observed in SCLC cell lines suggests its potential as a valuable marker for neoplastic cells [ 87 ]. GLRB is also considered as a potential drug target for colorectal cancer [ 88 ]. LINC00578, a recently identified long noncoding RNA (lncRNA), has been linked to prognosis in lung, breast, and pancreatic cancers [ [89] , [90] , [91] ]. Furthermore, research suggests that LINC00578 promotes the development and growth of pancreatic cancer by inhibiting ferroptosis, a type of cell death induced by iron-dependent oxidation [ 91 ]. Long intergenic non-protein coding RNA 968 ( LINC00968 ) may function as an oncogene in certain cancers, such as osteosarcoma [ 92 ], while acting as a tumor suppressor in others, such as lung adenocarcinoma [ 93 ]. Further research is needed to elucidate the mechanisms by which LINC00968 influences cancer progression. miR-145 plays a critical role in regulating gene expression and is notably downregulated in gastric cancer cells. Restoration of miR-145 activity has been linked to reduced cell proliferation and migration, alongside an increase in apoptosis in gastric cancer cells. These findings suggest that miR-145 holds significant potential as a therapeutic target for treating gastric cancer [ 94 , 95 ]. Nbea, encoded by the NBEA gene, belongs to a large and diverse group of A-kinase anchor proteins that direct protein kinase A activity to specific subcellular locations by binding to its type II regulatory subunits. Nbea is involved in cargo transport and endomembrane compartmentalization, as well as in the formation of electrical and chemical synapses [ 96 ]. Research has demonstrated an increase in NBEA mutations in gastric cancer [ 85 ], and its overexpression has been associated with a poorer prognosis in gastric cancer patients [ 97 ]. Negr1 is a protein that belongs to the immunoglobulin LON family and functions as a GPI-anchored cell adhesion molecule. It plays a critical role in regulating synapse formation and neurite outgrowth [ 98 ]. Analysis of gene expression profiles from tumor biopsies and normal tissues revealed that NEGR1 is commonly downregulated in various human cancer tissues. Overexpression of NEGR1 has been shown to suppress proliferation, anchorage-independent growth, and migration of human ovarian cancer cells [ 99 ]. Conversely, another study identified Negr1 as a target of TGFβ, essential for maintaining the tumorigenic activity of metastatic breast cancer cells [ 100 ]. Given the controversial results regarding the role of NEGR1 in cancer, further research is necessary to elucidate its precise function and therapeutic potential in cancer progression. RUNX1T1 functions as a transcriptional repressor and is commonly regarded as a biomarker for primary pancreatic endocrine tumors, breast cancer, and colorectal cancer, as well as being a reliable predictor of patient prognosis [ 101 ]. It is reported that RUNX1T1 promoter is hypermethylated in GC, and its ectopic expression in GC inhibits cancer cells’ proliferat ion and overall is considered a tumor suppressor in GC [ 102 ]. This study highlights several promising areas for future research. Firstly, prospective studies could offer more control over sampling and data quality, addressing the limitations inherent in using retrospective data like TCGA. Additionally, given the heterogeneous nature of gastric cancer, future research should aim to minimize sampling bias by incorporating a diverse array of samples. Furthermore, the functionality of the proposed decision tree model warrants assessment on larger datasets, and further experimental analyses are necessary to evaluate the model's efficiency. The present study has several limitations. First, the use of retrospective data from the TCGA and GEO databases introduces potential biases and limits control over data quality, sampling, and cohort characteristics. Gastric cancer's inherent heterogeneity further complicates the generalizability of our findings. While the survival analysis of the TCGA cohort demonstrated clinical potential, the borderline P -value of 0.06 suggests that the observed association may lack sufficient statistical power. In contrast, the GEO-derived clusters, which exhibited similar hypoxia and stemness patterns, showed a statistically significant difference in survival rates ( P -value: 0.01). This discrepancy could arise from differences in patient demographics, treatment regimens, or tumor stages between the TCGA and GEO cohorts, which may introduce noise and obscure survival differences in the TCGA dataset. To address this, future studies should aim to increase the cohort size or incorporate additional factors such as specific genetic mutations or tumor microenvironment characteristics to enhance statistical power and robustness. Furthermore, the proposed decision tree model requires validation on larger, more diverse datasets to confirm its predictive accuracy and clinical utility. Finally, experimental validation of the bioinformatics findings is essential to ensure the reliability and effectiveness of the identified biomarkers and prognostic models in clinical settings.

Introduction

Gastric cancer (GC) is the third cause of global cancer-related fatalities and is responsible for approximately one out of every twelve oncological deaths. It is also the fifth most frequently occurring form of cancer, accounting for 5.7 % of all new cases [ 1 ]. Despite a decrease in incidence rates, which is a result of better food preservation practices and improved living conditions associated with economic development, increasing the aging population will lead to a rise in GC cases in the future [ 2 , 3 ]. The prognosis of GC patients is influenced by various independent factors, including age, surgical procedures, chemotherapy, and the tumor-node-metastasis(TNM) staging system [ 4 ]. In addition, biomarker status is crucial for guiding treatment choices, particularly before initiating first-line systemic therapy for advanced GC. Key biomarkers include human epidermal growth factor receptor 2 (HER2), programmed death-ligand 1 (PD-L1), and mismatch repair (MMR)/microsatellite instability (MSI) status. Despite their clinical utility, these biomarkers have limitations in capturing the complexity and heterogeneity of GC [ 5 ]. For example, while HER2 monoclonal antibodies added to chemotherapy regimens show promise for patients with overexpressed HER2, the prognostic value of HER2 in GC remains uncertain, as outlined in the National Comprehensive Cancer Network (NCCN) guidelines (version 4.2024). Similarly, while Epstein-Barr virus (EBV)-positive GC patients may benefit from PD-1/PD-L1 immunotherapies due to PD-L1 overexpression, the prognostic role of EBV status is still debated, and further data are needed to establish its clinical significance. These challenges highlight the limitations of existing biomarkers and the need for innovative approaches. Discovering new biomarkers that better capture the complexity of GC could enable more precise prognostication and personalized treatment strategies. This underscores the relevance of developing methods such as the prognostic decision tree presented in this study, which leverages gene expression data to address these unmet needs. Intratumor heterogeneity is a major attribute of GC which contributes to poor outcomes such as tumor relapse and chemo resistance [ 6 ]. One of the factors associated with heterogeneity is hypoxia; higher hypoxia is linked with augmented mutational load across various cancer types [ 7 ]. Hypoxia, a condition in which the amount of Oxygen in the tissues is less than 2 % [ 8 ], promotes aggressive and metastatic cancer by enhancing tumor cell proliferation, survival, immune evasion, inflammation, angiogenesis, and invasion [ 9 ]. Importantly, hypoxia activates hypoxia-inducible factors (HIFs), which regulate genes crucial for maintaining cancer stem cells (CSCs) – a small but significant tumor subpopulation with stem cell-like properties, such as self-renewal and differentiation. CSCs are critical to cancer progression and therapy resistance due to their ability to survive treatments and drive metastasis [ 10 ]. In hypoxic conditions, CSCs remain poorly differentiated, generating malignant clones that increase survival, self-renewal, and metastasis [ 11 , 12 ]. Hypoxia also boosts CSC marker expression and influences stem cell phenotypes [ 13 , 14 ]. HIFs play a central role in regulating CSCs [ 15 ], while stemness-related pathways enable these cells to evade conventional therapies [ 16 ]. By integrating hypoxia and stemness in our model, we aim to identify key genetic markers and enhance the prediction of patient outcomes. Understanding the molecular connection between hypoxia and stemness could lead to therapies that target the fundamental mechanisms driving tumor survival and resistance. Numerous investigations have been conducted regarding the hand-in-hand effects of hypoxia and stemness in cancer progression. For example, Zhao et al. showed that chemical induction of hypoxia results in an increase in lung cancer cell stemness and drug resistance in CD166-positive stem cells. Of note, patients with higher expression of CD166 showed poorer prognoses as well [ 17 ]. It is also demonstrated that post-hypoxic breast cancer cells with the ability to metastasize, exhibit an enrichment in pathways that are associated with the expression of cancer stem cell-related genes. These cells are less sensitive to chemotherapeutics and lead to recurrence after treatment [ 18 ]. In the past decade, The Cancer Genome Atlas (TCGA) has provided a detailed understanding of primary tumor landscapes through the generation of comprehensive multi-omics characteristics and annotations of pathophysiological features and clinical information [ 19 ]. Recently, many studies have employed TCGA data along with bioinformatics tools and machine learning algorithms to provide insight into underlying cellular mechanisms of GC and suggest tools to predict the patients’ prognoses, However, despite conducting an extensive search for relevant studies, we were unable to find any that report a machine learning model based on hypoxia and stemness indices to predict the prognosis of GC patients. In the present study, RNA-seq data of stomach adenocarcinoma (STAD) patients were downloaded from TCGA and normalized using the edgeR package. Hypoxia and stemness scores of the samples were calculated using GSVA (gene set variation analysis) package and mRNAsi (mRNA stemness index) methods in R, respectively. Next, based on these scores, two-dimensional hierarchical clustering and survival analysis were performed on TCGA-STAD samples and differentially expressed genes (DEGs) were then identified between the two clusters exhibiting the most significant variations in survival rates. Meanwhile, WGCNA (weighted gene co-expression network analysis) was carried out and hub genes of the survival-associated clusters were selected. The shared genes between the WGCNA hub genes and the DEGs were then identified. Survival analysis was performed on these shared genes and using the genes that were significantly associated with survival, a prognostic decision tree model was constructed and the related area under the curves (AUCs) were calculated. The flowchart of the current study is presented in Fig. 1 . Fig. 1 The workflow of the present study In the workflow above, the study is divided into three parts that are identified with different colors, the first part of the study includes clustering the TCGA-STAD RNA-seq data based on hypoxia and stemness scores (green boxes and arrows), the second part includes WGCNA and identifying the hub genes (the blue box and arrow) and the third part includes identification of the shared genes between the first and the second parts and constructing a decision tree model using the survival-related shared genes (pink boxes and arrows). DEGs: Differentially-expressed genes, ROC: receiver operating characteristic, STAD: Stomach adenocarcinoma, TCGA: The Cancer Genome Atlas, WGCNA: Weighted gene co-expression network analysis. Fig. 1 The workflow of the present study In the workflow above, the study is divided into three parts that are identified with different colors, the first part of the study includes clustering the TCGA-STAD RNA-seq data based on hypoxia and stemness scores (green boxes and arrows), the second part includes WGCNA and identifying the hub genes (the blue box and arrow) and the third part includes identification of the shared genes between the first and the second parts and constructing a decision tree model using the survival-related shared genes (pink boxes and arrows). DEGs: Differentially-expressed genes, ROC: receiver operating characteristic, STAD: Stomach adenocarcinoma, TCGA: The Cancer Genome Atlas, WGCNA: Weighted gene co-expression network analysis.

Coi Statement

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.

Data Availability

The results published or shown here are in whole or part based upon data generated by the TCGA Research Network ( https://portal.gdc.cancer.gov/legacy-archive/search/f ) and GEO database ( https://www.ncbi.nlm.nih.gov/geo/ ). Data are available upon request to the corresponding author.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-23T09:30:01.253652+00:00