{"paper_id":"2b7cc980-72c6-4c0e-93ef-5dc907f936f3","body_text":"The prognostic risk model of ESCA patients was constructed based on intercellular-related genes | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article The prognostic risk model of ESCA patients was constructed based on intercellular-related genes Wei Cao, Dacheng Jin, Weirun Min, Haochi Li, Rong Wang, Jinlong Zhang, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4460813/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 20 Jan, 2025 Read the published version in BMC Cancer → Version 1 posted 4 You are reading this latest preprint version Abstract Background Esophageal cancer is a serious malignant tumor disease. Radiotherapy is the standard treatment, but treatment tolerance often leads to failure. Cell-in-cell are observed in a variety of tumors and have been shown to correlate with prognosis. Therefore, it is particularly important to study the prognostic value and regulatory mechanism of intracellular structure-related genes in esophageal cancer. Methods TCGA Esophageal Cancer (ESCA) was included in the analysis as the training set. The differentially expressed genes in ESCA samples in the training set were analyzed, and the differentially expressed intercellular-related genes were recorded as CIC-related DEGs. Cox analysis was used to screen prognostic genes. Samples were divided into high-low-risk groups according to the median value of the ESCA sample risk score. Validation was performed in the risk model GSE53624. Morphological mapping, enrichment analysis, immune infiltration analysis, prognostic gene expression verification, molecular docking, and RT-PCR verification were established. Results A total of 38 intersection genes were obtained between the disease group and the normal group of ESCA samples. After stepwise multivariate COX analysis, three prognostic genes (AR, CXCL8, EGFR) were selected. The applicability of the risk model was verified in the GSE53624 dataset. The analysis revealed eight significantly different immune-related gene sets. The prognostic gene expression validation found that the prognostic genes reached significant differences between the disease group and the normal group in both datasets. The corresponding proteins of the three prognostic genes all interacted with Gefitinib and osimertinib. The results of PCR confirmed the differential expression of prognostic genes in esophageal cancer tissues. Conclusions Three prognostic genes, AR, CXCL8, and EGFR, were obtained in this study, and the molecular docking of prognostic genes with Gefitinib and osimertinib showed that there were interactions between them, which provided a basis for the diagnosis and treatment of ESCA. esophageal cancer bioinformatics Cell-in-cell molecular docking prognostic risk model Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 1. Introduction The exercise of Esophageal cancer (ESCA) is the seventh most common cancer and the sixth cause of cancer-related death worldwide. Despite recent advances in comprehensive treatment, the prognosis of esophageal cancer remains poor 1 . Radiotherapy is one of the standard treatment options for esophageal cancer, which can effectively kill cancer cells, relieve symptoms, and prolong the survival of patients. However, radiotherapy tolerance often occurs in patients with esophageal cancer during radical surgery, perioperative or palliative radiotherapy, which eventually leads to treatment failure and becomes one of the major obstacles in the clinical treatment of esophageal cancer 2 . Therefore, there is an urgent need to understand the molecular mechanisms underlying the development of radiotherapy tolerance in esophageal cancer and subsequently to identify the molecular targets of radiotherapy sensitization to develop new therapeutic approaches. Cell-in-cell (CIC) structure refers to the internalization of one or more living cells into another living cell to form a \"bird's eye cell\". The CIC structure has been observed in various types of tumors and is associated with poor prognosis, such as breast cancer, lung cancer, and pancreatic cancer 3 , 4 . Evaluation of CIC status is an effective method to assess prognosis. However, few studies have investigated the value of CIC-related genes and their prognostic function in ESCA. The aim of this study is to investigate the prognostic value and potential molecular regulatory mechanism of cell-in-cell-related genes (CIC-related genes) in ESCA. 2. Materials and Methods 2.1 Data extraction In this study from the UCSC Xena database ( https://xenabrowser.net/datapages/ ) to download the TCGA esophageal cancer (ESCA) RNAseq - HTSeq - Counts matrix, as a training set, At the same time, the corresponding clinical phenotype data and survival information data were downloaded. The data set contained 173 samples, including 162 disease samples, 11 normal control samples, and 161 disease samples with survival information. Data was downloaded in December 2022. From the GEO database ( https://www.ncbi.nlm.nih.gov/geo/ ) to download GSE53624 data sets, the 238 samples the transcriptome data, including 119 disease samples (have corresponding survival information), 119 samples of normal, as a validation set. A total of 101 cell-to-cell related genes were downloaded from the existing literature 4 . 2.2 Differential gene analysis between disease group and normal group The R software package \"DESeq2\" was used to analyze the differentially expressed genes between ESCA samples and normal group samples in the training set. The screening threshold was | log2FoldChange|>0.5 and pvalue < 0.05.Using the online jvenn website ( http://jvenn.toulouse.inra.fr/app/example.html ) to obtain the above steps of different genes (DEGs) and related genes between literature downloaded 101 cells take intersection, the difference between related gene expression of the cell, These were denoted as CIC-related DEGs. 2.3. Risk model construction and prognostic gene analysis In the disease group of the training set, univariate COX and stepwise multivariate COX regression analysis were performed on CIC-related DEGs to obtain prognostic genes. Correlation analysis was performed on prognostic genes, and a risk model was constructed. The patients were grouped according to the median value of risk score, and K-M survival curve analysis was performed on the high and low-risk groups. The ROC curve was drawn by survivalROC to verify the accuracy of the survival curve scoring model, the risk curve and the expression heat map of prognostic genes in the high and low-risk groups were drawn, and the validation set GSE53624 dataset was used to verify the results. 2.4. Construction of nomogram model for prognostic factors The \"RMS\" package was used to construct nomograms for risk score, age, sex, TNM stage, and 1-year and 3-year survival rates. The calibration curve, decision curve, and ROC of the nomogram were used to evaluate the predictive ability of the nomogram model. 2.5. Enrichment analysis of high and low-risk groups Based on the above grouping information of high and low-risk groups, DESEq2 software was used to analyze the differentially expressed genes between high and low-risk groups in the training set, and the screening threshold was | log2FoldChange|>0.5, pvalue < 0.05. ClusterProfiler package was used to perform GO and KEGG enrichment analysis of DEGs between high and low-risk groups, and R package \"ggplot2\" was used to visualize the enrichment results. 2.6. Analysis of immune infiltration in the high and low-risk groups Based on the grouping information of high and low-risk groups, in the training set, the ssGSEA package was used to analyze the immune infiltration between high and low-risk groups, the composition of immune cells in the training set, and the differentially expressed immune cells between high and low-risk groups were observed. 2.7. Prognostic gene expression validation The expression matrix of prognostic genes was extracted in the training set and validation set, and t.test was used to check the expression of prognostic genes. 2.8. Molecular DOCKING Analysis The crystal structures of the prognostic genes were downloaded from the PDB database ( https://www.rcsb.org/ ), and the molecular docking of the crystal structures with Gefitinib and osimertinib was performed by AutoDock4.2 software. 2.9. RT-PCR Five pairs of frozen esophageal cancer samples, 5 pairs of cancer tissues, and 5 pairs of adjacent tissues were collected from the Department of Thoracic Surgery of Gansu Provincial People's Hospital. 50mg of tissue was taken from each sample for total RNA extraction. The SureScript-First-strand-cDNA-synthesis-kit was used for reverse transcription and detection. The relative RNA quantity was calculated using the relative standard method (2 −△△CT ). The primers and their sequences are shown in Table 1 Table 1 Primers and their sequences Primers Sequences AR F AGGCAGTGTCGGTGTCCATG AR R CCTTTGGTGTAACCTCCCTTGA CXCL8 F TCTGCAGCTCTGTGTGAAGG CXCL8 R TTCTCAGCCCTCTTCAAAAACT EGFR F GCCAAGGCACGAGTAACAAGC EGFR R AGGGCAATGAGGACATAACCAG GAPDH F CGAAGGTGGAGTCAACGGATTT GAPDH R ATGGGTGGAATCATATTGGAAC 2.10 Statistical Analysis Statistical analysis and graphics were performed and obtained using R software (version 4.2.0) and GraphPad Prism (version 8.0). Data were presented as mean ± stan-dard deviation. Differences were analyzed via the Wilcoxon rank-sum test. The p < 0.05 represented a significant difference. 3. Results 3.1 Analysis of differentially expressed genes between disease and normal groups in the training set The distribution of differential genes is shown in Fig. 1 A- 1 B. A total of 12985 differentially expressed genes (DEGs) were obtained, including 6617 up-regulated genes and 6368 down-regulated genes. 3.2 Identification of CIC-related DEGs The results are shown in Fig. 1 C. A total of 38 intersection genes were obtained, which were as follows: AR, AURKA, CDC20, CDH3, CDKN2A, CEBPB, CTSB, CTSL, CTTN, CXCL8, EGFR, ERI3, EZR, GULP1, GZMB, HTT, KIF2C, KRT7, MAD2L1, MAPT, MLANA, MYLK, NTS, PCDH7, PDPN, PRKAA2, PRNP, RAC1, RNF146, SFTPB, SPRN, STMN2, TF, TGFB1, TNFSF10, TP53, TP63, WT1. These were denoted as CIC-related DEGs. 3.3 Construction of risk models 3.3.1 Univariate COX regression analysis We used the TCGA-ESCA dataset as the training set and performed univariate Cox proportional hazards regression analysis using 38 CIC-related DEGs, with the univariate p-value set to 0.2 and a total of 5 genes; the results and forest plots are shown in Table 2 and Fig. 2 A, respectively. 3.3.2 Multivariate COX regression analysis Three prognostic genes (AR, CXCL8, EGFR) were obtained by stepwise multivariate COX analysis using the step function, and the multivariate coefficient (coef) of each prognostic gene was calculated to construct the survival risk model. The multivariate results and forest plots are shown in Table 3 and Fig. 2 B, respectively. Then in the ESCA group of the training set, the expression matrix of prognostic genes (AR, CXCL8, EGFR) was extracted, the correlation between prognostic genes was calculated by the \"spearman\" method, and correlation Chord diagram was plotted using the R software \"circlize\" package, and the results were shown in Fig. 2 C. 3.3.3 Effect of prognostic genes on survival To evaluate the prognostic value of the risk model, the expression levels of three prognostic genes were obtained in the TCGA-ESCA dataset, and the formula was used: $${\\mathbf{R}\\mathbf{i}\\mathbf{s}\\mathbf{k}\\mathbf{s}\\mathbf{c}\\mathbf{o}\\mathbf{r}\\mathbf{e}}_{\\mathbf{s}\\mathbf{a}\\mathbf{m}\\mathbf{p}\\mathbf{l}\\mathbf{e}}={\\sum }_{\\varvec{n}=1}^{\\varvec{n}}(\\varvec{c}\\varvec{o}\\varvec{e}{\\varvec{f}}_{\\varvec{i}}\\ast {\\varvec{x}}_{\\varvec{i}})$$ (Note: Coefi represents the multivariate regression coefficient of the ith gene, xi represents the expression value of the ith gene, and n represents the number of model genes.) The risk value of each patient was calculated, and 161 patients with survival information in the training set were divided into high and low-risk groups according to the median value of risk value (1.047788). There were 80 samples in the high-risk group and 81 samples in the low-risk group. The results of the survival analysis are shown in Fig. 2 D. There was a significant survival difference between the high and low-risk groups in the TCGA-ESCA data set (P < 0.05). 3.3.4 ROC curve and risk curve The multivariate COX model was used to calculate the RiskScore, the survival ROC package was used to calculate the false positive and true positive, the ROC curve was drawn using the results, and the AUC was calculated as shown in Fig. 2 E. The AUC values of 1-year and 3-year were 0.74 and 0.62, respectively, indicating that the risk regression model constructed could be effectively used as a prognostic model. The risk curve is shown in Fig. 2 G, with samples divided into high and low-risk groups according to the median value. 3.3.5 Expression heatmap of prognostic genes in the high and low-risk groups The R package \"pheatmap\" was used to map the expression of prognostic genes in the high and low-risk groups, and the results are shown in Fig. 2 F. Table 2 Results of univariate analysis id pvalue HR HR.95L HR.95H EGFR 0.025465674 0.816222106 0.683036621 0.975377463 AR 0.032093467 0.816018744 0.677565483 0.982763448 CXCL8 0.051096995 1.154841387 0.999313898 1.334574283 HTT 0.06904872 0.699816164 0.476290936 1.028242669 TP53 0.164783007 0.867781691 0.710418302 1.060002342 Table 3 Results of multivariate analysis id coef HR HR.95L HR.95H pvalue AR -0.178187505 0.836785507 0.699135927 1.001536265 0.051989207 CXCL8 0.137065624 1.14690341 0.985768463 1.334377678 0.075996366 EGFR -0.199620941 0.819041159 0.688642196 0.97413203 0.024059355 3.4 Survival risk model validation 3.4.1 Effect of prognostic genes on survival To verify the applicability of the risk model, we obtained the expression levels of three prognostic genes in the GSE53624 dataset and used the formula: $${\\mathbf{R}\\mathbf{i}\\mathbf{s}\\mathbf{k}\\mathbf{s}\\mathbf{c}\\mathbf{o}\\mathbf{r}\\mathbf{e}}_{\\mathbf{s}\\mathbf{a}\\mathbf{m}\\mathbf{p}\\mathbf{l}\\mathbf{e}}={\\sum }_{\\varvec{n}=1}^{\\varvec{n}}(\\varvec{c}\\varvec{o}\\varvec{e}{\\varvec{f}}_{\\varvec{i}}\\ast {\\varvec{x}}_{\\varvec{i}})$$ (Note: Coefi represents the multivariate regression coefficient of the ith gene, xi represents the expression value of the ith gene, and n represents the number of model genes.) The risk value of each patient was calculated, and 119 patients in the training set were divided into high and low-risk groups according to the median value of risk value (0.9625565), including 59 samples in the high-risk group and 60 samples in the low-risk group. The results of the survival analysis of the high and low-risk groups are shown in Fig. 3 A. It can be seen that there is a significant survival difference between the high and low-risk groups in the GSE53624 data set (p < 0.05). 3.4.2 ROC curve and risk curve The multivariate COX model was used to calculate the RiskScore, the survival ROC package was used to calculate the false positive and true positive, and the results were used to draw the ROC curve, and the AUC was calculated as shown in Fig. 3 B. The AUC values of 1-year and 3-year were both above 0.6, indicating that the risk regression model constructed could be effectively used as a prognostic model. As shown in Fig. 3 C, samples were divided into high and low-risk groups according to the median value. 3.4.3 Expression heatmap of prognostic genes in the high and low risk groups The R package \"pheatmap\" was used to map the expression of prognostic genes in the high and low-risk groups, and the results are shown in Fig. 3 D. 3.5 Construction of nomogram model for prognostic factors Based on risk score, age, sex, and TNM stage, the \"RMS\" package was used to construct survival nomograms at 1 and 3 years for several clinical factors in the risk model. Each factor corresponds to a score. The sum of the total scores of each factor corresponds to the total score. The 1-year and 3-year survival rates were predicted according to the total score. The higher the score, the lower the survival rate. The nomogram (Fig. 4 A) and calibration curve (Fig. 4 B) were constructed. The 1 - and 3-year ROC curves of the nomogram were drawn using the \"pROC\" package (Fig. 4 C), and the decision curve (DCA) was drawn using the \"rmda\" package (Fig. 4 D). The calibration curve showed that the error between the actual risk of disease and the predicted risk was small, indicating that the nomogram model had high prediction accuracy for ESCA disease. The AUC areas corresponding to the 1-year and 3-year ROC curves of the nomogram were both greater than 0.7, indicating that the nomogram model had high prediction accuracy for prognosis. Decision curve analysis (DCA) showed that within the high-risk threshold range of 0–1, the nomogram model could benefit, and the clinical benefit of the nomogram model was higher than that of the \"age\" curve, \"riskScore\" curve, \"T\" curve, \"N\" curve, \"M\" curve and \"gender\" curve. 3.6 Enrichment analysis of high and low-risk groups Based on the expression matrix in the TCGA-ESCA dataset, 80 samples from the high-risk group and 81 samples from the low-risk group were extracted (TCGA-ESCA), and the \"DESeq2\" package was used for differential gene analysis. The screening conditions for differential genes were | log2FoldChange|>0.5. p-value < 0.05. There were 5727 significantly differentially expressed genes in high-vs-low, including 2965 up-regulated genes and 2762 down-regulated genes. As shown in Fig. 5 A, the volcano plot is used to show the distribution of differential genes, and as shown in Fig. 5 B, the heat map is used to show the expression of differential genes among sample groups. ClusterProfiler package in R was used to perform (GO and KEGG) functional enrichment analysis of differential genes in high and low-risk groups, and the screening threshold was as follows: p-value < 0.05, a total of 887 GO BP, 67 GO CC, 178 GO MF and 53 KEGG pathways were screened. The enrichment results were visualized by the R package \"ggplot2\", as shown in Fig. 5 C- 5 D. We ranked according to p-value value. Each TOP10 Descriptions of GO and KEGG were selected for display. 3.7 Analysis of immune infiltration in the high and low risk groups Based on the expression matrix and grouping information of high and low-risk groups in the training set (TCGA-ESCA), the expression levels of 29 immune gene sets (16 immune cells and 13 immune-related pathways gene sets) in each sample of the training set were drawn by ssGSEA (single sample GSEA) analysis. And using t.test() to test, it was found that there were 8 significantly different immune-related gene sets in the training set, which were as follows: B_cells, iDCs, Macrophages, Mast_cells, NK_cells, T_cell_co − inhibition, T_helper_cells, Type_II_IFN_Reponse, the results are shown in Fig. 6 A. 3.8 Prognostic gene expression validation The expression matrices of prognostic genes (AR, CXCL8, EGFR) were extracted from the training set (TCGA-ESCA) and the validation set GSE53624, respectively. Based on the grouping information in the respective data sets, the t.test() test was used. The R software package \"ggplot2\" was used to draw the boxplot of prognostic gene expression in the two datasets, and it was found that the three prognostic genes reached significant differences between the disease group and the normal group in the two datasets, and the expression trend was consistent. The results are shown in Fig. 6 B and Fig. 6 C. 3.9 Molecular DOCKING Analysis Docking analysis was performed using three prognostic genes (AR, CXCL8, EGFR) and two drugs (Gefitinib and osimertinib). Firstly, the corresponding crystal structures of AR, CXCL8, and EGFR were downloaded from the PDB database ( https://www.rcsb.org/ ), respectively. The corresponding relationship between genes and PDB ID was AR-1T7R, CXCL8-7JNY, and EGFR-1IP0. Molecular docking using AutoDock revealed two hydrogen bonds and two residues in the 1t7r and Gefitinib docking. There was one hydrogen bond and one residue in the docking between 1t7r and osimertinib. There were three hydrogen bonds and two residues in the docking between 7jny and Gefitinib. There are three hydrogen bonds and two residues in the docking of 7jny and osimertinib. There were two hydrogen bonds and two residues in the docking between 1ip0 and Gefitinib. There were 4 hydrogen bonds and 4 residues in the docking of 1ip0 and osimertinib, and all three prognostic gene pairs interacted with the two drugs. Figure 7 A- 7 B shows the docking results of 1ip0 with Gefitinib and osimertinib, and other molecular docking results are detailed in the Supplementary material. 3.10 RT-PCR Five pairs of frozen esophageal cancer samples were obtained from the Department of Thoracic Surgery, Gansu Provincial People's Hospital for RT-pcr experiments. Results As shown in Fig. 9, AR was lowly expressed in esophageal cancer tissues (p < 0.0001), while CXCL8 and EGFR were highly expressed in esophageal cancer tissues (p < 0.05). The RT-pcr results were consistent with the previous analysis. 4. Discussion Esophageal cancer is a malignant lesion formed by abnormal proliferation of esophageal squamous epithelium or glandular epithelium, which is one of the most common malignant tumors in the world 5 .It ranks sixth in cancer mortality 6 .The early symptoms of esophageal cancer are often not obvious, so the patients have progressed to the middle and late stages once diagnosed, and 38% of the cases are diagnosed in the late stage, missing the best opportunity for treatment. Currently, the 5-year survival rate for patients with esophageal cancer is 10% 7 . Therefore, it is particularly important to understand the mechanism of the occurrence and development of esophageal cancer and to develop new treatment methods. During tumor cell culture, some large cells with internal vacuoles and multiple nuclei are found, that is, the situation of one cell inside another, and are called cell-in-cell are observed in a variety of tumors and have been shown to correlate with prognosis. Therefore, it is particularly important to study the prognostic value and regulatory mechanism of intracellular structure-related genes in esophageal cancer 8 . CIC structure has been observed in various tumor types and is associated with a poor prognosis, with one clinical study showing an association of CIC structure with aggressive features and cancer-related mortality in oral tongue cancer 9 . However, the role of CIC structure in esophageal cancer remains unclear. We used the GDC TCGA Esophageal Cancer (ESCA) dataset as the training set for this study. The GSE53624 dataset was downloaded from the GEO database and used as the validation set. A total of 101 cell-related genes were downloaded from the existing literature 4 . Based on the above data, a series of bioinformatics analyses were carried out in this study. First, we performed differential gene analysis between the disease group and the normal group, and took the intersection of the obtained differential genes and 101 cell-related genes, and a total of 38 intersection genes were obtained, which were named CIC-related DEGs. The following 38 genes were analyzed by univariate and multi-step multivariate analysis, and 3 prognostic genes were obtained, namely AR, CXCL8 and EGFR. We found that CXCL8 and EGFR were up-regulated in esophageal cancer. CXCL8, also known as interleukin-8 (IL-8), is a neutrophil-activating factor, which is mainly produced by neutrophils, monocytes, macrophages, T cells, epithelial cells, and endothelial cells. CXCL8 has the same expression trend in ovarian cancer, breast cancer, liver cancer 10 – 12 , head and neck squamous cell carcinoma, and has prognostic and diagnostic value 13 . Silencing CXCL8 in esophageal cancer cells can inhibit the proliferation and invasion of esophageal cancer cells 14 . Helen et al. found that CXCL8 gene expression is closely related to epithelial-mesenchymal transition and cell-matrix vascularization, thereby promoting tumor cell invasion and metastasis, which explains to a certain extent the ability of CXCL8 gene to predict tumor metastasis 15 . EGFR mainly exists in the stroma, epidermis, and some smooth muscle cells. After activation, EGFR can activate the downstream mitogen-activated protein kinase (MAPK) signaling pathway, inhibit cell apoptosis and DNA repair, accelerate cell invasion 16 , and participate in the occurrence and development of lung cancer, colorectal cancer, breast cancer, and other malignant tumors 17 – 19 . Subsequently, 29 immune-related gene sets were selected in this study. Through immune infiltration analysis of high and low-risk groups, 8 immune-related gene sets with significant differences were found in the training set, which were as follows: B_cells, iDCs, Macrophages, Mast_cells, NK_cells, T_cell_co − inhibition, T_helper_cells, Type_II_IFN_Reponse, Then, the expression levels of the three prognostic genes in the training set and validation set were analyzed, and the expression trends of the three prognostic genes in the training set and validation set were consistent. In recent years, molecular docking methods have become an important technique in the field of computer-aided drug research. Molecular docking is a method for drug design through the characteristics of receptors and the way of interaction between receptors and drug molecules. A theoretical simulation method that mainly studies the interaction between molecules (such as ligands and receptors) and predicts their binding modes and affinities 20 . Autodock is an open-source molecular simulation software, which is mainly used to perform ligand-protein molecular docking. AutoDock uses a semi-flexible docking method to allow conformational changes of small molecules, and the binding free energy is used as the basis for the evaluation of docking results 21 . We performed molecular docking of the three prognostic genes with Gefitinib and osimertinib, and found that there were interactions between them, indicating that the prognostic genes may affect the efficacy of drugs or the drugs may affect the expression of prognostic genes. Osimertinib belongs to the third generation of EG-FR-TKIs class and has been used in the treatment of lung cancer 22 . Some scholars have found that the use of osimertinib can lead to severe complications including lung injury, especially interstitial lung disease (ILD). Due to impaired epithelial healing of lung injury, pulmonary inflammation may worsen after EGFR-TKI treatment 23 . Gefitinib is a specific small molecule EGFR tyrosine kinase inhibitor, which can inhibit the growth of esophageal cancer cells by inhibiting the signal transduction pathway of EGFR tyrosine kinase and inhibiting the phosphorylation of EGFR. It also has the effect of inhibiting tumor angiogenesis and metastasis and plays the purpose of treating tumors together 24 . It also has the effect of inhibiting tumor angiogenesis and metastasis and plays the purpose of treating tumors together 24 . The inhibitory effect of gefitinib on metastatic non-small cell lung cancer, ovarian cancer, and other tumors has been widely recognized 25 . Our study found that Gefitinib and osimertinib not only interacted with EGFR but also interacted with AR and CXCL8, which may have implications for the development of targeted drugs in the future. 5. Conclusion In this study, the TCGA-ESCA transcriptome data were analyzed by bioinformatics methods, and three prognostic genes were obtained, namely AR, CXCL8, and EGFR, which were verified by RT-PCR. Moreover, molecular docking of prognostic genes with Gefitinib and osimertinib indicated that there were interactions between them. It is suggested that more attention and in-depth study are needed in the research of prognostic genes and drugs to better understand the pathogenesis of the disease and improve the therapeutic effect. Therefore, this study provides an important reference for the diagnosis, mechanism research, and treatment of esophageal cancer in the future. However, the influence of the \"intracellular cell\" structure on the biological behavior of esophageal cancer cells and the influence on the body's immunity still needs to be confirmed by in vivo studies, and the molecular docking results still need to be further confirmed, and verified by subsequent related experiments. Declarations Funding This study was supported by grants from the Gansu Provincial People's Hospital Research Funding (22GSSYD-30; 22GSSYD-25, 22GSSYC-9), Gansu Youth Science and Technology Fund (22JR11RA241)， and Science and Technology Department of Gansu Province (22YF7FA095) Declaration of competing interest The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. Data Availability Statements The data that supports the findings of this study are available in the supplementary material of this article. Ethics approval statement This topic project has passed the Gansu Province People's Hospital Ethics Committee for examination and approval. Author contribution statement Wei Cao: Conceptualization (Equal), Data curation (Equal), Funding acquisition (Equal), Investigation (Equal), Project administration (Equal), Resources (Equal), Writing – original draft (Equal) Dacheng Jin: Conceptualization (Equal), Data curation (Equal), Funding acquisition (Equal), Methodology (Equal), Project administration (Equal), Resources (Equal), Writing – original draft (Equal) WeiRun Min: Data curation (Supporting), Formal analysis (Equal), Project administration (Supporting), Resources (Supporting), Software (Supporting), Supervision (Supporting), Writing – review & editing (Equal) Haochi Li: Conceptualization (Supporting), Funding acquisition (Equal), Methodology (Supporting), Project administration (Supporting), Resources (Supporting), Supervision (Supporting), Writing – review & editing (Supporting) Rong Wang: Formal analysis (Supporting), Investigation (Supporting), Methodology (Supporting), Project administration (Supporting), Supervision (Supporting), Visualization (Equal) Jinlong Zhang: Data curation (Lead), Formal analysis (Supporting), Software (Supporting), Supervision (Supporting), Writing – review & editing (Supporting) YunJiu Gou (Corresponding Author): Data curation (Lead), Formal analysis (Lead), Funding acquisition (Lead), Methodology (Lead), Project administration (Lead), Software (Lead), Supervision (Lead), Writing – review & editing (Lead) *Author Wei Cao and Dacheng Jin share first authorship. References Wang M, Sun X, Xin H, Wen Z, Cheng Y. SPP1 promotes radiation resistance through JAK2/STAT3 pathway in esophageal carcinoma. Cancer Med. 2022;11:4526–43. Shitara K, et al. Nivolumab plus chemotherapy or ipilimumab in gastro-oesophageal cancer. Nature. 2022;603:942–8. Song J, et al. Cell-in-Cell-Mediated Entosis Reveals a Progressive Mechanism in Pancreatic Cancer. Gastroenterology. 2023;165:1505–e152120. Song J, et al. Construction of a novel model based on cell-in-cell-related genes and validation of KRT7 as a biomarker for predicting survival and immune microenvironment in pancreatic cancer. BMC Cancer. 2022;22:894. Rogers JE, Sewastjanow-Silva M, Waters RE, Ajani JA. Esophageal cancer: emerging therapeutics. Expert Opin Ther Targets. 2022;26:107–17. Siegel RL, Miller KD, Fuchs HE, Jemal A. Cancer statistics, 2022. CA Cancer J Clin. 2022;72:7–33. Huang F-L, Yu S-J. Esophageal cancer: Risk factors, genetic association, and treatment. Asian J Surg. 2018;41:210–5. Fais S, Overholtzer M. Cell-in-cell phenomena in cancer. Nat Rev Cancer. 2018;18:758–66. Siquara da Rocha L, de Souza O, de Lambert BS. Gurgel Rocha, C. de A. Cell-in-Cell Events in Oral Squamous Cell Carcinoma. Front Oncol. 2022;12:931092. Fu X, Wang Q, Du H, Hao H. CXCL8 and the peritoneal metastasis of ovarian and gastric cancer. Front Immunol. 2023;14:1159061. Mishra A, Suman KH, Nair N, Majeed J, Tripathi V. An updated review on the role of the CXCL8-CXCR1/2 axis in the progression and metastasis of breast cancer. Mol Biol Rep. 2021;48:6551–61. Yang S, Wang H, Qin C, Sun H, Han Y. Up-regulation of CXCL8 expression is associated with a poor prognosis and enhances tumor cell malignant behaviors in liver cancer. Biosci Rep. 2020;40:BSR20201169. Li Y, et al. Analysis of the Prognosis and Therapeutic Value of the CXC Chemokine Family in Head and Neck Squamous Cell Carcinoma. Front Oncol. 2020;10:570736. Yue D, et al. NEDD9 promotes cancer stemness by recruiting myeloid-derived suppressor cells via CXCL8 in esophageal squamous cell carcinoma. Cancer Biol Med. 2021;18:705–20. Ha H, Debnath B, Neamati N. Role of the CXCL8-CXCR1/2 Axis in Cancer and Inflammatory Diseases. Theranostics. 2017;7:1543–88. Levantini E, Maroni G, Del Re M, Tenen DG. EGFR signaling pathway as therapeutic target in human cancers. Semin Cancer Biol. 2022;85:253–75. Oxnard GR, et al. Germline EGFR Mutations and Familial Lung Cancer. J Clin Oncol. 2023;41:5274–84. Cheng W-L, et al. The Role of EREG/EGFR Pathway in Tumor Progression. Int J Mol Sci. 2021;22:12828. M N, V, M., F, M., P P. Crosstalk between CXCR4/ACKR3 and EGFR Signaling in Breast Cancer Cells. Int J Mol Sci 23, (2022). Crampon K, Giorkallos A, Deldossi M, Baud S, Steffenel LA. Machine-learning methods for ligand-protein molecular docking. Drug Discov Today. 2022;27:151–64. Eberhardt J, Santos-Martins D, Tillack AF, Forli S. AutoDock Vina 1.2.0: New Docking Methods, Expanded Force Field, and Python Bindings. J Chem Inf Model. 2021;61:3891–8. Cheng Y, et al. Osimertinib Versus Comparator EGFR TKI as First-Line Treatment for EGFR-Mutated Advanced NSCLC: FLAURA China, A Randomized Study. Target Oncol. 2021;16:165–76. Ohmori T, et al. Molecular and Clinical Features of EGFR-TKI-Associated Lung Injury. Int J Mol Sci. 2021;22:792. Noronha V, et al. Gefitinib Versus Gefitinib Plus Pemetrexed and Carboplatin Chemotherapy in EGFR-Mutated Lung Cancer. J Clin Oncol. 2020;38:124–36. Murphy M, Stordal B. Erlotinib or gefitinib for the treatment of relapsed platinum pretreated non-small cell lung cancer and ovarian cancer: a systematic review. Drug Resist Updat. 2011;14:177–90. Additional Declarations No competing interests reported. Supplementary Files Additionalmaterial.zip Cite Share Download PDF Status: Published Journal Publication published 20 Jan, 2025 Read the published version in BMC Cancer → Version 1 posted Editorial decision: Revision requested 29 May, 2024 Submission checks completed at journal 23 May, 2024 Editor assigned by journal 23 May, 2024 First submitted to journal 22 May, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {\"props\":{\"pageProps\":{\"initialData\":{\"identity\":\"rs-4460813\",\"acceptedTermsAndConditions\":true,\"allowDirectSubmit\":false,\"archivedVersions\":[],\"articleType\":\"Research Article\",\"associatedPublications\":[],\"authors\":[{\"id\":308090329,\"identity\":\"038b5904-bba9-42ef-9b21-35b22f77ecae\",\"order_by\":0,\"name\":\"Wei Cao\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Gansu University of Chinese Medicine\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Wei\",\"middleName\":\"\",\"lastName\":\"Cao\",\"suffix\":\"\"},{\"id\":308090330,\"identity\":\"745d1c56-353e-4130-9f74-e3ae65a8d8fb\",\"order_by\":1,\"name\":\"Dacheng Jin\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Chest Clinic Center, Gansu Provincial People's Hospital\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Dacheng\",\"middleName\":\"\",\"lastName\":\"Jin\",\"suffix\":\"\"},{\"id\":308090331,\"identity\":\"5cb0c0fc-019d-4de5-b9f0-ab08793b416a\",\"order_by\":2,\"name\":\"Weirun Min\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Gansu University of Chinese Medicine\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Weirun\",\"middleName\":\"\",\"lastName\":\"Min\",\"suffix\":\"\"},{\"id\":308090332,\"identity\":\"5d6df6ee-9a9f-472b-9014-0666da49a0b2\",\"order_by\":3,\"name\":\"Haochi Li\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Gansu University of Chinese Medicine\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Haochi\",\"middleName\":\"\",\"lastName\":\"Li\",\"suffix\":\"\"},{\"id\":308090333,\"identity\":\"9a10d7b2-46c1-4b85-8af5-0adb33f8ae22\",\"order_by\":4,\"name\":\"Rong Wang\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Gansu University of Chinese Medicine\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Rong\",\"middleName\":\"\",\"lastName\":\"Wang\",\"suffix\":\"\"},{\"id\":308090334,\"identity\":\"40253355-590c-41bd-9503-b8e318bb4bdc\",\"order_by\":5,\"name\":\"Jinlong Zhang\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"Gansu University of Chinese Medicine\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Jinlong\",\"middleName\":\"\",\"lastName\":\"Zhang\",\"suffix\":\"\"},{\"id\":308090335,\"identity\":\"d3cb3b9a-d468-46bb-a12f-59100f7b225f\",\"order_by\":6,\"name\":\"Yunjiu Gou\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAvElEQVRIiWNgGAWjYPACNh4G9sbGhx9I08JzuNlYgjSLJNLbBHiIUSjv3mMmzbuDT8bg5sM2BgkGOzndBgJaDM+cAWo5w8ZjcDux7UEBQ7Kx2QFCWmbkALW0gbW0G0gwHEjcRryWmwfbJHiI0SIvAdNyg5FILQY8x4ot5wK1SJ5JBAayARF+kW9v3njjbdsxe77jxx8+/FBhJ0dQi8EBBhZgBB6DcQkoB9vSwMAMTCY1RCgdBaNgFIyCEQsAbCM9brLDV1IAAAAASUVORK5CYII=\",\"orcid\":\"\",\"institution\":\"Chest Clinic Center, Gansu Provincial People's Hospital\",\"correspondingAuthor\":true,\"prefix\":\"\",\"firstName\":\"Yunjiu\",\"middleName\":\"\",\"lastName\":\"Gou\",\"suffix\":\"\"}],\"badges\":[],\"createdAt\":\"2024-05-22 12:15:28\",\"currentVersionCode\":1,\"declarations\":\"\",\"doi\":\"10.21203/rs.3.rs-4460813/v1\",\"doiUrl\":\"https://doi.org/10.21203/rs.3.rs-4460813/v1\",\"draftVersion\":[],\"editorialEvents\":[{\"content\":\"https://doi.org/10.1186/s12885-025-13483-8\",\"type\":\"published\",\"date\":\"2025-01-20T15:58:21+00:00\"}],\"editorialNote\":\"\",\"failedWorkflow\":false,\"files\":[{\"id\":57953359,\"identity\":\"8e9637d8-65e8-4378-abf4-243cbfdf82ba\",\"added_by\":\"auto\",\"created_at\":\"2024-06-07 23:01:49\",\"extension\":\"png\",\"order_by\":1,\"title\":\"Figure 1\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":1053852,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eAnalysis of differentially expressed genes. (A) Volcano plot of differentially expressed genes. (B) Heat map of differentially expressed genes between sample groups. (C) Intersection Venn diagram of differentially expressed genes and related genes between cells.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"image1.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-4460813/v1/23178b0c289c7e631cb51092.png\"},{\"id\":57953357,\"identity\":\"f40810db-7bca-4860-afc4-86c8c82ddcc2\",\"added_by\":\"auto\",\"created_at\":\"2024-06-07 23:01:49\",\"extension\":\"png\",\"order_by\":2,\"title\":\"Figure 2\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":379084,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eConstruction of the risk model (A) forest plot of univariate COX results. (B) forest plot of multivariate COX results. (C) Chord diagram of correlation between prognostic genes, the color of the lines represents the strength of the correlation, and the redder the more positive the correlation. (D) K-M survival curve of RiskScore, the difference between the high-risk group and the low-risk group was less than 0.05, showing a significant difference. (E) ROC curve to evaluate the effectiveness of risk model (F) heatmap of expression of prognostic genes between high and low-risk groups (G) risk curve of the validation set for high and low-risk groups.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"image2.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-4460813/v1/2162c2c336706a8b15bcf36a.png\"},{\"id\":57953364,\"identity\":\"98a03aed-198f-480e-8c15-d7927394a1d7\",\"added_by\":\"auto\",\"created_at\":\"2024-06-07 23:01:49\",\"extension\":\"png\",\"order_by\":3,\"title\":\"Figure 3\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":1028741,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eSurvival risk model validation. (A) K-M survival curve of RiskScore in the validation set. The significance of the difference between the high and low-risk groups was less than 0.05, indicating a significant difference. (B) The ROC curve of the validation set was used to evaluate the effectiveness of the risk model. (C) The risk curve of the validation set for the high and low-risk groups and (D) the expression heatmap of prognostic genes between the high and low-risk groups.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"image3.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-4460813/v1/bd47b5ee2063d808a4fcfd37.png\"},{\"id\":57953362,\"identity\":\"c2c3b70d-ff01-43d3-8604-78297dc4d809\",\"added_by\":\"auto\",\"created_at\":\"2024-06-07 23:01:49\",\"extension\":\"png\",\"order_by\":4,\"title\":\"Figure 4\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":989453,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003e\\u003cstrong\\u003eConstruction of a nomogram model for prognostic factors.\\u003c/strong\\u003e (A) Nomogram model for clinical factors. (B) Calibration curve to evaluate the predictive power of the nomogram model. The closer the slope of the calibration curve is to 1, The more accurate the prediction was, the more accurate the prediction was. (C) ROC curves of 1 and 3 years survival of the nomogram model. (D) DCA curves to evaluate the clinical application value of the nomogram model. The nomogram curve was higher than the gray line, the \\\"age\\\" curve, the \\\"riskScore\\\" curve, the \\\"T\\\" curve, the \\\"N\\\" curve, the \\\"M\\\" curve, and the \\\"gender\\\" curve, indicating that the nomogram model could benefit from the high-risk threshold range of 0-1.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"image4.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-4460813/v1/b6c756e24567c929ce5624ad.png\"},{\"id\":57953361,\"identity\":\"76bf4e60-fbc0-45d5-a50d-b38f09b7387c\",\"added_by\":\"auto\",\"created_at\":\"2024-06-07 23:01:49\",\"extension\":\"png\",\"order_by\":5,\"title\":\"Figure 5\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":1572695,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eEnrichment analysis of the high and low-risk groups. (A) Volcano plot of differentially expressed genes between high and low-risk groups in the training set. (B) Heat map of differentially expressed genes between high and low-risk groups in the training set. (C) Bubble plot of GO TOP10 enrichment results of differentially expressed genes between high and low-risk groups. The vertical axis is the enriched GOTerm, and the horizontal axis is GeneRatio, which refers to the ratio of the number of genes related to the Term in the differential gene to the total number of differential genes. (D) Bubble plot of KEGGTOP10 enrichment results of differential genes in high and low-risk groups. The ordinate is the enriched KEGG Pathway, and the abscissa is GeneRatio.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"image5.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-4460813/v1/72bf95939f3038d9fd6a5fc3.png\"},{\"id\":57953660,\"identity\":\"5fef64d2-bb4e-4fa4-9ad9-ca53708b1162\",\"added_by\":\"auto\",\"created_at\":\"2024-06-07 23:09:49\",\"extension\":\"png\",\"order_by\":6,\"title\":\"Figure 6\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":769162,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eImmune infiltration analysis and prognostic gene expression validation. (A) Box plot of the abundance of 29 immune gene sets in the high and low-risk groups of the training set. The vertical axis represents the infiltration level, and the horizontal axis represents the immune cell gene set. (B) Boxplot of prognostic gene expression in the training set. (C) Boxplot of prognostic gene expression in the validation set. (*P\\u0026lt;0.05, **P\\u0026lt;0.01, *** P\\u0026lt;0.001, **** P\\u0026lt;0.0001)\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"image6.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-4460813/v1/3cbab47b5c38ccdafc52be44.png\"},{\"id\":57953363,\"identity\":\"5755fca9-b227-46b1-868d-ad905bdbe38b\",\"added_by\":\"auto\",\"created_at\":\"2024-06-07 23:01:49\",\"extension\":\"png\",\"order_by\":7,\"title\":\"Figure 7\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":1269588,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eMolecular docking analysis. (A) docking results of 1ip0 with Gefitinib and (B) molecular docking results of 1ip0 with osimertinib. In the figure, the cyan ring model is the active molecule 1ip0, the stick structure near 1ip0 is the amino acid residue that has hydrogen bond interaction with the active molecule, the yellow dotted line is the hydrogen bond formed between the active molecule and the amino acid residue, and the red part is the amino acid residue.\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"image7.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-4460813/v1/71b3fbc25402ecd256cb8582.png\"},{\"id\":57953661,\"identity\":\"55a4faaa-03b0-4f86-b96c-12c36ddd28e2\",\"added_by\":\"auto\",\"created_at\":\"2024-06-07 23:09:49\",\"extension\":\"png\",\"order_by\":8,\"title\":\"Figure 8\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":66559,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003ePCR results. (A) AR expression level in cancer tissues and adjacent tissues (B) CXCL8 expression level in cancer tissues and adjacent tissues (C) EGFR expression level in cancer tissues and adjacent tissues (*P\\u0026lt;0.05, **** p\\u0026lt;0.0001)\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"image8.png\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-4460813/v1/901c162d81ea5cea8aaec3f2.png\"},{\"id\":74858515,\"identity\":\"2c5220fe-8595-4eb5-b948-2206569ae6e8\",\"added_by\":\"auto\",\"created_at\":\"2025-01-27 16:10:39\",\"extension\":\"pdf\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"manuscript-pdf\",\"size\":7581213,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"manuscript.pdf\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-4460813/v1/dd20fb3f-7d29-4099-88ab-945832e75cf5.pdf\"},{\"id\":57953366,\"identity\":\"47c2e16c-9811-4ce6-affd-098235a1d44d\",\"added_by\":\"auto\",\"created_at\":\"2024-06-07 23:01:50\",\"extension\":\"zip\",\"order_by\":4,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"supplement\",\"size\":54578634,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"Additionalmaterial.zip\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-4460813/v1/95418cb0cced24df0366c263.zip\"}],\"financialInterests\":\"No competing interests reported.\",\"formattedTitle\":\"The prognostic risk model of ESCA patients was constructed based on intercellular-related genes\",\"fulltext\":[{\"header\":\"1. Introduction\",\"content\":\"\\u003cp\\u003eThe exercise of Esophageal cancer (ESCA) is the seventh most common cancer and the sixth cause of cancer-related death worldwide. Despite recent advances in comprehensive treatment, the prognosis of esophageal cancer remains poor\\u003csup\\u003e\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e\\u003c/sup\\u003e. Radiotherapy is one of the standard treatment options for esophageal cancer, which can effectively kill cancer cells, relieve symptoms, and prolong the survival of patients. However, radiotherapy tolerance often occurs in patients with esophageal cancer during radical surgery, perioperative or palliative radiotherapy, which eventually leads to treatment failure and becomes one of the major obstacles in the clinical treatment of esophageal cancer\\u003csup\\u003e\\u003cspan citationid=\\\"CR2\\\" class=\\\"CitationRef\\\"\\u003e2\\u003c/span\\u003e\\u003c/sup\\u003e. Therefore, there is an urgent need to understand the molecular mechanisms underlying the development of radiotherapy tolerance in esophageal cancer and subsequently to identify the molecular targets of radiotherapy sensitization to develop new therapeutic approaches.\\u003c/p\\u003e \\u003cp\\u003eCell-in-cell (CIC) structure refers to the internalization of one or more living cells into another living cell to form a \\\"bird's eye cell\\\". The CIC structure has been observed in various types of tumors and is associated with poor prognosis, such as breast cancer, lung cancer, and pancreatic cancer\\u003csup\\u003e\\u003cspan citationid=\\\"CR3\\\" class=\\\"CitationRef\\\"\\u003e3\\u003c/span\\u003e,\\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e4\\u003c/span\\u003e\\u003c/sup\\u003e. Evaluation of CIC status is an effective method to assess prognosis. However, few studies have investigated the value of CIC-related genes and their prognostic function in ESCA. The aim of this study is to investigate the prognostic value and potential molecular regulatory mechanism of cell-in-cell-related genes (CIC-related genes) in ESCA.\\u003c/p\\u003e\"},{\"header\":\"2. Materials and Methods\",\"content\":\"\\u003cdiv id=\\\"Sec3\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.1 Data extraction\\u003c/h2\\u003e \\u003cp\\u003eIn this study from the UCSC Xena database (\\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://xenabrowser.net/datapages/\\u003c/span\\u003e\\u003cspan address=\\\"https://xenabrowser.net/datapages/\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e) to download the TCGA esophageal cancer (ESCA) RNAseq - HTSeq - Counts matrix, as a training set, At the same time, the corresponding clinical phenotype data and survival information data were downloaded. The data set contained 173 samples, including 162 disease samples, 11 normal control samples, and 161 disease samples with survival information. Data was downloaded in December 2022.\\u003c/p\\u003e \\u003cp\\u003eFrom the GEO database (\\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://www.ncbi.nlm.nih.gov/geo/\\u003c/span\\u003e\\u003cspan address=\\\"https://www.ncbi.nlm.nih.gov/geo/\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e) to download GSE53624 data sets, the 238 samples the transcriptome data, including 119 disease samples (have corresponding survival information), 119 samples of normal, as a validation set. A total of 101 cell-to-cell related genes were downloaded from the existing literature\\u003csup\\u003e\\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e4\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec4\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.2 Differential gene analysis between disease group and normal group\\u003c/h2\\u003e \\u003cp\\u003eThe R software package \\\"DESeq2\\\" was used to analyze the differentially expressed genes between ESCA samples and normal group samples in the training set. The screening threshold was | log2FoldChange|\\u0026gt;0.5 and pvalue\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.05.Using the online jvenn website (\\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttp://jvenn.toulouse.inra.fr/app/example.html\\u003c/span\\u003e\\u003cspan address=\\\"http://jvenn.toulouse.inra.fr/app/example.html\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e) to obtain the above steps of different genes (DEGs) and related genes between literature downloaded 101 cells take intersection, the difference between related gene expression of the cell, These were denoted as CIC-related DEGs.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec5\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.3. Risk model construction and prognostic gene analysis\\u003c/h2\\u003e \\u003cp\\u003eIn the disease group of the training set, univariate COX and stepwise multivariate COX regression analysis were performed on CIC-related DEGs to obtain prognostic genes. Correlation analysis was performed on prognostic genes, and a risk model was constructed. The patients were grouped according to the median value of risk score, and K-M survival curve analysis was performed on the high and low-risk groups. The ROC curve was drawn by survivalROC to verify the accuracy of the survival curve scoring model, the risk curve and the expression heat map of prognostic genes in the high and low-risk groups were drawn, and the validation set GSE53624 dataset was used to verify the results.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec6\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.4. Construction of nomogram model for prognostic factors\\u003c/h2\\u003e \\u003cp\\u003eThe \\\"RMS\\\" package was used to construct nomograms for risk score, age, sex, TNM stage, and 1-year and 3-year survival rates. The calibration curve, decision curve, and ROC of the nomogram were used to evaluate the predictive ability of the nomogram model.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec7\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.5. Enrichment analysis of high and low-risk groups\\u003c/h2\\u003e \\u003cp\\u003eBased on the above grouping information of high and low-risk groups, DESEq2 software was used to analyze the differentially expressed genes between high and low-risk groups in the training set, and the screening threshold was | log2FoldChange|\\u0026gt;0.5, pvalue\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.05.\\u003c/p\\u003e \\u003cp\\u003eClusterProfiler package was used to perform GO and KEGG enrichment analysis of DEGs between high and low-risk groups, and R package \\\"ggplot2\\\" was used to visualize the enrichment results.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec8\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.6. Analysis of immune infiltration in the high and low-risk groups\\u003c/h2\\u003e \\u003cp\\u003eBased on the grouping information of high and low-risk groups, in the training set, the ssGSEA package was used to analyze the immune infiltration between high and low-risk groups, the composition of immune cells in the training set, and the differentially expressed immune cells between high and low-risk groups were observed.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec9\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.7. Prognostic gene expression validation\\u003c/h2\\u003e \\u003cp\\u003eThe expression matrix of prognostic genes was extracted in the training set and validation set, and t.test was used to check the expression of prognostic genes.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec10\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.8. Molecular DOCKING Analysis\\u003c/h2\\u003e \\u003cp\\u003eThe crystal structures of the prognostic genes were downloaded from the PDB database (\\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://www.rcsb.org/\\u003c/span\\u003e\\u003cspan address=\\\"https://www.rcsb.org/\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e), and the molecular docking of the crystal structures with Gefitinib and osimertinib was performed by AutoDock4.2 software.\\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec11\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.9. RT-PCR\\u003c/h2\\u003e \\u003cp\\u003eFive pairs of frozen esophageal cancer samples, 5 pairs of cancer tissues, and 5 pairs of adjacent tissues were collected from the Department of Thoracic Surgery of Gansu Provincial People's Hospital. 50mg of tissue was taken from each sample for total RNA extraction. The SureScript-First-strand-cDNA-synthesis-kit was used for reverse transcription and detection. The relative RNA quantity was calculated using the relative standard method (2\\u003csup\\u003e\\u0026minus;△△CT\\u003c/sup\\u003e). The primers and their sequences are shown in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e\\u003c/p\\u003e \\u003cp\\u003e \\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab1\\\" border=\\\"1\\\"\\u003e \\u003ccaption language=\\\"En\\\"\\u003e \\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 1\\u003c/div\\u003e \\u003cdiv class=\\\"CaptionContent\\\"\\u003e \\u003cp\\u003ePrimers and their sequences\\u003c/p\\u003e \\u003c/div\\u003e \\u003c/caption\\u003e \\u003ccolgroup cols=\\\"2\\\"\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e \\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e \\u003cthead\\u003e \\u003ctr\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003ePrimers\\u003c/p\\u003e \\u003c/th\\u003e \\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eSequences\\u003c/p\\u003e \\u003c/th\\u003e \\u003c/tr\\u003e \\u003c/thead\\u003e \\u003ctbody\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eAR F\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eAGGCAGTGTCGGTGTCCATG\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eAR R\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eCCTTTGGTGTAACCTCCCTTGA\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eCXCL8 F\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eTCTGCAGCTCTGTGTGAAGG\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eCXCL8 R\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eTTCTCAGCCCTCTTCAAAAACT\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eEGFR F\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eGCCAAGGCACGAGTAACAAGC\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eEGFR R\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eAGGGCAATGAGGACATAACCAG\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eGAPDH F\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eCGAAGGTGGAGTCAACGGATTT\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003ctr\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e \\u003cp\\u003eGAPDH R\\u003c/p\\u003e \\u003c/td\\u003e \\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e \\u003cp\\u003eATGGGTGGAATCATATTGGAAC\\u003c/p\\u003e \\u003c/td\\u003e \\u003c/tr\\u003e \\u003c/tbody\\u003e \\u003c/colgroup\\u003e \\u003c/table\\u003e\\u003c/div\\u003e \\u003c/p\\u003e \\u003c/div\\u003e \\u003cdiv id=\\\"Sec12\\\" class=\\\"Section2\\\"\\u003e \\u003ch2\\u003e2.10 Statistical Analysis\\u003c/h2\\u003e \\u003cp\\u003eStatistical analysis and graphics were performed and obtained using R software (version 4.2.0) and GraphPad Prism (version 8.0). Data were presented as mean\\u0026thinsp;\\u0026plusmn;\\u0026thinsp;stan-dard deviation. Differences were analyzed via the Wilcoxon rank-sum test. The p\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.05 represented a significant difference.\\u003c/p\\u003e \"},{\"header\":\"3. Results\",\"content\":\"\\u003ch2\\u003e3.1 Analysis of differentially expressed genes between disease and normal groups in the training set\\u003c/h2\\u003e\\n\\u003cp\\u003eThe distribution of differential genes is shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003eA-\\u003cspan class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003eB. A total of 12985 differentially expressed genes (DEGs) were obtained, including 6617 up-regulated genes and 6368 down-regulated genes.\\u003c/p\\u003e\\n\\u003cdiv id=\\\"Sec14\\\" class=\\\"Section2\\\"\\u003e\\n\\u003ch2\\u003e3.2 Identification of CIC-related DEGs\\u003c/h2\\u003e\\n\\u003cp\\u003eThe results are shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003eC. A total of 38 intersection genes were obtained, which were as follows: AR, AURKA, CDC20, CDH3, CDKN2A, CEBPB, CTSB, CTSL, CTTN, CXCL8, EGFR, ERI3, EZR, GULP1, GZMB, HTT, KIF2C, KRT7, MAD2L1, MAPT, MLANA, MYLK, NTS, PCDH7, PDPN, PRKAA2, PRNP, RAC1, RNF146, SFTPB, SPRN, STMN2, TF, TGFB1, TNFSF10, TP53, TP63, WT1. These were denoted as CIC-related DEGs.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec15\\\" class=\\\"Section2\\\"\\u003e\\n\\u003ch2\\u003e3.3 Construction of risk models\\u003c/h2\\u003e\\n\\u003cdiv id=\\\"Sec16\\\" class=\\\"Section3\\\"\\u003e\\n\\u003ch2\\u003e3.3.1 Univariate COX regression analysis\\u003c/h2\\u003e\\n\\u003cp\\u003eWe used the TCGA-ESCA dataset as the training set and performed univariate Cox proportional hazards regression analysis using 38 CIC-related DEGs, with the univariate p-value set to 0.2 and a total of 5 genes; the results and forest plots are shown in Table\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e and Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003eA, respectively.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec17\\\" class=\\\"Section3\\\"\\u003e\\n\\u003ch2\\u003e3.3.2 Multivariate COX regression analysis\\u003c/h2\\u003e\\n\\u003cp\\u003eThree prognostic genes (AR, CXCL8, EGFR) were obtained by stepwise multivariate COX analysis using the step function, and the multivariate coefficient (coef) of each prognostic gene was calculated to construct the survival risk model. The multivariate results and forest plots are shown in Table\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e and Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003eB, respectively.\\u003c/p\\u003e\\n\\u003cp\\u003eThen in the ESCA group of the training set, the expression matrix of prognostic genes (AR, CXCL8, EGFR) was extracted, the correlation between prognostic genes was calculated by the \\\"spearman\\\" method, and correlation Chord diagram was plotted using the R software \\\"circlize\\\" package, and the results were shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003eC.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec18\\\" class=\\\"Section3\\\"\\u003e\\n\\u003ch2\\u003e3.3.3 Effect of prognostic genes on survival\\u003c/h2\\u003e\\n\\u003cp\\u003eTo evaluate the prognostic value of the risk model, the expression levels of three prognostic genes were obtained in the TCGA-ESCA dataset, and the formula was used:\\u003c/p\\u003e\\n\\u003cdiv id=\\\"Equa\\\" class=\\\"Equation\\\"\\u003e\\n\\u003cdiv id=\\\"FileID_Equa\\\" class=\\\"mathdisplay\\\"\\u003e$${\\\\mathbf{R}\\\\mathbf{i}\\\\mathbf{s}\\\\mathbf{k}\\\\mathbf{s}\\\\mathbf{c}\\\\mathbf{o}\\\\mathbf{r}\\\\mathbf{e}}_{\\\\mathbf{s}\\\\mathbf{a}\\\\mathbf{m}\\\\mathbf{p}\\\\mathbf{l}\\\\mathbf{e}}={\\\\sum }_{\\\\varvec{n}=1}^{\\\\varvec{n}}(\\\\varvec{c}\\\\varvec{o}\\\\varvec{e}{\\\\varvec{f}}_{\\\\varvec{i}}\\\\ast {\\\\varvec{x}}_{\\\\varvec{i}})$$\\u003c/div\\u003e\\n\\u003c/div\\u003e\\n\\u003cp\\u003e(Note: Coefi represents the multivariate regression coefficient of the ith gene, xi represents the expression value of the ith gene, and n represents the number of model genes.)\\u003c/p\\u003e\\n\\u003cp\\u003eThe risk value of each patient was calculated, and 161 patients with survival information in the training set were divided into high and low-risk groups according to the median value of risk value (1.047788). There were 80 samples in the high-risk group and 81 samples in the low-risk group. The results of the survival analysis are shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003eD. There was a significant survival difference between the high and low-risk groups in the TCGA-ESCA data set (P\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.05).\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec19\\\" class=\\\"Section3\\\"\\u003e\\n\\u003ch2\\u003e3.3.4 ROC curve and risk curve\\u003c/h2\\u003e\\n\\u003cp\\u003eThe multivariate COX model was used to calculate the RiskScore, the survival ROC package was used to calculate the false positive and true positive, the ROC curve was drawn using the results, and the AUC was calculated as shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003eE. The AUC values of 1-year and 3-year were 0.74 and 0.62, respectively, indicating that the risk regression model constructed could be effectively used as a prognostic model. The risk curve is shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003eG, with samples divided into high and low-risk groups according to the median value.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec20\\\" class=\\\"Section3\\\"\\u003e\\n\\u003ch2\\u003e3.3.5 Expression heatmap of prognostic genes in the high and low-risk groups\\u003c/h2\\u003e\\n\\u003cp\\u003eThe R package \\\"pheatmap\\\" was used to map the expression of prognostic genes in the high and low-risk groups, and the results are shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003eF.\\u003c/p\\u003e\\n\\u003cdiv class=\\\"gridtable\\\"\\u003e\\n\\u003cdiv class=\\\"colspec\\\" align=\\\"left\\\"\\u003e\\u0026nbsp;\\u0026nbsp;\\u003c/div\\u003e\\n\\u003ctable id=\\\"Tab2\\\" border=\\\"1\\\"\\u003e\\u003ccaption\\u003e\\n\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 2\\u003c/div\\u003e\\n\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\n\\u003cp\\u003eResults of univariate analysis\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003c/caption\\u003e\\n\\u003cthead\\u003e\\n\\u003ctr\\u003e\\n\\u003cth align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eid\\u003c/p\\u003e\\n\\u003c/th\\u003e\\n\\u003cth align=\\\"left\\\"\\u003e\\n\\u003cp\\u003epvalue\\u003c/p\\u003e\\n\\u003c/th\\u003e\\n\\u003cth align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eHR\\u003c/p\\u003e\\n\\u003c/th\\u003e\\n\\u003cth align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eHR.95L\\u003c/p\\u003e\\n\\u003c/th\\u003e\\n\\u003cth align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eHR.95H\\u003c/p\\u003e\\n\\u003c/th\\u003e\\n\\u003c/tr\\u003e\\n\\u003c/thead\\u003e\\n\\u003ctbody\\u003e\\n\\u003ctr\\u003e\\n\\u003ctd align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eEGFR\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.025465674\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.816222106\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.683036621\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.975377463\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003c/tr\\u003e\\n\\u003ctr\\u003e\\n\\u003ctd align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eAR\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.032093467\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.816018744\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.677565483\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.982763448\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003c/tr\\u003e\\n\\u003ctr\\u003e\\n\\u003ctd align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eCXCL8\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.051096995\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e1.154841387\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.999313898\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e1.334574283\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003c/tr\\u003e\\n\\u003ctr\\u003e\\n\\u003ctd align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eHTT\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.06904872\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.699816164\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.476290936\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e1.028242669\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003c/tr\\u003e\\n\\u003ctr\\u003e\\n\\u003ctd align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eTP53\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.164783007\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.867781691\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.710418302\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e1.060002342\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003c/tr\\u003e\\n\\u003c/tbody\\u003e\\n\\u003c/table\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv class=\\\"gridtable\\\"\\u003e\\n\\u003cdiv class=\\\"colspec\\\" align=\\\"left\\\"\\u003e\\u0026nbsp;\\u003c/div\\u003e\\n\\u003cdiv class=\\\"colspec\\\" align=\\\"char\\\"\\u003e\\u0026nbsp;\\u003c/div\\u003e\\n\\u003ctable id=\\\"Tab3\\\" border=\\\"1\\\"\\u003e\\u003ccaption\\u003e\\n\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 3\\u003c/div\\u003e\\n\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\n\\u003cp\\u003eResults of multivariate analysis\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003c/caption\\u003e\\n\\u003cthead\\u003e\\n\\u003ctr\\u003e\\n\\u003cth align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eid\\u003c/p\\u003e\\n\\u003c/th\\u003e\\n\\u003cth align=\\\"left\\\"\\u003e\\n\\u003cp\\u003ecoef\\u003c/p\\u003e\\n\\u003c/th\\u003e\\n\\u003cth align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eHR\\u003c/p\\u003e\\n\\u003c/th\\u003e\\n\\u003cth align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eHR.95L\\u003c/p\\u003e\\n\\u003c/th\\u003e\\n\\u003cth align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eHR.95H\\u003c/p\\u003e\\n\\u003c/th\\u003e\\n\\u003cth align=\\\"left\\\"\\u003e\\n\\u003cp\\u003epvalue\\u003c/p\\u003e\\n\\u003c/th\\u003e\\n\\u003c/tr\\u003e\\n\\u003c/thead\\u003e\\n\\u003ctbody\\u003e\\n\\u003ctr\\u003e\\n\\u003ctd align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eAR\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e-0.178187505\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.836785507\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.699135927\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e1.001536265\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.051989207\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003c/tr\\u003e\\n\\u003ctr\\u003e\\n\\u003ctd align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eCXCL8\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.137065624\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e1.14690341\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.985768463\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e1.334377678\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.075996366\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003c/tr\\u003e\\n\\u003ctr\\u003e\\n\\u003ctd align=\\\"left\\\"\\u003e\\n\\u003cp\\u003eEGFR\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e-0.199620941\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.819041159\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.688642196\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.97413203\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003ctd align=\\\"char\\\" char=\\\".\\\"\\u003e\\n\\u003cp\\u003e0.024059355\\u003c/p\\u003e\\n\\u003c/td\\u003e\\n\\u003c/tr\\u003e\\n\\u003c/tbody\\u003e\\n\\u003c/table\\u003e\\n\\u003c/div\\u003e\\n\\u003cp\\u003e\\u0026nbsp;\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec21\\\" class=\\\"Section2\\\"\\u003e\\n\\u003ch2\\u003e3.4 Survival risk model validation\\u003c/h2\\u003e\\n\\u003cdiv id=\\\"Sec22\\\" class=\\\"Section3\\\"\\u003e\\n\\u003ch2\\u003e3.4.1 Effect of prognostic genes on survival\\u003c/h2\\u003e\\n\\u003cp\\u003eTo verify the applicability of the risk model, we obtained the expression levels of three prognostic genes in the GSE53624 dataset and used the formula:\\u003c/p\\u003e\\n\\u003cdiv id=\\\"Equb\\\" class=\\\"Equation\\\"\\u003e\\n\\u003cdiv id=\\\"FileID_Equb\\\" class=\\\"mathdisplay\\\"\\u003e$${\\\\mathbf{R}\\\\mathbf{i}\\\\mathbf{s}\\\\mathbf{k}\\\\mathbf{s}\\\\mathbf{c}\\\\mathbf{o}\\\\mathbf{r}\\\\mathbf{e}}_{\\\\mathbf{s}\\\\mathbf{a}\\\\mathbf{m}\\\\mathbf{p}\\\\mathbf{l}\\\\mathbf{e}}={\\\\sum }_{\\\\varvec{n}=1}^{\\\\varvec{n}}(\\\\varvec{c}\\\\varvec{o}\\\\varvec{e}{\\\\varvec{f}}_{\\\\varvec{i}}\\\\ast {\\\\varvec{x}}_{\\\\varvec{i}})$$\\u003c/div\\u003e\\n\\u003c/div\\u003e\\n\\u003cp\\u003e(Note: Coefi represents the multivariate regression coefficient of the ith gene, xi represents the expression value of the ith gene, and n represents the number of model genes.)\\u003c/p\\u003e\\n\\u003cp\\u003eThe risk value of each patient was calculated, and 119 patients in the training set were divided into high and low-risk groups according to the median value of risk value (0.9625565), including 59 samples in the high-risk group and 60 samples in the low-risk group. The results of the survival analysis of the high and low-risk groups are shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003eA. It can be seen that there is a significant survival difference between the high and low-risk groups in the GSE53624 data set (p\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.05).\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec23\\\" class=\\\"Section3\\\"\\u003e\\n\\u003ch2\\u003e3.4.2 ROC curve and risk curve\\u003c/h2\\u003e\\n\\u003cp\\u003eThe multivariate COX model was used to calculate the RiskScore, the survival ROC package was used to calculate the false positive and true positive, and the results were used to draw the ROC curve, and the AUC was calculated as shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003eB. The AUC values of 1-year and 3-year were both above 0.6, indicating that the risk regression model constructed could be effectively used as a prognostic model. As shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003eC, samples were divided into high and low-risk groups according to the median value.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec24\\\" class=\\\"Section3\\\"\\u003e\\n\\u003ch2\\u003e3.4.3 Expression heatmap of prognostic genes in the high and low risk groups\\u003c/h2\\u003e\\n\\u003cp\\u003eThe R package \\\"pheatmap\\\" was used to map the expression of prognostic genes in the high and low-risk groups, and the results are shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003eD.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec25\\\" class=\\\"Section2\\\"\\u003e\\n\\u003ch2\\u003e3.5 Construction of nomogram model for prognostic factors\\u003c/h2\\u003e\\n\\u003cp\\u003eBased on risk score, age, sex, and TNM stage, the \\\"RMS\\\" package was used to construct survival nomograms at 1 and 3 years for several clinical factors in the risk model. Each factor corresponds to a score. The sum of the total scores of each factor corresponds to the total score. The 1-year and 3-year survival rates were predicted according to the total score. The higher the score, the lower the survival rate. The nomogram (Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e4\\u003c/span\\u003eA) and calibration curve (Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e4\\u003c/span\\u003eB) were constructed. The 1 - and 3-year ROC curves of the nomogram were drawn using the \\\"pROC\\\" package (Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e4\\u003c/span\\u003eC), and the decision curve (DCA) was drawn using the \\\"rmda\\\" package (Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e4\\u003c/span\\u003eD). The calibration curve showed that the error between the actual risk of disease and the predicted risk was small, indicating that the nomogram model had high prediction accuracy for ESCA disease.\\u003c/p\\u003e\\n\\u003cp\\u003eThe AUC areas corresponding to the 1-year and 3-year ROC curves of the nomogram were both greater than 0.7, indicating that the nomogram model had high prediction accuracy for prognosis. Decision curve analysis (DCA) showed that within the high-risk threshold range of 0\\u0026ndash;1, the nomogram model could benefit, and the clinical benefit of the nomogram model was higher than that of the \\\"age\\\" curve, \\\"riskScore\\\" curve, \\\"T\\\" curve, \\\"N\\\" curve, \\\"M\\\" curve and \\\"gender\\\" curve.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec26\\\" class=\\\"Section2\\\"\\u003e\\n\\u003ch2\\u003e3.6 Enrichment analysis of high and low-risk groups\\u003c/h2\\u003e\\n\\u003cp\\u003eBased on the expression matrix in the TCGA-ESCA dataset, 80 samples from the high-risk group and 81 samples from the low-risk group were extracted (TCGA-ESCA), and the \\\"DESeq2\\\" package was used for differential gene analysis. The screening conditions for differential genes were | log2FoldChange|\\u0026gt;0.5. p-value\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.05. There were 5727 significantly differentially expressed genes in high-vs-low, including 2965 up-regulated genes and 2762 down-regulated genes. As shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e5\\u003c/span\\u003eA, the volcano plot is used to show the distribution of differential genes, and as shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e5\\u003c/span\\u003eB, the heat map is used to show the expression of differential genes among sample groups.\\u003c/p\\u003e\\n\\u003cp\\u003eClusterProfiler package in R was used to perform (GO and KEGG) functional enrichment analysis of differential genes in high and low-risk groups, and the screening threshold was as follows: p-value\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.05, a total of 887 GO BP, 67 GO CC, 178 GO MF and 53 KEGG pathways were screened. The enrichment results were visualized by the R package \\\"ggplot2\\\", as shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e5\\u003c/span\\u003eC-\\u003cspan class=\\\"InternalRef\\\"\\u003e5\\u003c/span\\u003eD. We ranked according to p-value value. Each TOP10 Descriptions of GO and KEGG were selected for display.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec27\\\" class=\\\"Section2\\\"\\u003e\\n\\u003ch2\\u003e3.7 Analysis of immune infiltration in the high and low risk groups\\u003c/h2\\u003e\\n\\u003cp\\u003eBased on the expression matrix and grouping information of high and low-risk groups in the training set (TCGA-ESCA), the expression levels of 29 immune gene sets (16 immune cells and 13 immune-related pathways gene sets) in each sample of the training set were drawn by ssGSEA (single sample GSEA) analysis. And using t.test() to test, it was found that there were 8 significantly different immune-related gene sets in the training set, which were as follows: B_cells, iDCs, Macrophages, Mast_cells, NK_cells, T_cell_co\\u0026thinsp;\\u0026minus;\\u0026thinsp;inhibition, T_helper_cells, Type_II_IFN_Reponse, the results are shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e6\\u003c/span\\u003eA.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec28\\\" class=\\\"Section2\\\"\\u003e\\n\\u003ch2\\u003e3.8 Prognostic gene expression validation\\u003c/h2\\u003e\\n\\u003cp\\u003eThe expression matrices of prognostic genes (AR, CXCL8, EGFR) were extracted from the training set (TCGA-ESCA) and the validation set GSE53624, respectively. Based on the grouping information in the respective data sets, the t.test() test was used. The R software package \\\"ggplot2\\\" was used to draw the boxplot of prognostic gene expression in the two datasets, and it was found that the three prognostic genes reached significant differences between the disease group and the normal group in the two datasets, and the expression trend was consistent. The results are shown in Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e6\\u003c/span\\u003eB and Fig.\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e6\\u003c/span\\u003eC.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec29\\\" class=\\\"Section2\\\"\\u003e\\n\\u003ch2\\u003e3.9 Molecular DOCKING Analysis\\u003c/h2\\u003e\\n\\u003cp\\u003eDocking analysis was performed using three prognostic genes (AR, CXCL8, EGFR) and two drugs (Gefitinib and osimertinib). Firstly, the corresponding crystal structures of AR, CXCL8, and EGFR were downloaded from the PDB database (\\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://www.rcsb.org/\\u003c/span\\u003e\\u003c/span\\u003e), respectively. The corresponding relationship between genes and PDB ID was AR-1T7R, CXCL8-7JNY, and EGFR-1IP0.\\u003c/p\\u003e\\n\\u003cp\\u003eMolecular docking using AutoDock revealed two hydrogen bonds and two residues in the 1t7r and Gefitinib docking. There was one hydrogen bond and one residue in the docking between 1t7r and osimertinib. There were three hydrogen bonds and two residues in the docking between 7jny and Gefitinib. There are three hydrogen bonds and two residues in the docking of 7jny and osimertinib.\\u003c/p\\u003e\\n\\u003cp\\u003eThere were two hydrogen bonds and two residues in the docking between 1ip0 and Gefitinib. There were 4 hydrogen bonds and 4 residues in the docking of 1ip0 and osimertinib, and all three prognostic gene pairs interacted with the two drugs. Figure\\u0026nbsp;\\u003cspan class=\\\"InternalRef\\\"\\u003e7\\u003c/span\\u003eA-\\u003cspan class=\\\"InternalRef\\\"\\u003e7\\u003c/span\\u003eB shows the docking results of 1ip0 with Gefitinib and osimertinib, and other molecular docking results are detailed in the Supplementary material.\\u003c/p\\u003e\\n\\u003c/div\\u003e\\n\\u003cdiv id=\\\"Sec30\\\" class=\\\"Section2\\\"\\u003e\\n\\u003ch2\\u003e3.10 RT-PCR\\u003c/h2\\u003e\\n\\u003cp\\u003eFive pairs of frozen esophageal cancer samples were obtained from the Department of Thoracic Surgery, Gansu Provincial People's Hospital for RT-pcr experiments. Results As shown in Fig.\\u0026nbsp;9, AR was lowly expressed in esophageal cancer tissues (p\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.0001), while CXCL8 and EGFR were highly expressed in esophageal cancer tissues (p\\u0026thinsp;\\u0026lt;\\u0026thinsp;0.05). The RT-pcr results were consistent with the previous analysis.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u0026nbsp;\\u003c/p\\u003e\\n\\u003c/div\\u003e\"},{\"header\":\"4. Discussion\",\"content\":\"\\u003cp\\u003eEsophageal cancer is a malignant lesion formed by abnormal proliferation of esophageal squamous epithelium or glandular epithelium, which is one of the most common malignant tumors in the world\\u003csup\\u003e\\u003cspan citationid=\\\"CR5\\\" class=\\\"CitationRef\\\"\\u003e5\\u003c/span\\u003e\\u003c/sup\\u003e.It ranks sixth in cancer mortality\\u003csup\\u003e\\u003cspan citationid=\\\"CR6\\\" class=\\\"CitationRef\\\"\\u003e6\\u003c/span\\u003e\\u003c/sup\\u003e.The early symptoms of esophageal cancer are often not obvious, so the patients have progressed to the middle and late stages once diagnosed, and 38% of the cases are diagnosed in the late stage, missing the best opportunity for treatment. Currently, the 5-year survival rate for patients with esophageal cancer is 10%\\u003csup\\u003e7\\u003c/sup\\u003e. Therefore, it is particularly important to understand the mechanism of the occurrence and development of esophageal cancer and to develop new treatment methods. During tumor cell culture, some large cells with internal vacuoles and multiple nuclei are found, that is, the situation of one cell inside another, and are called cell-in-cell are observed in a variety of tumors and have been shown to correlate with prognosis. Therefore, it is particularly important to study the prognostic value and regulatory mechanism of intracellular structure-related genes in esophageal cancer\\u003csup\\u003e\\u003cspan citationid=\\\"CR8\\\" class=\\\"CitationRef\\\"\\u003e8\\u003c/span\\u003e\\u003c/sup\\u003e. CIC structure has been observed in various tumor types and is associated with a poor prognosis, with one clinical study showing an association of CIC structure with aggressive features and cancer-related mortality in oral tongue cancer\\u003csup\\u003e\\u003cspan citationid=\\\"CR9\\\" class=\\\"CitationRef\\\"\\u003e9\\u003c/span\\u003e\\u003c/sup\\u003e. However, the role of CIC structure in esophageal cancer remains unclear.\\u003c/p\\u003e \\u003cp\\u003eWe used the GDC TCGA Esophageal Cancer (ESCA) dataset as the training set for this study. The GSE53624 dataset was downloaded from the GEO database and used as the validation set. A total of 101 cell-related genes were downloaded from the existing literature\\u003csup\\u003e\\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e4\\u003c/span\\u003e\\u003c/sup\\u003e. Based on the above data, a series of bioinformatics analyses were carried out in this study.\\u003c/p\\u003e \\u003cp\\u003eFirst, we performed differential gene analysis between the disease group and the normal group, and took the intersection of the obtained differential genes and 101 cell-related genes, and a total of 38 intersection genes were obtained, which were named CIC-related DEGs. The following 38 genes were analyzed by univariate and multi-step multivariate analysis, and 3 prognostic genes were obtained, namely AR, CXCL8 and EGFR.\\u003c/p\\u003e \\u003cp\\u003eWe found that CXCL8 and EGFR were up-regulated in esophageal cancer. CXCL8, also known as interleukin-8 (IL-8), is a neutrophil-activating factor, which is mainly produced by neutrophils, monocytes, macrophages, T cells, epithelial cells, and endothelial cells. CXCL8 has the same expression trend in ovarian cancer, breast cancer, liver cancer\\u003csup\\u003e\\u003cspan additionalcitationids=\\\"CR11\\\" citationid=\\\"CR10\\\" class=\\\"CitationRef\\\"\\u003e10\\u003c/span\\u003e\\u0026ndash;\\u003cspan citationid=\\\"CR12\\\" class=\\\"CitationRef\\\"\\u003e12\\u003c/span\\u003e\\u003c/sup\\u003e, head and neck squamous cell carcinoma, and has prognostic and diagnostic value\\u003csup\\u003e\\u003cspan citationid=\\\"CR13\\\" class=\\\"CitationRef\\\"\\u003e13\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e \\u003cp\\u003eSilencing CXCL8 in esophageal cancer cells can inhibit the proliferation and invasion of esophageal cancer cells\\u003csup\\u003e\\u003cspan citationid=\\\"CR14\\\" class=\\\"CitationRef\\\"\\u003e14\\u003c/span\\u003e\\u003c/sup\\u003e. Helen et al. found that CXCL8 gene expression is closely related to epithelial-mesenchymal transition and cell-matrix vascularization, thereby promoting tumor cell invasion and metastasis, which explains to a certain extent the ability of CXCL8 gene to predict tumor metastasis\\u003csup\\u003e\\u003cspan citationid=\\\"CR15\\\" class=\\\"CitationRef\\\"\\u003e15\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e \\u003cp\\u003eEGFR mainly exists in the stroma, epidermis, and some smooth muscle cells. After activation, EGFR can activate the downstream mitogen-activated protein kinase (MAPK) signaling pathway, inhibit cell apoptosis and DNA repair, accelerate cell invasion\\u003csup\\u003e\\u003cspan citationid=\\\"CR16\\\" class=\\\"CitationRef\\\"\\u003e16\\u003c/span\\u003e\\u003c/sup\\u003e, and participate in the occurrence and development of lung cancer, colorectal cancer, breast cancer, and other malignant tumors\\u003csup\\u003e\\u003cspan additionalcitationids=\\\"CR18\\\" citationid=\\\"CR17\\\" class=\\\"CitationRef\\\"\\u003e17\\u003c/span\\u003e\\u0026ndash;\\u003cspan citationid=\\\"CR19\\\" class=\\\"CitationRef\\\"\\u003e19\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e \\u003cp\\u003eSubsequently, 29 immune-related gene sets were selected in this study. Through immune infiltration analysis of high and low-risk groups, 8 immune-related gene sets with significant differences were found in the training set, which were as follows: B_cells, iDCs, Macrophages, Mast_cells, NK_cells, T_cell_co\\u0026thinsp;\\u0026minus;\\u0026thinsp;inhibition, T_helper_cells, Type_II_IFN_Reponse, Then, the expression levels of the three prognostic genes in the training set and validation set were analyzed, and the expression trends of the three prognostic genes in the training set and validation set were consistent.\\u003c/p\\u003e \\u003cp\\u003eIn recent years, molecular docking methods have become an important technique in the field of computer-aided drug research. Molecular docking is a method for drug design through the characteristics of receptors and the way of interaction between receptors and drug molecules. A theoretical simulation method that mainly studies the interaction between molecules (such as ligands and receptors) and predicts their binding modes and affinities\\u003csup\\u003e\\u003cspan citationid=\\\"CR20\\\" class=\\\"CitationRef\\\"\\u003e20\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e \\u003cp\\u003eAutodock is an open-source molecular simulation software, which is mainly used to perform ligand-protein molecular docking. AutoDock uses a semi-flexible docking method to allow conformational changes of small molecules, and the binding free energy is used as the basis for the evaluation of docking results\\u003csup\\u003e\\u003cspan citationid=\\\"CR21\\\" class=\\\"CitationRef\\\"\\u003e21\\u003c/span\\u003e\\u003c/sup\\u003e. We performed molecular docking of the three prognostic genes with Gefitinib and osimertinib, and found that there were interactions between them, indicating that the prognostic genes may affect the efficacy of drugs or the drugs may affect the expression of prognostic genes.\\u003c/p\\u003e \\u003cp\\u003eOsimertinib belongs to the third generation of EG-FR-TKIs class and has been used in the treatment of lung cancer\\u003csup\\u003e\\u003cspan citationid=\\\"CR22\\\" class=\\\"CitationRef\\\"\\u003e22\\u003c/span\\u003e\\u003c/sup\\u003e. Some scholars have found that the use of osimertinib can lead to severe complications including lung injury, especially interstitial lung disease (ILD). Due to impaired epithelial healing of lung injury, pulmonary inflammation may worsen after EGFR-TKI treatment\\u003csup\\u003e\\u003cspan citationid=\\\"CR23\\\" class=\\\"CitationRef\\\"\\u003e23\\u003c/span\\u003e\\u003c/sup\\u003e. Gefitinib is a specific small molecule EGFR tyrosine kinase inhibitor, which can inhibit the growth of esophageal cancer cells by inhibiting the signal transduction pathway of EGFR tyrosine kinase and inhibiting the phosphorylation of EGFR. It also has the effect of inhibiting tumor angiogenesis and metastasis and plays the purpose of treating tumors together\\u003csup\\u003e\\u003cspan citationid=\\\"CR24\\\" class=\\\"CitationRef\\\"\\u003e24\\u003c/span\\u003e\\u003c/sup\\u003e. It also has the effect of inhibiting tumor angiogenesis and metastasis and plays the purpose of treating tumors together\\u003csup\\u003e\\u003cspan citationid=\\\"CR24\\\" class=\\\"CitationRef\\\"\\u003e24\\u003c/span\\u003e\\u003c/sup\\u003e. The inhibitory effect of gefitinib on metastatic non-small cell lung cancer, ovarian cancer, and other tumors has been widely recognized\\u003csup\\u003e\\u003cspan citationid=\\\"CR25\\\" class=\\\"CitationRef\\\"\\u003e25\\u003c/span\\u003e\\u003c/sup\\u003e. Our study found that Gefitinib and osimertinib not only interacted with EGFR but also interacted with AR and CXCL8, which may have implications for the development of targeted drugs in the future.\\u003c/p\\u003e\"},{\"header\":\"5. Conclusion\",\"content\":\"\\u003cp\\u003eIn this study, the TCGA-ESCA transcriptome data were analyzed by bioinformatics methods, and three prognostic genes were obtained, namely AR, CXCL8, and EGFR, which were verified by RT-PCR. Moreover, molecular docking of prognostic genes with Gefitinib and osimertinib indicated that there were interactions between them. It is suggested that more attention and in-depth study are needed in the research of prognostic genes and drugs to better understand the pathogenesis of the disease and improve the therapeutic effect. Therefore, this study provides an important reference for the diagnosis, mechanism research, and treatment of esophageal cancer in the future. However, the influence of the \\\"intracellular cell\\\" structure on the biological behavior of esophageal cancer cells and the influence on the body's immunity still needs to be confirmed by in vivo studies, and the molecular docking results still need to be further confirmed, and verified by subsequent related experiments.\\u003c/p\\u003e\"},{\"header\":\"Declarations\",\"content\":\"\\u003cp\\u003eFunding\\u003c/p\\u003e\\n\\u003cp\\u003eThis study was supported by grants from the Gansu Provincial People\\u0026apos;s Hospital\\u0026nbsp;\\u003c/p\\u003e\\n\\u003cp\\u003eResearch Funding (22GSSYD-30; 22GSSYD-25, 22GSSYC-9), Gansu Youth Science and Technology Fund (22JR11RA241)，\\u0026nbsp;and Science and Technology Department of Gansu Province (22YF7FA095)\\u003c/p\\u003e\\n\\u003cp\\u003eDeclaration of competing interest\\u0026nbsp;\\u003c/p\\u003e\\n\\u003cp\\u003eThe authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.\\u0026nbsp;\\u003c/p\\u003e\\n\\u003cp\\u003eData Availability Statements\\u003c/p\\u003e\\n\\u003cp\\u003eThe data that supports the findings of this study are available in the supplementary material of this article.\\u003c/p\\u003e\\n\\u003cp\\u003eEthics approval statement\\u003c/p\\u003e\\n\\u003cp\\u003eThis topic project has passed the Gansu Province People\\u0026apos;s Hospital Ethics Committee for examination and approval.\\u003c/p\\u003e\\n\\u003cp\\u003eAuthor contribution statement\\u003c/p\\u003e\\n\\u003cp\\u003eWei Cao: Conceptualization (Equal), Data curation (Equal), Funding acquisition (Equal), Investigation (Equal), Project administration (Equal), Resources (Equal), Writing\\u0026nbsp;\\u0026ndash;\\u0026nbsp;original draft (Equal)\\u003c/p\\u003e\\n\\u003cp\\u003eDacheng Jin: Conceptualization (Equal), Data curation (Equal), Funding acquisition (Equal), Methodology (Equal), Project administration (Equal), Resources (Equal), Writing\\u0026nbsp;\\u0026ndash;\\u0026nbsp;original draft (Equal)\\u0026nbsp;\\u003c/p\\u003e\\n\\u003cp\\u003eWeiRun Min: Data curation (Supporting), Formal analysis (Equal), Project administration (Supporting), Resources (Supporting), Software (Supporting), Supervision (Supporting), Writing\\u0026nbsp;\\u0026ndash;\\u0026nbsp;review \\u0026amp; editing (Equal)\\u003c/p\\u003e\\n\\u003cp\\u003eHaochi Li: Conceptualization (Supporting), Funding acquisition (Equal), Methodology (Supporting), Project administration (Supporting), Resources (Supporting), Supervision (Supporting), Writing\\u0026nbsp;\\u0026ndash;\\u0026nbsp;review \\u0026amp; editing (Supporting)\\u003c/p\\u003e\\n\\u003cp\\u003eRong Wang: Formal analysis (Supporting), Investigation (Supporting), Methodology (Supporting), Project administration (Supporting), Supervision (Supporting), Visualization (Equal)\\u003c/p\\u003e\\n\\u003cp\\u003eJinlong Zhang: Data curation (Lead), Formal analysis (Supporting), Software (Supporting), Supervision (Supporting), Writing\\u0026nbsp;\\u0026ndash;\\u0026nbsp;review \\u0026amp; editing (Supporting)\\u003c/p\\u003e\\n\\u003cp\\u003eYunJiu Gou (Corresponding Author): Data curation (Lead), Formal analysis (Lead), Funding acquisition (Lead), Methodology (Lead), Project administration (Lead), Software (Lead), Supervision (Lead), Writing\\u0026nbsp;\\u0026ndash;\\u0026nbsp;review \\u0026amp; editing (Lead)\\u003c/p\\u003e\\n\\u003cp\\u003e*Author Wei Cao and Dacheng Jin share first authorship.\\u003c/p\\u003e\"},{\"header\":\"References\",\"content\":\"\\u003col\\u003e\\u003cli\\u003e\\u003cspan\\u003eWang M, Sun X, Xin H, Wen Z, Cheng Y. SPP1 promotes radiation resistance through JAK2/STAT3 pathway in esophageal carcinoma. Cancer Med. 2022;11:4526\\u0026ndash;43.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eShitara K, et al. Nivolumab plus chemotherapy or ipilimumab in gastro-oesophageal cancer. Nature. 2022;603:942\\u0026ndash;8.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eSong J, et al. Cell-in-Cell-Mediated Entosis Reveals a Progressive Mechanism in Pancreatic Cancer. Gastroenterology. 2023;165:1505\\u0026ndash;e152120.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eSong J, et al. Construction of a novel model based on cell-in-cell-related genes and validation of KRT7 as a biomarker for predicting survival and immune microenvironment in pancreatic cancer. BMC Cancer. 2022;22:894.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eRogers JE, Sewastjanow-Silva M, Waters RE, Ajani JA. Esophageal cancer: emerging therapeutics. Expert Opin Ther Targets. 2022;26:107\\u0026ndash;17.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eSiegel RL, Miller KD, Fuchs HE, Jemal A. Cancer statistics, 2022. CA Cancer J Clin. 2022;72:7\\u0026ndash;33.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eHuang F-L, Yu S-J. Esophageal cancer: Risk factors, genetic association, and treatment. Asian J Surg. 2018;41:210\\u0026ndash;5.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eFais S, Overholtzer M. Cell-in-cell phenomena in cancer. Nat Rev Cancer. 2018;18:758\\u0026ndash;66.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eSiquara da Rocha L, de Souza O, de Lambert BS. Gurgel Rocha, C. de A. Cell-in-Cell Events in Oral Squamous Cell Carcinoma. Front Oncol. 2022;12:931092.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eFu X, Wang Q, Du H, Hao H. CXCL8 and the peritoneal metastasis of ovarian and gastric cancer. Front Immunol. 2023;14:1159061.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eMishra A, Suman KH, Nair N, Majeed J, Tripathi V. An updated review on the role of the CXCL8-CXCR1/2 axis in the progression and metastasis of breast cancer. Mol Biol Rep. 2021;48:6551\\u0026ndash;61.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eYang S, Wang H, Qin C, Sun H, Han Y. Up-regulation of CXCL8 expression is associated with a poor prognosis and enhances tumor cell malignant behaviors in liver cancer. Biosci Rep. 2020;40:BSR20201169.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eLi Y, et al. Analysis of the Prognosis and Therapeutic Value of the CXC Chemokine Family in Head and Neck Squamous Cell Carcinoma. Front Oncol. 2020;10:570736.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eYue D, et al. NEDD9 promotes cancer stemness by recruiting myeloid-derived suppressor cells via CXCL8 in esophageal squamous cell carcinoma. Cancer Biol Med. 2021;18:705\\u0026ndash;20.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eHa H, Debnath B, Neamati N. Role of the CXCL8-CXCR1/2 Axis in Cancer and Inflammatory Diseases. Theranostics. 2017;7:1543\\u0026ndash;88.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eLevantini E, Maroni G, Del Re M, Tenen DG. EGFR signaling pathway as therapeutic target in human cancers. Semin Cancer Biol. 2022;85:253\\u0026ndash;75.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eOxnard GR, et al. Germline EGFR Mutations and Familial Lung Cancer. J Clin Oncol. 2023;41:5274\\u0026ndash;84.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eCheng W-L, et al. The Role of EREG/EGFR Pathway in Tumor Progression. Int J Mol Sci. 2021;22:12828.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eM N, V, M., F, M., P P. Crosstalk between CXCR4/ACKR3 and EGFR Signaling in Breast Cancer Cells. Int J Mol Sci 23, (2022).\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eCrampon K, Giorkallos A, Deldossi M, Baud S, Steffenel LA. Machine-learning methods for ligand-protein molecular docking. Drug Discov Today. 2022;27:151\\u0026ndash;64.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eEberhardt J, Santos-Martins D, Tillack AF, Forli S. AutoDock Vina 1.2.0: New Docking Methods, Expanded Force Field, and Python Bindings. J Chem Inf Model. 2021;61:3891\\u0026ndash;8.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eCheng Y, et al. Osimertinib Versus Comparator EGFR TKI as First-Line Treatment for EGFR-Mutated Advanced NSCLC: FLAURA China, A Randomized Study. Target Oncol. 2021;16:165\\u0026ndash;76.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eOhmori T, et al. Molecular and Clinical Features of EGFR-TKI-Associated Lung Injury. Int J Mol Sci. 2021;22:792.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eNoronha V, et al. Gefitinib Versus Gefitinib Plus Pemetrexed and Carboplatin Chemotherapy in EGFR-Mutated Lung Cancer. J Clin Oncol. 2020;38:124\\u0026ndash;36.\\u003c/span\\u003e\\u003c/li\\u003e \\u003cli\\u003e\\u003cspan\\u003eMurphy M, Stordal B. Erlotinib or gefitinib for the treatment of relapsed platinum pretreated non-small cell lung cancer and ovarian cancer: a systematic review. Drug Resist Updat. 2011;14:177\\u0026ndash;90.\\u003c/span\\u003e\\u003c/li\\u003e\\u003c/ol\\u003e\"}],\"fulltextSource\":\"\",\"fullText\":\"\",\"funders\":[],\"hasAdminPriorityOnWorkflow\":false,\"hasManuscriptDocX\":true,\"hasOptedInToPreprint\":true,\"hasPassedJournalQc\":\"\",\"hasAnyPriority\":false,\"hideJournal\":false,\"highlight\":\"\",\"institution\":\"\",\"isAcceptedByJournal\":true,\"isAuthorSuppliedPdf\":false,\"isDeskRejected\":\"\",\"isHiddenFromSearch\":false,\"isInQc\":false,\"isInWorkflow\":false,\"isPdf\":false,\"isPdfUpToDate\":true,\"isWithdrawnOrRetracted\":false,\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"bmc-cancer\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":false,\"externalIdentity\":\"bcan\",\"sideBox\":\"Learn more about [BMC Cancer](http://bmccancer.biomedcentral.com/)\",\"snPcode\":\"\",\"submissionUrl\":\"https://www.editorialmanager.com/bcan/default.aspx\",\"title\":\"BMC Cancer\",\"twitterHandle\":\"BMC_series\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"em\",\"reportingPortfolio\":\"BMC Series\",\"inReviewEnabled\":true,\"inReviewRevisionsEnabled\":true},\"keywords\":\"esophageal cancer, bioinformatics, Cell-in-cell, molecular docking, prognostic risk model\",\"lastPublishedDoi\":\"10.21203/rs.3.rs-4460813/v1\",\"lastPublishedDoiUrl\":\"https://doi.org/10.21203/rs.3.rs-4460813/v1\",\"license\":{\"name\":\"CC BY 4.0\",\"url\":\"https://creativecommons.org/licenses/by/4.0/\"},\"manuscriptAbstract\":\"\\u003ch2\\u003eBackground\\u003c/h2\\u003e \\u003cp\\u003eEsophageal cancer is a serious malignant tumor disease. Radiotherapy is the standard treatment, but treatment tolerance often leads to failure. Cell-in-cell are observed in a variety of tumors and have been shown to correlate with prognosis. Therefore, it is particularly important to study the prognostic value and regulatory mechanism of intracellular structure-related genes in esophageal cancer.\\u003c/p\\u003e\\u003ch2\\u003eMethods\\u003c/h2\\u003e \\u003cp\\u003eTCGA Esophageal Cancer (ESCA) was included in the analysis as the training set. The differentially expressed genes in ESCA samples in the training set were analyzed, and the differentially expressed intercellular-related genes were recorded as CIC-related DEGs. Cox analysis was used to screen prognostic genes. Samples were divided into high-low-risk groups according to the median value of the ESCA sample risk score. Validation was performed in the risk model GSE53624. Morphological mapping, enrichment analysis, immune infiltration analysis, prognostic gene expression verification, molecular docking, and RT-PCR verification were established.\\u003c/p\\u003e\\u003ch2\\u003eResults\\u003c/h2\\u003e \\u003cp\\u003eA total of 38 intersection genes were obtained between the disease group and the normal group of ESCA samples. After stepwise multivariate COX analysis, three prognostic genes (AR, CXCL8, EGFR) were selected. The applicability of the risk model was verified in the GSE53624 dataset. The analysis revealed eight significantly different immune-related gene sets. The prognostic gene expression validation found that the prognostic genes reached significant differences between the disease group and the normal group in both datasets. The corresponding proteins of the three prognostic genes all interacted with Gefitinib and osimertinib. The results of PCR confirmed the differential expression of prognostic genes in esophageal cancer tissues.\\u003c/p\\u003e\\u003ch2\\u003eConclusions\\u003c/h2\\u003e \\u003cp\\u003eThree prognostic genes, AR, CXCL8, and EGFR, were obtained in this study, and the molecular docking of prognostic genes with Gefitinib and osimertinib showed that there were interactions between them, which provided a basis for the diagnosis and treatment of ESCA.\\u003c/p\\u003e\",\"manuscriptTitle\":\"The prognostic risk model of ESCA patients was constructed based on intercellular-related genes\",\"msid\":\"\",\"msnumber\":\"\",\"nonDraftVersions\":[{\"code\":1,\"date\":\"2024-06-07 23:01:44\",\"doi\":\"10.21203/rs.3.rs-4460813/v1\",\"editorialEvents\":[{\"type\":\"communityComments\",\"content\":0},{\"type\":\"decision\",\"content\":\"Revision requested\",\"date\":\"2024-05-29T08:10:01+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"checksComplete\",\"content\":\"\",\"date\":\"2024-05-23T14:05:41+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"editorAssigned\",\"content\":\"\",\"date\":\"2024-05-23T14:05:41+00:00\",\"index\":\"\",\"fulltext\":\"\"},{\"type\":\"submitted\",\"content\":\"BMC Cancer\",\"date\":\"2024-05-22T12:14:03+00:00\",\"index\":\"\",\"fulltext\":\"\"}],\"status\":\"published\",\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"bmc-cancer\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":false,\"externalIdentity\":\"bcan\",\"sideBox\":\"Learn more about [BMC Cancer](http://bmccancer.biomedcentral.com/)\",\"snPcode\":\"\",\"submissionUrl\":\"https://www.editorialmanager.com/bcan/default.aspx\",\"title\":\"BMC Cancer\",\"twitterHandle\":\"BMC_series\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"em\",\"reportingPortfolio\":\"BMC Series\",\"inReviewEnabled\":true,\"inReviewRevisionsEnabled\":true}}],\"origin\":\"\",\"ownerIdentity\":\"60bd3fec-3bf2-4c32-a8c1-e8c147af62ea\",\"owner\":[],\"postedDate\":\"June 7th, 2024\",\"published\":true,\"recentEditorialEvents\":[],\"rejectedJournal\":[],\"revision\":\"\",\"amendment\":\"\",\"status\":\"published-in-journal\",\"subjectAreas\":[],\"tags\":[],\"updatedAt\":\"2025-01-27T16:04:50+00:00\",\"versionOfRecord\":{\"articleIdentity\":\"rs-4460813\",\"link\":\"https://doi.org/10.1186/s12885-025-13483-8\",\"journal\":{\"identity\":\"bmc-cancer\",\"isVorOnly\":false,\"title\":\"BMC Cancer\"},\"publishedOn\":\"2025-01-20 15:58:21\",\"publishedOnDateReadable\":\"January 20th, 2025\"},\"versionCreatedAt\":\"2024-06-07 23:01:44\",\"video\":\"\",\"vorDoi\":\"10.1186/s12885-025-13483-8\",\"vorDoiUrl\":\"https://doi.org/10.1186/s12885-025-13483-8\",\"workflowStages\":[]},\"version\":\"v1\",\"identity\":\"rs-4460813\",\"journalConfig\":\"researchsquare\"},\"__N_SSP\":true},\"page\":\"/article/[identity]/[[...version]]\",\"query\":{\"redirect\":\"/article/rs-4460813\",\"identity\":\"rs-4460813\",\"version\":[\"v1\"]},\"buildId\":\"8U1c8b4HqxoKbykW_rLl7\",\"isFallback\":false,\"isExperimentalCompile\":false,\"dynamicIds\":[84888],\"gssp\":true,\"scriptLoader\":[]}","source_license":"CC-BY-4.0","license_restricted":false}