Dissecting heterogeneity and immune cell populations in non-small cell lung cancer by single cell RNA sequencing | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Dissecting heterogeneity and immune cell populations in non-small cell lung cancer by single cell RNA sequencing Tao Yu, Xuehan Gao, Jueyi Zhou, Liping Zhao, Jihong Feng This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3174725/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Lung cancer is the most common and aggressive cancer and the leading cause of cancer-related death worldwide, with non-smallcell lung cancer (NSCLC) being the most common type. Although traditional therapies include chemotherapy, radiation therapy, molecularly targeted therapy, and immunotherapy, 5-year survival rates for lung cancer patients have improved little. With the rapid development of targeted therapeutic drugs and immunotherapy, the clinical therapeutic effect of non-small cell lung cancer has been greatly improved. However, the issue of tumor heterogeneity in non-small cell lung cancer has received increasing attention and is not currently addressed at single-cell resolution. Therefore, exploring the impact of highly heterogeneous cells on diseases from the genomic and transcriptome levels respectively, and identifying the main influencing cell subsets, could provide a basis for the diagnosis and treatment of diseases. Methods In this study, integrated single-cell RNA sequencing (scRNA-seq) samples from Non-Small-Cell Lung Cancer (NSCLC) samples and paracancerous control samples were downloaded from the high-throughput Gene Expression Omnibus (GEO) data and batch RNA-seq data for analysis. Three NSCLC cell subsets in different differentiation states were compared and analyzed. GSEA-GO analysis predicts the biological functions and pathways of differentiation-related genes. Results The sequencing results of a total of 4320 cells from 11 NSCLC samples and 5 paracancerous lung tissue sample were obtained from the GEO database. After data standardization and data filtering, all cells were subjected to unsupervised clustering to obtain 3 different clusters, which were visualized after dimensionality reduction through T-SNE, and 10 differential marker genes were analyzed and screened, which can be clustered in different clusters. Gene set enrichment analysis found that CDRG was significantly associated with immune regulation and immune response, and 278 NSCLC cell differentiation related genes (CDRG) were identified. Conclusion Our study identified NSCLC cells with distinct differentiation characteristics based on single-cell sequencing data from GEO, emphasizing the important role of cell differentiation in predicting the clinical outcome of NSCLC patients and their potential response to immunotherapy. non-small cell lung cancer tumor heterogeneity immune prognosis single-cell sequencing GEO database GSEA analysis Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Non-small cell lung cancer (NSCLC) is the main cause of lung cancer-related death. About 85% of new lung cancer patients are pathologically classified as NSCLC, and the 5-year survival rate of this pathological type is less than 16% [1]. NSCLC not only has unique pathological characteristics at the tissue level, but also exhibits obvious tumor heterogeneity at the cellular and molecular levels [2]. During the occurrence and development of NSCLC, the evolution of tumor cells gradually diversifies, and the corresponding tumor microenvironment is infiltrated by a variety of immune-related components, which leads to more immune heterogeneity in NSCLC, and this heterogeneity varies with tumor progression and immunotherapy intervention. And there are changes in different dimensions [3,4]. Therefore, in-depth study of tumor immune heterogeneity is crucial for the development of effective treatment of NSCLC [5]. Although, lung cancer treatment entered the immune era as early as 5 years ago [6]. However, biomarkers and predictors that affect the prognosis of patients with immunotherapy have not been fully elucidated, and existing predictive models are far from satisfactory. The rapid development of related detection technologies such as single-cell multi-omics sequencing has prompted RNA-based detection to gradually become a powerful method for studying tumor immune heterogeneity. This technology mainly describes cell state transformation and functional clustering by comprehensively studying the characteristics of tumor sample genomes [7]. Single-cell sequencing technology can effectively predict the differentiation trajectories of tumor-related cells. Based on the predicted cell differentiation status, researchers further study tumor cell subsets, and in-depth reveal cell-related tumorigenic mechanisms and pathways [8]. According to current reports, single-cell scRNA-seq-based methods have accelerated the study of immune heterogeneity in various tumors, such as acute myeloid leukemia, breast cancer, pancreatic ductal adenocarcinoma, and malignant melanoma [9-11]. In this study, we downloaded and screened transcriptome data of study-compliant NSCLC samples and paracancerous control samples. First, we used single-cell RNA-sequencing (scRNA-seq) data to identify cell subsets in different differentiation states by trajectory analysis and to identify important NSCLC cell differentiation-related genes (CDRGs). Second, we explored the biological functions of CDRGs and found that they are involved in tumor immune regulation and immune responses. We predict and verify the different differentiation trajectories and tumor-promoting mechanisms of NSCLC cells, providing an important data basis for further modeling to predict tumor immunotherapy response and patient survival. Materials and methods Data download and collection This study mainly downloads data from the GEO database (https://www.ncbi.nlm.nih.gov/geo/), and the retrieval conditions are set as follows: (1) Keywords: lung cancer (set the pathological type as non-small cell lung cancer), (2) File options: the data type is "single-cell sequencing", and the data category is "transcriptome analysis". Download the single-cell RNA sequencing information of primary lung cancer patients uploaded as of June 1, 2020, and download 16 case samples, of which 11 samples are primary non-small cell lung cancer, and 5 sample is from adjacent normal lung tissue as a control group, a total of 4320 cells were detected, using Smart-seq2 technology, and the sequencing platform was Illumina Next Seq 500. Identify highly variable genes The feature subsets showing high inter-cell variation in the dataset were selected by the "Find Variable Features" function of the "Seurat" function package to select highly variable genes (HVGs), and the relationship between the mean and variance of expression was calculated. Using the vst method: linearly fit the log (variance) and log (mean) with loess, standardize the mean of the detected target gene and the variance of the expected target gene, and calculate the variance of the target gene with the maximum standardized gene expression. Principal Component (PCA) Analysis and Nonlinear Dimensionality Reduction (T-SNE) To reduce the effect of tumor-to-tumor variability on the comprehensive analysis of tumor cells, we refocused the data within each tumor individually so that the mean value for each gene in each tumor cell was zero. The variance matrix for PCA was generated using the method outlined by Shalek et al [12]. To reduce the weight of less reliable "missing" values in the data. The purpose of the T-distributed Stochastic Neighbor Embedding (T-SNE) algorithm is to understand the diversity of data and perform low-dimensional clustering. Calling the T-SNE function to map the data in the high-dimensional space to the low-dimensional space on the basis of retaining the local features of the data set is one of the better methods for data dimensionality reduction and visualization. The disadvantage is that it takes up a lot of memory and calculates the running time of the function for a long time. Orbital Analysis and CDRG Identification Differentiation trajectories of tumor cells in NCSLC-scRNAseq samples were described using the Monocle 2 technology [13]. Individual cells are projected into this space and arranged in trajectories with branch points. Cells in the same branch are generally considered to be in the same state of differentiation, while cells located in different branches are considered to have different cellular differentiation characteristics. Then, gene differential expression analysis was performed between each branch, and the significantly differentially expressed genes were defined as branch-dependent genes, and were captured and marked by the software. These differentially expressed marker genes located in different clades were defined as cell differentiation-related genes (CDRGs). Gene Set Enrichment Analysis (GSEA) GSEA (http://software.broadinstitute.org/gsea/index.jsp) is a gene probe enrichment assay based on the evaluation of various levels of gene probes from microarray data. GSEA generates an ordered list of all genes based on the correlation of selected gene expression, and performs KEGG signaling pathway analysis and GO biological function enrichment analysis, which are used to identify the relevant molecular mechanisms of NSCLC cells in different differentiation states. Statistical analysis All data of the study were conducted using R package software (version 3.6.1) and Perl scripting tools (version 5.30.0). The mean and standard deviation represent continuous variables, while the frequency and percentage represent categorical variables. The results with the default P < 0.05 were statistically significant and could be included in the next analysis. Results Identification of related genes based on single-cell sequencing data Following quality control criteria and normalization of NSCLC-scRNAseq data, 191 low-quality cells were excluded and 4129 cells from NSCLC cores were included in the analysis, including a total of 14901 corresponding genes. The ratio of ribosomal genes and the number of genes in the cells are calculated separately by the "Seurat" package of the R software, "nFeature RNA" is the number of detected genes whose expression level is greater than 0, and nCount is the sum of the expression levels of all genes. The number of genes, the number of cells and the proportion of genes were marked and their distribution frequencies were counted (Fig 1). Data normalization For standard selection, we filtered cells based on gene count and mitochondrial ratio, filtering for the lowest gene count, and selecting cells with a ratio greater than 50% of the characteristic standard count. In our selected dataset, no mitochondrial highly expressed genes were detected, indicating that all cells were in good condition, and all cells were reserved for further analysis. The results showed that the number of detected genes was significantly correlated with the sequencing depth, with a Pearson correlation coefficient of 0.63 (Fig 2). Identify and screen hypervariable genes ANOVA plot showing 1500 highly variable genes out of 14901 genes from NSCLC samples. Red dots represent highly variable genes and black dots represent immutable genes. CCL4, GZMK, TNFRSF4, KLRG1, GNLY, C12orf5, FGFBP2, MYO1G, SLC25A, and 14APEH with gene names marked in the figure are the 10 genes with the highest degree of variation (Fig 3). PCA principal component analysis and T-SNE dimensionality reduction Principal component analysis (PCA) was used to determine available dimensions and screen for related genes. PCA results did not show clear separation between cells in NSCLC (Fig. 4A). We selected the top 20 principal components (PCs) with P <0.05 for subsequent analysis (Fig. 4B). Applying the T-SNE algorithm to the dimensionality reduction of 20 PCs successfully classified 3 cell clusters, and the clustering results are shown in Fig 4C. Using the Wilcox method, we set the screening index to logFCfilter to 0.5 and adjPvalFilter to 0.05 to find significantly high-expressed genes in each cluster, and screened out 9876 differentially expressed genes. Then we performed cluster analysis on the Top10 differential genes. The clustering results showed that 10 differential marker genes could be clustered in different clusters. The colors from purple to yellow indicate gene expression levels from low to high (Fig. 4D). Tumor differentiation trajectory analysis and determination of biological functions of NSCLC differentiation-related genes Using Monocle 2 technology to analyze the cell differentiation trajectory, it can be seen that the tumor cells differentiate into 3 branches, and each branch has different NSCLC immune-related cells. Among them, branch I distributed 410 cells, branch II distributed 404 cells, and branch III distributed 438 cells (Fig. 5A). The three branches are defined by type I, II, and III cell subsets, respectively. Gene difference analysis obtained 275 type I CDRGs, 191 type II CDRGs and 198 type III CDRGs, and the differences in the degree of differentiation of the three types of cell subsets were statistically significant. Finally, GSEA functional enrichment analysis found that type I CDRGs were significantly associated with immune response modulation, while type II and III CDRGs were significantly associated with immune response pathways (Fig. 5B). Discussion NSCLC is the most common tumor, and the number of new cases and deaths still ranks first among malignant tumors. This tumor has significant tumor heterogeneity in the process of diagnosis and treatment [14]. Heterogeneity is characteristic of many cancers, such as lung cancer, and is associated with clinical progression and an important driver of drug resistance, so the analysis of different cell species in tumors is extremely important to reveal the mechanism of drug resistance. Single-cell sequencing technology enables specific analysis of cell populations at the single-cell level. The process mainly includes single-cell isolation, cell lysis and genomic DNA acquisition, whole genome amplification, sequencing, and data analysis. Single-cell sequencing includes single-cell whole genome sequencing and single-cell transcriptome sequencing, which explore the impact of highly heterogeneous cells on diseases from the genomic and transcriptome levels, respectively, and identify the main cell subsets that affect them, so as to provide a basis for the diagnosis and treatment of diseases. During the development of NSCLC, cells are constantly differentiated and mutated into different cell subsets to enhance immune suppression and immune evasion. As intratumor heterogeneity is increasingly recognized as one of the main reasons for tumor therapy resistance, there is an urgent need to develop new technologies to deeply study cellular heterogeneity in NSCLC [15]. However, up to now, studies on the cellular heterogeneity, microenvironmental heterogeneity, and tumor immune heterogeneity of NSCLC are limited. Therefore, we preliminarily explored the heterogeneity results of NSCLC in genetic testing from the level of single-cell RNA, in order to provide a data basis for the study of tumor heterogeneity in NSCLC. With the advent of the era of lung cancer immunotherapy, it is increasingly recognized that metabolic changes in cancer cells can affect immune cell function and lead to tumor immune evasion. It has a certain impact on the effect of immunotherapy [16]. Immune cells in the tumor stroma sometimes colonize an environment with different cell subsets and nutrients as they tour the body, and the cross-talk between tumor cells and immune cells ultimately results in an environment that promotes tumor growth and metastasis, In medicine, it is called the tumor immune microenvironment [17,18]. A large number of studies have shown that the heterogeneity of the tumor immune microenvironment affects the immunotherapy effect of tumor patients from many aspects such as genetics and immunity [19]. An in-depth understanding of the heterogeneity of this environment will facilitate the development of therapeutic approaches that simultaneously target multiple components of the immune microenvironment, thereby increasing the likelihood of good clinical outcomes [20]. Single-cell transcriptome sequencing enables a dynamic representation of gene lineage and heterogeneity to better define the cell types examined [21]. Almost all studies of predictive biomarkers associated with clinical prognosis in tumors are based on gene-level analysis of a single biopsy sample [22]. Based on single-cell scRNA-seq sequencing technology, this study compared and analyzed different cell subsets in tumor samples, and predicted the differentiation trajectory of tumor cells and their differentially expressed cell differentiation-related genes. The early stages of occurrence have already emerged. In this study, we identified 3 cell clusters from 11 NSCLC samples, and based on cell trajectory analysis, NSCLC cells were projected into three subpopulations with significantly different differentiation characteristics. Screening identified subpopulation-dependent Cell Differentiation-Related Genes (CDRGs). Through GSEA-GO biological function correlation analysis, we found that this differentiation model was significantly associated with tumor immune regulation and immune response, implying an intrinsic correlation between NSCLC cell differentiation and intratumoral immune and metabolic biology. Of course, the current research still has certain limitations. On the one hand, the patient details obtained by downloading are not complete enough, and some clinical parameters, such as tumor imaging results, medical records and medical history, and details of surgical records, cannot be downloaded, so the nomogram cannot be input. , on the other hand, has a limited number of cases available for download. To make this study more clinically meaningful, the predictive model needs to be further validated in future large-scale cohorts. In summary, we predicted NSCLC cells with different differentiation characteristics based on the scRNA-seq data of the GEO database, and then performed differential expression analysis to find cell differentiation-related genes (CDRGs). Its biological functions and metabolic pathways involved. This study highlights the unique cellular differentiation trajectories of NSCLC cells and their important role in predicting clinical outcome and tumor immunotherapy response in predicting clinical outcome and tumor immunotherapy response in lung cancer patients. Declarations Funding This work was Supported by the Program of Natural Science Foundation of Zhejiang Province (LY20H160017) , Chinese Medicine Study Foundation of Zhejiang Province (2020ZB292) and PhD research startup foundation of Lishui People's Hospital (2020bs01). Competing interests The authors have no relevant financial or non-financial interests to disclose. Author contributions Tao Yu and Xuehan Gao conceived and designed the study, participated with editing the drafted manuscript and submission of the article. Jueyi Zhou, Liping Zhao, Jihong Feng conducted the database mining. Tao Yu and Xuehan Gao analyzed data and compiled charts. All authors participated in the design of the study. All authors read and approved the final manuscript. Data Availability The datasets generated during and/or analysed during the current study are available from the corresponding author on reasonable request. Ethics approval This is an observational study. People's Hospital of Inner Mongolia Autonomous Region has confirmed that no ethical approval is required. Consent to publish The authors affirm that human research participants provided informed consent for publication of the images in Figure(s) 1, 2, 3, 4, and 5. References Torre LA, Siegel RL, Jemal A (2016) Lung cancer statistics. Adv Exp Med Biol 893(1):1–19 Seow WJ, Matsuo K, Hsiung CA (2017) Association between GWAS-identified lung adenocarcinoma susceptibility loci and EGFR mutations in never-smoking Asian women, and comparison withfindings from Western populations. Hum Mol Genet 26(1):454–465 Goveia J, Rohlenova K, Taverna F (2020) An Integrated Gene Expression Landscape Profiling Approach to Identify Lung Tumor Endothelial Cell Heterogeneity and Angiogenic Candidates. Cancer Cell 37(1):21–36 Kron A, Scheffler M, Heydt C (2021) Genetic Heterogeneity of MET-Aberrant NSCLC and Its Impact on the Outcome of Immunotherapy. J Thorac Oncol 16(4):572–582 Jia Q, Wu W, Wang Y (2018) Local mutational diversity drives intratumoral immune heterogeneity in non-small cell lung cancer. Nat Commun 9(1):53–61 Sui H, Ma N, Wang Y (2018) Anti-PD-1/PD-L1 Therapy for Non-Small-Cell Lung Cancer: Toward Personalized Medicine and Combination Strategies. J Immunol Res 8(1):20–69 Klein AM, Mazutis L, Akartuna I, Tallapragada N, Veres A, Li V et al (2015) Droplet barcoding for single-cell transcriptomics applied to embryonic stem cells. Cell 161(3):1187–1201 Wagner A, Regev A, Yosef N (2016) Revealing the vectors of cellular identity with single-cell genomics. Nat Biotechnol Nat Biotechnol 34(4):1145–1160 Zheng GX, Terry JM, Belgrader P (2019) Massively parallel digital transcriptional profiling of single cells. Cancer Manag Res 11(1):7197–7210 Azizi E, Carr AJ, Plitas G (2018) Single-cell map of diverse immune phenotypes in the breast tumor microenvironment. Cell 174(1):1293–1308 Peng J, Sun BF, Chen CY (2019) Single-cell RNA-seq highlights intra-tumoral heterogeneity and malignant progression in pancreatic ductal adenocarcinoma. Cell Res 29(1):725–738 Shalek AK, Satija R, Shuga J (2014) Single-cell RNA-seq reveals dynamic paracrine control cellular variation. Nature 510(7):363–369 Qiu X, Mao Q, Tang Y (2017) Reversed graph embedding resolves complex single-cell trajectories. Nat Methods 14(2):979–982 Jonna S, Subramaniam DS (2019) Molecular diagnostics and targeted therapies in non-small cell lung cancer (NSCLC): an update. Discov Med 27(148):167–170 Pe'er D, Ogawa S, Elhanani O (2021) Tumor heterogeneity. Cancer Cell 39(8):1015–1017 Ren X, Zhang L, Zhang Y (2021) Insights Gained from Single-Cell Analysis of Immune Cells in the Tumor. Microenvironment Annu Rev Immunol 26(39):583–609 Sun YF, Wu L, Liu SP (2021) Dissecting spatial heterogeneity and the immune-evasion mechanism of CTCs by single-cell RNA-seq in hepatocellular carcinoma. Nat Commun 12(1):40–91 Cooper LA, Demicco EG, Saltz JH (2018) PanCancer insights from The Cancer Genome Atlas: the pathologist's perspective. J Pathol 244(5):512–524 Chen YP, Lv JW, Mao YP (2021) Unraveling tumour microenvironment heterogeneity in nasopharyngeal carcinoma identifies biologically distinct immune subtypes predicting prognosis and immunotherapy responses. Mol Cancer 20(1):14–31 Havel JJ, Chowell D, Chan TA (2019) The evolving landscape of biomarkers for checkpoint inhibitor immunotherapy. Nat Rev Cancer 19(3):133–150 Gao S (2018) Data Analysis in Single-Cell Transcriptome Sequencing. Methods Mol Biol 1754(1):311–326 Stepan H, Hund M, Andraczek T (2020) Combining Biomarkers to Predict Pregnancy Complications and Redefine Preeclampsia: The Angiogenic-Placental Syndrome. Hypertension 75(4):918–926 .ung cancer patients Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3174725","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":219378841,"identity":"e316380a-11bd-4822-8808-a2338ac3474b","order_by":0,"name":"Tao Yu","email":"","orcid":"","institution":"People's Hospital of Inner Mongolia Autonomous Region","correspondingAuthor":false,"prefix":"","firstName":"Tao","middleName":"","lastName":"Yu","suffix":""},{"id":219378842,"identity":"d6ff2a4f-2710-4f6d-9f18-2bddf6e9804c","order_by":1,"name":"Xuehan Gao","email":"","orcid":"","institution":"Zunyi Medical University","correspondingAuthor":false,"prefix":"","firstName":"Xuehan","middleName":"","lastName":"Gao","suffix":""},{"id":219378843,"identity":"2fe639de-f43d-4197-b8af-df98aa1a8e03","order_by":2,"name":"Jueyi Zhou","email":"","orcid":"","institution":"The Sixth Affiliated Hospital of Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Jueyi","middleName":"","lastName":"Zhou","suffix":""},{"id":219378844,"identity":"8c4c5827-c768-4c26-abd3-b9f95cfea7e9","order_by":3,"name":"Liping Zhao","email":"","orcid":"","institution":"The Sixth Affiliated Hospital of Wenzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Liping","middleName":"","lastName":"Zhao","suffix":""},{"id":219378845,"identity":"9fd33f5b-c099-4fa0-a327-3a0e8d6ad1af","order_by":4,"name":"Jihong Feng","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAmklEQVRIiWNgGAWjYBACxhkMDB8+GNjYkaSFceaMgrRkEqyRYGCczfPhEGMD0TqYZzcwNtsYHGBmYD98dANxDptzgLE5x+AOHwNPWtoN4rTMSGB/nGPwjJlBgseMaC2MzRYGhxkbSNPCQLKWxh6DtGQ2ov1iCNTS8OOPjR0/++FjRGpp4P8AZrARpRwE5IlWOQpGwSgYBSMXAABCBS10k/QHfQAAAABJRU5ErkJggg==","orcid":"","institution":"The Sixth Affiliated Hospital of Wenzhou Medical University","correspondingAuthor":true,"prefix":"","firstName":"Jihong","middleName":"","lastName":"Feng","suffix":""}],"badges":[],"createdAt":"2023-07-16 09:44:23","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3174725/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3174725/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":40376710,"identity":"4614ee7a-c41e-4d71-9d2c-faa5c841bbd4","added_by":"auto","created_at":"2023-07-21 13:32:47","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":143559,"visible":true,"origin":"","legend":"\u003cp\u003eThe number of genes and the proportion of gene expression markers and their distribution frequencies from 6 samples.\u003c/p\u003e","description":"","filename":"Fig.1.png","url":"https://assets-eu.researchsquare.com/files/rs-3174725/v1/74ace3ed50b7c4d540ca2cd8.png"},{"id":40377332,"identity":"19d7bb75-e84a-4a69-a4b1-7b70c87eb171","added_by":"auto","created_at":"2023-07-21 13:40:47","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":168718,"visible":true,"origin":"","legend":"\u003cp\u003eNumber of genes significantly correlated with sequencing depth.\u003c/p\u003e","description":"","filename":"Fig.2.png","url":"https://assets-eu.researchsquare.com/files/rs-3174725/v1/2cc7b6fc04ce3c2b6eb6c8f0.png"},{"id":40377331,"identity":"034d2d5a-0fe5-4f4f-a9e8-be1005b26337","added_by":"auto","created_at":"2023-07-21 13:40:47","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":89521,"visible":true,"origin":"","legend":"\u003cp\u003eVariance plot of highly variable gene signatures between cells.\u003c/p\u003e","description":"","filename":"Fig.3.png","url":"https://assets-eu.researchsquare.com/files/rs-3174725/v1/988ff9f2f738b2f5fe4dff70.png"},{"id":40376712,"identity":"b1043921-29a6-4a70-b689-8dfedd2609fb","added_by":"auto","created_at":"2023-07-21 13:32:47","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":760747,"visible":true,"origin":"","legend":"\u003cp\u003eRelated genes after PCA and T-SNE dimensionality reduction.\u003c/p\u003e\n\u003cp\u003e(A) PCA does not show clear separation of cells in NSCLC. (B) PCA identified 20 PCs, estimated \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.05. (C) T-SNE dimensionality reduction analysis classified 3 cell clusters. (D) The top 10 marker genes for each cell cluster are shown in the heatmap.\u003c/p\u003e","description":"","filename":"Fig.4.png","url":"https://assets-eu.researchsquare.com/files/rs-3174725/v1/6e03b26746aa014dd349b772.png"},{"id":40376714,"identity":"bb8486fd-a74c-4ff3-a6a0-c956f42e1cf9","added_by":"auto","created_at":"2023-07-21 13:32:47","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":709946,"visible":true,"origin":"","legend":"\u003cp\u003eCell annotation, trajectory analysis, and GSEA analysis of three GBM cell subsets with different differentiation patterns.\u003c/p\u003e\n\u003cp\u003e(A) Trajectory analysis showing three subpopulations of NSCLC cells with distinct differentiation patterns. NSCLC CSCs were mainly distributed in roots, whereas NSCLC cells were distributed in three branches. (B) GSEA-GO analysis showed that the three subtypes of NSCLC cell subsets were significantly associated with immune regulation and immune response function.\u003c/p\u003e","description":"","filename":"Fig.5.png","url":"https://assets-eu.researchsquare.com/files/rs-3174725/v1/a970d7deff05c448b6fb7ae1.png"},{"id":43340727,"identity":"22b9c898-e309-4364-b504-b1c6524a8b4a","added_by":"auto","created_at":"2023-09-19 06:52:35","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1210374,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3174725/v1/3f0b8426-ed74-43ac-b581-34cbfa453add.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Dissecting heterogeneity and immune cell populations in non-small cell lung cancer by single cell RNA sequencing","fulltext":[{"header":"Introduction","content":"\u003cp\u003eNon-small cell lung cancer (NSCLC) is the main cause of lung cancer-related death. About 85% of new lung cancer patients are pathologically classified as NSCLC, and the 5-year survival rate of this pathological type is less than 16% [1]. NSCLC not only has unique pathological characteristics at the tissue level, but also exhibits obvious tumor heterogeneity at the cellular and molecular levels [2]. During the occurrence and development of NSCLC, the evolution of tumor cells gradually diversifies, and the corresponding tumor microenvironment is infiltrated by a variety of immune-related components, which leads to more immune heterogeneity in NSCLC, and this heterogeneity varies with tumor progression and immunotherapy intervention. And there are changes in different dimensions [3,4]. Therefore, in-depth study of tumor immune heterogeneity is crucial for the development of effective treatment of NSCLC [5]. Although, lung cancer treatment entered the immune era as early as 5 years ago [6]. However, biomarkers and predictors that affect the prognosis of patients with immunotherapy have not been fully elucidated, and existing predictive models are far from satisfactory.\u003c/p\u003e\n\u003cp\u003eThe rapid development of related detection technologies such as single-cell multi-omics sequencing has prompted RNA-based detection to gradually become a powerful method for studying tumor immune heterogeneity. This technology mainly describes cell state transformation and functional clustering by comprehensively studying the characteristics of tumor sample genomes [7]. Single-cell sequencing technology can effectively predict the differentiation trajectories of tumor-related cells. Based on the predicted cell differentiation status, researchers further study tumor cell subsets, and in-depth reveal cell-related tumorigenic mechanisms and pathways [8]. According to current reports, single-cell scRNA-seq-based methods have accelerated the study of immune heterogeneity in various tumors, such as acute myeloid leukemia, breast cancer, pancreatic ductal adenocarcinoma, and malignant melanoma [9-11].\u003c/p\u003e\n\u003cp\u003eIn this study, we downloaded and screened transcriptome data of study-compliant NSCLC samples and paracancerous control samples. First, we used single-cell RNA-sequencing (scRNA-seq) data to identify cell subsets in different differentiation states by trajectory analysis and to identify important NSCLC cell differentiation-related genes (CDRGs). Second, we explored the biological functions of CDRGs and found that they are involved in tumor immune regulation and immune responses. We predict and verify the different differentiation trajectories and tumor-promoting mechanisms of NSCLC cells, providing an important data basis for further modeling to predict tumor immunotherapy response and patient survival.\u003c/p\u003e"},{"header":"Materials and methods","content":"\u003cp\u003eData download and collection\u003c/p\u003e\n\u003cp\u003eThis study mainly downloads data from the GEO database (https://www.ncbi.nlm.nih.gov/geo/), and the retrieval conditions are set as follows: (1) Keywords: lung cancer (set the pathological type as non-small cell lung cancer), (2) File options: the data type is \u0026quot;single-cell sequencing\u0026quot;, and the data category is \u0026quot;transcriptome analysis\u0026quot;. Download the single-cell RNA sequencing information of primary lung cancer patients uploaded as of June 1, 2020, and download 16 case samples, of which 11 samples are primary non-small cell lung cancer, and 5 sample is from adjacent normal lung tissue as a control group, a total of 4320 cells were detected, using Smart-seq2 technology, and the sequencing platform was Illumina Next Seq 500.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIdentify highly variable genes\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe feature subsets showing high inter-cell variation in the dataset were selected by the \u0026quot;Find Variable Features\u0026quot; function of the \u0026quot;Seurat\u0026quot; function package to select highly variable genes (HVGs), and the relationship between the mean and variance of expression was calculated. Using the vst method: linearly fit the log (variance) and log (mean) with loess, standardize the mean of the detected target gene and the variance of the expected target gene, and calculate the variance of the target gene with the maximum standardized gene expression.\u003c/p\u003e\n\u003cp\u003ePrincipal Component (PCA) Analysis and Nonlinear Dimensionality Reduction (T-SNE)\u003c/p\u003e\n\u003cp\u003eTo reduce the effect of tumor-to-tumor variability on the comprehensive analysis of tumor cells, we refocused the data within each tumor individually so that the mean value for each gene in each tumor cell was zero. The variance matrix for PCA was generated using the method outlined by Shalek et al [12]. To reduce the weight of less reliable \u0026quot;missing\u0026quot; values in the data. The purpose of the T-distributed Stochastic Neighbor Embedding (T-SNE) algorithm is to understand the diversity of data and perform low-dimensional clustering. Calling the T-SNE function to map the data in the high-dimensional space to the low-dimensional space on the basis of retaining the local features of the data set is one of the better methods for data dimensionality reduction and visualization. The disadvantage is that it takes up a lot of memory and calculates the running time of the function for a long time.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eOrbital Analysis and CDRG Identification\u003c/p\u003e\n\u003cp\u003eDifferentiation trajectories of tumor cells in NCSLC-scRNAseq samples were described using the Monocle 2 technology [13]. Individual cells are projected into this space and arranged in trajectories with branch points. Cells in the same branch are generally considered to be in the same state of differentiation, while cells located in different branches are considered to have different cellular differentiation characteristics. Then, gene differential expression analysis was performed between each branch, and the significantly differentially expressed genes were defined as branch-dependent genes, and were captured and marked by the software. These differentially expressed marker genes located in different clades were defined as cell differentiation-related genes (CDRGs).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eGene Set Enrichment Analysis (GSEA)\u003c/p\u003e\n\u003cp\u003eGSEA (http://software.broadinstitute.org/gsea/index.jsp) is a gene probe enrichment assay based on the evaluation of various levels of gene probes from microarray data. GSEA generates an ordered list of all genes based on the correlation of selected gene expression, and performs KEGG signaling pathway analysis and GO biological function enrichment analysis, which are used to identify the relevant molecular mechanisms of NSCLC cells in different differentiation states.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eStatistical analysis\u003c/p\u003e\n\u003cp\u003eAll data of the study were conducted using R package software (version 3.6.1) and Perl scripting tools (version 5.30.0). The mean and standard deviation represent continuous variables, while the frequency and percentage represent categorical variables. The results with the default \u003cem\u003eP\u003c/em\u003e \u0026lt; 0.05 were statistically significant and could be included in the next analysis.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eIdentification of related genes based on single-cell sequencing data\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFollowing quality control criteria and normalization of NSCLC-scRNAseq data, 191 low-quality cells were excluded and 4129 cells from NSCLC cores were included in the analysis, including a total of 14901 corresponding genes. The ratio of ribosomal genes and the number of genes in the cells are calculated separately by the \u0026quot;Seurat\u0026quot; package of the R software, \u0026quot;nFeature RNA\u0026quot; is the number of detected genes whose expression level is greater than 0, and nCount is the sum of the expression levels of all genes. The number of genes, the number of cells and the proportion of genes were marked and their distribution frequencies were counted (Fig 1).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eData normalization\u003c/p\u003e\n\u003cp\u003eFor standard selection, we filtered cells based on gene count and mitochondrial ratio, filtering for the lowest gene count, and selecting cells with a ratio greater than 50% of the characteristic standard count. In our selected dataset, no mitochondrial highly expressed genes were detected, indicating that all cells were in good condition, and all cells were reserved for further analysis. The results showed that the number of detected genes was significantly correlated with the sequencing depth, with a Pearson correlation coefficient of 0.63 (Fig 2).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eIdentify and screen hypervariable genes\u003c/p\u003e\n\u003cp\u003eANOVA plot showing 1500 highly variable genes out of 14901 genes from NSCLC samples. Red dots represent highly variable genes and black dots represent immutable genes. CCL4, GZMK, TNFRSF4, KLRG1, GNLY, C12orf5, FGFBP2, MYO1G, SLC25A, and 14APEH with gene names marked in the figure are the 10 genes with the highest degree of variation (Fig 3).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003ePCA principal component analysis and T-SNE dimensionality reduction\u003c/p\u003e\n\u003cp\u003ePrincipal component analysis (PCA) was used to determine available dimensions and screen for related genes. PCA results did not show clear separation between cells in NSCLC (Fig. 4A). We selected the top 20 principal components (PCs) with \u003cem\u003eP\u003c/em\u003e\u0026lt;0.05 for subsequent analysis (Fig. 4B). Applying the T-SNE algorithm to the dimensionality reduction of 20 PCs successfully classified 3 cell clusters, and the clustering results are shown in Fig 4C. Using the Wilcox method, we set the screening index to logFCfilter to 0.5 and adjPvalFilter to 0.05 to find significantly high-expressed genes in each cluster, and screened out 9876 differentially expressed genes. Then we performed cluster analysis on the Top10 differential genes. The clustering results showed that 10 differential marker genes could be clustered in different clusters. The colors from purple to yellow indicate gene expression levels from low to high (Fig. 4D).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTumor differentiation trajectory analysis and determination of biological functions of NSCLC differentiation-related genes\u003c/p\u003e\n\u003cp\u003eUsing Monocle 2 technology to analyze the cell differentiation trajectory, it can be seen that the tumor cells differentiate into 3 branches, and each branch has different NSCLC immune-related cells. Among them, branch I distributed 410 cells, branch II distributed 404 cells, and branch III distributed 438 cells (Fig. 5A). The three branches are defined by type I, II, and III cell subsets, respectively. Gene difference analysis obtained 275 type I CDRGs, 191 type II CDRGs and 198 type III CDRGs, and the differences in the degree of differentiation of the three types of cell subsets were statistically significant. Finally, GSEA functional enrichment analysis found that type I CDRGs were significantly associated with immune response modulation, while type II and III CDRGs were significantly associated with immune response pathways (Fig. 5B).\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eNSCLC is the most common tumor, and the number of new cases and deaths still ranks first among malignant tumors. This tumor has significant tumor heterogeneity in the process of diagnosis and treatment [14]. Heterogeneity is characteristic of many cancers, such as lung cancer, and is associated with clinical progression and an important driver of drug resistance, so the analysis of different cell species in tumors is extremely important to reveal the mechanism of drug resistance. Single-cell sequencing technology enables specific analysis of cell populations at the single-cell level. The process mainly includes single-cell isolation, cell lysis and genomic DNA acquisition, whole genome amplification, sequencing, and data analysis. Single-cell sequencing includes single-cell whole genome sequencing and single-cell transcriptome sequencing, which explore the impact of highly heterogeneous cells on diseases from the genomic and transcriptome levels, respectively, and identify the main cell subsets that affect them, so as to provide a basis for the diagnosis and treatment of diseases. During the development of NSCLC, cells are constantly differentiated and mutated into different cell subsets to enhance immune suppression and immune evasion. As intratumor heterogeneity is increasingly recognized as one of the main reasons for tumor therapy resistance, there is an urgent need to develop new technologies to deeply study cellular heterogeneity in NSCLC [15]. However, up to now, studies on the cellular heterogeneity, microenvironmental heterogeneity, and tumor immune heterogeneity of NSCLC are limited. Therefore, we preliminarily explored the heterogeneity results of NSCLC in genetic testing from the level of single-cell RNA, in order to provide a data basis for the study of tumor heterogeneity in NSCLC.\u003c/p\u003e\n\u003cp\u003eWith the advent of the era of lung cancer immunotherapy, it is increasingly recognized that metabolic changes in cancer cells can affect immune cell function and lead to tumor immune evasion. It has a certain impact on the effect of immunotherapy [16]. Immune cells in the tumor stroma sometimes colonize an environment with different cell subsets and nutrients as they tour the body, and the cross-talk between tumor cells and immune cells ultimately results in an environment that promotes tumor growth and metastasis, In medicine, it is called the tumor immune microenvironment [17,18]. A large number of studies have shown that the heterogeneity of the tumor immune microenvironment affects the immunotherapy effect of tumor patients from many aspects such as genetics and immunity [19]. An in-depth understanding of the heterogeneity of this environment will facilitate the development of therapeutic approaches that simultaneously target multiple components of the immune microenvironment, thereby increasing the likelihood of good clinical outcomes [20].\u003c/p\u003e\n\u003cp\u003eSingle-cell transcriptome sequencing enables a dynamic representation of gene lineage and heterogeneity to better define the cell types examined [21]. Almost all studies of predictive biomarkers associated with clinical prognosis in tumors are based on gene-level analysis of a single biopsy sample [22]. Based on single-cell scRNA-seq sequencing technology, this study compared and analyzed different cell subsets in tumor samples, and predicted the differentiation trajectory of tumor cells and their differentially expressed cell differentiation-related genes. The early stages of occurrence have already emerged. In this study, we identified 3 cell clusters from 11 NSCLC samples, and based on cell trajectory analysis, NSCLC cells were projected into three subpopulations with significantly different differentiation characteristics. Screening identified subpopulation-dependent Cell Differentiation-Related Genes (CDRGs). Through GSEA-GO biological function correlation analysis, we found that this differentiation model was significantly associated with tumor immune regulation and immune response, implying an intrinsic correlation between NSCLC cell differentiation and intratumoral immune and metabolic biology. Of course, the current research still has certain limitations. On the one hand, the patient details obtained by downloading are not complete enough, and some clinical parameters, such as tumor imaging results, medical records and medical history, and details of surgical records, cannot be downloaded, so the nomogram cannot be input. , on the other hand, has a limited number of cases available for download. To make this study more clinically meaningful, the predictive model needs to be further validated in future large-scale cohorts.\u003c/p\u003e\n\u003cp\u003eIn summary, we predicted NSCLC cells with different differentiation characteristics based on the scRNA-seq data of the GEO database, and then performed differential expression analysis to find cell differentiation-related genes (CDRGs). Its biological functions and metabolic pathways involved. This study highlights the unique cellular differentiation trajectories of NSCLC cells and their important role in predicting clinical outcome and tumor immunotherapy response in predicting clinical outcome and tumor immunotherapy response in lung cancer patients.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eFunding\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThis work was Supported by the Program of Natural Science Foundation of Zhejiang Province (LY20H160017) , Chinese Medicine Study Foundation of Zhejiang Province (2020ZB292) and PhD research startup foundation of Lishui People\u0026apos;s Hospital (2020bs01).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eCompeting interests\u003c/p\u003e\n\u003cp\u003eThe authors have no relevant financial or non-financial interests to disclose.\u003c/p\u003e\n\u003cp\u003eAuthor contributions\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTao Yu and Xuehan Gao conceived and designed the study, participated with editing the drafted manuscript and submission of the article. Jueyi Zhou, Liping Zhao, Jihong Feng conducted the database mining. Tao Yu and Xuehan Gao analyzed data and compiled charts. All authors participated in the design of the study. All authors read and approved the final manuscript.\u0026nbsp;\u003c/p\u003e\n\u003ch4\u003eData Availability\u003c/h4\u003e\n\u003cp\u003eThe datasets generated during and/or analysed during the current study are available from the corresponding author on reasonable request.\u0026nbsp;\u003c/p\u003e\n\u003ch4\u003eEthics approval\u003c/h4\u003e\n\u003cp\u003eThis is an observational study.\u0026nbsp;People\u0026apos;s Hospital of Inner Mongolia Autonomous Region\u0026nbsp;has confirmed that no ethical approval is required.\u003c/p\u003e\n\u003ch4\u003eConsent to publish\u003c/h4\u003e\n\u003cp\u003eThe authors affirm that human research participants provided informed consent for publication of the images in Figure(s) 1, 2, 3, 4, and 5.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eTorre LA, Siegel RL, Jemal A (2016) Lung cancer statistics. Adv Exp Med Biol 893(1):1\u0026ndash;19\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSeow WJ, Matsuo K, Hsiung CA (2017) Association between GWAS-identified lung adenocarcinoma susceptibility loci and EGFR mutations in never-smoking Asian women, and comparison withfindings from Western populations. Hum Mol Genet 26(1):454\u0026ndash;465\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoveia J, Rohlenova K, Taverna F (2020) An Integrated Gene Expression Landscape Profiling Approach to Identify Lung Tumor Endothelial Cell Heterogeneity and Angiogenic Candidates. Cancer Cell 37(1):21\u0026ndash;36\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKron A, Scheffler M, Heydt C (2021) Genetic Heterogeneity of MET-Aberrant NSCLC and Its Impact on the Outcome of Immunotherapy. J Thorac Oncol 16(4):572\u0026ndash;582\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJia Q, Wu W, Wang Y (2018) Local mutational diversity drives intratumoral immune heterogeneity in non-small cell lung cancer. Nat Commun 9(1):53\u0026ndash;61\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSui H, Ma N, Wang Y (2018) Anti-PD-1/PD-L1 Therapy for Non-Small-Cell Lung Cancer: Toward Personalized Medicine and Combination Strategies. J Immunol Res 8(1):20\u0026ndash;69\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKlein AM, Mazutis L, Akartuna I, Tallapragada N, Veres A, Li V et al (2015) Droplet barcoding for single-cell transcriptomics applied to embryonic stem cells. Cell 161(3):1187\u0026ndash;1201\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWagner A, Regev A, Yosef N (2016) Revealing the vectors of cellular identity with single-cell genomics. Nat Biotechnol Nat Biotechnol 34(4):1145\u0026ndash;1160\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZheng GX, Terry JM, Belgrader P (2019) Massively parallel digital transcriptional profiling of single cells. Cancer Manag Res 11(1):7197\u0026ndash;7210\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAzizi E, Carr AJ, Plitas G (2018) Single-cell map of diverse immune phenotypes in the breast tumor microenvironment. Cell 174(1):1293\u0026ndash;1308\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeng J, Sun BF, Chen CY (2019) Single-cell RNA-seq highlights intra-tumoral heterogeneity and malignant progression in pancreatic ductal adenocarcinoma. Cell Res 29(1):725\u0026ndash;738\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShalek AK, Satija R, Shuga J (2014) Single-cell RNA-seq reveals dynamic paracrine control cellular variation. Nature 510(7):363\u0026ndash;369\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQiu X, Mao Q, Tang Y (2017) Reversed graph embedding resolves complex single-cell trajectories. Nat Methods 14(2):979\u0026ndash;982\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJonna S, Subramaniam DS (2019) Molecular diagnostics and targeted therapies in non-small cell lung cancer (NSCLC): an update. Discov Med 27(148):167\u0026ndash;170\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePe'er D, Ogawa S, Elhanani O (2021) Tumor heterogeneity. Cancer Cell 39(8):1015\u0026ndash;1017\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRen X, Zhang L, Zhang Y (2021) Insights Gained from Single-Cell Analysis of Immune Cells in the Tumor. Microenvironment Annu Rev Immunol 26(39):583\u0026ndash;609\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSun YF, Wu L, Liu SP (2021) Dissecting spatial heterogeneity and the immune-evasion mechanism of CTCs by single-cell RNA-seq in hepatocellular carcinoma. Nat Commun 12(1):40\u0026ndash;91\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCooper LA, Demicco EG, Saltz JH (2018) PanCancer insights from The Cancer Genome Atlas: the pathologist's perspective. J Pathol 244(5):512\u0026ndash;524\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen YP, Lv JW, Mao YP (2021) Unraveling tumour microenvironment heterogeneity in nasopharyngeal carcinoma identifies biologically distinct immune subtypes predicting prognosis and immunotherapy responses. Mol Cancer 20(1):14\u0026ndash;31\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHavel JJ, Chowell D, Chan TA (2019) The evolving landscape of biomarkers for checkpoint inhibitor immunotherapy. Nat Rev Cancer 19(3):133\u0026ndash;150\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao S (2018) Data Analysis in Single-Cell Transcriptome Sequencing. Methods Mol Biol 1754(1):311\u0026ndash;326\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStepan H, Hund M, Andraczek T (2020) Combining Biomarkers to Predict Pregnancy Complications and Redefine Preeclampsia: The Angiogenic-Placental Syndrome. Hypertension 75(4):918\u0026ndash;926 .ung cancer patients\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e "}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"non-small cell lung cancer, tumor heterogeneity, immune prognosis, single-cell sequencing, GEO database, GSEA analysis","lastPublishedDoi":"10.21203/rs.3.rs-3174725/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3174725/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eLung cancer is the most common and aggressive cancer and the leading cause of cancer-related death worldwide, with non-smallcell lung cancer (NSCLC) being the most common type. Although traditional therapies include chemotherapy, radiation therapy, molecularly targeted therapy, and immunotherapy, 5-year survival rates for lung cancer patients have improved little. With the rapid development of targeted therapeutic drugs and immunotherapy, the clinical therapeutic effect of non-small cell lung cancer has been greatly improved. However, the issue of tumor heterogeneity in non-small cell lung cancer has received increasing attention and is not currently addressed at single-cell resolution. Therefore, exploring the impact of highly heterogeneous cells on diseases from the genomic and transcriptome levels respectively, and identifying the main influencing cell subsets, could provide a basis for the diagnosis and treatment of diseases.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn this study, integrated single-cell RNA sequencing (scRNA-seq) samples from Non-Small-Cell Lung Cancer (NSCLC) samples and paracancerous control samples were downloaded from the high-throughput Gene Expression Omnibus (GEO) data and batch RNA-seq data for analysis. Three NSCLC cell subsets in different differentiation states were compared and analyzed. GSEA-GO analysis predicts the biological functions and pathways of differentiation-related genes.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe sequencing results of a total of 4320 cells from 11 NSCLC samples and 5 paracancerous lung tissue sample were obtained from the GEO database. After data standardization and data filtering, all cells were subjected to unsupervised clustering to obtain 3 different clusters, which were visualized after dimensionality reduction through T-SNE, and 10 differential marker genes were analyzed and screened, which can be clustered in different clusters. Gene set enrichment analysis found that CDRG was significantly associated with immune regulation and immune response, and 278 NSCLC cell differentiation related genes (CDRG) were identified.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eOur study identified NSCLC cells with distinct differentiation characteristics based on single-cell sequencing data from GEO, emphasizing the important role of cell differentiation in predicting the clinical outcome of NSCLC patients and their potential response to immunotherapy.\u003c/p\u003e","manuscriptTitle":"Dissecting heterogeneity and immune cell populations in non-small cell lung cancer by single cell RNA sequencing","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-07-21 13:32:43","doi":"10.21203/rs.3.rs-3174725/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2cd8a2c9-3e2f-49e0-a188-dab6719ad889","owner":[],"postedDate":"July 21st, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2023-09-19T06:44:28+00:00","versionOfRecord":[],"versionCreatedAt":"2023-07-21 13:32:43","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3174725","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3174725","identity":"rs-3174725","version":["v1"]},"buildId":"_2-kVJe1T_tPrBINL-cwx","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.