Identifying critical genes of breast cancer and corresponding leading compounds of potential therapeutic targets

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Background: In 2020, there were 2.26 million new breast cancer cases, accounting for 24.5% of the total 9.23 million new cancer cases in women, far exceeding other cancer types in women. And for the death of cancer patients, there were 4.43 million female cancer deaths, among them, about 15.5% cancer deaths were caused by breast cancer. Breast cancer is the number one morbidity and mortality among women in the world, and breast cancer has seriously endangered the health and life of women around the world. Therefore, to address the growing public health problem of breast cancer, we must identify the critical genes and additional treatment targets of breast cancer. Methods: The Weighted Gene Co-Expression Network Analysis (WGCNA) was used to explore the hub genes of breast cancer patients. The regulation network of these hub genes was constructed with reanalyzing Chromatin Immunoprecipitation sequencing (Chip-seq) of the breast cancer cells. With the single-cell RNA sequencing and spatial transcriptome dataset of breast cancer patients, the hub gene expression abundance of each cell cluster and associates of the hub genes and immune cell was estimated. To find the genes that could be a prognosis factor or a potential treatment target, we conducted survival analysis based on each gene’s mRNA level and protein level. Finally, we used virtual screening of natural product molecules to find the leading compounds of our predicted target. Results: 128 hub genes were found in breast cancer patients. Among these, Squalene Epoxidase (SQLE) can be a potential drug target, 17 molecules were ranked the top and the ZINC263585481 small molecule was the most possible as a leading compound of SQLE. Conclusion: Our study provides a whole critical genes of the development of breast cancer and found amounts of leading compounds, which will facilitate the curing of breast cancer.
Full text 120,848 characters · extracted from preprint-html · click to expand
Identifying critical genes of breast cancer and corresponding leading compounds of potential therapeutic targets | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Identifying critical genes of breast cancer and corresponding leading compounds of potential therapeutic targets Xiaokai Fan, Xuan Yu, Liang Chen This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4835618/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 10 Dec, 2024 Read the published version in Molecular Diversity → Version 1 posted 4 You are reading this latest preprint version Abstract Background: In 2020, there were 2.26 million new breast cancer cases, accounting for 24.5% of the total 9.23 million new cancer cases in women, far exceeding other cancer types in women. And for the death of cancer patients, there were 4.43 million female cancer deaths, among them, about 15.5% cancer deaths were caused by breast cancer. Breast cancer is the number one morbidity and mortality among women in the world, and breast cancer has seriously endangered the health and life of women around the world. Therefore, to address the growing public health problem of breast cancer, we must identify the critical genes and additional treatment targets of breast cancer. Methods: The Weighted Gene Co-Expression Network Analysis (WGCNA) was used to explore the hub genes of breast cancer patients. The regulation network of these hub genes was constructed with reanalyzing Chromatin Immunoprecipitation sequencing (Chip-seq) of the breast cancer cells. With the single-cell RNA sequencing and spatial transcriptome dataset of breast cancer patients, the hub gene expression abundance of each cell cluster and associates of the hub genes and immune cell was estimated. To find the genes that could be a prognosis factor or a potential treatment target, we conducted survival analysis based on each gene’s mRNA level and protein level. Finally, we used virtual screening of natural product molecules to find the leading compounds of our predicted target. Results: 128 hub genes were found in breast cancer patients. Among these, Squalene Epoxidase (SQLE) can be a potential drug target, 17 molecules were ranked the top and the ZINC263585481 small molecule was the most possible as a leading compound of SQLE. Conclusion: Our study provides a whole critical genes of the development of breast cancer and found amounts of leading compounds, which will facilitate the curing of breast cancer. Breast cancer Hub genes Leading compounds Targeted therapy Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 INTRODUCTION Breast cancer is the most common type of malignant tumor in women. According to GLOBOCAN 2020, breast cancer is the world’s most common cancer and the fifth leading cause of death [ 1 ]. The number of breast cancer cases is expected to reach 4.4 million in 2070 [ 2 ]. Considering the significant public health challenge caused by breast cancer, it is urgent to discover novel prognostic biomarkers, potential treatment targets, and corresponding drugs. With the development of high throughput "omics" approaches, a large number of patients’ sequencing datasets are being generated, and these data provide sufficient information that can for tracking and diagnosing disease progression. WGCNA is a classical weighted gene co-expression network analysis to identify candidate biomarkers or therapeutic targets[ 3 , 4 ]. Using this method, we can calculate the probability of tumor occurrence with the expression of genes, then find the maximum-likelihood caused-tumor/normal-related genes. Because of the maturation of single-cell RNA sequencing technology, the genes of each cell of patients’ tumor tissues were resolved in a high-resolution way, we can acquire a higher level of information compared with the bulk RNA sequencing[ 5 – 8 ]. Virtual screening of drug molecules and its ADMET prediction based on machine learning methods are increasingly being used in drug discovery programs with a growing number of successful applications[ 9 – 11 ]. In this study, we analyzed the bulk RNA sequencing datasets of breast cancer patients of The Cancer Genome Atlas Program (TCGA) and a microarray breast cancer dataset, to find the critical genes in the occurrence and progression of breast cancer. By analyzing a single-cell RNA sequencing dataset of breast cancer, we calculated the correlation between our annotated immune cells and hub genes, then after survival and expression level analysis, we found that SQLE could be a potential treatment target for breast cancer. Therefore, finally, we virtually screened the candidate inhibitors of SQLE with a library of natural product molecules and predicted the ADMET properties of candidate compounds, the workflow of the current study is shown in Fig. 1 . RESULTS Identification of critical genes involved in breast cancer development and formation To investigate the vital genes involved in the development and formation of breast cancer, we used the TCGA-GDC breast cancer (BRCA) datasets (1092 tumor tissues and 113 para-normal tissues)[ 12 ] and the GSE29174 breast cancer datasets (150 tumor tissues and 11 normal tissues)[ 13 ]. With calculating differential genes in each dataset between normal tissues and tumor tissues, and then using a classical gene network analysis algorithm WGCNA (the soft threshold power of TCGA BRCA patients’ datasets was selected 3, the soft threshold power of GSE29174 cohort was selected 9 (Supplementary Fig. 1A-D))[ 14 ] to find the high positive or negative correlation genes with breast cancer in each dataset, then overlap the differential genes and the high correlation genes with breast cancer (Fig. 2 A). Finally, 128 overlapping genes were found (Fig. 2 B), which may play important roles in the development and formation of breast carcinoma. We found that most of these hub genes were on chromosome 17, and the GTSE1 , SRPX , MAOA , KIF4A , SLC7A3 , ITM2A , GPC3 , FHL1 , HMGB3 genes located on chromosome X (Fig. 2 C), which are possibly playing significant roles in the occurrence and development of breast cancer since that nearly 100% of the breast cancer patients are female. To further investigate the potential functions associated with the identified key genes, we performed GO biological function and KEGG pathways function enrichment analysis, and found that these genes' functions primarily focused on the mitotic cell cycle process, lipid storage, microtubule cytoskeleton organization involved in mitosis, regulation of protein kinase activity, and some other important physiological functions which are critical for tumor cell survival and proliferation (Fig. 2 D). The single-cell atlas of breast cancer patients and the correlation between the hub genes and immune cells To investigate the cell components and heterogeneity in breast cancer patients and the general relationship between the tumor-infiltrating lymphocyte and hub genes, we re-analyzed the cell composition of patients with different subtypes from Wu et al’s breast cancer single-cell RNA sequencing datasets. Following multiple quality control and filtering steps as mentioned in Material and Methods, a total of 10,064 cells were analyzed. With dimension reduction and the unsupervised uniform manifold approximation and projection (UMAP) clustering method, we separated these cells into 48 clusters using inferCNV and distinguished neoplastic from epithelial cells (Supplementary Fig. 2A-C). The cell clusters were then annotated with well-established cell marker genes from the CellMarker database[ 15 ]. We grouped those cells into four major clusters which include cancer cells, cancer stem cells, stromal cells (perivascular-like cells, cancer-associated fibroblasts, vascular smooth muscle cells, endothelial cells), and immune cells (T cells, B cells, plasmablasts, dendritic cells, myeloid monocyte, myeloid macrophage I, myeloid macrophage II) (Fig. 3 A, B). The infiltrations of lymphocytes are directly related to the disease progression and the effect of immune therapy[ 16 , 17 ]. We further investigated the population of tumor-infiltrating T cells and innate immune cells in breast cancer patients. We extracted the T cells and used the UMAP-clustering method to analyze the major subtypes of T cells in breast cancer patients. Our data revealed that these T cells and innate immune cells can be clustered into twelve subtypes. Based on their specific cell marker genes, we annotated these cells as IL7R + CD4 + T cell, FOXP3 + Treg, CCR7 + CD4 + T cell, ZFP36 + CD8 + T cell, CXCL13 + follicular helper T cell, IFNG + CD8 + T cell, IFIT1 + T cell, LAG3 + CD8 + T cell, AREG + NK, FCGR3A + NKT cell, MKI67 + T cell, MKI67 + cycling CD8 + T cell (Fig. 3 C, D). To better understand the relationships between the hub genes and tumor-infiltrating lymphocytes, we used single sample gene set enrichment analysis (ssGSEA) to estimate the scores of these eighteen immune cell infiltrations of TCGA BRCA patients’ tumor tissue, then used the Spearman rank coefficient method to calculate the correlation of these hub genes and immune cells (Fig. 3 E). The results showed that expression levels of PPP1R6B , ICAM2 , and ITM2A have a stronger correlation with the infiltration of IFNG + CD8 + T cell, LAG3 + CD8 + T cell, ZFP36 + CD8 + T cell. CD8 + T cells was reported to take the primary role in the surveillance and clearance of cancer cells[ 18 – 20 ], therefore, high correlation between these genes’ expression and infiltration of CD8 + T cell indicates that these genes may play important roles in enhancing CD8 + T infiltration in breast cancer patients. Interestingly, KIFC1 , BIRC5 , and TROAP have a stronger negative correlation with the CXCL13 + follicular helper T cell. Previous studies have showed that CXCL13 + follicular helper T cell facilitates the formation of the tertiary lymphoid structure in breast cancer patients’ tumor tissues, which enhances the response of the immune therapy[ 21 ]. Hence, low expression levels of KIFC1 , BIRC5 , and TROAP might enhance the infiltration of the CXCL13 + follicular helper T cell. The expression levels of ITM2A , ABCB1 and PPP1R6B positively correlate with infiltration of the CXCL13 + follicular helper T cell in the tumor tissues of breast cancer patients, indicating that high expression level of these genes should facilitate infiltration of the CXCL13 + follicular helper T cell to improve the situation of the tumor microenvironment, to prevent the prefoliation of tumor cells (Supplementary Fig. 3). The NK cells play important roles in the response of immune therapy[ 16 , 22 ], the results of correlation analysis showed that ITM2A , ICAM2 , and PPP1R16B have stronger positive correlations with FCGR3A + NKT cell and AREG + NK cell. It is well-known that the dendritic cells participate in the recognition of antigens of tumor cells to function in immune surveillance[ 23 ]. We found that MAF , ICAM2 , and PP1R16B have stronger correlations with the dendritic cells. High expression MKI67 in the T cells is always standing for exhausted T cells, these T cells always highly express the PD-1, CTLA-4, and TIMP3 proteins which are inhibitors of activation of the T cells to cause the immune escape of tumor cells[ 24 ]. The results showed that CRY2 , OGN , and ATP1A2 have stronger negative correlations with MKI67 cycling CD8 + T cell (Supplementary Fig. 4). These results indicate that high expression of these genes decreases infiltration of the MKI67 cycling CD8 + T cell and MKI67 + T cell, which possibly ameliorates activation of the T cells. Transcription factor-mediated regulation of the hub genes expression in breast cancer Transcriptional factors (TFs) regulation is an important part of controlling gene expression. The construction of a TFs-hub genes regulation network can deepen our understanding of the contributions of partial TFs for regulating the expression of these hub genes and allow us to develop appropriate methods to regulate the hub genes’ expression for treatment purposes. Hence, we integrated 2504 datasets of Chip-seq data of breast cancer cell lines in the CistomeDB database to build the regulation network (Fig. 4 A, Supplementary Fig. 5, Supplementary Material 1). The genes not expressed in the breast cancer cell line were not included in the construction of the TFs-hub genes network. By construction of the TFs-hub genes network, we discovered that transcription factors: ESR1, BRD4, CTCF, FOXA1, POLR2A, EP300, MYC, GATA3, RAD21, CEBPB, and PR strongly regulate many of the hub genes in this regulation network (Fig. 4 B), indicating that these genes play important roles in the formation of breast cancer and are at the core of regulating hub genes. SQLE is a potential pharmacological target for the breast cancer Given the importance of the hub genes in the formation of the breast cancer, we investigated whether the hub gene expression level can be used as a predictor of patient survival. We performed survival analysis on TCGA-BRCA and METABRIC datasets using Kaplan-Meier method. We discovered that seven genes SLAC7A3 , TF , HMGB3 , IL17B , UBE2C , SQLE , and CDO1 could be prognosis factors (Fig. 5A, B). Further, we found that SQLE and HMGB3 were specifically highly expressed in the breast cancer cell and cancer stem cell, but not other cell types (Fig. 5C). SQLE is the second rate-limiting enzyme in cholesterol biosynthesis, and mainly catalyzes the stereospecific conversion of the non-sterol intermediate squalene to 2,3(S)-oxidosqualene, which is the first oxygenation step in cholesterol biosynthesis[ 25 ]. Cholesterol is an important biomolecule that is required for the structural integrity of cell membranes, cell signaling transduction, and production of steroid hormones[ 26 ]. It is obvious that inhibition of SQLE results in the cholesterol biosynthesis pathway inhibition, which possibly influences the abnormal cholesterol metabolism in cancer cells. Cancer stem cells possess unique contributions to cancer cell survival and distant metastasis or relapse of the tumor, in early-stage breast cancer patients, high SQLE expression is associated with decreased distant metastasis-free survival [ 27 ]. And according to Katerina Rohlenova et.al.’s study of endothelial metabolic plasticity in pathological angiogenesis, SQLE could be a novel metabolic angiogenic target, which is critical to prevent or cure not only breast cancer but also metastasis cancers[ 7 ]. HMGB3 is a protein family member, possessing one or more high mobility group DNA-binding motifs, and has important functions in sustaining the stem cell populations. Considering the pocket of these two proteins, the activity pocket of SQLE protein exists as the catalytic site whereas the activity pocket of HMGB3 protein mainly functions in binding DNA, lacking a deep pocket. Based on comparison of characteristics of these two proteins, we selected SQLE as a final pharmacal target. The copy number variation analysis (CNV) also showed that SQLE with most of the DNA copy numbers compared with other hub genes (Fig. 5D). Using TCGA-BRCA protein expression data, we found that SQLE influences patient survival at the protein level with the Kaplan-Meier method (Fig. 5E). Comparing the SQLE immunohistochemical expression results between normal breast tissue with breast cancer tissue, SQLE highly expresses in the tumor tissues of breast cancer patients (Fig. 5F). A spatial transcriptome analysis also shows that SQLE was highly expressed in the cancer region of four breast cancer patients (CID4290, CID4565, CID4535, CID44977), these results are consistent with the analysis of the SQLE expression levels found in single-cell RNA sequencing (Supplementary Fig. 6A-D). These results indicated that SQLE can be a potential treatment target for breast cancer, developing drug-inhibitors or inactivators of SQLE could increase the chance of survival or extension of patient survival. To understand the important roles of SQLE in breast cancer development, we focused on the regulation of the SQLE expression and found that FAIRE, POLR2A, NELFE, and ESR1 could bind to the promoter region of SQLE and reach the goal of facilitating curing the breast cancer patients (Supplementary Fig. 7). Virtual natural product molecules screening for SQLE To find novel potential leading compounds of SQLE, we performed a virtual screening of the ZINC library of 241,208 natural product compounds and compare the binding energy with control compounds (NB-598, Cmpd-4″) which have currently been used as the SQLE inhibitors[ 28 ]. We found 17 molecules as potential leading compounds using the ligand-docking analysis, these top 17 molecules all possess stronger binding energy with SQLE compared with control SQLE inhibitors (Fig. 6 A, B). Among them, four molecules, ZINC263585481, ZINC263585264, ZINC263584775, and ZINC2134358 showed the highest binding energy (-15.5kcal/mol, -15.0kcal/mol, -14.9kcal/mol and − 14.9kcal/mol respectively). We also analyzed the connection between the ZINC263585481 and the ligand-binding pocket of SQLE, the ligand formed hydrogen bonds with the ALA-284, LEU-134, VAL-133, ASP-408, GLY-164, ILE-162, TYR-335 amino acids (Fig. 6 C). These amino acids are critical for the formation of the binding between ZINC263585481 and the pocket of SQLE protein. From the point of view of binding energy, the natural product ZINC263585481 has the potential to be a novel leading compound as the SQLE inhibitor. Predicting ADMET properties of screened molecules In drug development, the undesirable properties of a compound such as unreasonable pharmacokinetics and strong cytotoxic activity always cause the failure of the whole project [ 29 ]. Therefore, predicting the properties (especially absorption, distribution, metabolism, excretion, and toxicity (ADMET)) of candidate compounds is critical for the optimization of compound structures or choosing a appropriate candidate compounds[ 30 , 31 ]. To provide more characteristics of our screened candidate nature product molecules, we used ADMETlab2.0, which is a machine-learning platform for the predictions of pharmacokinetics and toxicity properties of candidate compounds and the control drugs[ 32 ] (NB-598, Cmpd-4″) (Supplementary Fig. 8–11). In the preclinical research phase of potential drugs, nearly 24% of tested drugs were discontinued due to cardiovascular side effects, and approximately 45% of drugs were withdrawn from the market due to cardiotoxic side effects[ 33 ]. The main cause of drug-induced cardiotoxicity is the result of blocking the rapid delayed rectifier current (IKr) of the heart, resulting in prolongation of the QT interval in the cardiac action potential duration, and then inducing torsades de pointes (TdP) which can be severe in sudden death cases. IKr is conducted by the Kv11.1 potassium channel encoded by the hERG gene and plays a crucial role in the entire action potential time course[ 34 ]. Early and effective prediction, evaluation, optimization, and avoidance of the inhibitory activity of drugs on hERG potassium channels can reduce the cost of and improve the success rate of drug development. The results showed that ZINC263585236, ZINC60091384, ZINC263585264, ZINC263586907, ZINC263586908, ZINC263584775, and ZINC263585481 were all less likely to produce hERG toxicity than NB-598 and Cmpd4''. NB-598 and Cmpd4'' inhibit the hERG protein by almost 100% in the prediction of the hERG toxicity, which may cause sudden cardiac death of patients in the clinical application (Supplementary Fig. 10). LogP or LogD properties of small molecules are two indicators used to evaluate the lipophilicity of and the permeability of small molecule compounds in vivo [ 30 ]. LogP refers to the compound that exists in a certain pH condition, the partition coefficient of dispersion of the compound in the organic phase, and the aqueous phase. The partition coefficient of LogP of a small molecule between 0 and 3 is considered the most suitable. LogD is a descriptor describing the lipophilicity of an ionizable compound at a given pH, and a small molecule LogD between 1 and 3 is considered as the most suitable[ 35 ]. The results from our prediction analysis showed that the LogP values of ZINC217842260, ZINC15707890, ZINC263585481, ZINC253414682, ZINC263585264, and ZINC20609761 are all between 0 and 3, indicating that these compounds have a benefit distribution coefficient of the organic phase and the aqueous phase, while the LogP of NB-598 and Cmpd4'' values ​​are greater than 6, showing that the control drugs have poor dissolution property. ZIN263585264, ZINC253414682, ZINC263586908, ZINC217842260, ZINC263586907, ZINC263585236, ZINC60081384, ZINC15707890, ZINC20609761, ZINC263584775 and ZINC12858112 at the pH7.4 have better distribution coefficient of both the organic phase and the aqueous phase, suggesting that these molecules can enter the blood circulation system and has good membrane permeability (Supplementary Fig. 10). LC50FM refers to the lowest concentration by determining a critical amount of the drug that causes 50% of fathead minnows to die after 96 hours[ 36 ]. Our prediction results showed that the lethality of known SQLE-targeted drugs is much higher than that of 17 potential leading compounds, and the LC50FM of ZINC20609761 is lower than that of other leading compounds. LC50DM refers to the minimum concentration of the drug that kills half of Daphnia magna after 48 hours of drug treatment[ 36 ]. The results showed that 15 small molecules were less lethal than NB-598 and Cmpd4'' (Supplementary Fig. 10). The properties of SR-p53 mean that the drug tested activates the p53 tumor suppressor protein to regulate DNA repair and cell autophagy, and inhibit the growth of tumor cells[ 37 ]. The results show these small molecules may not only inhibit SQLE activity but also activate p53 to further inhibit the proliferation of tumor cells (Supplementary Fig. 11). For other ADMET properties, these leading compounds and control drugs each have advantages and disadvantages, therefore in the next optimization of leading compounds’ structures are important for the practical applications. DISCUSSION Despite significant improvements in overall survival and quality of life for women with breast cancer, breast cancer remains the leading cause of cancer-related deaths[ 38 ]. Therefore, the development of new prognostic and therapeutic markers for breast cancer patients is necessary. With the finding of the hub genes of breast cancer, we can acquire a summary of the critical genes during the development and formation of breast cancer. This gene list will give benefit the future the discovery of the novel treatment target for breast cancer and deepen our understanding of the critical genes of breast cancer. Here, we from different points of view to prove that SQLE could be a potential treatment target for breast cancer patients. There have evidence shows that inhibition of SQLE affects the synthesis of sterols and cell membranes construction, and even cell growth[ 39 – 41 ]. Cholesterol is a unique lipid that is essential for membrane formation, cell proliferation, and cell differentiation. Studies have shown that inhibition of cholesterol synthesis at different steps leads to human cancer cell death in both in vitro and in vivo models[ 42 – 44 ]. In addition, cholesterol metabolism has been found to play an important role in cancer regulation and is considered a new therapeutic approach[ 44 , 45 ]. These evidences reinforce the validity of our findings that SQLE is a suitable treatment target for breast cancer patients. Using virtual screening of a natural products molecules’ library, we find that the binding energy top1 molecule ZINC263585481 is possibly a candidate inhibitor for the SQLE protein. In the future, according to the true bioactivity and other drug properties test based on the real biological experiments, optimize the structure of ZINC263585481 possibly bringing to patients more suitable drugs. In conclusion, this study provides information for subsequent personalized diagnosis and treatment of breast cancer. Conclusion Our work presented the most critical 128 genes of breast cancer and construct the regulation network of these 128 hub genes. Using multiple datasets and methods, finally we find that SQLE can be a potential treatment target of breast cancer. Then based on this finding, we apply molecular docking and ADMET predicting tools to screen suitable leading compounds of SQLE. In the end, we discovery that ZINC263585481 natural product molecule possible be as a suitable leading compound for target SQLE to cure breast cancer patients. MATERIALS AND METHODS Data Acquisition The RNA sequencing datasets of TCGA breast cancer patients were extracted from the University of California Santa Cruz (UCSC) Xena ( https://xena.ucsc.edu/ ) database, along with clinical information, and included 1211 tumor samples and 113 normal samples, after filtering the unpaired and repeat data, we finally acquire 1092 tumor samples. Sample sequencing and clinical information datasets from 1998 breast cancer patients of the METABRIC program were also obtained from cBioPortal ( https://www.cbioportal.org/ ). The microarray RNA sequencing datasets for GSE29174, which include 15 normal breast tissues and 150 breast tumor tissues, were obtained from the GEO database ( https://www.ncbi.nlm.nih.gov/geo ). The probes were converted into gene symbols by the annotation files of the GPL3676 platform and repetitive symbols were averaged. The proteomics datasets of TCGA breast cancer patients were acquired from CPTAC ( https://cptac-data-portal.georgetown.edu/datasets ), which include 129 tumor tissue samples with their corresponding clinical information. The single-cell RNA sequencing datasets of 26 breast cancer samples (GSE176078)[ 17 ] were obtained from the GEO database ( https://www.ncbi.nlm.nih.gov/geo ), and the corresponding three samples of space transcriptomics datasets were obtained from the Zenodo data repository ( https://doi.org/10.5281/zenodo.4739739 ). Differential expression gene (Mori, #67) analysis, weighted gene co-expression network analysis (WGCNA), and functional enrichment analysis for bulk RNA-seq dataset and microarray dataset of breast cancer DESeq2 and limma were used to examine gene expression differences between tumor tissues and para-normal tissues, as well as tumor tissues and normal tissues[ 46 , 47 ]. FDR value less than 0.05 and absolute log2 fold change greater than one are used to define differential genes. To identify gene correlations with tumor or normal phenotypes, we use the WGCNA classical gene network analysis algorithm. Absolute correlation values between genes and phenotype greater than 0.4 were considered to be a strong link between genes and breast cancer phenotype. To intersect the WGCNA results and the differential expression gene results of these two datasets, we get the hub genes of breast cancer. The gene ontology and KEGG pathway enrichment of hub genes were analyzed by Metascape ( https://metascape.org/gp/index.html#/main/step1 )[ 48 ], a web-based functional enrichment tool. The functions were considered significantly enriched when the adjusted p-value or FDR was minus 0.05. Virtual screening of natural product compounds and visualization of dock results The mol2 format of natural product molecules was acquired from the ZINC15 database ( https://zinc15.docking.org/ ). The molecules were separated into single files by the OpenBabel program and converted into PDBQT format molecules by a python script of Autodock MGLTools. Then, the SQLE structure was prepared using the prepare_recpetor4.py script. The docking box of SQLE was chosen by the inhibitor (NB-598, Cpmd-4’’) and FDA binding sites. The molecular docking was done by Autodock Vina1.1.2. Quality Control, Cell Cluster, and Cell Type Annotation of single-cell RNA sequencing dataset and spatial transcriptome dataset The Wu et al., single-cell RNA sequencing datasets of breast cancer included samples from 11 ER + , 5 HER2 + , and 10 TNBC patients, and we performed the quality control process with Seurat (version 4.0.1). We filtered single cells with less than 500 UMIs and mitochondrial gene percentages greater than 20% in a cell. Using the CCA package and Seurat's IntegrateData function to eliminate batch effects among patients, following that major cell groups were identified by employing the dimensionality reduction algorithm. Finally, we used FindAllMakers in Seurat to calculate the specific markers of each cell cluster, and the major cell types were annotated based on marker genes of cell types in the CellMaker database and our cell makers gene collections. The standard of quality control of the spatial transcriptome data and spot clusters was the same as the methods described above. Inferring neoplastic from breast epithelial cells by inferCNV The inferCNV was used for predicting the CNV of high expression EPCAM marker cell clusters, we run inferCNV by setting following the parameters: cutoff = 0.1, denoise = TRUE, HMM = TRUE, and we project results to the Seurat object, then visualize of results by UMAP plot. Analysis of the correlations between immune cells and hub genes The ssGSEA method was used to estimate the proportion of immune cells in the TCGA BRCA patients through the marker genes calculated by the FindAllMarkers function of Seurat. Then the correlation between immune cells and hub genes is then calculated using the Spearman rank coefficient method. Gene Regulatory Network Analysis We used CistromeDB ( http://cistrome.org/db/#/ ) to collect the Chip-seq results of multiple breast cancer cell lines to build the hub genes regulatory network. We filtered out normal breast cell lines to construct only contained breast cancer cell lines regulation network of our mined hub genes in Cytoscape software. Survival and other statistical analysis To perform survival analysis, we use the KM method for categorical variables to determine whether gene expression influences patient survival, log-rank tests were used to calculate the significance of differences in the overall survival of patients. Unless otherwise stated, p-values were two-side tested, and p-values < 0.05 were considered statistically significant. Abbreviations KM Kaplan-Meier CNV copy number variation analysis ER estrogen receptor PR progesterone receptor HER2 human epidermal growth factor receptor-2 ADMET absorption, distribution, metabolism, excretion, and toxicity OS overall survival DCs dendritic cells TCGA The Cancer Genome Atlas ssGSEA single sample gene set enrichment analysis. Declarations Acknowledgments We sincerely thank Alexander Swarbrick (Garvan Institute of Medical Research) for sharing their single-cell RNA sequencing data in the GEO database and spatial transcriptome data in the Zenodo database, and the all participants of TCGA program. Author s’ contributions L.C., X.Y, and XKF. conceived and designed the study and drafted the manuscript. XKF, X.Y performed the analysis of the data. The author(s) read and approved the final manuscript. Funding Not applicable. Availability of data and materials Our re-analysis single-cell RNA data were deposited at the National Center for Biotechnology Information (NCBI) GEO and are accessible through the accession number GSE176078. The spatial transcriptome data of our used were deposited at the Zenodo database, the website is (https://doi.org/10.5281/zenodo.4739739). The RNA-sequencing, proteomics and whole genome sequencing data of breast cancer patients were acquired from TCGA program. The METABRIC data and microarray data of breast cancer were separately download from cBioPortal and GEO through the accession number GSE29174. The code of this paper is provided with reasonable request. Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Competing interests The authors declare no potential conflicts of interest. References Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F: Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA: A Cancer Journal for Clinicians 2021, 71: 209-249. Soerjomataram I, Bray F: Planning for tomorrow: global cancer incidence and the role of prevention 2020–2070. Nature Reviews Clinical Oncology 2021, 18: 663-672. Fuller TF, Ghazalpour A, Aten JE, Drake TA, Lusis AJ, Horvath S: Weighted gene coexpression network analysis strategies applied to mouse weight. Mamm Genome 2007, 18: 463-472. Presson AP, Sobel EM, Papp JC, Suarez CJ, Whistler T, Rajeevan MS, Vernon SD, Horvath S: Integrated weighted gene co-expression network analysis with an application to chronic fatigue syndrome. BMC Syst Biol 2008, 2: 95. Picelli S: Single-cell RNA-sequencing: The future of genome biology is now. RNA Biol 2017, 14: 637-650. Ding S, Chen X, Shen K: Single-cell RNA sequencing in breast cancer: Understanding tumor heterogeneity and paving roads to individualized therapy. Cancer Commun (Lond) 2020, 40: 329-344. Rohlenova K, Goveia J, García-Caballero M, Subramanian A, Kalucka J, Treps L, Falkenberg KD, de Rooij L, Zheng Y, Lin L, et al: Single-Cell RNA Sequencing Maps Endothelial Metabolic Plasticity in Pathological Angiogenesis. Cell Metab 2020, 31: 862-877.e814. Zhang Y, Wang D, Peng M, Tang L, Ouyang J, Xiong F, Guo C, Tang Y, Zhou Y, Liao Q, et al: Single-cell RNA sequencing in cancer research. J Exp Clin Cancer Res 2021, 40: 81. Kitchen DB, Decornez H, Furr JR, Bajorath J: Docking and scoring in virtual screening for drug discovery: methods and applications. Nat Rev Drug Discov 2004, 3: 935-949. López-Vallejo F, Caulfield T, Martínez-Mayorga K, Giulianotti MA, Nefzi A, Houghten RA, Medina-Franco JL: Integrating virtual screening and combinatorial chemistry for accelerated drug discovery. Comb Chem High Throughput Screen 2011, 14: 475-487. Polgár T, Keseru GM: Integration of virtual and high throughput screening in lead discovery settings. Comb Chem High Throughput Screen 2011, 14: 889-897. Weinstein JN, Collisson EA, Mills GB, Shaw KR, Ozenberger BA, Ellrott K, Shmulevich I, Sander C, Stuart JM: The Cancer Genome Atlas Pan-Cancer analysis project. Nat Genet 2013, 45: 1113-1120. Farazi TA, Horlings HM, Ten Hoeve JJ, Mihailovic A, Halfwerk H, Morozov P, Brown M, Hafner M, Reyal F, van Kouwenhove M, et al: MicroRNA sequence and expression analysis in breast tumors by deep sequencing. Cancer Res 2011, 71: 4443-4453. Langfelder P, Horvath S: WGCNA: an R package for weighted correlation network analysis. BMC Bioinformatics 2008, 9: 559. Zhang X, Lan Y, Xu J, Quan F, Zhao E, Deng C, Luo T, Xu L, Liao G, Yan M, et al: CellMarker: a manually curated resource of cell markers in human and mouse. Nucleic Acids Res 2019, 47: D721-d728. Liu S, Galat V, Galat Y, Lee YKA, Wainwright D, Wu J: NK cell-based cancer immunotherapy: from basic biology to clinical development. J Hematol Oncol 2021, 14: 7. Wu SZ, Al-Eryani G, Roden DL, Junankar S, Harvey K, Andersson A, Thennavan A, Wang C, Torpy JR, Bartonicek N, et al: A single-cell and spatially resolved atlas of human breast cancers. Nat Genet 2021, 53: 1334-1347. Ali HR, Provenzano E, Dawson SJ, Blows FM, Liu B, Shah M, Earl HM, Poole CJ, Hiller L, Dunn JA, et al: Association between CD8+ T-cell infiltration and breast cancer survival in 12,439 patients. Ann Oncol 2014, 25: 1536-1543. Afolabi LO, Adeshakin AO, Sani MM, Bi J, Wan X: Genetic reprogramming for NK cell cancer immunotherapy with CRISPR/Cas9. Immunology 2019, 158: 63-69. Bruni D, Angell HK, Galon J: The immune contexture and Immunoscore in cancer prognosis and therapeutic efficacy. Nat Rev Cancer 2020, 20: 662-680. Niogret J, Berger H, Rebe C, Mary R, Ballot E, Truntzer C, Thibaudin M, Derangère V, Hibos C, Hampe L, et al: Follicular helper-T cells restore CD8(+)-dependent antitumor immunity and anti-PD-L1/PD-1 efficacy. J Immunother Cancer 2021, 9 . Afolabi LO, Bi J, Li X, Adeshakin AO, Adeshakin FO, Wu H, Yan D, Chen L, Wan X: Synergistic Tumor Cytolysis by NK Cells in Combination With a Pan-HDAC Inhibitor, Panobinostat. Front Immunol 2021, 12: 701671. Wculek SK, Cueto FJ, Mujal AM, Melero I, Krummel MF, Sancho D: Dendritic cells in cancer immunology and immunotherapy. Nat Rev Immunol 2020, 20: 7-24. Wu SY, Liao P, Yan LY, Zhao QY, Xie ZY, Dong J, Sun HT: Correlation of MKI67 with prognosis, immune infiltration, and T cell exhaustion in hepatocellular carcinoma. BMC Gastroenterol 2021, 21: 416. Brown AJ, Chua NK, Yan N: The shape of human squalene epoxidase expands the arsenal against cancer. Nat Commun 2019, 10: 888. Hryniewicz-Jankowska A, Augoff K, Sikorski AF: The role of cholesterol and cholesterol-driven membrane raft domains in prostate cancer. Exp Biol Med (Maywood) 2019, 244: 1053-1061. Qin Y, Hou Y, Liu S, Zhu P, Wan X, Zhao M, Peng M, Zeng H, Li Q, Jin T, et al: A Novel Long Non-Coding RNA lnc030 Maintains Breast Cancer Stem Cell Stemness by Stabilizing SQLE mRNA and Increasing Cholesterol Synthesis. Adv Sci (Weinh) 2021, 8: 2002232. Nagaraja R, Olaharski A, Narayanaswamy R, Mahoney C, Pirman D, Gross S, Roddy TP, Popovici-Muller J, Smolen GA, Silverman L: Preclinical toxicology profile of squalene epoxidase inhibitors. Toxicol Appl Pharmacol 2020, 401: 115103. Hou T, Wang J: Structure-ADME relationship: still a long way to go? Expert Opin Drug Metab Toxicol 2008, 4: 759-770. Bickerton GR, Paolini GV, Besnard J, Muresan S, Hopkins AL: Quantifying the chemical beauty of drugs. Nat Chem 2012, 4: 90-98. Cumming JG, Davis AM, Muresan S, Haeberlein M, Chen H: Chemical predictive modelling to improve compound quality. Nat Rev Drug Discov 2013, 12: 948-962. Xiong G, Wu Z, Yi J, Fu L, Yang Z, Hsieh C, Yin M, Zeng X, Wu C, Lu A, et al: ADMETlab 2.0: an integrated online platform for accurate and comprehensive predictions of ADMET properties. Nucleic Acids Res 2021, 49: W5-w14. Kalyaanamoorthy S, Barakat KH: Development of Safe Drugs: The hERG Challenge. Med Res Rev 2018, 38: 525-555. He S, Moutaoufik MT, Islam S, Persad A, Wu A, Aly KA, Fonge H, Babu M, Cayabyab FS: HERG channel and cancer: A mechanistic review of carcinogenic processes and therapeutic potential. Biochim Biophys Acta Rev Cancer 2020, 1873: 188355. Kah M, Brown CD: LogD: lipophilicity for ionisable compounds. Chemosphere 2008, 72: 1401-1408. Lipinski CA, Lombardo F, Dominy BW, Feeney PJ: Experimental and computational approaches to estimate solubility and permeability in drug discovery and development settings. Adv Drug Deliv Rev 2001, 46: 3-26. Hong B, van den Heuvel AP, Prabhu VV, Zhang S, El-Deiry WS: Targeting tumor suppressor p53 for cancer therapy: strategies, challenges and opportunities. Curr Drug Targets 2014, 15: 80-89. Akram M, Iqbal M, Daniyal M, Khan AU: Awareness and current knowledge of breast cancer. Biol Res 2017, 50: 33. Zhang HY, Li HM, Yu Z, Yu XY, Guo K: Expression and significance of squalene epoxidase in squamous lung cancerous tissues and pericarcinoma tissues. Thorac Cancer 2014, 5: 275-280. Qin Y, Zhang Y, Tang Q, Jin L, Chen Y: SQLE induces epithelial-to-mesenchymal transition by regulating of miR-133b in esophageal squamous cell carcinoma. Acta Biochim Biophys Sin (Shanghai) 2017, 49: 138-148. Cirmena G, Franceschelli P, Isnaldi E, Ferrando L, De Mariano M, Ballestrero A, Zoppoli G: Squalene epoxidase as a promising metabolic target in cancer treatment. Cancer Lett 2018, 425: 13-20. Cruz PM, Mo H, McConathy WJ, Sabnis N, Lacko AG: The role of cholesterol metabolism and cholesterol transport in carcinogenesis: a review of scientific findings, relevant to future cancer therapeutics. Front Pharmacol 2013, 4: 119. Stadler SC, Hacker U, Burkhardt R: Cholesterol metabolism and breast cancer. Curr Opin Lipidol 2016, 27: 200-201. Huang B, Song BL, Xu C: Cholesterol metabolism in cancer: mechanisms and therapeutic opportunities. Nat Metab 2020, 2: 132-141. Kopecka J, Godel M, Riganti C: Cholesterol metabolism: At the cross road between cancer cells and immune environment. Int J Biochem Cell Biol 2020, 129: 105876. Love MI, Huber W, Anders S: Moderated estimation of fold change and dispersion for RNA-seq data with DESeq2. Genome Biol 2014, 15: 550. Ritchie ME, Phipson B, Wu D, Hu Y, Law CW, Shi W, Smyth GK: limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res 2015, 43: e47. Zhou Y, Zhou B, Pache L, Chang M, Khodabakhshi AH, Tanaseichuk O, Benner C, Chanda SK: Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat Commun 2019, 10: 1523. Additional Declarations No competing interests reported. Supplementary Files SupplementaryFigures.docx Cite Share Download PDF Status: Published Journal Publication published 10 Dec, 2024 Read the published version in Molecular Diversity → Version 1 posted Editorial decision: Revision requested 05 Aug, 2024 Editor assigned by journal 02 Aug, 2024 Submission checks completed at journal 02 Aug, 2024 First submitted to journal 31 Jul, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4835618","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":336087425,"identity":"6af3e23e-5cd0-41b6-b217-9c7d8a130d35","order_by":0,"name":"Xiaokai Fan","email":"","orcid":"","institution":"Fudan University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xiaokai","middleName":"","lastName":"Fan","suffix":""},{"id":336087426,"identity":"57c457d8-c7fb-46d3-81e4-5fa8397d09fc","order_by":1,"name":"Xuan Yu","email":"","orcid":"","institution":"Chinese Academy of Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xuan","middleName":"","lastName":"Yu","suffix":""},{"id":336087429,"identity":"a35c5976-46ca-466f-99c5-13acc7bad8a4","order_by":2,"name":"Liang Chen","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5klEQVRIiWNgGAWjYBACPmYIncDAwHwARBIGbAgtbAlA0oAILQxwLTwg5cRoYWd+9vALg10ev3TP5w8Pav4w8LcfYPxcgNdhbObGMgzJxZJzzm6TSDhmwCBxJoFZegZ+v5hJSzAcSNxwI3cbQwIb0GE3gII8eLWwfwNr2X8j5/GHhH8GDPKEtfCYSX4A2SKRwyCR2GbAYECEljJpBoPkxBk30swkEvuMeQzPJDZL49PCz398m+SPCrvE/hnJjz/++CYnJ3f88MHP+LSAADMPUmwAFTM2ENAAVPKDoJJRMApGwSgY0QAAbmA/3xiUUqMAAAAASUVORK5CYII=","orcid":"","institution":"Chinese Academy of Sciences","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Liang","middleName":"","lastName":"Chen","suffix":""}],"badges":[],"createdAt":"2024-07-31 12:44:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4835618/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4835618/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s11030-024-11035-z","type":"published","date":"2024-12-10T15:56:57+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":64146232,"identity":"71cdc3f5-009b-4aad-921b-9a76d83dbab2","added_by":"auto","created_at":"2024-09-08 19:56:56","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":158503,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSchematic of the workflow analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe “datasets” part shows that our used data include bulk RNA sequencing data, microarray data, proteomics data, single-cell RNA sequencing data, spatial transcriptome data, and natural product molecules from the ZINC database. The “results” part shows that together using these data, we mined the potential hub genes of breast cancer with the WGCNA network analysis algorithm and build the transcription factor regulation network of these genes, following this we also analyzed the correlation between these genes and tumor infiltering lymphocytes, then after survival analysis and expression levels of hub gene we find that SQLE can be as a potential treatment target, virtually screen inhibitors of SQLE, we find 17 candidate compounds and we also use a machine learning-based platform to predict ADMET of these leading compounds and also compare the ADMET properties of leading compounds with control drug (NB-598, Cmpd-4″) of SQLE.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-4835618/v1/7dcc85a5d9290522b40bb84c.png"},{"id":64146233,"identity":"29c60f68-6568-46f4-a085-c06a6b915c92","added_by":"auto","created_at":"2024-09-08 19:56:56","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1880894,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDifferential gene expression analysis and gene network analysis to find hub genes and functional enrichment analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) The workflow of the finding of the hub genes, first, we calculate the up-regulated genes and down-regulated genes of these two datasets and then use WGCNA to calculate the genes’ correlation with the phenotype, finally, we intersect the found genes to acquire hub genes. The standard of differential expression genes is |Log2FC|\u0026gt;1 and FDR \u0026lt; 0.05, the high correlation of the WGCNA results is an absolute correlation value greater than 0.4.\u003c/p\u003e\n\u003cp\u003e(B) The circles on the Venn diagram represent the number of hub genes finally found, of which 96 are positively correlated with normal phenotype and 32 highly positive with tumor phenotype.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-4835618/v1/044d3bdb9c7ac68c22f29a34.png"},{"id":64145205,"identity":"2bd1bcf1-bfab-496b-ad09-3a706347b2c2","added_by":"auto","created_at":"2024-09-08 19:48:56","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":2019660,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eExpression profiling of 100,064 single cells in breast cancer patients\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) The major cell types in our reanalyzed breast cancer patients’ datasets. The cell types were annotated by the canonical marker genes derived from the CellMarker database and our collection.\u003c/p\u003e\n\u003cp\u003e(B) The specifically selected marker genes depict major classes of cells in breast cancer, cancer cell or cancer stem cell (\u003cem\u003eCD24\u003c/em\u003e), myeloid monocyte (\u003cem\u003eLYZ\u003c/em\u003e), cancer-associated fibroblasts (\u003cem\u003eDCN\u003c/em\u003e), epithelial cell (\u003cem\u003eEPCAM\u003c/em\u003e), perivascular-like cell (\u003cem\u003eACTA2\u003c/em\u003e), plasmablast (\u003cem\u003eIGHG1\u003c/em\u003e), T cell (\u003cem\u003eCD3E\u003c/em\u003e), endothelial cell (\u003cem\u003eRAMP2\u003c/em\u003e), B cell (\u003cem\u003eMS4A1\u003c/em\u003e), myeloid macrophage I (\u003cem\u003eCD209\u003c/em\u003e), myeloid macrophage II (\u003cem\u003eCD14\u003c/em\u003e), dendritic cell (\u003cem\u003eCLEC4C\u003c/em\u003e). The expression levels of these genes were Log-normalized.\u003c/p\u003e\n\u003cp\u003e(C) The major cell types of T cells and innate immune cells in our reanalyzed breast cancer patients’ datasets, the UMAP plot shows the cell groups.\u003c/p\u003e\n\u003cp\u003e(D) The specifically selected marker genes depict major T classes of cells in breast cancer and the selected genes were Log-normalized.\u003c/p\u003e\n\u003cp\u003e(E) Heatmap showing correlations of hub genes with inferred immune cells using the TGCA BRCA cohort, Spearman correlation test was applied. The p-value of the correlation test was greater than 0.05 and the correlation equals zero, the pane wascolored white.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-4835618/v1/7941f63d66e4731fce45df99.png"},{"id":64145202,"identity":"820d7617-a5b1-41f6-93a4-6999f9081a25","added_by":"auto","created_at":"2024-09-08 19:48:56","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":299992,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eRegulation of selected hub genes\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) The regulation network shows that RP scores greater than 0.9 of transpiration factors regulate hub genes, the hub genes were manually annotated into twelve clusters based on their major function. RP, Regulatory Potential.\u003c/p\u003e\n\u003cp\u003e(B) Bar plot shows that the number of hub genes is regulated by the transcription factors, the top six genes were ESR1, BRD4, CTCF, FOXA1, POLR2A, and EP300.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-4835618/v1/b209870b0aa18b5e4ca260ca.png"},{"id":64147110,"identity":"8c4f5012-c4e9-4e2e-bf59-08e21bbe6f97","added_by":"auto","created_at":"2024-09-08 20:04:56","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":759329,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eScreen potential prognosis genes and treatment target\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) KM survival of SQLE with TCGA BRCA and METABRIC BRCA data, the p-value was tested by log-rank test, and the p-value\u0026lt;0.05 was considered significant. BRCA, breast cancer.\u003c/p\u003e\n\u003cp\u003e(C) Prognosis gene expression representation using multiple groups of violin plots.\u003c/p\u003e\n\u003cp\u003e(D) Copy number variations of hub genes in the TCGA BRCA and METABRIC cohorts, and perform this analysis in the cbioportal database. The genes were copied in the patient, and the line representing this patient were colored red. No alteration genes in the patient, and the line representing this patient was colored grey. And deep deletion in a patient, the line represents this patient was colored blue. This heatmap shows the top20 CNV genes of 128 hub genes.\u003c/p\u003e\n\u003cp\u003e(E) Survival analysis of SQLE with CPTAC-TCGA protein data using KM method and log-rank test.\u003c/p\u003e\n\u003cp\u003e(F) Compare SQLE protein expression IHC slice of normal tissue and tumor tissue, the IHC slice images were acquired from the HPA database. IHC, Immunohistochemistry; HPA, Human Protein Atlas.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-4835618/v1/55d4df9fd630488a1353edc3.png"},{"id":64146230,"identity":"41ba76e8-fde1-4ace-8375-8e184f586528","added_by":"auto","created_at":"2024-09-08 19:56:56","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":2359409,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eScreening candidate inhibitors of SQLE\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) Virtually screening the inhibitors of SQLE with a 241,208 natural molecules library from the ZINC database. The top seventeen molecules based on binding energy and control molecules (NB-598, Cmpd-4″) are represented in the right bar plot.\u003c/p\u003e\n\u003cp\u003e(B) The molecular structures of the top seventeen molecules and control molecules (NB-598, Cmpd-4″).\u003c/p\u003e\n\u003cp\u003e(C) The binding site view of the top one molecule ZINC263585481. The ligand was shown in spheres in the left plot, and in the right plot the main residues that are part of the pocket interact with this ligand were shown in sticks, the stick structure colored purple represents the ligand (ZINC263585481), around the ligand inside 5 angstroms of receptor were colored.\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-4835618/v1/f123aa196ace7183e71aff32.png"},{"id":71552341,"identity":"4f7f6d8e-6d95-4e84-9792-c86600a6f5e9","added_by":"auto","created_at":"2024-12-16 16:04:45","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":8269745,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4835618/v1/c647b050-cf62-4b8e-a149-dcd24b7084d7.pdf"},{"id":64145199,"identity":"6ebe07c6-7343-49ef-ad99-669d35e85ce9","added_by":"auto","created_at":"2024-09-08 19:48:56","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":2707078,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigures.docx","url":"https://assets-eu.researchsquare.com/files/rs-4835618/v1/b446655f0d153f594f17b1be.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Identifying critical genes of breast cancer and corresponding leading compounds of potential therapeutic targets","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eBreast cancer is the most common type of malignant tumor in women. According to GLOBOCAN 2020, breast cancer is the world\u0026rsquo;s most common cancer and the fifth leading cause of death [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The number of breast cancer cases is expected to reach 4.4\u0026nbsp;million in 2070 [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Considering the significant public health challenge caused by breast cancer, it is urgent to discover novel prognostic biomarkers, potential treatment targets, and corresponding drugs.\u003c/p\u003e \u003cp\u003eWith the development of high throughput \"omics\" approaches, a large number of patients\u0026rsquo; sequencing datasets are being generated, and these data provide sufficient information that can for tracking and diagnosing disease progression. WGCNA is a classical weighted gene co-expression network analysis to identify candidate biomarkers or therapeutic targets[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Using this method, we can calculate the probability of tumor occurrence with the expression of genes, then find the maximum-likelihood caused-tumor/normal-related genes.\u003c/p\u003e \u003cp\u003eBecause of the maturation of single-cell RNA sequencing technology, the genes of each cell of patients\u0026rsquo; tumor tissues were resolved in a high-resolution way, we can acquire a higher level of information compared with the bulk RNA sequencing[\u003cspan additionalcitationids=\"CR6 CR7\" citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Virtual screening of drug molecules and its ADMET prediction based on machine learning methods are increasingly being used in drug discovery programs with a growing number of successful applications[\u003cspan additionalcitationids=\"CR10\" citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. In this study, we analyzed the bulk RNA sequencing datasets of breast cancer patients of The Cancer Genome Atlas Program (TCGA) and a microarray breast cancer dataset, to find the critical genes in the occurrence and progression of breast cancer. By analyzing a single-cell RNA sequencing dataset of breast cancer, we calculated the correlation between our annotated immune cells and hub genes, then after survival and expression level analysis, we found that SQLE could be a potential treatment target for breast cancer. Therefore, finally, we virtually screened the candidate inhibitors of SQLE with a library of natural product molecules and predicted the ADMET properties of candidate compounds, the workflow of the current study is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e"},{"header":"RESULTS","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eIdentification of critical genes involved in breast cancer development and formation\u003c/h2\u003e \u003cp\u003eTo investigate the vital genes involved in the development and formation of breast cancer, we used the TCGA-GDC breast cancer (BRCA) datasets (1092 tumor tissues and 113 para-normal tissues)[\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e] and the GSE29174 breast cancer datasets (150 tumor tissues and 11 normal tissues)[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. With calculating differential genes in each dataset between normal tissues and tumor tissues, and then using a classical gene network analysis algorithm WGCNA (the soft threshold power of TCGA BRCA patients\u0026rsquo; datasets was selected 3, the soft threshold power of GSE29174 cohort was selected 9 (Supplementary Fig.\u0026nbsp;1A-D))[\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] to find the high positive or negative correlation genes with breast cancer in each dataset, then overlap the differential genes and the high correlation genes with breast cancer (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA). Finally, 128 overlapping genes were found (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB), which may play important roles in the development and formation of breast carcinoma. We found that most of these hub genes were on chromosome 17, and the \u003cem\u003eGTSE1\u003c/em\u003e, \u003cem\u003eSRPX\u003c/em\u003e, \u003cem\u003eMAOA\u003c/em\u003e, \u003cem\u003eKIF4A\u003c/em\u003e, \u003cem\u003eSLC7A3\u003c/em\u003e, \u003cem\u003eITM2A\u003c/em\u003e, \u003cem\u003eGPC3\u003c/em\u003e, \u003cem\u003eFHL1\u003c/em\u003e, \u003cem\u003eHMGB3\u003c/em\u003e genes located on chromosome X (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC), which are possibly playing significant roles in the occurrence and development of breast cancer since that nearly 100% of the breast cancer patients are female.\u003c/p\u003e \u003cp\u003eTo further investigate the potential functions associated with the identified key genes, we performed GO biological function and KEGG pathways function enrichment analysis, and found that these genes' functions primarily focused on the mitotic cell cycle process, lipid storage, microtubule cytoskeleton organization involved in mitosis, regulation of protein kinase activity, and some other important physiological functions which are critical for tumor cell survival and proliferation (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD).\u003c/p\u003e \u003cp\u003e \u003cb\u003eThe single-cell atlas of breast cancer patients and the correlation between the hub genes and immune cells\u003c/b\u003e \u003c/p\u003e \u003cp\u003eTo investigate the cell components and heterogeneity in breast cancer patients and the general relationship between the tumor-infiltrating lymphocyte and hub genes, we re-analyzed the cell composition of patients with different subtypes from Wu et al\u0026rsquo;s breast cancer single-cell RNA sequencing datasets. Following multiple quality control and filtering steps as mentioned in Material and Methods, a total of 10,064 cells were analyzed. With dimension reduction and the unsupervised uniform manifold approximation and projection (UMAP) clustering method, we separated these cells into 48 clusters using inferCNV and distinguished neoplastic from epithelial cells (Supplementary Fig.\u0026nbsp;2A-C). The cell clusters were then annotated with well-established cell marker genes from the CellMarker database[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. We grouped those cells into four major clusters which include cancer cells, cancer stem cells, stromal cells (perivascular-like cells, cancer-associated fibroblasts, vascular smooth muscle cells, endothelial cells), and immune cells (T cells, B cells, plasmablasts, dendritic cells, myeloid monocyte, myeloid macrophage I, myeloid macrophage II) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA, B).\u003c/p\u003e \u003cp\u003eThe infiltrations of lymphocytes are directly related to the disease progression and the effect of immune therapy[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. We further investigated the population of tumor-infiltrating T cells and innate immune cells in breast cancer patients. We extracted the T cells and used the UMAP-clustering method to analyze the major subtypes of T cells in breast cancer patients. Our data revealed that these T cells and innate immune cells can be clustered into twelve subtypes. Based on their specific cell marker genes, we annotated these cells as \u003cem\u003eIL7R\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e CD4\u003csup\u003e+\u003c/sup\u003e T cell, \u003cem\u003eFOXP3\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e Treg, \u003cem\u003eCCR7\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e CD4\u003csup\u003e+\u003c/sup\u003e T cell, \u003cem\u003eZFP36\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e CD8\u003csup\u003e+\u003c/sup\u003e T cell, \u003cem\u003eCXCL13\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e follicular helper T cell, \u003cem\u003eIFNG\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e CD8\u003csup\u003e+\u003c/sup\u003e T cell, \u003cem\u003eIFIT1\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e T cell, \u003cem\u003eLAG3\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e CD8\u003csup\u003e+\u003c/sup\u003e T cell, \u003cem\u003eAREG\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e NK, \u003cem\u003eFCGR3A\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e NKT cell, \u003cem\u003eMKI67\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e T cell, \u003cem\u003eMKI67\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e cycling CD8\u003csup\u003e+\u003c/sup\u003e T cell (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC, D).\u003c/p\u003e \u003cp\u003eTo better understand the relationships between the hub genes and tumor-infiltrating lymphocytes, we used single sample gene set enrichment analysis (ssGSEA) to estimate the scores of these eighteen immune cell infiltrations of TCGA BRCA patients\u0026rsquo; tumor tissue, then used the Spearman rank coefficient method to calculate the correlation of these hub genes and immune cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eE).\u003c/p\u003e \u003cp\u003eThe results showed that expression levels of \u003cem\u003ePPP1R6B\u003c/em\u003e, \u003cem\u003eICAM2\u003c/em\u003e, and \u003cem\u003eITM2A\u003c/em\u003e have a stronger correlation with the infiltration of \u003cem\u003eIFNG\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e CD8\u003csup\u003e+\u003c/sup\u003e T cell, \u003cem\u003eLAG3\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e CD8\u003csup\u003e+\u003c/sup\u003e T cell, \u003cem\u003eZFP36\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e CD8\u003csup\u003e+\u003c/sup\u003e T cell. CD8\u003csup\u003e+\u003c/sup\u003e T cells was reported to take the primary role in the surveillance and clearance of cancer cells[\u003cspan additionalcitationids=\"CR19\" citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], therefore, high correlation between these genes\u0026rsquo; expression and infiltration of CD8\u003csup\u003e+\u003c/sup\u003e T cell indicates that these genes may play important roles in enhancing CD8\u003csup\u003e+\u003c/sup\u003e T infiltration in breast cancer patients. Interestingly, \u003cem\u003eKIFC1\u003c/em\u003e, \u003cem\u003eBIRC5\u003c/em\u003e, and \u003cem\u003eTROAP\u003c/em\u003e have a stronger negative correlation with the \u003cem\u003eCXCL13\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e follicular helper T cell. Previous studies have showed that \u003cem\u003eCXCL13\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e follicular helper T cell facilitates the formation of the tertiary lymphoid structure in breast cancer patients\u0026rsquo; tumor tissues, which enhances the response of the immune therapy[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Hence, low expression levels of \u003cem\u003eKIFC1\u003c/em\u003e, \u003cem\u003eBIRC5\u003c/em\u003e, and \u003cem\u003eTROAP\u003c/em\u003e might enhance the infiltration of the \u003cem\u003eCXCL13\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e follicular helper T cell. The expression levels of \u003cem\u003eITM2A\u003c/em\u003e, \u003cem\u003eABCB1\u003c/em\u003e and \u003cem\u003ePPP1R6B\u003c/em\u003e positively correlate with infiltration of the \u003cem\u003eCXCL13\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e follicular helper T cell in the tumor tissues of breast cancer patients, indicating that high expression level of these genes should facilitate infiltration of the \u003cem\u003eCXCL13\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e follicular helper T cell to improve the situation of the tumor microenvironment, to prevent the prefoliation of tumor cells (Supplementary Fig.\u0026nbsp;3). The NK cells play important roles in the response of immune therapy[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], the results of correlation analysis showed that \u003cem\u003eITM2A\u003c/em\u003e, \u003cem\u003eICAM2\u003c/em\u003e, and \u003cem\u003ePPP1R16B\u003c/em\u003e have stronger positive correlations with \u003cem\u003eFCGR3A\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e NKT cell and \u003cem\u003eAREG\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e NK cell. It is well-known that the dendritic cells participate in the recognition of antigens of tumor cells to function in immune surveillance[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. We found that \u003cem\u003eMAF\u003c/em\u003e, \u003cem\u003eICAM2\u003c/em\u003e, and \u003cem\u003ePP1R16B\u003c/em\u003e have stronger correlations with the dendritic cells. High expression \u003cem\u003eMKI67\u003c/em\u003e in the T cells is always standing for exhausted T cells, these T cells always highly express the PD-1, CTLA-4, and TIMP3 proteins which are inhibitors of activation of the T cells to cause the immune escape of tumor cells[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. The results showed that \u003cem\u003eCRY2\u003c/em\u003e, \u003cem\u003eOGN\u003c/em\u003e, and \u003cem\u003eATP1A2\u003c/em\u003e have stronger negative correlations with \u003cem\u003eMKI67\u003c/em\u003e cycling CD8\u003csup\u003e+\u003c/sup\u003e T cell (Supplementary Fig.\u0026nbsp;4). These results indicate that high expression of these genes decreases infiltration of the \u003cem\u003eMKI67\u003c/em\u003e cycling CD8\u003csup\u003e+\u003c/sup\u003e T cell and \u003cem\u003eMKI67\u003c/em\u003e\u003csup\u003e+\u003c/sup\u003e T cell, which possibly ameliorates activation of the T cells.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eTranscription factor-mediated regulation of the hub genes expression in breast cancer\u003c/h2\u003e \u003cp\u003eTranscriptional factors (TFs) regulation is an important part of controlling gene expression. The construction of a TFs-hub genes regulation network can deepen our understanding of the contributions of partial TFs for regulating the expression of these hub genes and allow us to develop appropriate methods to regulate the hub genes\u0026rsquo; expression for treatment purposes. Hence, we integrated 2504 datasets of Chip-seq data of breast cancer cell lines in the CistomeDB database to build the regulation network (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA, Supplementary Fig.\u0026nbsp;5, Supplementary Material 1). The genes not expressed in the breast cancer cell line were not included in the construction of the TFs-hub genes network. By construction of the TFs-hub genes network, we discovered that transcription factors: ESR1, BRD4, CTCF, FOXA1, POLR2A, EP300, MYC, GATA3, RAD21, CEBPB, and PR strongly regulate many of the hub genes in this regulation network (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB), indicating that these genes play important roles in the formation of breast cancer and are at the core of regulating hub genes.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eSQLE is a potential pharmacological target for the breast cancer\u003c/h2\u003e \u003cp\u003eGiven the importance of the hub genes in the formation of the breast cancer, we investigated whether the hub gene expression level can be used as a predictor of patient survival. We performed survival analysis on TCGA-BRCA and METABRIC datasets using Kaplan-Meier method. We discovered that seven genes \u003cem\u003eSLAC7A3\u003c/em\u003e, \u003cem\u003eTF\u003c/em\u003e, \u003cem\u003eHMGB3\u003c/em\u003e, \u003cem\u003eIL17B\u003c/em\u003e, \u003cem\u003eUBE2C\u003c/em\u003e, \u003cem\u003eSQLE\u003c/em\u003e, and \u003cem\u003eCDO1\u003c/em\u003e could be prognosis factors (Fig.\u0026nbsp;5A, B). Further, we found that \u003cem\u003eSQLE\u003c/em\u003e and \u003cem\u003eHMGB3\u003c/em\u003e were specifically highly expressed in the breast cancer cell and cancer stem cell, but not other cell types (Fig.\u0026nbsp;5C). SQLE is the second rate-limiting enzyme in cholesterol biosynthesis, and mainly catalyzes the stereospecific conversion of the non-sterol intermediate squalene to 2,3(S)-oxidosqualene, which is the first oxygenation step in cholesterol biosynthesis[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Cholesterol is an important biomolecule that is required for the structural integrity of cell membranes, cell signaling transduction, and production of steroid hormones[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. It is obvious that inhibition of SQLE results in the cholesterol biosynthesis pathway inhibition, which possibly influences the abnormal cholesterol metabolism in cancer cells. Cancer stem cells possess unique contributions to cancer cell survival and distant metastasis or relapse of the tumor, in early-stage breast cancer patients, high SQLE expression is associated with decreased distant metastasis-free survival [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. And according to Katerina Rohlenova et.al.\u0026rsquo;s study of endothelial metabolic plasticity in pathological angiogenesis, SQLE could be a novel metabolic angiogenic target, which is critical to prevent or cure not only breast cancer but also metastasis cancers[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. HMGB3 is a protein family member, possessing one or more high mobility group DNA-binding motifs, and has important functions in sustaining the stem cell populations. Considering the pocket of these two proteins, the activity pocket of SQLE protein exists as the catalytic site whereas the activity pocket of HMGB3 protein mainly functions in binding DNA, lacking a deep pocket. Based on comparison of characteristics of these two proteins, we selected SQLE as a final pharmacal target. The copy number variation analysis (CNV) also showed that \u003cem\u003eSQLE\u003c/em\u003e with most of the DNA copy numbers compared with other hub genes (Fig.\u0026nbsp;5D). Using TCGA-BRCA protein expression data, we found that SQLE influences patient survival at the protein level with the Kaplan-Meier method (Fig.\u0026nbsp;5E). Comparing the SQLE immunohistochemical expression results between normal breast tissue with breast cancer tissue, SQLE highly expresses in the tumor tissues of breast cancer patients (Fig.\u0026nbsp;5F). A spatial transcriptome analysis also shows that \u003cem\u003eSQLE\u003c/em\u003e was highly expressed in the cancer region of four breast cancer patients (CID4290, CID4565, CID4535, CID44977), these results are consistent with the analysis of the \u003cem\u003eSQLE\u003c/em\u003e expression levels found in single-cell RNA sequencing (Supplementary Fig.\u0026nbsp;6A-D). These results indicated that SQLE can be a potential treatment target for breast cancer, developing drug-inhibitors or inactivators of SQLE could increase the chance of survival or extension of patient survival.\u003c/p\u003e \u003cp\u003eTo understand the important roles of SQLE in breast cancer development, we focused on the regulation of the \u003cem\u003eSQLE\u003c/em\u003e expression and found that FAIRE, POLR2A, NELFE, and ESR1 could bind to the promoter region of \u003cem\u003eSQLE\u003c/em\u003e and reach the goal of facilitating curing the breast cancer patients (Supplementary Fig.\u0026nbsp;7).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eVirtual natural product molecules screening for SQLE\u003c/h2\u003e \u003cp\u003eTo find novel potential leading compounds of SQLE, we performed a virtual screening of the ZINC library of 241,208 natural product compounds and compare the binding energy with control compounds (NB-598, Cmpd-4\u0026Prime;) which have currently been used as the SQLE inhibitors[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. We found 17 molecules as potential leading compounds using the ligand-docking analysis, these top 17 molecules all possess stronger binding energy with SQLE compared with control SQLE inhibitors (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e6\u003c/span\u003eA, B). Among them, four molecules, ZINC263585481, ZINC263585264, ZINC263584775, and ZINC2134358 showed the highest binding energy (-15.5kcal/mol, -15.0kcal/mol, -14.9kcal/mol and \u0026minus;\u0026thinsp;14.9kcal/mol respectively). We also analyzed the connection between the ZINC263585481 and the ligand-binding pocket of SQLE, the ligand formed hydrogen bonds with the ALA-284, LEU-134, VAL-133, ASP-408, GLY-164, ILE-162, TYR-335 amino acids (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e6\u003c/span\u003eC). These amino acids are critical for the formation of the binding between ZINC263585481 and the pocket of SQLE protein. From the point of view of binding energy, the natural product ZINC263585481 has the potential to be a novel leading compound as the SQLE inhibitor.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003ePredicting ADMET properties of screened molecules\u003c/h2\u003e \u003cp\u003eIn drug development, the undesirable properties of a compound such as unreasonable pharmacokinetics and strong cytotoxic activity always cause the failure of the whole project [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. Therefore, predicting the properties (especially absorption, distribution, metabolism, excretion, and toxicity (ADMET)) of candidate compounds is critical for the optimization of compound structures or choosing a appropriate candidate compounds[\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eTo provide more characteristics of our screened candidate nature product molecules, we used ADMETlab2.0, which is a machine-learning platform for the predictions of pharmacokinetics and toxicity properties of candidate compounds and the control drugs[\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e] (NB-598, Cmpd-4\u0026Prime;) (Supplementary Fig.\u0026nbsp;8\u0026ndash;11).\u003c/p\u003e \u003cp\u003eIn the preclinical research phase of potential drugs, nearly 24% of tested drugs were discontinued due to cardiovascular side effects, and approximately 45% of drugs were withdrawn from the market due to cardiotoxic side effects[\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. The main cause of drug-induced cardiotoxicity is the result of blocking the rapid delayed rectifier current (IKr) of the heart, resulting in prolongation of the QT interval in the cardiac action potential duration, and then inducing torsades de pointes (TdP) which can be severe in sudden death cases. IKr is conducted by the Kv11.1 potassium channel encoded by the hERG gene and plays a crucial role in the entire action potential time course[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. Early and effective prediction, evaluation, optimization, and avoidance of the inhibitory activity of drugs on hERG potassium channels can reduce the cost of and improve the success rate of drug development. The results showed that ZINC263585236, ZINC60091384, ZINC263585264, ZINC263586907, ZINC263586908, ZINC263584775, and ZINC263585481 were all less likely to produce hERG toxicity than NB-598 and Cmpd4''. NB-598 and Cmpd4'' inhibit the hERG protein by almost 100% in the prediction of the hERG toxicity, which may cause sudden cardiac death of patients in the clinical application (Supplementary Fig.\u0026nbsp;10).\u003c/p\u003e \u003cp\u003eLogP or LogD properties of small molecules are two indicators used to evaluate the lipophilicity of and the permeability of small molecule compounds \u003cem\u003ein vivo\u003c/em\u003e[\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. LogP refers to the compound that exists in a certain pH condition, the partition coefficient of dispersion of the compound in the organic phase, and the aqueous phase. The partition coefficient of LogP of a small molecule between 0 and 3 is considered the most suitable. LogD is a descriptor describing the lipophilicity of an ionizable compound at a given pH, and a small molecule LogD between 1 and 3 is considered as the most suitable[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. The results from our prediction analysis showed that the LogP values of ZINC217842260, ZINC15707890, ZINC263585481, ZINC253414682, ZINC263585264, and ZINC20609761 are all between 0 and 3, indicating that these compounds have a benefit distribution coefficient of the organic phase and the aqueous phase, while the LogP of NB-598 and Cmpd4'' values ​​are greater than 6, showing that the control drugs have poor dissolution property. ZIN263585264, ZINC253414682, ZINC263586908, ZINC217842260, ZINC263586907, ZINC263585236, ZINC60081384, ZINC15707890, ZINC20609761, ZINC263584775 and ZINC12858112 at the pH7.4 have better distribution coefficient of both the organic phase and the aqueous phase, suggesting that these molecules can enter the blood circulation system and has good membrane permeability (Supplementary Fig.\u0026nbsp;10).\u003c/p\u003e \u003cp\u003eLC50FM refers to the lowest concentration by determining a critical amount of the drug that causes 50% of fathead minnows to die after 96 hours[\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. Our prediction results showed that the lethality of known SQLE-targeted drugs is much higher than that of 17 potential leading compounds, and the LC50FM of ZINC20609761 is lower than that of other leading compounds. LC50DM refers to the minimum concentration of the drug that kills half of Daphnia magna after 48 hours of drug treatment[\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. The results showed that 15 small molecules were less lethal than NB-598 and Cmpd4'' (Supplementary Fig.\u0026nbsp;10).\u003c/p\u003e \u003cp\u003eThe properties of SR-p53 mean that the drug tested activates the p53 tumor suppressor protein to regulate DNA repair and cell autophagy, and inhibit the growth of tumor cells[\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. The results show these small molecules may not only inhibit SQLE activity but also activate p53 to further inhibit the proliferation of tumor cells (Supplementary Fig.\u0026nbsp;11).\u003c/p\u003e \u003cp\u003eFor other ADMET properties, these leading compounds and control drugs each have advantages and disadvantages, therefore in the next optimization of leading compounds\u0026rsquo; structures are important for the practical applications.\u003c/p\u003e \u003c/div\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eDespite significant improvements in overall survival and quality of life for women with breast cancer, breast cancer remains the leading cause of cancer-related deaths[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. Therefore, the development of new prognostic and therapeutic markers for breast cancer patients is necessary. With the finding of the hub genes of breast cancer, we can acquire a summary of the critical genes during the development and formation of breast cancer. This gene list will give benefit the future the discovery of the novel treatment target for breast cancer and deepen our understanding of the critical genes of breast cancer. Here, we from different points of view to prove that SQLE could be a potential treatment target for breast cancer patients. There have evidence shows that inhibition of SQLE affects the synthesis of sterols and cell membranes construction, and even cell growth[\u003cspan additionalcitationids=\"CR40\" citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Cholesterol is a unique lipid that is essential for membrane formation, cell proliferation, and cell differentiation. Studies have shown that inhibition of cholesterol synthesis at different steps leads to human cancer cell death in both in vitro and in vivo models[\u003cspan additionalcitationids=\"CR43\" citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]. In addition, cholesterol metabolism has been found to play an important role in cancer regulation and is considered a new therapeutic approach[\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e, \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. These evidences reinforce the validity of our findings that SQLE is a suitable treatment target for breast cancer patients. Using virtual screening of a natural products molecules\u0026rsquo; library, we find that the binding energy top1 molecule ZINC263585481 is possibly a candidate inhibitor for the SQLE protein. In the future, according to the true bioactivity and other drug properties test based on the real biological experiments, optimize the structure of ZINC263585481 possibly bringing to patients more suitable drugs.\u003c/p\u003e \u003cp\u003eIn conclusion, this study provides information for subsequent personalized diagnosis and treatment of breast cancer.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eOur work presented the most critical 128 genes of breast cancer and construct the regulation network of these 128 hub genes. Using multiple datasets and methods, finally we find that SQLE can be a potential treatment target of breast cancer. Then based on this finding, we apply molecular docking and ADMET predicting tools to screen suitable leading compounds of SQLE. In the end, we discovery that ZINC263585481 natural product molecule possible be as a suitable leading compound for target SQLE to cure breast cancer patients.\u003c/p\u003e"},{"header":"MATERIALS AND METHODS","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eData Acquisition\u003c/h2\u003e \u003cp\u003eThe RNA sequencing datasets of TCGA breast cancer patients were extracted from the University of California Santa Cruz (UCSC) Xena (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://xena.ucsc.edu/\u003c/span\u003e\u003cspan address=\"https://xena.ucsc.edu/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) database, along with clinical information, and included 1211 tumor samples and 113 normal samples, after filtering the unpaired and repeat data, we finally acquire 1092 tumor samples. Sample sequencing and clinical information datasets from 1998 breast cancer patients of the METABRIC program were also obtained from cBioPortal (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.cbioportal.org/\u003c/span\u003e\u003cspan address=\"https://www.cbioportal.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The microarray RNA sequencing datasets for GSE29174, which include 15 normal breast tissues and 150 breast tumor tissues, were obtained from the GEO database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/geo\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/geo\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The probes were converted into gene symbols by the annotation files of the GPL3676 platform and repetitive symbols were averaged. The proteomics datasets of TCGA breast cancer patients were acquired from CPTAC (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://cptac-data-portal.georgetown.edu/datasets\u003c/span\u003e\u003cspan address=\"https://cptac-data-portal.georgetown.edu/datasets\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), which include 129 tumor tissue samples with their corresponding clinical information.\u003c/p\u003e \u003cp\u003eThe single-cell RNA sequencing datasets of 26 breast cancer samples (GSE176078)[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] were obtained from the GEO database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/geo\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/geo\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), and the corresponding three samples of space transcriptomics datasets were obtained from the Zenodo data repository (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/zenodo.4739739\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.4739739\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cb\u003eDifferential expression gene (Mori, #67) analysis, weighted gene co-expression network analysis (WGCNA), and functional enrichment analysis for bulk RNA-seq dataset and microarray dataset of breast cancer\u003c/b\u003e \u003c/p\u003e \u003cp\u003eDESeq2 and limma were used to examine gene expression differences between tumor tissues and para-normal tissues, as well as tumor tissues and normal tissues[\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e, \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e]. FDR value less than 0.05 and absolute log2 fold change greater than one are used to define differential genes. To identify gene correlations with tumor or normal phenotypes, we use the WGCNA classical gene network analysis algorithm. Absolute correlation values between genes and phenotype greater than 0.4 were considered to be a strong link between genes and breast cancer phenotype. To intersect the WGCNA results and the differential expression gene results of these two datasets, we get the hub genes of breast cancer.\u003c/p\u003e \u003cp\u003eThe gene ontology and KEGG pathway enrichment of hub genes were analyzed by Metascape (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://metascape.org/gp/index.html#/main/step1\u003c/span\u003e\u003cspan address=\"https://metascape.org/gp/index.html#/main/step1\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e)[\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e], a web-based functional enrichment tool. The functions were considered significantly enriched when the adjusted p-value or FDR was minus 0.05.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eVirtual screening of natural product compounds and visualization of dock results\u003c/h2\u003e \u003cp\u003eThe mol2 format of natural product molecules was acquired from the ZINC15 database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://zinc15.docking.org/\u003c/span\u003e\u003cspan address=\"https://zinc15.docking.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The molecules were separated into single files by the OpenBabel program and converted into PDBQT format molecules by a python script of Autodock MGLTools. Then, the SQLE structure was prepared using the prepare_recpetor4.py script. The docking box of SQLE was chosen by the inhibitor (NB-598, Cpmd-4\u0026rsquo;\u0026rsquo;) and FDA binding sites. The molecular docking was done by Autodock Vina1.1.2.\u003c/p\u003e \u003cp\u003e \u003cb\u003eQuality Control, Cell Cluster, and Cell Type Annotation of single-cell RNA sequencing dataset and spatial transcriptome dataset\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThe Wu et al., single-cell RNA sequencing datasets of breast cancer included samples from 11 ER\u003csup\u003e+\u003c/sup\u003e, 5 HER2\u003csup\u003e+\u003c/sup\u003e, and 10 TNBC patients, and we performed the quality control process with Seurat (version 4.0.1). We filtered single cells with less than 500 UMIs and mitochondrial gene percentages greater than 20% in a cell. Using the CCA package and Seurat's IntegrateData function to eliminate batch effects among patients, following that major cell groups were identified by employing the dimensionality reduction algorithm. Finally, we used FindAllMakers in Seurat to calculate the specific markers of each cell cluster, and the major cell types were annotated based on marker genes of cell types in the CellMaker database and our cell makers gene collections. The standard of quality control of the spatial transcriptome data and spot clusters was the same as the methods described above.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eInferring neoplastic from breast epithelial cells by inferCNV\u003c/h2\u003e \u003cp\u003eThe inferCNV was used for predicting the CNV of high expression EPCAM marker cell clusters, we run inferCNV by setting following the parameters: cutoff\u0026thinsp;=\u0026thinsp;0.1, denoise\u0026thinsp;=\u0026thinsp;TRUE, HMM\u0026thinsp;=\u0026thinsp;TRUE, and we project results to the Seurat object, then visualize of results by UMAP plot.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eAnalysis of the correlations between immune cells and hub genes\u003c/h2\u003e \u003cp\u003eThe ssGSEA method was used to estimate the proportion of immune cells in the TCGA BRCA patients through the marker genes calculated by the FindAllMarkers function of Seurat. Then the correlation between immune cells and hub genes is then calculated using the Spearman rank coefficient method.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eGene Regulatory Network Analysis\u003c/h2\u003e \u003cp\u003eWe used CistromeDB (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://cistrome.org/db/#/\u003c/span\u003e\u003cspan address=\"http://cistrome.org/db/#/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) to collect the Chip-seq results of multiple breast cancer cell lines to build the hub genes regulatory network. We filtered out normal breast cell lines to construct only contained breast cancer cell lines regulation network of our mined hub genes in Cytoscape software.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eSurvival and other statistical analysis\u003c/h2\u003e \u003cp\u003eTo perform survival analysis, we use the KM method for categorical variables to determine whether gene expression influences patient survival, log-rank tests were used to calculate the significance of differences in the overall survival of patients. Unless otherwise stated, p-values were two-side tested, and p-values\u0026thinsp;\u0026lt;\u0026thinsp;0.05 were considered statistically significant.\u003c/p\u003e \u003c/div\u003e "},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eKM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eKaplan-Meier\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCNV\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ecopy number variation analysis\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eER\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eestrogen receptor\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003ePR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eprogesterone receptor\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eHER2\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ehuman epidermal growth factor receptor-2\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eADMET\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eabsorption, distribution, metabolism, excretion, and toxicity\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eOS\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eoverall survival\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eDCs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003edendritic cells\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTCGA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eThe Cancer Genome Atlas\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003essGSEA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003esingle sample gene set enrichment analysis.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe sincerely thank Alexander Swarbrick (Garvan Institute of Medical Research) for sharing their single-cell RNA sequencing data in the GEO database and spatial transcriptome data in the Zenodo database, and the all participants of TCGA program.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor\u003c/strong\u003e\u003cstrong\u003es\u0026rsquo;\u003c/strong\u003e\u003cstrong\u003e\u0026nbsp;contributions\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eL.C.,\u0026nbsp;X.Y, and XKF. conceived and designed the study and drafted the manuscript. XKF,\u0026nbsp;X.Y\u0026nbsp;performed the analysis of the data. The author(s) read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eOur re-analysis single-cell RNA data were deposited at the National Center for Biotechnology Information (NCBI) GEO and are accessible through the accession number GSE176078. The spatial transcriptome data of our used were deposited at the Zenodo database, the website is (https://doi.org/10.5281/zenodo.4739739). The RNA-sequencing, proteomics and whole genome sequencing data of breast cancer patients were acquired from TCGA program. The METABRIC data and microarray data of breast cancer were separately download from cBioPortal and GEO through the accession number GSE29174. The code of this paper is provided with reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no potential conflicts of interest.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eSung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, Bray F: \u003cstrong\u003eGlobal Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries.\u003c/strong\u003e \u003cem\u003eCA: A Cancer Journal for Clinicians \u003c/em\u003e2021, \u003cstrong\u003e71:\u003c/strong\u003e209-249.\u003c/li\u003e\n\u003cli\u003eSoerjomataram I, Bray F: \u003cstrong\u003ePlanning for tomorrow: global cancer incidence and the role of prevention 2020\u0026ndash;2070.\u003c/strong\u003e \u003cem\u003eNature Reviews Clinical Oncology \u003c/em\u003e2021, \u003cstrong\u003e18:\u003c/strong\u003e663-672.\u003c/li\u003e\n\u003cli\u003eFuller TF, Ghazalpour A, Aten JE, Drake TA, Lusis AJ, Horvath S: \u003cstrong\u003eWeighted gene coexpression network analysis strategies applied to mouse weight.\u003c/strong\u003e \u003cem\u003eMamm Genome \u003c/em\u003e2007, \u003cstrong\u003e18:\u003c/strong\u003e463-472.\u003c/li\u003e\n\u003cli\u003ePresson AP, Sobel EM, Papp JC, Suarez CJ, Whistler T, Rajeevan MS, Vernon SD, Horvath S: \u003cstrong\u003eIntegrated weighted gene co-expression network analysis with an application to chronic fatigue syndrome.\u003c/strong\u003e \u003cem\u003eBMC Syst Biol \u003c/em\u003e2008, \u003cstrong\u003e2:\u003c/strong\u003e95.\u003c/li\u003e\n\u003cli\u003ePicelli S: \u003cstrong\u003eSingle-cell RNA-sequencing: The future of genome biology is now.\u003c/strong\u003e \u003cem\u003eRNA Biol \u003c/em\u003e2017, \u003cstrong\u003e14:\u003c/strong\u003e637-650.\u003c/li\u003e\n\u003cli\u003eDing S, Chen X, Shen K: \u003cstrong\u003eSingle-cell RNA sequencing in breast cancer: Understanding tumor heterogeneity and paving roads to individualized therapy.\u003c/strong\u003e \u003cem\u003eCancer Commun (Lond) \u003c/em\u003e2020, \u003cstrong\u003e40:\u003c/strong\u003e329-344.\u003c/li\u003e\n\u003cli\u003eRohlenova K, Goveia J, Garc\u0026iacute;a-Caballero M, Subramanian A, Kalucka J, Treps L, Falkenberg KD, de Rooij L, Zheng Y, Lin L, et al: \u003cstrong\u003eSingle-Cell RNA Sequencing Maps Endothelial Metabolic Plasticity in Pathological Angiogenesis.\u003c/strong\u003e \u003cem\u003eCell Metab \u003c/em\u003e2020, \u003cstrong\u003e31:\u003c/strong\u003e862-877.e814.\u003c/li\u003e\n\u003cli\u003eZhang Y, Wang D, Peng M, Tang L, Ouyang J, Xiong F, Guo C, Tang Y, Zhou Y, Liao Q, et al: \u003cstrong\u003eSingle-cell RNA sequencing in cancer research.\u003c/strong\u003e \u003cem\u003eJ Exp Clin Cancer Res \u003c/em\u003e2021, \u003cstrong\u003e40:\u003c/strong\u003e81.\u003c/li\u003e\n\u003cli\u003eKitchen DB, Decornez H, Furr JR, Bajorath J: \u003cstrong\u003eDocking and scoring in virtual screening for drug discovery: methods and applications.\u003c/strong\u003e \u003cem\u003eNat Rev Drug Discov \u003c/em\u003e2004, \u003cstrong\u003e3:\u003c/strong\u003e935-949.\u003c/li\u003e\n\u003cli\u003eL\u0026oacute;pez-Vallejo F, Caulfield T, Mart\u0026iacute;nez-Mayorga K, Giulianotti MA, Nefzi A, Houghten RA, Medina-Franco JL: \u003cstrong\u003eIntegrating virtual screening and combinatorial chemistry for accelerated drug discovery.\u003c/strong\u003e \u003cem\u003eComb Chem High Throughput Screen \u003c/em\u003e2011, \u003cstrong\u003e14:\u003c/strong\u003e475-487.\u003c/li\u003e\n\u003cli\u003ePolg\u0026aacute;r T, Keseru GM: \u003cstrong\u003eIntegration of virtual and high throughput screening in lead discovery settings.\u003c/strong\u003e \u003cem\u003eComb Chem High Throughput Screen \u003c/em\u003e2011, \u003cstrong\u003e14:\u003c/strong\u003e889-897.\u003c/li\u003e\n\u003cli\u003eWeinstein JN, Collisson EA, Mills GB, Shaw KR, Ozenberger BA, Ellrott K, Shmulevich I, Sander C, Stuart JM: \u003cstrong\u003eThe Cancer Genome Atlas Pan-Cancer analysis project.\u003c/strong\u003e \u003cem\u003eNat Genet \u003c/em\u003e2013, \u003cstrong\u003e45:\u003c/strong\u003e1113-1120.\u003c/li\u003e\n\u003cli\u003eFarazi TA, Horlings HM, Ten Hoeve JJ, Mihailovic A, Halfwerk H, Morozov P, Brown M, Hafner M, Reyal F, van Kouwenhove M, et al: \u003cstrong\u003eMicroRNA sequence and expression analysis in breast tumors by deep sequencing.\u003c/strong\u003e \u003cem\u003eCancer Res \u003c/em\u003e2011, \u003cstrong\u003e71:\u003c/strong\u003e4443-4453.\u003c/li\u003e\n\u003cli\u003eLangfelder P, Horvath S: \u003cstrong\u003eWGCNA: an R package for weighted correlation network analysis.\u003c/strong\u003e \u003cem\u003eBMC Bioinformatics \u003c/em\u003e2008, \u003cstrong\u003e9:\u003c/strong\u003e559.\u003c/li\u003e\n\u003cli\u003eZhang X, Lan Y, Xu J, Quan F, Zhao E, Deng C, Luo T, Xu L, Liao G, Yan M, et al: \u003cstrong\u003eCellMarker: a manually curated resource of cell markers in human and mouse.\u003c/strong\u003e \u003cem\u003eNucleic Acids Res \u003c/em\u003e2019, \u003cstrong\u003e47:\u003c/strong\u003eD721-d728.\u003c/li\u003e\n\u003cli\u003eLiu S, Galat V, Galat Y, Lee YKA, Wainwright D, Wu J: \u003cstrong\u003eNK cell-based cancer immunotherapy: from basic biology to clinical development.\u003c/strong\u003e \u003cem\u003eJ Hematol Oncol \u003c/em\u003e2021, \u003cstrong\u003e14:\u003c/strong\u003e7.\u003c/li\u003e\n\u003cli\u003eWu SZ, Al-Eryani G, Roden DL, Junankar S, Harvey K, Andersson A, Thennavan A, Wang C, Torpy JR, Bartonicek N, et al: \u003cstrong\u003eA single-cell and spatially resolved atlas of human breast cancers.\u003c/strong\u003e \u003cem\u003eNat Genet \u003c/em\u003e2021, \u003cstrong\u003e53:\u003c/strong\u003e1334-1347.\u003c/li\u003e\n\u003cli\u003eAli HR, Provenzano E, Dawson SJ, Blows FM, Liu B, Shah M, Earl HM, Poole CJ, Hiller L, Dunn JA, et al: \u003cstrong\u003eAssociation between CD8+ T-cell infiltration and breast cancer survival in 12,439 patients.\u003c/strong\u003e \u003cem\u003eAnn Oncol \u003c/em\u003e2014, \u003cstrong\u003e25:\u003c/strong\u003e1536-1543.\u003c/li\u003e\n\u003cli\u003eAfolabi LO, Adeshakin AO, Sani MM, Bi J, Wan X: \u003cstrong\u003eGenetic reprogramming for NK cell cancer immunotherapy with CRISPR/Cas9.\u003c/strong\u003e \u003cem\u003eImmunology \u003c/em\u003e2019, \u003cstrong\u003e158:\u003c/strong\u003e63-69.\u003c/li\u003e\n\u003cli\u003eBruni D, Angell HK, Galon J: \u003cstrong\u003eThe immune contexture and Immunoscore in cancer prognosis and therapeutic efficacy.\u003c/strong\u003e \u003cem\u003eNat Rev Cancer \u003c/em\u003e2020, \u003cstrong\u003e20:\u003c/strong\u003e662-680.\u003c/li\u003e\n\u003cli\u003eNiogret J, Berger H, Rebe C, Mary R, Ballot E, Truntzer C, Thibaudin M, Derang\u0026egrave;re V, Hibos C, Hampe L, et al: \u003cstrong\u003eFollicular helper-T cells restore CD8(+)-dependent antitumor immunity and anti-PD-L1/PD-1 efficacy.\u003c/strong\u003e \u003cem\u003eJ Immunother Cancer \u003c/em\u003e2021, \u003cstrong\u003e9\u003c/strong\u003e.\u003c/li\u003e\n\u003cli\u003eAfolabi LO, Bi J, Li X, Adeshakin AO, Adeshakin FO, Wu H, Yan D, Chen L, Wan X: \u003cstrong\u003eSynergistic Tumor Cytolysis by NK Cells in Combination With a Pan-HDAC Inhibitor, Panobinostat.\u003c/strong\u003e \u003cem\u003eFront Immunol \u003c/em\u003e2021, \u003cstrong\u003e12:\u003c/strong\u003e701671.\u003c/li\u003e\n\u003cli\u003eWculek SK, Cueto FJ, Mujal AM, Melero I, Krummel MF, Sancho D: \u003cstrong\u003eDendritic cells in cancer immunology and immunotherapy.\u003c/strong\u003e \u003cem\u003eNat Rev Immunol \u003c/em\u003e2020, \u003cstrong\u003e20:\u003c/strong\u003e7-24.\u003c/li\u003e\n\u003cli\u003eWu SY, Liao P, Yan LY, Zhao QY, Xie ZY, Dong J, Sun HT: \u003cstrong\u003eCorrelation of MKI67 with prognosis, immune infiltration, and T cell exhaustion in hepatocellular carcinoma.\u003c/strong\u003e \u003cem\u003eBMC Gastroenterol \u003c/em\u003e2021, \u003cstrong\u003e21:\u003c/strong\u003e416.\u003c/li\u003e\n\u003cli\u003eBrown AJ, Chua NK, Yan N: \u003cstrong\u003eThe shape of human squalene epoxidase expands the arsenal against cancer.\u003c/strong\u003e \u003cem\u003eNat Commun \u003c/em\u003e2019, \u003cstrong\u003e10:\u003c/strong\u003e888.\u003c/li\u003e\n\u003cli\u003eHryniewicz-Jankowska A, Augoff K, Sikorski AF: \u003cstrong\u003eThe role of cholesterol and cholesterol-driven membrane raft domains in prostate cancer.\u003c/strong\u003e \u003cem\u003eExp Biol Med (Maywood) \u003c/em\u003e2019, \u003cstrong\u003e244:\u003c/strong\u003e1053-1061.\u003c/li\u003e\n\u003cli\u003eQin Y, Hou Y, Liu S, Zhu P, Wan X, Zhao M, Peng M, Zeng H, Li Q, Jin T, et al: \u003cstrong\u003eA Novel Long Non-Coding RNA lnc030 Maintains Breast Cancer Stem Cell Stemness by Stabilizing SQLE mRNA and Increasing Cholesterol Synthesis.\u003c/strong\u003e \u003cem\u003eAdv Sci (Weinh) \u003c/em\u003e2021, \u003cstrong\u003e8:\u003c/strong\u003e2002232.\u003c/li\u003e\n\u003cli\u003eNagaraja R, Olaharski A, Narayanaswamy R, Mahoney C, Pirman D, Gross S, Roddy TP, Popovici-Muller J, Smolen GA, Silverman L: \u003cstrong\u003ePreclinical toxicology profile of squalene epoxidase inhibitors.\u003c/strong\u003e \u003cem\u003eToxicol Appl Pharmacol \u003c/em\u003e2020, \u003cstrong\u003e401:\u003c/strong\u003e115103.\u003c/li\u003e\n\u003cli\u003eHou T, Wang J: \u003cstrong\u003eStructure-ADME relationship: still a long way to go?\u003c/strong\u003e \u003cem\u003eExpert Opin Drug Metab Toxicol \u003c/em\u003e2008, \u003cstrong\u003e4:\u003c/strong\u003e759-770.\u003c/li\u003e\n\u003cli\u003eBickerton GR, Paolini GV, Besnard J, Muresan S, Hopkins AL: \u003cstrong\u003eQuantifying the chemical beauty of drugs.\u003c/strong\u003e \u003cem\u003eNat Chem \u003c/em\u003e2012, \u003cstrong\u003e4:\u003c/strong\u003e90-98.\u003c/li\u003e\n\u003cli\u003eCumming JG, Davis AM, Muresan S, Haeberlein M, Chen H: \u003cstrong\u003eChemical predictive modelling to improve compound quality.\u003c/strong\u003e \u003cem\u003eNat Rev Drug Discov \u003c/em\u003e2013, \u003cstrong\u003e12:\u003c/strong\u003e948-962.\u003c/li\u003e\n\u003cli\u003eXiong G, Wu Z, Yi J, Fu L, Yang Z, Hsieh C, Yin M, Zeng X, Wu C, Lu A, et al: \u003cstrong\u003eADMETlab 2.0: an integrated online platform for accurate and comprehensive predictions of ADMET properties.\u003c/strong\u003e \u003cem\u003eNucleic Acids Res \u003c/em\u003e2021, \u003cstrong\u003e49:\u003c/strong\u003eW5-w14.\u003c/li\u003e\n\u003cli\u003eKalyaanamoorthy S, Barakat KH: \u003cstrong\u003eDevelopment of Safe Drugs: The hERG Challenge.\u003c/strong\u003e \u003cem\u003eMed Res Rev \u003c/em\u003e2018, \u003cstrong\u003e38:\u003c/strong\u003e525-555.\u003c/li\u003e\n\u003cli\u003eHe S, Moutaoufik MT, Islam S, Persad A, Wu A, Aly KA, Fonge H, Babu M, Cayabyab FS: \u003cstrong\u003eHERG channel and cancer: A mechanistic review of carcinogenic processes and therapeutic potential.\u003c/strong\u003e \u003cem\u003eBiochim Biophys Acta Rev Cancer \u003c/em\u003e2020, \u003cstrong\u003e1873:\u003c/strong\u003e188355.\u003c/li\u003e\n\u003cli\u003eKah M, Brown CD: \u003cstrong\u003eLogD: lipophilicity for ionisable compounds.\u003c/strong\u003e \u003cem\u003eChemosphere \u003c/em\u003e2008, \u003cstrong\u003e72:\u003c/strong\u003e1401-1408.\u003c/li\u003e\n\u003cli\u003eLipinski CA, Lombardo F, Dominy BW, Feeney PJ: \u003cstrong\u003eExperimental and computational approaches to estimate solubility and permeability in drug discovery and development settings.\u003c/strong\u003e \u003cem\u003eAdv Drug Deliv Rev \u003c/em\u003e2001, \u003cstrong\u003e46:\u003c/strong\u003e3-26.\u003c/li\u003e\n\u003cli\u003eHong B, van den Heuvel AP, Prabhu VV, Zhang S, El-Deiry WS: \u003cstrong\u003eTargeting tumor suppressor p53 for cancer therapy: strategies, challenges and opportunities.\u003c/strong\u003e \u003cem\u003eCurr Drug Targets \u003c/em\u003e2014, \u003cstrong\u003e15:\u003c/strong\u003e80-89.\u003c/li\u003e\n\u003cli\u003eAkram M, Iqbal M, Daniyal M, Khan AU: \u003cstrong\u003eAwareness and current knowledge of breast cancer.\u003c/strong\u003e \u003cem\u003eBiol Res \u003c/em\u003e2017, \u003cstrong\u003e50:\u003c/strong\u003e33.\u003c/li\u003e\n\u003cli\u003eZhang HY, Li HM, Yu Z, Yu XY, Guo K: \u003cstrong\u003eExpression and significance of squalene epoxidase in squamous lung cancerous tissues and pericarcinoma tissues.\u003c/strong\u003e \u003cem\u003eThorac Cancer \u003c/em\u003e2014, \u003cstrong\u003e5:\u003c/strong\u003e275-280.\u003c/li\u003e\n\u003cli\u003eQin Y, Zhang Y, Tang Q, Jin L, Chen Y: \u003cstrong\u003eSQLE induces epithelial-to-mesenchymal transition by regulating of miR-133b in esophageal squamous cell carcinoma.\u003c/strong\u003e \u003cem\u003eActa Biochim Biophys Sin (Shanghai) \u003c/em\u003e2017, \u003cstrong\u003e49:\u003c/strong\u003e138-148.\u003c/li\u003e\n\u003cli\u003eCirmena G, Franceschelli P, Isnaldi E, Ferrando L, De Mariano M, Ballestrero A, Zoppoli G: \u003cstrong\u003eSqualene epoxidase as a promising metabolic target in cancer treatment.\u003c/strong\u003e \u003cem\u003eCancer Lett \u003c/em\u003e2018, \u003cstrong\u003e425:\u003c/strong\u003e13-20.\u003c/li\u003e\n\u003cli\u003eCruz PM, Mo H, McConathy WJ, Sabnis N, Lacko AG: \u003cstrong\u003eThe role of cholesterol metabolism and cholesterol transport in carcinogenesis: a review of scientific findings, relevant to future cancer therapeutics.\u003c/strong\u003e \u003cem\u003eFront Pharmacol \u003c/em\u003e2013, \u003cstrong\u003e4:\u003c/strong\u003e119.\u003c/li\u003e\n\u003cli\u003eStadler SC, Hacker U, Burkhardt R: \u003cstrong\u003eCholesterol metabolism and breast cancer.\u003c/strong\u003e \u003cem\u003eCurr Opin Lipidol \u003c/em\u003e2016, \u003cstrong\u003e27:\u003c/strong\u003e200-201.\u003c/li\u003e\n\u003cli\u003eHuang B, Song BL, Xu C: \u003cstrong\u003eCholesterol metabolism in cancer: mechanisms and therapeutic opportunities.\u003c/strong\u003e \u003cem\u003eNat Metab \u003c/em\u003e2020, \u003cstrong\u003e2:\u003c/strong\u003e132-141.\u003c/li\u003e\n\u003cli\u003eKopecka J, Godel M, Riganti C: \u003cstrong\u003eCholesterol metabolism: At the cross road between cancer cells and immune environment.\u003c/strong\u003e \u003cem\u003eInt J Biochem Cell Biol \u003c/em\u003e2020, \u003cstrong\u003e129:\u003c/strong\u003e105876.\u003c/li\u003e\n\u003cli\u003eLove MI, Huber W, Anders S: \u003cstrong\u003eModerated estimation of fold change and dispersion for RNA-seq data with DESeq2.\u003c/strong\u003e \u003cem\u003eGenome Biol \u003c/em\u003e2014, \u003cstrong\u003e15:\u003c/strong\u003e550.\u003c/li\u003e\n\u003cli\u003eRitchie ME, Phipson B, Wu D, Hu Y, Law CW, Shi W, Smyth GK: \u003cstrong\u003elimma powers differential expression analyses for RNA-sequencing and microarray studies.\u003c/strong\u003e \u003cem\u003eNucleic Acids Res \u003c/em\u003e2015, \u003cstrong\u003e43:\u003c/strong\u003ee47.\u003c/li\u003e\n\u003cli\u003eZhou Y, Zhou B, Pache L, Chang M, Khodabakhshi AH, Tanaseichuk O, Benner C, Chanda SK: \u003cstrong\u003eMetascape provides a biologist-oriented resource for the analysis of systems-level datasets.\u003c/strong\u003e \u003cem\u003eNat Commun \u003c/em\u003e2019, \u003cstrong\u003e10:\u003c/strong\u003e1523.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"molecular-diversity","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"modi","sideBox":"Learn more about [Molecular Diversity](http://link.springer.com/journal/11030)","snPcode":"11030","submissionUrl":"https://submission.nature.com/new-submission/11030/3","title":"Molecular Diversity","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Breast cancer, Hub genes, Leading compounds, Targeted therapy","lastPublishedDoi":"10.21203/rs.3.rs-4835618/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4835618/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground:\u003c/h2\u003e \u003cp\u003eIn 2020, there were 2.26\u0026nbsp;million new breast cancer cases, accounting for 24.5% of the total 9.23\u0026nbsp;million new cancer cases in women, far exceeding other cancer types in women. And for the death of cancer patients, there were 4.43\u0026nbsp;million female cancer deaths, among them, about 15.5% cancer deaths were caused by breast cancer. Breast cancer is the number one morbidity and mortality among women in the world, and breast cancer has seriously endangered the health and life of women around the world. Therefore, to address the growing public health problem of breast cancer, we must identify the critical genes and additional treatment targets of breast cancer.\u003c/p\u003e\u003ch2\u003eMethods:\u003c/h2\u003e \u003cp\u003eThe Weighted Gene Co-Expression Network Analysis (WGCNA) was used to explore the hub genes of breast cancer patients. The regulation network of these hub genes was constructed with reanalyzing Chromatin Immunoprecipitation sequencing (Chip-seq) of the breast cancer cells. With the single-cell RNA sequencing and spatial transcriptome dataset of breast cancer patients, the hub gene expression abundance of each cell cluster and associates of the hub genes and immune cell was estimated. To find the genes that could be a prognosis factor or a potential treatment target, we conducted survival analysis based on each gene\u0026rsquo;s mRNA level and protein level. Finally, we used virtual screening of natural product molecules to find the leading compounds of our predicted target.\u003c/p\u003e\u003ch2\u003eResults:\u003c/h2\u003e \u003cp\u003e128 hub genes were found in breast cancer patients. Among these, Squalene Epoxidase (SQLE) can be a potential drug target, 17 molecules were ranked the top and the ZINC263585481 small molecule was the most possible as a leading compound of SQLE.\u003c/p\u003e\u003ch2\u003eConclusion:\u003c/h2\u003e \u003cp\u003eOur study provides a whole critical genes of the development of breast cancer and found amounts of leading compounds, which will facilitate the curing of breast cancer.\u003c/p\u003e","manuscriptTitle":"Identifying critical genes of breast cancer and corresponding leading compounds of potential therapeutic targets","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-09-08 19:48:51","doi":"10.21203/rs.3.rs-4835618/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-08-05T10:23:36+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-08-02T15:35:02+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-08-02T07:02:45+00:00","index":"","fulltext":""},{"type":"submitted","content":"Molecular Diversity","date":"2024-07-31T12:41:19+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"molecular-diversity","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"modi","sideBox":"Learn more about [Molecular Diversity](http://link.springer.com/journal/11030)","snPcode":"11030","submissionUrl":"https://submission.nature.com/new-submission/11030/3","title":"Molecular Diversity","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"0ca027bd-506a-4946-80e7-73cde1e33c31","owner":[],"postedDate":"September 8th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-12-16T15:58:36+00:00","versionOfRecord":{"articleIdentity":"rs-4835618","link":"https://doi.org/10.1007/s11030-024-11035-z","journal":{"identity":"molecular-diversity","isVorOnly":false,"title":"Molecular Diversity"},"publishedOn":"2024-12-10 15:56:57","publishedOnDateReadable":"December 10th, 2024"},"versionCreatedAt":"2024-09-08 19:48:51","video":"","vorDoi":"10.1007/s11030-024-11035-z","vorDoiUrl":"https://doi.org/10.1007/s11030-024-11035-z","workflowStages":[]},"version":"v1","identity":"rs-4835618","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4835618","identity":"rs-4835618","version":["v1"]},"buildId":"cTy_lsJlmDsVRNrSptgXS","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-08-12T06:43:03.944938+00:00
License: CC-BY-4.0