LASSO Based Analysis for Prediction of Prognostic Signature Genes Associated with Breast Cancer | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article LASSO Based Analysis for Prediction of Prognostic Signature Genes Associated with Breast Cancer Souvik Guha, Soumita Seth, Tapas Bhadra, Anirban Mukhopadhyay, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4363199/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background Cancer is a genetic disease, where gene alterations play a significant role in the disease onset and pathogenesis. Analysis of the underlying gene interaction pathways could reveal new biomarkers and could also potentially help in the development of targeted drugs for therapeutics. Microarray techniques have emerged as powerful tools capable of simultaneously measuring the expression levels of thousands of genes, making them invaluable in cancer biology research. However, the processing of the resultant datasets poses significant challenges due to their high dimensionality. Also, feature extraction becomes essential to discern the crucial features within these extensive datasets. To mitigate these difficulties advanced computational techniques like Machine Learning (ML) could be instrumental. LASSO- regression-based classification is an advanced ML technique that can help in feature selection by evaluating individual parameters like genes. Methods This study focuses on uncovering key prognostic genes for breast cancer using a combination of LASSO regression-based classifier and statistical bioinformatics models. Differentially expressed genes (DEGs) were identified using the "Limma" package in R, and significant genes were further filtered using the LASSO-based classifier significance coefficient. Genes common to both methods were considered as the focus of this study. Additionally, Protein-Protein Interaction (PPI) networks of these key genes were constructed using STRING, and hub genes, significant modules, and associated genes were identified using Cytoscape. Results This study identified CCR8, CXCL11, CCL23, CCL24, CCL28, and CCL21 as signature prognostic genes for breast cancer, revealing a strong association between chemokines and breast cancer pathogenesis. Extensive literature searches were conducted to validate and confirm their prognostic significance in the disease. Conclusion These findings are pivotal for enhancing our comprehension of the pathways involved in breast cancer. Additionally, they hold promise as novel biomarkers for diagnostic purposes and may also reveal significant therapeutic targets for the management of breast cancer. The codes are available in the following GitHub repository: https://github.com/guhasouvik/LASSO_BRCA.git Breast Cancer Machine Learning LASSO Regression Microarray Prognosis PPI Network Signature Genes Biomarkers Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 1. Introduction Breast cancer is one of the most diagnosed cancers among females globally. In 2020 alone more than 2.26 million breast cancer cases were diagnosed, accounting for 11.7% of total cancer cases. Breast cancer is also one of the major causes of mortality among females worldwide [ 1 ]. Breast cancer has surpassed lung cancer to account for 1 in 8 cancer diagnoses and 2.3 million new cases in both sexes [ 2 ]. Breast cancer may be classified into five subgroups based on molecular diagnosis namely basal-like, HER2, luminal A, and luminal B [ 3 ]. Numerous risk factors contribute to the onset of breast cancer. Gender is often the primary unmodifiable factor since women are more likely than males to acquire breast cancer due to the significant association between specific sex hormones and elevated risk. With over 40% of patients over 65 and over 60% of all breast cancer fatalities occurring in this age range, age is also a key factor in the incidence of breast cancer [ 5 ]. A family history of the disease also significantly raises the likelihood of developing breast cancer; around 13–19% of newly diagnosed patients report having a first-degree relative with the disease. On the other hand, modifiable factors like taking certain medications like diethylstilbestrol or undergoing hormone replacement therapy (HRT), significantly raise the probability of developing breast cancer [ 6 ]. Studies have confirmed that the consumption of alcohol affects the level of estrogen and thus causes a hormonal imbalance that can elevate the probability of carcinogenesis in females [ 7 ]. Like other diseases, early diagnosis of breast cancer is very crucial for effective treatment and positive prognosis, with significantly lower mortality rates and higher survival rates observed in patients with smaller tumors at the time of diagnosis [ 8 ]. Medical imaging techniques such as Mammography, ultrasonography, magnetic resonance imaging (MRI), scintimammography, single photon emission computed tomography (SPECT), and positron emission tomography (PET) are extensively used for diagnosis [ 9 ]. Precision critical care medicine now has more prospects due to recent developments in high-throughput transcriptomic technology, which allow for quick and therapeutically useful gene expression profiling in a matter of hours. In diseases like cancer analysis of the transcriptomic data could reveal intricate molecular pathways and could help in the development of targeted therapeutics. Differential gene expression analysis is a popular computational method for identifying genes whose expressions are significantly different between two phenotypes. This study provides a novel approach for improving the statistical transcriptomic analysis techniques' outcomes by fusion of Machine learning-based expression profile analysis. In the proposed pipeline [Figure 1] implements a LASSO regression-based identification of the significant genes. The gene expression profiles from the tumor samples belonging to 111 breast cancer patients and 84 control samples. Further studies of the overlapping genes using the PPI network and module analysis five signature genes were identified which demonstrated significant involvement in the disease. Figure 1. The proposed pipeline of Prognostic Gene analysis from mRNA expression data. 2. Material and Methodology 2.1 TGCA BRCA mRNA-Seq Data Selection UCSC Xena browser ( https://xenabrowser.net/datapages/?dataset=TCGA-BRCA.htseq_counts.tsv&host=https%3A%2F%2Fgdc.xenahubs.net&removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443 ) was queried for the extraction of The Cancer Genome Atlas (TCGA) breast cancer gene expression dataset. The data was subjected to back log transformation since the count data were \(lo{g}_{2}(x+1)\) transformed. The dataset was then verified with the samples present in the TGCA Dataset. 2.2 Data Validation and pre-processing Data pre-processing operations included checking for missing values. The edgeR package in R was employed on the mRNA data to obtain the \(lo{g}_{2}\) transformation and the upper quartile(normalization) of the data. The ARSyNseq function of the NOISeq package was applied to this normalized data to perform batch correction. The probe-id mapping with the respective gene symbols was performed using Synergizer [ 10 ], gProfile id converter [ 11 ], DB2DB conversion [ 12 ] and GEO2R [ 13 ]. It is typically observed that for genes with several splice variants, many probes match a certain gene symbol. For gene mapping, the average expression value of these probes was employed. 2.3 Identification of the Differentially Expressed Genes (DEGs) Following the data preprocessing a log2 transformation was applied. Then the DEGs were selected using the “limma” [ 14 ] package of the R ( http://www.r-project.org/ ). Differentially expressed genes (DEGs) between normal and cancerous samples were defined as those with a fold change greater than two and a p-value less than 0.05. 2.4 LASSO-regression Classifier-based Significant Genes (LCSG) Extraction Lasso regression uses the linear regression model as a foundation, but it also applies a technique known as L1 regularization, which adds extra data to the model to avoid overfitting. As a result, we may employ lasso to pick variables by fitting a model with every potential predictor and employing a regularization strategy that reduces the coefficient estimates to zero. Specifically, the reduction target includes the sum of the absolute values of the coefficients in addition to the residual sum of squares (RSS), as in the OLS regression setting. The residual sum of squares (RSS) is calculated as follows: $$RSS=\stackrel{n}{\sum _{i=1}}({y}_{i}\hat -{{y}_{i}}{)}^{2}$$ 1 This can also be stated as: $$RSS=\stackrel{n}{\sum _{i=1}}({y}_{i}-({\beta }_{0}+\stackrel{p}{\sum _{j=1}}{\beta }_{j}{x}_{ij}){)}^{2}$$ 2 n represents the number of observations, p denotes the number of genes \({x}_{ij}\) and represents the value of the jth variable for the ith observation (gene expression level), where i = 1, 2,. . ., n and j = 1, 2,. . ., p. Within the lasso regression, the goal of minimization is transformed into: $$\stackrel{n}{\sum _{i=1}}({y}_{i}-({\beta }_{0}+\stackrel{p}{\sum _{j=1}}{\beta }_{j}{x}_{ij}){)}^{2}+\alpha \stackrel{p}{\sum _{j=1}}|{\beta }_{j}|$$ 3 Or, $$RSS+\alpha \stackrel{p}{\sum _{j=1}}\left|{\beta }_{j}\right|$$ 4 Where \(\alpha\) can take various values, depending on the dataset the value is estimated. The bias-variance trade-off underlies the benefit of Lasso regression over least squares linear regression. The lasso regression fit becomes less flexible as rises, which results in lower variance but higher bias. We labelled the cancer samples with “1” and the normal samples with “0”. Hence, we obtained a binary class dataset. We divided our dataset into a training set and a validation set in a 7:3 ratio respectively using the sklearn train-test split module ( https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html ). The procedure for the LASSO-regression-based gene selection is discussed below: Step 1 : The optimal \(\alpha\) value was estimated using the GridSearchCV module ( https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GridSearchCV.html ) and was implemented into the Lasso module of Scikit ( https://scikit-learn.org/stable/modules/generated/sklearn.linear_model.Lasso.html ). Step 2 The model was trained with the training set. Step 3 The coefficients of all the genes present in the database were calculated. Step 4 The coefficients were sorted in descending order. Step 5 The genes having non-zero coefficients were extracted and other genes were removed. Step 6 The model was evaluated using the validation set and its performance was evaluated, Step 1 was to be repeated if required. 2.5 Overlapping Genes Identification between the DEGs and LCSG To find out the common genes between the differentially expressed genes identified by “limma” and the genes identified as significant by the LASSO classifier, we checked for their overlap, by using the online available tool Venny 2.1.0 ( http://bioinfogp.cnb.csic.es/tools/venny/ ). 2.6 Enrichment Analysis of the Overlapping Genes Based on the GO hierarchy and Reactome 2022 pathway database, we categorized the DEGs into several functional categories of cellular integrals, biological systems, and molecular activities using the online enrichment analysis tool EnrichR ( https://maayanlab.cloud/Enrichr/ ). The top 10 terms of each of these categories were extracted and their overlapping was further studied using Venny 2.1.0 ( http://bioinfogp.cnb.csic.es/tools/venny/ ). 2.7 PPI Network Analysis and Significant Module Analysis The STRING 11.5 ( https://string-db.org/ ) web-based service was employed to create the PPI network on the significant genes [ 15 ]. The confidence score was set to greater than 0.70 (very high confidence) and a maximum number of interactions of 0 was set to construct the network [ 16 ]. TSV data file was exported from STRING into Cytoscape to construct the PPI network. The proteins are represented by nodes, and protein-protein interactions are represented by edges. A higher node shape signifies a greater degree of connection. The significant modules were identified using MCODE with the default settings: node score = 0.2, degree = 2, K-score = 2 and max depth = 100, respectively. The genes in these modules are the signature prognostic genes. Prognostic Signature Genes \(=\stackrel{n}{\bigcup _{i=1}}\) Genes Present in Modules (5) Where n is the number of modules obtained. 2.8 Survival Analysis through Kaplan–Meier (KM) plot The Kaplan-Meier plotter ( www.kmplot.com ) was employed to assess the predictive significance of signature prognostic genes. This online database has gene expression profiles and survival data of cancer patients. Based on the median gene expression, samples were categorized into high and low-expression groups for OS and recurrence-free survival (RFS) analysis. The log-rank p-value and hazard ratios (HR) were computed along with the matching 95% confidence interval (CI). Significance was considered as a value of p < 0.05. 3. RESULTS 3.1 Study Setup Python version 3.6.9 in the Google Colab environment was used in all the analyses. The operating system was Apple Sonoma 14.2.1 was used in a hardware setup of MacBook Air 2023 based on the Apple M2 chip as the processor and with 8 GB of RAM. 3.2 Data Extraction and Identification of the DEGs Initially, the gene expression dataset contains 30,628 genes and 1,217 samples. Among those samples, we found 1,001 common samples in the phenotype data (patients’ survival data). Thus, we proceed with the subset of 30,628 x 1,217. Among 1,001 samples, 916 were breast cancer patients and remaining 85 samples were normal individuals. After all the preprocessing operations and gene index mapping the p-value was calculated for the dataset using “limma”. To reduce the dimensionality of the dataset a filter was applied to the p-value to take only the genes with p-value \(<0.05\) after the DEGs extraction a total of 3,858 genes were found to be up-regulated, 8,122 genes remained unaltered, and the remaining 2,020 genes were also up-regulated, based on the screening criteria of BH-corrected p-value, less than 0.05, and FC (fold change) of more than 2. Volcano plot highlighting the DEGs is shown in Fig. 2(a). The variation in the expression levels between normal and cancer patients can be visualized in Fig. 2(b). 3.3 LASSO-regression Classifier-based Significant Genes (LCSG) The LASSO regression classifier was trained with the training set, and the model was evaluated on the validation set. The confusion matrix of the model and the ROC curve were built using the sklearn Confusion Matrix package. The accuracy of the model was evaluated using the accuracy score module of the sklearn library ( https://scikit- learn.org/stable/modules/generated/sklearn.metrics.accuracy_score.html ). Accuracy of a model is defined as follows: $$Accuracy=\frac{TP+TN}{TP+TN+FP+FN}.100\%$$ 6 where TP and TN represent true positive and true negative cases while FP and FN represent false positive and false negative cases. The maximum accuracy of the model was 96.61% and an AUC score of 0.966 for \(\alpha\) value of 1e-5. Hence \(\alpha\) value was fixed at 1e-5 for which the classification accuracy was maximum. The coefficients of the genes were extracted using the “coef_” module of sklearn ( https://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LinearRegression.html ). The absolute value of the coefficients was sorted and all the genes with zero coefficients were removed. Out of the total 140000 genes, 1182 genes had an absolute coefficient value not equal to 0. These 1182 genes were termed LCSG. 3.4 Overlapping Genes between DEGs and LCSG As observed in Fig. 4 total of 952 genes were exclusively included in LASSO CLASSIFIED and 1499 genes were included exclusively in the DEGs set identified by “limma”. 230 genes were found to be overlapping between the two sets. These common 230 genes were characterized based on their biological function. 3.5 Characterization of the genes of interest We utilized the GO hierarchy and Reactome 2022 pathway database to categorize the Differentially Expressed Genes (DEGs) into functional categories of cellular integrals, biological systems, and molecular activities. This categorization was performed using the online enrichment analysis tool EnrichR ( https://maayanlab.cloud/Enrichr/ ), which revealed significant enrichment in the following top three GO terms under each following category: Molecular Function category (terms arranged in descending order of their significance in each category): Cytokine Activity (GO:0005125), Chemokine Activity (GO:0008009) and Chemokine Receptor Binding (GO:0042379). In Biological Process category: Phasic Smooth Muscle Contraction (GO:0014821), Regulation Of Peptidyl-Serine Phosphorylation (GO:0033135) and Positive Regulation Of MAPK Cascade (GO:0043410). In the Cellular component category (descending order): External Side Of Apical Plasma Membrane (GO:0098591), Exocytic Vesicle Membrane (GO:0099501) and Synaptic Vesicle Membrane (GO:0030672). The significantly enriched Reactome pathways (in descending order) are GPCR Ligand Binding (R-HSA-500792), ADORA2B Mediated Anti-Inflammatory Cytokine Production ( R-HSA-9660821) and Signaling By GPCR (R-HSA-372790). 3.6 PPI Network Analysis and Significant Modules The PPI Network was created from 55 genes extracted from the top ten GO terms in the three categories and the top ten Reactome pathways were constructed in STRING version 12.0. The network was then exported as a TSV file which was imported into Cytoscape version 3.10.1. Module analysis was done using MCODE. Two Modules each having a score of 3 having 3 nodes, 3 edges and were identified and the genes are termed as the prognostic signature genes on which further studies were carried out on their expression. The genes present in the modules were CCL24, CCL21, CCR8, CXCL11, CCL28 and CCL23. 4. Discussion The morbidity and mortality rates of breast cancer have risen drastically over the last few decades, and there is an urgent necessity to tailor an appropriate management and treatment strategy. A sustained decrease in breast cancer death rates could be achieved with an accurate diagnosis of cancer in the early stages. The most crucial step in achieving the best prognosis is to identify the cells that are cancerous in the early stages. Scientists have investigated numerous methods for the diagnosis of breast cancer, including mammography, positron emission tomography, magnetic resonance imaging, computed tomography, ultrasound, and biopsy. These methods are not suitable for young women and have several drawbacks, including cost and time commitment [ 17 ]. For the prediction of diseases like cancer, it becomes crucial to identify the genes that are involved in the disease development and progression. In the analysis of such genes, a Machine Learning technique like the Lasso regression method capable of handling large genetic data like the microarray sequences and adequate statistical knowledge capable of interpreting these data are required. The analysis of the microarray gene expression dataset utilized the methodology of Differential expression analysis and an efficient ML approach to find the common set of significant genes. The DEG analysis compares the degree of expression in the diseased and control groups using a variety of statistical techniques, including the t-test of cohorts. On the other hand, the LASSO-regression classifies the data using its statistical machine-learning approach. While this approach works well, it is limited by the presence of a large amount of noise in the gene expression data, the repeatability of the results, and individual differences due to age, gender, genotype, stage of illness, and other variables. This disadvantage can be removed by implementing a statistical meta-analysis of combined different studies which would enable us to find the distinct disease signatures that are typically consistent in several studies. [ 18 ]. Another drawback of using the LASSO-based classifier is that it requires a large set of data for initial training purposes. The dataset should also be balanced which means there should not be any large difference between the counts of the two categories. The LASSO-based classifier implemented in this study demonstrated a significant classification accuracy of 96.6%. From the DEGs identification through “limma” we got a total of 5879 differentially expressed genes. The intersection of these two sets obtained from LASSO and LIMMA gave 230 genes of interest which were used in further analysis. The genes of interest were subjected to Reactome and GO pathway analysis through Enrichr. The genes were enriched in significant GO terms like Cytokine Activity (GO:0005125), Chemokine Activity (GO:0008009), Chemokine Receptor Binding (GO:0042379), Positive Regulation Of MAPK Cascade (GO:0043410), Neutrophil Chemotaxis (GO:0030593). Earlier studies have shown that Cytokines actively take part in processes involved in tumor onset, promotion, angiogenesis, and metastasis in breast cancer [ 19 ]. Angiogenesis creates blood vessels that give the necessary nutrients and oxygen for tumor development in solid malignancies. This mechanism appears to be a well-defined characteristic of cancer and has been shown in several studies to play a critical role in the origin of cancer metastasis. This process is highly controlled, most notably by angiogenesis-promoting or angiogenesis-inhibiting cytokines [ 20 ]. Breast cancer is largely caused by persistent inflammation, which is the initial step of malignant growth [ 21 ]. Research has indicated that persistent inflammation could increase the probability of developing breast cancer. Additionally, inflammatory regulatory factors secreted by breast cancer cells might facilitate the advancement of inflammation, so establishing a vicious cycle whereby inflammatory tumors enhance the evolution of cancer [ 22 ]. These inflammations could trigger inflammatory mediators which result in the production of nonspecific proinflammatory cytokines (TNF-α [tumor necrosis factor-α], IL-6, and IFN-α), which can then trigger the development of chemokines to further inflammation [ 23 ]. Studies have shown that in the inflammatory microenvironment, the chemokines and chemokine receptors can support tumor development. According to clinical research, there is a strong correlation between the development of tumours and the overexpression of certain chemokines in breast cancer [ 25 ]. Central signaling pathways known as MAPK cascades control a broad range of stimulated cellular processes, such as stress response, apoptosis, differentiation, and proliferation. Consequently, dysregulation, or improper functioning of these cascades, plays a role in the onset and development of illnesses like diabetes, cancer, autoimmune disorders, and faulty developmental processes [ 26 ]. Human tissues contain three main MAP kinase pathways; however, the one that involves ERK-1 and − 2 is more pertinent to breast cancer. The primary regulators of ERK-1 and − 2 are peptide growth factors that function through receptors that include tyrosine kinase. A variety of additional ligands can also function non-genomically by activating MAP kinase through heterotrimeric G protein receptors, including progesterone, testosterone, and estradiol. According to recent research, there is often a higher percentage of cells with active MAP kinase in breast tumors [ 27 ]. The most prevalent leukocytes in the body, neutrophils, are becoming more well-acknowledged for their potential to actively modulate cancer. In the bloodstream, neutrophils escort circulating tumor cells to promote their survival and stimulate their proliferation and metastasis. Tumor-infiltrating neutrophils (TINs) are the neutrophils that are found in the tumor microenvironment (TME) and interact with cancer cells [ 28 ]. High TIN levels have been linked to advanced histologic grade, tumor stage, and the TNBC subtype [ 29 ]. In the REACTOME Pathways, some of the significant terms were: GPCR Ligand Binding (R-HSA-500792), Signaling by GPCR (R-HSA-372790), Chemokine Receptors Bind Chemokines (R-HSA-380108). The biggest family of cell-surface receptors, the G protein-coupled receptors (GPCRs), play a role in the initiation and progression of several malignancies, including breast cancer. Due to their abnormal activation and overexpression, GPCRs are commonly linked to several features of cancer, such as angiogenesis, metastasis, tumour development, invasion, migration, and survival [ 30 ]. The chemokine activities also come under significance in the pathway analysis, brief discussion on their activities has been previously discussed. This study found 6 prognostic signature genes namely CCL24, CCL21, CCR8, CXCL11, CCL28 and CCL23 which all belong to the chemokine family. The relative expression profiles are shown in Fig. 7. CCR8 which is a cell surface receptor that belongs to Class A of the G protein-coupled receptor (GPCR) family is observed to be upregulated in breast cancer patients. Studies show that it is well CCR8 is essential for recruiting Tregs to the tumor site and creating an immunosuppressive environment that facilitates tumor escape. This recruiting mechanism can be interfered with by inhibiting CCR8, which may enhance anti-tumor immune responses and slow tumor development [ 31 ]. According to current studies, anti-CCR8 antibodies can impair CCR8 activity, which lowers the number of Treg cells accumulating in tumors and interferes with their immunosuppressive role [ 32 ]. The KM plot analysis [Figure 8] though shows conflicting results. The Overall Survival (OS) analysis shows that under expression of CCR8 decreases OS. This could be partially attributed to the fact that at the initial stages of the tumor development, the cytokines could show anti-tumour activity [ 33 ]. Further studies in this regard are required. CXCL11 is a chemokine superfamily member of the CXC family and has been found to play a significant role in the development of breast cancer [ 34 ]. This molecule has been suggested to be the major ligand for CXCR3 and encodes the protein that activated T-cells need to trigger the chemotactic response. Research shows that CXCL11 activates ERK, which in turn promotes the migratory, invasive, and proliferative activities of MDA-MB-231 cells [ 35 ]. It is also exhibited by the KM plot that overexpression of CXCL11 decreases OS. From the analysis of the expression profiles [Figure 7], it is evident that the transcription levels of CCL 21/24/23/28 in breast cancer samples were decreased significantly [ 36 ]. This study similar to the previous ones proves that low expression of CCL21 is associated with worst OS. CCL21 is linked to enhanced immunogenicity in breast cancer. CCL24 which exhibits a high expression level in several types of cancer, this study found that it is under-expressed in breast cancer with no significant relation with OS. Again, previous works have shown that a reduction in the chemokine CCL23 in hepatic tumors is linked to a poor prognosis for HCC patients and may be a strategy used by the tumor cells to avoid the immune system [ 37 ]. The current study also confirms these results evident from the expression profile and Survival analysis. Studies have exhibited Oral squamous cell carcinoma cells with detectable RUNX3 expression levels, CCL28 prevented invasion and the epithelial-mesenchymal transition (EMT). Its suppression of EMT was characterized by enhanced E-cadherin expression and reduced nuclear localization of β-catenin. RARβ expression was elevated by CCL28 signaling through CCR10, which also decreased the interaction between RARα and HDAC1 [ 38 ]. In the present study, it has been observed that CCL28 is under expressed, confirming previous studies [ 39 ]. Studies have demonstrated the involvement of CCL28 and CCL27 in the immune system's anticancer response. These chemokines cause anticancer NK cells to infiltrate the tumor, improving the prognosis through a higher expression of these chemokines in the tumor [ 40 ]. The results highly support the fact that the chemokines that modulate immune cell trafficking play a crucial role in breast cancer tumor microenvironment perturbations. They play a significant role in immune cell infiltration in the tumor progression and could show anti-cancer properties, as well as some pro-cancer characteristics and thus they play an important role in neoplasia. 5. Conclusion This study has successfully identified several prognostic signature genes and associated mechanisms linked to breast cancer development. While we have pinpointed significant prognostic genes, it’s crucial to recognize that changes in gene expression are governed by regulatory pathways. Understanding these regulatory networks is essential for developing effective disease prediction models. Through the application of various statistical methods and a Machine Learning approach using a Lasso-regression-based classifier, we have unearthed biologically relevant components essential for modelling the intricate associations between genotype and phenotype in breast cancer. These findings will be invaluable to physicians and researchers interested in elucidating correlated pathways and underlying mechanisms in disease development. Further investigations into the roles of these six identified genes in breast cancer development and progression are necessary for enhancing our understanding of the underlying mechanisms [ 41 – 43 ]. 6. Limitations (1) Our analysis was conducted only using breast cancer data. In future, we will extend our work for other tissue-specific cancer data. (2) Our resultant gene signatures have not been tested in wet lab regarding their patients’ prognosis study. Declarations Data Availability The dataset generated and/or analysed during the current study are available in UCSC Xenabrowser repository, https://xenabrowser.net/datapages/?dataset=TCGA-BRCA.htseq_counts.tsv&host=https%3A%2F%2Fgdc.xenahubs.net&removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443 Ethics approval and consent to participate Not Applicable. Consent for publication Not Applicable. Competing interests The authors declare that he has no competing interests. Authors contribution S.G. and S.S. designed the algorithm. S.G. implemented the framework and wrote the manuscript. A.L., T.B. and S.M. validated the result. S.M., A.M. and M.A.S. reviewed and edited the manuscript. All authors approved the manuscript. References Arnold, M., Morgan, E., Rumgay, H., Mafra, A., Singh, D., Laversanne, M., Vignat, J., Gralow, J. R., Cardoso, F., Siesling, S., & Soerjomataram, I. (2022). Current and future burden of breast cancer: Global statistics for 2020 and 2040. Breast (Edinburgh, Scotland) , 66 , 15–23. https://doi.org/10.1016/j.breast.2022.08.010 Sung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, et al. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA A Cancer J Clin 2021;71:209–49 https://doi.org/10.3322/caac.21660[3]Orrantia-Borunda, E., Anchondo-Nuñez, P., Acuña-Aguilar, L. E., Gómez-Valles, F. O., & Ramírez-Valdespino, C. A. (2022). Subtypes of Breast Cancer. In H. N. Mayrovitz (Ed.), Breast Cancer . Exon Publications. Smolarz, B., Nowak, A. Z., & Romanowicz, H. (2022). Breast Cancer-Epidemiology, Classification, Pathogenesis and Treatment (Review of Literature). Cancers, 14(10), 2569. https://doi.org/10.3390/cancers14102569 Folkerd, E., & Dowsett, M. (2013). Sex hormones and breast cancer risk and prognosis. Breast (Edinburgh, Scotland) , 22 Suppl 2 , S38–S43. https://doi.org/10.1016/j.breast.2013.07.007Authors McGuire, A., Brown, J. A., Malone, C., McLaughlin, R., & Kerin, M. J. (2015). Effects of age on the detection and management of breast cancer. Cancers , 7 (2), 908–929. https://doi.org/10.3390/cancers7020815 Narod S. A. (2011). Hormone replacement therapy and the risk of breast cancer. Nature reviews. Clinical oncology, 8(11), 669–676. https://doi.org/10.1038/nrclinonc.2011.110 Zeinomar, N., Knight, J. A., Genkinger, J. M., Phillips, K. A., Daly, M. B., Milne, R. L., Dite, G. S., Kehm, R. D., Liao, Y., Southey, M. C., Chung, W. K., Giles, G. G., McLachlan, S. A., Friedlander, M. L., Weideman, P. C., Glendon, G., Nesci, S., kConFab Investigators, Andrulis, I. L., Buys, S. S., … Terry, M. B. (2019). Alcohol consumption, cigarette smoking, and familial breast cancer risk: findings from the Prospective Family Study Cohort (ProF-SC). Breast cancer research : BCR , 21 (1), 128. https://doi.org/10.1186/s13058-019-1213-1 Duncan, W., & Kerr, G. R. (1976). The curability of breast cancer. British medical journal , 2 (6039), 781–783. https://doi.org/10.1136/bmj.2.6039.781 Iranmakani, S., Mortezazadeh, T., Sajadian, F. et al. A review of various modalities in breast imaging: technical aspects and clinical outcomes. Egypt J Radiol Nucl Med 51, 57 (2020). https://doi.org/10.1186/s43055-020-00175-5 Gabriel F. Berriz, Frederick P. Roth, The Synergizer service for translating gene, protein and other biological identifiers, Bioinformatics , Volume 24, Issue 19, October 2008, Pages 2272–2273, https://doi.org/10.1093/bioinformatics/btn424 Jüri Reimand, Tambet Arak, Priit Adler, Liis Kolberg, Sulev Reisberg, Hedi Peterson, Jaak Vilo, g:Profiler—a web server for functional interpretation of gene lists (2016 update), Nucleic Acids Research , Volume 44, Issue W1, 8 July 2016, Pages W83–W89, https://doi.org/10.1093/nar/gkw199 Uma Mudunuri, Anney Che, Ming Yi, Robert M. Stephens, bioDBnet: the biological database network, Bioinformatics , Volume 25, Issue 4, February 2009, Pages 555–556, https://doi.org/10.1093/bioinformatics/btn654 Tanya Barrett, Stephen E. Wilhite, Pierre Ledoux, Carlos Evangelista, Irene F. Kim, Maxim Tomashevsky, Kimberly A. Marshall, Katherine H. Phillippy, Patti M. Sherman, Michelle Holko, Andrey Yefanov, Hyeseung Lee, Naigong Zhang, Cynthia L. Robertson, Nadezhda Serova, Sean Davis, Alexandra Soboleva, NCBI GEO: archive for functional genomics data sets—update, Nucleic Acids Research , Volume 41, Issue D1, 1 January 2013, Pages D991–D995, https://doi.org/10.1093/nar/gks1193 Matthew E. Ritchie, Belinda Phipson, Di Wu, Yifang Hu, Charity W. Law, Wei Shi, Gordon K. Smyth, limma powers differential expression analyses for RNA-sequencing and microarray studies, Nucleic Acids Research, Volume 43, Issue 7, 20 April 2015, Page e47 https://doi.org/10.1093/nar/gkv007 Jiang, X., Xu, Z., Du, Y. et al. Bioinformatics analysis reveals novel hub gene pathways associated with IgA nephropathy. Eur J Med Res 25 , 40 (2020). https://doi.org/10.1186/s40001-020-00441-2 Song, X., Du, R., Gui, H., Zhou, M., Zhong, W., Mao, C., & Ma, J. (2020). Identification of potential hub genes related to the progression and prognosis of hepatocellular carcinoma through integrated bioinformatics analysis. Oncology reports , 43 (1), 133–146. https://doi.org/10.3892/or.2019.7400 Wang, Lulu. 2017. "Early Diagnosis of Breast Cancer" Sensors 17, no. 7: 1572. https://doi.org/10.3390/s17071572 Haidich A. B. (2010). Meta-analysis in medical research. Hippokratia , 14 (Suppl 1), 29–37. Esquivel-Velázquez, M., Ostoa-Saloma, P., Palacios-Arreola, M. I., Nava-Castro, K. E., Castro, J. I., & Morales-Montor, J. (2015). The role of cytokines in breast cancer development and progression. Journal of interferon & cytokine research : the official journal of the International Society for Interferon and Cytokine Research, 35(1), 1–16. https://doi.org/10.1089/jir.2014.0026 Geindreau, M., Bruchard, M., & Vegran, F. (2022). Role of Cytokines and Chemokines in Angiogenesis in a Tumor Context. Cancers, 14(10), 2446. https://doi.org/10.3390/cancers14102446 Shacter, E., & Weitzman, S. A. (2002). Chronic inflammation and cancer. Oncology (Williston Park, N.Y.), 16(2), 217–232. Liu, L., Wang, Y., Miao, L., Liu, Q., Musetti, S., Li, J., & Huang, L. (2018). Combination Immunotherapy of MUC1 mRNA Nano-vaccine and CTLA-4 Blockade Effectively Inhibits Growth of Triple Negative Breast Cancer. Molecular therapy : the journal of the American Society of Gene Therapy, 26(1), 45–55. https://doi.org/10.1016/j.ymthe.2017.10.020 Burke, D. L., Frid, M. G., Kunrath, C. L., Karoor, V., Anwar, A., Wagner, B. D., Strassheim, D., & Stenmark, K. R. (2009). Sustained hypoxia promotes the development of a pulmonary artery-specific chronic inflammatory microenvironment. American journal of physiology. Lung cellular and molecular physiology, 297(2), L238–L250. https://doi.org/10.1152/ajplung.90591.2008 Liu, H., Yang, Z., Lu, W., Chen, Z., Chen, L., Han, S., Wu, X., Cai, T., & Cai, Y. (2020). Chemokines and chemokine receptors: A new strategy for breast cancer therapy. Cancer medicine, 9(11), 3786–3799. https://doi.org/10.1002/cam4.3014 Aldinucci, D., & Colombatti, A. (2014). The inflammatory chemokine CCL5 and cancer progression. Mediators of inflammation, 2014, 292376. https://doi.org/10.1155/2014/292376 Alexander Plotnikov, Eldar Zehorai, Shiri Procaccia, Rony Seger, The MAPK cascades: Signaling components, nuclear roles and mechanisms of nuclear translocation, Biochimica et BiophysicaActa (BBA) - Molecular Cell Research, Volume 1813, Issue 9,2011,Pages 1619-1633, ISSN 0167-4889,https://doi.org/10.1016/j.bbamcr.2010.12.012. Santen, R. J., Song, R. X., McPherson, R., Kumar, R., Adam, L., Jeng, M. H., & Yue, W. (2002). The role of mitogen-activated protein (MAP) kinase in breast cancer. The Journal of steroid biochemistry and molecular biology, 80(2), 239–256. https://doi.org/10.1016/s0960-0760(01)00189-3 Coffelt, S. B., Wellenstein, M. D., & de Visser, K. E. (2016). Neutrophils in cancer: neutral no more. Nature reviews. Cancer, 16(7), 431–446. https://doi.org/10.1038/nrc.2016.52 Wang, Y., Chen, J., Yang, L., Li, J., Wu, W., Huang, M., Lin, L., & Su, S. (2019). Tumor-Contacted Neutrophils Promote Metastasis by a CD90-TIMP-1 Juxtacrine-Paracrine Loop. Clinical cancer research: an official journal of the American Association for Cancer Research, 25(6), 1957–1969. https://doi.org/10.1158/1078-0432.CCR-18-2544 Dorsam, R. T., & Gutkind, J. S. (2007). G-protein-coupled receptors and cancer. Nature reviews. Cancer, 7(2), 79–94. https://doi.org/10.1038/nrc2069 Kidani, Y., Nogami, W., Yasumizu, Y., Kawashima, A., Tanaka, A., Sonoda, Y., Tona, Y., Nashiki, K., Matsumoto, R., Hagiwara, M., Osaki, M., Dohi, K., Kanazawa, T., Ueyama, A., Yoshikawa, M., Yoshida, T., Matsumoto, M., Hojo, K., Shinonome, S., Yoshida, H., … Sakaguchi, S. (2022). CCR8-targeted specific depletion of clonally expanded Treg cells in tumor tissues evokes potent tumour immunity with long-lasting memory. Proceedings of the National Academy of Sciences of the United States of America, 119(7), e2114282119. https://doi.org/10.1073/pnas.2114282119 Tanaka, T., Nanamiya, R., Takei, J., Nakamura, T., Yanaka, M., Hosono, H., Sano, M., Asano, T., Kaneko, M. K., & Kato, Y. (2021). Development of Anti-Mouse CC Chemokine Receptor 8 Monoclonal Antibodies for Flow Cytometry. Monoclonal antibodies in immunodiagnosis and immunotherapy, 40(2), 65–70. https://doi.org/10.1089/mab.2021.0005 Korbecki J, Grochans S, Gutowska I, Barczak K, Baranowska-Bosiacka I. CC Chemokines in a Tumor: A Review of Pro-Cancer and Anti-Cancer Properties of Receptors CCR5, CCR6, CCR7, CCR8, CCR9, and CCR10 Ligands. International Journal of Molecular Sciences. 2020; 21(20):7619. https://doi.org/10.3390/ijms21207619 Jie Zhang & Dawei Hu (2021) miR-1298-5p Influences the Malignancy Phenotypes of Breast Cancer Cells by Inhibiting CXCL11, Cancer Management and Research, 13:, 133-145, DOI: 10.2147/CMAR.S279121 Hyun Jung Hwang, Ye-Rim Lee, Donghee Kang, Hyung Chul Lee, Haeng Ran Seo, Ji-Kan Ryu, Yong-Nyun Kim, Young-Gyu Ko, Heon Joo Park, Jae-Seon Lee, Endothelial cells under therapy-induced senescence secrete CXCL11, which increases aggressiveness of breast cancer cells, Cancer Letters,Volume 490,2020,Pages 100-110,ISSN 0304-3835,https://doi.org/10.1016/j.canlet.2020.06.019. Bowen Chen, Shuyuan Zhang, Qiuyu Li, Shiting Wu, Han He, Jinbo Huang; Bioinformatics identification of CCL8/21 as potential prognostic biomarkers in breast cancer microenvironment. Biosci Rep 27 November 2020; 40 (11): BSR20202042. doi: https://doi.org/10.1042/BSR20202042 Karan Dev, CCL23 in Balancing the Act of Endoplasmic Reticulum Stress and Antitumor Immunity in Hepatocellular Carcinoma, Frontiers in Oncology,volume 11, 2021,doi:10.3389/fonc.2021.727583 Park, J., Zhang, X., Lee, S. K., Song, N. Y., Son, S. H., Kim, K. R., Shim, J. H., Park, K. K., & Chung, W. Y. (2019). CCL28-induced RARβ expression inhibits oral squamous cell carcinoma bone invasion. The Journal of clinical investigation, 129(12), 5381–5399. https://doi.org/10.1172/JCI125336 Mickanin, C. S., Bhatia, U., & Labow, M. (2001). Identification of a novel beta-chemokine, MEC, down-regulated in primary breast tumors. International journal of oncology, 18(5), 939–944. https://doi.org/10.3892/ijo.18.5.939 Degos, C., Heinemann, M., Barrou, J., Boucherit, N., Lambaudie, E., Savina, A., Gorvel, L., & Olive, D. (2019). Endometrial Tumor Microenvironment Alters Human NK Cell Recruitment, and Resident NK Cell Phenotype and Function. Frontiers in immunology, 10, 877. https://doi.org/10.3389/fimmu.2019.00877 Hossain, S.M.M., Khatun, L., Ray, S. et al. Pan-cancer classification by regularized multi-task learning. Sci Rep 11, 24252 (2021). https://doi.org/10.1038/s41598-021-03554-8 Bhar, A., Haubrock, M., Mukhopadhyay, A. et al. Coexpression and coregulation analysis of time-series gene expression data in estrogen-induced breast cancer cell. Algorithms Mol Biol 8, 9 (2013). https://doi.org/10.1186/1748-7188-8-9 A. Mukhopadhyay and M. Mandal, "Identifying Non-Redundant Gene Markers from Microarray Data: A Multiobjective Variable Length PSO-Based Approach," in IEEE/ACM Transactions on Computational Biology and Bioinformatics, vol. 11, no. 6, pp. 1170-1183, 1 Nov.-Dec. 2014, doi: 10.1109/TCBB.2014.2323065 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4363199","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":301387202,"identity":"a78685eb-c5d9-48c9-8d90-dbafc94d63a9","order_by":0,"name":"Souvik Guha","email":"","orcid":"","institution":"Jamia Millia Islamia (A Central University)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Souvik","middleName":"","lastName":"Guha","suffix":""},{"id":301387203,"identity":"81c96f6b-9a41-46ef-a0cc-8f4ad433ce0c","order_by":1,"name":"Soumita Seth","email":"","orcid":"","institution":"Aliah University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Soumita","middleName":"","lastName":"Seth","suffix":""},{"id":301387204,"identity":"50c4c1f7-76a0-48ba-9c05-51d4a14d828c","order_by":2,"name":"Tapas Bhadra","email":"","orcid":"","institution":"Aliah University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Tapas","middleName":"","lastName":"Bhadra","suffix":""},{"id":301387208,"identity":"82e90170-1709-441e-a1ae-4a5a8ec0f4a4","order_by":3,"name":"Anirban Mukhopadhyay","email":"","orcid":"","institution":"University of Kalyani","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Anirban","middleName":"","lastName":"Mukhopadhyay","suffix":""},{"id":301387210,"identity":"7d9126c5-fe12-49e0-9886-9a46e6a1c536","order_by":4,"name":"Aimin Li","email":"","orcid":"","institution":"Xi'an University of Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Aimin","middleName":"","lastName":"Li","suffix":""},{"id":301387211,"identity":"df71491f-0687-4854-95ee-50f444bebc8b","order_by":5,"name":"Saurav Mallik","email":"","orcid":"","institution":"Harvard T H Chan School of Public Health","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Saurav","middleName":"","lastName":"Mallik","suffix":""},{"id":301387213,"identity":"2bdf8bcc-f949-4854-a4b5-1c44a6106719","order_by":6,"name":"Mohd Asif Shah","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA30lEQVRIie3OMQrCMBSA4SeFujxwFUrxBEJKIA4GzxIRMiuCOHZzc1YEz9DiBQIZugQ9QBfFC1RcBAeNowgNOjnkhyQk8PEC4PP9a4JzYo/XgkZqt8pNpHwjjZV7jNRfkG4anKujONDWRufXMfA4U8EprSNMhawtRMnaezmNViBppsLEQYDB8FZyMEgiBD3MFPSO9aR5rYTY845Bekd4WNK8OKYgsR9TjBhkdoqyBB0f0zizZEQTE876SEZ0rXFST4rF7nITg2Rrgl2J80G8LBZ5LYHg7UY+Xnw+n8/3S08ubkrTnLI70AAAAABJRU5ErkJggg==","orcid":"","institution":"Kardan University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Mohd","middleName":"Asif","lastName":"Shah","suffix":""}],"badges":[],"createdAt":"2024-05-03 09:33:17","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4363199/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4363199/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":56547977,"identity":"bc0087f4-b521-4f93-9483-c3bd572eab57","added_by":"auto","created_at":"2024-05-15 15:39:35","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":314787,"visible":true,"origin":"","legend":"\u003cp\u003eThe proposed pipeline of Prognostic Gene analysis from mRNA expression data.\u003c/p\u003e","description":"","filename":"1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4363199/v1/b6fa29808be83672843dd009.jpeg"},{"id":56547979,"identity":"e0b9e353-5ab7-4f44-b628-c0cae3c57e19","added_by":"auto","created_at":"2024-05-15 15:39:36","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1541474,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eVisualization of the DEGs identified using “limma”.\u003c/strong\u003e (a) The volcano plot indicates that there are more down-regulated DEGs (shown in blue) than up-regulated DEGs (shown in red), with the unaltered genes (shown in grey). The fold change is shown on the X-axis (log2 scale), and the p-value is shown on the Y-axis (-log10) scale. (b) A heatmap matrix of the DEGs displaying each sample's expression levels. The sample name is shown on the horizontal axis, while expression level \u0026nbsp;is shown on the vertical axis.\u003c/p\u003e","description":"","filename":"2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4363199/v1/8b149970448145d455101b8c.jpeg"},{"id":56547981,"identity":"093d07fa-755e-46dc-bbcc-023380366240","added_by":"auto","created_at":"2024-05-15 15:39:37","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":407674,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eVarious evaluation parameters the LASSO-regression classifier.\u003c/strong\u003e (a) Confusion matrices representation of the results obtained during the validation of the classifier. All the model parameters like accuracy, specificity ROC curves, and AUC value can be calculated using this matrix. (b) The ROC curve of the classifier for different gene subsets(right). This graph's AUC (Area Under the Curve) score offers an overall performance indicator across all potential classification thresholds. In other words, the AUC score is the probability that the classifier ranks a random positive example higher than a random negative example. (c) Bar plot of the absolute value of the coefficients of the genes in descending order, a higher coefficient signifies a greater weightage during classification. (d) Model coefficients in a regression analysis, with an increase in the penalty parameter (alpha), all weights converge toward zero.\u003c/p\u003e","description":"","filename":"3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4363199/v1/68a8132fa93bcd275de3a5df.jpeg"},{"id":56547975,"identity":"db362408-1cdd-4c26-8ce7-de5257698f73","added_by":"auto","created_at":"2024-05-15 15:39:35","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":201666,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverlap analysis between various obtained genes sets.\u003c/strong\u003e (a) Venn diagram representation illustrating the overlap between the gene count classified to be significant by the LASSO Classifier and the genes identified as DEGs using “limma”. As 952 genes were exclusively placed in “LASSO CLASSIFIED”, 1499 genes were exclusively placed in “LIMMA CLASSIFIED”, and 230 genes were mutual to both . Blue circle illustrates the gene count of DEGs in “LASSO CLASSIFIED” group and yellow circle illustrates the gene count in “LIMMA CLASSIFIED” group. (b) Venn diagram representation shows the overlap between the genes associated with the top ten terms of GO Molecular Function (MF), Cellular Component (CC), Biological Process (BP) and the Reactome pathway. As shown, 1 gene was common in all the four sets. The four coloured ovals represent the different sets. Blue represents the genes belonging to the Reactome pathways, yellow represents GO MF genes, green denotes the GO CC genes and red belongs to GO BP genes.\u003c/p\u003e","description":"","filename":"4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4363199/v1/6003675fab6f8c7d1a9ec77d.jpeg"},{"id":56547983,"identity":"787627fd-a381-4cb5-9fae-e72baef80bae","added_by":"auto","created_at":"2024-05-15 15:39:37","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":223086,"visible":true,"origin":"","legend":"\u003cp\u003eCharacterization of the overlapping genes from “limma” and LASSO classifier. (a) Pathways of the genes of interest using Reactome. In EnrichR, the Conbined Scores of the different pathways are obtained by multiplying the z-score of the deviation from the predicted rank with the the log of the p-value from the Fisher exact test . (b) GO analysis of the overlapping genes from top Biological Processes(BP), Cellular Components(CC), and Molecular Functions.Counts represent the number of genes present in each category.\u003c/p\u003e","description":"","filename":"6.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4363199/v1/902d1cc9317cfd334684f3a2.jpeg"},{"id":56547976,"identity":"bf01167f-d3f3-492a-9383-6acd0d858147","added_by":"auto","created_at":"2024-05-15 15:39:35","extension":"jpeg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":311351,"visible":true,"origin":"","legend":"\u003cp\u003eProtein-Protein interaction (PPI) network. This PPI was constructed with the genes present in the top 10 terms of the Reactome, GO-MF, CC and BP respectively using the Venn Illustration in Figure 4. The nodes are the proteins, and the edges represents interaction. (a) PPI network as obtained from STRING web module visualized in Cytoscape software. (b) The modules obtained from the network in (a) obtained by applying MCODE on the network. The genes pesent in the modules are the signature prognostic genes.\u003c/p\u003e","description":"","filename":"5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4363199/v1/dac355400c6a72f4ddce9ca6.jpeg"},{"id":56549781,"identity":"1f2a4a0f-571b-49a7-84a0-2c509255a013","added_by":"auto","created_at":"2024-05-15 15:47:37","extension":"jpeg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":765580,"visible":true,"origin":"","legend":"\u003cp\u003eBox Plot depicting Expression levels of the Signature Prognostic genes. The ‘T’ group represents the gene expression in breast cancer patients, while the ‘N’ group represents gene expression in individuals without cancer. The statistical analysis of differential expression between these groups was conducted using the non-parametric Wilcoxon-signed rank test. The expression profile reveals that CCR8 and CXCL11 were notably overexpressed in breast cancer patients, whereas CCL21, CCL24, CCL23, and CCL28 were underexpressed (downregulated) in breast cancer patients. The low p-values obtained from the statistical analysis suggest the potential use of these genes as biomarkers associated with breast cancer.\u003c/p\u003e","description":"","filename":"7.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4363199/v1/0d6a62281aced58ee38848c1.jpeg"},{"id":56547980,"identity":"b931f7a4-9483-4626-9d18-93de0f779275","added_by":"auto","created_at":"2024-05-15 15:39:37","extension":"jpeg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":609276,"visible":true,"origin":"","legend":"\u003cp\u003eThe prognostic value of the CC and CXC chemokines mRNA expression (Kaplan–Meier plotter). In the plots, red and black lines depict the survival curves of patient groups with mRNA expression levels higher and lower than the median levels in the target genes, respectively. The confidence intervals are represented in brackets, and the hazard ratio (HR) is indicated. This analysis provides insights into the potential prognostic significance of the identified signature genes.\u003c/p\u003e","description":"","filename":"8.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-4363199/v1/ad87d1fd500ce1b2511bbe0b.jpeg"},{"id":63439072,"identity":"61215e48-2cd2-4f9a-a4ff-289081356600","added_by":"auto","created_at":"2024-08-28 07:03:06","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5035336,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4363199/v1/0bcb57a6-a82a-4f80-8471-cc3c6bb00d81.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"LASSO Based Analysis for Prediction of Prognostic Signature Genes Associated with Breast Cancer","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eBreast cancer is one of the most diagnosed cancers among females globally. In 2020 alone more than 2.26\u0026nbsp;million breast cancer cases were diagnosed, accounting for 11.7% of total cancer cases. Breast cancer is also one of the major causes of mortality among females worldwide [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Breast cancer has surpassed lung cancer to account for 1 in 8 cancer diagnoses and 2.3\u0026nbsp;million new cases in both sexes [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Breast cancer may be classified into five subgroups based on molecular diagnosis namely basal-like, HER2, luminal A, and luminal B [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Numerous risk factors contribute to the onset of breast cancer. Gender is often the primary unmodifiable factor since women are more likely than males to acquire breast cancer due to the significant association between specific sex hormones and elevated risk. With over 40% of patients over 65 and over 60% of all breast cancer fatalities occurring in this age range, age is also a key factor in the incidence of breast cancer [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. A family history of the disease also significantly raises the likelihood of developing breast cancer; around 13\u0026ndash;19% of newly diagnosed patients report having a first-degree relative with the disease. On the other hand, modifiable factors like taking certain medications like diethylstilbestrol or undergoing hormone replacement therapy (HRT), significantly raise the probability of developing breast cancer [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Studies have confirmed that the consumption of alcohol affects the level of estrogen and thus causes a hormonal imbalance that can elevate the probability of carcinogenesis in females [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Like other diseases, early diagnosis of breast cancer is very crucial for effective treatment and positive prognosis, with significantly lower mortality rates and higher survival rates observed in patients with smaller tumors at the time of diagnosis [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Medical imaging techniques such as Mammography, ultrasonography, magnetic resonance imaging (MRI), scintimammography, single photon emission computed tomography (SPECT), and positron emission tomography (PET) are extensively used for diagnosis [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Precision critical care medicine now has more prospects due to recent developments in high-throughput transcriptomic technology, which allow for quick and therapeutically useful gene expression profiling in a matter of hours. In diseases like cancer analysis of the transcriptomic data could reveal intricate molecular pathways and could help in the development of targeted therapeutics. Differential gene expression analysis is a popular computational method for identifying genes whose expressions are significantly different between two phenotypes. This study provides a novel approach for improving the statistical transcriptomic analysis techniques' outcomes by fusion of Machine learning-based expression profile analysis. In the proposed pipeline [Figure 1] implements a LASSO regression-based identification of the significant genes. The gene expression profiles from the tumor samples belonging to 111 breast cancer patients and 84 control samples. Further studies of the overlapping genes using the PPI network and module analysis five signature genes were identified which demonstrated significant involvement in the disease.\u003c/p\u003e \u003cp\u003e \u003cb\u003eFigure 1.\u003c/b\u003e The proposed pipeline of Prognostic Gene analysis from mRNA expression data.\u003c/p\u003e"},{"header":"2. Material and Methodology","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 TGCA BRCA mRNA-Seq Data Selection\u003c/h2\u003e \u003cp\u003eUCSC Xena browser (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://xenabrowser.net/datapages/?dataset=TCGA-BRCA.htseq_counts.tsv\u0026amp;host=https%3A%2F%2Fgdc.xenahubs.net\u0026amp;removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443\u003c/span\u003e\u003cspan address=\"https://xenabrowser.net/datapages/?dataset=TCGA-BRCA.htseq_counts.tsv\u0026amp;host=https%3A%2F%2Fgdc.xenahubs.net\u0026amp;removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) was queried for the extraction of The Cancer Genome Atlas (TCGA) breast cancer gene expression dataset. The data was subjected to back log transformation since the count data were \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(lo{g}_{2}(x+1)\\)\u003c/span\u003e\u003c/span\u003e transformed. The dataset was then verified with the samples present in the TGCA Dataset.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Data Validation and pre-processing\u003c/h2\u003e \u003cp\u003eData pre-processing operations included checking for missing values. The edgeR package in R was employed on the mRNA data to obtain the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(lo{g}_{2}\\)\u003c/span\u003e\u003c/span\u003e transformation and the upper quartile(normalization) of the data. The ARSyNseq function of the NOISeq package was applied to this normalized data to perform batch correction. The probe-id mapping with the respective gene symbols was performed using Synergizer [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], gProfile id converter [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e], DB2DB conversion [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e] and GEO2R [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. It is typically observed that for genes with several splice variants, many probes match a certain gene symbol. For gene mapping, the average expression value of these probes was employed.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Identification of the Differentially Expressed Genes (DEGs)\u003c/h2\u003e \u003cp\u003eFollowing the data preprocessing a \u003cem\u003elog2\u003c/em\u003e transformation was applied. Then the DEGs were selected using the \u0026ldquo;limma\u0026rdquo; [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] package of the R (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.r-project.org/\u003c/span\u003e\u003cspan address=\"http://www.r-project.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e).\u003c/span\u003e Differentially expressed genes (DEGs) between normal and cancerous samples were defined as those with a fold change greater than two and a p-value less than 0.05.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 LASSO-regression Classifier-based Significant Genes (LCSG) Extraction\u003c/h2\u003e \u003cp\u003eLasso regression uses the linear regression model as a foundation, but it also applies a technique known as L1 regularization, which adds extra data to the model to avoid overfitting. As a result, we may employ lasso to pick variables by fitting a model with every potential predictor and employing a regularization strategy that reduces the coefficient estimates to zero. Specifically, the reduction target includes the sum of the absolute values of the coefficients in addition to the residual sum of squares (RSS), as in the OLS regression setting.\u003c/p\u003e \u003cp\u003eThe residual sum of squares (RSS) is calculated as follows:\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$RSS=\\stackrel{n}{\\sum _{i=1}}({y}_{i}\\hat -{{y}_{i}}{)}^{2}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThis can also be stated as:\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$$RSS=\\stackrel{n}{\\sum _{i=1}}({y}_{i}-({\\beta }_{0}+\\stackrel{p}{\\sum _{j=1}}{\\beta }_{j}{x}_{ij}){)}^{2}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003en represents the number of observations, p denotes the number of genes \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({x}_{ij}\\)\u003c/span\u003e\u003c/span\u003e and represents the value of the jth variable for the ith observation (gene expression level), where i\u0026thinsp;=\u0026thinsp;1, 2,. . ., n and j\u0026thinsp;=\u0026thinsp;1, 2,. . ., p.\u003c/p\u003e \u003cp\u003eWithin the lasso regression, the goal of minimization is transformed into:\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e\n$$\\stackrel{n}{\\sum _{i=1}}({y}_{i}-({\\beta }_{0}+\\stackrel{p}{\\sum _{j=1}}{\\beta }_{j}{x}_{ij}){)}^{2}+\\alpha \\stackrel{p}{\\sum _{j=1}}|{\\beta }_{j}|$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eOr,\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ4\" name=\"EquationSource\"\u003e\n$$RSS+\\alpha \\stackrel{p}{\\sum _{j=1}}\\left|{\\beta }_{j}\\right|$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eWhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\alpha\\)\u003c/span\u003e\u003c/span\u003e can take various values, depending on the dataset the value is estimated.\u003c/p\u003e \u003cp\u003eThe bias-variance trade-off underlies the benefit of Lasso regression over least squares linear regression. The lasso regression fit becomes less flexible as rises, which results in lower variance but higher bias.\u003c/p\u003e \u003cp\u003eWe labelled the cancer samples with \u0026ldquo;1\u0026rdquo; and the normal samples with \u0026ldquo;0\u0026rdquo;. Hence, we obtained a binary class dataset. We divided our dataset into a training set and a validation set in a 7:3 ratio respectively using the sklearn train-test split module (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html\u003c/span\u003e\u003cspan address=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.train_test_split.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The procedure for the LASSO-regression-based gene selection is discussed below:\u003c/p\u003e \u003cp\u003e \u003cb\u003eStep 1\u003c/b\u003e: The optimal \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\alpha\\)\u003c/span\u003e\u003c/span\u003e value was estimated using the GridSearchCV module (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GridSearchCV.html\u003c/span\u003e\u003cspan address=\"https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GridSearchCV.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) and was implemented into the Lasso module of Scikit (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://scikit-learn.org/stable/modules/generated/sklearn.linear_model.Lasso.html\u003c/span\u003e\u003cspan address=\"https://scikit-learn.org/stable/modules/generated/sklearn.linear_model.Lasso.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStep 2\u003c/strong\u003e \u003cp\u003eThe model was trained with the training set.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStep 3\u003c/strong\u003e \u003cp\u003eThe coefficients of all the genes present in the database were calculated.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStep 4\u003c/strong\u003e \u003cp\u003eThe coefficients were sorted in descending order.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStep 5\u003c/strong\u003e \u003cp\u003eThe genes having non-zero coefficients were extracted and other genes were removed.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStep 6\u003c/strong\u003e \u003cp\u003eThe model was evaluated using the validation set and its performance was evaluated, Step 1 was to be repeated if required.\u003c/p\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5 Overlapping Genes Identification between the DEGs and LCSG\u003c/h2\u003e \u003cp\u003eTo find out the common genes between the differentially expressed genes identified by \u0026ldquo;limma\u0026rdquo; and the genes identified as significant by the LASSO classifier, we checked for their overlap, by using the online available tool Venny 2.1.0 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://bioinfogp.cnb.csic.es/tools/venny/\u003c/span\u003e\u003cspan address=\"http://bioinfogp.cnb.csic.es/tools/venny/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.6 Enrichment Analysis of the Overlapping Genes\u003c/h2\u003e \u003cp\u003eBased on the GO hierarchy and Reactome 2022 pathway database, we categorized the DEGs into several functional categories of cellular integrals, biological systems, and molecular activities using the online enrichment analysis tool EnrichR (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://maayanlab.cloud/Enrichr/\u003c/span\u003e\u003cspan address=\"https://maayanlab.cloud/Enrichr/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The top 10 terms of each of these categories were extracted and their overlapping was further studied using Venny 2.1.0 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://bioinfogp.cnb.csic.es/tools/venny/\u003c/span\u003e\u003cspan address=\"http://bioinfogp.cnb.csic.es/tools/venny/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.7 PPI Network Analysis and Significant Module Analysis\u003c/h2\u003e \u003cp\u003eThe STRING 11.5 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://string-db.org/\u003c/span\u003e\u003cspan address=\"https://string-db.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e web-based service was employed to create the PPI network on the significant genes [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. The confidence score was set to greater than 0.70 (very high confidence) and a maximum number of interactions of 0 was set to construct the network [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. TSV data file was exported from STRING into Cytoscape to construct the PPI network. The proteins are represented by nodes, and protein-protein interactions are represented by edges. A higher node shape signifies a greater degree of connection. The significant modules were identified using MCODE with the default settings: node score\u0026thinsp;=\u0026thinsp;0.2, degree\u0026thinsp;=\u0026thinsp;2, K-score\u0026thinsp;=\u0026thinsp;2 and max depth\u0026thinsp;=\u0026thinsp;100, respectively. The genes in these modules are the signature prognostic genes.\u003c/p\u003e \u003cp\u003ePrognostic Signature Genes\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(=\\stackrel{n}{\\bigcup _{i=1}}\\)\u003c/span\u003e\u003c/span\u003eGenes Present in Modules (5)\u003c/p\u003e \u003cp\u003eWhere n is the number of modules obtained.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e2.8 Survival Analysis through Kaplan\u0026ndash;Meier (KM) plot\u003c/h2\u003e \u003cp\u003eThe Kaplan-Meier plotter (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e\u003ca href=\"https://github.com/guhasouvik/LASSO_BRCA.git\" target=\"_blank\"\u003ewww.kmplot.com\u003c/a\u003e\u003c/span\u003e\u003cspan address=\"http://www.kmplot.com\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) was employed to assess the predictive significance of signature prognostic genes. This online database has gene expression profiles and survival data of cancer patients. Based on the median gene expression, samples were categorized into high and low-expression groups for OS and recurrence-free survival (RFS) analysis. The log-rank p-value and hazard ratios (HR) were computed along with the matching 95% confidence interval (CI). Significance was considered as a value of p\u0026thinsp;\u0026lt;\u0026thinsp;0.05.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. RESULTS","content":"\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\n\u003ch2\u003e3.1 Study Setup\u003c/h2\u003e\n\u003cp\u003ePython version 3.6.9 in the Google Colab environment was used in all the analyses. The operating system was Apple Sonoma 14.2.1 was used in a hardware setup of MacBook Air 2023 based on the Apple M2 chip as the processor and with 8 GB of RAM.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\n\u003ch2\u003e3.2 Data Extraction and Identification of the DEGs\u003c/h2\u003e\n\u003cp\u003eInitially, the gene expression dataset contains 30,628 genes and 1,217 samples. Among those samples, we found 1,001 common samples in the phenotype data (patients\u0026rsquo; survival data). Thus, we proceed with the subset of 30,628 x 1,217. Among 1,001 samples, 916 were breast cancer patients and remaining 85 samples were normal individuals. After all the preprocessing operations and gene index mapping the p-value was calculated for the dataset using \u0026ldquo;limma\u0026rdquo;. To reduce the dimensionality of the dataset a filter was applied to the p-value to take only the genes with p-value\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\u0026lt;0.05\\)\u003c/span\u003e\u003c/span\u003e after the DEGs extraction a total of 3,858 genes were found to be up-regulated, 8,122 genes remained unaltered, and the remaining 2,020 genes were also up-regulated, based on the screening criteria of BH-corrected p-value, less than 0.05, and FC (fold change) of more than 2. Volcano plot highlighting the DEGs is shown in Fig.\u0026nbsp;2(a). The variation in the expression levels between normal\u0026nbsp;and cancer patients can be visualized in Fig.\u0026nbsp;2(b).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\n\u003ch2\u003e3.3 LASSO-regression Classifier-based Significant Genes (LCSG)\u003c/h2\u003e\n\u003cp\u003eThe LASSO regression classifier was trained with the training set, and the model was evaluated on the validation set. The confusion matrix of the model and the ROC curve were built using the sklearn Confusion Matrix package. The accuracy of the model was evaluated using the accuracy score module of the sklearn library (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://scikit-\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"Underline\"\u003elearn.org/stable/modules/generated/sklearn.metrics.accuracy_score.html\u003c/span\u003e). Accuracy of a model is defined as follows:\u003c/p\u003e\n\u003cdiv id=\"Equ5\" class=\"Equation\"\u003e\n\u003cdiv id=\"FileID_Equ5\" class=\"mathdisplay\"\u003e$$Accuracy=\\frac{TP+TN}{TP+TN+FP+FN}.100\\%$$\u003c/div\u003e\n\u003cdiv class=\"EquationNumber\"\u003e6\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003ewhere TP and TN represent true positive and true negative cases while FP and FN represent false positive and false negative cases.\u003c/p\u003e\n\u003cp\u003eThe maximum accuracy of the model was 96.61% and an AUC score of 0.966 for \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\alpha\\)\u003c/span\u003e\u003c/span\u003e value of 1e-5. Hence \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\alpha\\)\u003c/span\u003e\u003c/span\u003e value was fixed at 1e-5 for which the classification accuracy was maximum. The coefficients of the genes were extracted using the \u0026ldquo;coef_\u0026rdquo; module of sklearn (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://scikit-learn.org/stable/modules/generated/sklearn.linear_model.LinearRegression.html\u003c/span\u003e\u003c/span\u003e\u003cspan class=\"Underline\"\u003e).\u003c/span\u003e The absolute value of the coefficients was sorted and all the genes with zero coefficients were removed. Out of the total 140000 genes, 1182 genes had an absolute coefficient value not equal to 0. These 1182 genes were termed LCSG.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\n\u003ch2\u003e3.4 Overlapping Genes between DEGs and LCSG\u003c/h2\u003e\n\u003cp\u003eAs observed in Fig.\u0026nbsp;4 total of 952 genes were exclusively included in LASSO CLASSIFIED and 1499 genes were included exclusively in the DEGs set identified by \u0026ldquo;limma\u0026rdquo;. 230 genes were found to be overlapping between the two sets. These common 230 genes were characterized based on their biological function.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e\n\u003ch2\u003e3.5 Characterization of the genes of interest\u003c/h2\u003e\n\u003cp\u003eWe utilized the GO hierarchy and Reactome 2022 pathway database to categorize the Differentially Expressed Genes (DEGs) into functional categories of cellular integrals, biological systems, and molecular activities. This categorization was performed using the online enrichment analysis tool EnrichR (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://maayanlab.cloud/Enrichr/\u003c/span\u003e\u003c/span\u003e), which revealed significant enrichment in the following top three GO terms under each following category: Molecular Function category (terms arranged in descending order of their significance in each category): Cytokine Activity (GO:0005125), Chemokine Activity (GO:0008009) and Chemokine Receptor Binding (GO:0042379). In Biological Process category: Phasic Smooth Muscle Contraction (GO:0014821), Regulation Of Peptidyl-Serine Phosphorylation (GO:0033135) and Positive Regulation Of MAPK Cascade (GO:0043410). In the Cellular component category (descending order): External Side Of Apical Plasma Membrane (GO:0098591), Exocytic Vesicle Membrane (GO:0099501) and Synaptic Vesicle Membrane (GO:0030672). The significantly enriched Reactome pathways (in descending order) are GPCR Ligand Binding (R-HSA-500792), ADORA2B Mediated Anti-Inflammatory Cytokine Production ( R-HSA-9660821) and Signaling By GPCR (R-HSA-372790).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e\n\u003ch2\u003e3.6 PPI Network Analysis and Significant Modules\u003c/h2\u003e\n\u003cp\u003eThe PPI Network was created from 55 genes extracted from the top ten GO terms in the three categories and the top ten Reactome pathways were constructed in STRING version 12.0. The network was then exported as a TSV file which was imported into Cytoscape version 3.10.1. Module analysis was done using MCODE. Two Modules each having a score of 3 having 3 nodes, 3 edges and were identified and the genes are termed as the prognostic signature genes on which further studies were carried out on their expression. The genes present in the modules were CCL24, CCL21, CCR8, CXCL11, CCL28 and CCL23.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eThe morbidity and mortality rates of breast cancer have risen drastically over the last few decades, and there is an urgent necessity to tailor an appropriate management and treatment strategy. A sustained decrease in breast cancer death rates could be achieved with an accurate diagnosis of cancer in the early stages. The most crucial step in achieving the best prognosis is to identify the cells that are cancerous in the early stages. Scientists have investigated numerous methods for the diagnosis of breast cancer, including mammography, positron emission tomography, magnetic resonance imaging, computed tomography, ultrasound, and biopsy. These methods are not suitable for young women and have several drawbacks, including cost and time commitment [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. For the prediction of diseases like cancer, it becomes crucial to identify the genes that are involved in the disease development and progression. In the analysis of such genes, a Machine Learning technique like the Lasso regression method capable of handling large genetic data like the microarray sequences and adequate statistical knowledge capable of interpreting these data are required. The analysis of the microarray gene expression dataset utilized the methodology of Differential expression analysis and an efficient ML approach to find the common set of significant genes. The DEG analysis compares the degree of expression in the diseased and control groups using a variety of statistical techniques, including the t-test of cohorts. On the other hand, the LASSO-regression classifies the data using its statistical machine-learning approach. While this approach works well, it is limited by the presence of a large amount of noise in the gene expression data, the repeatability of the results, and individual differences due to age, gender, genotype, stage of illness, and other variables. This disadvantage can be removed by implementing a statistical meta-analysis of combined different studies which would enable us to find the distinct disease signatures that are typically consistent in several studies. [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Another drawback of using the LASSO-based classifier is that it requires a large set of data for initial training purposes. The dataset should also be balanced which means there should not be any large difference between the counts of the two categories. The LASSO-based classifier implemented in this study demonstrated a significant classification accuracy of 96.6%. From the DEGs identification through \u0026ldquo;limma\u0026rdquo; we got a total of 5879 differentially expressed genes. The intersection of these two sets obtained from LASSO and LIMMA gave 230 genes of interest which were used in further analysis. The genes of interest were subjected to Reactome and GO pathway analysis through Enrichr. The genes were enriched in significant GO terms like Cytokine Activity (GO:0005125), Chemokine Activity (GO:0008009), Chemokine Receptor Binding (GO:0042379), Positive Regulation Of MAPK Cascade (GO:0043410), Neutrophil Chemotaxis (GO:0030593). Earlier studies have shown that Cytokines actively take part in processes involved in tumor onset, promotion, angiogenesis, and metastasis in breast cancer [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Angiogenesis creates blood vessels that give the necessary nutrients and oxygen for tumor development in solid malignancies. This mechanism appears to be a well-defined characteristic of cancer and has been shown in several studies to play a critical role in the origin of cancer metastasis. This process is highly controlled, most notably by angiogenesis-promoting or angiogenesis-inhibiting cytokines [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Breast cancer is largely caused by persistent inflammation, which is the initial step of malignant growth [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Research has indicated that persistent inflammation could increase the probability of developing breast cancer. Additionally, inflammatory regulatory factors secreted by breast cancer cells might facilitate the advancement of inflammation, so establishing a vicious cycle whereby inflammatory tumors enhance the evolution of cancer [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. These inflammations could trigger inflammatory mediators which result in the production of nonspecific proinflammatory cytokines (TNF-α [tumor necrosis factor-α], IL-6, and IFN-α), which can then trigger the development of chemokines to further inflammation [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. Studies have shown that in the inflammatory microenvironment, the chemokines and chemokine receptors can support tumor development. According to clinical research, there is a strong correlation between the development of tumours and the overexpression of certain chemokines in breast cancer [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Central signaling pathways known as MAPK cascades control a broad range of stimulated cellular processes, such as stress response, apoptosis, differentiation, and proliferation. Consequently, dysregulation, or improper functioning of these cascades, plays a role in the onset and development of illnesses like diabetes, cancer, autoimmune disorders, and faulty developmental processes [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. Human tissues contain three main MAP kinase pathways; however, the one that involves ERK-1 and \u0026minus;\u0026thinsp;2 is more pertinent to breast cancer. The primary regulators of ERK-1 and \u0026minus;\u0026thinsp;2 are peptide growth factors that function through receptors that include tyrosine kinase. A variety of additional ligands can also function non-genomically by activating MAP kinase through heterotrimeric G protein receptors, including progesterone, testosterone, and estradiol. According to recent research, there is often a higher percentage of cells with active MAP kinase in breast tumors [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. The most prevalent leukocytes in the body, neutrophils, are becoming more well-acknowledged for their potential to actively modulate cancer. In the bloodstream, neutrophils escort circulating tumor cells to promote their survival and stimulate their proliferation and metastasis. Tumor-infiltrating neutrophils (TINs) are the neutrophils that are found in the tumor microenvironment (TME) and interact with cancer cells [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. High TIN levels have been linked to advanced histologic grade, tumor stage, and the TNBC subtype [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn the REACTOME Pathways, some of the significant terms were: GPCR Ligand Binding (R-HSA-500792), Signaling by GPCR (R-HSA-372790), Chemokine Receptors Bind Chemokines (R-HSA-380108). The biggest family of cell-surface receptors, the G protein-coupled receptors (GPCRs), play a role in the initiation and progression of several malignancies, including breast cancer. Due to their abnormal activation and overexpression, GPCRs are commonly linked to several features of cancer, such as angiogenesis, metastasis, tumour development, invasion, migration, and survival [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. The chemokine activities also come under significance in the pathway analysis, brief discussion on their activities has been previously discussed.\u003c/p\u003e \u003cp\u003eThis study found 6 prognostic signature genes namely CCL24, CCL21, CCR8, CXCL11, CCL28 and CCL23 which all belong to the chemokine family. The relative expression profiles are shown in Fig.\u0026nbsp;7. CCR8 which is a cell surface receptor that belongs to Class A of the G protein-coupled receptor (GPCR) family is observed to be upregulated in breast cancer patients. Studies show that it is well CCR8 is essential for recruiting Tregs to the tumor site and creating an immunosuppressive environment that facilitates tumor escape. This recruiting mechanism can be interfered with by inhibiting CCR8, which may enhance anti-tumor immune responses and slow tumor development [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. According to current studies, anti-CCR8 antibodies can impair CCR8 activity, which lowers the number of Treg cells accumulating in tumors and interferes with their immunosuppressive role [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. The KM plot analysis [Figure 8] though shows conflicting results. The Overall Survival (OS) analysis shows that under expression of CCR8 decreases OS. This could be partially attributed to the fact that at the initial stages of the tumor development, the cytokines could show anti-tumour activity [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Further studies in this regard are required. CXCL11 is a chemokine superfamily member of the CXC family and has been found to play a significant role in the development of breast cancer [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. This molecule has been suggested to be the major ligand for CXCR3 and encodes the protein that activated T-cells need to trigger the chemotactic response. Research shows that CXCL11 activates ERK, which in turn promotes the migratory, invasive, and proliferative activities of MDA-MB-231 cells [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. It is also exhibited by the KM plot that overexpression of CXCL11 decreases OS. From the analysis of the expression profiles [Figure 7], it is evident that the transcription levels of CCL 21/24/23/28 in breast cancer samples were decreased significantly [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. This study similar to the previous ones proves that low expression of CCL21 is associated with worst OS. CCL21 is linked to enhanced immunogenicity in breast cancer. CCL24 which exhibits a high expression level in several types of cancer, this study found that it is under-expressed in breast cancer with no significant relation with OS. Again, previous works have shown that a reduction in the chemokine CCL23 in hepatic tumors is linked to a poor prognosis for HCC patients and may be a strategy used by the tumor cells to avoid the immune system [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. The current study also confirms these results evident from the expression profile and Survival analysis. Studies have exhibited Oral squamous cell carcinoma cells with detectable RUNX3 expression levels, CCL28 prevented invasion and the epithelial-mesenchymal transition (EMT). Its suppression of EMT was characterized by enhanced E-cadherin expression and reduced nuclear localization of β-catenin. RARβ expression was elevated by CCL28 signaling through CCR10, which also decreased the interaction between RARα and HDAC1 [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. In the present study, it has been observed that CCL28 is under expressed, confirming previous studies [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. Studies have demonstrated the involvement of CCL28 and CCL27 in the immune system's anticancer response. These chemokines cause anticancer NK cells to infiltrate the tumor, improving the prognosis through a higher expression of these chemokines in the tumor [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. The results highly support the fact that the chemokines that modulate immune cell trafficking play a crucial role in breast cancer tumor microenvironment perturbations. They play a significant role in immune cell infiltration in the tumor progression and could show anti-cancer properties, as well as some pro-cancer characteristics and thus they play an important role in neoplasia.\u003c/p\u003e"},{"header":"5. Conclusion","content":"\u003cp\u003eThis study has successfully identified several prognostic signature genes and associated mechanisms linked to breast cancer development. While we have pinpointed significant prognostic genes, it\u0026rsquo;s crucial to recognize that changes in gene expression are governed by regulatory pathways. Understanding these regulatory networks is essential for developing effective disease prediction models. Through the application of various statistical methods and a Machine Learning approach using a Lasso-regression-based classifier, we have unearthed biologically relevant components essential for modelling the intricate associations between genotype and phenotype in breast cancer. These findings will be invaluable to physicians and researchers interested in elucidating correlated pathways and underlying mechanisms in disease development. Further investigations into the roles of these six identified genes in breast cancer development and progression are necessary for enhancing our understanding of the underlying mechanisms [\u003cspan additionalcitationids=\"CR42\" citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e].\u003c/p\u003e"},{"header":"6. Limitations","content":"\u003cp\u003e(1) Our analysis was conducted only using breast cancer data. In future, we will extend our work for other tissue-specific cancer data. (2) Our resultant gene signatures have not been tested in wet lab regarding their patients\u0026rsquo; prognosis study.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData Availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe dataset generated and/or analysed during the current study are available in UCSC Xenabrowser repository, https://xenabrowser.net/datapages/?dataset=TCGA-BRCA.htseq_counts.tsv\u0026amp;host=https%3A%2F%2Fgdc.xenahubs.net\u0026amp;removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot Applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot Applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that he has no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors contribution\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eS.G. and S.S. designed the algorithm. S.G. implemented the framework and wrote the manuscript. A.L., T.B. and S.M. validated the result. S.M., A.M. and M.A.S. reviewed and edited the manuscript. All authors approved the manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eArnold, M., Morgan, E., Rumgay, H., Mafra, A., Singh, D., Laversanne, M., Vignat, J., Gralow, J. R., Cardoso, F., Siesling, S., \u0026amp; Soerjomataram, I. (2022). Current and future burden of breast cancer: Global statistics for 2020 and 2040. \u003cem\u003eBreast (Edinburgh, Scotland)\u003c/em\u003e, \u003cem\u003e 66\u003c/em\u003e, 15\u0026ndash;23. https://doi.org/10.1016/j.breast.2022.08.010\u003c/li\u003e\n\u003cli\u003eSung H, Ferlay J, Siegel RL, Laversanne M, Soerjomataram I, Jemal A, et al. Global cancer statistics 2020: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA A Cancer J Clin 2021;71:209\u0026ndash;49 https://doi.org/10.3322/caac.21660[3]Orrantia-Borunda, E., Anchondo-Nu\u0026ntilde;ez, P., Acu\u0026ntilde;a-Aguilar, L. E., G\u0026oacute;mez-Valles, F. O., \u0026amp; Ram\u0026iacute;rez-Valdespino, C. A. (2022). Subtypes of Breast Cancer. In H. N. Mayrovitz (Ed.), \u003cem\u003eBreast Cancer\u003c/em\u003e. Exon Publications.\u003c/li\u003e\n\u003cli\u003eSmolarz, B., Nowak, A. Z., \u0026amp; Romanowicz, H. (2022). Breast Cancer-Epidemiology, Classification, Pathogenesis and Treatment (Review of Literature). Cancers, 14(10), 2569. https://doi.org/10.3390/cancers14102569\u003c/li\u003e\n\u003cli\u003eFolkerd, E., \u0026amp; Dowsett, M. (2013). Sex hormones and breast cancer risk and prognosis. \u003cem\u003eBreast (Edinburgh, Scotland)\u003c/em\u003e, \u003cem\u003e22 Suppl 2\u003c/em\u003e, S38\u0026ndash;S43. https://doi.org/10.1016/j.breast.2013.07.007Authors\u003c/li\u003e\n\u003cli\u003eMcGuire, A., Brown, J. A., Malone, C., McLaughlin, R., \u0026amp; Kerin, M. J. (2015). Effects of age on the detection and management of breast cancer. \u003cem\u003eCancers\u003c/em\u003e, \u003cem\u003e7\u003c/em\u003e(2), 908\u0026ndash;929. https://doi.org/10.3390/cancers7020815\u003c/li\u003e\n\u003cli\u003eNarod S. A. (2011). Hormone replacement therapy and the risk of breast cancer. Nature reviews. Clinical oncology, 8(11), 669\u0026ndash;676. https://doi.org/10.1038/nrclinonc.2011.110\u003c/li\u003e\n\u003cli\u003eZeinomar, N., Knight, J. A., Genkinger, J. M., Phillips, K. A., Daly, M. B., Milne, R. L., Dite, G. S., Kehm, R. D., Liao, Y., Southey, M. C., Chung, W. K., Giles, G. G., McLachlan, S. A., Friedlander, M. L., Weideman, P. C., Glendon, G., Nesci, S., kConFab Investigators, Andrulis, I. L., Buys, S. S., \u0026hellip; Terry, M. B. (2019). Alcohol consumption, cigarette smoking, and familial breast cancer risk: findings from the Prospective Family Study Cohort (ProF-SC). \u003cem\u003eBreast cancer research : BCR\u003c/em\u003e, \u003cem\u003e21\u003c/em\u003e(1), 128. https://doi.org/10.1186/s13058-019-1213-1\u003c/li\u003e\n\u003cli\u003eDuncan, W., \u0026amp; Kerr, G. R. (1976). The curability of breast cancer. \u003cem\u003eBritish medical journal\u003c/em\u003e, \u003cem\u003e2\u003c/em\u003e(6039), 781\u0026ndash;783. https://doi.org/10.1136/bmj.2.6039.781\u003c/li\u003e\n\u003cli\u003eIranmakani, S., Mortezazadeh, T., Sajadian, F. \u003cem\u003eet al.\u003c/em\u003e A review of various modalities in breast imaging: technical aspects and clinical outcomes. \u003cem\u003eEgypt J Radiol Nucl Med\u003c/em\u003e 51, 57 (2020). https://doi.org/10.1186/s43055-020-00175-5\u003c/li\u003e\n\u003cli\u003eGabriel F. Berriz, Frederick P. Roth, The Synergizer service for translating gene, protein and other biological identifiers, \u003cem\u003eBioinformatics\u003c/em\u003e, Volume 24, Issue 19, October 2008, Pages 2272\u0026ndash;2273, \u003cu\u003ehttps://doi.org/10.1093/bioinformatics/btn424\u003c/u\u003e\u003c/li\u003e\n\u003cli\u003eJ\u0026uuml;ri Reimand, Tambet Arak, Priit Adler, Liis Kolberg, Sulev Reisberg, Hedi Peterson, Jaak Vilo, g:Profiler\u0026mdash;a web server for functional interpretation of gene lists (2016 update), \u003cem\u003eNucleic Acids Research\u003c/em\u003e, Volume 44, Issue W1, 8 July 2016, Pages W83\u0026ndash;W89, https://doi.org/10.1093/nar/gkw199\u003c/li\u003e\n\u003cli\u003eUma Mudunuri, Anney Che, Ming Yi, Robert M. Stephens, bioDBnet: the biological database network, \u003cem\u003eBioinformatics\u003c/em\u003e, Volume 25, Issue 4, February 2009, Pages 555\u0026ndash;556, \u003cu\u003ehttps://doi.org/10.1093/bioinformatics/btn654\u003c/u\u003e\u003c/li\u003e\n\u003cli\u003eTanya Barrett, Stephen E. Wilhite, Pierre Ledoux, Carlos Evangelista, Irene F. Kim, Maxim Tomashevsky, Kimberly A. Marshall, Katherine H. Phillippy, Patti M. Sherman, Michelle Holko, Andrey Yefanov, Hyeseung Lee, Naigong Zhang, Cynthia L. Robertson, Nadezhda Serova, Sean Davis, Alexandra Soboleva, NCBI GEO: archive for functional genomics data sets\u0026mdash;update, \u003cem\u003eNucleic Acids Research\u003c/em\u003e, Volume 41, Issue D1, 1 January 2013, Pages D991\u0026ndash;D995, \u003cu\u003ehttps://doi.org/10.1093/nar/gks1193\u003c/u\u003e\u003c/li\u003e\n\u003cli\u003eMatthew E. Ritchie, Belinda Phipson, Di Wu, Yifang Hu, Charity W. Law, Wei Shi, Gordon K. Smyth, limma powers differential expression analyses for RNA-sequencing and microarray studies, Nucleic Acids Research, Volume 43, Issue 7, 20 April 2015, Page e47 https://doi.org/10.1093/nar/gkv007\u003c/li\u003e\n\u003cli\u003eJiang, X., Xu, Z., Du, Y. \u003cem\u003eet al.\u003c/em\u003e Bioinformatics analysis reveals novel hub gene pathways associated with IgA nephropathy. \u003cem\u003eEur J Med Res\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, 40 (2020). https://doi.org/10.1186/s40001-020-00441-2\u003c/li\u003e\n\u003cli\u003eSong, X., Du, R., Gui, H., Zhou, M., Zhong, W., Mao, C., \u0026amp; Ma, J. (2020). Identification of potential hub genes related to the progression and prognosis of hepatocellular carcinoma through integrated bioinformatics analysis. \u003cem\u003eOncology reports\u003c/em\u003e, \u003cem\u003e43\u003c/em\u003e(1), 133\u0026ndash;146. https://doi.org/10.3892/or.2019.7400\u003c/li\u003e\n\u003cli\u003eWang, Lulu. 2017. \u0026quot;Early Diagnosis of Breast Cancer\u0026quot; \u003cem\u003eSensors\u003c/em\u003e 17, no. 7: 1572. https://doi.org/10.3390/s17071572\u003c/li\u003e\n\u003cli\u003eHaidich A. B. (2010). Meta-analysis in medical research. \u003cem\u003eHippokratia\u003c/em\u003e, \u003cem\u003e14\u003c/em\u003e(Suppl 1), 29\u0026ndash;37.\u003c/li\u003e\n\u003cli\u003eEsquivel-Vel\u0026aacute;zquez, M., Ostoa-Saloma, P., Palacios-Arreola, M. I., Nava-Castro, K. E., Castro, J. I., \u0026amp; Morales-Montor, J. (2015). The role of cytokines in breast cancer development and progression. Journal of interferon \u0026amp; cytokine research : the official journal of the International Society for Interferon and Cytokine Research, 35(1), 1\u0026ndash;16. https://doi.org/10.1089/jir.2014.0026\u003c/li\u003e\n\u003cli\u003eGeindreau, M., Bruchard, M., \u0026amp; Vegran, F. (2022). Role of Cytokines and Chemokines in Angiogenesis in a Tumor Context. Cancers, 14(10), 2446. https://doi.org/10.3390/cancers14102446\u003c/li\u003e\n\u003cli\u003eShacter, E., \u0026amp; Weitzman, S. A. (2002). Chronic inflammation and cancer. Oncology (Williston Park, N.Y.), 16(2), 217\u0026ndash;232.\u003c/li\u003e\n\u003cli\u003eLiu, L., Wang, Y., Miao, L., Liu, Q., Musetti, S., Li, J., \u0026amp; Huang, L. (2018). Combination Immunotherapy of MUC1 mRNA Nano-vaccine and CTLA-4 Blockade Effectively Inhibits Growth of Triple Negative Breast Cancer. Molecular therapy : the journal of the American Society of Gene Therapy, 26(1), 45\u0026ndash;55. https://doi.org/10.1016/j.ymthe.2017.10.020\u003c/li\u003e\n\u003cli\u003eBurke, D. L., Frid, M. G., Kunrath, C. L., Karoor, V., Anwar, A., Wagner, B. D., Strassheim, D., \u0026amp; Stenmark, K. R. (2009). Sustained hypoxia promotes the development of a pulmonary artery-specific chronic inflammatory microenvironment. American journal of physiology. Lung cellular and molecular physiology, 297(2), L238\u0026ndash;L250. https://doi.org/10.1152/ajplung.90591.2008\u003c/li\u003e\n\u003cli\u003eLiu, H., Yang, Z., Lu, W., Chen, Z., Chen, L., Han, S., Wu, X., Cai, T., \u0026amp; Cai, Y. (2020). Chemokines and chemokine receptors: A new strategy for breast cancer therapy. Cancer medicine, 9(11), 3786\u0026ndash;3799. https://doi.org/10.1002/cam4.3014\u003c/li\u003e\n\u003cli\u003eAldinucci, D., \u0026amp; Colombatti, A. (2014). The inflammatory chemokine CCL5 and cancer progression. Mediators of inflammation, 2014, 292376. https://doi.org/10.1155/2014/292376\u003c/li\u003e\n\u003cli\u003eAlexander Plotnikov, Eldar Zehorai, Shiri Procaccia, Rony Seger, The MAPK cascades: Signaling components, nuclear roles and mechanisms of nuclear translocation, Biochimica et BiophysicaActa (BBA) - Molecular Cell Research, Volume 1813, Issue 9,2011,Pages 1619-1633, ISSN 0167-4889,https://doi.org/10.1016/j.bbamcr.2010.12.012.\u003c/li\u003e\n\u003cli\u003eSanten, R. J., Song, R. X., McPherson, R., Kumar, R., Adam, L., Jeng, M. H., \u0026amp; Yue, W. (2002). The role of mitogen-activated protein (MAP) kinase in breast cancer. The Journal of steroid biochemistry and molecular biology, 80(2), 239\u0026ndash;256. https://doi.org/10.1016/s0960-0760(01)00189-3\u003c/li\u003e\n\u003cli\u003eCoffelt, S. B., Wellenstein, M. D., \u0026amp; de Visser, K. E. (2016). Neutrophils in cancer: neutral no more. Nature reviews. Cancer, 16(7), 431\u0026ndash;446. https://doi.org/10.1038/nrc.2016.52\u003c/li\u003e\n\u003cli\u003eWang, Y., Chen, J., Yang, L., Li, J., Wu, W., Huang, M., Lin, L., \u0026amp; Su, S. (2019). Tumor-Contacted Neutrophils Promote Metastasis by a CD90-TIMP-1 Juxtacrine-Paracrine Loop. Clinical cancer research: an official journal of the American Association for Cancer Research, 25(6), 1957\u0026ndash;1969. https://doi.org/10.1158/1078-0432.CCR-18-2544\u003c/li\u003e\n\u003cli\u003eDorsam, R. T., \u0026amp; Gutkind, J. S. (2007). G-protein-coupled receptors and cancer. Nature reviews. Cancer, 7(2), 79\u0026ndash;94. https://doi.org/10.1038/nrc2069\u003c/li\u003e\n\u003cli\u003eKidani, Y., Nogami, W., Yasumizu, Y., Kawashima, A., Tanaka, A., Sonoda, Y., Tona, Y., Nashiki, K., Matsumoto, R., Hagiwara, M., Osaki, M., Dohi, K., Kanazawa, T., Ueyama, A., Yoshikawa, M., Yoshida, T., Matsumoto, M., Hojo, K., Shinonome, S., Yoshida, H., \u0026hellip; Sakaguchi, S. (2022). CCR8-targeted specific depletion of clonally expanded Treg cells in tumor tissues evokes potent tumour immunity with long-lasting memory. Proceedings of the National Academy of Sciences of the United States of America, 119(7), e2114282119. https://doi.org/10.1073/pnas.2114282119\u003c/li\u003e\n\u003cli\u003eTanaka, T., Nanamiya, R., Takei, J., Nakamura, T., Yanaka, M., Hosono, H., Sano, M., Asano, T., Kaneko, M. K., \u0026amp; Kato, Y. (2021). Development of Anti-Mouse CC Chemokine Receptor 8 Monoclonal Antibodies for Flow Cytometry. Monoclonal antibodies in immunodiagnosis and immunotherapy, 40(2), 65\u0026ndash;70. https://doi.org/10.1089/mab.2021.0005\u003c/li\u003e\n\u003cli\u003eKorbecki J, Grochans S, Gutowska I, Barczak K, Baranowska-Bosiacka I. CC Chemokines in a Tumor: A Review of Pro-Cancer and Anti-Cancer Properties of Receptors CCR5, CCR6, CCR7, CCR8, CCR9, and CCR10 Ligands. International Journal of Molecular Sciences. 2020; 21(20):7619. https://doi.org/10.3390/ijms21207619\u003c/li\u003e\n\u003cli\u003eJie Zhang \u0026amp; Dawei Hu (2021) miR-1298-5p Influences the Malignancy Phenotypes of Breast Cancer Cells by Inhibiting CXCL11, Cancer Management and Research, 13:, 133-145, DOI: 10.2147/CMAR.S279121\u003c/li\u003e\n\u003cli\u003eHyun Jung Hwang, Ye-Rim Lee, Donghee Kang, Hyung Chul Lee, Haeng Ran Seo, Ji-Kan Ryu, Yong-Nyun Kim, Young-Gyu Ko, Heon Joo Park, Jae-Seon Lee, Endothelial cells under therapy-induced senescence secrete CXCL11, which increases aggressiveness of breast cancer cells, Cancer Letters,Volume 490,2020,Pages 100-110,ISSN 0304-3835,https://doi.org/10.1016/j.canlet.2020.06.019.\u003c/li\u003e\n\u003cli\u003eBowen Chen, Shuyuan Zhang, Qiuyu Li, Shiting Wu, Han He, Jinbo Huang; Bioinformatics identification of CCL8/21 as potential prognostic biomarkers in breast cancer microenvironment. Biosci Rep 27 November 2020; 40 (11): BSR20202042. doi: \u003cu\u003ehttps://doi.org/10.1042/BSR20202042\u003c/u\u003e\u003c/li\u003e\n\u003cli\u003eKaran Dev, CCL23 in Balancing the Act of Endoplasmic Reticulum Stress and Antitumor Immunity in Hepatocellular Carcinoma, Frontiers in Oncology,volume 11, 2021,doi:10.3389/fonc.2021.727583\u003c/li\u003e\n\u003cli\u003ePark, J., Zhang, X., Lee, S. K., Song, N. Y., Son, S. H., Kim, K. R., Shim, J. H., Park, K. K., \u0026amp; Chung, W. Y. (2019). CCL28-induced RAR\u0026beta; expression inhibits oral squamous cell carcinoma bone invasion. The Journal of clinical investigation, 129(12), 5381\u0026ndash;5399. https://doi.org/10.1172/JCI125336\u003c/li\u003e\n\u003cli\u003eMickanin, C. S., Bhatia, U., \u0026amp; Labow, M. (2001). Identification of a novel beta-chemokine, MEC, down-regulated in primary breast tumors. International journal of oncology, 18(5), 939\u0026ndash;944. https://doi.org/10.3892/ijo.18.5.939\u003c/li\u003e\n\u003cli\u003eDegos, C., Heinemann, M., Barrou, J., Boucherit, N., Lambaudie, E., Savina, A., Gorvel, L., \u0026amp; Olive, D. (2019). Endometrial Tumor Microenvironment Alters Human NK Cell Recruitment, and Resident NK Cell Phenotype and Function. Frontiers in immunology, 10, 877. https://doi.org/10.3389/fimmu.2019.00877\u003c/li\u003e\n\u003cli\u003eHossain, S.M.M., Khatun, L., Ray, S. et al. Pan-cancer classification by regularized multi-task learning. Sci Rep 11, 24252 (2021). https://doi.org/10.1038/s41598-021-03554-8\u003c/li\u003e\n\u003cli\u003eBhar, A., Haubrock, M., Mukhopadhyay, A. et al. Coexpression and coregulation analysis of time-series gene expression data in estrogen-induced breast cancer cell. Algorithms Mol Biol 8, 9 (2013). https://doi.org/10.1186/1748-7188-8-9\u003c/li\u003e\n\u003cli\u003eA. Mukhopadhyay and M. Mandal, \u0026quot;Identifying Non-Redundant Gene Markers from Microarray Data: A Multiobjective Variable Length PSO-Based Approach,\u0026quot; in IEEE/ACM Transactions on Computational Biology and Bioinformatics, vol. 11, no. 6, pp. 1170-1183, 1 Nov.-Dec. 2014, doi: 10.1109/TCBB.2014.2323065\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Breast Cancer, Machine Learning, LASSO Regression, Microarray, Prognosis, PPI Network, Signature Genes, Biomarkers","lastPublishedDoi":"10.21203/rs.3.rs-4363199/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4363199/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eCancer is a genetic disease, where gene alterations play a significant role in the disease onset and pathogenesis. Analysis of the underlying gene interaction pathways could reveal new biomarkers and could also potentially help in the development of targeted drugs for therapeutics. Microarray techniques have emerged as powerful tools capable of simultaneously measuring the expression levels of thousands of genes, making them invaluable in cancer biology research. However, the processing of the resultant datasets poses significant challenges due to their high dimensionality. Also, feature extraction becomes essential to discern the crucial features within these extensive datasets. To mitigate these difficulties advanced computational techniques like Machine Learning (ML) could be instrumental. LASSO- regression-based classification is an advanced ML technique that can help in feature selection by evaluating individual parameters like genes.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eThis study focuses on uncovering key prognostic genes for breast cancer using a combination of LASSO regression-based classifier and statistical bioinformatics models. Differentially expressed genes (DEGs) were identified using the \"Limma\" package in R, and significant genes were further filtered using the LASSO-based classifier significance coefficient. Genes common to both methods were considered as the focus of this study. Additionally, Protein-Protein Interaction (PPI) networks of these key genes were constructed using STRING, and hub genes, significant modules, and associated genes were identified using Cytoscape.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eThis study identified CCR8, CXCL11, CCL23, CCL24, CCL28, and CCL21 as signature prognostic genes for breast cancer, revealing a strong association between chemokines and breast cancer pathogenesis. Extensive literature searches were conducted to validate and confirm their prognostic significance in the disease.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eThese findings are pivotal for enhancing our comprehension of the pathways involved in breast cancer. Additionally, they hold promise as novel biomarkers for diagnostic purposes and may also reveal significant therapeutic targets for the management of breast cancer. The codes are available in the following GitHub repository: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/guhasouvik/LASSO_BRCA.git\u003c/span\u003e\u003cspan address=\"https://github.com/guhasouvik/LASSO_BRCA.git\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e","manuscriptTitle":"LASSO Based Analysis for Prediction of Prognostic Signature Genes Associated with Breast Cancer","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-05-15 15:39:29","doi":"10.21203/rs.3.rs-4363199/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"a59b29ce-add2-4441-865a-03f52eaf51a1","owner":[],"postedDate":"May 15th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-08-28T06:54:56+00:00","versionOfRecord":[],"versionCreatedAt":"2024-05-15 15:39:29","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4363199","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4363199","identity":"rs-4363199","version":["v1"]},"buildId":"GqpaHPwrfC8PjnIFayRh5","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.