Integration of RNA-Seq and proteomics data identifies glioblastoma multiforme surfaceome signature

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background: Glioblastoma multiforme (GBM) is a highly lethal, stage IV brain tumour with a prevalence of approximately 2 per 10000 people globally. The cell surface proteins or surfaceome serve as an information gateway in many oncogenic signalling pathways and are important in modulating cancer phenotypes. Dysregulation of surfaceome expression and activity have been shown to promote tumorigenesis. The expression of GBM surfaceome is a case in point; OMICS screening in cell-based system identified that this sub-proteome is largely perturbed in GBM. Additionally, since these cell surface proteins have ‘direct’ access to drugs, they are appealing targets for cancer therapy. However, a comprehensive aberrant GBM surfaceome landscape has not been fully defined. Thus, this study aimed to define GBM-specific surfaceome genes and identify key cell-surface genes that could potentially be developed as novel GBM biomarkers for therapeutic purposes. Methods: We integrated the RNA-Seq data from TCGA GBM (n=166) and GTEx normal brain cortex (n=408) databases to identify the significantly dysregulated surfaceome in GBM. This was followed by integrative analysis that combines transcriptomics, proteomics and protein-protein interaction network data to prioritize the high-confidence GBM surfaceome signature. Results: Of the 2,381 significantly dysregulated genes in GBM, 395 genes were classified as surfaceome. Via the integrative analysis, we identified 6 high-confidence GBM molecular signature, HLA-DRA, CD44, SLC1A5, EGFR, ITGB2, PTPRJ, which were significantly upregulated in GBM. The expression of these genes were validated in an independent transcriptomics database, which confirmed their upregulated expression in GBM. Importantly, high expression of CD44, PTPRJ and HLA-DRA is significantly associated with poor disease-free survival. Last, using the Drugbank database, we identified several clinically-approved drugs targeting the GBM molecular signature suggesting potential drug repurposing. Conclusions: In summary, we identified and highlighted the key GBM surface-enriched repertoires that could be biologically relevant in supporting GBM pathogenesis. These genes could be further interrogated experimentally in future studies that could lead to efficient diagnostic/prognostic markers or potential treatment options for GBM.
Full text 144,433 characters · extracted from preprint-html · click to expand
Integration of RNA-Seq and proteomics data identifies glioblastoma multiforme surfaceome signature | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research article Integration of RNA-Seq and proteomics data identifies glioblastoma multiforme surfaceome signature Saiful Effendi Syafruddin, Wan Fahmi Wan Mohamad Nazarie, Nurshahirah Ashikin Moidu, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-46071/v2 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 23 Jul, 2021 Read the published version in BMC Cancer → Version 2 posted 7 You are reading this latest preprint version Show more versions Abstract Background: Glioblastoma multiforme (GBM) is a highly lethal, stage IV brain tumour with a prevalence of approximately 2 per 10000 people globally. The cell surface proteins or surfaceome serve as an information gateway in many oncogenic signalling pathways and are important in modulating cancer phenotypes. Dysregulation of surfaceome expression and activity have been shown to promote tumorigenesis. The expression of GBM surfaceome is a case in point; OMICS screening in cell-based system identified that this sub-proteome is largely perturbed in GBM. Additionally, since these cell surface proteins have ‘direct’ access to drugs, they are appealing targets for cancer therapy. However, a comprehensive aberrant GBM surfaceome landscape has not been fully defined. Thus, this study aimed to define GBM-specific surfaceome genes and identify key cell-surface genes that could potentially be developed as novel GBM biomarkers for therapeutic purposes. Methods: We integrated the RNA-Seq data from TCGA GBM (n=166) and GTEx normal brain cortex (n=408) databases to identify the significantly dysregulated surfaceome in GBM. This was followed by integrative analysis that combines transcriptomics, proteomics and protein-protein interaction network data to prioritize the high-confidence GBM surfaceome signature. Results: Of the 2,381 significantly dysregulated genes in GBM, 395 genes were classified as surfaceome. Via the integrative analysis, we identified 6 high-confidence GBM molecular signature, HLA-DRA, CD44, SLC1A5, EGFR, ITGB2, PTPRJ, which were significantly upregulated in GBM. The expression of these genes were validated in an independent transcriptomics database, which confirmed their upregulated expression in GBM. Importantly, high expression of CD44, PTPRJ and HLA-DRA is significantly associated with poor disease-free survival. Last, using the Drugbank database, we identified several clinically-approved drugs targeting the GBM molecular signature suggesting potential drug repurposing. Conclusions: In summary, we identified and highlighted the key GBM surface-enriched repertoires that could be biologically relevant in supporting GBM pathogenesis. These genes could be further interrogated experimentally in future studies that could lead to efficient diagnostic/prognostic markers or potential treatment options for GBM. Oncology Cancer Biology Differentially expressed genes protein-protein interaction cell surface proteins network analysis TCGA GTEx Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Background Glioblastoma multiforme (GBM) is the most common and lethal tumour of the central nervous system in adults [1]. Despite decades of efforts to tackle this disease, the median survival rate of GBM patients is still not improving [2]. GBM patients have an average life expectancy of 15 months post-diagnosis and the 5-years survival rate is less than 3% [3]. The standard-of-care GBM treatment generally consists of maximal safe surgical resection followed by radiotherapy and concomitant chemotherapy. However rapid post-treatment relapse and high intra-tumoral heterogeneity that could either arise naturally during disease progression or treatments-induced have made this disease intractable and more challenging to treat [4, 5]. Therefore, there is a pressing need for better and efficient diagnostic and therapeutic strategies for this disease. Temozolomide, an orally-administered DNA-alkylating drug, is the current and commonly used chemotherapy agent to treat GBM in the clinic [6]. This combination treatment of temozolomide and radiotherapy is referred to as the Stupp regimen and it is widely used as the standard-of-care for the treatment of GBM. The landmark study showed that the combination of radiotherapy and concomitant chemotherapy with temozolomide improve the patient's prognosis compared to radiotherapy alone (median survival of 14.6 months vs 12.1 months, respectively) [6]. Alternative GBM treatment options such as the VEGF-targeting monoclonal antibody Bevacizumab, other DNA alkylating agents such as lomustine and carmustine implants, alternating electric field therapy and the checkpoint blockade inhibitor have thus far yielded low efficacy in treating GBM [2, 7, 8]. The Cancer Genome Atlas (TCGA) comprehensive GBM molecular characterizations have identified significant genetic alterations in several important oncogenic signalling pathways such as the RTK/Ras/PI3K (88%), p53 (87%) and pRB signalling pathways (78%) in GBM patients [9]. Several clinical trials are currently ongoing that aim to target these altered GBM oncogenic signalling pathways components using small molecule inhibitors and/or monoclonal antibodies. However, the results thus far were far from satisfactory [10]. This seems to suggest that instead of using a single agent targeting a specific component or pathway, novel treatments should consider the administration of several inhibitors targeting multiple different pathways. The cell surface proteins or surfaceome serve as an information gateway that integrates and transduces extracellular cues into intracellular signalling cascades or vice versa . Surfaceome also play important role in cell adhesion and migration which are among the critical processes during tumorigenesis. Indeed, aberrant surfaceome expression and activity are frequently observed in many cancer types and therefore are good candidates for cancer diagnostic or biomarkers as well as therapeutic targets. Recent evidence has demonstrated that 56% of cell surface proteins are differentially expressed in GBM which are also present in cerebrospinal fluid or plasma, suggesting their potential use as biomarkers [11]. Of note, surfaceome expression is more dynamic than intracellular proteins and they could be sometimes cell type-specific [12, 13]. Mass spectrometry analysis showed that the average surfaceome size in brain cancer cell lines is higher than other cancer types [12]. Thus, surfaceome genes in GBM may hold the key to understand GBM pathogenesis and drug responsiveness, in which targeting these genes may unravel potential ‘ druggable ’ stage in GBM pathways. A comprehensive overview of GBM surfaceome landscape has not been fully defined. Therefore, this study aimed to characterize the GBM surfaceome genes expression profile by unifying the two large RNA-Seq datasets from the TCGA (GBM) and GTEx (normal brain). We integrated and performed differential gene expression analysis on these two datasets because of the low number of normal brain tissue samples in the TCGA database. Previously annotated surfaceome gene set was employed to filter and identify the significant differentially expressed surfaceome genes in GBM. To further prioritize the high-confidence GBM cell surface signature, we integrated our transcriptomics analysis with GBM tissues and cell surface proteomics, and PPI hub gene analysis. Collectively, we identified a list of upregulated surfaceome genes in GBM that include CD44, PTPRJ and HLA-DRA in which their biological relevance in supporting GBM pathogenesis could be comprehensively investigated in the future studies for the development of novel GBM diagnostic/prognostic or therapeutic strategies. Methods TCGA and GTEx data acquisition, normalization and quality control: The analysis combined of the TCGA-GBM and GTEx normal brain RNA-Seq read count data. The GBM RNA-Seq gene raw read counts from TCGA were downloaded from Genomics Data Commons Data Portal ( https://portal.gdc.cancer.gov ). GTEx data were used for the normal brain tissues. The GTEx data used for the analyses described in this manuscript were obtained from the GTEx Portal on 29/03/19. We downloaded RNA-Seq gene raw read counts (from the cortex, frontal cortex, anterior cingulate cortex) from GTEx portal ( https://gtexportal.org/home/datasets ). This allows us to perform the analysis of the differentially expressed gene on the 166 samples of GBM tumour from TCGA and 408 samples of normal brain tissues data from GTEx. The RNA-Seq raw read counts pre-processing steps involve are data filtering and data normalization. The normalization process of both data set was then performed by using mean as gene-level normalization using log 2 -counts per million where raw data are adjusted to account for factors that will prevent a direct comparison of expression measures and to safeguard the expression distributions are similar for each sample across the whole experiment. Data that unlikely to be informative or simply erroneous data will be removed by using variance filter (less than 15) and low abundance (less than 4). Cell surface gene set classification and analysis: The identified differentially expressed genes (DEGs) of glioblastoma were classified into cell-surface genes set as discussed in the main text (See Results 2.4). The classification of the gene sets was performed based on the mapping set of DEGs with this resource. Other genes, which did not map to this resource were removed from the final dataset. Differential gene expression: DEGs analysis was performed using NetworkAnalyst [14], a web-based application tool for visualizing molecular and entity interactions. This platform utilizes the statistical method on data comparison from R package, limma to identify genes whose expression is different. Genes that have adjusted p-value <0.05 and log2 fold change |2| were considered as statistically significant DEGs. Functional annotation and pathway analysis: The enrichment analysis of the identified glioblastoma associated genes was performed using DAVID ( https://david.ncifcrf.gov/ ), a web-based tool for analyzing functional gene analysis. The tool comprises databases from various public resources for biological analysis. The enrichment analysis such as GO and KEGG pathways were performed with top results as per gene counts. Identification of hub genes through PPI network analysis: A biological database for known and predicted protein-protein interactions called IMEx interactome database ( https://www.imexconsortium.org ) was used to construct the protein-protein interaction (PPI) of the DEGs. The network of interacting proteins was extracted and visualized using NetworkAnalyst. The top 87-gene modules of highly interacting gene clusters among the DEG were found with default parameters. For the classified gene sets, the PPI network was constructed and the network topological parameters i.e. degree and betweenness centrality were calculated. Co-expression network of CD44: Co-expression analysis was performed using Graphia Professional ( https://kajeka.com/graphia-professional/ ), previously known as BioLayout Express 3D [15] using raw read counts and then saved as an ‘‘.expression’’ file. This contains a unique identifier for each row of data. Following import into Graphia Professional, a pairwise Pearson correlation matrix was calculated thereby performing a gene vs. gene comparison of the expression profile of each gene. All Pearson correlations where r>0.7 were saved to a ‘‘.pearson’’ file. Based on a user-defined threshold of r>0.75, an undirected network graph of the data was generated. In this context, nodes represent individual genes and the edges between them represent Pearson correlation coefficients above the selected threshold (r>0.75). CD44 was selected along with its neighbour in the network, representing CD44 co-expression partners. The class set of CD44 co-expressed genes were visualized to compare the expression values in this class set with genes in normal samples. Results Patients’ characteristics of TCGA and GTEx: We utilized the publicly available TCGA and GTEx RNA-Seq database as our primary sources of GBM tumour and normal brain tissue transcriptomic data, respectively. We downloaded the datasets containing RNA-Seq gene expression profiles and clinical information of 166 patients from TCGA-GBM and 408 normal brain tissues from GTEx database. The combined data were stratified based on gender, age and treatment as shown in Table 1. Out of a total of 166 GBM cases, 104 cases (62.7%) were male and 56 cases (33.7%) were female. GBM is more prevalent in patients aged ≥ 60 years old which accounts for 42.8% of total cases in the TCGA GBM cohort. Fifty-two patients (31.3%) have undergone treatments whereas 62.1% of cases did not have any treatment data. Unfortunately, the clinical data for the GTEx normal brain samples are not publicly available. Identification of differentially expressed genes in glioblastoma: The analysis pipeline employed in this study is depicted in Figure 1. Briefly, the RNA-Seq raw read counts from the two large compendiums, TCGA and GTEx were utilized to identify the differentially expressed genes between GBM and normal brain tissues. Since most GBM cases are generally found in the supratentorial region of the brain such as the cerebral hemisphere [16], we only extracted the RNA-Seq profiles of this region namely the cortex, frontal cortex, anterior cingulate cortex as per GTEx description. We performed t-distributed stochastic neighbour embedding (t-SNE) analysis to reflect the directionality of transcripts expression among GBM tumour and normal brain tissues read count values. The t-SNE plot showed that all RNA-Seq profiles of all GTEx cortex region clustered together while the GBM RNA-Seq profiles form a separate cluster, thus confirming distinct expression patterns between these groups (Figure 2A). In total, RNA expression data from 18,021 genes were obtained from these combined TCGA and GTEx dataset but only 13,548 genes passed the quality control check. By applying the cut-off criteria log 2 fold change |2| and adjusted p-value <0.05, we identified 2381 genes as significantly differentially expressed genes (DEGs) in GBM, of which 648 genes were upregulated and 1733 genes were downregulated (Figure 2B). The detailed information of the differential gene expression analysis is listed in Supplementary Table S1. Functional enrichment analysis and classification of DEGs: The significant DEGs were then subjected to functional enrichment analysis using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) tools to define their properties and putative biological relevance in GBM. Interestingly, the GO cellular component analysis of both upregulated and downregulated DEGs showed enrichment of cell surface and membrane-associated proteins (Supplementary Figure S1A and S1B). The KEGG pathway enrichment indicated that the upregulated DEGs are involved in pathways related to infectious diseases, pathway in cancer and cell adhesion (Supplementary Figure S1C). Downregulated genes mainly involve in neuroactive ligand-receptor interaction and major cellular signalling pathways (Supplementary Figure S1D). Identification of GBM cell-surface antigen candidates: The DEGs were then further filtered and classified into the surfaceome gene set as previously defined by Bausch-Fluck et. al [13], Cunha et. al [17] and Lee et. al [18]. These studies utilized different criteria and stringency in curating the surfaceome gene list. From the overall DEGs in GBM, we identified 395 common cell surface genes within these three surfaceome definitions, including 124 upregulated and 271 downregulated genes (Supplementary Figure S2A and Supplementary Table S2). We further classified the surfaceome according to their main subclasses, which are receptors, transporters, enzymes, miscellaneous and unclassified, as previously reported by Almén et. al [19]. Among the defined surfaceome subclasses, 42.8% of the significant differentially expressed surfaceome in GBM belong to the receptor subclass (Supplementary Figure S2B). KEGG analysis of the GBM-enriched cell surface proteins identified pathways related to immune defence and infectious disease pathways while GBM-deficient cell surface genes are enriched in pathways related to in neuroactive ligand-receptor interaction and major cellular signalling pathways (Supplementary Figure S3A and S3B). These findings are almost similar to the enrichment analysis of overall DEGs in GBM (Supplementary Figure S1C and S1D) suggesting that surfaceome has significant roles in dictating GBM cellular activities. Identification of GBM cell-surface signature by integration of proteomics and transcriptomics data analysis: Thus far, we have (i) classified the overall DEGs in GBM using transcriptomics data and (ii) highlighted the differentially expressed cell-surface genes in GBM. Even though this transcriptomics analysis is very informative for biomarker discovery, we aimed to add another layer of analysis to select for a more high-confidence cell surface signature for GBM. To attain this, we integrated our transcriptomics analysis data with the publicly available proteomics data. This integration will validate the cell surface genes prediction and eliminate the possible discrepancy between the expression levels of mRNAs and proteins due to post-transcriptional and post-translational modifications. Thus, we gathered the publicly available quantitative mass spectrometry analysis data for both GBM tissues and cell lines. We postulated that GBM tissues and cell lines might have different cell surface repertoires and therefore it is important to stratify between these two sources. Additionally, GBM cell lines cell surface signature, as identified in this present study, could be validated experimentally in future functional studies. Mass spectrometry analysis of five GBM cell lines revealed the upregulation of EGFR, CD44, PTPRJ, SLC1A5, F2R, and TSPAN6 proteins in these samples [12], whereby the expression level of these proteins were in concordance with our transcriptomics data analysis (Figure 3). For tissue proteomics, we found several studies that performed comparative GBM vs. normal brain tissues proteome profiling [11, 20–23]. However, some of these studies either identified only a limited number of proteins or the data are not downloadable. Only one study by Polisetty et al. that has identified a large number of proteins in their proteome profiling study that included 1834 high-confidence membrane proteins with more than 2-fold change [11]. We, therefore, used this dataset where we performed integrative analysis with our analyzed transcriptomics data and identified 10 overlapped genes, MRC2, FCGR3A, HLA-DRA, CD44, CD74, MSR1, CD163, EGFR, ITGB2, PTPRZ1 (Figure 3). The mRNA expression levels correlated with the protein expression levels except the PTPRZ1 where the mRNA levels showed upregulation while proteomics data showed downregulation (Supplementary Table S3 and S4). In total, there are 14 genes from the combined tissues and cell lines proteomics that overlapped with our transcriptomics data (Figure 3). It is important to note that proteins identification in mass spectrometry can be limiting due to protein isolation methods, proteins solubility, and other intrinsic variations that affect the proteins abundance as well as the sensitivity and detection capability of the MS instrumentation [24, 25]. Thus, these limitations may underestimate the results between transcriptomics prediction and proteomics discoveries. Surfaceome protein-protein interaction network cluster analysis and prioritization of high-confidence GBM cell surface markers: We set out to further analyze the GBM-enriched cell surface markers using protein-protein interaction (PPI) network analysis. This is to better understand the interplay between the cell surface genes within the identified DEGs as well as with other genes. More importantly, this would enable us to further select the genes that are highly interconnected from the integrated proteomics and transcriptomics analysis. Network analysis of the identified differentially expressed cell surface protein genes was performed using NetworkAnalyst [14] to determine the relationship between genes according to the network topological parameters such as degree and betweenness. These parameters reflect the role and property of proteins within the network. The nodes and edges in the PPI network represent the proteins and their interactions, respectively. The GBM-enriched cell surface proteins network contains 1,321 nodes and 1,767 edges interactions based on a number of validated features including functional experiments, co-expression analysis, text mining, neighbourhood, gene fusion and databases (Figure 4A). We identified 87-gene modules of clusters and the top cluster genes with more than 30 interactions include VCAM1, EGFR, TGFBR1, CD44, NGFR, ITGB2, DCC, PTPRJ, ANBCA1, HLA-DRA, CCR5 and CSF1R (Figure 4A and Supplementary Table S5). Vascular Cell Adhesion Molecule 1 (VCAM1) has the highest interacting cluster as it was found to have 426 degree with 422,712.18 betweenness score. We subsequently mapped the 14 genes identified from the integrated transcriptomics and proteomics data analysis (Figure 3) with the top genes that have at least 20 interactions from the PPI network analysis. We found 6 genes that were in common between these two datasets which represent the high-confidence GBM predictive surfaceome markers (Figure 4B). Validation of high-confidence GBM signature gene and survival-expression correlation analysis: Next, we validated the expression profiles of the identified 6 high-confidence cell surface markers using an independent database, Gene Expression Profiling Interactive Analysis (GEPIA) [26]. GEPIA also combines the TCGA and GTEx gene expression data that were processed from raw reads count and unified using its own pipeline. In line with our findings, the identified GBM cell surface signature genes were confirmed to be significantly upregulated in the GBM GEPIA database (Supplementary Figure S4A – S4F). To investigate whether the expression level of these signature genes would modulate/influence GBM patients’ prognosis, we first performed the overall survival analyses on GBM patients who had high or low expression of each of these 6 genes (Supplementary Figure S5A – S5F). However, there were no significant differences in the overall survival between patients who had high or low expression of these 6 individual prioritized genes. Since GBM patients have low overall survival rate (average <2 years’ survival post-diagnosis), we postulated that it would be more appropriate to look at the disease-free survival endpoint rather than the overall survival. Moreover, the overall survival endpoint is more suited for a longer follow-up period (typically 5 years) for the data to be meaningful [27]. Hence, we examined the disease-free survival profile of the GBM patients in a similar fashion. We found that high expression of CD44 , PTPRJ and HLA-DRA were significantly correlated (p<0.05) with poor disease-free survival in GBM patients (Supplementary Figure S6A – S6F). In addition to performing survival analysis on the individual gene, we also assessed whether combining the level of all 6 GBM signature genes as a group could predict the GBM patients’ overall survival and disease-free survival. We observed that there was no statistically significant difference in the overall survival and disease-free survival between patients who had high expression and low expression of the signature group (Supplementary Figure S7A – S7B). Interestingly, by combining only CD44 , PTPRJ and HLA-DRA in the gene signature, we found that subjects with high expression of this signature group had significantly poor disease-free survival (p<0.0084) compared to patients who had low expression of these genes (Supplementary Figure S8B). However, there was still no significant difference in the overall survival between GBM patients in this signature group (Supplementary Figure S8A). Co-expression network of CD44 : CD44 is a transmembrane receptor and has multifaceted functions in both normal and disease physiology. OMICS studies have identified CD44 to be overexpressed in many types of cancer including glioblastoma [28, 29]. Based on our analysis, CD44 seems particularly important as it can be both identified in transcriptomics and proteomics-based approaches, among the top hub gene and whose high expression correlate with poor disease-free survival. We performed a co-expression network analysis to further interrogate its association with other genes using our transcriptomics. The nodes represent in the network analysis represent genes, while the edges represent Pearson correlation above r>0.75. The neighbouring genes connected to CD44 was extracted and shown in Figure 5A. There are 27 genes in this complex connected to CD44. Among the highly correlated genes are ELK3, CLIC4, GALNT2, TNC, and VIM. All genes in this CD44 co-expression cluster are highly expressed in GBM compared to normal brain samples (Figure 5B), further corroborating the biological relevance of CD44 in supporting GBM pathogenesis. Identification of drugs targeting GBM signature and CD44 network: We next determined whether there are any clinically approved drugs targeting the identified high confidence GBM cell surface markers (Supplementary Figure S4) and components of the constructed CD44 co-expression network (Figure 5A). To achieve this objective, we utilized the Drugbank database ( https://go.drugbank.com/ ) and our searches yielded several approved drugs that can be potentially effective or repurposed to target CD44, EGFR, C1R, CALR and TNFSR1A (Table 2). Hyaluronic acid, for example, is a clinically approved ligand for CD44 and this drug has been administered in the clinic to treat disease such as osteoarthritis [30]. Excessive hyaluronic acid administration has been demonstrated to inhibit tumour growth, possibly by impeding cell-cell interaction [31]. Besides, the use of nanomaterials to enhance the efficiency of hyaluronic acid delivery for cancer therapy is also actively being explored [32, 33]. Thus, the promising features of hyaluronic acid in mediating enhanced drugs or genes delivery to cancer cells via the overexpressed CD44 receptor could potentially be applied and developed for novel GBM therapeutic strategies. In regards to EGFR, several inhibitors and monoclonal antibodies have already been therapeutically approved to target this protein due to its roles as an important driver of tumorigenesis in many cancer types [34]. Moreover, of the 28 components of CD44 co-expression network (Figure 5A), only C1R, CALR, and TNFSR1A have drugs that can modulate them (Table 2). For instance, 3 drugs can be used or repurposed to target C1R. The pharmacological activity of Palivizumab to bind C1R subcomponent is under investigation, whereas the conestat alfa and human C1-esterase inhibitor can directly target C1R subcomponent and disrupt the complement system activation. Discussion The surfaceome comprise cellular frontiers that permit/inhibit signal transduction as well as playing important roles in modulating cells proliferation, migration and invasion, and cells-cells interaction. The surfaceome can organize themselves at a nanoscale resolution [35]. This spatiotemporal nanoscale organization could define the cell identity and phenotypes, and capacity to communicate with microenvironments such as the extracellular matrix, growth factors, hormones and drugs. Due to their accessibility on the cell membrane, surfaceome proteins are ideal candidates for biomarkers and often targeted for drugs development. Over 50% of drugs curated in the DrugBank target the surfaceome. In addition to their ubiquitous expression on the plasma membrane, the extracellular stalks of these cell surface proteins can be cleaved and released into the bloodstream, making them as suitable targets for blood-based diagnostics. Surfaceome can also be draped with glycans during post-translational modifications, which will mediate their interaction with other proteins that reside on either the same or neighbouring cells as well as with the microenvironments [35] . Dysregulated surfaceome expressions and functions have been shown to promote tumour formation and progression [36]. Therefore, scientists have begun profiling and cataloguing surfaceome in various types of cancers [37–40]. These cell surface proteins can be elevated in cancer cells in which they can respond to the increased level of growth factors, rendering cancer cells to sustain their infinite proliferative capabilities [41] and interact with the microenvironment that could either directly or indirectly modulate the tumour growth and metastatic capabilities [42]. The GBM transcriptomics dataset have been previously utilized to uncover genes that support GBM pathogenesis as well as genes that have potential prognostic values [43–45]. For example, Nicolasjilwan et al. analyzed the TCGA database to predict the survival of GBM patients based on clinical features, MRI images genomics alterations [43]. However, most TCGA GBM differential genes expression analyses either relied on low number of normal brain tissue samples, in which the TCGA GBM cohort contained only 5 normal brain tissues RNA-Seq data, or the data were combined with the GBM TCGA microarray data. This might create an imbalance that would lead to inaccuracy or bias in the downstream analysis. Hence, to increase the robustness of this study in identifying the significantly upregulated GBM surfaceome repertoire, we included the normal brain tissues GTEx RNA-Seq database TCGA in our analysis. On a similar scale, the GTEx studies have performed genes expression profiling in more than 11,000 samples across multiple human tissues from nearly 1,000 healthy donors. We compared the TCGA GBM and normal cortex GTEx RNA-seq data and identified 2,381 significant differentially expressed genes in GBM, in which 648 were upregulated and 1,733 downregulated genes. In agreement with previous GBM proteomics profiling study [12], the GO cellular compartment analysis showed that most of the dysregulated genes in GBM encode for the cell surface proteins, suggesting the importance of cell surface proteins in GBM pathogenesis. Of the 2,381 significant DEGs in GBM, 395 genes encode for cell surface proteins, in which 124 and 271 genes were found to be significantly upregulated and downregulated, respectively. Interestingly, receptor subclass was the predominant dysregulated genes in GBM, suggesting the crucial roles of cell surface receptors in supporting GBM pathogenesis. This was indeed in line with several studies reporting the implications of cell surface receptors dysregulation in the pathogenesis of many cancer types [46]. For this reason, the development of cancer treatment strategies have been revolved around targeting the cell surface receptors such as the receptor tyrosine kinases (RTKs) [47] and G protein-coupled receptors (GPCRs) [48]. Therefore, targeting the cell surface proteins particularly the receptor subclass could potentially be further explored as novel GBM therapeutic options. Robust cancer biomarkers are those that could be reproducibly identified by multi-omics platforms or reported in several different studies. To this end, we integrated the analyzed transcriptomics data with publicly available GBM proteomics data to prioritize for high-confidence cell surface proteins. Also, due to post-transcriptional and post-translational modifications, the mRNAs expression level are sometimes not correlated with their respective protein expression levels [49]. After mapping the prioritized genes from the transcriptomics-proteomics integrative analysis with the PPI network analysis data, we identified 6 genes; HLA-DRA, CD44, SLC1A5, EGFR, ITGB2, PTPRJ , whereby we considered these genes as the high-confidence GBM predictive surface markers. Overall survival analyses revealed that there were no significant difference in the overall survival between patients who had high and low expression of these 6 genes, either the genes were analyzed individually or when combined together. However, when looking at the disease-free survival, patients who had high expression of CD44, PTPRJ, and HLA- DRA, either individually or as a group, had significantly poor disease-free survival (Supplementary Figure 6 and 8B) compared to subjects with low expression of the genes. These findings indicate that these 3 genes, CD44, PTPRJ, and HLA-DRA, could potentially be developed as GBM prognostic markers in the clinic. In addition to identifying the already known GBM drivers like CD44 and EGFR, our integrative analysis approach has also enabled us to identify potential novel genes that have not either been reported or thoroughly discussed in the context of GBM. For instance, within the 6 GBM signature genes, ITGB2 has not been widely associated with the pathogenesis of GBM. ITGB2 encodes for cell surface protein that is important in regulating cell adhesion and cell-surface mediated signalling [50]. Hence overexpression of this protein is relevant in promoting cancer growth possibly by modulating cancer cells adhesive and migratory properties, and the pro-oncogenic signalling cascades. Though there are in-silico and in-vitro studies that associated the ITGB2 as one of the important genes in cancer, the exact mechanisms of how this gene promotes GBM remains elusive and worth to be investigated in the future [11, 51, 52]. Human leukocyte antigen (HLA)-DRA is a classical major histocompatibility complex (MHC) class II molecule that plays important role in immune responses modulation. High expression of the HLA-DR gene family has been associated with more aggressive tumour grade in gliomas and poor prognosis [53, 54]. Nonetheless, the functions of HLA-DRA in driving GBM growth has not been fully elucidated. PTPRJ gene is a member of the protein tyrosine phosphatase (PTP) family whose substrates include the RTKs such VEGFR, PDGFR and EGFR [55]. Since the RTKs pro-oncogenic properties are well-established in which their activation largely depends on phosphorylation, PTPRJ is thus deemed to function as tumour suppressor proteins due to its function as a phosphatase that can negatively regulate signalling pathway. This was also evidenced by the ectopic expression of PTPRJ in in-vitro models that resulted in cell growth inhibition [56, 57]. In contrast to these previous reports, we found that PTPRJ expression was upregulated in GBM and led us to suggest that PTPRJ might have a pro-oncogenic role in GBM pathogenesis. To our knowledge, there have been no previous reports linking PTPRJ expression and function with GBM pathogenesis. This notion of PTPRJ potential ‘double-edged sword’ and GBM-specific pro-oncogenic function needs to be investigated further. SLC1A5, another hit target from our analysis, is a neutral amino acid transporter in which its high expression has been implicated in many cancer types including GBM [58]. In GBM, SLC1A5 expression is under the control of pro-oncogenic c-Myc protein but how this transporter supports the tumour cells proliferation and growth remain poorly understood [59]. As highlighted above, the identification of CD44 and EGFR in this present study is expected because they have been previously described as one of the key targets for GBM [29, 60]. This validates the robustness of our approach in the sense that not only our analysis identified several novel genes, but also the findings overlap with previous studies. Since EGFR pro-oncogenic roles have been widely implicated in many cancer types and several drugs have been developed and clinically approved to target EGFR [34, 61, 62], we focused our analysis on CD44. The CD44 encodes for transmembrane glycoprotein that serves as the receptor for hyaluronic acid, a component of the extracellular matrix, and several other ligands including osteopontin, fibronectin and collagen [29]. The CD44 antigen has been implicated in modulating tumorigenesis in many cancer types in which high expression of this CD44 increases cancer cells proliferation, motility and survival as well as promoting cancer metastasis [63]. In GBM, high expression of CD44 was identified in the proteogenomic profiling of GBM tissues [23] and further classified as a GBM cell surface antigen in a systematic analysis [28]. Interestingly, this transmembrane glycoprotein can be cleaved and secreted into the vasculatures, suggesting its potential to be developed as a diagnostic marker [64]. It has been reported that the activation of CD44 by its ligand promotes cancer stem cell-like phenotypes in GBM and increased therapeutic resistance [65]. Consistent with this, drugs targeting CD44 are currently in clinical trials, and so far the results are promising in that CD44 inhibition impede GBM cells growth [66]. Our co-expression network analysis using a graph-based analytics [15] demonstrated that genes connected to CD44 were also highly co-expressed in GBM compared to normal brain tissues, suggesting that CD44 signalling axis is important in GBM tumorigenesis. The current approved therapies to treat GBM are far from satisfactory and have remained unchanged for more than a decade [67]. This includes the alkylating agent temozolomide, which is the first line of drug used in treating GBM. Therefore, there is a need for novel or alternative treatment strategies for GBM. Due to the upregulated expression of CD44 in GBM, drugs targeting CD44 are currently undergoing clinical trials and the results are thus far promising in that CD44 inhibition impedes GBM cells growth [66]. In addition to this, our drug mapping analysis revealed hyaluronic acid as an actionable CD44 binding molecule. It is therefore appealing to investigate the activity and potential use of this existing drug to treat GBM in the future, which has yet to be comprehensively studied. Within the CD44 co-expressed interactome, three additional targets already have drugs that can modulate them namely the C1R, CALR and TNFSR1A (Table 2). Based on our knowledge, the activity and efficacy of these drugs have not been tested in any in-vitro or in-vivo GBM models yet. Also, studying a combination of these available drugs targeting our GBM signature or the CD44 co-expression network could disrupt the aberrant hub gene interactome and potentially enhance GBM treatment efficacy. Conclusion In summary, we identified GBM surfaceome by combining RNA-seq data. Through an integrative multi-OMICS strategy, we highlighted 6 GBM surface-enriched genes that could be important in driving GBM development. Some of these genes can be targeted by clinically approved drugs for other diseases suggesting potential drug repurposing. Additionally, further studies of these genes could lead to potential GBM diagnostic/prognostic markers or a therapeutic regimen to treat GBM. Abbreviations GBM: Glioblastoma multiforme; TCGA: The cancer genome atlas; GTEx: Genotype-tissue expression; DEG: Differentially expressed pene; PPI: Protein-protein interaction; t-SNE: t-distributed stochastic neighbour embedding; GO: Gene ontology; KEGG: and Kyoto encyclopedia of genes and genomes; GEPIA: Gene expression profiling interactive analysis; MHC: Major histocompatibility complex; RTK: Receptor tyrosine kinase; GPCR: G protein-coupled receptor; PTP: Protein tyrosine kinase. Declarations Ethics approval and consent to participate Because the present study did not use any patient samples, this is not applicable. Consent for publication Not applicable. Availability of data and materials The data are included within the manuscript and in the supplementary files. The TCGA GBM data can be obtained from the Genomics Data Commons Data Portal (https://portal.gdc.cancer.gov). The normal brain tissues RNA-seq data were obtained from the GTEx Portal ( https://gtexportal.org/home/datasets ). Other data are available from the corresponding author upon reasonable request. Competing interests The authors declare that they have no competing interests. Funding This study is supported by the Fundamental Research Grant Scheme by the Ministry of Education, Malaysia (FRGS/1/2018/STG04/UKM/03/1) and Collaborative Research Programme - International Centre for Genetic Engineering and Biotechnology Grant (CRP/MYS19-04_EC). The funders had no role in this study. Authors’ contribution Conceptualization, MAM and SES; methodology, WFWMN; software, WFWN; formal analysis, WFWMN, MAM and SES; investigation, WFWMN, MAM and SES; resources, WFWMN and NAM; data curation, NAM and SBH.; writing—original draft preparation, SES, MAM and WFWMN; supervision, MAM; writing—review & editing, MAM and SES; funding acquisition, MAM. All authors have read and approved the manuscript. Acknowledgements The authors thank David Shorthouse (MRC Cancer Unit, University of Cambridge) and Low Teck Yew (UKM Medical Molecular Biology Institute) for discussion, critical insight and proofreading the manuscript. Author details 1 UKM Medical Molecular Biology Institute, UKM Medical Centre, Universiti Kebangsaan Malaysia, Bandar Tun Razak, 56000 Cheras, Kuala Lumpur, Malaysia. 2 Faculty of Science and Natural Resources, Universiti Malaysia Sabah, 88400 Kota Kinabalu, Sabah, Malaysia. 3 Neurosurgery Division, Department of Surgery, Faculty of Medicine, Universiti Kebangsaan Malaysia, Bandar Tun Razak, 56000 Cheras, Kuala Lumpur, Malaysia. References Siegel RL, Miller KD, Jemal A. Cancer statistics, 2016. CA Cancer J Clin. 2016;66:7–30. Kamiya-Matsuoka C, Gilbert MR. Treating recurrent glioblastoma: an update. CNS Oncol. 2015;4:91–104. Ohgaki H. Epidemiology of brain tumors. Methods Mol Biol Clifton NJ. 2009;472:323–42. Qazi MA, Vora P, Venugopal C, Sidhu SS, Moffat J, Swanton C, et al. Intratumoral heterogeneity: pathways to treatment resistance and relapse in human glioblastoma. Ann Oncol. 2017;28:1448–56. Shergalis A, Bankhead A, Luesakul U, Muangsin N, Neamati N. Current Challenges and Opportunities in Treating Glioblastoma. Pharmacol Rev. 2018;70:412–45. Stupp R, Mason WP, van den Bent MJ, Weller M, Fisher B, Taphoorn MJB, et al. Radiotherapy plus concomitant and adjuvant temozolomide for glioblastoma. N Engl J Med. 2005;352:987–96. Ito H, Nakashima H, Chiocca EA. Molecular responses to immune checkpoint blockade in glioblastoma. Nat Med. 2019;25:359. Nam JY, de Groot JF. Treatment of Glioblastoma. J Oncol Pract. 2017;13:629–38. Cancer Genome Atlas Research Network. Comprehensive genomic characterization defines human glioblastoma genes and core pathways. Nature. 2008;455:1061–8. Pearson JRD, Regad T. Targeting cellular pathways in glioblastoma multiforme. Signal Transduct Target Ther. 2017;2:17040. Polisetty RV, Gautam P, Sharma R, Harsha HC, Nair SC, Gupta MK, et al. LC-MS/MS Analysis of Differentially Expressed Glioblastoma Membrane Proteome Reveals Altered Calcium Signaling and Other Protein Groups of Regulatory Functions. Mol Cell Proteomics. 2012;11:M111.013565. Bausch-Fluck D, Hofmann A, Bock T, Frei AP, Cerciello F, Jacobs A, et al. A Mass Spectrometric-Derived Cell Surface Protein Atlas. PLOS ONE. 2015;10:e0121314. Bausch-Fluck D, Goldmann U, Müller S, Oostrum M van, Müller M, Schubert OT, et al. The in silico human surfaceome. Proc Natl Acad Sci. 2018;115:E10988–97. Xia J, Gill EE, Hancock REW. NetworkAnalyst for statistical, visual and network-based meta-analysis of gene expression data. Nat Protoc. 2015;10:823–44. Theocharidis A, van Dongen S, Enright AJ, Freeman TC. Network visualization and analysis of gene expression data using BioLayout Express(3D). Nat Protoc. 2009;4:1535–50. Nakada M, Kita D, Watanabe T, Hayashi Y, Teng L, Pyko IV, et al. Aberrant Signaling Pathways in Glioma. Cancers. 2011;3:3242–78. Cunha JPC da, Galante P a. F, Souza JE de, Souza RF de, Carvalho PM, Ohara DT, et al. Bioinformatics construction of the human cell surfaceome. Proc Natl Acad Sci. 2009;106:16752–7. Lee JK, Bangayan NJ, Chai T, Smith BA, Pariva TE, Yun S, et al. Systemic surfaceome profiling identifies target antigens for immune-based therapy in subtypes of advanced prostate cancer. Proc Natl Acad Sci. 2018;115:E4473–82. Almén MS, Nordström KJV, Fredriksson R, Schiöth HB. Mapping the human membrane proteome: a majority of the human membrane proteins can be classified according to function and evolutionary origin. BMC Biol. 2009;7:50. Banerjee HN, Mahaffey K, Riddick E, Banerjee A, Bhowmik N, Patra M. Search for a diagnostic/prognostic biomarker for the brain cancer glioblastoma multiforme by 2D-DIGE-MS technique. Mol Cell Biochem. 2012;367:59–63. Collet B, Guitton N, Saïkali S, Avril T, Pineau C, Hamlat A, et al. Differential analysis of glioblastoma multiforme proteome by a 2D-DIGE approach. Proteome Sci. 2011;9:16. Heroux MS, Chesnik MA, Halligan BD, Al-Gizawiy M, Connelly JM, Mueller WM, et al. Comprehensive characterization of glioblastoma tumor tissues for biomarker identification using mass spectrometry-based label-free quantitative proteomics. Physiol Genomics. 2014;46:467–81. Song Y-C, Lu G-X, Zhang H-W, Zhong X-M, Cong X-L, Xue S-B, et al. Proteogenomic characterization and integrative analysis of glioblastoma multiforme. Oncotarget. 2017;8:97304–12. Low TY, Mohtar MA, Ang MY, Jamal R. Connecting Proteomics to Next-Generation Sequencing: Proteogenomics and Its Current Applications in Biology. Proteomics. 2019;19:e1800235. Ang MY, Low TY, Lee PY, Wan Mohamad Nazarie WF, Guryev V, Jamal R. Proteogenomics: From next-generation sequencing (NGS) and mass spectrometry-based proteomics to precision medicine. Clin Chim Acta. 2019;498:38–46. Tang Z, Li C, Kang B, Gao G, Li C, Zhang Z. GEPIA: a web server for cancer and normal gene expression profiling and interactive analyses. Nucleic Acids Res. 2017;45:W98–102. Sargent DJ, Wieand HS, Haller DG, Gray R, Benedetti JK, Buyse M, et al. Disease-Free Survival Versus Overall Survival As a Primary End Point for Adjuvant Colon Cancer Studies: Individual Patient Data From 20,898 Patients on 18 Randomized Trials. J Clin Oncol. 2005;23:8664–70. Ghosh D, Funk CC, Caballero J, Shah N, Rouleau K, Earls JC, et al. A Cell-Surface Membrane Protein Signature for Glioblastoma. Cell Syst. 2017;4:516-529.e7. Chen C, Zhao S, Karnad A, Freeman JW. The biology and role of CD44 in cancer progression: therapeutic implications. J Hematol OncolJ Hematol Oncol. 2018;11:64. Bowman S, Awad ME, Hamrick MW, Hunter M, Fulzele S. Recent advances in hyaluronic acid based therapy for osteoarthritis. Clin Transl Med. 2018;7. doi:10.1186/s40169-017-0180-3. Misra S, Hascall VC, Markwald RR, Ghatak S. Interactions between Hyaluronan and Its Receptors (CD44, RHAMM) Regulate the Activities of Inflammation and Cancer. Front Immunol. 2015;6. doi:10.3389/fimmu.2015.00201. Kim JH, Moon MJ, Kim DY, Heo SH, Jeong YY. Hyaluronic Acid-Based Nanomaterials for Cancer Therapy. Polymers. 2018;10. doi:10.3390/polym10101133. Kim K, Choi H, Choi ES, Park M-H, Ryu J-H. Hyaluronic Acid-Coated Nanomedicine for Targeted Cancer Therapy. Pharmaceutics. 2019;11. Sigismund S, Avanzato D, Lanzetti L. Emerging functions of the EGFR in cancer. Mol Oncol. 2018;12:3–20. Bausch-Fluck D, Milani ES, Wollscheid B. Surfaceome nanoscale organization and extracellular interaction networks. Curr Opin Chem Biol. 2019;48:26–33. Teh JLF, Chen S. Glutamatergic signaling in cellular transformation. Pigment Cell Melanoma Res. 2012;25:331–42. Mirkowska P, Hofmann A, Sedek L, Slamova L, Mejstrikova E, Szczepanski T, et al. Leukemia surfaceome analysis reveals new disease-associated features. Blood. 2013;121:e149–59. Fenner A. Surfaceome profiling for NEPC target antigens. Nat Rev Urol. 2018;15:396–7. Ziegler A, Cerciello F, Bigosch C, Bausch-Fluck D, Felley-Bosco E, Ossola R, et al. Proteomic surfaceome analysis of mesothelioma. Lung Cancer. 2012;75:189–96. Pais H, Ruggero K, Zhang J, Al-Assar O, Bery N, Bhuller R, et al. Surfaceome interrogation using an RNA-seq approach highlights leukemia initiating cell biomarkers in an LMO2 T cell transgenic model. Sci Rep. 2019;9:1–16. Hanahan D, Weinberg RA. Hallmarks of Cancer: The Next Generation. Cell. 2011;144:646–74. Leth-Larsen R, Lund RR, Ditzel HJ. Plasma membrane proteomics and its application in clinical cancer biomarker discovery. Mol Cell Proteomics MCP. 2010;9:1369–82. Nicolasjilwan M, Hu Y, Yan C, Meerzaman D, Holder CA, Gutman D, et al. Addition of MR imaging features and genetic biomarkers strengthens glioblastoma survival prediction in TCGA patients. J Neuroradiol J Neuroradiol. 2015;42:212–21. Han J, Puri RK. Analysis of the cancer genome atlas (TCGA) database identifies an inverse relationship between interleukin-13 receptor α1 and α2 gene expression and poor prognosis and drug resistance in subjects with glioblastoma multiforme. J Neurooncol. 2018;136:463–74. Jia D, Li S, Li D, Xue H, Yang D, Liu Y. Mining TCGA database for genes of prognostic value in glioblastoma microenvironment. Aging. 2018;10:592–605. Sanchez-Vega F, Mina M, Armenia J, Chatila WK, Luna A, La KC, et al. Oncogenic Signaling Pathways in The Cancer Genome Atlas. Cell. 2018;173:321-337.e10. Regad T. Targeting RTK Signaling Pathways in Cancer. Cancers. 2015;7:1758–84. Lundstrom K. An Overview on GPCRs and Drug Discovery: Structure-Based Drug Design and Structural Biology on GPCRs. In: Leifert WR, editor. G Protein-Coupled Receptors in Drug Discovery. Totowa, NJ: Humana Press; 2009. p. 51–66. doi:10.1007/978-1-60327-317-6_4. Vogel C, Marcotte EM. Insights into the regulation of protein abundance from proteomic and transcriptomic analyses. Nat Rev Genet. 2012;13:227–32. Camponeschi A, Gerasimcik N, Wang Y, Fredriksson T, Chen D, Farroni C, et al. Dissecting Integrin Expression and Function on Memory B Cells in Mice and Humans in Autoimmunity. Front Immunol. 2019;10:534. Wang A, Chen M, Wang H, Huang J, Bao Y, Gan X, et al. Cell Adhesion-Related Molecules Play a Key Role in Renal Cancer Progression by Multinetwork Analysis. BioMed Res Int. 2019;2019:2325765. Dunwoodie LJ, Poehlman WL, Ficklin SP, Feltus FA. Discovery and validation of a glioblastoma co-expressed gene module. Oncotarget. 2018;9:10995–1008. Fan X, Liang J, Wu Z, Shan X, Qiao H, Jiang T. Expression of HLA-DR genes in gliomas: correlation with clinicopathological features and prognosis. Chin Neurosurg J. 2017;3:27. Diao J, Xia T, Zhao H, Liu J, Li B, Zhang Z. Overexpression of HLA-DR is associated with prognosis of glioma patients. Int J Clin Exp Pathol. 2015;8:5485–90. Godfrey R, Arora D, Bauer R, Stopp S, Müller JP, Heinrich T, et al. Cell transformation by FLT3 ITD in acute myeloid leukemia involves oxidative inactivation of the tumor suppressor protein-tyrosine phosphatase DEP-1/ PTPRJ. Blood. 2012;119:4499–511. Iuliano R, Trapasso F, Le Pera I, Schepis F, Samà I, Clodomiro A, et al. An adenovirus carrying the rat protein tyrosine phosphatase eta suppresses the growth of human thyroid carcinoma cell lines in vitro and in vivo. Cancer Res. 2003;63:882–6. Massa A, Barbieri F, Aiello C, Arena S, Pattarozzi A, Pirani P, et al. The expression of the phosphotyrosine phosphatase DEP-1/PTPeta dictates the responsivity of glioma cells to somatostatin inhibition of cell proliferation. J Biol Chem. 2004;279:29004–12. Bhutia YD, Ganapathy V. Glutamine transporters in mammalian cells and their functions in physiology and cancer. Biochim Biophys Acta BBA - Mol Cell Res. 2016;1863:2531–9. Wise DR, DeBerardinis RJ, Mancuso A, Sayed N, Zhang X-Y, Pfeiffer HK, et al. Myc regulates a transcriptional program that stimulates mitochondrial glutaminolysis and leads to glutamine addiction. Proc Natl Acad Sci. 2008;105:18782–7. Westphal M, Maire CL, Lamszus K. EGFR as a Target for Glioblastoma Treatment: An Unfulfilled Promise. CNS Drugs. 2017;31:723–35. Singh D, Attri BK, Gill RK, Bariwal J. Review on EGFR Inhibitors: Critical Updates. Mini Rev Med Chem. 2016;16:1134–66. Hynes NE, Lane HA. ERBB receptors and cancer: the complexity of targeted inhibitors. Nat Rev Cancer. 2005;5:341–54. Senbanjo LT, Chellaiah MA. CD44: A Multifunctional Cell Surface Adhesion Receptor Is a Regulator of Progression and Metastasis of Cancer Cells. Front Cell Dev Biol. 2017;5. doi:10.3389/fcell.2017.00018. Lim S, Kim D, Ju S, Shin S, Cho I, Park S-H, et al. Glioblastoma-secreted soluble CD44 activates tau pathology in the brain. Exp Mol Med. 2018;50:1–11. Pietras A, Katz AM, Ekström EJ, Wee B, Halliday JJ, Pitter KL, et al. Osteopontin-CD44 signaling in the glioma perivascular niche enhances cancer stem cell phenotypes and promotes aggressive tumor growth. Cell Stem Cell. 2014;14:357–69. Mooney KL, Choy W, Sidhu S, Pelargos P, Bui TT, Voth B, et al. The role of CD44 in glioblastoma multiforme. J Clin Neurosci Off J Neurosurg Soc Australas. 2016;34:1–5. Kazda T, Dziacky A, Burkon P, Pospisil P, Slavik M, Rehak Z, et al. Radiotherapy of Glioblastoma 15 Years after the Landmark Stupp’s Trial: More Controversies than Standards? Radiol Oncol. 2018;52:121–8. Tables Due to technical limitations, the tables are provided in the Supplementary Files section. Supplementary Materials Supplementary Tables Supplementary Table S1. Overall differentially expressed genes in TCGA GBM tissues vs. GTEx normal brain tissues. Supplementary Table S2. Significantly dysregulated cell surface genes in TCGA GBM tissues vs. GTEx normal brain tissues. Supplementary Table S3. GBM cell lines proteomics data from Bausch-Fluck et al. 2015. Supplementary Table S4. GBM tissue samples proteomics data from Polisetty et. al 2012. Supplementary Table S5. Protein-protein interaction network analysis of surfaceome. Supplementary Figures Figure S1. Gene ontology and deregulated pathways in GBM. (A-B) Gene ontology cellular component of the significantly (A) upregulated and (B) downregulated genes in GBM. (C-D) KEGG pathway analysis of the (C) upregulated and (D) downregulated genes in GBM. Figure S2. Significant differentially expressed cell surface genes in GBM. (A) GBM surfaceome classification using previously annotated cell surface genes dataset identifies 395 DEGs that belongs to surfaceome. (B) Cell surface genes stratification from (A) based on its subclass. Figure S3. KEGG pathway analysis of differentially expressed surfaceome in GBM. (A) Upregulated surfaceome and (B) Downregulated surfaceome. Figure S4. Significant upregulation of the prioritized GBM surfaceome signature in GBM patients. (A-F) Boxplot showing the RNA-Seq data (transcript per million) of (A) CD44 (B) PTPRJ (C) SLC1A5 (D) EGFR (E) HLA-DRA and (F) ITGB2 in GBM and GTEx normal brain tissue samples.. Figure S5. Overall survival analysis of the prioritized GBM surfaceome signature as potential GBM prognostic biomarker. (A-F) Overall survival analysis of GBM patients having high and low expression of (A) CD44 (B) PTPRJ (C) SLC1A5 (D) EGFR (E) HLA-DRA and (F) ITGB2. Figure S6. Disease-free survival analysis of the prioritized GBM surfaceome signature as potential GBM prognostic biomarker. (A-F) Disease-free survival analysis of GBM patients having high and low expression of (A) CD44 (B) PTPRJ (C) SLC1A5 (D) EGFR (E) HLA-DRA and (F) ITGB2. Figure S7. Survival analysis of the 6 GBM signature genes. (A) Overall survival and (B) disease-free survival analysis of GBM patients having high and low expression of all 6 genes; CD44, PTPRJ, SLC1A5, EGFR, HLA-DRA and ITGB2 Figure S8. Survival analysis of the 3 GBM signature genes. (A) Overall survival and (B) disease-free survival analysis of GBM patients having high and low expression of CD44, PTPRJ and HLA-DRA. Supplementary Files SyafruddinetalSuppTable.xlsx Tablerev.pdf Supplementaryfiguresrev.pdf Cite Share Download PDF Status: Published Journal Publication published 23 Jul, 2021 Read the published version in BMC Cancer → Version 2 posted Reviewer # 2 agreed at journal 16 Mar, 2021 Review # 1 received at journal 09 Nov, 2020 Reviewer # 1 agreed at journal 19 Oct, 2020 Editor assigned by journal 16 Oct, 2020 Reviewers invited by journal 16 Oct, 2020 Submission checks completed at journal 15 Oct, 2020 Editor invited by journal 15 Oct, 2020 You are reading this latest preprint version Show more versions Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-46071","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research article","associatedPublications":[],"authors":[{"id":3609856,"identity":"63e31d5d-88e4-4dc5-8257-b8e48bf4efb5","order_by":0,"name":"Saiful Effendi Syafruddin","email":"","orcid":"","institution":"Universiti Kebangsaan Malaysia","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Saiful","middleName":"Effendi","lastName":"Syafruddin","suffix":""},{"id":3609857,"identity":"33d6b3f4-d0b9-4ccd-903a-622c450e2e12","order_by":1,"name":"Wan Fahmi Wan Mohamad Nazarie","email":"","orcid":"","institution":"Universiti Kebangsaan Malaysia","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wan","middleName":"Fahmi Wan Mohamad","lastName":"Nazarie","suffix":""},{"id":3609858,"identity":"bc5ddcd5-94cc-46fc-899b-5ecd99dafc79","order_by":2,"name":"Nurshahirah Ashikin Moidu","email":"","orcid":"","institution":"Universiti Kebangsaan Malaysia","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Nurshahirah","middleName":"Ashikin","lastName":"Moidu","suffix":""},{"id":3609859,"identity":"4643a25a-27ef-407d-9733-00f911a05a99","order_by":3,"name":"Bee Hong Soon","email":"","orcid":"","institution":"Universiti Kebangsaan Malaysia","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Bee","middleName":"Hong","lastName":"Soon","suffix":""},{"id":3609860,"identity":"a290f6cb-b270-4d3e-8283-744972ccc2db","order_by":4,"name":"M. Aiman Mohtar","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABIklEQVRIiWNgGAWjYJCCA2BSAkRUADEPXCyBkBZmIHGGSC0McC2MbURo0W1gfnjoRs2daH7p/oOfK+fdyTM4czqBmefMNgZ+9hwDppttGFrMDrAZHM459ix35pzDzJJntz0rNjjbu4GZ58ZtBsmeNwbMudi0MAC1sB3O3XAjmUGycdvhxA3neYFaPtxmMLiRg0ML+4fDOf/AWph/Ns5B0mKPUwuPweHcNrAWNsnGBqAWmMMMJHBoOcxTcDi373DuzBnJZpYNx54VS545u+HgnDO3eSTOPCs4nHMOU8vx9s2fc74dzu2XSHx8s6HmTh7fmdyND94cuy3H35688XFOGWYoM6NyDySASQZY9DCyYWpBAxAtSOAPQS2jYBSMglEw7AEAvDWATByEjxYAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-9015-9802","institution":"Universiti Kebangsaan Malaysia","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"M.","middleName":"Aiman","lastName":"Mohtar","suffix":""}],"badges":[],"createdAt":"2020-07-20 11:27:15","currentVersionCode":2,"declarations":"","doi":"10.21203/rs.3.rs-46071/v2","doiUrl":"https://doi.org/10.21203/rs.3.rs-46071/v2","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12885-021-08591-0","type":"published","date":"2021-07-23T15:00:38+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":3111578,"identity":"f15a23a9-ec6a-4679-b039-eec9dfc5c10f","added_by":"auto","created_at":"2020-10-21 15:25:01","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":170711,"visible":true,"origin":"","legend":"Analysis pipeline to obtain the GBM predictive surfaceome markers applied from the initial TCGA GBM and GTEx data integration.","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-46071/v2/539913ecec5f9cb469aa6549.png"},{"id":3111580,"identity":"f9fa28f9-ea61-4d73-8dac-1d9c8a203609","added_by":"auto","created_at":"2020-10-21 15:25:02","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":581855,"visible":true,"origin":"","legend":"Identification of global differentially expressed genes in GBM. (A) t-SNE plots showing the GBM and GTEX data cluster. (B) Volcano plot of the differentially expressed genes in GBM versus normal brain tissues. Genes that are significantly dysregulated in GBM versus GTEx (log2 fold change |2|) were highlighted in red (downregulated) and green (upregulated).","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-46071/v2/b36a1713ca62a174337ed895.png"},{"id":3111582,"identity":"b506a999-0194-42c3-90fb-be4d6d4e0416","added_by":"auto","created_at":"2020-10-21 15:25:03","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":110816,"visible":true,"origin":"","legend":"Integration of TCGA GBM transcriptomics, GBM tissues proteomics and cell lines proteomics data.","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-46071/v2/2815009258459f717c9da5d7.png"},{"id":3111584,"identity":"a9055b2c-1c28-4552-97da-757bb5efb05b","added_by":"auto","created_at":"2020-10-21 15:25:03","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1093529,"visible":true,"origin":"","legend":"Prioritization of six high-confidence GBM surface marker genes. (A) Protein-protein interaction network analysis of the significantly upregulated GBM surfaceome genes. (B) Venn diagram showing the genes that are overlapped between the PPI network and transcriptomics-proteomics data integration analysis. ","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-46071/v2/e2dd75e95b000b08ac4bc93e.png"},{"id":3111585,"identity":"1b0ce1dd-4b7e-4c67-a20b-f34621be9d3d","added_by":"auto","created_at":"2020-10-21 15:25:04","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":576069,"visible":true,"origin":"","legend":"CD44 gene co-expressed network analysis. (A) CD44 gene co-expressed network with Pearson correlation value, r \u003e0.75. Nodes represent genes and edges are colored on a sliding scale according to the strength of the correlation (red, r = 1.0 and blue, r = 0.75). (B) Histograms of CD44 co-expression cluster from (A) showing the average expression of genes on GBM tumor (red bar) and normal (yellow bar). ","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-46071/v2/c26c74501150fb69f7fdeda5.png"},{"id":13604654,"identity":"0d95b403-e5dd-4170-bb89-5756b0d42e9e","added_by":"auto","created_at":"2021-09-17 06:00:52","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1860358,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-46071/v2/bc10ba12-8863-4726-a3b1-0c4fb99e5a57.pdf"},{"id":3111579,"identity":"1b1e6d4e-243b-490a-b840-f8fc9b7ed86f","added_by":"auto","created_at":"2020-10-21 15:25:01","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":948944,"visible":true,"origin":"","legend":"","description":"","filename":"SyafruddinetalSuppTable.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-46071/v2/44c15f2a861e842e36d8f3a8.xlsx"},{"id":3111581,"identity":"0930dbf9-8d1e-40aa-be08-1c402b6f929f","added_by":"auto","created_at":"2020-10-21 15:25:03","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":22676,"visible":true,"origin":"","legend":"","description":"","filename":"Tablerev.pdf","url":"https://assets-eu.researchsquare.com/files/rs-46071/v2/9544e1fa534f9098b9d11cbc.pdf"},{"id":3111583,"identity":"a108732d-b7c0-46e5-982d-bbc02e13c4fc","added_by":"auto","created_at":"2020-10-21 15:25:03","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":1364073,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementaryfiguresrev.pdf","url":"https://assets-eu.researchsquare.com/files/rs-46071/v2/9f94a38090116bdba796328f.pdf"}],"financialInterests":"","formattedTitle":"Integration of RNA-Seq and proteomics data identifies glioblastoma multiforme surfaceome signature","fulltext":[{"header":"Background","content":"\u003cp\u003eGlioblastoma multiforme (GBM) is the most common and lethal tumour of the central nervous system in adults [1]. Despite decades of efforts to tackle this disease, the median survival rate of GBM patients is still not improving [2]. GBM patients have an average life expectancy of 15 months post-diagnosis and the 5-years survival rate is less than 3% [3]. The standard-of-care GBM treatment generally consists of maximal safe surgical resection followed by radiotherapy and concomitant chemotherapy. However rapid post-treatment relapse and high intra-tumoral heterogeneity that could either arise naturally during disease progression or treatments-induced have made this disease intractable and more challenging to treat [4, 5]. Therefore, there is a pressing need for better and efficient diagnostic and therapeutic strategies for this disease.\u003c/p\u003e\n\u003cp\u003eTemozolomide, an orally-administered DNA-alkylating drug, is the current and commonly used chemotherapy agent to treat GBM in the clinic [6]. This combination treatment of temozolomide and radiotherapy is referred to as the Stupp regimen and it is widely used as the standard-of-care for the treatment of GBM. The landmark study showed that the combination of radiotherapy and concomitant chemotherapy with\u0026nbsp;\u003ca href=\"https://radiopaedia.org/articles/missing?article%5Btitle%5D=temozolomide\u0026amp;lang=gb\"\u003etemozolomide\u003c/a\u003e improve the patient's prognosis compared to radiotherapy alone (median survival of 14.6 months vs 12.1 months, respectively) [6].\u0026nbsp; Alternative GBM treatment options such as the VEGF-targeting monoclonal antibody Bevacizumab, other DNA alkylating agents such as lomustine and carmustine implants, alternating electric field therapy and the checkpoint blockade inhibitor have thus far yielded low efficacy in treating GBM [2, 7, 8]. The Cancer Genome Atlas (TCGA) comprehensive GBM molecular characterizations have identified significant genetic alterations in several important oncogenic signalling pathways such as the RTK/Ras/PI3K (88%), p53 (87%) and pRB signalling pathways (78%) in GBM patients [9]. Several clinical trials are currently ongoing that aim to target these altered GBM oncogenic signalling pathways components using small molecule inhibitors and/or monoclonal antibodies. However, the results thus far were far from satisfactory [10]. This seems to suggest that instead of using a single agent targeting a specific component or pathway, novel treatments should consider the administration of several inhibitors targeting multiple different pathways.\u003c/p\u003e\n\u003cp\u003eThe cell surface proteins or surfaceome serve as an information gateway that integrates and transduces extracellular cues into intracellular signalling cascades or \u003cem\u003evice versa\u003c/em\u003e. Surfaceome also play important role in cell adhesion and migration which are among the critical processes during tumorigenesis. Indeed, aberrant surfaceome expression and activity are frequently observed in many cancer types and therefore are good candidates for cancer diagnostic or biomarkers as well as therapeutic targets. Recent evidence has demonstrated that 56% of cell surface proteins are differentially expressed in GBM which are also present in cerebrospinal fluid or plasma, suggesting their potential use as biomarkers [11]. Of note, surfaceome expression is more dynamic than intracellular proteins and they could be sometimes cell type-specific [12, 13]. Mass spectrometry analysis showed that the average surfaceome size in brain cancer cell lines is higher than other cancer types [12]. Thus, surfaceome genes in GBM may hold the key to understand GBM pathogenesis and drug responsiveness, in which targeting these genes may unravel potential \u0026lsquo;\u003cem\u003edruggable\u003c/em\u003e\u0026rsquo; stage in GBM pathways.\u003c/p\u003e\n\u003cp\u003eA comprehensive overview of GBM surfaceome landscape has not been fully defined. Therefore, this study aimed to characterize the GBM surfaceome genes expression profile by unifying the two large RNA-Seq datasets from the TCGA (GBM) and GTEx (normal brain). We integrated and performed differential gene expression analysis on these two datasets because of the low number of normal brain tissue samples in the TCGA database. Previously annotated surfaceome gene set was employed to filter and identify the significant differentially expressed surfaceome genes in GBM. To further prioritize the high-confidence GBM cell surface signature, we integrated our transcriptomics analysis with GBM tissues and cell surface proteomics, and PPI hub gene analysis. Collectively, we identified a list of upregulated surfaceome genes in GBM that include \u003cem\u003eCD44, PTPRJ\u003c/em\u003e and \u003cem\u003eHLA-DRA\u003c/em\u003e in which their biological relevance in supporting GBM pathogenesis could be comprehensively investigated in the future studies for the development of novel GBM diagnostic/prognostic or therapeutic strategies.\u0026nbsp;\u0026nbsp;\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e\u003cstrong\u003e\u003cem\u003eTCGA and GTEx data acquisition, normalization and quality control:\u003c/em\u003e\u003c/strong\u003e The analysis combined of the TCGA-GBM and GTEx normal brain RNA-Seq read count data. The GBM RNA-Seq gene raw read counts from TCGA were downloaded from Genomics Data Commons Data Portal (\u003ca href=\"https://portal.gdc.cancer.gov/\"\u003ehttps://portal.gdc.cancer.gov\u003c/a\u003e). GTEx data were used for the normal brain tissues. The GTEx data used for the analyses described in this manuscript were obtained from the GTEx Portal on 29/03/19. We downloaded RNA-Seq gene raw read counts (from the cortex, frontal cortex, anterior cingulate cortex) from GTEx portal (\u003ca href=\"https://gtexportal.org/home/datasets\"\u003ehttps://gtexportal.org/home/datasets\u003c/a\u003e). This allows us to perform the analysis of the differentially expressed gene on the 166 samples of GBM tumour from TCGA and 408 samples of normal brain tissues data from GTEx. The RNA-Seq raw read counts pre-processing steps involve are data filtering and data normalization. The normalization process of both data set was then performed by using mean as gene-level normalization using log\u003csub\u003e2\u003c/sub\u003e-counts per million where raw data are adjusted to account for factors that will prevent a direct comparison of expression measures and to safeguard the expression distributions are similar for each sample across the whole experiment. Data that unlikely to be informative or simply erroneous data will be removed by using variance filter (less than 15) and low abundance (less than 4).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eCell surface gene set classification and analysis:\u003c/em\u003e\u003c/strong\u003e The identified differentially expressed genes (DEGs) of glioblastoma were classified into cell-surface genes set as discussed in the main text (See Results 2.4). The classification of the gene sets was performed based on the mapping set of DEGs with this resource. Other genes, which did not map to this resource were removed from the final dataset.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eDifferential gene expression:\u003c/em\u003e\u003c/strong\u003e DEGs analysis was performed using NetworkAnalyst [14], a web-based application tool for visualizing molecular and entity interactions. This platform utilizes the statistical method on data comparison from R package, limma to identify genes whose expression is different. Genes that have adjusted p-value \u0026lt;0.05 and log2 fold change |2| were considered as statistically significant DEGs.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eFunctional annotation and pathway analysis:\u003c/em\u003e\u003c/strong\u003e The enrichment analysis of the identified glioblastoma associated genes was performed using DAVID (\u003ca href=\"https://david.ncifcrf.gov/\"\u003ehttps://david.ncifcrf.gov/\u003c/a\u003e), a web-based tool for analyzing functional gene analysis. The tool comprises databases from various public resources for biological analysis. The enrichment analysis such as GO and KEGG pathways were performed with top results as per gene counts.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eIdentification of hub genes through PPI network analysis:\u003c/em\u003e\u003c/strong\u003e A biological database for known and predicted protein-protein interactions called IMEx interactome database (\u003ca href=\"https://www.imexconsortium.org\"\u003ehttps://www.imexconsortium.org\u003c/a\u003e) was used to construct the protein-protein interaction (PPI) of the DEGs. The network of interacting proteins was extracted and visualized using NetworkAnalyst. The top 87-gene modules of highly interacting gene clusters among the DEG were found with default parameters. For the classified gene sets, the PPI network was constructed and the network topological parameters i.e. degree and betweenness centrality were calculated.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eCo-expression network of CD44:\u003c/em\u003e\u003c/strong\u003e Co-expression analysis was performed using Graphia Professional (\u003ca href=\"https://kajeka.com/graphia-professional/\"\u003ehttps://kajeka.com/graphia-professional/\u003c/a\u003e), previously known as BioLayout \u003cem\u003eExpress\u003c/em\u003e\u003csup\u003e3D\u003c/sup\u003e [15] using raw read counts and then saved as an \u0026lsquo;\u0026lsquo;.expression\u0026rsquo;\u0026rsquo; file. This contains a unique identifier for each row of data. Following import into Graphia Professional, a pairwise Pearson correlation matrix was calculated thereby performing a gene vs. gene comparison of the expression profile of each gene. All Pearson correlations where r\u0026gt;0.7 were saved to a \u0026lsquo;\u0026lsquo;.pearson\u0026rsquo;\u0026rsquo; file. Based on a user-defined threshold of r\u0026gt;0.75, an undirected network graph of the data was generated. In this context, nodes represent individual genes and the edges between them represent Pearson correlation coefficients above the selected threshold (r\u0026gt;0.75). \u003cem\u003eCD44\u003c/em\u003e was selected along with its neighbour in the network, representing \u003cem\u003eCD44\u003c/em\u003e co-expression partners. The class set of \u003cem\u003eCD44\u003c/em\u003e co-expressed genes were visualized to compare the expression values in this class set with genes in normal samples.\u0026nbsp;\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003e\u003cem\u003ePatients\u0026rsquo; characteristics of TCGA and GTEx:\u003c/em\u003e\u003c/strong\u003e We utilized the publicly available TCGA and GTEx RNA-Seq database as our primary sources of GBM tumour and normal brain tissue transcriptomic data, respectively. We downloaded the datasets containing RNA-Seq gene expression profiles and clinical information of 166 patients from TCGA-GBM and 408 normal brain tissues from GTEx database. The combined data were stratified based on gender, age and treatment as shown in Table 1. Out of a total of 166 GBM cases, 104 cases (62.7%) were male and 56 cases (33.7%) were female. GBM is more prevalent in patients aged \u003cem\u003e\u0026ge; \u003c/em\u003e60 years old which accounts for 42.8% of total cases in the TCGA GBM cohort. Fifty-two patients (31.3%) have undergone treatments whereas 62.1% of cases did not have any treatment data. Unfortunately, the clinical data for the GTEx normal brain samples are not publicly available.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eIdentification of differentially expressed genes in glioblastoma:\u003c/em\u003e\u003c/strong\u003e The analysis pipeline employed in this study is depicted in Figure 1. Briefly, the RNA-Seq raw read counts from the two large compendiums, TCGA and GTEx were utilized to identify the differentially expressed genes between GBM and normal brain tissues. Since most GBM cases are generally found in the supratentorial region of the brain such as the cerebral hemisphere [16], we only extracted the RNA-Seq profiles of this region namely the cortex, frontal cortex, anterior cingulate cortex as per GTEx description. We performed t-distributed stochastic neighbour embedding (t-SNE) analysis to reflect the directionality of transcripts expression among GBM tumour and normal brain tissues read count values. The t-SNE plot showed that all RNA-Seq profiles of all GTEx cortex region clustered together while the GBM RNA-Seq profiles form a separate cluster, thus confirming distinct expression patterns between these groups (Figure 2A). In total, RNA expression data from 18,021 genes were obtained from these combined TCGA and GTEx dataset but only 13,548 genes passed the quality control check. By applying the cut-off criteria log\u003csub\u003e2 \u003c/sub\u003efold change |2| and adjusted p-value \u0026lt;0.05, we identified 2381 genes as significantly differentially expressed genes (DEGs) in GBM, of which 648 genes were upregulated and 1733 genes were downregulated (Figure 2B). The detailed information of the differential gene expression analysis is listed in Supplementary Table S1.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eFunctional enrichment analysis and classification of DEGs:\u003c/em\u003e\u003c/strong\u003e The significant DEGs were then subjected to functional enrichment analysis using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) tools to define their properties and putative biological relevance in GBM. Interestingly, the GO cellular component analysis of both upregulated and downregulated DEGs showed enrichment of cell surface and membrane-associated proteins (Supplementary Figure S1A and S1B). The KEGG pathway enrichment indicated that the upregulated DEGs are involved in pathways related to infectious diseases, pathway in cancer and cell adhesion (Supplementary Figure S1C). Downregulated genes mainly involve in neuroactive ligand-receptor interaction and major cellular signalling pathways (Supplementary Figure S1D).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eIdentification of GBM cell-surface antigen candidates:\u003c/em\u003e\u003c/strong\u003e The DEGs were then further filtered and classified into the surfaceome gene set as previously defined by Bausch-Fluck et. al [13], Cunha et. al [17] and Lee et. al [18]. These studies utilized different criteria and stringency in curating the surfaceome gene list. From the overall DEGs in GBM, we identified 395 common cell surface genes within these three surfaceome definitions, including 124 upregulated and 271 downregulated genes (Supplementary Figure S2A and Supplementary Table S2). We further classified the surfaceome according to their main subclasses, which are receptors, transporters, enzymes, miscellaneous and unclassified, as previously reported by Alm\u0026eacute;n et. al [19]. Among the defined surfaceome subclasses, 42.8% of the significant differentially expressed surfaceome in GBM belong to the receptor subclass (Supplementary Figure S2B). KEGG analysis of the GBM-enriched cell surface proteins identified pathways related to immune defence and infectious disease pathways while GBM-deficient cell surface genes are enriched in pathways related to in neuroactive ligand-receptor interaction and major cellular signalling pathways (Supplementary Figure S3A and S3B). These findings are almost similar to the enrichment analysis of overall DEGs in GBM (Supplementary Figure S1C and S1D) suggesting that surfaceome has significant roles in dictating GBM cellular activities.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eIdentification of GBM cell-surface signature by integration of proteomics and transcriptomics data analysis: \u003c/em\u003e\u003c/strong\u003eThus far, we have (i) classified the overall DEGs in GBM using transcriptomics data and (ii) highlighted the differentially expressed cell-surface genes in GBM. Even though this transcriptomics analysis is very informative for biomarker discovery, we aimed to add another layer of analysis to select for a more high-confidence cell surface signature for GBM. To attain this, we integrated our transcriptomics analysis data with the publicly available proteomics data. This integration will validate the cell surface genes prediction and eliminate the possible discrepancy between the expression levels of mRNAs and proteins due to post-transcriptional and post-translational modifications. Thus, we gathered the publicly available quantitative mass spectrometry analysis data for both GBM tissues and cell lines. We postulated that GBM tissues and cell lines might have different cell surface repertoires and therefore it is important to stratify between these two sources. Additionally, GBM cell lines cell surface signature, as identified in this present study, could be validated experimentally in future functional studies.\u003c/p\u003e\n\u003cp\u003eMass spectrometry analysis of five GBM cell lines revealed the upregulation of EGFR, CD44, PTPRJ, SLC1A5, F2R, and TSPAN6 proteins in these samples [12], whereby the expression level of these proteins were in concordance with our transcriptomics data analysis (Figure 3). For tissue proteomics, we found several studies that performed comparative GBM vs. normal brain tissues proteome profiling [11, 20\u0026ndash;23]. However, some of these studies either identified only a limited number of proteins or the data are not downloadable. Only one study by Polisetty et al. that has identified a large number of proteins in their proteome profiling study that included 1834 high-confidence membrane proteins with more than 2-fold change [11]. We, therefore, used this dataset where we performed integrative analysis with our analyzed transcriptomics data and identified 10 overlapped genes, \u003cem\u003eMRC2, FCGR3A, HLA-DRA, CD44, CD74, MSR1, CD163, EGFR, ITGB2, PTPRZ1 \u003c/em\u003e(Figure 3). The mRNA expression levels correlated with the protein expression levels except the \u003cem\u003ePTPRZ1\u003c/em\u003e where the mRNA levels showed upregulation while proteomics data showed downregulation (Supplementary Table S3 and S4). In total, there are 14 genes from the combined tissues and cell lines proteomics that overlapped with our transcriptomics data (Figure 3). It is important to note that proteins identification in mass spectrometry can be limiting due to protein isolation methods, proteins solubility, and other intrinsic variations that affect the proteins abundance as well as the sensitivity and detection capability of the MS instrumentation [24, 25]. Thus, these limitations may underestimate the results between transcriptomics prediction and proteomics discoveries.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eSurfaceome protein-protein interaction network cluster analysis and prioritization of high-confidence GBM cell surface markers:\u003c/em\u003e\u003c/strong\u003e We set out to further analyze the GBM-enriched cell surface markers using protein-protein interaction (PPI) network analysis. This is to better understand the interplay between the cell surface genes within the identified DEGs as well as with other genes. More importantly, this would enable us to further select the genes that are highly interconnected from the integrated proteomics and transcriptomics analysis. Network analysis of the identified differentially expressed cell surface protein genes was performed using NetworkAnalyst [14] to determine the relationship between genes according to the network topological parameters such as degree and betweenness. These parameters reflect the role and property of proteins within the network. The nodes and edges in the PPI network represent the proteins and their interactions, respectively. The GBM-enriched cell surface proteins network contains 1,321 nodes and 1,767 edges interactions based on a number of validated features including functional experiments, co-expression analysis, text mining, neighbourhood, gene fusion and databases (Figure 4A). We identified 87-gene modules of clusters and the top cluster genes with more than 30 interactions include \u003cem\u003eVCAM1, EGFR, TGFBR1, CD44, NGFR, ITGB2, DCC, PTPRJ, ANBCA1, HLA-DRA, CCR5\u003c/em\u003e and \u003cem\u003eCSF1R\u003c/em\u003e (Figure 4A and Supplementary Table S5). Vascular Cell Adhesion Molecule 1 (VCAM1) has the highest interacting cluster as it was found to have 426 degree with 422,712.18 betweenness score. We subsequently mapped the 14 genes identified from the integrated transcriptomics and proteomics data analysis (Figure 3) with the top genes that have at least 20 interactions from the PPI network analysis. We found 6 genes that were in common between these two datasets which represent the high-confidence GBM predictive surfaceome markers (Figure 4B).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eValidation of high-confidence GBM signature gene and survival-expression correlation analysis:\u003c/em\u003e\u003c/strong\u003e Next, we validated the expression profiles of the identified 6 high-confidence cell surface markers using an independent database, Gene Expression Profiling Interactive Analysis (GEPIA) [26]. GEPIA also combines the TCGA and GTEx gene expression data that were processed from raw reads count and unified using its own pipeline. In line with our findings, the identified GBM cell surface signature genes were confirmed to be significantly upregulated in the GBM GEPIA database (Supplementary Figure S4A \u0026ndash; S4F). To investigate whether the expression level of these signature genes would modulate/influence GBM patients\u0026rsquo; prognosis, we first performed the overall survival analyses on GBM patients who had high or low expression of each of these 6 genes (Supplementary Figure S5A \u0026ndash; S5F). \u0026nbsp;\u003cbr /\u003e However, there were no significant differences in the overall survival between patients who had high or low expression of these 6 individual prioritized genes. Since GBM patients have low overall survival rate (average \u0026lt;2 years\u0026rsquo; survival post-diagnosis), we postulated that it would be more appropriate to look at the disease-free survival endpoint rather than the overall survival. Moreover, \u0026nbsp;the overall survival endpoint is more suited for a longer follow-up period (typically 5 years) for the data to be meaningful [27]. Hence, we examined the disease-free survival profile of the GBM patients in a similar fashion. We found that high expression of \u003cem\u003eCD44\u003c/em\u003e, \u003cem\u003ePTPRJ\u003c/em\u003e and \u003cem\u003eHLA-DRA \u003c/em\u003ewere significantly correlated (p\u0026lt;0.05) with poor disease-free survival in GBM patients (Supplementary Figure S6A \u0026ndash; S6F).\u003c/p\u003e\n\u003cp\u003eIn addition to performing survival analysis on the individual gene, we also assessed whether combining the level of all 6 GBM signature genes as a group could predict the GBM patients\u0026rsquo; overall survival and disease-free survival. We observed that there was no statistically significant difference in the overall survival and disease-free survival between patients who had high expression and low expression of the signature group (Supplementary Figure S7A \u0026ndash; S7B). Interestingly, by combining only \u003cem\u003eCD44\u003c/em\u003e, \u003cem\u003ePTPRJ\u003c/em\u003e and \u003cem\u003eHLA-DRA\u003c/em\u003e in the gene signature, we found that subjects with high expression of this signature group had significantly poor disease-free survival (p\u0026lt;0.0084) compared to patients who had low expression of these genes (Supplementary Figure S8B). However, there was still no significant difference in the overall survival between GBM patients in this signature group (Supplementary Figure S8A). \u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eCo-expression network of CD44\u003c/em\u003e\u003c/strong\u003e: CD44 is a transmembrane receptor and has multifaceted functions in both normal and disease physiology. OMICS studies have identified CD44 to be overexpressed in many types of cancer including glioblastoma [28, 29]. Based on our analysis, CD44 seems particularly important as it can be both identified in transcriptomics and proteomics-based approaches, among the top hub gene and whose high expression correlate with poor disease-free survival. We performed a co-expression network analysis to further interrogate its association with other genes using our transcriptomics. The nodes represent in the network analysis represent genes, while the edges represent Pearson correlation above r\u0026gt;0.75. The neighbouring genes connected to CD44 was extracted and shown in Figure 5A. There are 27 genes in this complex connected to CD44. Among the highly correlated genes are ELK3, CLIC4, GALNT2, TNC, and VIM. All genes in this CD44 co-expression cluster are highly expressed in GBM compared to normal brain samples (Figure 5B), further corroborating the biological relevance of CD44 in supporting GBM pathogenesis. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eIdentification of drugs targeting GBM signature and CD44 network:\u003c/em\u003e\u003c/strong\u003e We next determined whether there are any clinically approved drugs targeting the identified high confidence GBM cell surface markers (Supplementary Figure S4) and components of the constructed CD44 co-expression network (Figure 5A). To achieve this objective, we utilized the Drugbank database (\u003ca href=\"https://go.drugbank.com/\"\u003ehttps://go.drugbank.com/\u003c/a\u003e) and our searches yielded several approved drugs that can be potentially effective or repurposed to target CD44, EGFR, C1R, CALR and TNFSR1A (Table 2). Hyaluronic acid, for example, is a clinically approved ligand for CD44 and this drug has been administered in the clinic to treat disease such as osteoarthritis [30]. Excessive hyaluronic acid administration has been demonstrated to inhibit tumour growth, possibly by impeding cell-cell interaction [31]. Besides, the use of nanomaterials to enhance the efficiency of hyaluronic acid delivery for cancer therapy is also actively being explored [32, 33]. Thus, the promising features of hyaluronic acid in mediating enhanced drugs or genes delivery to cancer cells via the overexpressed CD44 receptor could potentially be applied and developed for novel GBM therapeutic strategies. In regards to EGFR, several inhibitors and monoclonal antibodies have already been therapeutically approved to target this protein due to its roles as an important driver of tumorigenesis in many cancer types [34].\u003c/p\u003e\n\u003cp\u003eMoreover, of the 28 components of CD44 co-expression network (Figure 5A), only C1R, CALR, and TNFSR1A have drugs that can modulate them (Table 2). For instance, 3 drugs can be used or repurposed to target C1R. The pharmacological activity of Palivizumab to bind C1R subcomponent is under investigation, whereas the conestat alfa and human C1-esterase inhibitor can directly target C1R subcomponent and disrupt the complement system activation.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe surfaceome comprise cellular frontiers that permit/inhibit signal transduction as well as playing important roles in modulating cells proliferation, migration and invasion, and cells-cells interaction. The surfaceome can organize themselves at a nanoscale resolution [35]. This spatiotemporal nanoscale organization could define the cell identity and phenotypes, and capacity to communicate with microenvironments such as the extracellular matrix, growth factors, hormones and drugs. Due to their accessibility on the cell membrane, surfaceome proteins are ideal candidates for biomarkers and often targeted for drugs development. Over 50% of drugs curated in the DrugBank target the surfaceome. In addition to their ubiquitous expression on the plasma membrane, the extracellular stalks of these cell surface proteins can be cleaved and released into the bloodstream, making them as suitable targets for blood-based diagnostics. Surfaceome can also be draped with glycans during post-translational modifications, which will mediate their interaction with other proteins that reside on either the same or neighbouring cells as well as with the microenvironments [35]\u003cem\u003e. \u003c/em\u003eDysregulated surfaceome expressions and functions have been shown to promote tumour formation and progression [36]. Therefore, scientists have begun profiling and cataloguing surfaceome in various types of cancers [37\u0026ndash;40]. These cell surface proteins can be elevated in cancer cells in which they can respond to the increased level of growth factors, rendering cancer cells to sustain their infinite proliferative capabilities [41] and interact with the microenvironment that could either directly or indirectly modulate the tumour growth and metastatic capabilities [42].\u003c/p\u003e\n\u003cp\u003eThe GBM transcriptomics dataset have been previously utilized to uncover genes that support GBM pathogenesis as well as genes that have potential prognostic values [43\u0026ndash;45]. For example, Nicolasjilwan\u0026nbsp;et al.\u0026nbsp;analyzed the TCGA database to predict the survival of GBM patients based on clinical features, MRI images genomics alterations [43]. However, most TCGA GBM differential genes expression analyses either relied on low number of normal brain tissue samples, in which the TCGA GBM cohort contained only 5 normal brain tissues RNA-Seq data, or the data were combined with the GBM TCGA microarray data. This might create an imbalance that would lead to inaccuracy or bias in the downstream analysis. Hence, to increase the robustness of this study in identifying the significantly upregulated GBM surfaceome repertoire, we included the normal brain tissues GTEx RNA-Seq database TCGA in our analysis. On a similar scale, the GTEx studies have performed genes expression profiling in more than 11,000 samples across multiple human tissues from nearly 1,000 healthy donors. We compared the TCGA GBM and normal cortex GTEx RNA-seq data and identified 2,381 significant differentially expressed genes in GBM, in which 648 were upregulated and 1,733 downregulated genes. In agreement with previous GBM proteomics profiling study [12], the GO cellular compartment analysis showed that most of the dysregulated genes in GBM encode for the cell surface proteins, suggesting the importance of cell surface proteins in GBM pathogenesis.\u003c/p\u003e\n\u003cp\u003eOf the 2,381 significant DEGs in GBM, 395 genes encode for cell surface proteins, in which 124 and 271 genes were found to be significantly upregulated and downregulated, respectively. Interestingly, receptor subclass was the predominant dysregulated genes in GBM, suggesting the crucial roles of cell surface receptors in supporting GBM pathogenesis. This was indeed in line with several studies reporting the implications of cell surface receptors dysregulation in the pathogenesis of many cancer types [46]. For this reason, the development of cancer treatment strategies have been revolved around targeting the cell surface receptors such as the receptor tyrosine kinases (RTKs) [47] and G protein-coupled receptors (GPCRs) [48]. Therefore, targeting the cell surface proteins particularly the receptor subclass could potentially be further explored as novel GBM therapeutic options.\u003c/p\u003e\n\u003cp\u003eRobust cancer biomarkers are those that could be reproducibly identified by multi-omics platforms or reported in several different studies. To this end, we integrated the analyzed transcriptomics data with publicly available GBM proteomics data to prioritize for high-confidence cell surface proteins. Also, due to post-transcriptional and post-translational modifications, the mRNAs expression level are sometimes not correlated with their respective protein expression levels [49]. After mapping the prioritized genes from the transcriptomics-proteomics integrative analysis with the PPI network analysis data, we identified 6 genes; \u003cem\u003eHLA-DRA, CD44, SLC1A5, EGFR, ITGB2, PTPRJ\u003c/em\u003e, whereby we considered these genes as the high-confidence GBM predictive surface markers. Overall survival analyses revealed that there were no significant difference in the overall survival between patients who had high and low expression of these 6 genes, either the genes were analyzed individually or when combined together. However, when looking at the disease-free survival, patients who had high expression of \u003cem\u003eCD44, PTPRJ, \u003c/em\u003eand \u003cem\u003eHLA-\u003c/em\u003eDRA, either individually or as a group, had significantly poor disease-free survival (Supplementary Figure 6 and 8B) compared to subjects with low expression of the genes. These findings indicate that these 3 genes, \u003cem\u003eCD44, PTPRJ, \u003c/em\u003eand \u003cem\u003eHLA-DRA,\u003c/em\u003e could potentially be developed as GBM prognostic markers in the clinic.\u003c/p\u003e\n\u003cp\u003eIn addition to identifying the already known GBM drivers like CD44 and EGFR, our integrative analysis approach has also enabled us to identify potential novel genes that have not either been reported or thoroughly discussed in the context of GBM. For instance, within the 6 GBM signature genes, ITGB2 has not been widely associated with the pathogenesis of GBM. ITGB2 encodes for cell surface protein that is important in regulating cell adhesion and cell-surface mediated signalling [50]. Hence overexpression of this protein is relevant in promoting cancer growth possibly by modulating cancer cells adhesive and migratory properties, and the pro-oncogenic signalling cascades. Though there are \u003cem\u003ein-silico\u003c/em\u003e and \u003cem\u003ein-vitro\u003c/em\u003e studies that associated the ITGB2 as one of the important genes in cancer, the exact mechanisms of how this gene promotes GBM remains elusive and worth to be investigated in the future [11, 51, 52]. Human leukocyte antigen (HLA)-DRA is a classical major histocompatibility complex (MHC) class II molecule that plays important role in immune responses modulation. High expression of the HLA-DR gene family has been associated with more aggressive tumour grade in gliomas and poor prognosis [53, 54]. Nonetheless, the functions of HLA-DRA in driving GBM growth has not been fully elucidated.\u003c/p\u003e\n\u003cp\u003ePTPRJ gene is a member of the protein tyrosine phosphatase (PTP) family whose substrates include the RTKs such VEGFR, PDGFR and EGFR [55]. Since the RTKs pro-oncogenic properties are well-established in which their activation largely depends on phosphorylation, PTPRJ is thus deemed to function as tumour suppressor proteins due to its function as a phosphatase that can negatively regulate signalling pathway. This was also evidenced by the ectopic expression of PTPRJ in in-vitro models that resulted in cell growth inhibition [56, 57]. In contrast to these previous reports, we found that \u003cem\u003ePTPRJ \u003c/em\u003eexpression was upregulated in GBM and led us to suggest that PTPRJ might have a pro-oncogenic role in GBM pathogenesis. To our knowledge, there have been no previous reports linking \u003cem\u003ePTPRJ \u003c/em\u003eexpression and function with GBM pathogenesis. This notion of PTPRJ potential \u0026lsquo;double-edged sword\u0026rsquo; and GBM-specific pro-oncogenic function needs to be investigated further. SLC1A5, another hit target from our analysis, is a neutral amino acid transporter in which its high expression has been implicated in many cancer types including GBM [58]. In GBM, SLC1A5 expression is under the control of pro-oncogenic c-Myc protein but how this transporter supports the tumour cells proliferation and growth remain poorly understood [59].\u003c/p\u003e\n\u003cp\u003eAs highlighted above, the identification of CD44 and EGFR in this present study is expected because they have been previously described as one of the key targets for GBM [29, 60]. This validates the robustness of our approach in the sense that not only our analysis identified several novel genes, but also the findings overlap with previous studies. Since EGFR pro-oncogenic roles have been widely implicated in many cancer types and several drugs have been developed and clinically approved to target EGFR [34, 61, 62], we focused our analysis on CD44. The CD44 encodes for transmembrane glycoprotein that serves as the receptor for hyaluronic acid, a component of the extracellular matrix, and several other ligands including osteopontin, fibronectin and collagen [29]. The CD44 antigen has been implicated in modulating tumorigenesis in many cancer types in which high expression of this CD44 increases cancer cells proliferation, motility and survival as well as promoting cancer metastasis [63]. In GBM, high expression of CD44 was identified in the proteogenomic profiling of GBM tissues [23] and further classified as a GBM cell surface antigen in a systematic analysis [28]. Interestingly, this transmembrane glycoprotein can be cleaved and secreted into the vasculatures, suggesting its potential to be developed as a diagnostic marker [64]. It has been reported that the activation of CD44 by its ligand promotes cancer stem cell-like phenotypes in GBM and increased therapeutic resistance [65]. Consistent with this, drugs targeting CD44 are currently in clinical trials, and so far the results are promising in that CD44 inhibition impede GBM cells growth [66]. Our co-expression network analysis using a graph-based analytics [15] demonstrated that genes connected to CD44 were also highly co-expressed in GBM compared to normal brain tissues, suggesting that CD44 signalling axis is important in GBM tumorigenesis.\u003c/p\u003e\n\u003cp\u003eThe current approved therapies to treat GBM are far from satisfactory and have remained unchanged for more than a decade [67]. This includes the alkylating agent temozolomide, which is the first line of drug used in treating GBM. Therefore, there is a need for novel or alternative treatment strategies for GBM. Due to the upregulated expression of CD44 in GBM, drugs targeting CD44 are currently undergoing clinical trials and the results are thus far promising in that CD44 inhibition impedes GBM cells growth [66]. In addition to this, our drug mapping analysis revealed hyaluronic acid as an actionable CD44 binding molecule. It is therefore appealing to investigate the activity and potential use of this existing drug to treat GBM in the future, which has yet to be comprehensively studied. Within the CD44 co-expressed interactome, three additional targets already have drugs that can modulate them namely the C1R, CALR and TNFSR1A (Table 2). Based on our knowledge, the activity and efficacy of these drugs have not been tested in any \u003cem\u003ein-vitro\u003c/em\u003e or \u003cem\u003ein-vivo\u003c/em\u003e GBM models yet. Also, studying a combination of these available drugs targeting our GBM signature or the CD44 co-expression network could disrupt the aberrant hub gene interactome and potentially enhance GBM treatment efficacy.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn summary, we identified GBM surfaceome by combining RNA-seq data. Through an integrative multi-OMICS strategy, we highlighted 6 GBM surface-enriched genes that could be important in driving GBM development. Some of these genes can be targeted by clinically approved drugs for other diseases suggesting potential drug repurposing. Additionally, further studies of these genes could lead to potential GBM diagnostic/prognostic markers or a therapeutic regimen to treat GBM.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eGBM: Glioblastoma multiforme; TCGA: The cancer genome atlas; GTEx: Genotype-tissue expression; DEG: Differentially expressed pene; PPI:\u0026nbsp; Protein-protein interaction; t-SNE: t-distributed stochastic neighbour embedding; GO: Gene ontology; KEGG: and Kyoto encyclopedia of genes and genomes; GEPIA: Gene expression profiling interactive analysis; MHC: Major histocompatibility complex; RTK: Receptor tyrosine kinase; GPCR: G protein-coupled receptor; PTP: Protein tyrosine kinase.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBecause the present study did not use any patient samples, this is not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data are included within the manuscript and in the supplementary files. The TCGA GBM data can be obtained from the Genomics Data Commons Data Portal (https://portal.gdc.cancer.gov). The normal brain tissues RNA-seq data were obtained from the GTEx Portal (\u003ca href=\"https://gtexportal.org/home/datasets\"\u003ehttps://gtexportal.org/home/datasets\u003c/a\u003e). Other data are available from the corresponding author upon reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study is supported by the Fundamental Research Grant Scheme by the Ministry of Education, Malaysia (FRGS/1/2018/STG04/UKM/03/1) and Collaborative Research Programme - International Centre for Genetic Engineering and Biotechnology Grant (CRP/MYS19-04_EC). The funders had no role in this study.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; contribution\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConceptualization, MAM and SES; methodology, WFWMN; software, WFWN; formal analysis, WFWMN, MAM and SES; investigation, WFWMN, MAM and SES; resources, WFWMN and NAM; data curation, NAM and SBH.; writing\u0026mdash;original draft preparation, SES, MAM and WFWMN; supervision, MAM; writing\u0026mdash;review \u0026amp; editing, MAM and SES; funding acquisition, MAM. All authors have read and approved the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors thank David Shorthouse (MRC Cancer Unit, University of Cambridge) and Low Teck Yew (UKM Medical Molecular Biology Institute) for discussion, critical insight and proofreading the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor details\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e1 \u0026nbsp;\u0026nbsp;\u0026nbsp;\u003c/sup\u003eUKM Medical Molecular Biology Institute, UKM Medical Centre, Universiti Kebangsaan Malaysia, Bandar Tun Razak, 56000 Cheras, Kuala Lumpur, Malaysia.\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e2\u0026nbsp;\u0026nbsp; \u003c/sup\u003e\u0026nbsp;Faculty of Science and Natural Resources, Universiti Malaysia Sabah, 88400 Kota Kinabalu, Sabah, Malaysia.\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e3\u003c/sup\u003e Neurosurgery Division, Department of Surgery, Faculty of Medicine, Universiti Kebangsaan Malaysia, Bandar Tun Razak, 56000 Cheras, Kuala Lumpur, Malaysia.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eSiegel RL, Miller KD, Jemal A. Cancer statistics, 2016. CA Cancer J Clin. 2016;66:7\u0026ndash;30.\u003c/li\u003e\n\u003cli\u003eKamiya-Matsuoka C, Gilbert MR. Treating recurrent glioblastoma: an update. CNS Oncol. 2015;4:91\u0026ndash;104.\u003c/li\u003e\n\u003cli\u003eOhgaki H. Epidemiology of brain tumors. Methods Mol Biol Clifton NJ. 2009;472:323\u0026ndash;42.\u003c/li\u003e\n\u003cli\u003eQazi MA, Vora P, Venugopal C, Sidhu SS, Moffat J, Swanton C, et al. Intratumoral heterogeneity: pathways to treatment resistance and relapse in human glioblastoma. Ann Oncol. 2017;28:1448\u0026ndash;56.\u003c/li\u003e\n\u003cli\u003eShergalis A, Bankhead A, Luesakul U, Muangsin N, Neamati N. Current Challenges and Opportunities in Treating Glioblastoma. Pharmacol Rev. 2018;70:412\u0026ndash;45.\u003c/li\u003e\n\u003cli\u003eStupp R, Mason WP, van den Bent MJ, Weller M, Fisher B, Taphoorn MJB, et al. Radiotherapy plus concomitant and adjuvant temozolomide for glioblastoma. N Engl J Med. 2005;352:987\u0026ndash;96.\u003c/li\u003e\n\u003cli\u003eIto H, Nakashima H, Chiocca EA. Molecular responses to immune checkpoint blockade in glioblastoma. Nat Med. 2019;25:359.\u003c/li\u003e\n\u003cli\u003eNam JY, de Groot JF. Treatment of Glioblastoma. J Oncol Pract. 2017;13:629\u0026ndash;38.\u003c/li\u003e\n\u003cli\u003eCancer Genome Atlas Research Network. Comprehensive genomic characterization defines human glioblastoma genes and core pathways. Nature. 2008;455:1061\u0026ndash;8.\u003c/li\u003e\n\u003cli\u003ePearson JRD, Regad T. Targeting cellular pathways in glioblastoma multiforme. Signal Transduct Target Ther. 2017;2:17040.\u003c/li\u003e\n\u003cli\u003ePolisetty RV, Gautam P, Sharma R, Harsha HC, Nair SC, Gupta MK, et al. LC-MS/MS Analysis of Differentially Expressed Glioblastoma Membrane Proteome Reveals Altered Calcium Signaling and Other Protein Groups of Regulatory Functions. Mol Cell Proteomics. 2012;11:M111.013565.\u003c/li\u003e\n\u003cli\u003eBausch-Fluck D, Hofmann A, Bock T, Frei AP, Cerciello F, Jacobs A, et al. A Mass Spectrometric-Derived Cell Surface Protein Atlas. PLOS ONE. 2015;10:e0121314.\u003c/li\u003e\n\u003cli\u003eBausch-Fluck D, Goldmann U, M\u0026uuml;ller S, Oostrum M van, M\u0026uuml;ller M, Schubert OT, et al. The in silico human surfaceome. Proc Natl Acad Sci. 2018;115:E10988\u0026ndash;97.\u003c/li\u003e\n\u003cli\u003eXia J, Gill EE, Hancock REW. NetworkAnalyst for statistical, visual and network-based meta-analysis of gene expression data. Nat Protoc. 2015;10:823\u0026ndash;44.\u003c/li\u003e\n\u003cli\u003eTheocharidis A, van Dongen S, Enright AJ, Freeman TC. Network visualization and analysis of gene expression data using BioLayout Express(3D). Nat Protoc. 2009;4:1535\u0026ndash;50.\u003c/li\u003e\n\u003cli\u003eNakada M, Kita D, Watanabe T, Hayashi Y, Teng L, Pyko IV, et al. Aberrant Signaling Pathways in Glioma. Cancers. 2011;3:3242\u0026ndash;78.\u003c/li\u003e\n\u003cli\u003eCunha JPC da, Galante P a. F, Souza JE de, Souza RF de, Carvalho PM, Ohara DT, et al. Bioinformatics construction of the human cell surfaceome. Proc Natl Acad Sci. 2009;106:16752\u0026ndash;7.\u003c/li\u003e\n\u003cli\u003eLee JK, Bangayan NJ, Chai T, Smith BA, Pariva TE, Yun S, et al. Systemic surfaceome profiling identifies target antigens for immune-based therapy in subtypes of advanced prostate cancer. Proc Natl Acad Sci. 2018;115:E4473\u0026ndash;82.\u003c/li\u003e\n\u003cli\u003eAlm\u0026eacute;n MS, Nordstr\u0026ouml;m KJV, Fredriksson R, Schi\u0026ouml;th HB. Mapping the human membrane proteome: a majority of the human membrane proteins can be classified according to function and evolutionary origin. BMC Biol. 2009;7:50.\u003c/li\u003e\n\u003cli\u003eBanerjee HN, Mahaffey K, Riddick E, Banerjee A, Bhowmik N, Patra M. Search for a diagnostic/prognostic biomarker for the brain cancer glioblastoma multiforme by 2D-DIGE-MS technique. Mol Cell Biochem. 2012;367:59\u0026ndash;63.\u003c/li\u003e\n\u003cli\u003eCollet B, Guitton N, Sa\u0026iuml;kali S, Avril T, Pineau C, Hamlat A, et al. Differential analysis of glioblastoma multiforme proteome by a 2D-DIGE approach. Proteome Sci. 2011;9:16.\u003c/li\u003e\n\u003cli\u003eHeroux MS, Chesnik MA, Halligan BD, Al-Gizawiy M, Connelly JM, Mueller WM, et al. Comprehensive characterization of glioblastoma tumor tissues for biomarker identification using mass spectrometry-based label-free quantitative proteomics. Physiol Genomics. 2014;46:467\u0026ndash;81.\u003c/li\u003e\n\u003cli\u003eSong Y-C, Lu G-X, Zhang H-W, Zhong X-M, Cong X-L, Xue S-B, et al. Proteogenomic characterization and integrative analysis of glioblastoma multiforme. Oncotarget. 2017;8:97304\u0026ndash;12.\u003c/li\u003e\n\u003cli\u003eLow TY, Mohtar MA, Ang MY, Jamal R. Connecting Proteomics to Next-Generation Sequencing: Proteogenomics and Its Current Applications in Biology. Proteomics. 2019;19:e1800235.\u003c/li\u003e\n\u003cli\u003eAng MY, Low TY, Lee PY, Wan Mohamad Nazarie WF, Guryev V, Jamal R. Proteogenomics: From next-generation sequencing (NGS) and mass spectrometry-based proteomics to precision medicine. Clin Chim Acta. 2019;498:38\u0026ndash;46.\u003c/li\u003e\n\u003cli\u003eTang Z, Li C, Kang B, Gao G, Li C, Zhang Z. GEPIA: a web server for cancer and normal gene expression profiling and interactive analyses. Nucleic Acids Res. 2017;45:W98\u0026ndash;102.\u003c/li\u003e\n\u003cli\u003eSargent DJ, Wieand HS, Haller DG, Gray R, Benedetti JK, Buyse M, et al. Disease-Free Survival Versus Overall Survival As a Primary End Point for Adjuvant Colon Cancer Studies: Individual Patient Data From 20,898 Patients on 18 Randomized Trials. J Clin Oncol. 2005;23:8664\u0026ndash;70.\u003c/li\u003e\n\u003cli\u003eGhosh D, Funk CC, Caballero J, Shah N, Rouleau K, Earls JC, et al. A Cell-Surface Membrane Protein Signature for Glioblastoma. Cell Syst. 2017;4:516-529.e7.\u003c/li\u003e\n\u003cli\u003eChen C, Zhao S, Karnad A, Freeman JW. The biology and role of CD44 in cancer progression: therapeutic implications. J Hematol OncolJ Hematol Oncol. 2018;11:64.\u003c/li\u003e\n\u003cli\u003eBowman S, Awad ME, Hamrick MW, Hunter M, Fulzele S. Recent advances in hyaluronic acid based therapy for osteoarthritis. Clin Transl Med. 2018;7. doi:10.1186/s40169-017-0180-3.\u003c/li\u003e\n\u003cli\u003eMisra S, Hascall VC, Markwald RR, Ghatak S. Interactions between Hyaluronan and Its Receptors (CD44, RHAMM) Regulate the Activities of Inflammation and Cancer. Front Immunol. 2015;6. doi:10.3389/fimmu.2015.00201.\u003c/li\u003e\n\u003cli\u003eKim JH, Moon MJ, Kim DY, Heo SH, Jeong YY. Hyaluronic Acid-Based Nanomaterials for Cancer Therapy. Polymers. 2018;10. doi:10.3390/polym10101133.\u003c/li\u003e\n\u003cli\u003eKim K, Choi H, Choi ES, Park M-H, Ryu J-H. Hyaluronic Acid-Coated Nanomedicine for Targeted Cancer Therapy. Pharmaceutics. 2019;11.\u003c/li\u003e\n\u003cli\u003eSigismund S, Avanzato D, Lanzetti L. Emerging functions of the EGFR in cancer. Mol Oncol. 2018;12:3\u0026ndash;20.\u003c/li\u003e\n\u003cli\u003eBausch-Fluck D, Milani ES, Wollscheid B. Surfaceome nanoscale organization and extracellular interaction networks. Curr Opin Chem Biol. 2019;48:26\u0026ndash;33.\u003c/li\u003e\n\u003cli\u003eTeh JLF, Chen S. Glutamatergic signaling in cellular transformation. Pigment Cell Melanoma Res. 2012;25:331\u0026ndash;42.\u003c/li\u003e\n\u003cli\u003eMirkowska P, Hofmann A, Sedek L, Slamova L, Mejstrikova E, Szczepanski T, et al. Leukemia surfaceome analysis reveals new disease-associated features. Blood. 2013;121:e149\u0026ndash;59.\u003c/li\u003e\n\u003cli\u003eFenner A. Surfaceome profiling for NEPC target antigens. Nat Rev Urol. 2018;15:396\u0026ndash;7.\u003c/li\u003e\n\u003cli\u003eZiegler A, Cerciello F, Bigosch C, Bausch-Fluck D, Felley-Bosco E, Ossola R, et al. Proteomic surfaceome analysis of mesothelioma. Lung Cancer. 2012;75:189\u0026ndash;96.\u003c/li\u003e\n\u003cli\u003ePais H, Ruggero K, Zhang J, Al-Assar O, Bery N, Bhuller R, et al. Surfaceome interrogation using an RNA-seq approach highlights leukemia initiating cell biomarkers in an LMO2 T cell transgenic model. Sci Rep. 2019;9:1\u0026ndash;16.\u003c/li\u003e\n\u003cli\u003eHanahan D, Weinberg RA. Hallmarks of Cancer: The Next Generation. Cell. 2011;144:646\u0026ndash;74.\u003c/li\u003e\n\u003cli\u003eLeth-Larsen R, Lund RR, Ditzel HJ. Plasma membrane proteomics and its application in clinical cancer biomarker discovery. Mol Cell Proteomics MCP. 2010;9:1369\u0026ndash;82.\u003c/li\u003e\n\u003cli\u003eNicolasjilwan M, Hu Y, Yan C, Meerzaman D, Holder CA, Gutman D, et al. Addition of MR imaging features and genetic biomarkers strengthens glioblastoma survival prediction in TCGA patients. J Neuroradiol J Neuroradiol. 2015;42:212\u0026ndash;21.\u003c/li\u003e\n\u003cli\u003eHan J, Puri RK. Analysis of the cancer genome atlas (TCGA) database identifies an inverse relationship between interleukin-13 receptor \u0026alpha;1 and \u0026alpha;2 gene expression and poor prognosis and drug resistance in subjects with glioblastoma multiforme. J Neurooncol. 2018;136:463\u0026ndash;74.\u003c/li\u003e\n\u003cli\u003eJia D, Li S, Li D, Xue H, Yang D, Liu Y. Mining TCGA database for genes of prognostic value in glioblastoma microenvironment. Aging. 2018;10:592\u0026ndash;605.\u003c/li\u003e\n\u003cli\u003eSanchez-Vega F, Mina M, Armenia J, Chatila WK, Luna A, La KC, et al. Oncogenic Signaling Pathways in The Cancer Genome Atlas. Cell. 2018;173:321-337.e10.\u003c/li\u003e\n\u003cli\u003eRegad T. Targeting RTK Signaling Pathways in Cancer. Cancers. 2015;7:1758\u0026ndash;84.\u003c/li\u003e\n\u003cli\u003eLundstrom K. An Overview on GPCRs and Drug Discovery: Structure-Based Drug Design and Structural Biology on GPCRs. In: Leifert WR, editor. G Protein-Coupled Receptors in Drug Discovery. Totowa, NJ: Humana Press; 2009. p. 51\u0026ndash;66. doi:10.1007/978-1-60327-317-6_4.\u003c/li\u003e\n\u003cli\u003eVogel C, Marcotte EM. Insights into the regulation of protein abundance from proteomic and transcriptomic analyses. Nat Rev Genet. 2012;13:227\u0026ndash;32.\u003c/li\u003e\n\u003cli\u003eCamponeschi A, Gerasimcik N, Wang Y, Fredriksson T, Chen D, Farroni C, et al. Dissecting Integrin Expression and Function on Memory B Cells in Mice and Humans in Autoimmunity. Front Immunol. 2019;10:534.\u003c/li\u003e\n\u003cli\u003eWang A, Chen M, Wang H, Huang J, Bao Y, Gan X, et al. Cell Adhesion-Related Molecules Play a Key Role in Renal Cancer Progression by Multinetwork Analysis. BioMed Res Int. 2019;2019:2325765.\u003c/li\u003e\n\u003cli\u003eDunwoodie LJ, Poehlman WL, Ficklin SP, Feltus FA. Discovery and validation of a glioblastoma co-expressed gene module. Oncotarget. 2018;9:10995\u0026ndash;1008.\u003c/li\u003e\n\u003cli\u003eFan X, Liang J, Wu Z, Shan X, Qiao H, Jiang T. Expression of HLA-DR genes in gliomas: correlation with clinicopathological features and prognosis. Chin Neurosurg J. 2017;3:27.\u003c/li\u003e\n\u003cli\u003eDiao J, Xia T, Zhao H, Liu J, Li B, Zhang Z. Overexpression of HLA-DR is associated with prognosis of glioma patients. Int J Clin Exp Pathol. 2015;8:5485\u0026ndash;90.\u003c/li\u003e\n\u003cli\u003eGodfrey R, Arora D, Bauer R, Stopp S, M\u0026uuml;ller JP, Heinrich T, et al. Cell transformation by FLT3 ITD in acute myeloid leukemia involves oxidative inactivation of the tumor suppressor protein-tyrosine phosphatase DEP-1/ PTPRJ. Blood. 2012;119:4499\u0026ndash;511.\u003c/li\u003e\n\u003cli\u003eIuliano R, Trapasso F, Le Pera I, Schepis F, Sam\u0026agrave; I, Clodomiro A, et al. An adenovirus carrying the rat protein tyrosine phosphatase eta suppresses the growth of human thyroid carcinoma cell lines in vitro and in vivo. Cancer Res. 2003;63:882\u0026ndash;6.\u003c/li\u003e\n\u003cli\u003eMassa A, Barbieri F, Aiello C, Arena S, Pattarozzi A, Pirani P, et al. The expression of the phosphotyrosine phosphatase DEP-1/PTPeta dictates the responsivity of glioma cells to somatostatin inhibition of cell proliferation. J Biol Chem. 2004;279:29004\u0026ndash;12.\u003c/li\u003e\n\u003cli\u003eBhutia YD, Ganapathy V. Glutamine transporters in mammalian cells and their functions in physiology and cancer. Biochim Biophys Acta BBA - Mol Cell Res. 2016;1863:2531\u0026ndash;9.\u003c/li\u003e\n\u003cli\u003eWise DR, DeBerardinis RJ, Mancuso A, Sayed N, Zhang X-Y, Pfeiffer HK, et al. Myc regulates a transcriptional program that stimulates mitochondrial glutaminolysis and leads to glutamine addiction. Proc Natl Acad Sci. 2008;105:18782\u0026ndash;7.\u003c/li\u003e\n\u003cli\u003eWestphal M, Maire CL, Lamszus K. EGFR as a Target for Glioblastoma Treatment: An Unfulfilled Promise. CNS Drugs. 2017;31:723\u0026ndash;35.\u003c/li\u003e\n\u003cli\u003eSingh D, Attri BK, Gill RK, Bariwal J. Review on EGFR Inhibitors: Critical Updates. Mini Rev Med Chem. 2016;16:1134\u0026ndash;66.\u003c/li\u003e\n\u003cli\u003eHynes NE, Lane HA. ERBB receptors and cancer: the complexity of targeted inhibitors. Nat Rev Cancer. 2005;5:341\u0026ndash;54.\u003c/li\u003e\n\u003cli\u003eSenbanjo LT, Chellaiah MA. CD44: A Multifunctional Cell Surface Adhesion Receptor Is a Regulator of Progression and Metastasis of Cancer Cells. Front Cell Dev Biol. 2017;5. doi:10.3389/fcell.2017.00018.\u003c/li\u003e\n\u003cli\u003eLim S, Kim D, Ju S, Shin S, Cho I, Park S-H, et al. Glioblastoma-secreted soluble CD44 activates tau pathology in the brain. Exp Mol Med. 2018;50:1\u0026ndash;11.\u003c/li\u003e\n\u003cli\u003ePietras A, Katz AM, Ekstr\u0026ouml;m EJ, Wee B, Halliday JJ, Pitter KL, et al. Osteopontin-CD44 signaling in the glioma perivascular niche enhances cancer stem cell phenotypes and promotes aggressive tumor growth. Cell Stem Cell. 2014;14:357\u0026ndash;69.\u003c/li\u003e\n\u003cli\u003eMooney KL, Choy W, Sidhu S, Pelargos P, Bui TT, Voth B, et al. The role of CD44 in glioblastoma multiforme. J Clin Neurosci Off J Neurosurg Soc Australas. 2016;34:1\u0026ndash;5.\u003c/li\u003e\n\u003cli\u003eKazda T, Dziacky A, Burkon P, Pospisil P, Slavik M, Rehak Z, et al. Radiotherapy of Glioblastoma 15 Years after the Landmark Stupp\u0026rsquo;s Trial: More Controversies than Standards? Radiol Oncol. 2018;52:121\u0026ndash;8.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eDue to technical limitations, the tables are provided in the Supplementary Files section.\u003c/p\u003e\n"},{"header":"Supplementary Materials","content":"\u003cp\u003e\u003cstrong\u003eSupplementary Tables\u003c/strong\u003e\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003eSupplementary Table S1. Overall differentially expressed genes in TCGA GBM tissues vs. GTEx normal brain tissues.\u003c/li\u003e\n\u003cli\u003eSupplementary Table S2. Significantly dysregulated cell surface genes in TCGA GBM tissues vs. GTEx normal brain tissues.\u003c/li\u003e\n\u003cli\u003eSupplementary Table S3. GBM cell lines proteomics data from Bausch-Fluck et al. 2015.\u003c/li\u003e\n\u003cli\u003eSupplementary Table S4. GBM tissue samples proteomics data from Polisetty et. al 2012.\u003c/li\u003e\n\u003cli\u003eSupplementary Table S5. Protein-protein interaction network analysis of surfaceome.\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003e\u003cstrong\u003eSupplementary Figures\u003c/strong\u003e\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003eFigure S1. Gene ontology and deregulated pathways in GBM. (A-B) Gene ontology cellular component of the significantly (A) upregulated and (B) downregulated genes in GBM. (C-D) KEGG pathway analysis of the (C) upregulated and (D) downregulated genes in GBM.\u003c/li\u003e\n\u003cli\u003eFigure S2. Significant differentially expressed cell surface genes in GBM. (A) GBM surfaceome classification using previously annotated cell surface genes dataset identifies 395 DEGs that belongs to surfaceome. (B) Cell surface genes stratification from (A) based on its subclass.\u003c/li\u003e\n\u003cli\u003eFigure S3. KEGG pathway analysis of differentially expressed surfaceome in GBM. (A) Upregulated surfaceome and (B) Downregulated surfaceome.\u003c/li\u003e\n\u003cli\u003eFigure S4. Significant upregulation of the prioritized GBM surfaceome signature in GBM patients. (A-F) Boxplot showing the RNA-Seq data (transcript per million) of (A) CD44 (B) PTPRJ (C) SLC1A5 (D) EGFR (E) HLA-DRA and (F) ITGB2 in GBM and GTEx normal brain tissue samples..\u003c/li\u003e\n\u003cli\u003eFigure S5. Overall survival analysis of the prioritized GBM surfaceome signature as potential GBM prognostic biomarker. (A-F) Overall survival analysis of GBM patients having high and low expression of (A) CD44 (B) PTPRJ (C) SLC1A5 (D) EGFR (E) HLA-DRA and (F) ITGB2.\u003c/li\u003e\n\u003cli\u003eFigure S6. Disease-free survival analysis of the prioritized GBM surfaceome signature as potential GBM prognostic biomarker. (A-F) Disease-free survival analysis of GBM patients having high and low expression of (A) CD44 (B) PTPRJ (C) SLC1A5 (D) EGFR (E) HLA-DRA and (F) ITGB2.\u003c/li\u003e\n\u003cli\u003eFigure S7. Survival analysis of the 6 GBM signature genes. (A) Overall survival and (B) disease-free survival analysis of GBM patients having high and low expression of all 6 genes; CD44, PTPRJ, SLC1A5, EGFR, HLA-DRA and ITGB2\u003c/li\u003e\n\u003cli\u003eFigure S8. Survival analysis of the 3 GBM signature genes. (A) Overall survival and (B) disease-free survival analysis of GBM patients having high and low expression of CD44, PTPRJ and HLA-DRA.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-cancer","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bcan","sideBox":"Learn more about [BMC Cancer](http://bmccancer.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bcan/default.aspx","title":"BMC Cancer","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Differentially expressed genes, protein-protein interaction, cell surface proteins, network analysis, TCGA, GTEx","lastPublishedDoi":"10.21203/rs.3.rs-46071/v2","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-46071/v2","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground: \u003c/strong\u003eGlioblastoma multiforme (GBM) is a highly lethal, stage IV brain tumour with a prevalence of approximately 2 per 10000 people globally. The cell surface proteins or surfaceome serve as an information gateway in many oncogenic signalling pathways and are important in modulating cancer phenotypes. Dysregulation of surfaceome expression and activity have been shown to promote tumorigenesis. The expression of GBM surfaceome is a case in point; OMICS screening in cell-based system identified that this sub-proteome is largely perturbed in GBM. Additionally, since these cell surface proteins have ‘direct’ access to drugs, they are appealing targets for cancer therapy. However, a comprehensive aberrant GBM surfaceome landscape has not been fully defined. Thus, this study aimed to define GBM-specific surfaceome genes and identify key cell-surface genes that could potentially be developed as novel GBM biomarkers for therapeutic purposes. \u003c/p\u003e\u003cp\u003e\u003cstrong\u003eMethods: \u003c/strong\u003eWe integrated the RNA-Seq data from TCGA GBM (n=166) and GTEx normal brain cortex (n=408) databases to identify the significantly dysregulated surfaceome in GBM. This was followed by integrative analysis that combines transcriptomics, proteomics and protein-protein interaction network data to prioritize the high-confidence GBM surfaceome signature. \u003c/p\u003e\u003cp\u003e\u003cstrong\u003eResults: \u003c/strong\u003eOf the 2,381 significantly dysregulated genes in GBM, 395 genes were classified as surfaceome. Via the integrative analysis, we identified 6 high-confidence GBM molecular signature, HLA-DRA, CD44, SLC1A5, EGFR, ITGB2, PTPRJ, which were significantly upregulated in GBM. The expression of these genes were validated in an independent transcriptomics database, which confirmed their upregulated expression in GBM. Importantly, high expression of CD44, PTPRJ and HLA-DRA is significantly associated with poor disease-free survival. Last, using the Drugbank database, we identified several clinically-approved drugs targeting the GBM molecular signature suggesting potential drug repurposing. \u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConclusions: \u003c/strong\u003eIn summary, we identified and highlighted the key GBM surface-enriched repertoires that could be biologically relevant in supporting GBM pathogenesis. These genes could be further interrogated experimentally in future studies that could lead to efficient diagnostic/prognostic markers or potential treatment options for GBM.\u003c/p\u003e","manuscriptTitle":"Integration of RNA-Seq and proteomics data identifies glioblastoma multiforme surfaceome signature","msid":"","msnumber":"","nonDraftVersions":[{"code":2,"date":"2020-10-21 15:24:59","doi":"10.21203/rs.3.rs-46071/v2","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"","date":"2021-03-17T00:00:00+00:00","index":2,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2020-11-10T00:00:00+00:00","index":1,"fulltext":"Recommendation: Reviewer's comments unavailable pending editorial decision\n"},{"type":"reviewerAgreed","content":"","date":"2020-10-19T12:00:00+00:00","index":1,"fulltext":""},{"type":"editorAssigned","content":"","date":"2020-10-16T12:00:00+00:00","index":"","fulltext":""},{"type":"reviewersInvited","content":"","date":"2020-10-16T12:00:00+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2020-10-15T12:00:00+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2020-10-15T12:00:00+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-cancer","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bcan","sideBox":"Learn more about [BMC Cancer](http://bmccancer.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bcan/default.aspx","title":"BMC Cancer","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}},{"code":1,"date":"2020-07-28 01:12:10","doi":"10.21203/rs.3.rs-46071/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2020-09-15T12:00:00+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2020-09-14T12:00:00+00:00","index":2,"fulltext":"Recommendation: Major revisions required\nForm responses:\n---\n\nComments to Author:\n---\nDear authors,\n\nyour manuscript on the surfaceome of glioblastoma is, from my point of view, an interesting piece of work. I like the idea of specifically investigating the surfaceome and I think your analysis is comprehensive. I am particularly fond of the t-SNE calculations in the beginning and the drug database analysis at the end of the manuscript.\n\nHowever, before the manuscript is ready for publication, major revisions should be performed:\n\n1) Even though you seem to have checked the language by a native speaker, I still see sentences that do not seem impeccable gramatically, even in the abstract. One example is: \"The cell surface proteins or surfaceome play significant roles in modulating cancer phenotypes and acting as information gateway in many oncogenic signaling pathways.\"\n\n2) You bring some arguments why you focused on progression-free survival (or disease-free survival) and did not formally investigate overall survival. However, from a clinical point of view, overall survival is of course THE relevant parameter. I can not imagine that you did not look at it raising the suscipicion that you might have tried analysing it but did not find a significant association. So please either mention that in the manuscript - it is totally fine if nothing significant comes out of the overall survival analysis by the way - or please perform the overall survival analysis and add it to the manuscript. Importantly, this should also be discussed in the discussion section of the manuscript.\n\n3) The drugs you identify are not surprising and not an entirely novel suggestion in the context of glioblastoma. This should also be discussed more openly!\n\n4) Some minor corrections are necessary: E.g. in Figure 5a I guess it should say \"CD44\" and not \"CD4\".\n\n\n\n\nPlease include all comments for the authors in this box rather than uploading your report as an attachment. Please only upload as attachments annotated versions of manuscripts, graphs, supporting materials or other aspects of your report which cannot be included in a text format.\nPlease overwrite this text when adding your comments to the authors.* Publons Reviewer Recognition. Springer Nature can send verification of this review directly to Publons (a subsidiary of Clarivate Analytics). If you would like to take advantage of this service, please click on the “Yes” option below. Your name, email address, title of the reviewed manuscript, name of the journal, and date of your review submission (the “Review Data”) will then be transmitted to Publons upon publication of the manuscript. If you have already registered at Publons, they will notify you of the receipt of this review and update your profile as per your settings and their policy. If you are not registered with Publons, you will receive an email from them asking you to register in order for them to be able to recognize your review on your new profile page. Publons may use the Review Data to generate derivative metadata for the benefit of Publons and you as a reviewer, carefully considering the sensitivity of such information. For example, Publons may verify your record as a reviewer by updating your profile published on its webservice if you have registered for such service or help editors to identify candidate reviewers. Please find the details of processing in Publons’ privacy policy https://publons.com/about/terms: **No**\n* Declaration of competing interests: **I declare that I have no competing interests.**\n* Reviewer Publication Consent. I agree for my report to be made available under an Open Access Creative Commons CC-BY License (http://creativecommons.org/licenses/by/4.0) if this manuscript is accepted for publication. Any comments that I do not wish to be included in the published report have been included as confidential comments to the editor, which will not be published.: **I agree to the terms of the CC-BY 4.0 license; please do not publish my name with my report. (default)**\n* Is the study design appropriate to answer the research question (including the use of appropriate controls), and are the conclusions supported by the evidence presented?: **Yes**\n* Are the methods sufficiently described to allow the study to be repeated?: **Yes**\n* Is the use of statistics and treatment of uncertainties appropriate?: **Yes**\n* Is the presentation of the work clear?: **Yes**\n* Are the images in this manuscript (including electrophoretic gels and blots) free from apparent manipulation?: **Yes**\n"},{"type":"editorInvitedReview","content":"","date":"2020-09-02T12:00:00+00:00","index":1,"fulltext":"Recommendation: Major revisions required\nForm responses:\n---\n\nComments to Author:\n---\nThe authors have used TCGA and GTEx RNA-seq data to identify genes sets with altered expression in glioblastoma versus normal brain. Available proteomics data from tissues and cell lines was used to further filter candidates. Six genes were ultimately identified as being upregulated in glioblastoma with high confidence. The basic strategy used to identify these genes seems solid and the approach was explained clearly. There were several weaknesses in the submitted manuscript: TCGA data and tissue proteomics data have expression data on glioblastoma cells as well as other cell types in the samples such as activated microglia. The authors should check what cell types are expressing their hits. The sections on drug targets (last two paragraphs of the results and second last paragraph of discussion were poorly written. Cetuximab and bevacizumab do not bind C1R. The possible therapeutic use of hyaluronic acid in glioblastoma needs to described more carefully (probably limited to helping with drug delivery?) I would prefer to see a careful description of what is known about all six of the final genes identified - most of the discussion is focused on CD44. Most of the hits have been identified in earlier studies and, in some cases, there is supporting immunohistochemistry data for their overexpression. The relationship of the current study to this earlier work needs to be discussed more thoroughly.* Publons Reviewer Recognition. Springer Nature can send verification of this review directly to Publons (a subsidiary of Clarivate Analytics). If you would like to take advantage of this service, please click on the “Yes” option below. Your name, email address, title of the reviewed manuscript, name of the journal, and date of your review submission (the “Review Data”) will then be transmitted to Publons upon publication of the manuscript. If you have already registered at Publons, they will notify you of the receipt of this review and update your profile as per your settings and their policy. If you are not registered with Publons, you will receive an email from them asking you to register in order for them to be able to recognize your review on your new profile page. Publons may use the Review Data to generate derivative metadata for the benefit of Publons and you as a reviewer, carefully considering the sensitivity of such information. For example, Publons may verify your record as a reviewer by updating your profile published on its webservice if you have registered for such service or help editors to identify candidate reviewers. Please find the details of processing in Publons’ privacy policy https://publons.com/about/terms: **No**\n* Declaration of competing interests: **I declare that I have no competing interests**\n* Reviewer Publication Consent. I agree for my report to be made available under an Open Access Creative Commons CC-BY License (http://creativecommons.org/licenses/by/4.0) if this manuscript is accepted for publication. Any comments that I do not wish to be included in the published report have been included as confidential comments to the editor, which will not be published.: **I agree to the terms of the CC-BY 4.0 license; please publish my name with my report.**\n* Is the study design appropriate to answer the research question (including the use of appropriate controls), and are the conclusions supported by the evidence presented?: **Yes**\n* Are the methods sufficiently described to allow the study to be repeated?: **Yes**\n* Is the use of statistics and treatment of uncertainties appropriate?: **Yes**\n* Is the presentation of the work clear?: **Yes**\n* Are the images in this manuscript (including electrophoretic gels and blots) free from apparent manipulation?: **Yes**\n"},{"type":"reviewerAgreed","content":"","date":"2020-08-28T12:00:00+00:00","index":2,"fulltext":""},{"type":"reviewerAgreed","content":"","date":"2020-08-12T12:00:00+00:00","index":1,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2020-08-11T12:00:00+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2020-07-30T12:00:00+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2020-07-23T12:00:00+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2020-07-23T12:00:00+00:00","index":"","fulltext":""},{"type":"submitted","content":"","date":"2020-07-20T12:00:00+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-cancer","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bcan","sideBox":"Learn more about [BMC Cancer](http://bmccancer.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bcan/default.aspx","title":"BMC Cancer","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"912a8ed8-07c8-4042-b7f8-531f4826fe19","owner":[],"postedDate":"October 21st, 2020","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":209018,"name":"Oncology"},{"id":209019,"name":"Cancer Biology"}],"tags":[],"updatedAt":"2021-08-22T15:05:58+00:00","versionOfRecord":{"articleIdentity":"rs-46071","link":"https://doi.org/10.1186/s12885-021-08591-0","journal":{"identity":"bmc-cancer","isVorOnly":false,"title":"BMC Cancer"},"publishedOn":"2021-07-23 15:00:38","publishedOnDateReadable":"July 23rd, 2021"},"versionCreatedAt":"2020-10-21 15:24:59","video":"","vorDoi":"10.1186/s12885-021-08591-0","vorDoiUrl":"https://doi.org/10.1186/s12885-021-08591-0","workflowStages":[]},"version":"v2","identity":"rs-46071","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-46071","identity":"rs-46071","version":["v2"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00