An RNA seq-based reference landscape of human normal and neoplastic brain | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article An RNA seq-based reference landscape of human normal and neoplastic brain Sonali Arora, Frank Szulzewsky, Matt Jensen, Nicholas Nuechterlein, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2448083/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 14 Mar, 2023 Read the published version in Scientific Reports → Version 1 posted 8 You are reading this latest preprint version Abstract In order to better understand the relationship between normal and neoplastic brain, we combined five publicly available large-scale datasets, correcting for batch effects and applying Uniform Manifold Approximation and Projection (UMAP) to RNA-seq data. We assembled a reference Brain-UMAP including 702 adult gliomas, 802 pediatric tumors and 1409 healthy normal brain samples, which can be utilized to investigate the wealth of information obtained from combining several publicly available datasets to study a single organ site. Normal brain regions and tumor types create distinct clusters and because the landscape is generated by RNA seq, comparative gene expression profiles and gene ontology patterns are readily evident. To our knowledge, this is the first meta-analysis that allows for comparison of gene expression and pathways of interest across adult gliomas, pediatric brain tumors, and normal brain regions. We provide access to this resource via the open source, interactive online tool Oncoscape, where the scientific community can readily visualize clinical metadata, gene expression patterns, gene fusions, mutations, and copy number patterns for individual genes and pathway over this reference landscape. Biological sciences/Cancer/Cancer genomics Biological sciences/Cancer/Cns cancer Biological sciences/Computational biology and bioinformatics Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Introduction Over the past several decades the scientific community has characterized and cataloged individual genes for their function and involvement in development and disease. More recently, several single-cell atlases from different tissues and organ sites for various species (for example human and mouse) have been created that provide deep insights into the relationships between different cell types in development and in adult tissues. In this study, we have blended publicly available RNA seq datasets of normal and neoplastic brain to derive similar insights into the relationship between various central nervous system (CNS) tumors and between neoplastic vs. normal brain. Using this reference landscape, expression of specific genes and gene ontology groups can be compared across all tumor types and normal brain regions. A variety of omics approaches have been employed to characterize both tumor and healthy tissue at the molecular level by various large-scale international initiatives. The Cancer Genome Atlas (TCGA) 1 contains data from across 33 cancer types, including uniformly processed cancer genomic data (whole transcriptome RNA-Seq), microarray, gene fusions, gene mutations and copy number calls) from 702 glioma patients and 5 matched normal patients. The Chinese Glioma Genome Atlas (CGGA) 2 contains RNA-seq, whole genome sequencing, DNA methylation, microarray data from over 2000 brain tumor samples. The Children’s Brain Tumor Tissue Consortium (CBTTC) 3 contains whole genome sequencing and RNA-seq data across 23 different pediatric tumors. The Genotype Tissue Expression Project (GTEx) 4 , contains genomic data from 54 non-diseased tissue sites across nearly 1000 individuals, including 1409 brain tissue samples from 12 GTEx defined brain regions. Here, we present a visual integration approach for analyzing multiple diverse molecular datasets combining large numbers of samples across different brain regions and brain tumor subtypes. We combined transcriptional data from three different datasets for adult glioma, pediatric tumors and healthy normal brain regions, corrected for batch effects, and used a dimension reduction technique to construct a UMAP to find meaningful patterns in this pooled multi-disease dataset. This comparative study can help researchers and clinicians visualize similarities and differences in patient cohorts, study and compare alterations in gene expression, signaling pathways, gene fusions, copy number profiles and mutation calls across multiple tumor types. By adding normal healthy brain tissues from GTEx to our reference landscape, we also allow for comparisons between healthy and neoplastic states. Visualizing these similarities and differences in an opensource, interactive website, Oncoscape 5 ( https://oncoscape.sttrcancer.org/#project_bulkrnaseqbrainumap ), can aid in analysis for translational research. Results Constructing the Brain-UMAP / Clustering of gene expression data identifies diverse disease types To characterize and better understand the molecular intricacies of brain tumors, we downloaded uniformly processed RNA-seq abundances values from recount-brain, a curated repository for human brain RNA-seq datasets, for three different uniformly processed datasets − 702 adult glioma samples from TCGA 1 , 270 adult glioma samples from CGGA 5 , 6 , 1409 healthy normal brain samples from GTEx 4 across 12 GTEx-defined brain regions ( Supplementary Table 1a ). Retrieving data from recount 7 ensured that consistent bioinformatic pipelines were used for these three datasets thus resulting in no batch effects between the three datasets. The most common solid tumors in children are brain tumors with approximately 1.15 to 5.14 cases per 100,000 children in the United States alone 8 . To adequately represent a wide range of CNS tumors in our reference landscape, we additionally included 802 pediatric tumor samples ( Supplementary Table 1b) from the Children Brain Tumor Network (CBTN) 3 . Figure 1 a represents an overview of data sources (details in Supplementary Table 1c ) Gene expression data from each of the datasets was converted to units of Transcripts per million (TPM) to avoid inter-pipeline difference and were limited to a common set of 19142 protein-coding genes 9 . While both CBTN and recount used different bioinformatic pipelines ( Supplementary Table 1c ), in order to ensure that there were no batch effects we used combat 10 method from the R package “sva” to remove unwanted variation in our combined dataset. We explored three different dimension reduction techniques (Principal Component Analysis (PCA), t-distributed Stochastic Neighbor Embedding (tSNE) and uniform manifold approximation and projection (UMAP) for data visualization. We chose UMAP to build a Brain-UMAP (Fig. 1 b) on batch corrected TPM integrated counts, as UMAP segregated the mini clusters well and was very effective in visualizing clusters and their relative proximities ( Supplementary Fig. 1a ). While 2-dimenisonal representation of the data is helpful, we also provide a 3-dimensional representation of the data on the interactive web-based platform Oncoscape 5 ) where users can easily toggle among and compare different patient groups, while using a suite of interoperable tools. At a first glance, distinct clustering of samples is observed. The adult glioma samples from the different datasets, TCGA-LGG, TCGA-GBM and CGGA, co-localized closely within the same region of the UMAP, whereas the pediatric tumor samples clustered between the GTEx healthy normal brain and the adult glioma samples. (Fig. 1 b). Distinct gene expression profiles in normal human brain The 1409 normal brain samples segregated into two distinct clusters of multiple supratentorial regions and cerebellum, as we have previously shown 11 (Fig. 1 c). The supratentorial regions further revealed three anatomically distinct regions for basal ganglia (caudate, nucleus accumbens, putamen), cortex (amygdala, Brodmann Area 24, cerebral cortex), and the midline structures (spinal cord, substantia nigra and hypothalamus). We confirmed that different classifiers such as postmortem interval (PMI), age (years), sex, Hardy score and type of sample preparation ( Supplementary Fig. 1b ) were not associated with distinct clustering patterns, suggesting that the sample clusters we observe are based on actual biological differences between these brain regions rather than sample preparation parameters. However, we observed that a few samples with much lower RNA integrity number (RIN) compared to all the other samples from different brain regions converged at a point ( Supplementary Fig. 1b ). Clustering of transcriptomic glioma datasets reveals distinct glioma subtype Within the adult glioma clusters, we observe that while the samples from the two TCGA datasets cluster together, there is a more continuous pattern when looking at gene expression profiles for samples from TCGA-GBM and TCGA-LGG (Fig. 2 a). This contrasts with the distinct clusters observed by analyzing whole exome single nucleotide alterations (SNAs) and whole genome copy number alterations (CNAs) from the same patients, as shown by Bolouri 12 et al. In line with previous reports, we observed that the age of the patients at diagnosis for the TCGA-GBM and TCGA-LGG samples revealed a sharp gradient illustrating the known correlation between age and outcome (Fig. 2 b). By contrast, patient gender was not associated with any specific clusters ( Supplementary Fig. 2a ). Regarding chromosomal alterations, tumors with a gain of chromosome 7 and hemizygous deletion of chromosome 10 (Fig. 2 c) or co-gain of chromosome 19 and 20 ( Supplementary Fig. 2b ) co-localized in the top area of the UMAP containing the TCGA-GBM samples (Fig. 2 c). Tumors exhibiting a co-deletion of chromosome 1p and 19q and mutation in isocitrate dehydrogenase 1 and 2 genes (IDH-mut) (Fig. 2 d) were concentrated in the lower half of the UMAP containing the TCGA-LGG samples. Tumors containing mutation in IDH1 (Fig. 2 e), TP53 (Fig. 2 f) and ATRX (Fig. 2 g) were also found to be concentrated and clustered together in specific regions of the adult glioma UMAP. Using the supervised DNA methylation classification, transcriptional subtype, MGMT promoter status and TERT promoter status ( Supplementary Fig. 2c-f ) to color the adult gliomas from TCGA-LGG and TCGA-GBM also revealed distinct patterns across the adult glioma landscape. Selecting for common glioma mutations and copy number alterations clearly shows three distinct subtypes of glioma - IDHmut-1p19q co-deleted oligodendrogliomas, the IDH mutated astrocytomas with p53 and ATRX mutations, and the wild-type IDH (IDH-wt ) glioblastomas with gain of chromosome 7 and loss of chromosome 10 molecular GBM. (Fig. 2 h). Adult glioma subtypes from TCGA and CGGA show similar gene expression profiles We next assessed if glioma samples of similar molecular subtypes from the TCGA and CGGA datasets exerted similar gene expression profiles and co-localized in their respective clusters. Similar to the TCGA samples, using grade, IDH mutation status and chromosome 1p 19q codeletion status for the CGGA samples, three distinct subtypes of adult glioma from the CGGA samples were observed. Interestingly, we observed that the oligodendrogliomas, astrocytomas and IDH-wt tumors from both CGGA and TCGA colocalized (Fig. 3 a-f). Separated from the large cluster containing adult glioma subtypes, there were two small clusters of gliomas from the CGGA dataset. Based on their grade and IDH mutation status, one cluster consisted of a mix of IDH mutant grade 2 and grade 3 oligodendrogliomas while the second cluster consisted of grade 4 IDH-wt glioblastomas. For the remainder of the paper, we will refer to the adult glioma datasets from TCGA and CGGA by their molecular subtypes (oligodendrogliomas, astrocytomas and IDH-wt glioblastomas). We analyzed the survival for each subtype for all three glioma subtypes and observed similar survival rates between TCGA and CGGA datasets for the respective subtypes. (Fig. 3 g). We then used a nearest neighbor approach to predict the survival for different UMAP subregions. Survival was predicted for samples with the median survival of its nearest neighbors, present in close proximity on the UMAP landscape (Fig. 3 h). We observed a small number of glioma samples formed an isthmus connecting to the normal brain samples. These samples were characterized by a low amount of copy number alterations and were a mix of IDH-wt glioblastomas, IDH-mut oligodendrogliomas or astrocytoma. These tumors were characterized by longer survival compared to other gliomas of their respective molecular subtype. It is conceivable that in some samples this region may represent either early forms of gliomas or a mix of glioma and normal brain. Taken together, our results show that the transcriptomic data from different datasets can be combined to generate a population of adult gliomas, creating a landscape where the location of specific tumor sample can be predictive of subtype and outcome. Pediatric Tumors cluster by disease type We added the 802 pediatric tumors from the CBTN dataset to our analysis to compare their gene expression patterns to both normal brain and adult glioma samples (Fig. 4 a-b). We observed the formation of distinct subclusters for several tumor types that correlated with established molecular subgroups. As an example, we observed that medulloblastoma samples were split into three distinct clusters that correlated with known Medulloblastoma subtypes 13 (Wnt, Sonic hedgehog (SHH) and groups 3,4, Supplementary Fig. 4a ). Similarly, Ependymomas (EPN) samples formed several clusters that correlated with the anatomic tumor location (supratentorial (ST)-EPN, spinal-EPN, and posterior fossa (PF)-EPN) 13 (Fig. 4 a, Supplementary Fig. 4b ). Pediatric pilocytic astrocytomas (PAs) and pediatric low-grade gliomas clustered closely together suggesting that they exert similar gene expression patterns. The schwannomas separated from the pediatric tumors and were located near the neurofibromas, which were localized adjacent to the malignant peripheral nerve sheath tumors (MPNST). The subependymal giant cell astrocytoma (SEGA) form a tight cluster as do the meningiomas. Interestingly, we observed that neurocytomas, DNET, ganglioglioma clustered near normal brain samples, specifically the hypothalamus and amygdala samples. Building a UMAP with just normal brain samples and the pediatric tumors, showed similar clustering profile ( Supplementary Fig. 4c ). Taken together, these results suggest that pediatric tumors cluster by disease type and also form subclusters made by subtypes in the case of medulloblastomas and ependymomas. Using the reference landscape to understand pathway regulation Alterations in signaling pathways are a hallmark of cancer and understanding the extent to which these pathways are dysregulated in tumor samples compared to healthy normal brain can help inform researchers about the underlying mechanisms of different cancer types. Bulk gene expression from adult glioma, pediatric brain tumors and healthy brain samples was subjected to a Gene Set variation Analysis (GSVA) and the gene set variation scores for each pathway were used to color in the Brain-UMAP. A score closer to 1 represented up-regulation of pathway in the given samples, whereas a score closer to -1 represented down-regulation of the pathway. We calculated GSVA scores for all pathways present in Reactome 14 , KEGG 15 and biocarta 16 pathways. We then tested whether there is a difference between the GSVA enrichment scores from different pair of phenotypes using a linear model and moderated t-statistic. As examples, we found that 605, 589, and 529 ( Supplementary Table 2a-c ) pathways were up-regulated in IDH-wt glioblastomas, IDH-mut astrocytomas and oligodendrogliomas respectively compared to healthy brain samples from GTEx. A total of 456 pathways ( Supplementary Table 2d ) were up regulated in all three adult glioma subtypes compared to healthy brain. The top pathways which were enriched in all adult glioma subtypes were pathways enriched for cell cycle, DNA repair, translation, splicing, oncogenic signaling pathways such as RAS pathway, Notch pathway, MHC pathway, PI3K/AKT Signaling, Wnt pathway, SHH pathway (Fig. 5 , Supplementary Fig. 5a ). Additionally, neurotransmitter pathways, calcium channel, potassium channel opening pathways were up regulated in healthy brain regions compared to pediatric tumors and adult gliomas ( Supplementary Fig. 5b ). Interestingly, we noted that the two small clusters (IDH mutated grade 2 and grade3 oligodendroglioma and grade 4 IDH-wt glioblastomas) from the CGGA dataset were enriched in pathways related to olfaction, glucoronidation, ascorbate and aldarate metabolism and xenobiotics ( Supplementary Fig. 5c ) in comparison to the main adult glioma cluster. While visualizing pathways across the reference Brain-UMAP is extremely informative, researchers exploring targets for drug development may be also interested in investigating individual genes of a particular pathway. For example, Reactome’s mismatch repair pathway (R-HAS-5358508) and Biocarta RELA pathway (M10183) were both found to be up-regulated in all adult glioma subtypes compared to healthy normal brain, coloring in the gene expression for individual genes over the Brain-UMAP show different gene expression patterns across members of the same pathway. (Fig. 6 , Supplementary Fig. 6 ). For example, almost all the genes in the mismatch repair pathway have elevated gene expression levels in medulloblastoma, except RPA3 and POLD4. Studying candidate genes at multiple genomic levels While bulk gene expression exhibited illuminating patterns in our data, analyzing other genomic information such as copy number alteration, gene fusions, or somatic alteration for each of these tumors can further enhance our understanding for a given gene of interest. Since processed copy number, gene fusions and somatic variants were publicly available only for two out of five datasets (TCGA and CBTN), we first built a much smaller UMAP using only the bulk gene expression data from these two datasets. The resulting UMAP (Fig. 7 a) showed a similar clustering pattern as our original Brain-UMAP. Next, we downloaded the copy number calls, gene fusion calls and somatic variants for adult gliomas from Genomic Data Commons (GDC) and the pediatric tumors from CBTN. Of note, the IDH-mut oligodendrogliomas, astrocytoma and pediatric high-grade gliomas have the highest mutational burden and number of gene fusions across all brain diseases (Fig. 7 b-c, Supplementary Fig. 7a-b ) in comparison to copy number profiles (Fig. 7 d-e) which show a mixture of genes with amplified and lost copy number profiles across all brain diseases. Gene fusions are another class of potential oncogenic drivers in cancer, including pediatric cancers. 17 We observed that specific gene fusions and/or gene fusion partners were enriched in specific cancer subtypes. Adult IDH-wt glioblastoma frequently harbored gene fusions involving EGFR (17.7% of tumors, most commonly EGFR − PSPHP1, EGFR − LINC01445, EGFR − SEC61G − DT), whereas pediatric lower-grade gliomas frequently harbored BRAF fusions (KIAA1549 − BRAF in 32% of low-grade gliomas and in 60% of pilocytic astrocytomas). In addition, supratentorial ependymomas most commonly harbored the C11orf95 − RELA fusion (71% of tumors, also known as ZFTA-RELA) and meningiomas most frequently harbored fusions in NF2 and YAP1 (both 15% of tumors) (Supplementary Fig. 7c). Armed with additional genomic information such as gene fusions, copy number variation and somatic alternation, we investigated members of the Reactome mismatch repair pathway (R-HAS-5358508) ( Supplementary Fig. 8a-c ). We observe that different genes belonging to this pathway exhibit different trends, for example genes such as POLD1(chr19), RPA2(chr1), and LIG1(chr19) loose a copy in IDH-mut oligodendrogliomas, while genes RPA3(chr19), PMS2(chr19), PCNA(chr20) and POLD2(chr7) gain a copy in IDH-wt GBM. While most genes show elevated gene expression levels for all brain tumors, of interest is EXO1 and RPA3 which are only up-regulated in IDH-wt GBM samples. While members of this pathway do not form gene fusions, they get mutated in different brain tumors. For example, in pediatric high grade tumors, we observed mutations in all the members of the mismatch repair pathway - MSH2 (22%), MSH6 (15.71%), POLD3 (14.29%), MSH3 (12.86%), LIG1(11.43%), EXO1(10%), PMS2(10%), PCNA(10%), POLD1 (8.57%), RPA2(8.57%), MLH1 (7.14%), RPA1(7.14%), POLD2(5.71%), POLD4(5.71%) and RPA3(5.71%). Oncogenes show altered gene expression in tumor samples, leading to abnormal phenotype in samples. Understanding gene expression patterns across various cancers of the nervous system can further our understanding of the disease. When studying known oncogenes such as EGFR, PTEN and CIC at a gene level (Fig. 8 ), as expected, we observe high number of mutations, high number of gene fusions, amplified gene expression values and copy number gains for EGFR across IDH-wt GBM. This contrasts with PTEN which shows loss of 1 copy in IDH-wt GBM samples. CIC, a transcriptional repressor, shows high number of mutations and copy number loss in oligodendrogliomas. For pediatric tumors, we observe BRAF gene fusions (Fig. 8 ) in 63% pilocytic astrocytoma tumors and 34% of low grade pediatric tumors ( Supplementary Table 3a ) ALK (Fig. 8 ) mutations are also observed in 38% high-grade pediatric tumors, 33% spinal cord ependymomas, 22% ATRT and 21% of the medulloblastoma tumors ( Supplementary Table 3b ).This reference landscape is a useful research tool for the scientific community, where researchers can explore existing data to increase their understanding of oncogenic pathways and individual genes that make up these pathways, potentially uncovering candidates for novel therapeutic targets. By providing access to this reference landscape via an open source website like Oncoscape, we provide an interactive tool to researchers doing CNS research. Discussion As costs for performing RNASeq continues to decline with increasing technological advances, more and more tumors will be sequenced, and additional tumor banks will be created with the underlying goal to understanding cancer’s complexity. The wealth of knowledge that that already exists in publicly available datasets such as those described here (GTEx, TCGA, CGGA, CBTN) is remarkable. Individually these datasets comprise of well-defined biologically similar set of patient samples and allow the analysis in exquisite resolution of genetic changes from one sample to the next. For example, comparing medulloblastomas to other medulloblastomas allows for precise characterization of molecular subtypes. As researchers, we can get so focused on comparing like with like that we lose sight of the proverbial forest for focusing too much on the leaves of a single tree. By integrating multiple datasets while correcting for batch effects, such as with the Brain-UMAP presented here, we can harness the power of multiple datasets. This landscape can be used to prospectively compare patients, similar to work that has been done using methylation arrays. Unlike methylation arrays, however, RNA seq is based on gene expression and therefore each cluster represented on the landscape contains granular information about the underlying tumor biology of the samples it contains. Because the overall expression pattern is identifiable for each tumor, this allows for cross comparison of various kinds of cancer to allow for characterization of tumor types not previously known. For example, asking questions about expression of particular genes or pathways across a wide panel of samples and tumor cohorts may uncover previously unknown roles or similarities between adult and pediatric tumor subtypes, which can possibly open new avenues for therapeutic discovery. The landscape that we present reveals correlations between tumor type and outcome. We have learned how diagnosis of a given tumor type can be confirmed using various pieces of genomic information. For example – the adult glioma subtype of oligodendroglioma can be confirmed by presence of co-deletion of chr1p/19q and somatic alteration in IDH1. We are also able to compare tumor samples which were sequenced years apart (TCGA and CGGA) in two different continents and confirm that these differences did not contribute to expression changes in majority of the tumor samples. We are also able to identify novel entities and gain insight into genes that drive their unique character, such as in the case of the two adult glioma subtypes seen only in the CCGA data that appear to strongly and uniquely express genes involved in olfaction. Using this study as an example, we have seen how one diagnosis can be comprised of more than one expression entity, such as in the case of medulloblastoma and ependymoma. The UMAP also indicates room for better classification of tumors. For example, we found tumors which were documented as one type in the publicly available database, but in our UMAP were found to cluster with tumors of a different type. For example, some embryonal tumors ended up clustering with the medulloblastoma samples indicating that they may be medulloblastomas, or at least share many common features with medulloblastoma. We also have seen that multiple diagnoses may really be one entity as in the case of tumors diagnosed as either pilocytic astrocytoma (PA) or pediatric low-grade glioma. Reference landscapes like the Brain-UMAP, will be informative for researchers who wish to obtain a quick diagnosis or characterization of newly obtained tumor samples. For example, if we look at the medulloblastoma subtypes from CBTN, there are 14 group 3, 49 group 4, 30 SHH, 9 WNT and 19 unclassified medulloblastoma samples. For the 19 unclassified samples, we can predict which subtype they belong to, based on which medulloblastoma subtype samples they cluster with. (Supp Fig. 4 a). By combining results from both RNA-seq data (gene expression, gene fusions) as well as whole genome sequencing (copy number, mutation calls) and with the help of a few examples, we illustrate the knowledge that can be mined from this resource, at both a pathway level as well as a gene level. The work presented here demonstrates the utility of reference landscapes for combining genomic data across multiple tumor types for both diagnosis, prognosis and better understanding the biology of the tumors that are similar to a given patient collected prospectively. The power of visualizing gene expression changes, regulation of pathways, chromosomal alteration, gene fusions across multiple tumor types is informative to every researcher, especially those who may not have immediate access to computational biology experts. The reference landscape described here provides a useful tool for researchers interested in gene level questions across large scale patient data, while the methods used for integrating data sources highlight the tremendous potential for combining future datasets with existing resources to address complex biological questions. Methods All analyses were performed in R ( https://www.r-project.org/ ) using Bioconductor ( https://www.bioconductor.org/ ) packages. We have deposited all scripts, associated data, at: https://zenodo.org/badge/latestdoi/584982012 to maximize transparency and reproducibility. TCGA data from GDC was downloaded using R/Bioconductor package TCGAbiolinks 18 , 19 , SummarizedExperiment 20 was used to store adult glioma and pediatric tumor data. All plots were made using ggplot2 21 , RcolorBrewer 22 . Obtaining gene expression RNASeq Data RNA Seq gene expression was downloaded from two sources for adult gliomas, GTEx defined healthy brain samples and pediatric tumors respectively (Supplemental Table 1). Conversion of abundance estimates to transcripts per million (TPM) For consistency, we converted all FPKM gene expression data to TPM data using the formula $$TPM=\frac{\text{F}\text{P}\text{K}\text{M}}{\sum _{all genes}\text{F}\text{P}\text{K}\text{M}} \times {10}^{6}$$ as described by Collins et al 23 . Uniform Manifold Approximation and Project (UMAP) Uniform Manifold Approximation and Project (UMAP) values were generated using the R function umap() using all protein-coding genes and visualized with R package ggplot2 21 . Obtaining Copy Number Gistic2 thresholded gene level copy number variation estimated using the GSITIC2 method were downloaded from UCSC Xena( https://xenabrowser.net/datapages/?cohort=TCGA%20Lower%20Grade%20Glioma%20(LGG)&removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443 and https://xenabrowser.net/datapages/?cohort=TCGA%20Glioblastoma%20(GBM)&removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443 ) for adult gliomas. Copy Number calls for pediatric tumors estimated from the GISTIC2 pipeline were downloaded from https://github.com/AlexsLemonade/OpenPBTA-analysis . Obtaining Gene Fusions Gene fusions for adult gliomas estimated using Arriba and Star-Fusion as per the GDC RNA-Seq pipeline were downloaded from the GDC portal ( https://portal.gdc.cancer.gov/projects/TCGA-GBM , https://portal.gdc.cancer.gov/projects/TCGA-LGG ). Gene fusions for pediatric tumors were downloaded from https://github.com/AlexsLemonade/OpenPBTA-analysis . Only high-confidence gene fusion calls were retained from ARRIBA. Obtaining Mutations Annotated Variant call Format (VCF) files containing somatic variants from MuSE, VarScan2, MuTect2 and Somatic Sniper were downloaded from GDC portal for TCGA-GBM and TCGA-LGG. The variants for each patient were combined based on custom script present in our github repository. Somatic Variant calls for pediatric tumors was obtained from https://github.com/AlexsLemonade/OpenPBTA-analysis . GSVA Pathway Analysis Gene sets for all the pathways from Biocarta, Kyoto Encylcopedia of Genes and Genomes (KEGG) and Reactome pathways were downloaded from Molecular Signature Databases (MSigDB)(v7.2) 24 . https://www.gsea-msigdb.org/gsea/msigdb/collections.jsp#C2 Batch corrected log2(TPM) counts from each pipeline were used to conduct a Gene Set Variation Analysis (GSVA) 25 using Biocarta, KEGG and Reactome pathways. GSVA scores obtained from 1 and − 1 for each sample, were visualized using ggplot2. Kaplan-Meier curves for TCGA and CGGA patients To generate Kaplan-Meier curves for TCGA and CGGA patients, all samples labeled recurrent, secondary, or normal tissue were removed. Duplicate samples were also excluded so that each patient was represented by exactly one sample. Kaplan-Meier curves were drawn by the Python package lifelines (Python version 3.8.6, lifelines version 0.25.2) 26 using survival data and glioma subtype labels downloaded from GDC, where the world health organization (WHO) 2016 criteria for the classification of adult diffuse gliomas was used to determine glioma subtype 27 . Survival-Annotated UMAP To annotate the Brain-UMAP with survival data, each TCGA and CGGA sample was colored with the median survival of a cohort of patients (nearest neighbors) close to the sample in question on the UMAP landscape. The notion of nearest neighbors was defined as a number of samples within a radius of 2 from the sample in question under the constraint that all nearest neighbors must be of the same glioma subtype and from the same dataset (TCGA or CGGA) as the sample in question. The radius parameter was chosen qualitatively. The number of nearest neighbors was defined as 25% of the total number of samples of the same glioma subtype and dataset as the sample in question. If the median survival of a sample’s cohort of nearest neighbors is undefined, or if a sample has fewer than 10 nearest neighbors, the point is not colored in. Non-primary and duplicate samples were excluded in the same manner as was done for the Kaplan-Meier curve analysis. Oncoscape integration Gene expression and survival data analysis result files were prepared according to the oncoscape instructions here: https://github.com/FredHutch/OncoscapeV3/blob/master/docs/upload.md Data Analysis All statistical analyses and plots were done in R (v.3.3.1) as implemented in Rstudio (v.1.0.136). Plots were created using the R basic graphics. The following R packages were used: GenomicAlignments53 (v.1.24), reshape64 (v.0.8.8), png65 (v.0.1–7), ape66 (v.5.3) and seqinr67 (v.3.6–1). Declarations Acknowledgements We thank members of the Holland Lab at Fred Hutchinson Cancer Research Center for discussions. This research was supported by the R35 CA253119-01A1(E.C.H.), NIH U54 CA193461 (E.C.H.), National Institutes of Health R01 CA195718 (E.C.H.), R01 CA100688 (E.C.H.), T32 CA9657-25 (S.S.P.), U54 DK106829 (S.S.P.), R21 CA223531 (S.S.P.); K22 CA258953-01 (SSP) ,Jacobs Foundation Research Fellowship (S.S.P.) and National Science Foundation Graduate Research Fellowship Program DGE-1762114 (N.N.). Author Information Contributions S.A., S.S.P., F.S., and E.C.H. conceived and organized the study based on S.A.’s proposed design. S.A. and N.N. performed all the analysis. M.J. incorporated the data into oncoscape. S.A. and S.S.P. wrote the manuscript, with contributions from all authors. Corresponding Author Correspondence to Eric C Holland Competing Interests The authors have stated explicitly that there are no conflicts of interest in connection with this article. Data availability All data are available in the main text, supplementary materials and/or accompanying database is found at https://zenodo.org/badge/latestdoi/584982012. We used data from recount2 (https://jhubiostatistics.shinyapps.io/recount/) , CBTN (https://github.com/AlexsLemonade/OpenPBTA-analysis) , copy number data for TCGA-GBM ( https://xenabrowser.net/datapages/?cohort=TCGA%20Glioblastoma%20(GBM)&removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443 ) and TCGA-LGG(https://xenabrowser.net/datapages/?cohort=TCGA%20Lower%20Grade%20Glioma%20(LGG)&removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443 ) , Gene fusions from GDC Data portal (https://portal.gdc.cancer.gov/projects/TCGA-GBM , https://portal.gdc.cancer.gov/projects/TCGA-LGG ) Code availability All custom code used in this study is available at https://zenodo.org/badge/latestdoi/584982012 References Cancer Genome Atlas Research, N. , et al. The Cancer Genome Atlas Pan-Cancer analysis project. Nat Genet 45 , 1113-1120 (2013). Zhao, Z. , et al. Chinese Glioma Genome Atlas (CGGA): A Comprehensive Resource with Functional Genomic Data from Chinese Glioma Patients. Genomics Proteomics Bioinformatics 19 , 1-12 (2021). Ijaz, H. , et al. Pediatric high-grade glioma resources from the Children's Brain Tumor Tissue Consortium. Neuro Oncol 22 , 163-165 (2020). Carithers, L.J. , et al. A Novel Approach to High-Quality Postmortem Tissue Procurement: The GTEx Project. Biopreserv Biobank 13 , 311-319 (2015). McFerrin, L.G. , et al. Analysis and visualization of linked molecular and clinical cancer data by using Oncoscape. Nat Genet 50 , 1203-1204 (2018). Shapiro, J.A. , et al. OpenPBTA: An Open Pediatric Brain Tumor Atlas. bioRxiv (2022). Collado-Torres, L. , et al. Reproducible RNA-seq analysis using recount2. Nat Biotechnol 35 , 319-321 (2017). Subramanian S, A.T. Childhood Brain Tumors. In: StatPearls [Internet]. Treasure Island (FL): StatPearls Publishing , https://www.ncbi.nlm.nih.gov/books/NBK535415/ (2022). Arora, S., Pattwell, S.S., Holland, E.C. & Bolouri, H. Variability in estimated gene expression among commonly used RNA-seq pipelines. Sci Rep 10 , 2734 (2020). Johnson, W.E., Li, C. & Rabinovic, A. Adjusting batch effects in microarray expression data using empirical Bayes methods. Biostatistics 8 , 118-127 (2007). Pattwell, S.S. , et al. A kinase-deficient NTRK2 splice variant predominates in glioma and amplifies several oncogenic signaling pathways. Nat Commun 11 , 2977 (2020). Bolouri, H., Zhao, L.P. & Holland, E.C. Big data visualization identifies the multidimensional molecular landscape of human gliomas. Proc Natl Acad Sci U S A 113 , 5394-5399 (2016). Pollack, I.F., Agnihotri, S. & Broniscer, A. Childhood brain tumors: current management, biological insights, and future directions. J Neurosurg Pediatr 23 , 261-273 (2019). Gillespie, M. , et al. The reactome pathway knowledgebase 2022. Nucleic Acids Res 50 , D687-d692 (2022). Kanehisa, M. & Goto, S. KEGG: kyoto encyclopedia of genes and genomes. Nucleic Acids Res 28 , 27-30 (2000). BioCarta. Biotech Software & Internet Report 2 , 117-120 (2001). Szulzewsky, F. , et al. Both YAP1-MAML2 and constitutively active YAP1 drive the formation of tumors that resemble NF2 mutant meningiomas in mice. Genes Dev 36 , 857-870 (2022). Colaprico, A. , et al. TCGAbiolinks: an R/Bioconductor package for integrative analysis of TCGA data. Nucleic Acids Res 44 , e71 (2016). Silva, T.C. , et al. TCGA Workflow: Analyze cancer genomics and epigenomics data using Bioconductor packages. F1000Res 5 , 1542 (2016). Morgan M, O.V., Hester J, Pagès H. SummarizedExperiment: SummarizedExperiment container. R package version 1.16.0. (2019). Wickham, H. ggplot2: Elegant Graphics for Data Analysis , (Springer-Verlag New York, 2016). Neuwirth, E. Package ‘RColorBrewer’ , ColorBrewer Palettes. (2014). Bo Li, C.N.D. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics (2011). Aravind Subramanian, P.T., Vamsi K. Mootha, Sayan Mukherjee, Benjamin L. Ebert, Michael A. Gillette, Amanda Paulovich, Scott L. Pomeroy, Todd R. Golub, Eric S. Lander, Jill P. Mesirov. Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profiles. PNAS (2005). Sonja Hänzelmann, R.C.J.G. GSVA: gene set variation analysis for microarray and RNA-Seq data. BMC Bioinformatics (2013). Davidson-Pilon, C. lifelines: survival analysis in Python. Journal of Open Source Software 4 (2019). Louis, D.N. , et al. The 2016 World Health Organization Classification of Tumors of the Central Nervous System: a summary. Acta Neuropathol 131 , 803-820 (2016). Additional Declarations No competing interests reported. Supplementary Files Supptables.xlsx BrainUMAPSuppFigs132023.pdf Cite Share Download PDF Status: Published Journal Publication published 14 Mar, 2023 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Major revision 14 Feb, 2023 Reviews received at journal 27 Jan, 2023 Reviewers agreed at journal 27 Jan, 2023 Reviewers invited by journal 24 Jan, 2023 Editor assigned by journal 24 Jan, 2023 Editor invited by journal 06 Jan, 2023 Submission checks completed at journal 06 Jan, 2023 First submitted to journal 05 Jan, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2448083","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":165644036,"identity":"a3129380-8717-4ce6-a5e2-9c21a9d84ce6","order_by":0,"name":"Sonali Arora","email":"","orcid":"","institution":"Fred Hutchinson Cancer Center","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Sonali","middleName":"","lastName":"Arora","suffix":""},{"id":165644037,"identity":"d9618520-4e3d-462d-894e-688b014a5eec","order_by":1,"name":"Frank Szulzewsky","email":"","orcid":"","institution":"Fred Hutchinson Cancer Center","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Frank","middleName":"","lastName":"Szulzewsky","suffix":""},{"id":165644038,"identity":"48f0ad00-29b9-47e1-9683-c5ea2b21c5d7","order_by":2,"name":"Matt Jensen","email":"","orcid":"","institution":"Fred Hutchinson Cancer Center","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Matt","middleName":"","lastName":"Jensen","suffix":""},{"id":165644039,"identity":"88579def-e0a4-4c3a-891b-1e1bfc13d4d7","order_by":3,"name":"Nicholas Nuechterlein","email":"","orcid":"","institution":"University of Washington","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Nicholas","middleName":"","lastName":"Nuechterlein","suffix":""},{"id":165644040,"identity":"ff166f10-18d4-47e3-aca2-af3dae36dae4","order_by":4,"name":"Siobhan S Pattwell","email":"","orcid":"","institution":"Ben Towne Center for Childhood Cancer Research, Seattle Children’s Research Institute","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Siobhan","middleName":"S","lastName":"Pattwell","suffix":""},{"id":165644041,"identity":"48392b99-e39c-4beb-9038-0b25a55074cf","order_by":5,"name":"Eric C. Holland","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA9UlEQVRIie3NsUoDQRSF4Tss3FTLtrOEvMMuW1r4KjMI2vgAFiITAreSpN2geYdAIFheGZgq6bcIuEGwC6QSURGTLawcsulE5i8PfByAUOgvJkFwDSBGEQPEplkOEmC1I2NSDZHtydS1JUl3wKyuV72Zi9f15sHeJHdGPF+9+0l67xQr91LMXafIJwsr5YqjfDn0k6y6zFih1fMnwm5MVmZSYdq/9ZPThnxZPSPsfLQimdwRTVZPETH6IebNT2R1rlgPbVE6jNIJXaRlpQd53/hJUp491ttX2xsRiu2GTpL9sjaffvJbwoCg48i+I19CoVDoX/cNcuhbWWam80EAAAAASUVORK5CYII=","orcid":"","institution":"Fred Hutchinson Cancer Center","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Eric","middleName":"C.","lastName":"Holland","suffix":""}],"badges":[],"createdAt":"2023-01-05 21:14:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2448083/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2448083/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-023-31180-z","type":"published","date":"2023-03-14T20:02:43+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":31373908,"identity":"b4bdb8c9-5ca8-4cda-9624-a971c3f573fc","added_by":"auto","created_at":"2023-01-10 16:11:51","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1344741,"visible":true,"origin":"","legend":"\u003cp\u003eOverview of data analyzed (a) showing datasets used, batch correction and construction of Brain-UMAP. (b) UMAP of complete dataset including adult gliomas, pediatric tumors and GTEx-defined normal brain. (c) UMAP showing unique clustering of GTEx-defined brain-regions\u003c/p\u003e","description":"","filename":"BrainUMAPMainFIGS132023Fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-2448083/v1/6e252a3e4d6a7dd9a125e769.png"},{"id":31372071,"identity":"032fce48-8f7c-437e-aa7f-10cc131ef592","added_by":"auto","created_at":"2023-01-10 15:55:51","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1556842,"visible":true,"origin":"","legend":"\u003cp\u003eUMAP for adult glioma showing patients colored in by (a) TCGA-GBM and TCGA-LGG patients (b) age at diagnosis (c) Chr 7 gain/ Chr10 loss in patients. (d) Chr 1p/19q co-deletion status in patients (e) IDH1 mutation (f) TP53 mutation (g) ATRX mutation. (h) UMAP identifying the 3 distinct adult glioma subtypes – IDH wildtype, Astrocytoma and Oligodendrogliomas.\u003c/p\u003e","description":"","filename":"BrainUMAPMainFIGS132023Fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-2448083/v1/b1921034259f8b48db2b51c0.png"},{"id":31373090,"identity":"a8a12fe4-f6bd-421e-9641-eb352ac9c051","added_by":"auto","created_at":"2023-01-10 16:03:51","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":1722016,"visible":true,"origin":"","legend":"\u003cp\u003eCo-localization of adult gliomas from two publicly available datasets TCGA and CGGA. Top panel shows (a) grade, (c) IDH mutation status and (e) 1p19q co-deletion status for adult gliomas from TCGA colored in, and adult gliomas from CGGA greyed out. Bottom panel shows (b) grade, (d) IDH mutation status and (f) 1p19q co-deletion status for adult gliomas from CGGA colored in, and adult gliomas from TCGA greyed out. (g) Survival analysis for adult glioma subtypes IDH wildtype(red), Astrocytoma (blue) and Oligodendroglioma (green) from TCGA and CGGA shown in solid and dotted lines respectively. (h) Prediction of survival time (in years) using a nearest neighbor approach shows a gradient for adult gliomas.\u003c/p\u003e","description":"","filename":"BrainUMAPMainFIGS132023Fig3.png","url":"https://assets-eu.researchsquare.com/files/rs-2448083/v1/db7448c3fd715914c39af5e7.png"},{"id":31372070,"identity":"d9df3b28-462e-4d3a-af11-c969c996b81b","added_by":"auto","created_at":"2023-01-10 15:55:51","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1450163,"visible":true,"origin":"","legend":"\u003cp\u003e(a) UMAP of pediatric tumors (b) Updated coloring of the Brain-UMAP showing pediatric tumors and three subtypes for the adult gliomas.\u003c/p\u003e","description":"","filename":"BrainUMAPMainFIGS132023Fig4.png","url":"https://assets-eu.researchsquare.com/files/rs-2448083/v1/14b3043f353b64625dffdbdb.png"},{"id":31372075,"identity":"4aca8c89-32e6-41ad-8ef5-a2844b42fe42","added_by":"auto","created_at":"2023-01-10 15:55:51","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":4434738,"visible":true,"origin":"","legend":"\u003cp\u003eVisualization of GSVA Pathway scores across Brain-UMAP for cancer pathways and cellular processes.\u003c/p\u003e","description":"","filename":"BrainUMAPMainFIGS132023Fig5.png","url":"https://assets-eu.researchsquare.com/files/rs-2448083/v1/5b79319496fb85a74c7421a4.png"},{"id":31373092,"identity":"354fda7e-8843-4851-9de4-86770fd57ae2","added_by":"auto","created_at":"2023-01-10 16:03:51","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":3180346,"visible":true,"origin":"","legend":"\u003cp\u003eVisualization of gene expression profiles for genes from the Reactome mismatch repair pathway across the Brain-UMAP\u003c/p\u003e","description":"","filename":"BrainUMAPMainFIGS132023Fig6.png","url":"https://assets-eu.researchsquare.com/files/rs-2448083/v1/ea775e2ab0bd13c28b0e0aee.png"},{"id":31373094,"identity":"da688ff3-4004-4c9a-ad6d-8506eb068fc8","added_by":"auto","created_at":"2023-01-10 16:03:51","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":2652174,"visible":true,"origin":"","legend":"\u003cp\u003e(a) UMAP of pediatric tumors and adult glioma subtypes from TCGA. Coloring in UMAP of pediatric tumors and adult glioma subtypes from TCGA by (b) number of point mutations and (c) number of gene fusions per tumor (d) number of genes with copies gained per tumor and (e) number of genes with copies deleted per tumor.\u003c/p\u003e","description":"","filename":"BrainUMAPMainFIGS132023Fig7.png","url":"https://assets-eu.researchsquare.com/files/rs-2448083/v1/23d6acc8d42173c1c5a4c9a8.png"},{"id":31373909,"identity":"6a9177b1-9bc1-41ff-8cbb-76fb5dc62b57","added_by":"auto","created_at":"2023-01-10 16:11:51","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":2825387,"visible":true,"origin":"","legend":"\u003cp\u003eIntegration and visualization of genomic information such as gene expression, mutation, copy number and gene fusions at a single gene level across Brain-UMAP for 5 genes – EGFR, PTEN, CIC, BRAF and ALK.\u003c/p\u003e","description":"","filename":"BrainUMAPMainFIGS132023Fig8.png","url":"https://assets-eu.researchsquare.com/files/rs-2448083/v1/5c9162c7cb04f3b8321c5433.png"},{"id":44722839,"identity":"02005f59-4fcb-4f0a-9284-66cc4aa9876f","added_by":"auto","created_at":"2023-10-16 20:08:41","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":4238607,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2448083/v1/40198774-e550-40f5-8391-a7555efd91d4.pdf"},{"id":31372067,"identity":"6cf44fd8-f71d-48aa-a5d7-f1ddb7424ded","added_by":"auto","created_at":"2023-01-10 15:55:51","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":176934,"visible":true,"origin":"","legend":"","description":"","filename":"Supptables.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-2448083/v1/5e2fab4e9f14554c5d0b843b.xlsx"},{"id":31374333,"identity":"89a3a33d-1a04-4ca4-bee6-6a194d959ee2","added_by":"auto","created_at":"2023-01-10 16:19:51","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":3047161,"visible":true,"origin":"","legend":"","description":"","filename":"BrainUMAPSuppFigs132023.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2448083/v1/a5dedf923e3947dea3c31a31.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"An RNA seq-based reference landscape of human normal and neoplastic brain","fulltext":[{"header":"Introduction","content":"\u003cp\u003eOver the past several decades the scientific community has characterized and cataloged individual genes for their function and involvement in development and disease. More recently, several single-cell atlases from different tissues and organ sites for various species (for example human and mouse) have been created that provide deep insights into the relationships between different cell types in development and in adult tissues. In this study, we have blended publicly available RNA seq datasets of normal and neoplastic brain to derive similar insights into the relationship between various central nervous system (CNS) tumors and between neoplastic vs. normal brain. Using this reference landscape, expression of specific genes and gene ontology groups can be compared across all tumor types and normal brain regions.\u003c/p\u003e \u003cp\u003eA variety of omics approaches have been employed to characterize both tumor and healthy tissue at the molecular level by various large-scale international initiatives. The Cancer Genome Atlas (TCGA)\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e contains data from across 33 cancer types, including uniformly processed cancer genomic data (whole transcriptome RNA-Seq), microarray, gene fusions, gene mutations and copy number calls) from 702 glioma patients and 5 matched normal patients. The Chinese Glioma Genome Atlas (CGGA)\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e contains RNA-seq, whole genome sequencing, DNA methylation, microarray data from over 2000 brain tumor samples. The Children\u0026rsquo;s Brain Tumor Tissue Consortium (CBTTC)\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e contains whole genome sequencing and RNA-seq data across 23 different pediatric tumors. The Genotype Tissue Expression Project (GTEx)\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e, contains genomic data from 54 non-diseased tissue sites across nearly 1000 individuals, including 1409 brain tissue samples from 12 GTEx defined brain regions.\u003c/p\u003e \u003cp\u003eHere, we present a visual integration approach for analyzing multiple diverse molecular datasets combining large numbers of samples across different brain regions and brain tumor subtypes. We combined transcriptional data from three different datasets for adult glioma, pediatric tumors and healthy normal brain regions, corrected for batch effects, and used a dimension reduction technique to construct a UMAP to find meaningful patterns in this pooled multi-disease dataset.\u003c/p\u003e \u003cp\u003eThis comparative study can help researchers and clinicians visualize similarities and differences in patient cohorts, study and compare alterations in gene expression, signaling pathways, gene fusions, copy number profiles and mutation calls across multiple tumor types. By adding normal healthy brain tissues from GTEx to our reference landscape, we also allow for comparisons between healthy and neoplastic states. Visualizing these similarities and differences in an opensource, interactive website, Oncoscape \u003csup\u003e5\u003c/sup\u003e(\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://oncoscape.sttrcancer.org/#project_bulkrnaseqbrainumap\u003c/span\u003e\u003cspan address=\"https://oncoscape.sttrcancer.org/#project_bulkrnaseqbrainumap\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), can aid in analysis for translational research.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eConstructing the Brain-UMAP / Clustering of gene expression data identifies diverse disease types\u003c/h2\u003e \u003cp\u003eTo characterize and better understand the molecular intricacies of brain tumors, we downloaded uniformly processed RNA-seq abundances values from recount-brain, a curated repository for human brain RNA-seq datasets, for three different uniformly processed datasets \u0026minus;\u0026thinsp;702 adult glioma samples from TCGA\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e, 270 adult glioma samples from CGGA\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e, 1409 healthy normal brain samples from GTEx\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e across 12 GTEx-defined brain regions (\u003cb\u003eSupplementary Table\u0026nbsp;1a\u003c/b\u003e). Retrieving data from recount\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e ensured that consistent bioinformatic pipelines were used for these three datasets thus resulting in no batch effects between the three datasets.\u003c/p\u003e \u003cp\u003eThe most common solid tumors in children are brain tumors with approximately 1.15 to 5.14 cases per 100,000 children in the United States alone\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. To adequately represent a wide range of CNS tumors in our reference landscape, we additionally included 802 pediatric tumor samples (\u003cb\u003eSupplementary Table\u0026nbsp;1b)\u003c/b\u003e from the Children Brain Tumor Network (CBTN)\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea represents an overview of data sources (details in \u003cb\u003eSupplementary Table\u0026nbsp;1c\u003c/b\u003e)\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eGene expression data from each of the datasets was converted to units of Transcripts per million (TPM) to avoid inter-pipeline difference and were limited to a common set of 19142 protein-coding genes\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. While both CBTN and recount used different bioinformatic pipelines (\u003cb\u003eSupplementary Table\u0026nbsp;1c\u003c/b\u003e), in order to ensure that there were no batch effects we used combat\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e method from the R package \u0026ldquo;sva\u0026rdquo; to remove unwanted variation in our combined dataset.\u003c/p\u003e \u003cp\u003eWe explored three different dimension reduction techniques (Principal Component Analysis (PCA), t-distributed Stochastic Neighbor Embedding (tSNE) and uniform manifold approximation and projection (UMAP) for data visualization. We chose UMAP to build a Brain-UMAP (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb) on batch corrected TPM integrated counts, as UMAP segregated the mini clusters well and was very effective in visualizing clusters and their relative proximities (\u003cb\u003eSupplementary Fig.\u0026nbsp;1a\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eWhile 2-dimenisonal representation of the data is helpful, we also provide a 3-dimensional representation of the data on the interactive web-based platform Oncoscape\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e) where users can easily toggle among and compare different patient groups, while using a suite of interoperable tools.\u003c/p\u003e \u003cp\u003eAt a first glance, distinct clustering of samples is observed. The adult glioma samples from the different datasets, TCGA-LGG, TCGA-GBM and CGGA, co-localized closely within the same region of the UMAP, whereas the pediatric tumor samples clustered between the GTEx healthy normal brain and the adult glioma samples. (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eDistinct gene expression profiles in normal human brain\u003c/h2\u003e \u003cp\u003eThe 1409 normal brain samples segregated into two distinct clusters of multiple supratentorial regions and cerebellum, as we have previously shown\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ec). The supratentorial regions further revealed three anatomically distinct regions for basal ganglia (caudate, nucleus accumbens, putamen), cortex (amygdala, Brodmann Area 24, cerebral cortex), and the midline structures (spinal cord, substantia nigra and hypothalamus). We confirmed that different classifiers such as postmortem interval (PMI), age (years), sex, Hardy score and type of sample preparation (\u003cb\u003eSupplementary Fig.\u0026nbsp;1b\u003c/b\u003e) were not associated with distinct clustering patterns, suggesting that the sample clusters we observe are based on actual biological differences between these brain regions rather than sample preparation parameters. However, we observed that a few samples with much lower RNA integrity number (RIN) compared to all the other samples from different brain regions converged at a point (\u003cb\u003eSupplementary Fig.\u0026nbsp;1b\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eClustering of transcriptomic glioma datasets reveals distinct glioma subtype\u003c/h2\u003e \u003cp\u003eWithin the adult glioma clusters, we observe that while the samples from the two TCGA datasets cluster together, there is a more continuous pattern when looking at gene expression profiles for samples from TCGA-GBM and TCGA-LGG (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ea). This contrasts with the distinct clusters observed by analyzing whole exome single nucleotide alterations (SNAs) and whole genome copy number alterations (CNAs) from the same patients, as shown by Bolouri\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e et al. In line with previous reports, we observed that the age of the patients at diagnosis for the TCGA-GBM and TCGA-LGG samples revealed a sharp gradient illustrating the known correlation between age and outcome (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb). By contrast, patient gender was not associated with any specific clusters (\u003cb\u003eSupplementary Fig.\u0026nbsp;2a\u003c/b\u003e). Regarding chromosomal alterations, tumors with a gain of chromosome 7 and hemizygous deletion of chromosome 10 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ec) or co-gain of chromosome 19 and 20 (\u003cb\u003eSupplementary Fig.\u0026nbsp;2b\u003c/b\u003e) co-localized in the top area of the UMAP containing the TCGA-GBM samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ec). Tumors exhibiting a co-deletion of chromosome 1p and 19q and mutation in isocitrate dehydrogenase 1 and 2 genes (IDH-mut) (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ed) were concentrated in the lower half of the UMAP containing the TCGA-LGG samples. Tumors containing mutation in IDH1 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ee), TP53 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ef) and ATRX (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eg) were also found to be concentrated and clustered together in specific regions of the adult glioma UMAP. Using the supervised DNA methylation classification, transcriptional subtype, MGMT promoter status and TERT promoter status (\u003cb\u003eSupplementary Fig.\u0026nbsp;2c-f\u003c/b\u003e) to color the adult gliomas from TCGA-LGG and TCGA-GBM also revealed distinct patterns across the adult glioma landscape. Selecting for common glioma mutations and copy number alterations clearly shows three distinct subtypes of glioma - IDHmut-1p19q co-deleted oligodendrogliomas, the IDH mutated astrocytomas with p53 and ATRX mutations, and the wild-type IDH (IDH-wt\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e glioblastomas with gain of chromosome 7 and loss of chromosome 10 molecular GBM. (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eh).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eAdult glioma subtypes from TCGA and CGGA show similar gene expression profiles\u003c/h2\u003e \u003cp\u003eWe next assessed if glioma samples of similar molecular subtypes from the TCGA and CGGA datasets exerted similar gene expression profiles and co-localized in their respective clusters. Similar to the TCGA samples, using grade, IDH mutation status and chromosome 1p 19q codeletion status for the CGGA samples, three distinct subtypes of adult glioma from the CGGA samples were observed. Interestingly, we observed that the oligodendrogliomas, astrocytomas and IDH-wt tumors from both CGGA and TCGA colocalized (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea-f). Separated from the large cluster containing adult glioma subtypes, there were two small clusters of gliomas from the CGGA dataset. Based on their grade and IDH mutation status, one cluster consisted of a mix of IDH mutant grade 2 and grade 3 oligodendrogliomas while the second cluster consisted of grade 4 IDH-wt glioblastomas. For the remainder of the paper, we will refer to the adult glioma datasets from TCGA and CGGA by their molecular subtypes (oligodendrogliomas, astrocytomas and IDH-wt glioblastomas).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe analyzed the survival for each subtype for all three glioma subtypes and observed similar survival rates between TCGA and CGGA datasets for the respective subtypes. (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eg). We then used a nearest neighbor approach to predict the survival for different UMAP subregions. Survival was predicted for samples with the median survival of its nearest neighbors, present in close proximity on the UMAP landscape (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eh).\u003c/p\u003e \u003cp\u003eWe observed a small number of glioma samples formed an isthmus connecting to the normal brain samples. These samples were characterized by a low amount of copy number alterations and were a mix of IDH-wt glioblastomas, IDH-mut oligodendrogliomas or astrocytoma. These tumors were characterized by longer survival compared to other gliomas of their respective molecular subtype. It is conceivable that in some samples this region may represent either early forms of gliomas or a mix of glioma and normal brain.\u003c/p\u003e \u003cp\u003eTaken together, our results show that the transcriptomic data from different datasets can be combined to generate a population of adult gliomas, creating a landscape where the location of specific tumor sample can be predictive of subtype and outcome.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003ePediatric Tumors cluster by disease type\u003c/h2\u003e \u003cp\u003eWe added the 802 pediatric tumors from the CBTN dataset to our analysis to compare their gene expression patterns to both normal brain and adult glioma samples (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ea-b). We observed the formation of distinct subclusters for several tumor types that correlated with established molecular subgroups. As an example, we observed that medulloblastoma samples were split into three distinct clusters that correlated with known Medulloblastoma subtypes\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e (Wnt, Sonic hedgehog (SHH) and groups 3,4, \u003cb\u003eSupplementary Fig.\u0026nbsp;4a\u003c/b\u003e). Similarly, Ependymomas (EPN) samples formed several clusters that correlated with the anatomic tumor location (supratentorial (ST)-EPN, spinal-EPN, and posterior fossa (PF)-EPN) \u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e(Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ea, \u003cb\u003eSupplementary Fig.\u0026nbsp;4b\u003c/b\u003e). Pediatric pilocytic astrocytomas (PAs) and pediatric low-grade gliomas clustered closely together suggesting that they exert similar gene expression patterns. The schwannomas separated from the pediatric tumors and were located near the neurofibromas, which were localized adjacent to the malignant peripheral nerve sheath tumors (MPNST). The subependymal giant cell astrocytoma (SEGA) form a tight cluster as do the meningiomas. Interestingly, we observed that neurocytomas, DNET, ganglioglioma clustered near normal brain samples, specifically the hypothalamus and amygdala samples. Building a UMAP with just normal brain samples and the pediatric tumors, showed similar clustering profile (\u003cb\u003eSupplementary Fig.\u0026nbsp;4c\u003c/b\u003e). Taken together, these results suggest that pediatric tumors cluster by disease type and also form subclusters made by subtypes in the case of medulloblastomas and ependymomas.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eUsing the reference landscape to understand pathway regulation\u003c/h2\u003e \u003cp\u003eAlterations in signaling pathways are a hallmark of cancer and understanding the extent to which these pathways are dysregulated in tumor samples compared to healthy normal brain can help inform researchers about the underlying mechanisms of different cancer types.\u003c/p\u003e \u003cp\u003eBulk gene expression from adult glioma, pediatric brain tumors and healthy brain samples was subjected to a Gene Set variation Analysis (GSVA) and the gene set variation scores for each pathway were used to color in the Brain-UMAP. A score closer to 1 represented up-regulation of pathway in the given samples, whereas a score closer to -1 represented down-regulation of the pathway. We calculated GSVA scores for all pathways present in Reactome\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e, KEGG\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e and biocarta\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e pathways. We then tested whether there is a difference between the GSVA enrichment scores from different pair of phenotypes using a linear model and moderated t-statistic.\u003c/p\u003e \u003cp\u003eAs examples, we found that 605, 589, and 529 (\u003cb\u003eSupplementary Table\u0026nbsp;2a-c\u003c/b\u003e) pathways were up-regulated in IDH-wt glioblastomas, IDH-mut astrocytomas and oligodendrogliomas respectively compared to healthy brain samples from GTEx. A total of 456 pathways (\u003cb\u003eSupplementary Table\u0026nbsp;2d\u003c/b\u003e) were up regulated in all three adult glioma subtypes compared to healthy brain. The top pathways which were enriched in all adult glioma subtypes were pathways enriched for cell cycle, DNA repair, translation, splicing, oncogenic signaling pathways such as RAS pathway, Notch pathway, MHC pathway, PI3K/AKT Signaling, Wnt pathway, SHH pathway (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, \u003cb\u003eSupplementary Fig.\u0026nbsp;5a\u003c/b\u003e). Additionally, neurotransmitter pathways, calcium channel, potassium channel opening pathways were up regulated in healthy brain regions compared to pediatric tumors and adult gliomas (\u003cb\u003eSupplementary Fig.\u0026nbsp;5b\u003c/b\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eInterestingly, we noted that the two small clusters (IDH mutated grade 2 and grade3 oligodendroglioma and grade 4 IDH-wt glioblastomas) from the CGGA dataset were enriched in pathways related to olfaction, glucoronidation, ascorbate and aldarate metabolism and xenobiotics (\u003cb\u003eSupplementary Fig.\u0026nbsp;5c\u003c/b\u003e) in comparison to the main adult glioma cluster.\u003c/p\u003e \u003cp\u003eWhile visualizing pathways across the reference Brain-UMAP is extremely informative, researchers exploring targets for drug development may be also interested in investigating individual genes of a particular pathway. For example, Reactome\u0026rsquo;s mismatch repair pathway (R-HAS-5358508) and Biocarta RELA pathway (M10183) were both found to be up-regulated in all adult glioma subtypes compared to healthy normal brain, coloring in the gene expression for individual genes over the Brain-UMAP show different gene expression patterns across members of the same pathway. (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e, \u003cb\u003eSupplementary Fig.\u0026nbsp;6\u003c/b\u003e). For example, almost all the genes in the mismatch repair pathway have elevated gene expression levels in medulloblastoma, except RPA3 and POLD4.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003eStudying candidate genes at multiple genomic levels\u003c/h2\u003e \u003cp\u003eWhile bulk gene expression exhibited illuminating patterns in our data, analyzing other genomic information such as copy number alteration, gene fusions, or somatic alteration for each of these tumors can further enhance our understanding for a given gene of interest. Since processed copy number, gene fusions and somatic variants were publicly available only for two out of five datasets (TCGA and CBTN), we first built a much smaller UMAP using only the bulk gene expression data from these two datasets. The resulting UMAP (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003ea) showed a similar clustering pattern as our original Brain-UMAP. Next, we downloaded the copy number calls, gene fusion calls and somatic variants for adult gliomas from Genomic Data Commons (GDC) and the pediatric tumors from CBTN.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eOf note, the IDH-mut oligodendrogliomas, astrocytoma and pediatric high-grade gliomas have the highest mutational burden and number of gene fusions across all brain diseases (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003eb-c, \u003cb\u003eSupplementary Fig.\u0026nbsp;7a-b\u003c/b\u003e) in comparison to copy number profiles (Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003ed-e) which show a mixture of genes with amplified and lost copy number profiles across all brain diseases. Gene fusions are another class of potential oncogenic drivers in cancer, including pediatric cancers.\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e We observed that specific gene fusions and/or gene fusion partners were enriched in specific cancer subtypes. Adult IDH-wt glioblastoma frequently harbored gene fusions involving EGFR (17.7% of tumors, most commonly EGFR\u0026thinsp;\u0026minus;\u0026thinsp;PSPHP1, EGFR\u0026thinsp;\u0026minus;\u0026thinsp;LINC01445, EGFR\u0026thinsp;\u0026minus;\u0026thinsp;SEC61G\u0026thinsp;\u0026minus;\u0026thinsp;DT), whereas pediatric lower-grade gliomas frequently harbored BRAF fusions (KIAA1549\u0026thinsp;\u0026minus;\u0026thinsp;BRAF in 32% of low-grade gliomas and in 60% of pilocytic astrocytomas). In addition, supratentorial ependymomas most commonly harbored the C11orf95\u0026thinsp;\u0026minus;\u0026thinsp;RELA fusion (71% of tumors, also known as ZFTA-RELA) and meningiomas most frequently harbored fusions in NF2 and YAP1 (both 15% of tumors) \u003cb\u003e(Supplementary Fig.\u0026nbsp;7c).\u003c/b\u003e\u003c/p\u003e \u003cp\u003eArmed with additional genomic information such as gene fusions, copy number variation and somatic alternation, we investigated members of the Reactome mismatch repair pathway (R-HAS-5358508) (\u003cb\u003eSupplementary Fig.\u0026nbsp;8a-c\u003c/b\u003e). We observe that different genes belonging to this pathway exhibit different trends, for example genes such as POLD1(chr19), RPA2(chr1), and LIG1(chr19) loose a copy in IDH-mut oligodendrogliomas, while genes RPA3(chr19), PMS2(chr19), PCNA(chr20) and POLD2(chr7) gain a copy in IDH-wt GBM. While most genes show elevated gene expression levels for all brain tumors, of interest is EXO1 and RPA3 which are only up-regulated in IDH-wt GBM samples. While members of this pathway do not form gene fusions, they get mutated in different brain tumors. For example, in pediatric high grade tumors, we observed mutations in all the members of the mismatch repair pathway - MSH2 (22%), MSH6 (15.71%), POLD3 (14.29%), MSH3 (12.86%), LIG1(11.43%), EXO1(10%), PMS2(10%), PCNA(10%), POLD1 (8.57%), RPA2(8.57%), MLH1 (7.14%), RPA1(7.14%), POLD2(5.71%), POLD4(5.71%) and RPA3(5.71%).\u003c/p\u003e \u003cp\u003eOncogenes show altered gene expression in tumor samples, leading to abnormal phenotype in samples. Understanding gene expression patterns across various cancers of the nervous system can further our understanding of the disease. When studying known oncogenes such as EGFR, PTEN and CIC at a gene level (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e), as expected, we observe high number of mutations, high number of gene fusions, amplified gene expression values and copy number gains for EGFR across IDH-wt GBM. This contrasts with PTEN which shows loss of 1 copy in IDH-wt GBM samples. CIC, a transcriptional repressor, shows high number of mutations and copy number loss in oligodendrogliomas. For pediatric tumors, we observe BRAF gene fusions (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e) in 63% pilocytic astrocytoma tumors and 34% of low grade pediatric tumors (\u003cb\u003eSupplementary Table\u0026nbsp;3a\u003c/b\u003e) ALK (Fig.\u0026nbsp;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e) mutations are also observed in 38% high-grade pediatric tumors, 33% spinal cord ependymomas, 22% ATRT and 21% of the medulloblastoma tumors (\u003cb\u003eSupplementary Table\u0026nbsp;3b\u003c/b\u003e).This reference landscape is a useful research tool for the scientific community, where researchers can explore existing data to increase their understanding of oncogenic pathways and individual genes that make up these pathways, potentially uncovering candidates for novel therapeutic targets. By providing access to this reference landscape via an open source website like Oncoscape, we provide an interactive tool to researchers doing CNS research.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eAs costs for performing RNASeq continues to decline with increasing technological advances, more and more tumors will be sequenced, and additional tumor banks will be created with the underlying goal to understanding cancer\u0026rsquo;s complexity. The wealth of knowledge that that already exists in publicly available datasets such as those described here (GTEx, TCGA, CGGA, CBTN) is remarkable. Individually these datasets comprise of well-defined biologically similar set of patient samples and allow the analysis in exquisite resolution of genetic changes from one sample to the next. For example, comparing medulloblastomas to other medulloblastomas allows for precise characterization of molecular subtypes. As researchers, we can get so focused on comparing like with like that we lose sight of the proverbial forest for focusing too much on the leaves of a single tree. By integrating multiple datasets while correcting for batch effects, such as with the Brain-UMAP presented here, we can harness the power of multiple datasets.\u003c/p\u003e \u003cp\u003eThis landscape can be used to prospectively compare patients, similar to work that has been done using methylation arrays. Unlike methylation arrays, however, RNA seq is based on gene expression and therefore each cluster represented on the landscape contains granular information about the underlying tumor biology of the samples it contains. Because the overall expression pattern is identifiable for each tumor, this allows for cross comparison of various kinds of cancer to allow for characterization of tumor types not previously known. For example, asking questions about expression of particular genes or pathways across a wide panel of samples and tumor cohorts may uncover previously unknown roles or similarities between adult and pediatric tumor subtypes, which can possibly open new avenues for therapeutic discovery.\u003c/p\u003e \u003cp\u003eThe landscape that we present reveals correlations between tumor type and outcome. We have learned how diagnosis of a given tumor type can be confirmed using various pieces of genomic information. For example \u0026ndash; the adult glioma subtype of oligodendroglioma can be confirmed by presence of co-deletion of chr1p/19q and somatic alteration in IDH1.\u003c/p\u003e \u003cp\u003eWe are also able to compare tumor samples which were sequenced years apart (TCGA and CGGA) in two different continents and confirm that these differences did not contribute to expression changes in majority of the tumor samples. We are also able to identify novel entities and gain insight into genes that drive their unique character, such as in the case of the two adult glioma subtypes seen only in the CCGA data that appear to strongly and uniquely express genes involved in olfaction.\u003c/p\u003e \u003cp\u003eUsing this study as an example, we have seen how one diagnosis can be comprised of more than one expression entity, such as in the case of medulloblastoma and ependymoma. The UMAP also indicates room for better classification of tumors. For example, we found tumors which were documented as one type in the publicly available database, but in our UMAP were found to cluster with tumors of a different type. For example, some embryonal tumors ended up clustering with the medulloblastoma samples indicating that they may be medulloblastomas, or at least share many common features with medulloblastoma. We also have seen that multiple diagnoses may really be one entity as in the case of tumors diagnosed as either pilocytic astrocytoma (PA) or pediatric low-grade glioma.\u003c/p\u003e \u003cp\u003eReference landscapes like the Brain-UMAP, will be informative for researchers who wish to obtain a quick diagnosis or characterization of newly obtained tumor samples. For example, if we look at the medulloblastoma subtypes from CBTN, there are 14 group 3, 49 group 4, 30 SHH, 9 WNT and 19 unclassified medulloblastoma samples. For the 19 unclassified samples, we can predict which subtype they belong to, based on which medulloblastoma subtype samples they cluster with. (Supp Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ea).\u003c/p\u003e \u003cp\u003eBy combining results from both RNA-seq data (gene expression, gene fusions) as well as whole genome sequencing (copy number, mutation calls) and with the help of a few examples, we illustrate the knowledge that can be mined from this resource, at both a pathway level as well as a gene level. The work presented here demonstrates the utility of reference landscapes for combining genomic data across multiple tumor types for both diagnosis, prognosis and better understanding the biology of the tumors that are similar to a given patient collected prospectively. The power of visualizing gene expression changes, regulation of pathways, chromosomal alteration, gene fusions across multiple tumor types is informative to every researcher, especially those who may not have immediate access to computational biology experts. The reference landscape described here provides a useful tool for researchers interested in gene level questions across large scale patient data, while the methods used for integrating data sources highlight the tremendous potential for combining future datasets with existing resources to address complex biological questions.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eAll analyses were performed in R (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.r-project.org/\u003c/span\u003e\u003cspan address=\"https://www.r-project.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) using Bioconductor (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.bioconductor.org/\u003c/span\u003e\u003cspan address=\"https://www.bioconductor.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) packages. We have deposited all scripts, associated data, at: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://zenodo.org/badge/latestdoi/584982012\u003c/span\u003e\u003cspan address=\"https://zenodo.org/badge/latestdoi/584982012\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e to maximize transparency and reproducibility. TCGA data from GDC was downloaded using R/Bioconductor package TCGAbiolinks\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e,\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e, SummarizedExperiment\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e was used to store adult glioma and pediatric tumor data. All plots were made using ggplot2\u003csup\u003e21\u003c/sup\u003e, RcolorBrewer\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eObtaining gene expression RNASeq Data\u003c/h2\u003e \u003cp\u003eRNA Seq gene expression was downloaded from two sources for adult gliomas, GTEx defined healthy brain samples and pediatric tumors respectively (Supplemental Table\u0026nbsp;1).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eConversion of abundance estimates to transcripts per million (TPM)\u003c/h2\u003e \u003cp\u003eFor consistency, we converted all FPKM gene expression data to TPM data using the formula\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$TPM=\\frac{\\text{F}\\text{P}\\text{K}\\text{M}}{\\sum _{all genes}\\text{F}\\text{P}\\text{K}\\text{M}} \\times {10}^{6}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eas described by Collins et al\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eUniform Manifold Approximation and Project (UMAP)\u003c/h2\u003e \u003cp\u003eUniform Manifold Approximation and Project (UMAP) values were generated using the R function umap() using all protein-coding genes and visualized with R package ggplot2\u003csup\u003e21\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eObtaining Copy Number\u003c/h2\u003e \u003cp\u003eGistic2 thresholded gene level copy number variation estimated using the GSITIC2 method were downloaded from UCSC Xena(\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://xenabrowser.net/datapages/?cohort=TCGA%20Lower%20Grade%20Glioma%20(LGG)\u0026amp;removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443\u003c/span\u003e\u003cspan address=\"https://xenabrowser.net/datapages/?cohort=TCGA%20Lower%20Grade%20Glioma%20(LGG)\u0026amp;removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://xenabrowser.net/datapages/?cohort=TCGA%20Glioblastoma%20(GBM)\u0026amp;removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443\u003c/span\u003e\u003cspan address=\"https://xenabrowser.net/datapages/?cohort=TCGA%20Glioblastoma%20(GBM)\u0026amp;removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e ) for adult gliomas. Copy Number calls for pediatric tumors estimated from the GISTIC2 pipeline were downloaded from \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/AlexsLemonade/OpenPBTA-analysis\u003c/span\u003e\u003cspan address=\"https://github.com/AlexsLemonade/OpenPBTA-analysis\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eObtaining Gene Fusions\u003c/h2\u003e \u003cp\u003eGene fusions for adult gliomas estimated using Arriba and Star-Fusion as per the GDC RNA-Seq pipeline were downloaded from the GDC portal ( \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://portal.gdc.cancer.gov/projects/TCGA-GBM\u003c/span\u003e\u003cspan address=\"https://portal.gdc.cancer.gov/projects/TCGA-GBM\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://portal.gdc.cancer.gov/projects/TCGA-LGG\u003c/span\u003e\u003cspan address=\"https://portal.gdc.cancer.gov/projects/TCGA-LGG\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e ). Gene fusions for pediatric tumors were downloaded from \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/AlexsLemonade/OpenPBTA-analysis\u003c/span\u003e\u003cspan address=\"https://github.com/AlexsLemonade/OpenPBTA-analysis\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Only high-confidence gene fusion calls were retained from ARRIBA.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eObtaining Mutations\u003c/h2\u003e \u003cp\u003eAnnotated Variant call Format (VCF) files containing somatic variants from MuSE, VarScan2, MuTect2 and Somatic Sniper were downloaded from GDC portal for TCGA-GBM and TCGA-LGG. The variants for each patient were combined based on custom script present in our github repository. Somatic Variant calls for pediatric tumors was obtained from \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/AlexsLemonade/OpenPBTA-analysis\u003c/span\u003e\u003cspan address=\"https://github.com/AlexsLemonade/OpenPBTA-analysis\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eGSVA Pathway Analysis\u003c/h2\u003e \u003cp\u003eGene sets for all the pathways from Biocarta, Kyoto Encylcopedia of Genes and Genomes (KEGG) and Reactome pathways were downloaded from Molecular Signature Databases (MSigDB)(v7.2)\u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.gsea-msigdb.org/gsea/msigdb/collections.jsp#C2\u003c/span\u003e\u003cspan address=\"https://www.gsea-msigdb.org/gsea/msigdb/collections.jsp#C2\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e Batch corrected log2(TPM) counts from each pipeline were used to conduct a Gene Set Variation Analysis (GSVA)\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e using Biocarta, KEGG and Reactome pathways. GSVA scores obtained from 1 and \u0026minus;\u0026thinsp;1 for each sample, were visualized using ggplot2.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eKaplan-Meier curves for TCGA and CGGA patients\u003c/h2\u003e \u003cp\u003eTo generate Kaplan-Meier curves for TCGA and CGGA patients, all samples labeled recurrent, secondary, or normal tissue were removed. Duplicate samples were also excluded so that each patient was represented by exactly one sample. Kaplan-Meier curves were drawn by the Python package \u003cem\u003elifelines\u003c/em\u003e (Python version 3.8.6, \u003cem\u003elifelines\u003c/em\u003e version 0.25.2)\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e using survival data and glioma subtype labels downloaded from GDC, where the world health organization (WHO) 2016 criteria for the classification of adult diffuse gliomas was used to determine glioma subtype\u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eSurvival-Annotated UMAP\u003c/h2\u003e \u003cp\u003eTo annotate the Brain-UMAP with survival data, each TCGA and CGGA sample was colored with the median survival of a cohort of patients (nearest neighbors) close to the sample in question on the UMAP landscape. The notion of nearest neighbors was defined as a number of samples within a radius of 2 from the sample in question under the constraint that all nearest neighbors must be of the same glioma subtype and from the same dataset (TCGA or CGGA) as the sample in question. The radius parameter was chosen qualitatively. The number of nearest neighbors was defined as 25% of the total number of samples of the same glioma subtype and dataset as the sample in question. If the median survival of a sample\u0026rsquo;s cohort of nearest neighbors is undefined, or if a sample has fewer than 10 nearest neighbors, the point is not colored in. Non-primary and duplicate samples were excluded in the same manner as was done for the Kaplan-Meier curve analysis.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eOncoscape integration\u003c/h2\u003e \u003cp\u003eGene expression and survival data analysis result files were prepared according to the oncoscape instructions here: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/FredHutch/OncoscapeV3/blob/master/docs/upload.md\u003c/span\u003e\u003cspan address=\"https://github.com/FredHutch/OncoscapeV3/blob/master/docs/upload.md\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003eData Analysis\u003c/h2\u003e \u003cp\u003eAll statistical analyses and plots were done in R (v.3.3.1) as implemented in Rstudio (v.1.0.136). Plots were created using the R basic graphics. The following R packages were used: GenomicAlignments53 (v.1.24), reshape64 (v.0.8.8), png65 (v.0.1\u0026ndash;7), ape66 (v.5.3) and seqinr67 (v.3.6\u0026ndash;1).\u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003e Acknowledgements\u003c/strong\u003e \u003c/p\u003e\n\u003cp\u003eWe thank members of the Holland Lab at Fred Hutchinson Cancer Research Center for discussions. This research was supported by the R35 CA253119-01A1(E.C.H.), NIH U54 CA193461 (E.C.H.), National Institutes of Health R01 CA195718 (E.C.H.), R01 CA100688 (E.C.H.), T32 CA9657-25 (S.S.P.), U54 DK106829 (S.S.P.), R21 CA223531 (S.S.P.); K22 CA258953-01 (SSP) ,Jacobs Foundation Research Fellowship (S.S.P.) and National Science Foundation Graduate Research Fellowship Program DGE-1762114 (N.N.).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Information \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eContributions\u003c/p\u003e\n\u003cp\u003eS.A., S.S.P., F.S., and E.C.H. conceived and organized the study based on S.A.\u0026rsquo;s proposed design. S.A. and N.N. performed all the analysis. M.J. incorporated the data into oncoscape. S.A. and S.S.P. wrote the manuscript, with contributions from all authors. \u003c/p\u003e\n\u003cp\u003eCorresponding Author \u003c/p\u003e\n\u003cp\u003eCorrespondence to Eric C Holland\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors have stated explicitly that there are no conflicts of interest in connection with this article.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll data are available in the main text, supplementary materials and/or accompanying database is found at https://zenodo.org/badge/latestdoi/584982012. We used data from recount2 (https://jhubiostatistics.shinyapps.io/recount/) , CBTN (https://github.com/AlexsLemonade/OpenPBTA-analysis) , copy number data for TCGA-GBM ( https://xenabrowser.net/datapages/?cohort=TCGA%20Glioblastoma%20(GBM)\u0026amp;removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443 ) and TCGA-LGG(https://xenabrowser.net/datapages/?cohort=TCGA%20Lower%20Grade%20Glioma%20(LGG)\u0026amp;removeHub=https%3A%2F%2Fxena.treehouse.gi.ucsc.edu%3A443 ) , Gene fusions from GDC Data portal (https://portal.gdc.cancer.gov/projects/TCGA-GBM , https://portal.gdc.cancer.gov/projects/TCGA-LGG )\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll custom code used in this study is available at https://zenodo.org/badge/latestdoi/584982012\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eCancer Genome Atlas Research, N.\u003cem\u003e, et al.\u003c/em\u003e The Cancer Genome Atlas Pan-Cancer analysis project. \u003cem\u003eNat Genet\u003c/em\u003e \u003cstrong\u003e45\u003c/strong\u003e, 1113-1120 (2013).\u003c/li\u003e\n\u003cli\u003eZhao, Z.\u003cem\u003e, et al.\u003c/em\u003e Chinese Glioma Genome Atlas (CGGA): A Comprehensive Resource with Functional Genomic Data from Chinese Glioma Patients. \u003cem\u003eGenomics Proteomics Bioinformatics\u003c/em\u003e \u003cstrong\u003e19\u003c/strong\u003e, 1-12 (2021).\u003c/li\u003e\n\u003cli\u003eIjaz, H.\u003cem\u003e, et al.\u003c/em\u003e Pediatric high-grade glioma resources from the Children\u0026apos;s Brain Tumor Tissue Consortium. \u003cem\u003eNeuro Oncol\u003c/em\u003e \u003cstrong\u003e22\u003c/strong\u003e, 163-165 (2020).\u003c/li\u003e\n\u003cli\u003eCarithers, L.J.\u003cem\u003e, et al.\u003c/em\u003e A Novel Approach to High-Quality Postmortem Tissue Procurement: The GTEx Project. \u003cem\u003eBiopreserv Biobank\u003c/em\u003e \u003cstrong\u003e13\u003c/strong\u003e, 311-319 (2015).\u003c/li\u003e\n\u003cli\u003eMcFerrin, L.G.\u003cem\u003e, et al.\u003c/em\u003e Analysis and visualization of linked molecular and clinical cancer data by using Oncoscape. \u003cem\u003eNat Genet\u003c/em\u003e \u003cstrong\u003e50\u003c/strong\u003e, 1203-1204 (2018).\u003c/li\u003e\n\u003cli\u003eShapiro, J.A.\u003cem\u003e, et al.\u003c/em\u003e OpenPBTA: An Open Pediatric Brain Tumor Atlas. \u003cem\u003ebioRxiv\u003c/em\u003e (2022).\u003c/li\u003e\n\u003cli\u003eCollado-Torres, L.\u003cem\u003e, et al.\u003c/em\u003e Reproducible RNA-seq analysis using recount2. \u003cem\u003eNat Biotechnol\u003c/em\u003e \u003cstrong\u003e35\u003c/strong\u003e, 319-321 (2017).\u003c/li\u003e\n\u003cli\u003eSubramanian S, A.T. Childhood Brain Tumors. \u003cem\u003eIn: StatPearls [Internet]. Treasure Island (FL): StatPearls Publishing\u003c/em\u003e, https://www.ncbi.nlm.nih.gov/books/NBK535415/ (2022).\u003c/li\u003e\n\u003cli\u003eArora, S., Pattwell, S.S., Holland, E.C. \u0026amp; Bolouri, H. Variability in estimated gene expression among commonly used RNA-seq pipelines. \u003cem\u003eSci Rep\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 2734 (2020).\u003c/li\u003e\n\u003cli\u003eJohnson, W.E., Li, C. \u0026amp; Rabinovic, A. Adjusting batch effects in microarray expression data using empirical Bayes methods. \u003cem\u003eBiostatistics\u003c/em\u003e \u003cstrong\u003e8\u003c/strong\u003e, 118-127 (2007).\u003c/li\u003e\n\u003cli\u003ePattwell, S.S.\u003cem\u003e, et al.\u003c/em\u003e A kinase-deficient NTRK2 splice variant predominates in glioma and amplifies several oncogenic signaling pathways. \u003cem\u003eNat Commun\u003c/em\u003e \u003cstrong\u003e11\u003c/strong\u003e, 2977 (2020).\u003c/li\u003e\n\u003cli\u003eBolouri, H., Zhao, L.P. \u0026amp; Holland, E.C. Big data visualization identifies the multidimensional molecular landscape of human gliomas. \u003cem\u003eProc Natl Acad Sci U S A\u003c/em\u003e \u003cstrong\u003e113\u003c/strong\u003e, 5394-5399 (2016).\u003c/li\u003e\n\u003cli\u003ePollack, I.F., Agnihotri, S. \u0026amp; Broniscer, A. Childhood brain tumors: current management, biological insights, and future directions. \u003cem\u003eJ Neurosurg Pediatr\u003c/em\u003e \u003cstrong\u003e23\u003c/strong\u003e, 261-273 (2019).\u003c/li\u003e\n\u003cli\u003eGillespie, M.\u003cem\u003e, et al.\u003c/em\u003e The reactome pathway knowledgebase 2022. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e50\u003c/strong\u003e, D687-d692 (2022).\u003c/li\u003e\n\u003cli\u003eKanehisa, M. \u0026amp; Goto, S. KEGG: kyoto encyclopedia of genes and genomes. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e28\u003c/strong\u003e, 27-30 (2000).\u003c/li\u003e\n\u003cli\u003eBioCarta. \u003cem\u003eBiotech Software \u0026amp; Internet Report\u003c/em\u003e \u003cstrong\u003e2\u003c/strong\u003e, 117-120 (2001).\u003c/li\u003e\n\u003cli\u003eSzulzewsky, F.\u003cem\u003e, et al.\u003c/em\u003e Both YAP1-MAML2 and constitutively active YAP1 drive the formation of tumors that resemble NF2 mutant meningiomas in mice. \u003cem\u003eGenes Dev\u003c/em\u003e \u003cstrong\u003e36\u003c/strong\u003e, 857-870 (2022).\u003c/li\u003e\n\u003cli\u003eColaprico, A.\u003cem\u003e, et al.\u003c/em\u003e TCGAbiolinks: an R/Bioconductor package for integrative analysis of TCGA data. \u003cem\u003eNucleic Acids Res\u003c/em\u003e \u003cstrong\u003e44\u003c/strong\u003e, e71 (2016).\u003c/li\u003e\n\u003cli\u003eSilva, T.C.\u003cem\u003e, et al.\u003c/em\u003e TCGA Workflow: Analyze cancer genomics and epigenomics data using Bioconductor packages. \u003cem\u003eF1000Res\u003c/em\u003e \u003cstrong\u003e5\u003c/strong\u003e, 1542 (2016).\u003c/li\u003e\n\u003cli\u003eMorgan M, O.V., Hester J, Pag\u0026egrave;s H. SummarizedExperiment: SummarizedExperiment container. R package version 1.16.0. (2019).\u003c/li\u003e\n\u003cli\u003eWickham, H. \u003cem\u003eggplot2: Elegant Graphics for Data Analysis\u003c/em\u003e, (Springer-Verlag New York, 2016).\u003c/li\u003e\n\u003cli\u003eNeuwirth, E. Package \u0026lsquo;RColorBrewer\u0026rsquo; , ColorBrewer Palettes. (2014).\u003c/li\u003e\n\u003cli\u003eBo Li, C.N.D. RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. \u003cem\u003eBMC Bioinformatics\u003c/em\u003e (2011).\u003c/li\u003e\n\u003cli\u003eAravind Subramanian, P.T., Vamsi K. Mootha, Sayan Mukherjee, Benjamin L. Ebert, Michael A. Gillette, Amanda Paulovich, Scott L. Pomeroy, Todd R. Golub, Eric S. Lander, Jill P. Mesirov. Gene set enrichment analysis: A knowledge-based approach for interpreting genome-wide expression profiles. \u003cem\u003ePNAS\u003c/em\u003e (2005).\u003c/li\u003e\n\u003cli\u003eSonja H\u0026auml;nzelmann, R.C.J.G. GSVA: gene set variation analysis for microarray and RNA-Seq data. \u003cem\u003eBMC Bioinformatics \u003c/em\u003e(2013).\u003c/li\u003e\n\u003cli\u003eDavidson-Pilon, C. lifelines: survival analysis in Python. \u003cem\u003eJournal of Open Source Software\u003c/em\u003e \u003cstrong\u003e4\u003c/strong\u003e(2019).\u003c/li\u003e\n\u003cli\u003eLouis, D.N.\u003cem\u003e, et al.\u003c/em\u003e The 2016 World Health Organization Classification of Tumors of the Central Nervous System: a summary. \u003cem\u003eActa Neuropathol\u003c/em\u003e \u003cstrong\u003e131\u003c/strong\u003e, 803-820 (2016).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-2448083/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2448083/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eIn order to better understand the relationship between normal and neoplastic brain, we combined five publicly available large-scale datasets, correcting for batch effects and applying Uniform Manifold Approximation and Projection (UMAP) to RNA-seq data. We assembled a reference Brain-UMAP including 702 adult gliomas, 802 pediatric tumors and 1409 healthy normal brain samples, which can be utilized to investigate the wealth of information obtained from combining several publicly available datasets to study a single organ site. Normal brain regions and tumor types create distinct clusters and because the landscape is generated by RNA seq, comparative gene expression profiles and gene ontology patterns are readily evident. To our knowledge, this is the first meta-analysis that allows for comparison of gene expression and pathways of interest across adult gliomas, pediatric brain tumors, and normal brain regions. We provide access to this resource via the open source, interactive online tool Oncoscape, where the scientific community can readily visualize clinical metadata, gene expression patterns, gene fusions, mutations, and copy number patterns for individual genes and pathway over this reference landscape.\u003c/p\u003e","manuscriptTitle":"An RNA seq-based reference landscape of human normal and neoplastic brain","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-01-10 15:55:45","doi":"10.21203/rs.3.rs-2448083/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2023-02-14T06:03:08+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-01-27T17:36:29+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"7cb35d82-9723-47c3-ae10-78c4792e83a4_SNPRID","date":"2023-01-27T16:11:43+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-01-24T17:29:03+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-01-24T17:00:34+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2023-01-06T18:36:52+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2023-01-06T15:32:45+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2023-01-05T21:11:53+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"140ad354-acb2-4e7a-b72d-a0b649cba598","owner":[],"postedDate":"January 10th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":18211687,"name":"Biological sciences/Cancer/Cancer genomics"},{"id":18211688,"name":"Biological sciences/Cancer/Cns cancer"},{"id":18211689,"name":"Biological sciences/Computational biology and bioinformatics"}],"tags":[],"updatedAt":"2023-10-16T20:05:43+00:00","versionOfRecord":{"articleIdentity":"rs-2448083","link":"https://doi.org/10.1038/s41598-023-31180-z","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2023-03-14 20:02:43","publishedOnDateReadable":"March 14th, 2023"},"versionCreatedAt":"2023-01-10 15:55:45","video":"","vorDoi":"10.1038/s41598-023-31180-z","vorDoiUrl":"https://doi.org/10.1038/s41598-023-31180-z","workflowStages":[]},"version":"v1","identity":"rs-2448083","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2448083","identity":"rs-2448083","version":["v1"]},"buildId":"rHA-KDH7Qsr4HCuvH75dn","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.