{"paper_id":"357d1922-0e46-44b2-a608-6af91cd4bd8b","body_text":"The Open Pediatric Cancer Project \n This manuscript (permalink) was automatically generated from rokitalab/OpenPedCan-\nmanuscript@a19d489 on 2025-06-26.  \nAuthors \n• Zhuangzhuang Geng  0009-0007-6883-0691 ·  zzgeng   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA  \n Eric Wafula  0000-0001-8073-3797 ·  ewafula   Department of Biomedical and \nHealth Informatics, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Ryan J. Corbett  0000-0002-3478-0784 ·  rjcorb   Center for Cancer and \nImmunology Research, Children’s National Hospital, Washington, DC, 20010, USA; \nCenter for Data-Driven Discovery in Biomedicine, Children’s Hospital of Philadelphia, \nPhiladelphia, PA, 19104, USA; Division of Neurosurgery, Children’s Hospital of \nPhiladelphia, Philadelphia, PA, 19104, USA  \n Yuanchao Zhang  Department of Biomedical and Health Informatics, Children’s \nHospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Run Jin  0000-0002-8958-9266 ·  runjin326   Center for Data-Driven Discovery in \nBiomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA; Division \nof Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Krutika S. Gaonkar  0000-0003-0838-2405  Center for Data-Driven Discovery in \nBiomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA; Division \nof Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA; \nDepartment of Biomedical and Health Informatics, Children’s Hospital of Philadelphia, \nPhiladelphia, PA, 19104, USA  \n Sangeeta Shukla  0000-0002-3727-9602\n ·  sangeetashukla   Department of \nBiomedical and Health Informatics, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA  \n Komal S. Rathi  0000-0001-5534-6904 ·  komalsrathi   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Department of Biomedical and Health Informatics, Children’s Hospital of \nPhiladelphia, Philadelphia, PA, 19104, USA  \nn \nn \nA, \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\n Dave Hill  0000-0002-1337-1789 ·  atlas4213   Department of Biomedical and \nHealth Informatics, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Aditya Lahiri  0000-0001-9352-1312 ·  adilahiri   Department of Biomedical and \nHealth Informatics, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Daniel P. Miller  0000-0002-2032-4358 ·  dmiller15   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA  \n Alex Sickler  0000-0001-7830-7537 ·  sickler-alex   Center for Cancer and \nImmunology Research, Children’s National Hospital, Washington, DC, 20010, USA; \nCenter for Data-Driven Discovery in Biomedicine, Children’s Hospital of Philadelphia, \nPhiladelphia, PA, 19104, USA; Division of Neurosurgery, Children’s Hospital of \nPhiladelphia, Philadelphia, PA, 19104, USA  \n Kelsey Keith  0000-0002-7451-5117 ·  kelseykeith   Department of Biomedical \nand Health Informatics, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA  \n Christopher Blackden  0000-0002-4928-090X\n ·  devbyaccident   Center for Data-\nDriven Discovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, \nPA, 19104, USA  \n Antonia Chroni ·  AntoniaChroni   Center for Data-Driven Discovery in Biomedicine, \nChildren’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA; Division of \nNeurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Miguel A. Brown  0000-0001-6782-1442\n ·  migbro   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA  \n Adam A. Kraya  0000-0002-8526-5694 ·  aadamk   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA  \n Kaylyn L. Clark  0000-0002-7207-6608  Center for Cancer and Immunology \nResearch, Children’s National Hospital, Washington, DC, 20010, USA  \na, \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\n Brian R. Rood  Center for Cancer and Immunology Research, Children’s National \nHospital, Washington, DC, 20010, USA; George Washington University School of \nMedicine and Health Sciences, Washington, D.C., 20052, USA  \n Adam C. Resnick  0000-0003-0436-4189 ·  adamcresnick   Center for Data-\nDriven Discovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, \nPA, 19104, USA · Funded by Children’s Brain Tumor Network; NIH 3P30 CA016520-\n44S5, U2C HL138346-03, U24 CA220457-03; NCI/NIH Contract No. 75N91019D00024, \nTask Order No. 75N91020F00003; Children’s Hospital of Philadelphia Division of \nNeurosurgery  \n Nicholas Van Kuren  0000-0002-7414-9516 ·  nicholasvk   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA  \n John M. Maris  Division of Oncology, Children’s Hospital of Philadelphia, Philadelphia, \nPA, 19104, USA; Department of Pediatrics, University of Pennsylvania, Philadelphia, \nPA, 19104, USA  \n Alvin Farrel  0000-0003-1087-9840\n ·  afarrel   Department of Biomedical and \nHealth Informatics, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA; \nDivision of Oncology, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA; \nCenter for Childhood Cancer Research, Children’s Hospital of Philadelphia, \nPhiladelphia, PA, 19104, USA · Funded by NCI/NIH Contract No. 75N91019D00024, \nTask Order No. 75N91020F00003  \n Mateusz P. Koptyra  0000-0002-3857-6633 ·  mkoptyra   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA  \n Gerri R. Trooskin  Center for Data-Driven Discovery in Biomedicine, Children’s Hospital \nof Philadelphia, Philadelphia, PA, 19104, USA; Division of Neurosurgery, Children’s \nHospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Noel Coleman  0000-0001-6454-1285\n  Center for Data-Driven Discovery in \nBiomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA; Division \nof Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Yuankun Zhu  0000-0002-2455-9525 ·  yuankunzhu   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA  \na, \n4, \nn \n; \ntal \nn \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\n Stephanie Stefankiewicz  Center for Data-Driven Discovery in Biomedicine, Children’s \nHospital of Philadelphia, Philadelphia, PA, 19104, USA; Division of Neurosurgery, \nChildren’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Zied Abdullaev  Laboratory of Pathology, National Cancer Institute, Bethesda, MD, \n20892, USA  \n Asif T Chinwalla  0000-0001-7831-3996 ·  chinwallaa   Department of Biomedical \nand Health Informatics, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA  \n Mariarita Santi  0000-0002-6728-3450  Department of Pathology and Laboratory \nMedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA; \nDepartment of Pathology and Laboratory Medicine, University of Pennsylvania Perelman \nSchool of Medicine, Philadelphia, PA, 19104, USA  \n Ammar S. Naqvi ·  naqvia   Center for Data-Driven Discovery in Biomedicine, \nChildren’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA; Division of \nNeurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Jennifer L. Mason  Center for Data-Driven Discovery in Biomedicine, Children’s \nHospital of Philadelphia, Philadelphia, PA, 19104, USA; Division of Neurosurgery, \nChildren’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Carl J. Koschmann  0000-0002-0825-7615  Department of Pediatrics, University of \nMichigan Health, Ann Arbor, MI, 48105, USA; Pediatric Hematology Oncology, Mott \nChildren’s Hospital, Ann Arbor, MI, 48109, USA  \n Xiaoyan Huang  0000-0001-7267-4512\n ·  HuangXiaoyan0106   Center for Data-\nDriven Discovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, \nPA, 19104, USA  \n Sharon J. Diskin  0000-0002-7200-8939  Division of Oncology, Children’s Hospital \nof Philadelphia, Philadelphia, PA, 19104, USA; Department of Pediatrics, University of \nPennsylvania, Philadelphia, PA, 19104, USA  \n Kenneth Aldape  Laboratory of Pathology, National Cancer Institute, Bethesda, MD, \n20892, USA  \n Bailey K. Farrow  0000-0001-6727-6333 ·  baileyckelly   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA  \nal \nan \na, \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\n Weiping Ma  Department of Genetics and Genomic Sciences, Icahn School of Medicine \nat Mount Sinai, New York, NY 10029, USA; Tisch Cancer Institute, Icahn School of \nMedicine at Mount Sinai, New York, NY, 10029, USA  \n Bo Zhang  0000-0002-0743-5379 ·  zhangb1   Center for Data-Driven Discovery in \nBiomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA; Division \nof Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Brian M. Ennis  0000-0002-2653-5009 ·  bmennis   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA  \n Sarah Tasian  Division of Oncology, Children’s Hospital of Philadelphia, Philadelphia, \nPA, 19104, USA; Department of Pediatrics, University of Pennsylvania, Philadelphia, \nPA, 19104, USA  \n Saksham Phul  0000-0002-2771-2572\n ·  sakshamphul   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA  \n Matthew R. Lueder  0009-0002-7370-102X\n ·  luederm   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA; Department of Pathology and Laboratory Medicine, Children’s Hospital of \nPhiladelphia, Philadelphia, PA, 19104, USA  \n Chuwei Zhong  0000-0003-2406-2735 ·  zoomzoom1011   Center for Data-Driven \nDiscovery in Biomedicine, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, \nUSA; Division of Neurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, \n19104, USA  \n Joseph M. Dybas ·  JosephDybas   Center for Data-Driven Discovery in Biomedicine, \nChildren’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA; Division of \nNeurosurgery, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA  \n Pei Wang  \nDepartment of Genetics and Genomic Sciences, Icahn School of Medicine at \nMount Sinai, New York, NY 10029, USA; Tisch Cancer Institute, Icahn School of \nMedicine at Mount Sinai, New York, NY 10029, USA  \n Deanne Taylor  0000-0002-3302-4610\n ·  taylordm   Department of Biomedical and \nHealth Informatics, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA; \nDepartment of Pediatrics, University of Pennsylvania Perelman Medical School, \ne \n in \nn \nn \ne, \nat \nnd \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nPhiladelphia, PA, 19104, USA · Funded by NCI/NIH Contract No. 75N91019D00024, \nTask Order No. 75N91020F00003  \n Jo Lynne Rokita /i1  0000-0003-2171-3627 ·  jharenza   Center for Cancer and \nImmunology Research, Children’s National Hospital, Washington, DC, 20010, USA; \nCenter for Data-Driven Discovery in Biomedicine, Children’s Hospital of Philadelphia, \nPhiladelphia, PA, 19104, USA; Division of Neurosurgery, Children’s Hospital of \nPhiladelphia, Philadelphia, PA, 19104, USA; Department of Biomedical and Health \nInformatics, Children’s Hospital of Philadelphia, Philadelphia, PA, 19104, USA · Funded \nby NCI/NIH Contract No. 75N91019D00024, Task Order No. 75N91020F00003  \nContact information \n/i1Correspondence: Jo Lynne Rokita jrokita@childrensnational.org \nAbstract \nBackground: In 2019, the Open Pediatric Brain Tumor Atlas (OpenPBTA) was created as a \nglobal, collaborative open-science initiative to genomically characterize 1,074 pediatric brain \ntumors and 22 patient-derived cell lines. Here, we present an extension of the OpenPBTA called\nthe Open Pediatric Cancer (OpenPedCan) Project, a harmonized open-source multi-omic \ndataset from 6,112 pediatric cancer patients with 7,096 tumor events across more than 100 \nhistologies. Combined with RNA-Seq from the Genotype-Tissue Expression (GTEx) and The \nCancer Genome Atlas (TCGA), OpenPedCan contains nearly 48,000 total biospecimens \n(24,002 tumor and 23,893 normal specimens). \nFindings: We utilized Gabriella Miller Kids First (GMKF) workflows to harmonize WGS, WXS, \nRNA-seq, and Targeted Sequencing datasets to include somatic SNVs, InDels, CNVs, SVs, \nRNA expression, fusions, and splice variants. We integrated summarized CPTAC whole cell \nproteomics and phospho-proteomics data, miRNA-\nSeq data, and have developed a methylation \narray harmonization workflow to include m-values, beta-vales, and copy number calls. \nOpenPedCan contains reproducible, docker\nized workflows in GitHub, CAVATICA, and Amazon \nWeb Services (AWS) to deliver harmonized and processed data from over 60 scalable modules \nwhich can be leveraged both locally and on AWS. The processed data are released in a \nversioned manner and accessible th\nrough CAVATICA or AWS S3 download (from GitHub), and \nqueryable through PedcBioPortal and the NCI’s pediatric Molecular Targets Platform. Notably, \nwe have expanded PBTA molecular subtyping to include methylation information to align with \nthe WHO 2021 Central Nervous System Tumor classifications, allowing us to create research-\ngrade integrated diagnoses for these tumors. \nConclusions: OpenPedCan data and its reproducible analysis module framework are openly \navailable and can be utilized and/or adapted by researchers to accelerate discovery, validation, \nand clinical translation. \ned \nn \nn \ns \nd \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nKeywords \nPediatric cancer, open science, reproducibility, multi-omics, Docker, OpenPedCan \nData Description \nThe Open Pediatric Cancer (OpenPedCan) project is an iterative open analysis effort in which \nwe harmonize pediatric cancer data from multiple sources, perform downstream cancer \nanalyses on these data, and provide them through Amazon S3, CAVATICA, PedcBioPortal, and \nv2.1 of NCI’s Pediatric Molecular Targets Platform (MTP). We harmonized, aggregated, and \nanalyzed data from multiple pediatric and adult data sources, building upon the work of the \nOpenPBTA (Figure 1). All RNA-seq and DNA-seq data from OpenPBTA were updated from \nGENCODE v27 to GENCODE v39 as part of the OpenPedCan project. Further, all data within \nOpenPedCan is harmonized with GENCODE v39 annotations. Biospecimen-\nlevel metadata and \nclinical data are contained in Supplemental Table 1. \nFigure 1: OpenPedCan Data. A, OpenPedCan contains multi-omic data from seven cohorts of \npediatric tumors (A-B) with counts by tumor event, RNA-Seq from adult tumors from The Cancer \nGenome Atlas (TCGA) Program (C-D) and RNA-Seq from normal adult tissues from the \nGenotype-Tissue Expression (GTeX) project (E) with counts by specimen. (Abbreviations: \nTARGET = Therapeutically Applicable Research to Generate Effective Treatments , PPTC = \nPediatric Preclinical Testing Consortium, PBTA = Pediatric Brain Tumor Atlas, Maris = \nNeuroblastoma cell lines from the Maris Laboratory at CHOP, GMKF = Gabriella Miller Kids \nd \nd \n \ner \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nFirst, DGD = Division of Genomic Diagnostics at CHOP, CPTAC = Clinical Proteomic Tumor \nAnalysis Consortium) \nOpenPedCan currently include the following datasets, described more fully below: \n OpenPBTA \n TARGET \n Kids First Neuroblastoma (X01) \n Kids First PBTA (X01) \n Chordoma Foundation \n PPTC \n Maris \n MI-ONCOSEQ Study \n DGD \n GTEx \n TCGA \n CPTAC PBTA \n CPTAC GBM \n HOPE proteomics \nOpen Pediatric Brain Tumor Atlas (OpenPBTA) \nIn September of 2018, the Children’s Brain Tumor Network (CBTN)\n released the Pediatric Brain \nTumor Atlas (PBTA), a genomic dataset (whole genome sequencing, whole exome sequencing, \nRNA sequencing, proteomic, and clinical data) for nearly 1,000 tumors, available from the \nGabriella Miller Kids First Portal. In September of 2019, the Open Pediatric Brain Tumor Atlas \n(OpenPBTA) Project was launched. OpenPBTA was a global open science initiative to \ncomprehensively define the molecular landscape of tumors of 943 patients from the CBTN and \nthe PNOC003 DIPG clinical trial from the Pediatric Pacific Neuro-oncology Consortium through \nreal-time, collaborative analyses and collaborative manuscript writing on GitHub [1]. Additional \nPBTA data has been, and will be continually added to, OpenPedCan. \nTherapeutically Applicable Research to Generate Effective Treatments (TARGET) \nThe Therapeutically Applicable Research to Generate Effective Treatments (TARGET) Initiative \nis an NCI-funded collection of disease-specific projects that seeks to identify the genomic \nchanges of pediatric cancers. The overall goal is to collect genomic data to accelerate the \ndevelopment of more effective therapies. OpenPedCan analyses include newly harmonized, \nopen-access data associated with the seven diseases present in the TARGET dataset: Acute \nLymphoblastic Leukemia (ALL), Acute Myeloid Leukemia (AML), Clear cell sarcoma of the \nkidney, Neuroblastoma, Osteosarcoma, Rhabdoid tumor, and Wilm’s Tumor. \nGabriella Miller Kids First (Neuroblastoma)\n and PBTA \nThe Gabriella Miller Kids First Pediatric Research Program (Kids First) is a large-scale effort to \naccelerate research and gene discovery in pediatric cancers and structural birth defects. The \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nprogram includes whole genome sequencing (WGS) from patients with pediatric cancers and \nstructural birth defects and their families. OpenPedCan analyses include Neuroblastoma and \nPBTA data from the Kids First projects. \nChordoma Foundation \nThe Chordoma Foundation\n seeks to advance research and improve healthcare for patients \ndiagnosed with chordoma and has shared patient and model sequencing data with the CBTN. \nPediatric Preclinical Testing Consortium (PPTC)\n \nThe National Cancer Institute’s (NCI) former PPTC, now the Pediatric Preclinical in Vivo Testing \n(PIVOT) Program, molecularly and pharmacologically characterizes cell-derived and patient-\nderived xenograft (PDX) models. OpenPedCan includes re-harmonized RNA-Seq data for 244 \nmodels from the initial PPTC study [2]. A subset of PPTC includes neuroblastoma models; the \nMaris cohort includes re-harmonized RNA-Seq data for 39 neuroblastoma cell lines [3], some \nof which have corresponding PDX models within the PPTC. \nMI-ONCOSEQ Study [4] \nThese clinical sequencing data from the University of Michigan were donated to CBTN and \nadded to the PBTA cohort. \nDivision of Genomic Diagnostics at Children’s Hospital of Philadelphia (DGD) \nCHOP’s Division of Genomic Diagnostics has partnered with CCDI to add somatic panel \nsequencing data to OpenPedCan and the Molecular Targets Platform. \nThe Genotype-Tissue Expression Project (GTEx)\n \nThe GTEx project is an ongoing effort to build a comprehensive public data resource and tissue \nbank to study tissue-specific gene expression, regulation and their relationship with genetic \nvariants. Samples were collected from 54 non-diseased tissue sites across nearly 1000 \nindividuals, primarily for molecular assays including WGS, WXS, and RNA-Seq. OpenPedCan \nproject includes 17,382 GTEx RNA-Seq samples from GTEx v8 release, which span across 31 \nGTEx groups in the v12 release. \nThe Cancer Genome Atlas Program (TCGA)\n \nTCGA is a landmark cancer genomics program that molecularly characterized over 20,000 \nprimary cancer and matched normal samples spanning 33 cancer types. It is a joint effort \nbetween NCI and the National Human Genome Research Institute. OpenPedCan project \nincludes open-access 10,414 RNA-Seq for 716 normal and 9,698 TCGA tumor samples from 33 \ncancer types. \nClinical Proteomic Tumor Analysis Consortium (CPTAC) PBTA proteomics study \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nThe CPTAC pediatric pan-brain tumor study [5] contains 218 tumors profiled by proteogenomics \nand are included in OPC. \nCPTAC adult GBM proteomics study \nThis CPTAC adult GBM study [6] contains 99 tumors profiled by proteogenomics and are \nincluded in OPC. \nProject HOPE proteomics study \nProject HOPE is an adolescent and young adult high-grade glioma study (in preparation for \npublication) that contains 90 tumors profiled by proteogenomics and are included in OPC. \nOpenPedCan represents a substantial expansion since the OpenPBTA, both in cohort size and \nin data modality integration. By incorporating methylation, proteomics, splicing, and reference \ndatasets, and enabling reproducible analyses across more than 48,000 biospecimens, \nOpenPedCan delivers a uniquely scalable and reusable resource for pediatric cancer research. \nContext \nCreation of this dataset had multiple motivations. First, we sought to harmonize, summarize, \nand contextualize pediatric cancer genomics data among normal tissues (GTEx) and adult \ncancer tissues (TCGA) to enable the creation of the National Cancer Institute’s Molecular \nTargets Platform (MTP) at https://moleculartargets.ccdi.cancer.gov/. The inclusion of \nharmonized GTEx and adult TCGA data specifically allows for the identification of genes and/or \ntranscripts expressed in a tumor-specific and/or pediatric tumor-specific manner. Next, we \ncreated this resource for broad community use to promote rapid reuse and accelerate the \ndiscovery of additional mechanisms contributing to the pathogenesis of pediatric cancers and/or \nto identify novel candidate therapeutic targets for pediatric cancer. \nSimilar to OpenPBTA, OpenPedCan operates on a pull request model to accept contributions. \nWe set up continuous integration software via GitHub Actions to confirm the reproducibility of \nanalyses within the project’s Docker container. We maintained a data release folder on Amazon \nS3, downloadable directly from S3 or our open-access CAVATICA project, with merged files for \neach analysis. As we produced new results, identified data issues, or added additional data, we \ncreated new data releases in a versioned manner. The project maintainers have included \nengineers and scientists from the Children’s Hospital of Philadelphia\n and Children’s National \nHospital. \nMethods \nAn overview of the OpenPedCan methods is depicted in Figure 2. Briefly, most primary data \nharmonization analysis workflows were performed with Kids First pipelines written in Common \nWorkflow Language (CWL) using CAVATICA (detailed below). Alignment and expression \nquantification for GTEx and TCGA RNA-Seq was performed by the respective consortium. \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nCustom python, R, and/or bash scripts were then created in OpenPedCan using the primary \nharmonized output files. \nFigure 2: OpenPedCan Analysis Workflow. Depicted are the datasets (yellow, orange, and \ngrey) contained within OpenPedCan. These datasets are made available in a harmonized \nmanner through primary analysis workflows (blue) for DNA, RNA, and/or proteogenomics data. \nFiles derived from the primary analysis workflows (green) are released within OpenPedCan. \nAdditional analysis modules developed within OpenPedCan (red) also generate results files \n(green) which are released within OpenPedCan. \nSample Details \nA list of all biospecimens and associated metadata can be found in Supplemental Table 1. \nNucleic acids extraction and library preparation (PBTA X01 and miRNA-\nSeq) \nFor detailed methods about the OpenPBTA cohort, please refer to the manuscript [ 1]. For the \nPBTA X01 cohort, libraries were prepped using the Illumina TruSeq Strand-Specific Protocol to \npull out poly-adenylated transcripts. \ncDNA Library Construction \n \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nTotal RNA was quantified using the Quant-iT RiboGreen RNA Assay Kit and normalized to \n5ng/ul. Following plating, 2 uL of ERCC controls (using a 1:1000 dilution) were spiked into each \nsample. An aliquot of 325 ng for each sample was transferred into library preparation. The \nresultant 400bp cDNA went through dual-indexed library preparation: ‘A’ base addition, adapter \nligation using P7 adapters, and PCR enrichment using P5 adapters. After enrichment, the \nlibraries were quantified using Quant-iT PicoGreen (1:200 dilution). Samples were normalized to \n5 ng/uL. The sample set was pooled and quantified using the KAPA Library Quantification Kit for \nIllumina Sequencing Platforms. \nmiRNA Extraction and Library Preparation \nTotal RNA for CBTN samples was extracted as described in OpenPBTA [1\n] and prepared \naccording to the HTG Edge Seq protocol for the extracted RNA miRNA Whole transcriptome \nassay (WTA). 15ng of RNA were mixed in 25ul of lysis buffer, which were then loaded onto a \n96-well plate. Human Fetal Brain Total RNA (Takara Bio USA, #636526) and Human Brain Total \nRNA (Ambion, Inc., Austin, TX, USA) were used as controls. The plate was loaded into the HTG \nEdgeSeq processor along with the miRNA WTA assay reagent pack. Samples were processed \nfor 18-20 hours, then were barcoded and amplified using a unique forward and reverse primer \ncombination. PCR settings used for barcoding and amplification were 95C for 4 min, 16 cycles \nof (95C for 15 sec, 56C for 45 sec, 68C for 45 sec), and 68C for 10 min. Barcoded and amplified \nsamples were cleaned using AMPure magnetic beads (Ampure XP,Cat# A63881). Libraries \nwere quantified using the KAPA Biosystem assay qPCR kit (Kapa Biosystems Cat#KK4824) \nand CT values were used to determine the pM concentration of each library. \nData generation \nPBTA X01 Illumina Sequencing Pooled libraries were normalized to 2nM and denatured using \n0.1 N NaOH prior to sequencing. Flowcell cluster amplification and sequencing were performed \naccording to the manufacturer’s protocols using the NovaSeq 6000. Each run was a 151bp \npaired-end with an eight-base index barcode read. Data was analyzed using the Broad Picard \nPipeline which includes de-multiplexing and data aggregation. \nPBTA miRNA Sequencing Libraries were pooled, denatured, and loaded onto sequencing \ncartridge. Libraries were sequenced using an Illumina Nextseq 500 per manufacturer guidelines. \nFASTQ files were generated from raw sequencing data using Illumina BaseSpace and analyzed \nwith the HTG EdgeSeq Parser software v5.4.0.7543 to generate an excel file containing \nquantification of 2083 miRNAs per sample. Any sample that did not pass the quality control set \nby the HTG REVEAL software version 2.0.1 (Tuscon, AR, USA) was excluded from the \nanalysis. \nPrimary Workflows through Kids First \nDNA WGS Alignment and SNP Calling \nPlease refer to the OpenPBTA manuscript for details on DNA WGS Alignment, prediction of \nparticipants’ genetic sex, and SNP calling for B-allele Frequency (BAF) generation. [1\n]. \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nSomatic Mutation and INDEL Calling \nFor matched tumor/normal samples, we used the same mutation calling methods as described \nin OpenPBTA manuscript for details [1]. For tumor only samples, we ran Mutect2 from GATK \nv4.2.2.0 using the following workflow. \nVCF annotation and MAF creation \nSomatic variants were annotated by the Ensembl Variant Effect Predictor (VEP v105) [7]. From \ntumor only variant calls, we removed variants with alt_depth == 0 or t_depth < 4. \nConsensus SNV Calling (tumor/normal only) \nWe adopted the consensus SNV calling method described in OpenPBTA manuscript with \nadjustment [1\n]. For SNV calling, we combined four consensus SNV calling algorithms: \nStrelka2[8], Mutect2[9], Lancet[10], and VarDict[11]. \nStrelka2 outputs multi-nucleotide polymorphisms (MNPs) as consecutive single-nucleotide \npolymorphisms. In order to preserve MNPs, we gather MNP calls from the other caller inputs, \nand search for evidence supporting these consecutive SNP calls as MNP candidates. Once \nfound, the Strelka2 SNP calls supporting a MNP are converted to a single MNP call. This is \ndone to preserve the predicted gene model as accurately as possible in our consensus calls. \nConsensus SNV from all four callers were collected and by default, calls that were detected in at \nleast two calling algorithms or marked with “HotSpotAllele” were retained. \nFor all SNVs, potential non-hotspot germline variants were removed if they had a normal depth \n<= 7 and gnomAD allele frequency > 0.001. Final results were saved in MAF format. \nSomatic Copy Number Variant (CNV) Calling \nWe called copy number variants for tumor/normal samples using Control-FREEC [12,13] and \nCNVkit [14] as described in the OpenPBTA manuscript [1]. We used GATK [15] to call CNVs for \nmatched tumor/normal WGS samples when there were at least 30 male and 30 female normals \nfrom the same sequencing platform available for panel of normal creation. For tumor only \nsamples, we used Control-FREEC with the following modifications. Instead of the b-allele \nfrequency germline input file, we used the dbSNP_v153_ucsc-\ncompatible.converted.vt.decomp.norm.common_snps.vcf.gz dbSNP common snps \nfile and to avoid hard-to-call regions, utilized the hg38_canonical_150.mappability \nmappability file. Both are also linked in the public Kids First references CAVATICA project. The \nControl-FREEC tumor only workflow can be found here. \nSomatic Structural Variant Calling (WGS samples only) \nWe called structural variants (SVs) using Manta [16], restricting analysis to the same regions \nutilized by Strelka2. We annotated SVs using AnnotSV [17]. \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nGene Expression \nThe tumor-normal-differential-expression module performs differential expression \nanalyses for all sets of Disease (cancer_group) and Dataset (cohort) across all genes found \nin the gene-expression-rsem-tpm-collapsed.rds table. The purpose of this analysis is \nto highlight the correlation and understand the variability in gene expression in different cancer \nconditions across different histological tissues. For OpenPedCan v12 data release, this module \nperforms expression analysis over 102 cancer groups across 52 histological tissues for all \n54,346 genes found in the dataset. This analysis was performed on the Children’s Hospital of \nPhiladelphia HPC and was configured to use 96G of RAM per CPU, with one task (one iteration \nof expression analysis for each set of tissue and cancer group) per CPU (total 102x52=5304 \nCPUs) using the R/DESeq2\n package. Please refer to script run-tumor-normal-\ndifferential-expression.sh in the module for additional details on Slurm processing \nconfiguration. The same analysis can also be performed on CAVATICA, but requires further \noptimization. The module describes the steps for CAVATICA set up, and scripts to publish an \napplication on the portal. The required data files are also available publicly on CAVATICA under \nthe Open Pediatric Cancer (OpenPedCan) Open Access\n. Refer to the module for detailed \ndescription and scripts. \nAbundance Estimation \nAmong the data sources used for OpenPedCan, GTEx and TCGA used GENCODE v27 and \nv36, respectively. Therefore, the gene symbols had to be harmonized to GENCODE v39 for \ncompatibility with the rest of the dataset. The liftover process was done via a custom script. The \nscript first constructs an object detailing the gene symbol changes from the HGNC symbol \ndatabase. Using the symbol-change object, the script updates any columns containing gene \nsymbols. This liftover process was used on GTEx RNA-Seq, TCGA RNA-Seq, DGD fusions, \nand DNA hotspot files. \nAdditionally, the gene expression matrices had some instances where multiple Ensembl gene \nidentifiers mapped to the same gene symbol. This was dealt with by filtering the expression \nmatrix to only genes with [FPKM/TPM] > 0 and then selecting the instance of the gene symbol \nwith the maximum mean [FPKM/TPM/Expected_count] value across samples. This enabled \nmany downstream modules that require RNA-seq data have gene symbols as unique gene \nidentifiers. Refer to collapse-rnaseq\n module for scripts and details. \nGene fusion detection from RNA-Seq \nGene fusions were called using Arriba [18] and STAR-Fusion [19] as previously reported in \nOpenPBTA [1]. We updated the annoFuseData R package to liftover gene symbols to be \nconcordant with VEP v105. Fusions are now filtered with annoFuse [20] upstream and released \nin fusion-annoFuse.tsv.gz. \nGene fusion detection from fusion panels (DGD only) \nClinical RNA fusion calls from the CHOP DGD fusion panel are included in the data release in \nthe fusion-dgd.tsv.gz file. \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nSplicing quantification \nTo detect alternative splicing events, we utilized rMATS turbo (v. 4.1.0) with \nEnsembl/GENCODE v39 GFF annotations using the Kids First RNA-Seq workflow. We used --\nvariable-read-length and -t paired options and applied an additional filter to include \nonly splicing events with total junction read counts greater than 10. The OpenPedCan data \nrelease file splice-events-rmats.tsv.gz contains predicted single exon (SE), alternative \n5’ splice site (A5SS), alternative 3’ splice site (A3SS), and retained intron (RI) events. These are \nmade available for the community, but were not yet used in OpenPedCan analysis modules. \nProteomics data integration \nCPTAC PBTA, CPTAC GBM, and HOPE proteogenomics \nThe following methods are the general proteomics approaches used for the CPTAC PBTA [5], \nCPTAC GBM [6], and HOPE (pre-publication, correspondence with Dr. Pei Wang) studies. For \nspecific descriptions of sample preparation, mass spectrometry instrumentation and \napproaches, and data generation, processing, or analysis please refer to the relevant \npublications. \nTMT-11 Labeling and Phosphopeptide Enrichment \nProteome and phosphoproteome analysis of brain cancer samples in the CPTAC PBTA \n(pediatric), CPTAC GBM (adult), and HOPE (adolescent and young adult, AYA) cohort studies \nwere structured as TMT11-plex experiments. Tumor samples were digested with LysC and \ntrypsin. Digested peptides were labeled with TMT11-plex reagent and prepared for \nphosphopeptide enrichment. For each dataset, a common reference sample was compiled from \nrepresentative samples within the cohort. Phosphopeptides were enriched using Immobilized \nMetal Affinity Chromatography (IMAC) with Fe3+-NTA-agarose bead kits. \nLiquid Chromatography with Tandem Mass Spectrometry (LC-MS/MS) Analysis \nTo reduce sample complexity, peptide samples were separated by high pH reversed phase \nHPLC fractionation. For CPTAC PBTA a total of 96 fractions were consolidated into 12 final \nfractions for LC-MS/MS analysis. For CPTAC GBM and HOPE cohorts a total of 96 fractions \nwere consolidated into 24 fractions. For CPTAC PBTA, global proteome mass spectrometry \nanalyses were performed on an Orbitrap Fusion Tribrid Mass Spectrometer and \nphosphoproteome analyses were performed on an Orbitrap Fusion Lumos Tribrid Mass \nSpectrometer. For CPTAC GBM and HOPE studies, mass spectrometry analysis was \nperformed using an Orbitrap Fusion Lumos Mass Spectrometer. \nProtein Identification \nThe CPTAC PBTA spectra data were analyzed with MSFragger version 20190628 [21\n] \nsearching against a CPTAC harmonized RefSeq-based sequence database containing 41,457 \nproteins mapped to the human reference genome (GRCh38/hg38) obtained via the UCSC Table \nBrowser on June 29, 2018, with the addition of 13 proteins encoded in the human mitochondrial \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\ngenome, 264 common laboratory contaminant proteins, and an equal number of decoy \nsequences. The CPTAC GBM and HOPE spectra data were analyzed with MS-GF+ v9881 \n[22,23,24] searching against the RefSeq human protein sequence database downloaded on \nJune 29, 2018 (hg38; 41,734 proteins), combined with 264 contaminants, and a decoy database \ncomposed of the forward and reversed protein sequences. \nProtein Quantification and Data Analysis \nRelative protein (gene) abundance was calculated as the ratio of sample abundance to \nreference abundance using the summed reporter ion intensities from peptides mapped to the \nrespective gene. For phosphoproteomic datasets, data were not summarized by protein but left \nat the phosphopeptide level. Global normalization was performed on the gene-level abundance \nmatrix (log2 ratio) for global proteomic and on the site-level abundance matrix (log2 ratio) for \nphosphoproteomic data. The median, log2 relative protein or peptide abundance for each \nsample was calculated and used to normalize each sample to achieve a common median of 0. \nTo identify TMT outliers, inter-TMT t-tests were performed for each individual protein or \nphosphopeptide. Batch effects were checked using the log2 relative protein or phosphopeptide \nabundance and corrected using the Combat algorithm [25\n]. Imputation was performed after \nbatch effect correction for proteins or phosphopeptides with a missing rate < 50%. For the \nphosphopeptide datasets, 440 markers associated with cold-regulated ischemia genes were \nfiltered and removed. \nCreation of OpenPedCan Analysis modules \nA list of all modules, repository links, one line description, input, and output files can be found in \nSupplemental Table 2. \nMethylation Analysis \nMethylation array preprocessing \nWe preprocessed raw Illumina 450K and EPIC 850K Infinium Human Methylation Bead Array \nintensities using the array preprocessing methods implemented in the minfi Bioconductor \npackage [26]. We utilized either preprocessFunnorm when an array dataset had both tumor \nand normal samples or multiple OpenPedCan-defined cancer_groups and \npreprocessQuantile when an array dataset had only tumor samples from a single \nOpenPedCan-defined cancer_group to estimate usable methylation measurements (beta-\nvalues and m-values) and copy number (cn-values). Some Illumina Infinium array probes \ntargeting CpG loci contain single-nucleotide polymorphisms (SNPs) near or within the probe \n[27], which could affect DNA methylation measurements [28]. As the minfi preprocessing \nworkflow recommends, we dropped probes containing common SNPs in dbSNP (minor allele \nfrequency > 1%) at the CpG interrogation or the single nucleotide extensions. \nDetails of methylation array preprocessing are available in the OpenPedCan methylation-\npreprocessing module. \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nMethylation classification of brain tumor molecular subtypes \nThe Clinical Methylation Unit Laboratory of Pathology at the National Cancer Institute Center for \nCancer Research ran the DKFZ brain classifier version 12.6, a comprehensive DNA \nmethylation-based classification of CNS tumors across all entities and age groups [29] and/or \nthe NIH Bethesda Brain tumor classifier v2.0 (NIH_v2) and the combo reporter pipeline v2.0 on \ndocker container trust1/bethesda:latest. Unprocessed IDAT-files from the Children’s Brain \nTumor Network (CBTN) Infinium Human Methylation EPIC (850k) BeadChip \narrays were used as input and the following information was compiled into the \nhistologies.tsv file: dkfz_v12_methylation_subclass (predicted methylation \nsubtype), dkfz_v12_methylation_subclass_score (classification score), \ndkfz_v12_methylation_mgmt_status (MGMT methylation status), \ndkfz_v12_methylation_mgmt_estimated (estimated MGMT methylation fraction), \nNIH_v2_methylation_Superfamily, \nNIH_v2_methylation_Superfamily_mean_score, \nNIH_v2_methylation_Superfamily_Consistency_score, \nNIH_v2_methylation_Class, NIH_v2_methylation_Class_mean_score, \nNIH_v2_methylation_Class_consistency_score, \nNIH_v2_methylation_Superfamily_match, and NIH_v2_methylation_Class_match. \nGene Set Variation Analysis (gene-set-enrichment-analysis analysis \nmodule) \nWe performed Gene Set Variation Analysis (GSVA) for the Hallmark gene sets from MSigDB \n[30] on log2-transformed, gene-collapsed RSEM TPM expression values from RNA-Seq using \nthe GSVA package from Bioconductor [31]. GSVA was performed separately by RNA library \ntype to avoid batch effects. \nFusion prioritization (fusion_filtering analysis module) \nThe fusion_filtering module filters artifacts and annotates fusion calls, with prioritization \nfor oncogenic fusions, for the fusion calls from STAR-Fusion and Arriba. After artifact filtering, \nfusions were prioritized and annotated as “putative oncogenic fusions” when at least one gene \nwas a known kinase, oncogene, tumor suppressor, curated transcription factor, on the COSMIC \nCancer Gene Census List, or observed in TCGA. Fusions were retained in this module if they \nwere called by both callers, recurrent or specific to a cancer group, or annotated as a putative \noncogenic fusion. Please refer to the module linked above for more detailed documentation and \nscripts. \nConsensus CNV Calling (WGS samples only) (copy_number_consensus_call* \nanalysis modules) \nWe adopted the consensus CNV calling described in OpenPBTA manuscript [1] with minor \nadjustments. For each caller and sample with WGS performed, we called CNVs based on \nconsensus among Control-FREEC [12,13], CNVkit [14], and GATK [15]. Sample and consensus \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\ncaller files with more than 2,500 CNVs were removed to de-noise and increase data quality, \nbased on cutoffs used in GISTIC [32]. For each sample, we included the following regions in the \nfinal consensus set: 1) regions with reciprocal overlap of 50% or more between at least two of \nthe callers; 2) smaller CNV regions in which more than 90% of regions were covered by another \ncaller. For GATK, if a panel of normal was not able to be created (required 30 male and 30 \nfemale with the same sequencing platform), consensus was run for that tumor using Control-\nFREEC, CNVkit, and MantaSV. We defined copy number as NA for any regions that had a \nneutral call for the samples included in the consensus file. We merged CNV regions within \n10,000 bp of each other with the same direction of gain or loss into single region. \nAny CNVs that overlapped 50% or more with immunoglobulin, telomeric, centromeric, segment \nduplicated regions, or that were shorter than 3000 bp were filtered out. The CNVKit calls for \nWXS samples were appended to the consensus CNV file. \nFocal Copy Number Calling (focal-cn-file-preparation analysis module) \nPlease refer to the OpenPBTA manuscript for details on assignment of copy number status \nvalues to CNV segments, cytobands, and genes [1]. We applied criteria to resolve instances of \nmultiple conflicting status calls for the same gene and sample, which are described in detail in \nthe focal-cn-file-preparation\n module. Briefly, we prioritized 1) non-neutral status calls, 2) calls \nmade from dominant segments with respect to gene overlap, and 3) amplification and deep \ndeletion status calls over gain and loss calls, respectively, when selecting a dominant status call \nper gene and sample. These methods resolved >99% of duplicated gene-level status calls. \nMutational Signatures (mutational-signatures analysis module) \nWe obtained mutational signature weights (i.e., exposures) from consensus SNVs using the \ndeconstructSigs R package [33\n]. We estimated weights for single- and double-base substitution \n(SBS and DBS, respectively) signatures from the Catalogue of Somatic Mutations in Cancer \n(COSMIC) database versions 2 and 3.3, as well as SBS signatures from Alexandrov et al. 2013 \n[34\n]. The following COSMIC SBS signatures were excluded from weight estimation in all tumors: \n1) sequencing artifact signatures, 2) signatures associated with environmental exposure, and 3) \nsignatures with an unknown etiology. Additionally, we excluded therapy-associated signatures \nfrom mutational signature weight estimation in tumors collected prior to treatment (i.e. “Initial \nCNS Tumor” or “Primary Tumor”). \nTumor Mutation Burden [TMB] (tmb-calculation analysis module) \nRecent clinical studies have associated high TMB with improved patient response rates and \nsurvival benefit from immune checkpoint inhibitors [35]. \nThe Tumor Mutation Burden (TMB) tmb-calculation module was adapted from the snv-\ncallers module of the OpenPBTA project [1]. Here, we use mutations in the snv-\nconsensus-plus-hotspots.maf.tsv.gz file which is generated using Kids First DRC \nConsensus Calling Workflow and is included in the OpenPedCan data download. The \nconsensus MAF contains SNVs or MNVs called in at least 2 of the 4 callers (Mutect2, Strelka2, \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nLancet, and Vardict) plus hotspot mutations if called in 1 of the 4 callers. We calculated TMB for \ntumor samples sequenced with either WGS or WXS. Briefly, we split the SNV consensus MAF \ninto SNVs and multinucleotide variants (MNVs). We split the MNV subset into SNV calls, \nmerged those back with the SNVs subset, and then removed sample-specific redundant calls. \nThe resulting merged and non-redundant SNV consensus calls were used as input for the TMB \ncalculation. We tallied only nonsynonymous variants with classifications of high/moderate \nconsequence (“Missense_Mutation”, “Frame_Shift_Del”, “In_Frame_Ins”, “Frame_Shift_Ins”, \n“Splice_Site”, “Nonsense_Mutation”, “In_Frame_Del”, “Nonstop_Mutation”, and \n“Translation_Start_Site”) for the numerator. All BED files are provided in the data release. \nAll mutation TMB \nFor WGS samples, we calculated the size of the genome covered as the intersection of Strelka2 \nand Mutect2’s effectively surveyed areas, regions common to all variant callers, and used this \nas the denominator. WGS_all_mutations_TMB = (total # mutations in consensus \nMAF) / intersection_strelka_mutect_vardict_genome_size For WXS samples, \nwe used the size of the WXS bed region file as the denominator. WXS_all_mutations_TMB = \n(total # mutations in consensus MAF)) / wxs_genome_size \nCoding only TMB \nWe generated coding only TMB from the consensus MAF as well. We calculated the \nintersection for Strelka2 and Mutect2 surveyed regions using the coding sequence ranges in the \nGENCODE v39 gtf supplied in the OpenPedCan data download. We removed SNVs outside of \nthese coding sequences prior to implementing the TMB calculation below: \nWGS_coding_only_TMB = (total # coding mutations in consensus MAF) / \nintersection_wgs_strelka_mutect_vardict_CDS_genome_size For WXS samples, \nwe intersected each WXS bed region file with the GENCODE v39 coding sequence, sum only \nvariants within this region for the numerator, and calculate the size of this region as the \ndenominator. WXS_coding_only_TMB = (total # coding mutations in consensus \nMAF) / intersection_wxs_CDS_genome_size \nFinally, we include an option (nonsynfilter_focr) to use specific nonsynonymous mutation \nvariant classifications recommended from the TMB Harmonization Project\n. \nMolecular Subtyping \nHere, we build upon the molecular subtyping performed in OpenPBTA [1] to align with WHO \n2021 subtypes [36]. Molecular subtypes were generated per tumor event and are listed for each \nbiospecimen in Supplemental Table 1, with the number of tumors grouped by broad histology \nand molecular subtype in Supplemental Table 3. \nHigh-grade gliomas \nHigh-grade gliomas (HGG) were categorized based on a combination of clinical information, \nmolecular features, and DNA methylation data. H3 K28-altered diffuse midline gliomas (DMG) \nwere classified based on the presence of a p.K28M or p.K28I mutation in H3F3A, HIST1H3B, \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nHIST1H3C, or HIST2H3C, or a high-confidence DKFZ methylation score (>=0.8) in the \nappropriate subclass. Oligodendroglioma, IDH-mutant tumors were classified based on high-\nconfidence “O_IDH” methylation classifications, and oligosarcoma, IDH-mutant tumors were \ndefined as those with high-confidence “OLIGOSARC_IDH” methylation classifications. \nPleomorphic xanthoastrocytomas (PXA) were classified using the following criteria: 1) \nmethylation subtype is high-confidence “PXA” or pathology_free_text_diagnosis \ncontains “pleomorphic xanthoastrocytoma” or “pxa”, and 2) tumor contains a BRAF V600E \nmutation and a CDKN2A or CDKN2B homozygous deletion. Methylation classifications were \nused in classifying the following subtypes: \n1. DHG, H3 G35 (“DHG_G34” and “GBM_G34” classifications) \n2. HGG, IDH (“A_IDH_HG” and “GBM_IDH” classifications) \n3. HGG, H3 wild type (methylation classification contains “GBM_MES”, “GBM_RTK”, \n“HGG_”, “HGAP”, “AAP”, or “ped_”) \nA new high-grade glioma entity called infant-type hemispheric gliomas (IHGs), characterized by \ndistinct gene fusions enriched in receptor tyrosine kinase (RTK) genes including ALK, \nNTRK1/2/3, ROS1 or MET, was identified in 2021 [37\n]. To identify IHG tumors, first, tumors \nwhich were classified as “IHG” by the DKFZ methylation classifier or diagnosed as “infant type \nhemispheric glioma” from pathology_free_text_diagnosis were selected [29]. Then, the \ncorresponding tumor RNA-seq data were utilized to seek the evidence for RTK gene fusion. \nBased on the specific RTK gene fusion present in the samples, IHGs were further classified as \n“IHG, ALK-altered”, “IHG, NTRK-altered”, “IHG, ROS1-altered”, or “IHG, MET-altered”. If no \nfusion was observed, the samples were identified as “IHG, To be classified”. \nAtypical teratoid rhabdoid tumors \nAtypical teratoid rhabdoid tumors (ATRT) tumors were categorized into three subtypes: “ATRT, \nMYC”, “ATRT, SHH”, and “ATRT, TYR” [38]. In OpenPedCan, the molecular subtyping of ATRT \nwas based solely on the DNA methylation data. Briefly, ATRT samples with a high confidence \nDKFZ methylation subclass score (>= 0.8) were selected and subtypes were assigned based on \nthe DKFZ methylation subclass [29]. Samples with low confidence DKFZ methylation subclass \nscores (< 0.8) were identified as “ATRT, To be classified”. \nNeuroblastoma tumors \nNeuroblastoma (NBL) tumors with a pathology diagnosis of neuroblastoma, \nganglioneuroblastoma, or ganglioneuroma were subtyped based on their MYCN copy number \nstatus as either “NBL, MYCN amplified” or “NBL, MYCN non-amplified”. If \npathology_free_text_diagnosis was “NBL, MYCN non-amplified” and the genetic data \nsuggested MYCN amplification, the samples were subtyped as “NBL, MYCN amplified”. On the \nother hand, if pathology_free_text_diagnosis was “NBL, MYCN amplified” and the \ngenetic data suggested MYCN non-amplification, the RNA-Seq gene expression level of MYCN \nwas used as a prediction indicator. In those cases, samples with MYCN gene expression above \nor below the cutoff (TPM >= 140.83 based on visual inspection of MYCN CNV status) were \nsubtyped as “NBL, MYCN amplified” and “NBL, MYCN non-amplified”, respectively. MYCN gene \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nexpression was also used to subtype samples without DNA sequencing data. If a sample did not \nfit none of these situations, it was denoted as “NBL, To be classified”. \nCraniopharyngiomas \nIn addition to molecular criteria established in OpenPBTA [1], craniopharyngiomas (CRANIO) \nare now subtyped using DNA methylation classifiers. Craniopharyngiomas with a high-\nconfidence methylation subclass containing “CPH_PAP” were classified as papillary (CRANIO, \nPAP), and those with high-confidence methylation subclass containing “CPH_ADM” were \nclassified as adamantinomatous (CRANIO, ADAM), respectively. \nEpendymomas \nEpendymomas (EPN) are subtyped using the following criteria: \n1. Any spinal tumor with MYCN amplification or with a high-confidence “EPN, SP-MYCN” \nmethylation classification was subtyped as EPN, spinal and MYCN-amplified (SP-\nMYCN). \n2. EPN tumors containing one or more gene fusions of YAP1::MAMLD1, YAP1::MAML2, or \nYAP1::FAM118B, or else had a high-confidence “EPN, ST YAP1” methylation \nclassification were subtyped as EPN, ST YAP1. \n3. EPN tumors containing one or more gene fusions of ZFTA::RELA or ZFTA::MAML2, or \nelse had a high-confidence “EPN, ST ZFTA” methylation classification were subtyped as \nEPN, ST ZFTA. This reflects an update to WHO classifications that now characterizes \nthis subtype based on ZFTA fusions rather than RELA fusions. \n4. EPN tumors with 1) chromosome 1q gain and TKTL1 over-expression, or 2) EZHIP over-\nexpression, or 3) posterior fossa anatomical location and a histone H3 K28 mutation in \nH3F3A, HIST1H3B, HIST1H3C, or HIST2H3C, or 4) a high-confidence “EPN, PF A” \nmethylation classification were subtyped as posterior fossa group A ependymomas \n(EPN, PF A). \n5. Tumors with 1) chr 6p or 6q loss and GPBP1 or IFT46 over-expression, or 2) a high-\nconfidence “EPN, PF B” methylation classification were subtyped as posterior fossa \ngroup B ependymomas (EPN, PF B). \n6. EPN tumors with a high-confidence “EPN, MPE” methylation classification were \nsubtyped as myxopapillary ependymomas (EPN, MPE). \n7. EPN tumors with a high-confidence “EPN, PF SE” methylation classification were \nsubtyped as posterior fossa subependymomas (EPN, PF SE). \n8. EPN tumors with a high-confidence “EPN, SP SE” methylation classification were \nsubtyped as spinal subependymomas (EPN, SP SE). \n9. EPN tumors with a high-confidence “EPN, SP” methylation classification were subtyped \nas spinal ependymomas (EPN, SP). \n10. All other EPN tumors were classified as “EPN, To be classified”. \nLow-grade gliomas \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nIn addition to subtyping methods described in OpenPBTA [1], high-confidence methylation \nclassifications are now used in classifying the following low-grade glioma (LGG) subtypes: \n1. LGG, other MAPK-altered (methylation subclass “PA_MID” or “PLNTY”) \n2. LGG, FGFR-altered (methylation subclass “PA_INF_FGFR”) \n3. LGG, IDH-altered (methylation subclass “A_IDH_LG”) \n4. LGG, MYB/MYBL1 fusion (methylation subclass “AG_MYB” or “LGG_MYB”) \n5. LGG, MAPK-altered (methy lation subclass “LGG, MAPK”) \n6. LGG, BRAF- and MAPK-altered (methylation subclass “LGG, BRAF/MAPK”) \n7. SEGA, to be classified (methylation subclass “SEGA, To be classified”) \nMedulloblastomas (MBs) In addition to our previous work classifying MB tumors into the four \nmajor subtypes (WNT, SHH, Group 3, and Group 4) using the transcriptomic MedulloClassifier \n[39\n], we integrated high-confidence methylation classification, demographic, and molecular \ncriteria to molecularly subtype SHH tumors into one of four subgroups (alpha, beta, gamma, or \ndelta) (Figure 3). \n \nFigure 3: Medulloblastoma Sample Clustering. A, UMAP projection of 271 MB tumors and B, \n63 SHH-activated MB tumors using methylation beta values of the 20,000 most variable probes \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nfrom the Infinium MethylationEPIC array. C, UMAP projection of MB, SHH activated samples \nindicating copy number status of SHH subgroup known somatic driver genes CCND2, GLI2, \nMYCN, and PTEN. \nWe implemented molecular subtyping as follows: \n1. MB tumors with methylation classification that contains “MB_SHH” are subtyped as \nSHH-activated medulloblastoma (MB, SHH) \n2. MB tumors with “MB_G34_I”, “MB_G34_II”, “MB_G34_III”, and “MB_G334_IV” \nmethylation classifications are subtyped as medulloblastoma group 3 (MB, Group3) \n3. MB tumors with “MB_G34_V”, “MB_G34_VI”, “MB_G34_VII”, and “MB_G334_VIII” \nmethylation classifications are subtyped as medulloblastoma group 4 (MB, Group4) \n4. MB tumors with “MB_WNT” methylation classification are subtyped as WNT-activated \nMB (MB, WNT) \n5. MB tumors with “MB_MYO” methylation classification are subtyped as medulloblastomas \nwith myogenic differentiation (MB, MYO) \nWe classified MB, SHH subtype tumors using the following criteria: \n1. MB, SHH alpha: sample has a high-confidence “MB_SHH_3” methylation classification, \nor patient had an age at diagnosis >= 2 years and harbored one of the following \nmolecular alterations in tumor or germline: \n– MYCN, GLI2, or CCND2 amplification or sample TPM z-score >= 2 in tumor. \n– A pathogenic or likely pathogenic germline variant in ELP1  or TP53. \n– A TP53 hotspot mutation in tumor. \n– Chromosome 9p gain or chromosome 17p loss in tumor. \n2. MB, SHH beta: sample has a high-confidence “MB_SHH_1” methylation classification, or \npatient had an age at diagnosis < 5 years and harbored one of the following molecular \nalterations: \n– A KMT2D loss of function variant. \n– PTEN copy number loss or deep deletion, or sample TPM z-score < -2. \n– Chromosome 2p or 2q gain. \n3. MB, SHH gamma: sample has a high-confidence “MB_SHH_2” methylation \nclassification, or patient had an age at diagnosis < 5 years and tumor harbored a \nchromosome 2p arm gain. \n4. MB, SHH delta: sample has a high-confidence “MB_SHH_4” methylation classification, \nor patient had an age at diagnosis >= 10 years and harbored one of the following \nmolecular alterations in tumor: \n– a DDX3X or SMO loss-of-function mutation. \n– a hotspot TERT or U1 snRNA gene mutation. \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\n– Chromosome 14q arm loss. \nPineoblastomas \nPineoblastomas (PB) are classified as follows using high-confidence methylation classifications: \n1. Pineoblastoma, MYC/FOXR2-activated (“PB_FOXR2” methylation classification) \n2. Pineoblastoma, RB1-altered (“PB_RB1” methylation classification) \n3. Pineoblastoma, group 1 (“PB_GRP1A” and “PB_GRP1B” methylation classifications) \n4. Pineoblastoma, group 2 (“PB_GRP2” methylation classification) \n5. All other pineoblastomas were classified as “PB, To be classified” \nnon-MB, non-ATRT Embryonal Tumors \nUpdates were made to non-MB, non-ATRT embryonal tumor subtyping as follows: \n1. Embryonal tumors with multilayered rosettes and C19MC-altered (ETMR, C19MC-\naltered) were classified based on 1) high-confidence “ETMR_C19MC” methylation \nclassification or 2) TTYH1 gene fusion and either chromosome 19 amplification or \nLIN28A over-expression. \n2. ETMR, not otherwise specified (NOS) were classified based on LIN28A over-expression \nand no TTYH1 gene fusion. \nTP53 Alteration Annotation (tp53_nf1_score analysis module) \nWe classified TP53-altered high-grade glioma (HGG) samples as either TP53 lost or TP53 \nactivated and incorporated these annotations into the molecular subtype framework. To support \nthis classification, we used a previously published RNA-based TP53 inactivation signature \noriginally developed using TCGA pan-cancer cohorts [40]. We applied this to OpenPedCan \nRNA-seq data, stratified by library preparation type. This classifier was used in combination with \ngenomic variant data, including consensus SNVs, CNVs, and structural variants (SVs), as well \nas curated reference databases cataloging somatic TP53 hotspot mutations [41,42] and known \nfunctional domains [43] to annotate lost or activated status. Briefly, samples were annotated as \nTP53 activated if they harbored either of two known gain-of-function mutations: p.R273C or \np.R248W [44]. Samples were assigned TP53 lost status under any of the following conditions: \n(i) presence of a hotspot TP53 mutation listed in the IARC or MSKCC databases; (ii) detection \nof two distinct TP53 alterations (e.g., SNV, CNV, or SV) consistent with biallelic inactivation; (iii) \npresence of a single somatic TP53 variant or a pathogenic germline variant associated with Li-\nFraumeni syndrome (LFS) [45\n]; or (iv) presence of a germline TP53 variant linked to LFS \nalongside a TP53 inactivation classifier score >0.5 from matched RNA-seq data. \nClinical data harmonization \nTo remain consistent with the Kids First data model and our previous OpenPBTA study [1], all \nclinical metadata was harmonized using the same data model. TARGET and TCGA metadata \nfields (e.g., sample_type, composition, tumor_descriptor, etc.) were harmonized to \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nthose of Kids First. Additional histology-related fields were created through OpenPedCan, \nfollowing molecular subtyping: integrated_diagnosis, harmonized_diagnosis, and \ncancer_group. These fields were expanded from our previous study, to utilize the WHO 2021 \nCNS tumor classifications[36]. Any samples with molecular subtypes which did not match the \ninitial pathology_diagnosis were reviewed with a board-certified molecular pathologist and \nupdated accordingly. \nEFO, MONDO, and NCIT Mapping \nWe created a script to search ontology mappings by cancer_group. The efo_code \nrepresents the Experimental Factor Ontology (EFO) description available in European \nBioinformatics Institute database, the mondo_code represents the Mondo Disease Ontology \n(MONDO) from an independent resource that aims to harmonize disease definitions, and the \nncit_code represents the NCI Thesaurus (NCIt) reference terminology. Codes were \nautomatically pulled based on text matching, manually reviewed, and can be found in \nSupplemental Table 1\n \nSelection of independent samples (independent-samples analysis module) \nFor analyses that require all input biospecimens to be independent, we use the OpenPedCan-\nanalysis independent-samples module to select only one biospecimen from each input \nparticipant. For each input participant of an analysis, the independent biospecimen is selected \nbased on the analysis-specific filters and preferences for the biospecimen metadata, such as \nexperimental strategy, cancer group, and tumor descriptor. \nData Validation and Quality Control \nAll RNA-seq and WGS samples passed minimum quality thresholds, including ≥ 20 million total \nreads and ≥ 50% alignment for RNA-Seq, and ≥ 20X mean coverage for DNA sequencing \nSupplemental Table 4. Sample identity was confirmed using NGSCheckMate [46] and \nSomalier relate [47] to detect and exclude mismatched or contaminated samples. \nWe expanded upon the molecular subtyping modules from OpenPBTA to recover hallmark \ngenomic and transcriptomic features known in pediatric tumors. These include: \nKIAA1549::BRAF fusions in low-grade gliomas, H3 K28M/I mutations in diffuse midline gliomas, \nH3 G35R/V mutations in diffuse hemispheric gliomas, somatic TP53 mutations in high-grade \ngliomas, and MYCN amplification in neuroblastoma, for example. \nAll subtyping modules are version-controlled, containerized, and publicly available, and have \nundergone internal code review and validation by independent analysts. Where molecular \nfeatures conflicted with original pathology labels, cases were reviewed with board-certified \nmolecular pathologists, and integrated diagnoses were updated accordingly. This collaborative \nre-review process led to improved sample annotation and is fully documented in the molecular-\nsubtype-pathology module. \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nTo assess concordance between data types, we compared RNA-based and methylation-based \nmolecular subtypes in medulloblastoma. As shown in [Table  1], we observed nearly 100% \nconcordance, validating both experimental modalities and classifier accuracy. Notably, \nmethylation classification identified one rare case (MB, MYO) not captured by the transcriptome-\nbased MedulloClassifier. \nTable 1: Medulloblastoma subtype concordance across experimental strategies. \nComparison of medulloblastoma subtypes using methylation or RNA-Seq classification.  \nMethylation Subtype Group3 (RNA-Seq) Group4 (RNA-Seq) SHH (RNA-Seq) WNT (RNA-Seq) \nMB_G34_II 8  0 0 0 \nMB_G34_III 17  0 0 0 \nMB_G34_IV 8  0 0 0 \nMB_G34_V 0  6 0 0 \nMB_G34_VI 0  4 0 0 \nMB_G34_VII 0  30 0 0 \nMB_G34_VIII 0  34 0 0 \nMB_MYO 1  0 0 0 \nMB_SHH_1 0  0 11 0 \nMB_SHH_2 0  0 4 0 \nMB_SHH_3 0  0 2 0 \nMB_SHH_4 0  0 7 0 \nMB_WNT 0  0 0 18 \nIn addition to verifying known findings, OpenPedCan modules support pediatric cancer \ndiscovery and translation. The reproducibility of these results is further supported by their reuse \nacross studies, > 100 Zenodo downloads, GitHub forks, and independent analysis pipelines. \nTogether, these validation measures—spanning sample QC, molecular feature recovery, cross-\nplatform concordance, and expert review—ensure that OpenPedCan is a robust, reproducible, \nand reusable resource for the pediatric cancer research community. \nEthics and Consent Statement \nThis study did not generate new sequencing data. All previously-published raw data were \nobtained through Database of Genotypes and Phenotypes (dbGAP) access requests with \npatients consented as “General Research Use (GRU)” or “Disease-Specific (Pediatric Cancer \nResearch)”. OpenPedCan integrates only summary-level outputs (e.g., gene expression \nmatrices, mutation calls) that are designated for GRU by the data custodians. No protected \nhealth information (PHI), raw sequencing files, or individually identifiable clinical metadata are \ndistributed as part of this project. \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nRe-use potential \nOpenPedCan represents a valuable resource, not only by significantly extending OpenPBTA to \ninclude more than 5,000 additional patients and 6,000 tumors, but also by adding a number of \nnew “omic” data types not previously included, such as methylation arrays, miRNA-Seq, \nproteomics, and normal tissue RNA-Seq. OpenPedCan also serves as a community resource \nwhose outputs and/or code can be leveraged directly to ask research questions or serve as an \northogonal validation dataset. By providing this data in a harmonized manner, we enable \ninvestigators to reduce the financial and time-related costs associated with their analyses, which \nwould otherwise total years of project hours and over $50,000 in data analysis alone [48\n]. We \nencourage re-use of the data, ideas and suggestions for improving the data or adding analyses, \nand/or direct code contributions through a pull-request. \nAvailability of source code and requirements \nProject name: The Open Pediatric Cancer (OpenPedCan) Project \nProject home page: https://github.com/d3b-center/OpenPedCan-analysis \nArchived Source code: https://zenodo.org/records/15750097 \nOperating system(s): Platform independent \nProgramming languages: R, Python, bash \nOther requirements: CAVATICA is required to run all primary Kids First workflows. All \ndownstream OpenPedCan workflows can be run using the Docker image at pgc-\nimages.sbgenomics.com/d3b-bixu/openpedcanverse:latest. Most workflows run \nefficiently on local or cloud machines with 16–64 GB RAM. The most memory-intensive module \nruns on a 64 GB instance at <$2 per run. \nLicense: CC-BY 4.0 \nPrimary analyses were performed using Gabriella Miller Kids First pipelines and are listed in the \nmethods section. Analysis modules were either initially developed within \nhttps://github.com/AlexsLemonade/OpenPBTA-analysis\n [1], were modified, and/or created anew \nwithin the https://github.com/d3b-center/OpenPedCan-analysis publicly available repository. \nSoftware versions are documented in Supplemental Table 5. \nData Availability \nDatasets \nThe datasets supporting this study are available as follows: The TARGET dataset is available in \ndbGAP under phs000218.v23.p8 [49\n]. The GMKF Neuroblastoma dataset is available in dbGAP \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nunder phs001436.v1.p1[50]. The Pediatric Brain Tumor Atlas data (PBTA), containing the \nsubcohorts OpenPBTA, Kids First PBTA (X01), Chordoma Foundation, MI-ONCOSEQ Study, \nPNOC, and DGD is available in dbGAP under phs002517.v4.p2 [51] or in the Kids First Portal \n(kidsfirstdrc.org). The raw Genotype-Tissue Expression (GTEx) dataset is available in dbGAP \nunder phs000424.v9.p2 and publicly available at https://gtexportal.org/home/. The Cancer \nGenome Atlas (TCGA) dataset is available in dbGAP under phs000178.v11.p8 [52]. \nMerged summary files for the latest release of OpenPedCan are openly accessible in \nCAVATICA or via download-data.sh script in the https://github.com/d3b-\ncenter/OpenPedCan-analysis repository. Cancer group summary data from release v11 are \nvisible within the NCI’s pediatric Molecular Targets Platform. Cohort, cancer group, and \nindividual data are visible within PedcBioPortal. An overview of the OpenPedCan data \navailability is summarized in [Table 22]. \nTable 2: OpenPedCan Data Availability. OpenPedCan data is available on multiple platforms \nwith varying access requirements.  \nPlatform Data Type Access \nType \nAccess Requirement \nPedcBioPortal Individual and summary somatic \ndata \nQuery Gmail account \nMolecular Targets \nPlatform \nCancer group summary data Query Open Access \nGitHub Merged summary files Full access AWS S3 download script \nCAVATICA Merged summary files Full access CAVATICA account \ndbGAP - \nphs002517.v4.p2 \nRaw data Full access Access request via \ninstitution \nAcknowledgments \nWe are incredibly grateful to each patient and family for donating tissue and associated \nmetadata and clinical data to their respective consortia. This project has been funded in whole \nor in part with Federal funds from the National Cancer Institute, National Institutes of Health, \nunder Contract No. 75N91019D00024, Task Order No. 75N91020F00003 (DMT, JLR, SJD, \nJMM, ST, AF, ACR). The content of this publication does not necessarily reflect the views or \npolicies of the Department of Health and Human Services, nor does mention of trade names, \ncommercial products or organizations imply endorsement by the U.S. Government. The authors \nalso wish to thank the anonymous private investors to the Children’s National Hospital Brain \nTumor Institute who have supported this work. We thank Rocky Breslow for GitHub actions \ncontributions and Rust Turakulov for contributing to methylation data analysis. \nAuthor Contributions \nAuthor Contributions \nZhuangzhuang Data curation, Formal analysis, Investigation, Methodology, Software, Writing – \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nAuthor Contributions \nGeng Original draft \nEric Wafula Formal analysis, Software, Investigation, Writing – Original draft \nRyan J. Corbett Formal analysis, Writing - original draft \nYuanchao Zhang Software, Formal analysis, Methodology, Writing – Original draft \nRun Jin Formal analysis \nKrutika S. \nGaonkar \nData curation, Formal analysis, Investigation \nSangeeta Shukla Formal analysis, Investigation, Methodology, Writing – Original draft, Writing - \nReview and editing \nKomal S. Rathi Formal analysis, Investigation, Methodology \nDave Hill Formal analysis, Writing - original draft \nAditya Lahiri Formal analysis, Investigation, Methodology, Writing – Original draft \nDaniel P. Miller Formal analysis, Writing – Original draft \nAlex Sickler Methodology, Formal analysis \nKelsey Keith Writing - original draft, Formal Analysis \nChristopher \nBlackden \nSoftware \nAntonia Chroni Validation \nMiguel A. Brown Data curation, Methodology, Formal analysis, Investigation, Software, Supervision, \nWriting – Original draft \nAdam A. Kraya Methodology \nKaylyn L. Clark Data Curation, Writing - Review and editing \nBrian R. Rood Data Curation \nAdam C. Resnick Funding acquisition, Resources \nNicholas Van \nKuren \nData curation, Software \nJohn M. Maris Funding acquisition \nAlvin Farrel Supervision, Investigation, Methodology, Funding acquisition \nMateusz P. \nKoptyra \nData curation, Investigation, Methodology, Supervision \nGerri R. Trooskin Funding acquisition \nNoel Coleman Data curation \nYuankun Zhu Supervision \nStephanie \nStefankiewicz \nProject administration \nZied Abdullaev Formal Analysis, Investigation, Data curation \nAsif T Chinwalla Project administration, Supervision, Methodology, Investigation, Validation \nMariarita Santi Investigation, Validation \nAmmar S. Naqvi Methodology, Writing – Original draft \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nAuthor Contributions \nJennifer L. Mason Supervision \nCarl J. \nKoschmann \nData Curation \nXiaoyan Huang Formal analysis, Software \nSharon J. Diskin Funding acquisition \nKenneth Aldape Formal Analysis, Investigation, Data curation \nBailey K. Farrow Data curation, Software, Project administration, Supervision \nWeiping Ma Formal Analysis, Investigation, Data curation \nBo Zhang Data curation, Formal analysis \nBrian M. Ennis Formal analysis \nSarah Tasian Funding acquisition \nSaksham Phul Formal analysis \nMatthew R. \nLueder \nData curation \nChuwei Zhong Formal analysis \nJoseph M. Dybas Writing – Original draft, Methodology \nPei Wang Formal Analysis, Supervision \nDeanne Taylor Conceptualization, Data curation, Funding acquisition, Investigation, Methodology, \nSupervision, Project administration \nJo Lynne Rokita Conceptualization, Data curation, Formal analysis, Funding acquisition, Project \nadministration, Investigation, Methodology, Software, Supervision, Writing – Original \ndraft, Writing - Review and editing \nDeclarations of Interest \nThe authors declare no conflicts. \nSupplemental Information Titles and Legends \nSupplemental Table 1 README, metadata, and clinical data for each patient and biospecimen \nin OpenPedCan. \nSupplemental Table 2 Description of OpenPedCan analysis modules. Listed are the modules, \nshort descriptions, links, input files, whether the resulting files are contained in data releases, \nand which files are consumed in other analyses. \nSupplemental Table 3 Number of tumors and corresponding patients from which WHO 2021 \nmolecular subtypes were generated through OpenPedCan analysis modules are listed in Sheet \n1. Molecular subgroups (alpha, beta, gamma, or delta) for medulloblastoma SHH tumors are \nlisted in Sheet 2. \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nSupplemental Table 4 Read count and coverage for DNA- and RNA-sequencing \nbiospecimens. \nSupplemental Table 5 Listed are the software versions for all packages and workflows used in \nthis manuscript. \nReferences \n1.  OpenPBTA: The Open Pediatric Brain Tumor Atlas Joshua A Shapiro, Krutika S \nGaonkar, Stephanie J Spielman, Candace L Savonen, Chante J Bethell, Run Jin, Komal S \nRathi, Yuankun Zhu, Laura E Egolf, Bailey K Farrow, … Jaclyn N Taroni Cell Genomics (2023-\n07) https://doi.org/gr92p6 DOI: 10.1016/j.xgen.2023.100340 · PMID: 37492101 · PMCID: \nPMC10363844 \n2.  Genomic Profiling of Childhood Tumor Patient-Derived Xenograft Models to \nEnable Rational Clinical Trial Design Jo Lynne Rokita, Komal S Rathi, Maria F Cardenas, \nKristen A Upton, Joy Jayaseelan, Katherine L Cross, Jacob Pfeil, Laura E Egolf, Gregory P \nWay, Alvin Farrel, … John M Maris Cell Reports (2019-11) https://doi.org/gg596n DOI: \n10.1016/j.celrep.2019.09.071 · PMID: 31693904 · PMCID: PMC6880934 \n3.  Transcriptomic profiling of 39 commonly-used neuroblastoma cell lines Jo Lynne \nHarenza, Maura A Diamond, Rebecca N Adams, Michael M Song, Heather L Davidson, Lori S \nHart, Maiah H Dent, Paolo Fortina, CPatrick Reynolds, John M Maris Scientific Data (2017-03-\n28) https://doi.org/f9v8hh\n DOI: 10.1038/sdata.2017.33 · PMID: 28350380 · PMCID: \nPMC5369315 \n4.  Michigan Center for Translational Pathology https://mctp.med.umich.edu \n5.  Integrated Proteogenomic Characterization across Major Histological Types of \nPediatric Brain Cancer Francesca Petralia, Nicole Tignor, Boris Reva, Mateusz Koptyra, \nShrabanti Chowdhury, Dmitry Rykunov, Azra Krek, Weiping Ma, Yuankun Zhu, Jiayi Ji, … \nWilliam E Bocik Cell (2020-12) https://doi.org/ghqjkz DOI: 10.1016/j.cell.2020.10.044 · PMID: \n33242424 · PMCID: PMC8143193 \n6.  Proteogenomic and metabolomic characterization of human glioblastoma Liang-Bo \nWang, Alla Karpova, Marina A Gritsenko, Jennifer E Kyle, Song Cao, Yize Li, Dmitry Rykunov, \nAntonio Colaprico, Joseph H Rothstein, Runyu Hong, … Jun Zhu Cancer Cell (2021-04) \nhttps://doi.org/gh7whf\n DOI: 10.1016/j.ccell.2021.01.006 · PMID: 33577785 · PMCID: \nPMC8044053 \n7.  The Ensembl Variant Effect Predictor William McLaren, Laurent Gil, Sarah E Hunt, \nHarpreet Singh Riat, Graham RS Ritchie, Anja Thormann, Paul Flicek, Fiona Cunningham \nGenome Biology (2016-06-06) https://doi.org/gdz75c\n DOI: 10.1186/s13059-016-0974-4 · PMID: \n27268795 · PMCID: PMC4893825 \n8.  Strelka2: fast and accurate calling of germline and somatic variants Sangtae Kim, \nKonrad Scheffler, Aaron L Halpern, Mitchell A Bekritsky, Eunho Noh, Morten Källberg, Xiaoyu \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nChen, Yeonbin Kim, Doruk Beyter, Peter Krusche, Christopher T Saunders Nature Methods \n(2018-07-16) https://doi.org/gdwrp4 DOI: 10.1038/s41592-018-0051-x · PMID: 30013048 \n9.  Calling Somatic SNVs and Indels with Mutect2 David Benjamin, Takuto Sato, Kristian \nCibulskis, Gad Getz, Chip Stewart, Lee Lichtenstein Cold Spring Harbor Laboratory (2019-12-\n02) https://doi.org/ggntwv DOI: 10.1101/861054 \n10.  Genome-wide somatic variant calling using localized colored de Bruijn graphs \nGiuseppe Narzisi, André Corvelo, Kanika Arora, Ewa A Bergmann, Minita Shah, Rajeeva \nMusunuri, Anne-Katrin Emde, Nicolas Robine, Vladimir Vacic, Michael C Zody Communications \nBiology (2018-03-22) https://doi.org/gfcfr8 DOI: 10.1038/s42003-018-0023-9 · PMID: 30271907 \n· PMCID: PMC6123722 \n11.  VarDict: a novel and versatile variant caller for next-generation sequencing in \ncancer research Zhongwu Lai, Aleksandra Markovets, Miika Ahdesmaki, Brad Chapman, \nOliver Hofmann, Robert McEwen, Justin Johnson, Brian Dougherty, JCarl Barrett, Jonathan R \nDry Nucleic Acids Research (2016-04-07) https://doi.org/f8v6qz\n DOI: 10.1093/nar/gkw227 · \nPMID: 27060149 · PMCID: PMC4914105 \n12.  Control-FREEC: a tool for assessing copy number and allelic content using next-\ngeneration sequencing data Valentina Boeva, Tatiana Popova, Kevin Bleakley, Pierre Chiche, \nJulie Cappo, Gudrun Schleiermacher, Isabelle Janoueix-Lerosey, Olivier Delattre, Emmanuel \nBarillot Bioinformatics (2011-12-06) https://doi.org/ckt4vz\n DOI: 10.1093/bioinformatics/btr670 · \nPMID: 22155870 · PMCID: PMC3268243 \n13.  Control-free calling of copy number alterations in deep-sequencing data using \nGC-content normalization Valentina Boeva, Andrei Zinovyev, Kevin Bleakley, Jean-Philippe \nVert, Isabelle Janoueix-Lerosey, Olivier Delattre, Emmanuel Barillot Bioinformatics (2010-11-15) \nhttps://doi.org/c6bcps DOI: 10.1093/bioinformatics/btq635 · PMID: 21081509 · PMCID: \nPMC3018818 \n14.  CNVkit: Genome-Wide Copy Number Detection and Visualization from Targeted \nDNA Sequencing Eric Talevich, AHunter Shain, Thomas Botton, Boris C Bastian PLOS \nComputational Biology (2016-04-21) https://doi.org/c9pd DOI: 10.1371/journal.pcbi.1004873 · \nPMID: 27100738 · PMCID: PMC4839673 \n15.  The Genome Analysis Toolkit: A MapReduce framework for analyzing next-\ngeneration DNA sequencing data Aaron McKenna, Matthew Hanna, Eric Banks, Andrey \nSivachenko, Kristian Cibulskis, Andrew Kernytsky, Kiran Garimella, David Altshuler, Stacey \nGabriel, Mark Daly, Mark A DePristo Genome Research (2010-07-19) https://doi.org/bnzbn6\n \nDOI: 10.1101/gr.107524.110 · PMID: 20644199 · PMCID: PMC2928508 \n16.  Manta: rapid detection of structural variants and indels for germline and cancer \nsequencing applications Xiaoyu Chen, Ole Schulz-Trieglaff, Richard Shaw, Bret Barnes, Felix \nSchlesinger, Morten Källberg, Anthony J Cox, Semyon Kruglyak, Christopher T Saunders \nBioinformatics (2015-12-08) https://doi.org/gf3ggb\n DOI: 10.1093/bioinformatics/btv710 · PMID: \n26647377 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\n17.  AnnotSV: an integrated tool for structural variations annotation Véronique \nGeoffroy, Yvan Herenger, Arnaud Kress, Corinne Stoetzel, Amélie Piton, Hélène Dollfus, Jean \nMuller Bioinformatics (2018-04-14) https://doi.org/gdcsh3 DOI: 10.1093/bioinformatics/bty304 · \nPMID: 29669011 \n18.  Accurate and efficient detection of gene fusions from RNA sequencing data \nSebastian Uhrig, Julia Ellermann, Tatjana Walther, Pauline Burkhardt, Martina Fröhlich, Barbara \nHutter, Umut H Toprak, Olaf Neumann, Albrecht Stenzinger, Claudia Scholl, … Benedikt Brors \nGenome Research (2021-01-13) https://doi.org/gjvdvp\n DOI: 10.1101/gr.257246.119 · PMID: \n33441414 · PMCID: PMC7919457 \n19.  STAR-Fusion: Fast and Accurate Fusion Transcript Detection from RNA-Seq Brian \nJ Haas, Alex Dobin, Nicolas Stransky, Bo Li, Xiao Yang, Timothy Tickle, Asma Bankapur, Carrie \nGanote, Thomas G Doak, Nathalie Pochet, … Aviv Regev Cold Spring Harbor Laboratory \n(2017-03-24) https://doi.org/gf5pc5\n DOI: 10.1101/120295 \n20.  annoFuse: an R Package to annotate, prioritize, and interactively explore putative \noncogenic RNA fusions Krutika S Gaonkar, Federico Marini, Komal S Rathi, Payal Jain, \nYuankun Zhu, Nicholas A Chimicles, Miguel A Brown, Ammar S Naqvi, Bo Zhang, Phillip B \nStorm, … Jo Lynne Rokita BMC Bioinformatics (2020-12) https://doi.org/gm84mh\n DOI: \n10.1186/s12859-020-03922-7 · PMID: 33317447 · PMCID: PMC7737294 \n21.  MSFragger: ultrafast and comprehensive peptide identification in mass \nspectrometry–based proteomics Andy T Kong, Felipe V Leprevost, Dmitry M Avtonomov, \nDattatreya Mellacheruvu, Alexey I Nesvizhskii Nature Methods (2017-04-10) \nhttps://doi.org/f9z6p7 DOI: 10.1038/nmeth.4256 · PMID: 28394336 · PMCID: PMC5409104 \n22.  Correcting systematic bias and instrument measurement drift with mzRefinery \nBryson C Gibbons, Matthew C Chambers, Matthew E Monroe, David L Tabb, Samuel H Payne \nBioinformatics (2015-08-04) https://doi.org/gb5g57\n DOI: 10.1093/bioinformatics/btv437 · PMID: \n26243018 · PMCID: PMC4653383 \n23.  MS-GF+ makes progress towards a universal database search tool for proteomics \nSangtae Kim, Pavel A Pevzner Nature Communications (2014-10-31) https://doi.org/ggkdq8 \nDOI: 10.1038/ncomms6277 · PMID: 25358478 · PMCID: PMC5036525 \n24.  Spectral probabilities of top-down tandem mass spectra Xiaowen Liu, Matthew W \nSegar, Shuai Cheng Li, Sangtae Kim BMC Genomics (2014-01) https://doi.org/gb3gzt DOI: \n10.1186/1471-2164-15-s1-s9 · PMID: 24564718 · PMCID: PMC4046700 \n25.  A probability-based approach for high-throughput protein phosphorylation \nanalysis and site localization Sean A Beausoleil, Judit Villén, Scott A Gerber, John Rush, \nSteven P Gygi Nature Biotechnology (2006-09-10) https://doi.org/dbwqf4 DOI: 10.1038/nbt1240 \n· PMID: 16964243 \n26.  Preprocessing, normalization and integration of the Illumina \nHumanMethylationEPIC array with minfi Jean-Philippe Fortin, Timothy J Triche Jr, Kasper D \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nHansen Bioinformatics (2016-11-29) https://doi.org/f9x7kd DOI: 10.1093/bioinformatics/btw691 · \nPMID: 28035024 · PMCID: PMC5408810 \n27.  Review of processing and analysis methods for DNA methylation array data CS \nWilhelm-Benartzi, DC Koestler, MR Karagas, JM Flanagan, BC Christensen, KT Kelsey, CJ \nMarsit, EA Houseman, R Brown British Journal of Cancer (2013-08-27) https://doi.org/gb9qvv \nDOI: 10.1038/bjc.2013.496 · PMID: 23982603 · PMCID: PMC3777004 \n28.  Impact of SNPs on methylation readouts by Illumina Infinium \nHumanMethylation450 BeadChip Array: implications for comparative population studies \nPatrycja Daca-Roszak, Aleksandra Pfeifer, Jadwiga Ż ebracka-Gala, Dagmara Rusinek, \nAleksandra Szybiń ska, Barbara Jarzą b, Michał Witt, Ewa Zię tkiewicz BMC Genomics (2015-11-\n25) https://doi.org/gb3h5r DOI: 10.1186/s12864-015-2202-0 · PMID: 26607064 · PMCID: \nPMC4659175 \n29.  DNA methylation-based classification of central nervous system tumours David \nCapper, David TW Jones, Martin Sill, Volker Hovestadt, Daniel Schrimpf, Dominik Sturm, \nChristian Koelsche, Felix Sahm, Lukas Chavez, David E Reuss, … Stefan M Pfister Nature \n(2018-03-14) https://doi.org/gc5t36\n DOI: 10.1038/nature26000 · PMID: 29539639 · PMCID: \nPMC6093218 \n30.  The Molecular Signatures Database Hallmark Gene Set Collection Arthur Liberzon, \nChet Birger, Helga Thorvaldsdóttir, Mahmoud Ghandi, Jill P Mesirov, Pablo Tamayo Cell \nSystems (2015-12) https://doi.org/gf78hq DOI: 10.1016/j.cels.2015.12.004 · PMID: 26771021 · \nPMCID: PMC4707969 \n31.  GSVA: gene set variation analysis for microarray and RNA-Seq data Sonja \nHänzelmann, Robert Castelo, Justin Guinney BMC Bioinformatics (2013-01-16) \nhttps://doi.org/gb8vx5 DOI: 10.1186/1471-2105-14-7 · PMID: 23323831 · PMCID: PMC3618321 \n32.  GISTIC2.0 facilitates sensitive and confident localization of the targets of focal \nsomatic copy-number alteration in human cancers Craig H Mermel, Steven E Schumacher, \nBarbara Hill, Matthew L Meyerson, Rameen Beroukhim, Gad Getz Genome Biology (2011-04-\n28) https://doi.org/10.1186/gb-2011-12-4-r41 DOI: 10.1186/gb-2011-12-4-r41 \n33.  deconstructSigs: delineating mutational processes in single tumors distinguishes \nDNA repair deficiencies and patterns of carcinoma evolution Rachel Rosenthal, Nicholas \nMcGranahan, Javier Herrero, Barry S Taylor, Charles Swanton Genome Biology (2016-02-22) \nhttps://doi.org/f8bdsq DOI: 10.1186/s13059-016-0893-4 · PMID: 26899170 · PMCID: \nPMC4762164 \n34.  Signatures of mutational processes in human cancer, Ludmil B Alexandrov, Serena \nNik-Zainal, David C Wedge, Samuel AJR Aparicio, Sam Behjati, Andrew V Biankin, … Michael \nR Stratton Nature (2013-08-14) https://doi.org/f22m2q DOI: 10.1038/nature12477 · PMID: \n23945592 · PMCID: PMC3776390 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\n35.  Tumor mutational burden standardization initiatives: Recommendations for \nconsistent tumor mutational burden assessment in clinical samples to guide \nimmunotherapy treatment decisions Albrecht Stenzinger, Jeffrey D Allen, Jörg Maas, Mark D \nStewart, Diana M Merino, Madison M Wempe, Manfred Dietel Genes, Chromosomes and \nCancer (2019-03-07) https://doi.org/ggc8pj DOI: 10.1002/gcc.22733 · PMID: 30664300 · \nPMCID: PMC6618007 \n36.  The 2021 WHO Classification of Tumors of the Central Nervous System: a \nsummary David N Louis, Arie Perry, Pieter Wesseling, Daniel J Brat, Ian A Cree, Dominique \nFigarella-Branger, Cynthia Hawkins, HK Ng, Stefan M Pfister, Guido Reifenberger, … David W \nEllison Neuro-Oncology (2021-06-29) https://doi.org/gmqhbf DOI: 10.1093/neuonc/noab106 · \nPMID: 34185076 · PMCID: PMC8328013 \n37.  Alterations in ALK/ROS1/NTRK/MET drive a group of infantile hemispheric \ngliomas Ana S Guerreiro Stucklin, Scott Ryall, Kohei Fukuoka, Michal Zapotocky, Alvaro \nLassaletta, Christopher Li, Taylor Bridge, Byungjin Kim, Anthony Arnoldo, Paul E Kowalski, … \nCynthia Hawkins Nature Communications (2019-09-25) https://doi.org/gh7bg8 DOI: \n10.1038/s41467-019-12187-5 · PMID: 31554817 · PMCID: PMC6761184 \n38.  Molecular subgrouping of atypical teratoid/rhabdoid tumors—a reinvestigation \nand current consensus Ben Ho, Pascal D Johann, Yura Grabovska, Mamy Jean De Dieu \nAndrianteranagna, Fupan Yao, Michael Frühwald, Martin Hasselblatt, Franck Bourdeaut, Daniel \nWilliamson, Annie Huang, Marcel Kool Neuro-Oncology (2019-12-31) https://doi.org/gn3kcm\n \nDOI: 10.1093/neuonc/noz235 · PMID: 31889194 · PMCID: PMC7229260 \n39.  A transcriptome-based classifier to determine molecular subtypes in \nmedulloblastoma Komal S Rathi, Sherjeel Arif, Mateusz Koptyra, Ammar S Naqvi, Deanne M \nTaylor, Phillip B Storm, Adam C Resnick, Jo Lynne Rokita, Pichai Raman PLOS Computational \nBiology (2020-10-29) https://doi.org/gm84kq DOI: 10.1371/journal.pcbi.1008263 · PMID: \n33119584 · PMCID: PMC7654754 \n40.  Genomic and Molecular Landscape of DNA Damage Repair Deficiency across The \nCancer Genome Atlas Theo A Knijnenburg, Linghua Wang, Michael T Zimmermann, Nyasha \nChambwe, Galen F Gao, Andrew D Cherniack, Huihui Fan, Hui Shen, Gregory P Way, Casey S \nGreene, … Armaz Mariamidze Cell Reports (2018-04) https://doi.org/gfspsc DOI: \n10.1016/j.celrep.2018.03.076 · PMID: 29617664 · PMCID: PMC5961503 \n41.  Accelerating Discovery of Functional Mutant Alleles in Cancer Matthew T Chang, \nTripti Shrestha Bhattarai, Alison M Schram, Craig M Bielski, Mark TA Donoghue, Philip \nJonsson, Debyani Chakravarty, Sarah Phillips, Cyriac Kandoth, Alexander Penson, … Barry S \nTaylor Cancer Discovery (2018-02-01) https://doi.org/gf9twp DOI: 10.1158/2159-8290.cd-17-\n0321 · PMID: 29247016 · PMCID: PMC5809279 \n42.  Identifying recurrent mutations in cancer reveals widespread lineage diversity and \nmutational specificity Matthew T Chang, Saurabh Asthana, Sizhi Paul Gao, Byron H Lee, \nJocelyn S Chapman, Cyriac Kandoth, JianJiong Gao, Nicholas D Socci, David B Solit, Adam B \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint \n\nOlshen, … Barry S Taylor Nature Biotechnology (2015-11-30) https://doi.org/gf7vxg DOI: \n10.1038/nbt.3391 · PMID: 26619011 · PMCID: PMC4744099 \n43.  The functional domains in p53 family proteins exhibit both common and distinct \nproperties KL Harms, X Chen Cell Death &amp; Differentiation (2006-03-17) \nhttps://doi.org/fwgrrt DOI: 10.1038/sj.cdd.4401904 · PMID: 16543939 \n44.  Gain of function mutations in p53 Dirk Dittmer, Sibani Pati, Gerard Zambetti, Shelley \nChu, Angelika K Teresky, Mary Moore, Cathy Finlay, Arnold J Levine Nature Genetics (1993-\n05) https://doi.org/crqst7 DOI: 10.1038/ng0593-42 · PMID: 8099841 \n45.  InheritedTP53Mutations and the Li–Fraumeni Syndrome Tanya Guha, David Malkin \nCold Spring Harbor Perspectives in Medicine (2017-03-07) https://doi.org/f9s4h3 DOI: \n10.1101/cshperspect.a026187 · PMID: 28270529 · PMCID: PMC5378014 \n46.  NGSCheckMate: software for validating sample identity in next-generation \nsequencing studies within and across data types Sejoon Lee, Soohyun Lee, Scott Ouellette, \nWoong-Yang Park, Eunjung A Lee, Peter J Park Nucleic Acids Research (2017-03-23) \nhttps://doi.org/f9xrq4 DOI: 10.1093/nar/gkx193 · PMID: 28369524 · PMCID: PMC5499645 \n47.  Somalier: rapid relatedness estimation for cancer and germline studies using \nefficient genome sketches Brent S Pedersen, Preetida J Bhetariya, Joe Brown, Stephanie N \nKravitz, Gabor Marth, Randy L Jensen, Mary P Bronner, Hunter R Underhill, Aaron R Quinlan \nGenome Medicine (2020-07-14) https://doi.org/gtsm62 DOI: 10.1186/s13073-020-00761-2 · \nPMID: 32664994 · PMCID: PMC7362544 \n48.  Barriers to accessing public cancer genomic data Katrina Learned, Ann Durbin, \nRobert Currie, Ellen Towle Kephart, Holly C Beale, Lauren M Sanders, Jacob Pfeil, Theodore C \nGoldstein, Sofie R Salama, David Haussler, … Isabel M Bjork Scientific Data (2019-06-20) \nhttps://doi.org/gjfrhz DOI: 10.1038/s41597-019-0096-4 · PMID: 31222016 · PMCID: \nPMC6586850 \n49.  dbGaP Study https://www.ncbi.nlm.nih.gov/projects/gap/cgi-\nbin/study.cgi?study_id=phs000218.v23.p8 \n50.  dbGaP Study https://www.ncbi.nlm.nih.gov/projects/gap/cgi-\nbin/study.cgi?study_id=phs001436.v1.p1 \n51.  dbGaP Study https://www.ncbi.nlm.nih.gov/projects/gap/cgi-\nbin/study.cgi?study_id=phs002517.v4.p2 \n52.  dbGaP Study https://www.ncbi.nlm.nih.gov/projects/gap/cgi-\nbin/study.cgi?study_id=phs000178.v11.p8 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted June 28, 2025. ; https://doi.org/10.1101/2024.07.09.599086doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}