Pan-cancer gene set discovery via scRNA-seq for optimal deep learning based downstream tasks | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Pan-cancer gene set discovery via scRNA-seq for optimal deep learning based downstream tasks Jongseong Jang, JONG HYUN KIM This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4785088/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 02 Dec, 2025 Read the published version in Scientific Reports → Version 1 posted 10 You are reading this latest preprint version Abstract The application of machine learning to transcriptomics data has led to significant advances in cancer research. However, the high dimensionality and complexity of RNA sequencing (RNA-seq) data pose significant challenges in pan-cancer studies. This study hypothesizes that gene sets derived from single-cell RNA sequencing (scRNA-seq) data will outperform those selected using bulk RNA-seq in pan-cancer downstream tasks. We analyzed scRNA-seq data from 181 tumor biopsies across 13 cancer types. High-dimensional weighted gene co-expression network analysis (hdWGCNA) was performed to identify relevant gene sets, which were further refined using XGBoost for feature selection. These gene sets were applied to downstream tasks using TCGA pan-cancer RNA-seq data and compared to six reference gene sets and oncogenes from OncoKB evaluated with deep learning models, including multilayer perceptrons (MLPs) and graph neural networks (GNNs). The XGBoost-refined hdWGCNA gene set demonstrated higher performance in most tasks, including tumor mutation burden assessment, microsatellite instability classification, mutation prediction, cancer subtyping, and grading. In particular, genes such as DPM1, BAD, and FKBP4 emerged as important pan-cancer biomarkers, with DPM1 consistently significant across tasks. This study presents a robust approach for feature selection in cancer genomics by integrating scRNA-seq data and advanced analysis techniques. Biological sciences/Cancer/Cancer genomics Biological sciences/Biological techniques/Bioinformatics Physical sciences/Mathematics and computing/Information technology Physical sciences/Engineering/Biomedical engineering Figures Figure 1 Figure 2 Figure 3 Full Text Additional Declarations No competing interests reported. Supplementary Files SupplementaryData.xlsx Cite Share Download PDF Status: Published Journal Publication published 02 Dec, 2025 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Revision requested 28 Jul, 2025 Reviews received at journal 01 Jun, 2025 Reviews received at journal 25 May, 2025 Reviewers agreed at journal 22 May, 2025 Reviewers agreed at journal 22 May, 2025 Reviewers invited by journal 14 Aug, 2024 Editor assigned by journal 14 Aug, 2024 Editor invited by journal 05 Aug, 2024 Submission checks completed at journal 30 Jul, 2024 First submitted to journal 22 Jul, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4785088","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":342326082,"identity":"680e9aaa-4b8d-4238-a21b-4fb90605ffa8","order_by":0,"name":"Jongseong Jang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5UlEQVRIiWNgGAWjYBACxgaGBCBlIwMXMQDiA0RoSeMhXgsUHEbVghcwNzA8/Fzw6zyPfPThwx8+1NyTN2dgfkjIYcnSM/tu8xieS0uTnHGs2HBnA5sBIS0J0rw9QC09PGbMPGwJCUD1BLUk/+btOQfUwv/5859/IC3sHwhpSZPm+XGAR56Hh0GasQ2khYeALc0Mada8Dck8BjxsZpK9fQmGGw7zFODVYtjek3yb54+dnHwP8+MPP74lyBscb9/8Aa+WZp4EBsY2BiQvM+NTDwTyDOxAtX+AjAYCKkfBKBgFo2DkAgB1ZEpnbWAY/gAAAABJRU5ErkJggg==","orcid":"","institution":"LG AI Research","correspondingAuthor":true,"prefix":"","firstName":"Jongseong","middleName":"","lastName":"Jang","suffix":""},{"id":342326085,"identity":"51fa12f0-00c9-431e-874c-615e01ad8a72","order_by":1,"name":"JONG HYUN KIM","email":"","orcid":"","institution":"LG AI Research","correspondingAuthor":false,"prefix":"","firstName":"JONG","middleName":"HYUN","lastName":"KIM","suffix":""}],"badges":[],"createdAt":"2024-07-23 02:32:42","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4785088/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4785088/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-025-27296-z","type":"published","date":"2025-12-02T15:56:51+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":63290656,"identity":"da6a9aa1-3e2e-44f8-897c-284f147487e4","added_by":"auto","created_at":"2024-08-26 14:15:46","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":2577216,"visible":true,"origin":"","legend":"\u003cp\u003eSchematic overview of the study workflow. a. Each cancer type is indicated along with the number of samples analyzed: Uveal melanoma (UM, n=9), esophageal adenocarcinoma (EA, n=3), non-small cell lung cancer (NSCLC, n=8), breast invasive carcinoma (BRCA, n=8), hepatocellular carcinoma (HCC, n=18), intrahepatic cholangiocarcinoma (ICC, n=11), renal cell carcinoma (RCC, n=3), pancreatic adenocarcinoma (PAAD, n=24), colorectal cancer (CRC, n=27), ovarian cancer (OC, n=4), squamous cell carcinoma (SCC, n=4), cutaneous melanoma (CM, n=52), basal cell carcinoma (BCC, n=10). b. Proportional distribution of the 13 cancer types within the scRNA-seq dataset, highlighting the diversity and representation of different tumor types used in the study. The dataset includes 317,111 tumor immune cells classified into 25 distinct cell types. c. Schematic UMAP visualization of single-cell expression profiles, which were prepared for hdWGCNA. d. Schematic of the co-expression gene network derived from hdWGCNA. The genes comprising this biological module were used for various downstream analyses through the target gene process. e. (e) Application of selected co-expressed gene modules to various pan-cancer downstream tasks. These tasks include mutation prediction, cancer subtyping, tumor mutation burden (TMB) assessment, microsatellite instability (MSI) classification, cancer subtyping, and grading.\u003c/p\u003e","description":"","filename":"Fig1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4785088/v1/296718c3af962f6e8b95e04c.jpg"},{"id":63290658,"identity":"76efe108-139c-465a-872b-c1a6aafec512","added_by":"auto","created_at":"2024-08-26 14:15:47","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":3743412,"visible":true,"origin":"","legend":"\u003cp\u003eComprehensive gene co-expression networks reveal key functional modules in the tumor immune microenvironment. a. UMAP embeddings display the distribution of 25 immune cell types across 181 tumor biopsy samples, visualizing all 317,111 cells. Each color represents a distinct immune cell type. b. UMAP representation of the gene co-expression network. A total of 11 modules were identified. Nodes indicate individual genes, and edges signify co-expression relationships between genes and hub genes within modules. Node sizes are proportional to their kME (eigengene-based connectivity) values. Colors denote different co-expression modules. c. Visualisation of hub gene networks for each spatial co-expression module. The 25 highest-ranked hub genes based on kME are presented. In the network, nodes represent genes, while edges indicate co-expression links. d. Heatmap summarizing GO pathway enrichment analysis for each module. Each row represents a biological process, and columns correspond to modules, with the color scale indicating Z-score values.\u003c/p\u003e","description":"","filename":"Fig2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4785088/v1/d27a01c3927ccd5dc25079de.jpg"},{"id":63290657,"identity":"d66eba02-4694-4e1b-b9f2-ce58ae38f5dc","added_by":"auto","created_at":"2024-08-26 14:15:46","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":2595860,"visible":true,"origin":"","legend":"\u003cp\u003eFeature importance analysis and overlap of key genes across downstream tasks. a-d. Feature importance scores for the top 10 genes were extracted for each downstream task, with figures representing: a. mutation prediction (MUT), b. microsatellite instability (MSI), c. tumor mutation burden (TMB), and d. cancer grading (GRAD). Each panel displays the importance scores for the respective genes, highlighting those with the highest relevance to each task. e. The upset plot illustrates the overlap of key genes among the different downstream tasks. Genes consistently ranking high in feature importance across multiple tasks are shown, with red-bolded genes indicating the most significant overlaps. The left bar plot presents the number of features selected for each task.\u003c/p\u003e","description":"","filename":"Fig3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4785088/v1/d8e22dc719156c87cd9ecfc7.jpg"},{"id":97723880,"identity":"5d0a61e5-c3d0-4e96-8cfe-49f7186acdac","added_by":"auto","created_at":"2025-12-08 16:09:05","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":9223657,"visible":true,"origin":"","legend":"","description":"","filename":"Manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4785088/v1_covered_0e7bd964-0286-4300-a47a-6f5e8d02bb31.pdf"},{"id":63290660,"identity":"f9f34ed6-c16d-4122-adcb-0b9bd4ac105a","added_by":"auto","created_at":"2024-08-26 14:15:47","extension":"xlsx","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":219032,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryData.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4785088/v1/e6b7c537e572521a88be0cb7.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Pan-cancer gene set discovery via scRNA-seq for optimal deep learning based downstream tasks","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-4785088/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4785088/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The application of machine learning to transcriptomics data has led to significant advances in cancer research. However, the high dimensionality and complexity of RNA sequencing (RNA-seq) data pose significant challenges in pan-cancer studies. This study hypothesizes that gene sets derived from single-cell RNA sequencing (scRNA-seq) data will outperform those selected using bulk RNA-seq in pan-cancer downstream tasks. We analyzed scRNA-seq data from 181 tumor biopsies across 13 cancer types. High-dimensional weighted gene co-expression network analysis (hdWGCNA) was performed to identify relevant gene sets, which were further refined using XGBoost for feature selection. These gene sets were applied to downstream tasks using TCGA pan-cancer RNA-seq data and compared to six reference gene sets and oncogenes from OncoKB evaluated with deep learning models, including multilayer perceptrons (MLPs) and graph neural networks (GNNs). The XGBoost-refined hdWGCNA gene set demonstrated higher performance in most tasks, including tumor mutation burden assessment, microsatellite instability classification, mutation prediction, cancer subtyping, and grading. In particular, genes such as DPM1, BAD, and FKBP4 emerged as important pan-cancer biomarkers, with DPM1 consistently significant across tasks. This study presents a robust approach for feature selection in cancer genomics by integrating scRNA-seq data and advanced analysis techniques.","manuscriptTitle":"Pan-cancer gene set discovery via scRNA-seq for optimal deep learning based downstream tasks","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-08-26 14:15:42","doi":"10.21203/rs.3.rs-4785088/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-07-28T08:19:05+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-06-01T17:50:31+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-05-25T19:24:07+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"247598535017580976699274319144270999492","date":"2025-05-22T22:31:25+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"299255662978156535708581267792219842174","date":"2025-05-22T13:20:01+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-08-14T20:25:29+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-08-14T20:08:11+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2024-08-05T08:04:08+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-07-30T04:38:06+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2024-07-23T02:31:20+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"79ba3d77-478d-4391-913b-6eec2515915e","owner":[],"postedDate":"August 26th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":36266925,"name":"Biological sciences/Cancer/Cancer genomics"},{"id":36266926,"name":"Biological sciences/Biological techniques/Bioinformatics"},{"id":36266927,"name":"Physical sciences/Mathematics and computing/Information technology"},{"id":36266928,"name":"Physical sciences/Engineering/Biomedical engineering"}],"tags":[],"updatedAt":"2025-12-08T16:02:04+00:00","versionOfRecord":{"articleIdentity":"rs-4785088","link":"https://doi.org/10.1038/s41598-025-27296-z","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2025-12-02 15:56:51","publishedOnDateReadable":"December 2nd, 2025"},"versionCreatedAt":"2024-08-26 14:15:42","video":"","vorDoi":"10.1038/s41598-025-27296-z","vorDoiUrl":"https://doi.org/10.1038/s41598-025-27296-z","workflowStages":[]},"version":"v1","identity":"rs-4785088","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4785088","identity":"rs-4785088","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.