ClusterDE: a post-clustering differential expression (DE) method robust to false-positive inflation caused by double dipping

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract In typical single-cell RNA-seq (scRNA-seq) data analysis, a clustering algorithm is applied to find putative cell types as clusters, and then a statistical differential expression (DE) test is employed to identify the differentially expressed (DE) genes between the cell clusters. However, this common procedure uses the same data twice, an issue known as “double dipping”: the same data is used twice to define cell clusters as potential cell types and DE genes as potential cell-type marker genes, leading to false-positive cell-type marker genes even when the cell clusters are spurious. To overcome this challenge, we propose ClusterDE, a post-clustering DE method for controlling the false discovery rate (FDR) of identified DE genes regardless of clustering quality, which can work as an add-on to popular pipelines such as Seurat. The core idea of ClusterDE is to generate real-data-based synthetic null data containing only one cluster, as contrast to the real data, for evaluating the whole procedure of clustering followed by a DE test. Using comprehensive simulation and real data analysis, we show that ClusterDE has not only solid FDR control but also the ability to identify cell-type marker genes as top DE genes and distinguish them from housekeeping genes. ClusterDE is fast, transparent, and adaptive to a wide range of clustering algorithms and DE tests. Besides scRNA-seq data, ClusterDE is generally applicable to post-clustering DE analysis, including single-cell multi-omics data analysis.
Full text 11,439 characters · extracted from preprint-html · click to expand
ClusterDE: a post-clustering differential expression (DE) method robust to false-positive inflation caused by double dipping | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Brief Communication ClusterDE: a post-clustering differential expression (DE) method robust to false-positive inflation caused by double dipping Jingyi Jessica Li, Dongyuan Song, Kexin Li, Xinzhou Ge This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3211191/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract In typical single-cell RNA-seq (scRNA-seq) data analysis, a clustering algorithm is applied to find putative cell types as clusters, and then a statistical differential expression (DE) test is employed to identify the differentially expressed (DE) genes between the cell clusters. However, this common procedure uses the same data twice, an issue known as “double dipping”: the same data is used twice to define cell clusters as potential cell types and DE genes as potential cell-type marker genes, leading to false-positive cell-type marker genes even when the cell clusters are spurious. To overcome this challenge, we propose ClusterDE, a post-clustering DE method for controlling the false discovery rate (FDR) of identified DE genes regardless of clustering quality, which can work as an add-on to popular pipelines such as Seurat. The core idea of ClusterDE is to generate real-data-based synthetic null data containing only one cluster, as contrast to the real data, for evaluating the whole procedure of clustering followed by a DE test. Using comprehensive simulation and real data analysis, we show that ClusterDE has not only solid FDR control but also the ability to identify cell-type marker genes as top DE genes and distinguish them from housekeeping genes. ClusterDE is fast, transparent, and adaptive to a wide range of clustering algorithms and DE tests. Besides scRNA-seq data, ClusterDE is generally applicable to post-clustering DE analysis, including single-cell multi-omics data analysis. Biological sciences/Computational biology and bioinformatics/Statistical methods Biological sciences/Computational biology and bioinformatics/Software Full Text Additional Declarations There is NO Competing Interest. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3211191","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Brief Communication","associatedPublications":[],"authors":[{"id":223106582,"identity":"5361ba73-8016-4f41-8c4c-d8fcb21e4d12","order_by":0,"name":"Jingyi Jessica Li","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAyklEQVRIiWNgGAWjYBACPmaGBIaEHzYgNuMBENFASAsbSMvHnjQwh0gtIFUz2A6TooWd4Zk0D895uwaJ5AcHfjDYyG44QNhhadI8FreTGyTSDA72MKQZE6mF53Yyg0QOwwEehsOJRGphOwfWcvAPw3/itEjOYDtgB9JymIfhAFFaki0+9iQnsPE8MzgsY5BsPJOQFn7+M4k3En7Y2fOzJz98+KbCTraPkBYGBp4UCSCZ2AbmGBBUDgLshz8ASXui1I6CUTAKRsHIBADqYT4pEiQF2AAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0002-9288-5648","institution":"University of California, Los Angeles","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Jingyi","middleName":"Jessica","lastName":"Li","suffix":""},{"id":223106583,"identity":"8cd7ab36-6143-4833-8692-5c2d3d4b25b7","order_by":1,"name":"Dongyuan Song","email":"","orcid":"https://orcid.org/0000-0003-1114-1215","institution":"University of California, Los Angeles","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Dongyuan","middleName":"","lastName":"Song","suffix":""},{"id":223106584,"identity":"520411e3-aeca-4ce8-903a-f59dc9e46969","order_by":2,"name":"Kexin Li","email":"","orcid":"","institution":"University of California, Los Angeles","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Kexin","middleName":"","lastName":"Li","suffix":""},{"id":223106585,"identity":"9d32b43d-d774-47a0-8629-0c2fe0ec3105","order_by":3,"name":"Xinzhou Ge","email":"","orcid":"","institution":"University of California, Los Angeles","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xinzhou","middleName":"","lastName":"Ge","suffix":""}],"badges":[],"createdAt":"2023-07-28 00:15:35","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3211191/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3211191/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":47430213,"identity":"ee363f91-d5b8-4abb-bbba-c2885cb77f43","added_by":"auto","created_at":"2023-12-01 09:30:35","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":23213282,"visible":true,"origin":"","legend":"","description":"","filename":"ClusterDEapplication2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3211191/v1_covered_1f3fa0b9-29e7-4e9d-9134-cd030b42e66b.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"ClusterDE: a post-clustering differential expression (DE) method robust to false-positive inflation caused by double dipping","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-3211191/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3211191/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"In typical single-cell RNA-seq (scRNA-seq) data analysis, a clustering algorithm is applied to find putative cell types as clusters, and then a statistical differential expression (DE) test is employed to identify the differentially expressed (DE) genes between the cell clusters. However, this common procedure uses the same data twice, an issue known as “double dipping”: the same data is used twice to define cell clusters as potential cell types and DE genes as potential cell-type marker genes, leading to false-positive cell-type marker genes even when the cell clusters are spurious. To overcome this challenge, we propose ClusterDE, a post-clustering DE method for controlling the false discovery rate (FDR) of identified DE genes regardless of clustering quality, which can work as an add-on to popular pipelines such as Seurat. The core idea of ClusterDE is to generate real-data-based synthetic null data containing only one cluster, as contrast to the real data, for evaluating the whole procedure of clustering followed by a DE test. Using comprehensive simulation and real data analysis, we show that ClusterDE has not only solid FDR control but also the ability to identify cell-type marker genes as top DE genes and distinguish them from housekeeping genes. ClusterDE is fast, transparent, and adaptive to a wide range of clustering algorithms and DE tests. Besides scRNA-seq data, ClusterDE is generally applicable to post-clustering DE analysis, including single-cell multi-omics data analysis.","manuscriptTitle":"ClusterDE: a post-clustering differential expression (DE) method robust to false-positive inflation caused by double dipping","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-08-02 05:40:27","doi":"10.21203/rs.3.rs-3211191/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"4bc51a31-ecc6-4324-b237-96852e8ae8f2","owner":[],"postedDate":"August 2nd, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":23708426,"name":"Biological sciences/Computational biology and bioinformatics/Statistical methods"},{"id":23708427,"name":"Biological sciences/Computational biology and bioinformatics/Software"}],"tags":[],"updatedAt":"2023-12-01T09:21:18+00:00","versionOfRecord":[],"versionCreatedAt":"2023-08-02 05:40:27","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3211191","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3211191","identity":"rs-3211191","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00