Recovering undersampled single-cell transcriptomes with HyperCell

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Single-cell transcriptomic technology has now matured, allowing quantification of mRNA transcripts corresponding to tens of thousands of genes within a cell. However, still only a small fraction of these mRNA is captured and measured by today’s single-cell assays. There are likely hundreds of thousands of mRNA copies present within a typical human cell, yet these assays omit a majority of the transcripts that are actually present. This introduces technical noise, especially non-biological variability and excessive sparsity, which frustrates downstream analysis and potentially skews biological conclusions. To overcome these challenges, we here develop HyperCell, a probabilistic deep learning approach that explicitly models this undersampling to produce estimates of each cell’s original gene transcript abundances across the whole transcriptome. We demonstrate that our framework offers benefits in various mRNA modeling settings, by i) correctly differentiating between spurious sampling-induced and real biological zeros, outperforming existing approaches, ii) estimating the total mRNA content of cells across states to reduce contamination due to background transcripts, iii) reducing contamination due to background transcripts, and iv) helping to counteract biases that may appear during typical differential gene expression analyses using widespread normalization approaches. Our approach to correcting for the technical noise introduced by the single-cell experimental process brings us closer to studying biology, starting from the true transcriptome of cells.
Full text 11,476 characters · extracted from preprint-html · click to expand
Recovering undersampled single-cell transcriptomes with HyperCell | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Recovering undersampled single-cell transcriptomes with HyperCell Danilo Bzdok, Liam Hodgson, Smita Krishnaswamy This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6436032/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Single-cell transcriptomic technology has now matured, allowing quantification of mRNA transcripts corresponding to tens of thousands of genes within a cell. However, still only a small fraction of these mRNA is captured and measured by today’s single-cell assays. There are likely hundreds of thousands of mRNA copies present within a typical human cell, yet these assays omit a majority of the transcripts that are actually present. This introduces technical noise, especially non-biological variability and excessive sparsity, which frustrates downstream analysis and potentially skews biological conclusions. To overcome these challenges, we here develop HyperCell, a probabilistic deep learning approach that explicitly models this undersampling to produce estimates of each cell’s original gene transcript abundances across the whole transcriptome. We demonstrate that our framework offers benefits in various mRNA modeling settings, by i) correctly differentiating between spurious sampling-induced and real biological zeros, outperforming existing approaches, ii) estimating the total mRNA content of cells across states to reduce contamination due to background transcripts, iii) reducing contamination due to background transcripts, and iv) helping to counteract biases that may appear during typical differential gene expression analyses using widespread normalization approaches. Our approach to correcting for the technical noise introduced by the single-cell experimental process brings us closer to studying biology, starting from the true transcriptome of cells. Biological sciences/Computational biology and bioinformatics Health sciences/Health care Full Text Additional Declarations There is NO Competing Interest. Supplementary Files Supplementfinal.docx SOM Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6436032","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":446127313,"identity":"3f33e6c2-23e1-4285-a1fa-c0b6db54c13f","order_by":0,"name":"Danilo Bzdok","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA30lEQVRIiWNgGAWjYDCCA0D0AEyyMT4A8nn4iNKSANHCbADSwkaMFgaoFjYJkABBLXzHzx48kFBxx57veFta5dccOxk2BuaHj27g0SJ5Ji/hQMKZZ8ySZ44duy27LRnoMDZj4xw8WgwO5BgcSGw7zGZwI73ttuQ2ZqAWHjZpvFrOvwFr4QFpKZbcVk+ElhsQWyQMbqQdY/y47TBhLZI3gLYA/WIA9EuyNOO24zxszAT8wnc+x/jDB0iIGX78ua3anp+9+eFjfFpQADMPmCRWOQgw/iBF9SgYBaNgFIwYAACmwlLxilBtaAAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0003-3466-6620","institution":"McGill University \u0026 Mila Quebec AI Institute","correspondingAuthor":true,"prefix":"","firstName":"Danilo","middleName":"","lastName":"Bzdok","suffix":""},{"id":446127314,"identity":"29c0dd2a-7425-4728-8eec-c885a31abf16","order_by":1,"name":"Liam Hodgson","email":"","orcid":"https://orcid.org/0009-0001-8462-9863","institution":"McGill University","correspondingAuthor":false,"prefix":"","firstName":"Liam","middleName":"","lastName":"Hodgson","suffix":""},{"id":446127315,"identity":"98d3024e-194c-4a99-915b-2d08ada51793","order_by":2,"name":"Smita Krishnaswamy","email":"","orcid":"https://orcid.org/0000-0001-5823-1985","institution":"Yale University","correspondingAuthor":false,"prefix":"","firstName":"Smita","middleName":"","lastName":"Krishnaswamy","suffix":""}],"badges":[],"createdAt":"2025-04-12 18:15:09","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6436032/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6436032/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":106289406,"identity":"cbd0f511-0245-4958-897f-a3b01de3e500","added_by":"auto","created_at":"2026-04-07 07:32:19","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1451018,"visible":true,"origin":"","legend":"","description":"","filename":"Liammanuscript17april2025final.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6436032/v1_covered_7bfae9b0-c30b-4829-901d-ce2a2b7104d0.pdf"},{"id":81172021,"identity":"4b1c642d-d5b3-4e9c-afa1-b90cd648b2ca","added_by":"auto","created_at":"2025-04-23 05:33:17","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":1135896,"visible":true,"origin":"","legend":"SOM","description":"","filename":"Supplementfinal.docx","url":"https://assets-eu.researchsquare.com/files/rs-6436032/v1/b8bac9b59e24ba51b83dc440.docx"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Recovering undersampled single-cell transcriptomes with HyperCell","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-6436032/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6436032/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eSingle-cell transcriptomic technology has now matured, allowing quantification of mRNA transcripts corresponding to tens of thousands of genes within a cell. However, still only a small fraction of these mRNA is captured and measured by today\u0026rsquo;s single-cell assays. There are likely hundreds of thousands of mRNA copies present within a typical human cell, yet these assays omit a majority of the transcripts that are actually present. This introduces technical noise, especially non-biological variability and excessive sparsity, which frustrates downstream analysis and potentially skews biological conclusions. To overcome these challenges, we here develop HyperCell, a probabilistic deep learning approach that explicitly models this undersampling to produce estimates of each cell\u0026rsquo;s original gene transcript abundances across the whole transcriptome. We demonstrate that our framework offers benefits in various mRNA modeling settings, by i) correctly differentiating between spurious sampling-induced and real biological zeros, outperforming existing approaches, ii) estimating the total mRNA content of cells across states to reduce contamination due to background transcripts, iii) reducing contamination due to background transcripts, and iv) helping to counteract biases that may appear during typical differential gene expression analyses using widespread normalization approaches. Our approach to correcting for the technical noise introduced by the single-cell experimental process brings us closer to studying biology, starting from the true transcriptome of cells.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e","manuscriptTitle":"Recovering undersampled single-cell transcriptomes with HyperCell","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-23 05:33:13","doi":"10.21203/rs.3.rs-6436032/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"1880ba57-75bb-4cf3-91d8-dd0175522e0a","owner":[],"postedDate":"April 23rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":47484915,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":47484916,"name":"Health sciences/Health care"}],"tags":[],"updatedAt":"2026-04-07T07:31:59+00:00","versionOfRecord":[],"versionCreatedAt":"2025-04-23 05:33:13","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6436032","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6436032","identity":"rs-6436032","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00