Recovering Single-cell Heterogeneity Through Information-based Dimensionality Reduction
preprint
OA: closed
CC-BY-NC-ND-4.0
Abstract
Dimensionality reduction is crucial to summarizing the complex transcriptomic landscape of single cell datasets for downstream analyses. However, current dimensionality reduction approaches favor large cellular populations defined by many genes, at the expense of smaller and more subtly-defined populations. Here, we present surprisal component analysis (SCA), a technique that leverages the information-theoretic notion of surprisal for dimensionality reduction, and demonstrate its ability to improve the representation of clinically important populations that are indistinguishable using existing pipelines. For example, in cytotoxic T-cell data, SCA cleanly separates the gamma-delta and MAIT cell subpopulations, which are not detectable via PCA, ICA, scVI, or a wide array of specialized rare cell recovery tools. We also show that, when used instead of PCA, SCA improves downstream imputation to more accurately restore mRNA dropouts and recover important gene-gene relationships. SCA’s information-theoretic paradigm opens the door to more meaningful signal extraction, with broad applications to the study of complex biological tissues in health and disease.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00
- unpaywall
- last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-NC-ND-4.0