scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks

preprint OA: closed CC-BY-NC-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Large single-cell RNA-sequencing (scRNA-seq) datasets offer unprecedented biological insights but pose major computational challenges for visualisation and analysis. Existing subsampling methods can improve efficiency yet may not guarantee downstream machine and deep learning (ML/DL) performance. Here, we propose scValue, a conceptually distinct approach that ranks individual cells by “data value” based on out-of-bag estimates from a random forest. scValue prioritises higher-value cells and allocates more representation to cell types displaying greater value variability, preserving essential biological signals in subsamples. We benchmarked scValue in automatic cell-type annotation tasks on four large datasets (human peripheral blood mononuclear cells, mouse brain cells, human cross-tissue atlas, and mouse aging cell atlas), paired with distinct ML/DL models (scANVI, scPoli, CellTypist, and ACTINN). Our method consistently outperformed existing subsampling methods, closely matching full-data performance in all annotation tasks. Furthermore, in two additional case studies of label transfer learning (via CellTypist) and cross-study label harmonisation (via CellHint), scValue better preserved T-cell annotations across human gut-colon datasets and more accurately reproduced T-cell subtype relationships in a human spleen dataset. Finally, using 16 public datasets ranging from tens of thousands to millions of cells, we compared subsampling quality of scValue and its counterparts on computational time, Gini coefficient, and Hausdorff distance. The method demonstrated fast execution, balanced cell-type representation, and near-random subsampling distributional characteristics. Overall, scValue provides an efficient and accurate solution for subsampling large scRNA-seq data for ML/DL tasks. It is implemented as an open-source Python package installable via pip, with source code available at https://github.com/LHBCB/scvalue .
Full text 2,176 characters · extracted from oa-doi-fallback · click to expand
Abstract Large single-cell RNA-sequencing (scRNA-seq) datasets offer unprecedented biological insights but pose major computational challenges for visualisation and analysis. Existing subsampling methods can improve efficiency yet may not guarantee downstream machine and deep learning (ML/DL) performance. Here, we propose scValue, a conceptually distinct approach that ranks individual cells by “data value” based on out-of-bag estimates from a random forest. scValue prioritises higher-value cells and allocates more representation to cell types displaying greater value variability, preserving essential biological signals in subsamples. We benchmarked scValue in automatic cell-type annotation tasks on four large datasets (human peripheral blood mononuclear cells, mouse brain cells, human cross-tissue atlas, and mouse aging cell atlas), paired with distinct ML/DL models (scANVI, scPoli, CellTypist, and ACTINN). Our method consistently outperformed existing subsampling methods, closely matching full-data performance in all annotation tasks. Furthermore, in two additional case studies of label transfer learning (via CellTypist) and cross-study label harmonisation (via CellHint), scValue better preserved T-cell annotations across human gut-colon datasets and more accurately reproduced T-cell subtype relationships in a human spleen dataset. Finally, using 16 public datasets ranging from tens of thousands to millions of cells, we compared subsampling quality of scValue and its counterparts on computational time, Gini coefficient, and Hausdorff distance. The method demonstrated fast execution, balanced cell-type representation, and near-random subsampling distributional characteristics. Overall, scValue provides an efficient and accurate solution for subsampling large scRNA-seq data for ML/DL tasks. It is implemented as an open-source Python package installable via pip, with source code available at https://github.com/LHBCB/scvalue. Competing Interest Statement The authors have declared no competing interest. Footnotes Article's content reorganised; figures and tables revised; two additional cell-type annotation experiments added; discussion added.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-02T02:00:03.124865+00:00
License: CC-BY-NC-4.0