Single-Cell Foundation Models: Architectures, Pretraining Strategies, and Benchmarking Challenges
preprint
OA: closed
CC-BY-4.0
Abstract
The recent proliferation of single-cell RNA sequencing (scRNA-seq) data has catalyzed the development of single-cell foundation models (scFMs) — large-scale deep learning architectures pretrained on millions of cells to learn generalizable representations of transcriptional programs. Unlike conventional single-cell analysis tools designed for specific tasks (e.g., clustering, differential expression, batch correction), scFMs such as scGPT, Geneformer, UCE, and scFoundation adopt a pretrain-then-finetune paradigm, enabling transfer learning across diverse tissues, species, and conditions. This review provides a systematic methodological survey of existing scFMs, focusing on three core pillars. First, we compare architectural choices, including transformer-based models with self-attention mechanisms, graph neural networks that leverage gene interaction priors, and variational autoencoders adapted for count-based expression data. Second, we analyze pretraining strategies: masked language modeling of gene tokens, contrastive learning across perturbed or multimodal inputs (e.g., CITE-seq), and generative reconstruction of expression values. Third, we critically evaluate current benchmarking practices, highlighting inconsistencies in data splitting (e.g., cell-type leakage), evaluation metrics that correlate poorly with downstream task performance, and the absence of standardized tests for robustness to batch effects and dropout noise. We conclude by identifying open challenges, including scaling laws for single-cell data, biological interpretability of learned embeddings, and the need for community-led benchmark suites. This review targets computational biologists and machine learning researchers seeking to understand, apply, or improve scFMs for purely in silico analyses.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00
- unpaywall
- last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0