Abstract
Generating realistic single-cell transcriptomic profiles from structured biological descriptions would enable controlled simulation, data augmentation, and hypothesis-driven cell-state creation—yet no existing method combines text–cell alignment with conditional generation. We present CLOP-DiT, a modular three-stage pipeline: (1) a contrastive aligner (CLOP) maps BiomedBERT text embeddings and scGPT cell embeddings into a shared 512-dimensional space; (2) a conditional Diffusion Transformer (DiT) generates scGPT-compatible latent states via flow matching, steered by a five-field biological template (cell type, tissue, organism, marker genes, disease); and (3) a frozen scGPT decoder maps latents to gene expression. Across 69 cell types from 80 GEO datasets (220,304 cells), a high-fidelity regime (CFG = 2.0) achieves 36.9% KNN accuracy (25× chance) and 81.0% steering, while a high-diversity regime (CFG = 1.0) reaches diversity ratio 0.93 at 80.7% steering. Conditioning field ablation and swap-label permutation tests confirm that marker genes are the dominant steering signal (steering accuracy drops from 99.8% to 62.4% when only metadata fields are retained). Key limitations are identified transparently: in-distribution per-gene variance structure is well preserved ( r = 0.98) but cross-dataset variance correlation drops to near zero, the discriminator AUC of 0.656 indicates residual distinguishability, and a pilot rare-cell augmentation study was negative. The modular architecture enables targeted remediation of each limitation without full retraining. CLOP-DiT establishes the feasibility of structured-metadata-conditioned single-cell generation and provides a composable framework for iterative improvement. Simple Summary CLOP-DiT is a computational pipeline that generates synthetic single-cell gene expression profiles from structured biological descriptions (cell type, tissue, organism, marker genes, and disease context). It first learns to align text descriptions with real cell data in a shared mathematical space, then uses a diffusion model to generate new cell states matching a given description. The generated cells capture correct cell-type identity and marker gene patterns, but do not yet reproduce the full cell-to-cell variability seen in real single-cell experiments. This work demonstrates that text-guided single-cell generation is feasible as a proof of concept, opening directions for future simulation and hypothesis-generation tools in biology.
Full text
3,272 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
Generating realistic single-cell transcriptomic profiles from structured biological descriptions would enable controlled simulation, data augmentation, and hypothesis-driven cell-state creation—yet no existing method combines text–cell alignment with conditional generation. We present CLOP-DiT, a modular three-stage pipeline: (1) a contrastive aligner (CLOP) maps BiomedBERT text embeddings and scGPT cell embeddings into a shared 512-dimensional space; (2) a conditional Diffusion Transformer (DiT) generates scGPT-compatible latent states via flow matching, steered by a five-field biological template (cell type, tissue, organism, marker genes, disease); and (3) a frozen scGPT decoder maps latents to gene expression. Across 69 cell types from 80 GEO datasets (220,304 cells), a high-fidelity regime (CFG = 2.0) achieves 36.9% KNN accuracy (25× chance) and 81.0% steering, while a high-diversity regime (CFG = 1.0) reaches diversity ratio 0.93 at 80.7% steering. Conditioning field ablation and swap-label permutation tests confirm that marker genes are the dominant steering signal (steering accuracy drops from 99.8% to 62.4% when only metadata fields are retained). Key limitations are identified transparently: in-distribution per-gene variance structure is well preserved (r = 0.98) but cross-dataset variance correlation drops to near zero, the discriminator AUC of 0.656 indicates residual distinguishability, and a pilot rare-cell augmentation study was negative. The modular architecture enables targeted remediation of each limitation without full retraining. CLOP-DiT establishes the feasibility of structured-metadata-conditioned single-cell generation and provides a composable framework for iterative improvement.
Simple Summary CLOP-DiT is a computational pipeline that generates synthetic single-cell gene expression profiles from structured biological descriptions (cell type, tissue, organism, marker genes, and disease context). It first learns to align text descriptions with real cell data in a shared mathematical space, then uses a diffusion model to generate new cell states matching a given description. The generated cells capture correct cell-type identity and marker gene patterns, but do not yet reproduce the full cell-to-cell variability seen in real single-cell experiments. This work demonstrates that text-guided single-cell generation is feasible as a proof of concept, opening directions for future simulation and hypothesis-generation tools in biology.
Competing Interest Statement
The authors have declared no competing interest.
Abbreviations
- AdaLN
- Adaptive Layer Normalization
- CFG
- Classifier-free guidance
- CLOP
- Contrastive Language-Omics Pretraining
- DE
- Differential expression
- DiT
- Diffusion Transformer
- DivR
- Diversity Ratio
- EMA
- Exponential Moving Average
- FD
- Fréchet Distance
- GEO
- Gene Expression Omnibus
- KNN
- k-Nearest Neighbor
- LinAcc
- Linear Classifier Accuracy
- logFC
- Log Fold Change
- LoRA
- Low-Rank Adaptation
- ODE
- Ordinary Differential Equation
- OOD
- Out-of-Distribution
- scGPT
- Single-cell Generative Pre-trained Transformer
- scRNA-seq
- Single-cell RNA sequencing
- SigLIP
- Sigmoid Loss for Language-Image Pre-training
- ZCA
- Zero-phase Component Analysis
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.