CLOP-DiT: Structured-Metadata-Conditioned Single-Cell Latent Generation via Contrastive Language-Omics Pretraining and Diffusion Transformers

preprint OA: closed
Full text JSON View at publisher

Abstract

Generating realistic single-cell transcriptomic profiles from structured biological descriptions would enable controlled simulation, data augmentation, and hypothesis-driven cell-state creation—yet no existing method combines text–cell alignment with conditional generation. We present CLOP-DiT, a modular three-stage pipeline: (1) a contrastive aligner (CLOP) maps BiomedBERT text embeddings and scGPT cell embeddings into a shared 512-dimensional space; (2) a conditional Diffusion Transformer (DiT) generates scGPT-compatible latent states via flow matching, steered by a five-field biological template (cell type, tissue, organism, marker genes, disease); and (3) a frozen scGPT decoder maps latents to gene expression. Across 69 cell types from 80 GEO datasets (220,304 cells), a high-fidelity regime (CFG = 2.0) achieves 36.9% KNN accuracy (25× chance) and 81.0% steering, while a high-diversity regime (CFG = 1.0) reaches diversity ratio 0.93 at 80.7% steering. Conditioning field ablation and swap-label permutation tests confirm that marker genes are the dominant steering signal (steering accuracy drops from 99.8% to 62.4% when only metadata fields are retained). Key limitations are identified transparently: in-distribution per-gene variance structure is well preserved ( r = 0.98) but cross-dataset variance correlation drops to near zero, the discriminator AUC of 0.656 indicates residual distinguishability, and a pilot rare-cell augmentation study was negative. The modular architecture enables targeted remediation of each limitation without full retraining. CLOP-DiT establishes the feasibility of structured-metadata-conditioned single-cell generation and provides a composable framework for iterative improvement. Simple Summary CLOP-DiT is a computational pipeline that generates synthetic single-cell gene expression profiles from structured biological descriptions (cell type, tissue, organism, marker genes, and disease context). It first learns to align text descriptions with real cell data in a shared mathematical space, then uses a diffusion model to generate new cell states matching a given description. The generated cells capture correct cell-type identity and marker gene patterns, but do not yet reproduce the full cell-to-cell variability seen in real single-cell experiments. This work demonstrates that text-guided single-cell generation is feasible as a proof of concept, opening directions for future simulation and hypothesis-generation tools in biology.
Full text 3,272 characters · extracted from oa-doi-fallback · click to expand
Abstract Generating realistic single-cell transcriptomic profiles from structured biological descriptions would enable controlled simulation, data augmentation, and hypothesis-driven cell-state creation—yet no existing method combines text–cell alignment with conditional generation. We present CLOP-DiT, a modular three-stage pipeline: (1) a contrastive aligner (CLOP) maps BiomedBERT text embeddings and scGPT cell embeddings into a shared 512-dimensional space; (2) a conditional Diffusion Transformer (DiT) generates scGPT-compatible latent states via flow matching, steered by a five-field biological template (cell type, tissue, organism, marker genes, disease); and (3) a frozen scGPT decoder maps latents to gene expression. Across 69 cell types from 80 GEO datasets (220,304 cells), a high-fidelity regime (CFG = 2.0) achieves 36.9% KNN accuracy (25× chance) and 81.0% steering, while a high-diversity regime (CFG = 1.0) reaches diversity ratio 0.93 at 80.7% steering. Conditioning field ablation and swap-label permutation tests confirm that marker genes are the dominant steering signal (steering accuracy drops from 99.8% to 62.4% when only metadata fields are retained). Key limitations are identified transparently: in-distribution per-gene variance structure is well preserved (r = 0.98) but cross-dataset variance correlation drops to near zero, the discriminator AUC of 0.656 indicates residual distinguishability, and a pilot rare-cell augmentation study was negative. The modular architecture enables targeted remediation of each limitation without full retraining. CLOP-DiT establishes the feasibility of structured-metadata-conditioned single-cell generation and provides a composable framework for iterative improvement. Simple Summary CLOP-DiT is a computational pipeline that generates synthetic single-cell gene expression profiles from structured biological descriptions (cell type, tissue, organism, marker genes, and disease context). It first learns to align text descriptions with real cell data in a shared mathematical space, then uses a diffusion model to generate new cell states matching a given description. The generated cells capture correct cell-type identity and marker gene patterns, but do not yet reproduce the full cell-to-cell variability seen in real single-cell experiments. This work demonstrates that text-guided single-cell generation is feasible as a proof of concept, opening directions for future simulation and hypothesis-generation tools in biology. Competing Interest Statement The authors have declared no competing interest. Abbreviations - AdaLN - Adaptive Layer Normalization - CFG - Classifier-free guidance - CLOP - Contrastive Language-Omics Pretraining - DE - Differential expression - DiT - Diffusion Transformer - DivR - Diversity Ratio - EMA - Exponential Moving Average - FD - Fréchet Distance - GEO - Gene Expression Omnibus - KNN - k-Nearest Neighbor - LinAcc - Linear Classifier Accuracy - logFC - Log Fold Change - LoRA - Low-Rank Adaptation - ODE - Ordinary Differential Equation - OOD - Out-of-Distribution - scGPT - Single-cell Generative Pre-trained Transformer - scRNA-seq - Single-cell RNA sequencing - SigLIP - Sigmoid Loss for Language-Image Pre-training - ZCA - Zero-phase Component Analysis

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00