Abstract
The spatial organization of cells within tissues is critical for understanding biological function and disease, and spatial transcriptomics enables genome-wide mapping of this organization. Numerous computational methods aim to identify spatial domains, yet their performance is often evaluated on limited datasets, leading to conflicting conclusions. Here, we present an explanatory benchmark of 26 spatial domain detection methods across 63 tissue sections from six spatial transcriptomics technologies, supplemented by over 1,000 semi-synthetic datasets that systematically vary resolution, gene panel size, and tissue architecture. By jointly analyzing real and semi-synthetic data across this broad parameter space, our benchmark uncovers systematic performance differences and sources of variability that are obscured in standard evaluations. Although most spatial methods outperform non-spatial baselines, their performance depends strongly on data resolution and cellular heterogeneity. To enable systematic analysis beyond individual methods, we introduce a modular, plug-and-play benchmarking framework that facilitates method refinement and component exchange. Using this framework, an ablation study of neural network-based approaches shows that the choice of preprocessing and clustering often has a larger impact on performance than ar-chitectural novelty alone. Together, these results provide a principled foundation for informed method selection and offer guidance for the development of robust and scalable spatial domain detection tools as spatial transcriptomics technologies continue to advance.
Full text
1,703 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
The spatial organization of cells within tissues is critical for understanding biological function and disease, and spatial transcriptomics enables genome-wide mapping of this organization. Numerous computational methods aim to identify spatial domains, yet their performance is often evaluated on limited datasets, leading to conflicting conclusions. Here, we present an explanatory benchmark of 26 spatial domain detection methods across 63 tissue sections from six spatial transcriptomics technologies, supplemented by over 1,000 semi-synthetic datasets that systematically vary resolution, gene panel size, and tissue architecture. By jointly analyzing real and semi-synthetic data across this broad parameter space, our benchmark uncovers systematic performance differences and sources of variability that are obscured in standard evaluations. Although most spatial methods outperform non-spatial baselines, their performance depends strongly on data resolution and cellular heterogeneity. To enable systematic analysis beyond individual methods, we introduce a modular, plug-and-play benchmarking framework that facilitates method refinement and component exchange. Using this framework, an ablation study of neural network-based approaches shows that the choice of preprocessing and clustering often has a larger impact on performance than ar-chitectural novelty alone. Together, these results provide a principled foundation for informed method selection and offer guidance for the development of robust and scalable spatial domain detection tools as spatial transcriptomics technologies continue to advance.
Competing Interest Statement
The authors have declared no competing interest.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.