oCELLoc: Automated Cell Type Assignment in Transcriptomics Data Using Reference Filtering

preprint OA: closed CC-BY-NC-ND-4.0
AI-generated deep summary by claude@2026-07, 2026-07-05 · read from full text

The paper presents oCELLoc, an automated cell-type assignment method for single-cell RNA-seq and spatial transcriptomics that addresses how prediction accuracy depends on the reference cell types used. Using pseudobulk gene expression from ST or scRNA-seq together with a large reference atlas, it applies regularized regression with cross-validation to select a limited subset of relevant reference cell types tailored to each new sample. Across toy datasets, additional scRNA-seq data, and 2,144 Visium samples from diverse tissues and conditions, filtered reference cell types improved the biological meaningfulness of downstream predictions. The study’s main limitation is that evaluation is framed around the method’s performance on those datasets rather than establishing guarantees for every possible tissue, reference atlas composition, or experimental setting. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Interpreting single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) data requires accurate cell-type prediction, which strongly depends on the quality of the reference used. However, prediction accuracy is highly dependent on reference quality; missing relevant cell types or including irrelevant ones can substantially impair performance. To address this challenge, we developed oCELLoc, a regression-based method that selects the most appropriate reference cell types from a large atlas and tailors them to each new sample. oCELLoc takes pseudobulk gene expression from ST or scRNA-seq data together with a broad reference matrix and uses regularized regression with cross-validation to identify a limited number of essential cell types. We applied oCELLoc to toy datasets, scRNA-seq data, and 2,144 Visium samples across diverse tissues and conditions, demonstrating that using the filtered cell types leads to more biologically meaningful downstream predictions. oCELLoc is available as an R package on GitHub and CRAN.
Full text 1,127 characters · extracted from oa-doi-fallback · click to expand
Abstract Interpreting single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics (ST) data requires accurate cell-type prediction, which strongly depends on the quality of the reference used. However, prediction accuracy is highly dependent on reference quality; missing relevant cell types or including irrelevant ones can substantially impair performance. To address this challenge, we developed oCELLoc, a regression-based method that selects the most appropriate reference cell types from a large atlas and tailors them to each new sample. oCELLoc takes pseudobulk gene expression from ST or scRNA-seq data together with a broad reference matrix and uses regularized regression with cross-validation to identify a limited number of essential cell types. We applied oCELLoc to toy datasets, scRNA-seq data, and 2,144 Visium samples across diverse tissues and conditions, demonstrating that using the filtered cell types leads to more biologically meaningful downstream predictions. oCELLoc is available as an R package on GitHub and CRAN. Competing Interest Statement The authors have declared no competing interest.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-02T02:00:03.124865+00:00
License: CC-BY-NC-ND-4.0