Abstract
Deep learning-based segmentation models can accelerate the analysis of high-throughput microscopy data by automatically identifying and classifying cells in images. However, the datasets needed to train these models are typically assembled via laborious hand-annotation. This limits their scale and diversity, which in turn limits model performance. We present Cell-APP (Cellular Annotation and Perception Pipeline), a tool that automates the annotation of high-quality training data for transmitted-light (TL) cell segmentation. Cell-APP uses two inputs—paired TL and nuclear fluorescence images—and operates in two main steps. First, it extracts each cell’s location from the nuclear fluorescence channel and provides these locations to promptable deep learning models to generate cell masks. Then, it classifies each cell as mitotic or non-mitotic based on nuclear features. Together, these masks and classifications form the basis for cell segmentation training data. By training vision-transformer-based models on Cell-APP-generated datasets, we demonstrate how Cell-APP enables the creation of both cell line-specific and multi-cell line segmentation models. Cell-APP thus empowers laboratories to tailor cell segmentation models to their needs, and outlines a scalable path to creating general models for the research community. Significance Statement Deep learning-based cell segmentation models are typically trained on manually annotated datasets. Manual annotation limits dataset scalability and, consequently, the ability of trained models to generalize across cell types. Cell-APP automates mask generation and cell classification to rapidly create large, custom training datasets. It uses Meta AI’s SAM for mask generation and extracts information from fluorescence signals for classification. By reducing the need for hand-annotation, Cell-APP lowers the cost associated with building cell line-specific and generalist segmentation models. It may accelerate high-throughput image analysis and democratize model development across research labs.
Full text
2,744 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
Deep learning-based segmentation models can accelerate the analysis of high-throughput microscopy data by automatically identifying and classifying cells in images. However, the datasets needed to train these models are typically assembled via laborious hand-annotation. This limits their scale and diversity, which in turn limits model performance. We present Cell-APP (Cellular Annotation and Perception Pipeline), a tool that automates the annotation of high-quality training data for transmitted-light (TL) cell segmentation. Cell-APP uses two inputs—paired TL and nuclear fluorescence images—and operates in two main steps. First, it extracts each cell’s location from the nuclear fluorescence channel and provides these locations to promptable deep learning models to generate cell masks. Then, it classifies each cell as mitotic or non-mitotic based on nuclear features. Together, these masks and classifications form the basis for cell segmentation training data. By training vision-transformer-based models on Cell-APP-generated datasets, we demonstrate how Cell-APP enables the creation of both cell line-specific and multi-cell line segmentation models. Cell-APP thus empowers laboratories to tailor cell segmentation models to their needs, and outlines a scalable path to creating general models for the research community.
Significance Statement
Deep learning-based cell segmentation models are typically trained on manually annotated datasets. Manual annotation limits dataset scalability and, consequently, the ability of trained models to generalize across cell types.
Cell-APP automates mask generation and cell classification to rapidly create large, custom training datasets. It uses Meta AI’s SAM for mask generation and extracts information from fluorescence signals for classification.
By reducing the need for hand-annotation, Cell-APP lowers the cost associated with building cell line-specific and generalist segmentation models. It may accelerate high-throughput image analysis and democratize model development across research labs.
Competing Interest Statement
The authors have declared no competing interest.
Footnotes
Entirely redesigned figures; evaluation of the method and resulting models on hand-annotations (Figure 2b; Figure S2b); application of the method on two additional cell lines: HT1080 and RPE-1; in-depth analysis of the model's failure modes (Figure S5); investigation of how dataset class imbalance influences model performance (Figure 3d; Figure S5); the development of a general model that performantly segments HeLa, U2OS, HT1080, and RPE-1 cells (Figure 3e) and generalizes to unseen cell lines (Figure 4c); investigation on how dataset size influences model performance (Figure 3f).
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.