anndata: Annotated data

preprint OA: closed
📄 Open PDF View at publisher

Abstract

Summary anndata is a Python package for handling annotated data matrices in memory and on disk ( github.com/theislab/anndata ), positioned between pandas and xarray. anndata offers a broad range of computationally efficient features including, among others, sparse data support, lazy operations, and a PyTorch interface. Statement of need Generating insight from high-dimensional data matrices typically works through training models that annotate observations and variables via low-dimensional representations. In exploratory data analysis, this involves iterative training and analysis using original and learned annotations and task-associated representations. anndata offers a canonical data structure for book-keeping these, which is neither addressed by pandas (McKinney, 2010), nor xarray (Hoyer & Hamman, 2017), nor commonly-used modeling packages like scikit-learn (Pedregosa et al., 2011).

My notes (saved in your browser only)

Citation neighborhood (sparse)

Too few in-corpus citations on either side for a chart; here are the lists.

Cited by (2)

Cited by (2)

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00