Key
The Human Cell Atlas (HCA) is an international, open-science consortium initiative that aspires to map all human cell types with single-cell resolution comparable in scope to the Human Genome Project [ 13 ]. Through single-cell or single-nucleus molecular profiling and spatial analysis, the HCA aims to characterize human cells at the single-cell level based on distinctive gene expression patterns, physiological states, development trajectories, and spatial locations, thereby generating a comprehensive reference map of the human body [ 13 ].
The HCA serves as a fundamental resource for understanding human health, disease mechanisms, regenerative biology, drug discovery, drug efficacy and toxicities, monitoring, and diagnostics [ 13 ]. Its translational relevance became evident during the COVID-19 pandemic, particularly in understanding severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection, which initiates at the single-cell level [ 77 ]. This approach proved essential for identifying targeted cells, viral entry sites, transmission, risk factors, epidemiology, and pathogenesis [ 77 ].
By leveraging its existing healthy single-cell reference data, the HCA community pooled millions of single-cell profiles to investigate SARS-CoV-2 pathology, comparing these with COVID-19 patient data to create a new COVID-19 atlas [ 67 , 77 ]. These efforts revealed that the SARS-CoV-2 viral entry receptor gene, angiotensin-converting enzyme 2 ( ACE2) , is highly expressed in type II alveolar epithelial cells, corneal epithelial cells, and two nasal-specific epithelial cell subtypes, identifying key portals of infection [ 77 ]. It was also detected at the uterine-placental interface and in gastrointestinal enterocytes, suggesting potential mother-to-baby and oral-faecal transmission routes, respectively. Furthermore, single-cell analysis discovered rare ACE2 -expressing cells in multiple organs, such as the heart, olfactory system, kidneys, and brain, which might be overlooked in bulk RNA profiling. HCA researchers also found that ACE2 expression increases with age, which is higher in men, and is elevated in the airway cells of smokers, providing insights into the epidemiological factors associated with COVID-19 susceptibility [ 77 ]. Moreover, single-cell analysis has contributed to drug discovery by identifying neutralizing antibodies from recovered patients and advancing vaccine research through monitoring immune cell dynamics during disease progression and post-vaccination [ 67 ].
Additionally, HCA fosters global collaboration by enabling scientists worldwide to contribute locally generated single-cell data while promoting open data sharing through the publicly accessible HCA Data Portal, a cloud-based platform that supports open data access and integrative analysis [ 78 ]. It advances scientific endeavour and addresses diverse ethical and legal challenges, such as human tissue sampling, data protection, and privacy considerations for other large-scale biomedical research, through the development of a harmonized, international, and interoperable ethics toolkit [ 78 ].
The Human BioMolecular Atlas Program (HuBMAP), funded by the National Institutes of Health (NIH), aims to develop an open-access framework of spatially resolved multi-omics maps (e.g., genomic, epigenomic, transcriptomic, proteomic, and metabolomic) of healthy human cells at the single-cell level from a diverse group of adults for biomedical research [ 79 ]. In contrast to the HCA, which primarily focuses on defining cellular profiles and their biological roles across the human body [ 13 ], HuBMAP specializes in high-resolution three-dimensional (3D) tissue mapping to create a spatial atlas. Central to this effort is delineating tissue architecture and identifying functional tissue units (FTUs), which provide spatial context for interpreting cellular organization and function within intact tissue microenvironments [ 80 ].
To construct detailed spatial maps of human tissues, the HuBMAP employs a two-step methodology [ 79 ]. In the first stage, single-cell omics assays, such as single-cell transcriptomics and chromatin accessibility assays, are utilized to generate genome-wide gene expression profiles, reveal the molecular states of individual cells, and uncover their regulatory mechanisms across distinct cell types. In the second stage, spatially resolved techniques, including fluorescence microscopy, sequential fluorescence in situ hybridization (seqFISH), imaging mass spectrometry, and imaging mass cytometry, are applied to capture biomolecule localization within intact tissues. This integrated approach reveals important biological details such as protein localization, post-translational modifications (e.g., phosphorylation), and microenvironment states [ 79 ].
To support this comprehensive mapping, HuBMAP employs approximately 40 different analytical technologies and has developed advanced tools for 3D tissue mapping and spatial data integration [ 80 ]. These include subcellular spatial mapping via imaging mass spectrometry and co-detection by indexing (CODEX), multi-scale data harmonization and standardization through Common Coordinate Framework (CCF), automated annotation using the Azimuth platform for scRNA-seq analysis, and interactive visualization using Vitessce. During the production phase, new spatial mapping technologies, such as HiFi-Slide sequencing and CosMx, are continuously expanded and integrated [ 80 ]. Notably, the program has pushed the development of new technologies and innovations, facilitating the study of cell interactions and complex tissue organizations.
HuBMAP provides biological insights into disease mechanisms and potential therapeutic targets by enabling the analysis of cellular interactions and spatial relationships within tissues [ 81 ]. For instance, in the studies of children with bronchopulmonary dysplasia (BPD), HuBMAP data and analytical tools have been used to compare the distance distribution of parenchymal cells and immune cells, as well as the cellularity of specific lung regions between healthy and BPD individuals [ 81 ]. Similarly, one of HuBMAP’s intestinal atlases discovered the cellular diversity of the intestine, spatial organizations of gut immune cells, and gene regulation and differentiation pathways, highlighting key clinical connections [ 82 ]. The study found that individuals with a high body mass index (BMI) exhibited an increase in pro-inflammatory M1 macrophages. At the same time, those with a history of hypertension showed a decrease in endothelial and CD8 + T cells. These findings may serve as a potential early sign of gastrointestinal disease and suggest potential associations with diseases such as colon cancer [ 82 ].
To encourage an open and collaborative team science approach, HuBMAP actively collaborates with other initiatives, including the HCA, Human Protein Atlas, and organ-specific projects (e.g., brain, lung, kidney, and genitourinary NIH-sponsored mapping consortia) [ 79 ]. One notable outcome of the collaboration between the NIH and the HCA was the formation of the Human Reference Atlas (HRA) Working Group, which has built a shared, interactive 3D atlas of the human body that provides a standardized framework for integrating multi-scale biological data [ 81 ]. To further support open data sharing, the HuBMAP Portal represents a comprehensive public gateway to all HuBMAP data, including raw experimental outputs, processed data, downstream analyses, and extensive metadata [ 80 ]. In addition to data accessibility, HuBMAP contributes to community education by publishing over 215 experimental protocols, more than 20 standard operating procedures, and a wide range of public talks, demonstrations, and other videos to support knowledge dissemination and reproducibility in the broader biomedical research community [ 80 ].
In contrast to previous atlases that predominantly relied on tissues from different donors, Tabula Sapiens serves as a groundbreaking and large-scale single-cell transcriptomic atlas that profiles multiple human tissues from the same individuals [ 83 ]. In this atlas, 17 tissues were collected from one donor, 14 from a second, five from two additional donors, and smaller numbers of tissues from 11 other donors. Instead of using isolated nuclei, Tabula Sapiens worked with live cells containing all mRNA, including both spliced and unspliced transcripts, facilitating the studies of alternative splicing. To understand the association between scRNA-seq data and traditional histology, hematoxylin and eosin (H&E) stained tissue sections were classified based on morphology, providing a reference for estimating relative cell abundance and assessing spatial heterogeneity. These comparisons revealed broad concordance between histology and transcriptomics, highlighting discrepancies in cell-type proportions, allowing more accurate estimations of true cell abundances while addressing dissociation-induced biases inherent in single-cell RNA-seq methodologies [ 83 ].
Through systematic cross-tissue comparisons using Smart-seq and 10 × droplet-based sequencing technologies, this atlas, developed by the HCA Consortium researchers, offers valuable insights into both conserved and tissue-specific characteristics, as well as disease associations in specific cell types or states [ 84 ]. An example of cellular specialization can be seen in macrophages, which exhibited tissue-specific gene expression patterns: while spleen macrophages expressed high levels of CD5L, macrophages in solid tissues such as the skin and uterus showed elevated epiregulin expression [ 83 ]. The researchers also identified shared T-cell lineages across organs and tissue-specific hypermutation rates in B cells [ 83 ]. Beyond its contributions to basic research, Tabula Sapiens served as a comprehensive multi-tissue single-cell reference atlas for cell-free RNA study, enabling researchers to analyze cell-free RNA signals and identify cellular sources correlating with disease [ 85 ]. By linking cfRNA profiles to their cellular sources, this resource facilitates more precise interpretations of circulating RNA biomarkers, thereby advancing mechanistic understanding of human physiology and the development of novel diagnostic tools.
Recently, the Tabula Sapiens 2.0 consortium significantly expanded its dataset by adding samples from nine additional donors and four new tissues. This has led to vital discoveries into transcription factor (TF) activity, cellular senescence patterns, and sex-specific gene expression across tissues [ 86 ]. The study identified universal active TFs involved in stress response, proliferation, and metabolism (e.g., JUN , FOS , and STAT1 ) and cell-type-specific TFs such as FOXP3 , which plays a role in T cell regulation. In cellular senescence, this expanded atlas revealed heterogeneity in the proportion of senescent cells across tissues, with the highest burden in the eye, bladder, and tongue and the lowest in the heart, muscle, and ovary. These senescent cells displayed complex and diverse phenotypes, with distinct patterns in metabolic reprogramming, cellular stress, immune responses, and inflammation. In addition, Tabula Sapiens 2.0 discovered sex-biased gene expression at the single-cell level, varying by cell type and tissue. Notably, HLA-DPB1 expression was enriched in female samples, potentially contributing to increased susceptibility to immune-mediated diseases, especially autoimmune diseases. These findings underscore the potential impact of sex-specific gene expression in influencing disease susceptibilities and physiological traits. Furthermore, the consortium developed ChatTS, a web-based application powered by a large language model for medical record search. ChatTS achieved an 83.2% agreement and provided more informative answers than manually curated data, improving time efficiency in data retrieval [ 86 ].
To enhance the utility of single-cell atlases, future developments should prioritize the integration of diverse datasets, improved methodologies for capturing cellular heterogeneity, and the incorporation of machine learning techniques for data analysis. This will enable researchers to derive more accurate insights into cellular dynamics and disease mechanisms, ultimately advancing the field of precision medicine. Furthermore, a focus on enhancing cross-species comparisons and integrating multi-omics data will be crucial for a more comprehensive understanding of cellular diversity and its implications for health and disease. Future efforts should emphasize the inclusion of underrepresented populations in single-cell atlases to address health disparities and improve the applicability of findings across diverse demographics. The continuous evolution of single-cell atlases is essential for advancing our understanding of cellular diversity and its implications for health, disease mechanisms, and therapeutic strategies. In summary, the comparative analysis of single-cell atlases reveals significant insights into cellular diversity, disease mechanisms, and the potential for precision medicine advancements. As the field of single-cell genomics progresses, integrating diverse datasets and improved methodologies will be essential for enhancing our understanding of cellular dynamics and advancing precision medicine.
Scope
Based on the single-cell atlas examined in this study (Tables 1 , 2 , 3 and 4 ), the biological focus of these resources can be broadly categorized into four classes: holistic (whole-organism), cross-tissue (multi-tissue or multi-organ without full-body coverage), cross-cancer, and cell-type/tissue-specific atlases. Holistic atlases provide a broader perspective on the whole organism, allowing comparisons across multiple tissues and organ systems for a more representative understanding of biological systems. Notable examples include the HCA [ 13 , 17 ] and MCA [ 14 ]. Among the 36 atlases reviewed, only two are classified as holistic single-cell atlases: the human Ensemble Cell Atlas (hECA) [ 18 ] and the Single Cell Atlas (SCA) [ 19 ] (Table 1 ). Table 1 Summary of the holistic single-cell atlases ( n = 2) Biological focus Size Species Conditions Modalities Integration strategy Key contributions Reference Human cells (integumentary, endocrine, urinary, cardiovascular, lymphatic, nervous, respiratory, digestive, muscular, reproductive, and skeletal systems) 1,093,299 cells, 116 published datasets covering 38 human organs and 11 systems Human Healthy adult and fetal Transcriptomic, spatial transcriptomics Integration with the extended Harmony method on cells of the same organ and batch-effect correction separately for each organ. The corrected expression matrix was derived from the Harmony embedding with inverse transformation Developed a unified informatics framework for cell-centric data organization. Introduced three novel atlas applications: targeted data retrieval using logical expressions within “in data” cell sorting, multi-perspective visualization of biological entities through “quantitative portraiture,” and customizable reference generation for automatic annotations [ 18 ] Multi-omics analysis of human tissues 8 omics types, 125 adults and fetal tissues, 3.8 M cells post-QC scRNA-seq, 773 K scATAC-seq cells, 209 K scImmune profiling cells, 2.2 M cells for CyTOF, 192.9 M cells for flow cytometry Human Healthy adult and foetal tissues Transcriptomic, chromatin accessibility, spatial transcriptomics, single-cell immune profiling, whole genome sequencing (scRNA-seq, scATAC-seq, scImmune profiling, spatial transcriptomics, CyTOF, flow cytometry, genotyping, RNA-seq) Multi-omics integration across 125 adult and foetal tissues with omics-layer comparisons Established a comprehensive human cell encyclopedia integrating multi-omics data [ 19 ] scRNA-seq , Single-cell ribonucleic acid sequencing; scATAC-seq , Single-cell assay for transposase-accessible chromatin sequencing; scImmune , Single-cell immune; CyTOF , Cytometry by time-of-flight Table 2 Summary of cross-tissue single-cell atlases ( n = 4) Biological focus Size Species Organ/tissue involved Conditions Modalities Integration strategy Key contributions Reference Cross-tissue immune cells (primary and secondary lymphoid organs, gut, lung, blood, liver, thymus, skeletal muscle and omentum) 16 tissues, 12 adult donors, 360,000 cells Human Lungs, liver, omentum, transverse colon, duodenum, jejunum, caecum, ileum, sigmoid colon, blood, skeletal muscle, bone marrow, mesenteric lymph node, thymus, lung-draining lymph nodes, spleen Deceased adults Transcriptomic, VDJ sequencing Integration with BBKNN and scVI Developed a cross-tissue immune cell atlas and introduced CellTypist, a machine learning-based cell type annotation tool [ 20 ] Developing immune system across organs 7 studies, 25 subjects, 221 samples, 908 k cells Human Prenatal hematopoietic (yolk sac, liver, and bone marrow), lymphoid (thymus, spleen, and lymph node), and non-lymphoid peripheral organs (skin, kidney, and gut) Post-conception weeks 4–17 Transcriptomic, spatial transcriptomic, T/B cell receptor sequencing Integration with scVI, 5’ vs 3’ sequencing and donors as batches Described acquisition of immune cell characteristics over developmental time. Discovered system-wide development of immune cells contrary to previous understanding of fetal hematopoiesis. Identified cellular processes conserved in immune cell types across organs [ 21 ] Endoderm-derived organoids across tissues 55 datasets, 218 samples, 800 k cells Human Small and large intestine, stomach, liver, lung, biliary system, pancreas, prostate, salivary glands Various organoid differentiation protocols Transcriptomic 3000 HVGs. Integration with scPoli, sample as batch Described how variations in media composition affect cell differentiation and maturation. Identified limitations of organoid models [ 23 ] Fibroblast lineage across tissues Healthy mouse: 28 datasets, 16 tissues, 121 k cells; Diseased/perturbed mouse: 17 datasets, 16 tissues, 100 k cells; Diseased human: 5 datasets, 3 tissues, 10 k cells Human, mouse (separate atlases) Artery, lymph node, pancreas, skin, muscle, tendon, omentum, adipose, bone, heart, intestine, mesentery, lung, liver, spleen, joint Healthy and 11 diseased/perturbed states (separate atlases for healthy and diseased/perturbed) Transcriptomic 2000 HVGs. Separate integration for healthy and diseased data with Harmony; dataset as batch Identified universal and tissue-specific fibroblasts, and their developmental relationships. Compared fibroblast transcriptional states across species and associated cell subtypes with diseases [ 22 ] VDJ , Variable–diversity–joining; BBKNN , Batch balanced k nearest neighbors; scVI , Single-cell variational inference; HVGs , Highly variable genes; scPoli , Single-cell population level integration Table 3 Summary of cross-cancer single-cell atlases ( n = 6) Biological focus Size Species Conditions Modalities Integration strategy Key contributions Reference Natural killer cells 61 datasets, 575 donors, 90 k cells Human Healthy (peripheral blood), 7 cancer types (SARC, PAAD, NSCLC, BRCA, SKCM, PRAD, GBM) Transcriptomic Integration of healthy data with scANVI. Reference mapping of tumor cells with scArches. Batch covariate not reported Characterized healthy NK cell states and GRNs associated with NK cell fates. Associated tumor NK cell subtypes to therapy response and survival [ 30 ] Natural killer cells 71 datasets, 776 subjects, 1223 samples, 160 k cells Human Healthy, 24 cancer types (HNSCC, THCA, LC, HCC, ICC, NB, PRAD, UCEC, MM, AML, ALL, CLL, NPC, ESCA, BRCA, GC, PACA, RC, CRC, OV, FTC, MELA, BCC, SCC) Transcriptomic HVG selection per cancer type. Integration with BBKNN, donor as batch Described NK cells across cancer types. Associated NK subtypes with prognosis and therapy response [ 29 ] Single-cell metabolic heterogeneity across tumors types 10 scRNA-seq datasets, 6 tumor types, 197 tumor samples and 99 normal samples, 764432 cells Human 6 tumor types (BRCA, CRC, LUAD, PAAD, PRAD and STAD) and normal tissue Transcriptomic, spatial transcriptomic 2000 HVGs. Integration with Seurat (v4.3.0). NMF was used to identify MMPs across multiple cancer types Discovered a multilayered landscape of metabolic heterogeneity across cancer, identifying MMPs, key regulators, and tumor-immune interactions [ 31 ] T cells 27 datasets, 324 donors, 486 samples, 308 k cells Human Healthy (9 tissues), 16 cancer types (AML, GBM, LGG, BRCA, CRC, HNSC, HCC, iCCA, NSCLC, DLBCL, FL, OV, PDAC, BCC, SKCM, STAD) Transcriptomic 1000 HVGs. Integration with Seurat rPCA, sample as batch Identified known and new T cell states. Associated cell states with cancer genomic, pathologic, and clinical features, including therapy response [ 28 ] Tumor immune cells 13 studies, 217 subjects, 317 k cells Human 13 cancer types (NSCLC, CM, CRC, BCC, BC, SCC, EA, RCC, HCC, PDAC, UM, ICC, OC) Transcriptomic, spatial transcriptomic Removal of dataset-specific genes. Integration with Seurat CCA, dataset as batch Stratified tumors across cancer types based on immune cell composition. Analyzed immune cell type and state colocalization patterns in tumors by combining single-cell and spatial transcriptomic data [ 27 ] Tumor-infiltrating myeloid cells 10 studies, 194 subjects, 338 samples, 138 k cells Human 15 cancer types (LYM, THCA, LUNG, HCC, PAAD, UCEC, MYE, NPC, ESCA, BRCA, STAD, KIDNEY, CRC, OV-FTC, MEL) Transcriptomic 2000 HVGs. Integration with Scanorama, 5' vs 3' sequencing and single-cell platform as batches Found tumor-infiltrating myeloid cells across cancer types to be diverse and cancer type-dependent, possibly affecting responsiveness to immunotherapy. Linked tumor myeloid cell composition to genomic and transcriptomic features of tumors [ 26 ] scRNA-seq , Single-cell ribonucleic acid sequencing; SARC , Sarcoma; PAAD , Pancreatic adenocarcinoma; NSCLC , non-small cell lung cancer; BRCA/BC , Breast cancer; SKCM , Skin cutaneous melanoma; PRAD , Prostate adenocarcinoma; GBM , Glioblastoma; HNSCC/HNSC , Head and neck squamous cell carcinoma; THCA , Thyroid cancer; LC/LUNG , Lung cancer; HCC , Hepatocellular carcinoma; ICC/iCCA , Intrahepatic cholangiocarcinoma; NB , Neuroblastoma; UCEC , Uterine corpus endometrial carcinoma; MM/MYE , Multiple myeloma; AML , Acute myeloid leukemia; ALL , Acute lymphoblastic leukemia; CLL , Chronic lymphocytic leukemia; NPC , Nasopharyngeal carcinoma; ESCA , Esophageal carcinoma; GC , Gastric cancer; PACA , Pancreatic cancer; RC , Rectal cancer; CRC , Colorectal cancer; OV/OC , Ovarian cancer; OV-FTC , Ovarian and fallopian tube cancer; MELA/MEL , Melanoma; BCC , Basal cell carcinoma; SCC , Squamous cell carcinoma; LUAD , Lung adenocarcinoma; STAD , Stomach adenocarcinoma; LGG , Low-grade glioma; DLBCL , Diffuse large B-cell lymphoma; FL , Follicular lymphoma; CM , Cutaneous melanoma; EA , Endometrial adenocarcinoma; RCC , Renal cell carcinoma; PDAC , Pancreatic ductal adenocarcinoma; UM , Uveal melanoma; LYM , Lymphoma; KIDNEY , Kidney cancer; scANVI , Single-cell annotation using variational inference; scArches – Single-cell architectural surgery; HVGs , Highly variable genes; BBKNN , Batch balanced k nearest neighbors; NMF , Non-negative matrix factorization; MMPs , Metabolic meta-programs; rPCA , Reciprocal principal component analysis; CCA , Canonical correlation analysis; NK , Natural killer; GRNs , Gene-regulatory networks Table 4 Single-cell atlases of the central nervous system ( n = 6) Biological focus Size Species Conditions Modalities Integration strategy Key contributions Reference Brain 70 studies, 6577 samples, 11,3 M cells Human Healthy (adult and fetal) and disease Transcriptomic Integration of adult, fetal, tumor and organoid data separately with scANVI, sample as batch Developed an automated hierarchical annotation method. Discovered previously unidentified cell types [ 34 ] Brain Human: 19 studies, 1.5 M cells; Mouse: 9 studies, 2.3 M cells Human, mouse (separate atlases) Range of developmental stages Transcriptomic Integration with Seurat CCA, dataset as a batch Described cell populations across brain subregions and developmental stages [ 33 ] Glioblastoma 26 studies, 240 subjects, 1,136 k cells Human Glioblastoma Transcriptomic, spatial transcriptomic 5000 HVGs. Integration with scANVI, study as batch Associated gene expression with clinical and diagnostic metadata. Elucidated tumor architecture and cell–cell interactions based on integration with spatial transcriptomic data [ 39 ] Human brain vasculature 117 samples, 68 human fetuses and adult patients, 606380 cells Human Healthy (fetal and adult) and disease (adult) Transcriptomic, spatial transcriptomic rPCA-based batch correction. Separate integration for fetal, adult/control, and pathological endothelial cell datasets; dataset as batch Identified molecular characteristics of human brain cell types and their variations by brain developmental stage and pathology, revealing organizational principles of ECs and PVCs that form the human brain vasculature [ 40 ] Hypothalamus 18 datasets, 126 samples, 385 k cells Mouse Healthy, perturbed Transcriptomic HVG selection, integration with scVI, dataset as batch Identified previously unknown rare cell types. Described cell types and transcriptomic patterns in the hypothalamus related to feeding behavior [ 41 ] Neural organoid 36 datasets, 396 samples, 1,77 M cells Human 26 organoid differentiation protocols Transcriptomic 3000 HVGs. Integration with scPoli, sample as batch Identify cell types that are under-represented in neural organoids generated with currently available protocols. Compare transcriptomic profiles of organoids to those of cell types in the developing brain [ 42 ] scANVI , Single-cell annotation using variational inference; CCA , Canonical correlation analysis; HVGs , Highly variable genes; rPCA , Reciprocal principal component analysis; scVI , Single-cell variational inference; scPoli , Single-cell population level integration; ECs , Endothelial cells; PVCs , Perivascular cells
Summary of the holistic single-cell atlases ( n = 2)
scRNA-seq , Single-cell ribonucleic acid sequencing; scATAC-seq , Single-cell assay for transposase-accessible chromatin sequencing; scImmune , Single-cell immune; CyTOF , Cytometry by time-of-flight
Summary of cross-tissue single-cell atlases ( n = 4)
Prenatal hematopoietic (yolk sac, liver, and bone marrow), lymphoid (thymus,
spleen, and lymph node), and non-lymphoid peripheral organs (skin, kidney, and gut)
VDJ , Variable–diversity–joining; BBKNN , Batch balanced k nearest neighbors; scVI , Single-cell variational inference; HVGs , Highly variable genes; scPoli , Single-cell population level integration
Summary of cross-cancer single-cell atlases ( n = 6)
scRNA-seq , Single-cell ribonucleic acid sequencing; SARC , Sarcoma; PAAD , Pancreatic adenocarcinoma; NSCLC , non-small cell lung cancer; BRCA/BC , Breast cancer; SKCM , Skin cutaneous melanoma; PRAD , Prostate adenocarcinoma; GBM , Glioblastoma; HNSCC/HNSC , Head and neck squamous cell carcinoma; THCA , Thyroid cancer; LC/LUNG , Lung cancer; HCC , Hepatocellular carcinoma; ICC/iCCA , Intrahepatic cholangiocarcinoma; NB , Neuroblastoma; UCEC , Uterine corpus endometrial carcinoma; MM/MYE , Multiple myeloma; AML , Acute myeloid leukemia; ALL , Acute lymphoblastic leukemia; CLL , Chronic lymphocytic leukemia; NPC , Nasopharyngeal carcinoma; ESCA , Esophageal carcinoma; GC , Gastric cancer; PACA , Pancreatic cancer; RC , Rectal cancer; CRC , Colorectal cancer; OV/OC , Ovarian cancer; OV-FTC , Ovarian and fallopian tube cancer; MELA/MEL , Melanoma; BCC , Basal cell carcinoma; SCC , Squamous cell carcinoma; LUAD , Lung adenocarcinoma; STAD , Stomach adenocarcinoma; LGG , Low-grade glioma; DLBCL , Diffuse large B-cell lymphoma; FL , Follicular lymphoma; CM , Cutaneous melanoma; EA , Endometrial adenocarcinoma; RCC , Renal cell carcinoma; PDAC , Pancreatic ductal adenocarcinoma; UM , Uveal melanoma; LYM , Lymphoma; KIDNEY , Kidney cancer; scANVI , Single-cell annotation using variational inference; scArches – Single-cell architectural surgery; HVGs , Highly variable genes; BBKNN , Batch balanced k nearest neighbors; NMF , Non-negative matrix factorization; MMPs , Metabolic meta-programs; rPCA , Reciprocal principal component analysis; CCA , Canonical correlation analysis; NK , Natural killer; GRNs , Gene-regulatory networks
Single-cell atlases of the central nervous system ( n = 6)
scANVI , Single-cell annotation using variational inference; CCA , Canonical correlation analysis; HVGs , Highly variable genes; rPCA , Reciprocal principal component analysis; scVI , Single-cell variational inference; scPoli , Single-cell population level integration; ECs , Endothelial cells; PVCs , Perivascular cells
The hECA is a comprehensive single-cell reference that integrates over 1.09 million annotated human cells from 116 datasets, covering 38 organs and 11 physiological systems [ 18 ]. Distinguished from the traditional dataset-centric single-cell atlases, hECA employs a unique cell-centric assembly framework, enabling flexible and customizable data integration, retrieval, and organization. This approach conceptually positions hECA as a virtual human framework, facilitating cross-organ analyses and in silico computational experiments while expanding opportunities for biomedical research, particularly in disease studies and virtual drug experiments [ 18 ].
Similarly, Pan et al. [ 19 ] introduced the SCA as a single-cell multi-omics reference integrating data from 125 healthy adult and fetal tissues spanning nearly all human tissues and organs. This atlas incorporates eight omics platforms, including scRNA-seq, single-cell ATAC sequencing, single-cell immune profiling, mass cytometry, flow cytometry, spatial transcriptomics, bulk RNA sequencing, and whole-genome sequencing, to provide a multidimensional characterization of human cellular and molecular heterogeneity. By offering a holistic characterization of molecular phenotypic variations and cellular complexity across multiple tissues and omics layers, SCA enables comprehensive cross-tissue analyses, supports advanced developmental biology research, and contributes to developing precision medicine and personalized treatments [ 19 ].
While holistic atlases aim to cover nearly all human tissues or organs, cross-tissue atlases provide a more focused perspective by profiling a subset of cell types across multiple tissues or organs without encompassing the entire organism. Among the 36 atlases, four atlases in the dataset fall into this category [ 20 – 23 ] (Table 2 ).
For instance, instead of characterizing all cell types, Domínguez Conde et al. [ 20 ] focused on describing the tissue-specific features and clonal architecture of immune cells within a subset of tissues, while Suo et al. [ 21 ] similarly reconstructed the developing human immune system by profiling immune cells across nine prenatal tissues. In another study, Buechler et al. [ 22 ] characterize fibroblast heterogeneity across multiple tissues, highlighting inter-tissue variation in fibroblast states rather than focusing on a single organ or tissue type.
Additionally, distinct from traditional single-cell atlases, the Human Endoderm Organoid Cell Atlas (HEOCA) focuses on organoids rather than directly profiling primary tissue samples, integrating data from nine endoderm-derived tissues [ 23 ]. Although lineage-specific, incorporating multiple tissue types categorizes it as a cross-tissue atlas within the endodermal lineage [ 23 ]. Cross-tissue atlases enable comparative analysis of single-cell data across various tissues from the same individual, allowing researchers to control for confounding variables such as age, sex, and sampling background. This approach facilitates the identification of conserved cell states, tissue-specific gene expression patterns, and cell heterogeneity among tissues [ 20 , 24 ]. While this approach has revealed conserved cell features, tissue-specific adaptations, and disease-relevant patterns, challenges in cross-tissue single-cell data generation and integration must be addressed to ensure robust biological insights [ 25 ].
In addition to cross-tissue atlases, several cross-cancer studies have focused on tumor-associated cell populations within the tumor microenvironment (TME), including tumor-infiltrating myeloid cells [ 26 ], immune cells [ 27 ], T cells [ 28 ], and natural killer (NK) cells [ 29 , 30 ] (Table 3 ). Although these studies primarily examine specific cell types within the cancer context, they analyze multiple cancer types originating from different tissues. This approach enables intratumor and intertumoral comparisons, enabling the identification of shared cellular states as well as cancer-specific characteristics and functional features. Such insights are crucial for developing personalized therapy and effective patient stratification strategies [ 27 , 31 ].
In contrast to holistic and cross-tissue atlases, the majority of single-cell resources are cell-type or tissue-specific, providing high-resolution characterizations of particular cell types or individual organs (Tables 4 , 5 , 6 , 7 , 8 and 9 ). These atlases offer a comprehensive reference framework for organ-specific biology and disease modelling. Notable examples include atlases of the kidney [ 32 ], brain [ 33 , 34 ], liver [ 35 ], heart [ 36 ], and lung [ 37 , 38 ]. For instance, the Human Lung Cell Atlas offers a detailed overview of all cell types within the healthy human lung, serving as a foundational resource for understanding pulmonary physiology and the cellular alterations associated with respiratory diseases [ 37 ]. These whole-organ single-cell atlases offer a comprehensive reference of organ structure and cellular and molecular composition, identify conserved and context-dependent cell states, and support comparative analyses across health and disease conditions [ 36 , 37 ]. Table 5 Single-cell atlases of the visual system ( n = 4) Biological focus Size Species Conditions Modalities Integration strategy Key contributions Reference Eye (update of Swamy 2021 retinal atlas) 35 studies, 44 datasets, 1.1 M cells Human, mouse, macaque, chicken Unclear Transcriptomic 4000 HVGs. Integration with scANVI, study as batch Identified markers shared across species [ 43 ] Retina 7 studies, 16 samples, 330 k cells Mouse Healthy Transcriptomics FASTQs realigned. Integration with scVI, sample as batch Identified previously unknown cell types [ 44 ] Retina scATAC: 1 dataset, 26 donors, 52 samples, 373 k cells; snRNA: 2 datasets, 32 donors, 93 samples, 1.8 M cells; scRNA: 6 datasets, 20 donors, 51 samples, 266 k cells Human Healthy Transcriptomic (separate atlases for single cell and single nucleus), chromatin accessibility FASTQs realigned, 10,000 HVGs. Separate integration of scRNA-seq and snRNA-seq with scVI, sample as batch. Integration of snRNA-seq and snATAC-seq with scGLUE. For reference mapping, a separate reference of only healthy donors was integrated with scANVI Characterized transcriptomic regulatory elements and GRNs at the level of cell subtypes, as annotated using joint mapping of transcriptomic and chromatin accessibility data. Associated GWAS variants and eQTLs with cell types [ 45 ] Retina 26 studies, 33 datasets, 86 samples, 767 k cells Human, mouse, macaque Healthy Transcriptomic 5000 HVGs. Integration of human data with scVI, study as batch. Reference mapping of other species with scArches Provided a pipeline for data integration benchmarking. Showed that large-scale integration and reference mapping can bridge multiple species and in vitro models [ 46 ] HVGs , Highly variable genes; scANVI , Single-cell annotation using variational inference; scVI , Single-cell variational inference; scRNA , Single-cell ribonucleic acid; seq , Sequencing; snRNA , Single-nucleotide ribonucleic acid; scGLUE , Single-cell graph-linked unified embedding; scArches , Single-cell architectural surgery; GRNs , Gene regulatory networks; GWAS , Genome-wide association studies; eQTLs , Expression quantitative trait loci Table 6 Single-cell atlases of the respiratory system ( n = 4) Biological focus Size Species Conditions Modalities Integration strategy Key contributions Reference (Developing) lung Human: 10 studies, 104 subjects, 148 samples, 505 k cells; Mouse: 1 study, 17 samples, 96 k cells Human, mouse (separate atlases) Human: healthy. Mouse: embryonic day 16.5—postnatal day 28 Transcriptomic Integration with MNN, donor as batch Harmonized annotation of cell types for both human and mouse lungs [ 38 ] Lung endothelial cells Human: 6 studies, 73 subjects, 73 samples, 15 k cells; Mouse: 6 studies, 18 subjects, 13 k cells Human, mouse (separate atlases) Healthy Transcriptomic Integration with Seurat CCA, dataset as batch Described previously unidentified cell populations in both humans and mice [ 47 ] Lung and nose 36 studies, 49 datasets, 487 subjects, 743 samples, 2,4 M cells Human Healthy, 15 lung diseases Transcriptomic 2000 HVGs. Integration of healthy data with scANVI, reference mapping of healthy and disease datasets with scArches, dataset as batch Provided consensus on lung cell types and identified new cell types. Associated cell-type specific gene expression patterns with demographic variables and anatomical location. Identified cell states shared across diseases [ 37 ] Non-small cell lung cancer 19 studies, 29 datasets, 318 subjects, 556 samples, 1,284 k cells Human Control, NSCLC (different stages, primary tumors and metastases) Transcriptomic 6000 HVGs. Integration with scVI followed by scANVI, sample as batch Discovered plasticity of TRN and TRN gene signature associated with treatment failure [ 48 ] NSCLC , Non-small cell lung cancer; MNN , Mutual nearest neighbour; CCA , Canonical correlation analysis; HVGs , Highly variable genes; scANVI , Single-cell annotation using variational inference; scArches , Single-cell architectural surgery; scVI , Single-cell variational inference; TRN , Tumor-resident neutrophils Table 7 Single-cell atlases of the integumentary and skeletal systems ( n = 2) Biological focus Size Species Conditions Modalities Integration strategy Key contributions Reference Human prenatal skin 534,581 cells, of which 433,961 cells passed quality control Human 7 to 17 post-conception weeks (prenatal) Transcriptomic, spatial transcriptomics Integration with Harmony, datasets as batches, dataset batches as covariates Provided a comprehensive multi-omics reference atlas of prenatal human skin. Revealed interactions between immune and non-immune cells in skin morphogenesis. Demonstrated the contributions of macrophages in endothelial development and vascular remodeling in a skin organoid model [ 49 ] (Developing) skeleton 10 studies, 133 k cells Mouse Embryonic to adult Transcriptomic 5000 HVGs. Integration with scANVI, dataset as batch Leveraged the availability of time-series data across datasets to predict the effect of transcription factor knock-outs in silico . Identified cell–cell communication patterns driving early limb formation [ 50 ] HVGs , Highly variable genes; scANVI , Single-cell annotation using variational inference Table 8 Single-cell atlases of the metabolic, endocrine and digestive systems ( n = 5) Biological focus Size Species Conditions Modalities Integration strategy Key contributions Reference Gastric cancer 48 samples, 31 individuals, > 200,000 cells Human Patients with gastric cancer Transcriptomic, spatial transcriptomic, CNV analysis PrepSCTIntegration was used to select features, FindIntegrationAnchors was used for anchor gene identification, and Seurat was used for data integration Discovered 34 distinct cell-lineage states in gastric cancer, including rare tumor-associated fibroblast and plasma cell subtypes. Revealed an increase in plasma cell proportions as a novel characteristic of diffuse-type tumors. Identified INHBA–FAP -high fibroblast populations as potential poor prognostic biomarkers. Provided insights into potential origins of abnormal cellular states within gastric tumors [ 51 ] Kidney 8 studies, 59 samples, 140 k cells Mouse Healthy Transcriptomic 3000 HVGs. Integration with scANVI, dataset as batch Characterized rare cell types and identified robust markers of kidney cell types [ 32 ] Liver 10,372 cells, 9 donors Human Non-diseased liver tissue from patients undergoing liver resection due to colorectal cancer metastases or cholangiocarcinoma without a history of chronic liver disease Transcriptomic Integration using RaceID3 for clustering and DPT/SOMs for hepatocyte zonation. Batch integration of fresh versus cryopreserved liver samples Identified novel subtypes of liver endothelial cells, Kupffer cells, and hepatocytes, mapped transcriptome-wide hepatocyte zonation, and discovered EPCAM + TROP2 int cells as potential bipotent liver progenitors [ 52 ] Liver 18 datasets, 518 samples, 2.2 M cells Human Healthy, 4 liver diseases Transcriptomic Integration of healthy reference with scANVI, donor as batch; query-to-reference mapping of disease data with scArches Identified altered cell type states in hepatocellular carcinoma [ 35 ] Pancreatic islets 9 datasets, 56 samples, 302 k cells Mouse Healthy across ages, diabetes models, treatments Transcriptomic Removal of most ambiently expressed genes, 2000 HVGs. Integration with scArches-cVAE, sequencing sample as batch Revealed misconceptions about a commonly used mouse diabetes model. Identified new beta cell states in diabetes and molecular pathways shared in health and diabetes. Corrected and improved existing beta cell state markers across datasets [ 53 ] CNV , Copy-number variations; HVGs , Highly variable genes; scANVI , Single-cell annotation using variational inference; DPT/SOMs , Diffusion pseudo time/self-organizing maps; scArches , Single-cell architectural surgery; scArches-cVAE , Single-cell architectural surgery using conditional variational autoencoder Table 9 Single-cell atlases of the cardiovascular and reproductive systems ( n = 3) Biological focus Size Species Conditions Modalities Integration strategy Key contributions Reference Heart 8 studies, 65 samples, 1 M cells Human Healthy fetal and adult (separate atlases) Transcriptomic Integration with scANVI, dataset as batch Identified potential drug targets and associations between drug side effects and molecular profiles of heart cells based on expression level changes [ 36 ] Breast 55 donors, 161 samples, 800 k cells Human Healthy, healthy with high risk of breast cancer, breast cancer Transcriptomic Integration with scVI, sequencing sample as batch Described the influence of BRCA1/2 mutations, age, and reproductive history on breast tissue, and identified markers related to immune responses and cellular changes for early breast cancer detection [ 54 ] Endometrium 7 studies, 63 subjects, 314 k cells Human Healthy, endometriosis Transcriptomic FASTQs realigned, 2000 HVGs. Integration with scVI, donor and study as batch. Repeated per lineage for lineage-specific integrations. For compatibility with scArches, a new integration with the sample as the only batch covariate was performed with scANVI Identified new cell types and cell types associated with endometriosis. Described cell interactions across space and time to explain differentiation and regeneration [ 55 ] scANVI , Single-cell annotation using variational inference; scVI , Single-cell variational inference; HVGs , Highly variable genes; scArches , Single-cell architectural surgery; BRCA1/2 , Breast cancer gene 1/2
Single-cell atlases of the visual system ( n = 4)
HVGs , Highly variable genes; scANVI , Single-cell annotation using variational inference; scVI , Single-cell variational inference; scRNA , Single-cell ribonucleic acid; seq , Sequencing; snRNA , Single-nucleotide ribonucleic acid; scGLUE , Single-cell graph-linked unified embedding; scArches , Single-cell architectural surgery; GRNs , Gene regulatory networks; GWAS , Genome-wide association studies; eQTLs , Expression quantitative trait loci
Single-cell atlases of the respiratory system ( n = 4)
NSCLC , Non-small cell lung cancer; MNN , Mutual nearest neighbour; CCA , Canonical correlation analysis; HVGs , Highly variable genes; scANVI , Single-cell annotation using variational inference; scArches , Single-cell architectural surgery; scVI , Single-cell variational inference; TRN , Tumor-resident neutrophils
Single-cell atlases of the integumentary and skeletal systems ( n = 2)
HVGs , Highly variable genes; scANVI , Single-cell annotation using variational inference
Single-cell atlases of the metabolic, endocrine and digestive systems ( n = 5)
CNV , Copy-number variations; HVGs , Highly variable genes; scANVI , Single-cell annotation using variational inference; DPT/SOMs , Diffusion pseudo time/self-organizing maps; scArches , Single-cell architectural surgery; scArches-cVAE , Single-cell architectural surgery using conditional variational autoencoder
Single-cell atlases of the cardiovascular and reproductive systems ( n = 3)
scANVI , Single-cell annotation using variational inference; scVI , Single-cell variational inference; HVGs , Highly variable genes; scArches , Single-cell architectural surgery; BRCA1/2 , Breast cancer gene 1/2
Current single-cell atlases demonstrate a strong focus on human biology and diseases, with 24 human-centric atlases identified in our analysis (Table 10 ). While these atlases reflect the growing emphasis on human health, mouse models remain crucial in biomedical research, as evidenced by the presence of five mouse-centric atlases. In addition to these organism-centric atlases, we identified four studies presenting separate human and mouse atlases, one fully integrated cross-species atlas, and two meta-atlases. Several studies, including some organism-centric atlases, incorporated cross-species analyses through different approaches, including computational alignment tools and xenograft models. Table 10 Classification of single-cell atlases by species focus and cross-species analysis methods ( n = 36) Atlas type Species Cross-species analysis methods Key findings of cross-species analysis Reference Analysis strategy Comparison Validation Human-centric Human N/A N/A N/A N/A [ 18 – 20 , 26 , 28 – 31 , 35 – 37 , 39 , 42 , 48 , 51 , 54 , 55 ] Separate Cell state alignment (logistic regression, scVI); pseudotime comparison (CellRank, Genes2Genes); marker gene analysis (Wilcoxon test, Visium) Spatial co-localization (RNAscope, immunofluorescence); functional pathway validation (organoid experiments) Human skin showed slower hair follicle maturation, delayed fibroblast differentiation, and macrophage functional divergence compared to mice. Early human fibroblasts mirrored scarless reindeer antler fibroblasts, while late-gestation fibroblasts resembled scar-forming back skin fibroblasts [ 49 ] Separate Cell/gene alignment (SATURN, MetaNeighbor); marker gene comparison; hierarchical clustering Functional validation (MPRAs); label transfer from public datasets (Seurat, scPred) Bipolar cells exhibited primate-specific subtypes (DB4a/b) that aligned with mouse BC5A, while BC5D was unique to rodents. Amacrine cells showed greater conservation between macaques and humans (94%) than between mice and humans (83%). Most human CREs were inactive in mice, with only 27.3% functioning as enhancers and 6.6% as silencers. RGCs were more diverse in rodents (45 types) than in primates (~ 18 types) [ 45 ] Separate Cell alignment (Seurat reference projection) Functional validation (TCR clonality analysis) Approximately 70% of mouse immune cells mapped to human immune cells with high confidence. Mouse T cells aligned with human exhausted, cytotoxic, and proliferative states showed clonal expansion [ 27 ] Separate N/A Literature-based validations (marker gene conservation, functional pathway comparison, developmental biology comparison, ATOs) Human prenatal B1 cells were identified using conserved murine markers (CD5, CD27, CD43, IgM/IgD). Monocyte maturation (CXCR4 → CCR2 pathway) and hematopoietic progression (yolk sac → liver → bone marrow) matched murine biology. Human iPSCs in mouse-modelled ATOs generated unconventional T cells via thymocyte-thymocyte selection [ 21 ] Separate Transcriptomic alignment (gene ortholog mapping, clustering analysis); functional pathway analysis (GSEA, RNA velocity) Human Protein Atlas validation; comparative cell-type mapping While the overall AV-zonation structure is conserved between humans and mice, gene expression varies across AV compartments, with the highest human-specific markers in small vessels. ECs and PVCs show high transcriptomic conservation, unlike neurons and astrocytes. Both species shared some BBB dysregulation pathways [ 40 ] Separate Orthologous gene mapping (gene symbol matching, blastn); zonation comparison (dpt, Pearson's correlation); gene expression module identification (SOMs) Xenograft model (FRG-NOD mice); indirect comparative literature-based validation (immunostaining) human and mouse liver zonation showed limited evolutionary conservation. Engrafted human liver cells conserved core functions (e.g., ALB , CLEC4G ) but exhibited injury/disease-like responses ( AKR1B10 upregulation, WNT activation) in mice [ 52 ] Separate N/A Xenograft model (NSG mice) Human intestinal organoids developed greater cellular diversity and closer similarity to human primary enterocytes after xenografting into a mouse host [ 23 ] Mouse-centric Mouse N/A N/A N/A N/A [ 41 , 44 , 50 ] Separate Transcriptomic alignment (UMAP/PCA co-embedding, reference mapping via scArches); pathway conservation analysis (differential gene expression, BioMart ortholog mapping, GO/KEGG enrichment); label transfer ( k -NN classification) Ortholog-based marker testing (BioMart ortholog mapping, human scRNA-seq validation); negative validation (non-conserved markers) The mouse NOD model matched human T1D immune pathways, while db/db and mSTZ models aligned with human T2D metabolic stress pathways. Hallmark diabetes genes like Dgkb were conserved, but DEG patterns differed between mice and humans. Some markers like Spock3 showed species specificity (δ-cell in mice, α-cell in humans) [ 53 ] Separate N/A Label transfer (Azimuth’s human kidney reference) Species-matched references like Mouse Kidney Reference provided better classification accuracy than human reference [ 32 ] Separate atlases Human, mouse (separate atlases) Separate Orthologous gene mapping; gene signature projection (Seurat’s AddModuleScore) Multi-disease functional alignment; GTEx co-expression analysis; ClusterMap algorithmic matching Universal ( Pi16 + /Col15a1 + ) and activated ( Lrrc15 + ) fibroblast states are conserved between mice and humans. Analysis of disease datasets revealed that human fibroblast subtypes have corresponding mouse fibroblast orthologues. However, a distinct COL3A1 + myofibroblast population was exclusively observed in COVID-19 patients, indicating the presence of species-specific or condition-specific differences [ 22 ] Separate Unsupervised integration (joint UMAP space); orthologous gene mapping (gene symbol matching, conserved marker identification); differential expression analysis (Wilcoxon rank-sum test) Subpopulation identification; spatial localization (immunofluorescent); functional conservation test (ligand-receptor pair analysis) Conserved and species-specific marker genes were identified across endothelial cell subpopulations in both mice and humans. Notable differences were observed in the ratio of aerocytes to general capillary endothelial cells, with aerocytes being less common in mice. Ligand-receptor expression patterns also varied between species [ 47 ] N/A N/A N/A N/A [ 33 , 38 ] Fully integrated cross-species atlas Human, Mouse Combined Orthologous gene mapping (Ensembl database, SCANPY); batch correction (Harmony); logistic regression; hierarchical annotation (scANVI) Putative human NPCs validation (NPC gene module score); protein-level NPC confirmation (immunostaining); trajectory analysis (RNA velocity, pseudo time comparison); microglia differential expression analysis Adult human NPCs expressed conserved markers and developmental pathways found in fetal and mouse NPCs, although adult neurogenesis is less active [ 34 ] Meta-atlas Human, mouse, macaque, chicken (excluded due to limited samples) Combined Ortholog gene mapping (Ensembl BioMart); differential expression testing (pseudobulk DESeq2 analysis); UMAP visualization (scANVI); heatmap visualization Consistency filters; automated literature validation (PubMed database); gene-level replication of prior studies; machine learning label accuracy (xgboost, confusion matrices); independent dataset projection Conserved gene expression patterns were identified across human, mouse, and macaque retinas. Both known and novel genes involved in the transition from retinal progenitor cells to neurogenic cells were discovered in mice and humans, including the downregulation of ZBTB38 and the upregulation of 11 genes [ 43 ] Meta-atlas Human, mouse, macaque Combined Orthologous gene mapping (Ensembl BioMart); custom macaque reference quantification (Ensembl, Gencode human reference); integrated UMAP visualization (scVI); differential expression testing (Wilcoxon rank-sum) Integration quality control (scPOP metrics: LISI/Silhouette); machine learning label transfer (xgboost, precision-recall curves); manual literature checks of markers Major retinal cell types overlapped across species. Machine learning models trained on one species (human or mouse) accurately predicted cell types in another species, particularly for terminal cell types [ 46 ] N/A , Not applicable; scVI , Single-cell variational inference; SATURN , Species alignment through unification of RNA and proteins; MPRAs , Massively parallel reporter assays; CREs , Cis-regulatory elements; RGCs , Retinal ganglion cells; TCR , T cell receptor; ATOs , Artificial thymic organoids; iPSCs , Induced pluripotent stem cells; GSEA , Gene set enrichment analysis; AV , Arteriovenous; ECs , Endothelial cells; PVCs , Perivascular cells; BBB , Blood–brain barrier; dpt , Diffusion pseudo time; SOMs , Self-organizing maps; FRG-NOD , Fah −/− /Rag2 −/− /Il2rg −/− non-obese diabetic; NSG , NOD-SCID IL2Rg null; UMAP , Uniform manifold approximation and projection; PCA , Principal component analysis; scArches – Single-cell architectural surgery; GO , Gene ontology; KEGG , Kyoto Encyclopedia of Genes and Genomes; k-NN , k -nearest neighbors; scRNA-seq , Single-cell ribonucleic acid sequencing; T1D , Type 1 diabetes; T2D , Type 2 diabetes; NOD , Non-obese diabetic; db/db , Glucotoxicity/lipotoxicity type 2 diabetes; mSTZ , Multiple low-dose streptozotocin; NPCs , Neural progenitor cells; scANVI , Single-cell annotation using variational inference; scPOP , Single-cell pick optimal parameters; LISI , Local Inverse Simpson’s Index
Classification of single-cell atlases by species focus and cross-species analysis methods ( n = 36)
N/A , Not applicable; scVI , Single-cell variational inference; SATURN , Species alignment through unification of RNA and proteins; MPRAs , Massively parallel reporter assays; CREs , Cis-regulatory elements; RGCs , Retinal ganglion cells; TCR , T cell receptor; ATOs , Artificial thymic organoids; iPSCs , Induced pluripotent stem cells; GSEA , Gene set enrichment analysis; AV , Arteriovenous; ECs , Endothelial cells; PVCs , Perivascular cells; BBB , Blood–brain barrier; dpt , Diffusion pseudo time; SOMs , Self-organizing maps; FRG-NOD , Fah −/− /Rag2 −/− /Il2rg −/− non-obese diabetic; NSG , NOD-SCID IL2Rg null; UMAP , Uniform manifold approximation and projection; PCA , Principal component analysis; scArches – Single-cell architectural surgery; GO , Gene ontology; KEGG , Kyoto Encyclopedia of Genes and Genomes; k-NN , k -nearest neighbors; scRNA-seq , Single-cell ribonucleic acid sequencing; T1D , Type 1 diabetes; T2D , Type 2 diabetes; NOD , Non-obese diabetic; db/db , Glucotoxicity/lipotoxicity type 2 diabetes; mSTZ , Multiple low-dose streptozotocin; NPCs , Neural progenitor cells; scANVI , Single-cell annotation using variational inference; scPOP , Single-cell pick optimal parameters; LISI , Local Inverse Simpson’s Index
The mouse remains a cornerstone of experimental biology due to its high anatomical and physiological similarities to humans [ 56 , 57 ]. However, despite sharing similar gene activity in some tissues, significant species-specific differences in gene regulation and expression pose challenges for the applicability of the mouse model in certain areas of human biology research [ 56 , 58 ]. For example, a co-expression network-based study comparing human and mouse gene regulation revealed strong brain, bone, and metabolic disorder-related gene conservation [ 58 ]. Nonetheless, greater divergence was observed in tissues such as the testis, eye, and skin, as well as in immune response pathways and tumor-related genes [ 58 ]. These discrepancies highlight the need for context-aware interpretation when extrapolating findings from mouse models to human systems. Furthermore, nonhuman primate models, especially macaques, serve as another promising model for human research due to their greater genetic similarity, sharing over 95% DNA sequence homology with humans [ 59 , 60 ]. Their relevance is especially pronounced in the study of model-specific diseases and immune responses. However, this model comes with higher costs, handling difficulties, and ethical considerations [ 60 ]. These findings underscore that the choice of animal models in human biology is context-dependent and requires cautious interpretation. Hence, cross-species analysis at the single-cell level is essential to improving model selection by providing higher-resolution gene expression data, refining cross-species relationships, and ensuring the relevance of findings to human biology.
Cross-species analyses of single-cell atlases are vital in bridging the biological gap between humans and other organisms, discovering the shared and unique cellular and molecular features, gaining insights into cellular heterogeneity and rare cell populations, and enhancing our understanding of evolutionary and developmental biology [ 61 , 62 ]. Such analyses can be conducted using separate or combined approaches, each with distinct advantages and limitations [ 63 ]. While the separate analysis preserves species-specific variations and dataset-specific heterogeneity, it relies on manual cross-annotation of homologous cell types, which can be labor-intensive and prone to annotation inconsistencies. In contrast, the combined analysis merges datasets from multiple species into a unified framework with batch correction, increasing the overall cell numbers for clustering, which enhances rare cell type detection. However, this approach is computationally intensive and may obscure species-specific cell identities due to overcorrection [ 63 ]. Therefore, selecting an appropriate analytical strategy necessitates careful consideration of the research objectives and the trade-offs between preserving biological specificity and maximizing comparative resolution.
Computational tools play several roles in integrating and analyzing cross-species single-cell datasets. Technical and biological batch effects arising from experimental protocols or biological variations such as sex, age, and genetic background may obscure actual biological differences [ 63 ]. To overcome this, various computational methods have been generated to mitigate these batch effects and enable more accurate cross-species comparisons [ 63 – 65 ]. However, careful consideration is needed to avoid overcorrection, which may eliminate biologically meaningful differences, or undercorrection, which may inadequately mix cell populations across species clusters [ 64 , 65 ]. Although computational tools facilitate cross-species gene mapping and alignments, several challenges persist in homolog annotation, cell type annotation and transfer, and trajectory alignment. These challenges result from evolutionary divergence, inaccurate or incomplete gene homology annotation, and developmental variations across species, which may lead to information loss during integration. As a result, future computational tools should incorporate multi-omics data integration and leverage machine learning for better-integrated embeddings [ 64 , 65 ].
Beyond computational comparisons and predictions, cross-species validations are also crucial for confirming biological insights. Among the current single-cell atlases, one widely adopted validation approach is the xenograft model, which provides valuable insights into the functionality, maturation, and fidelity of human cells or organoids in living organisms [ 23 , 52 ]. These studies examined the impact of the cross-species environment on human cells or organoids transplanted into murine hosts [ 23 , 52 ]. These models offer critical insights into how human cells respond to and adapt within a foreign biological environment, such as the murine host. For example, Aizarani et al. demonstrated that human hepatocytes and non-parenchymal liver cells were engrafted into mice and retained core liver-specific gene expressions, including ALB and CLEC4G [ 52 ]. However, the altered gene expression of AKR1B10 was observed between engrafted and non-engrafted cells, indicating environmental influence on cellular gene regulation [ 52 ]. Similarly, Xu et al. observed different maturation patterns in human endoderm-derived organoids before and after xenografting into a mouse host, highlighting the impact of the host environment on cellular development [ 23 ]. These findings underscore the importance of xenograft models in translational research, providing a valuable platform for studying conserved and cross-species adaptation mechanisms, validating therapeutic models, and investigating disease mechanisms in vivo while also facilitating preclinical validation of regenerative and therapeutic models [ 23 , 52 ].
Understanding the dataset size and coverage is crucial for evaluating the comprehensiveness and applicability of single-cell atlases in biological research. The breadth and depth of data within single-cell atlases directly impact their utility in research, particularly in elucidating cellular diversity and disease mechanisms. The dataset size and coverage of single-cell atlases vary significantly, influencing their capacity to provide insights into cellular heterogeneity and disease mechanisms. In summary, the integration of diverse and extensive datasets is essential for enhancing the understanding of cellular diversity and the underlying mechanisms of diseases across various biological contexts. This comparative analysis emphasizes the significance of dataset size and biological coverage in enhancing understanding of cellular diversity and disease mechanisms, ultimately guiding future research directions. In conclusion, a systematic comparative assessment of single-cell atlases is vital for advancing our understanding of cellular diversity and enhancing the relevance of research findings across species.