Inferring spatial gene expression from tissue images using large-scale histology foundation model with SpaFoundation

preprint OA: closed CC-BY-NC-ND-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-06 · read from full text

The paper introduces SpaFoundation, a large-scale histology foundation model that uses a teacher-student Vision Transformer framework to predict spatial transcriptomics gene expression from histological tissue images. It pretrains on 1.79 million image patches across 26 tissue types and 117 validation samples, using joint self-distillation and masked image modeling to capture both semantic and fine-grained structural features, and reports improved spatial gene expression prediction performance with flexibility across spatial resolutions and transferability to downstream tasks like tumor detection and spatial domain clustering. A key limitation is that the work is designed around computational inference from histology images rather than addressing the underlying high cost and time requirements of spatial transcriptomics technologies directly. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Spatial transcriptomics (ST) has revolutionized biological research by enabling the joint profiling of gene expression and spatial context, along with histological images. However, current existing ST technologies remain high-cost and time-consuming, hindering their broader clinical applications. Although computational methods have been developed to infer gene expression directly from histology images, these methods still suffer from limited accuracy and spatial resolution due to insufficient training data and model capacity. Here, we introduce SpaFoundation, a large-scale histology foundation model designed to accurately predict spatial gene expression from tissue images. SpaFoundation employs a teacher-student Vision Transformer (ViT) architecture to learn generalizable histological representations by modeling potential dependencies among image patches. Notably, we incorporated self-distillation and masked image modeling (MIM) jointly to capture high-level semantic representations and fine-grained structural features, enriching spots’ representations. The model is pretrained on 1.79 million patches spanning 26 tissue types, with 80 million parameters. We validated SpaFoundation using 117 samples, demonstrating its flexibility across different spatial resolutions and superior performance in spatial gene expression prediction, as well as strong transferability to downstream tasks such as tumor detection and spatial domain clustering. Our results highlight the potential of large-scale foundation model to learn informative histological representations and underscore the benefits of domain-specific pretraining in extracting task-relevant representations, paving the way for foundation model-driven spatial gene expression inference. The implementation and pre-trained weights of SpaFoundation are publicly available at https://github.com/NingZhangCSUBio/SpaFoundation .
Full text 2,058 characters · extracted from oa-html · click to expand
Abstract Spatial transcriptomics (ST) has revolutionized biological research by enabling the joint profiling of gene expression and spatial context, along with histological images. However, current existing ST technologies remain high-cost and time-consuming, hindering their broader clinical applications. Although computational methods have been developed to infer gene expression directly from histology images, these methods still suffer from limited accuracy and spatial resolution due to insufficient training data and model capacity. Here, we introduce SpaFoundation, a large-scale histology foundation model designed to accurately predict spatial gene expression from tissue images. SpaFoundation employs a teacher-student Vision Transformer (ViT) architecture to learn generalizable histological representations by modeling potential dependencies among image patches. Notably, we incorporated self-distillation and masked image modeling (MIM) jointly to capture high-level semantic representations and fine-grained structural features, enriching spots’ representations. The model is pretrained on 1.79 million patches spanning 26 tissue types, with 80 million parameters. We validated SpaFoundation using 117 samples, demonstrating its flexibility across different spatial resolutions and superior performance in spatial gene expression prediction, as well as strong transferability to downstream tasks such as tumor detection and spatial domain clustering. Our results highlight the potential of large-scale foundation model to learn informative histological representations and underscore the benefits of domain-specific pretraining in extracting task-relevant representations, paving the way for foundation model-driven spatial gene expression inference. The implementation and pre-trained weights of SpaFoundation are publicly available at https://github.com/NingZhangCSUBio/SpaFoundation. Competing Interest Statement The authors have declared no competing interest. Footnotes Some experimental results and analytical content have been updated.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-NC-ND-4.0