Mosaic integration of spatial multi-omics with SpaMosaic

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by qwen3.7-flash, 2026-09-10

SpaMosaic integrates spatial multi-omics datasets to identify domains and impute missing modalities, revealing accurate transcriptome-epigenome correlations in mouse brain tissue.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by qwen3.7-flash, 2026-09-07 · read from full text

This paper introduces SpaMosaic, a computational tool designed for the mosaic integration of spatial multi-omics datasets that possess partially overlapping modalities. Utilizing contrastive learning and graph neural networks, the method constructs a modality-agnostic latent space to identify spatial domains and impute missing omic layers across various tissues including brain, embryo, tonsil, and lymph node. The authors demonstrate through benchmarking that SpaMosaic outperforms existing methods in reducing batch effects and enhancing resolution without requiring image registration or co-profiling from the same tissue section. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract With the advent of spatial multi-omics, we can mosaic integrate diverse datasets with partially overlapping modalities to construct consensus multi-modal spatial atlases of the source tissue. SpaMosaic is a spatial multi-omics mosaic integration tool that employs contrastive learning and graph neural networks to construct a modality-agnostic and batch-corrected latent space suited for analyses like spatial domain identification and imputing missing omes. Using simulated and experimentally acquired datasets, we benchmarked SpaMosaic against single-cell multi-omics mosaic integration methods. The experimental data encompassed RNA and protein abundance, chromatin accessibility or histone modifications, acquired from brain, embryo, tonsil, and lymph node tissues. SpaMosaic achieved superior performance over existing methods in identifying known spatial domains with enhanced resolution and clarity while reducing noise and batch effects. It also ranked top for modality alignment, enabling seamless diagonal integration without the need for image registration. After integration, SpaMosaic can also impute missing modalities. With a mosaic set of mouse brain data with RNA and different epigenomic modalities, we integrated and imputed the missing omics. There we found the imputed gene activity scores of activating and silencing histone marks show the correct correlation with RNA. Moreover, we uncovered more region-specific genes and pathways showing the correct transcriptome-epigenome correlations in the imputed histone modification than in the measured chromatin accessibility modalities. Lastly, SpaMosaic’s imputation also allows the inference of relationships between different modalities without requiring co-profiling from the same section.
Full text 50,056 characters · extracted from preprint-html · click to expand
Mosaic integration of spatial multi-omics with SpaMosaic | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Article Mosaic integration of spatial multi-omics with SpaMosaic Jinmiao Chen, Xuhua Yan, Kok Siong Ang, Lynn van Olst, Alex Edwards, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5507983/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 24 Apr, 2026 Read the published version in Nature Genetics → Version 1 posted You are reading this latest preprint version Abstract With the advent of spatial multi-omics, we can mosaic integrate diverse datasets with partially overlapping modalities to construct consensus multi-modal spatial atlases of the source tissue. SpaMosaic is a spatial multi-omics mosaic integration tool that employs contrastive learning and graph neural networks to construct a modality-agnostic and batch-corrected latent space suited for analyses like spatial domain identification and imputing missing omes. Using simulated and experimentally acquired datasets, we benchmarked SpaMosaic against single-cell multi-omics mosaic integration methods. The experimental data encompassed RNA and protein abundance, chromatin accessibility or histone modifications, acquired from brain, embryo, tonsil, and lymph node tissues. SpaMosaic achieved superior performance over existing methods in identifying known spatial domains with enhanced resolution and clarity while reducing noise and batch effects. It also ranked top for modality alignment, enabling seamless diagonal integration without the need for image registration. After integration, SpaMosaic can also impute missing modalities. With a mosaic set of mouse brain data with RNA and different epigenomic modalities, we integrated and imputed the missing omics. There we found the imputed gene activity scores of activating and silencing histone marks show the correct correlation with RNA. Moreover, we uncovered more region-specific genes and pathways showing the correct transcriptome-epigenome correlations in the imputed histone modification than in the measured chromatin accessibility modalities. Lastly, SpaMosaic’s imputation also allows the inference of relationships between different modalities without requiring co-profiling from the same section. Biological sciences/Computational biology and bioinformatics/Data mining Biological sciences/Computational biology and bioinformatics/Data processing spatial multi-omics mosaic integration graph neural network contrastive learning Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Full Text Additional Declarations There is NO Competing Interest. Supplementary Files SupplementaryInformation.docx Supplementary material for the manuscript SuppFigS1.pdf Supp. Fig. S1 Graph neural network architecture and imputation strategy. a Aggregation and propagation processes of weighted LGCN. The weight for spatial neighbors is set to 1 while the weight for expression neighbors is set to 0.8. In each layer, each node aggregates the information from its neighbors using the weighted average. After two layers, the information from each node is concatenated across three output layers and then processed by a multiple layer perceptron. bImputation process. Taking section 3 as an example, which only has ATAC modality, we aim to impute its RNA and protein profiles. To impute the RNA profile of each spot in section 3, we find its RNA-specific neighbors from the other sections and then compute the average target expression profiles of these neighbors. The same approach is also taken for other modalities. SuppFigS2.pdf Supp. Fig. S2 Simulation settings and plots for competing methods. a Five sets of simulation parameters used for data generation. b Mosaic configuration of each dataset. A dotted box denotes that this profile is missing. c ARI scores of clusters obtained by k- means and mclust clustering of the embeddings from the different methods and the raw data. d UMAP plots of the embeddings from all benchmarked methods. Spots are colored by section (first row), cluster (second row) and factor (third row). e Spatial plots of the clustering results of all methods. All methods and sections share the same color map. SuppFigS3.pdf Supp. Fig. S3 Additional results from benchmarking with an embryonic mouse brain (RNA+ATAC) dataset. a UMAP plots of the raw RNA and ATAC profiles. Spots are colored by section. b UMAP plots of the embeddings from all tested methods. Spots in the first row are colored by section, the second row by cluster and the third row by manual annotation. c Mosaic dataset for evaluation with modality alignment metrics. The RNA and ATAC profiles of the E18.5 section were split into two separate sections to simulate originating from different experiments. Cluster labels were obtained from the original study. dUMAP plots of the raw RNA and ATAC profiles of the E15.5 and E18.5 sections. Spots are colored by section. SuppFigS4.pdf Supp. Fig. S4 Additional results of a postnatal mouse brain dataset (RNA+H3K4me3) benchmarking. a UMAP plots of the raw RNA and H3K4me3 profiles. Spots are colored by section. b UMAP plots of the embeddings from all methods. Spots in the first row are colored by section, the second row by cluster. SuppFigS5.pdf Supp. Fig. S5 Downstream analysis results obtained with SpaMosaic and quantitative benchmarking on a postnatal mouse brain (RNA+H3K27me3) dataset. a Mosaic dataset configuration and clustering patterns obtained for all three sections. The RNA and H3K27me3 cluster labels were obtained by performing k -means clustering on the batch-corrected low-dimensional embeddings from Harmony. The multi-modal cluster labels of section 1 were obtained by performing k -means clustering on the concatenated batch-corrected low-dimensional embeddings of the RNA and H3K27me3 profiles. b Spatial plots of the clusters from the methods for all three sections. All methods share the color map. Manual annotation was performed to identify SpaMosaic’s clusters. c RNA expression of marker genes of the six brain regions in SpaMosaic’s clusters. d Clusters overlaid on the image of the third section (top row) and spatial heatmap showing the gene activity scores of selected markers on the third section (middle and bottom rows). e Moran’s I scores of all methods across the three sections. We varied the clustering numbers (n=4, 6, 8, 10) and calculated the Moran’s I scores for each clustering result. f iLISI scores of all methods. SuppFigS6.pdf Supp. Fig. S6 Additional results of a postnatal mouse brain dataset (RNA+H3K27me3) benchmarking. a UMAP plots of the raw RNA and H3K27me3 profiles. Spots are colored by section. b UMAP plots of the embeddings from all methods. Spots in the first row are colored by section, the second row by cluster. SuppFigS7.pdf Supp. Fig. S7 Downstream analysis results obtained with SpaMosaic and quantitative benchmarking on a postnatal mouse brain dataset (RNA+ATAC). a Mosaic dataset configuration and clustering patterns obtained for all three sections. The RNA and ATAC cluster labels were obtained by performing k -means clustering on the batch-corrected low-dimensional embeddings from Harmony. The multi-modal cluster labels of section 1 were obtained by performing k -means clustering on the concatenated batch-corrected low-dimensional embeddings of the RNA and ATAC profiles. b Spatial plots of the clusters obtained from the methods for all three sections. All methods share the color map. Manual annotation was performed to identify SpaMosaic. c Heatmap of the DEGs of SpaMosaic’s clusters in section 1 and section 2. d Clusters overlaid on the image of the third section (left) and spatial heatmap depicting the gene activity scores of selected markers on the third section (middle and right). e Moran’s I scores of all methods across the three sections. We varied the clustering numbers (n=4, 6, 8, 10) and calculated the Moran’s I scores for each clustering result. f iLISI scores of all methods. g Mosaic dataset for evaluation with modality alignment metrics. Raw RNA and ATAC profiles of the second section were split into two separate sections to simulate originating from different experiments. Cluster labels were obtained by performing k -means clustering on the batch-corrected low-dimensional embeddings from Harmony. hUMAP plots of the raw RNA profiles and ATAC profiles in g. Spots are colored by section. i Modality alignment scores (FOSCTTM and MS) of all methods. SuppFigS8.pdf Supp. Fig. S8 Additional results of a postnatal mouse brain dataset (RNA+ATAC) benchmarking. a UMAP plots of the raw RNA and ATAC profiles. Spots are colored by section. b UMAP plots of the embeddings from all methods. Spots in the first row are colored by section, the second row by cluster. SuppFigS9.pdf Supp. Fig. S9 Downstream analysis results obtained with SpaMosaic and quantitative benchmarking on a postnatal mouse brain (RNA+H3K27ac) dataset. a Mosaic dataset configuration and clustering patterns obtained for all three sections. The RNA and H3K27ac cluster labels were obtained by performing k -means clustering on the batch-corrected low-dimensional embeddings from Harmony. The multi-modal cluster labels of section 1 were obtained by performing k -means clustering on the concatenated batch-corrected low-dimensional embeddings of the RNA and H3K27ac profiles. b Spatial plots of the clustering labels from the methods for all three sections. All methods share the color map. Manual annotation was performed to identify SpaMosaic’s clusters. c Heatmap of the DEG of SpaMosaic’s clusters in section 1 and section 2. d Clusters overlaid on the image of the third section (top left) and spatial heatmap showing the gene activity scores of selected markers on the third section. e Moran’s I scores of all methods across the three sections. We varied the clustering numbers (n=4, 6, 8, 10) and calculated the Moran’s I scores for each clustering result. f iLISI scores of all methods. g Mosaic dataset for evaluation with modality alignment metrics. Raw RNA and H3K27ac profiles of the second section were split into two separate sections to simulate originating from different experiments. Cluster labels were obtained by performing k -means clustering on the batch-corrected low-dimensional embeddings from Harmony. hUMAP plots of the raw RNA and H3K27ac profiles in g. Spots are colored by section. i Modality alignment scores of all methods (FOSCTTM and MS). Cite Share Download PDF Status: Published Journal Publication published 24 Apr, 2026 Read the published version in Nature Genetics → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5507983","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":388798657,"identity":"02d8c050-8a56-48e8-8a62-52224d9ae923","order_by":0,"name":"Jinmiao Chen","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAt0lEQVRIiWNgGAWjYBACxmYgkcDAIAfiHHhAihZjsJYEUmxLbGCA6CUMmNt5j254UGOTPj/s8EOgLXZyug0EHcaXdiPhWFruxttpBkAtycZmBwhq4TG7kcB2OHfj7ASQlgOJ24jT8u9wuuHs9A8kaElsO5wgL51Dii2JfWmGG6RzCg4kGBDhF8P+M2Y3f3yzkZefnb75w4cKOznCWhqgDAOwSgMCykFAHs5owKNqFIyCUTAKRjYAAHypRxNOWRIgAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0001-7547-6423","institution":"Centre for Computational Biology and Program in Cancer and Stem Cell Biology, Duke-NUS Medical School","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Jinmiao","middleName":"","lastName":"Chen","suffix":""},{"id":388798658,"identity":"0d838aaa-462d-4f6f-9b17-7ce2403230a7","order_by":1,"name":"Xuhua Yan","email":"","orcid":"","institution":"Hunan Provincial Key Lab on Bioinformatics, School of Computer Science and Engineering, Central South University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xuhua","middleName":"","lastName":"Yan","suffix":""},{"id":388798659,"identity":"0227d7d6-9774-4c08-bb8f-9c99504378d6","order_by":2,"name":"Kok Siong Ang","email":"","orcid":"","institution":"Bioinformatics Institute (BII), Agency for Science, Technology and Research (A*STAR)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Kok","middleName":"Siong","lastName":"Ang","suffix":""},{"id":388798660,"identity":"452f8962-5140-46cc-bf92-97671b56abbd","order_by":3,"name":"Lynn van Olst","email":"","orcid":"","institution":"The Ken \u0026 Ruth Davee Department of Neurology, Northwestern University Feinberg School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Lynn","middleName":"van","lastName":"Olst","suffix":""},{"id":388798661,"identity":"0d9c5e56-f5e1-45d6-9479-ba631d72308e","order_by":4,"name":"Alex Edwards","email":"","orcid":"","institution":"The Ken \u0026 Ruth Davee Department of Neurology, Northwestern University Feinberg School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Alex","middleName":"","lastName":"Edwards","suffix":""},{"id":388798662,"identity":"7ee98c08-6fbc-412c-9c49-c071820c7589","order_by":5,"name":"Thomas Watson","email":"","orcid":"","institution":"The Ken \u0026 Ruth Davee Department of Neurology, Northwestern University Feinberg School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Thomas","middleName":"","lastName":"Watson","suffix":""},{"id":388798663,"identity":"1b311d17-448b-4262-a233-a8da901a70f3","order_by":6,"name":"Ruiqing Zheng","email":"","orcid":"https://orcid.org/0000-0001-6372-6798","institution":"Central South University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ruiqing","middleName":"","lastName":"Zheng","suffix":""},{"id":388798664,"identity":"ef58d760-842e-462b-84be-5d0cd499f411","order_by":7,"name":"Rong Fan","email":"","orcid":"https://orcid.org/0000-0001-7805-8059","institution":"Yale School of Engineering and Applied Science","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Rong","middleName":"","lastName":"Fan","suffix":""},{"id":388798665,"identity":"06fd8885-c85d-40d4-b53f-a7556ad820ec","order_by":8,"name":"Min Li","email":"","orcid":"https://orcid.org/0000-0002-0188-1394","institution":"Central South University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Min","middleName":"","lastName":"Li","suffix":""},{"id":388798666,"identity":"ff67f8ca-bd46-4d3c-b76a-96c03f6537be","order_by":9,"name":"David Gate","email":"","orcid":"","institution":"The Ken \u0026 Ruth Davee Department of Neurology, Northwestern University Feinberg School of Medicine","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"David","middleName":"","lastName":"Gate","suffix":""}],"badges":[],"createdAt":"2024-11-23 05:00:32","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5507983/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5507983/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41588-026-02573-3","type":"published","date":"2026-04-24T04:00:00+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":71220627,"identity":"08e7560b-a619-4ff6-98f9-09a13c5a6cac","added_by":"auto","created_at":"2024-12-12 09:12:30","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":64190259,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverview of SpaMosaic framework. a \u003c/strong\u003eSpaMosaic is designed to integrate multiple tissue sections with different modality compositions to construct a more comprehensive view of tissues and impute missing modalities. \u003cstrong\u003eb\u003c/strong\u003e SpaMosaic first performs dimension reduction and applies Harmony to obtain batch-corrected low-dimension representations for each modality. Thereafter, a heterogeneous adjacency graph is constructed between spots within each modality. SpaMosaic employs graph neural networks to embed the modality-specific representations into modality-specific embeddings. The paired multi-modal embeddings are input for contrastive learning which aims to align embeddings from different modalities and constructs a modality-agnostic latent space. Finally, SpaMosaic infers the aligned embeddings for each section for downstream analysis (such as spatial domain identification) and imputation.\u003c/p\u003e","description":"","filename":"Fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/fb589842aa63b32b13762192.png"},{"id":71218987,"identity":"46a726ea-d527-4e55-b586-e2c1c3a027a6","added_by":"auto","created_at":"2024-12-12 09:04:30","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":37735713,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eBenchmarking on simulated datasets. a\u003c/strong\u003e Ground truth of spatial factors and background of the three sections. \u003cstrong\u003eb\u003c/strong\u003e Clustering patterns obtained with the concatenated RNA and protein profiles, RNA profiles and protein profiles. Each section was clustered individually using \u003cem\u003ek\u003c/em\u003e-means. ARI scores were evaluated by comparing the clustering with the ground truth. \u003cstrong\u003ec\u003c/strong\u003e UMAP plots for RNA profiles of sections 1 and 2, protein profiles of sections 1 and 3. Spots are colored by section and factor in each panel. \u003cstrong\u003ed\u003c/strong\u003e Clustering patterns obtained using SpaMosaic’s embeddings for all three sections. \u003cstrong\u003ee\u003c/strong\u003e UMAP plots of SpaMosaic’s embeddings. From left to right, spots are colored by section, clustering and factor. \u003cstrong\u003ef\u003c/strong\u003e iLISI scores of all benchmarked methods across five simulation sets. The error bars represent the 95% confidence interval around the mean of the scores.\u003cstrong\u003e g \u003c/strong\u003eARI scores on the whole dataset, section 1, section 2 and section 3 for all methods. ‘Raw’ indicates the ARI scores that were calculated on the raw data of each section. In the box plot, the center lines indicate the median, boxes indicate the interquartile range, and whiskers indicate 1.5× interquartile range. \u003cstrong\u003eh\u003c/strong\u003e Average Moran’s \u003cem\u003eI\u003c/em\u003e scores on each dataset for all methods across the five datasets. \u003cstrong\u003ei \u003c/strong\u003eAverage Moran’s \u003cem\u003eI\u003c/em\u003e score on individual sections\u003cstrong\u003e \u003c/strong\u003efor all methods. \u003cstrong\u003ej\u003c/strong\u003e Label transfer F1-scores of all methods. \u003cstrong\u003ek\u003c/strong\u003e Sankey plots showing the comparison of prediction and ground truth of factors on the first replicate of the first simulation setting.\u003c/p\u003e","description":"","filename":"Fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/d2fa42d699f305f62cb570c7.png"},{"id":71218977,"identity":"99d62d69-3363-4d45-bd07-6397bde74696","added_by":"auto","created_at":"2024-12-12 09:04:29","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":685081,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDownstream analysis results obtained with SpaMosaic and quantitative benchmarking on an embryonic mouse brain (RNA+ATAC) dataset. a \u003c/strong\u003eMosaic dataset and clustering patterns for all three sections: E13.5 without the ATAC modality, E15.5 with both RNA and ATAC expression profiles, and E18.5 without the RNA modality. The modality-specific clustering labels and manual annotations were obtained from the original study. The rightmost column shows the Allen Developing Mouse Brain Atlas reference annotations for the corresponding three developmental stages. \u003cstrong\u003eb \u003c/strong\u003eSpatial plots of the clusters obtained from all methods for the three sections. All methods share the color map. Manual annotation was performed to identify SpaMosaic’s clusters. \u003cstrong\u003ec\u003c/strong\u003eRNA expression and gene activity scores of marker genes \u003cem\u003eEmx1\u003c/em\u003e, \u003cem\u003eEmx2\u003c/em\u003e, \u003cem\u003eId4\u003c/em\u003e, \u003cem\u003eNr2f1\u003c/em\u003e and \u003cem\u003ePax6\u003c/em\u003e in SpaMosaic’s clusters. \u003cstrong\u003ed \u003c/strong\u003eSpatial heatmap of gene expression and gene activity scores of the marker genes on the E13.5 section. The bottom row shows the ISH images from the Allen Mouse Developing Brain Atlas for the corresponding markers. \u003cstrong\u003ee\u003c/strong\u003e Reference annotations of the E13.5 embryonic mouse brain obtained from the Allen Developing Mouse Brain Atlas. \u003cstrong\u003ef\u003c/strong\u003e ARI scores of the whole dataset, E15.5, E13.5 and E18.5 sections, respectively, for all methods. \u003cstrong\u003eg\u003c/strong\u003e Moran’s I scores of all methods on each section. \u003cstrong\u003eh \u003c/strong\u003eiLISI scores of all methods. \u003cstrong\u003ee\u003c/strong\u003e Modality alignment scores (FOSCTTM and MS) of all methods.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/f283c9887acfb23464a16f1a.png"},{"id":71218980,"identity":"7b529cc8-d6e6-4204-a1a1-69c6d5984d99","added_by":"auto","created_at":"2024-12-12 09:04:29","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1423847,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDownstream analysis results obtained with SpaMosaic and quantitative benchmarking on a postnatal mouse brain (RNA+H3K4me3) dataset. a \u003c/strong\u003eMosaic dataset configuration and clustering patterns obtained for all three sections. The RNA and H3K4me3 clusters were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on their respective batch-corrected low-dimensional embeddings from Harmony. The multi-modal cluster labels of section 1 were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the concatenated batch-corrected low-dimensional embeddings of the RNA and H3K4me3 profiles. \u003cstrong\u003eb \u003c/strong\u003eSpatial plots of the clusters obtained from all methods for all three sections. All methods share the color map. Manual annotation was performed to identify SpaMosaic’s clusters. \u003cstrong\u003ec\u003c/strong\u003e Allen Brain Atlas reference annotation for the postnatal mouse brain. \u003cstrong\u003ed\u003c/strong\u003e RNA expression of marker genes for the six brain regions in SpaMosaic’s clusters.\u003cstrong\u003e e\u003c/strong\u003e Spatial heatmap showing the gene activity scores of selected markers on the third section. \u003cstrong\u003ef \u003c/strong\u003eClusters overlaid on the image of the third section and cell types identified by label transfer. \u003cstrong\u003eg\u003c/strong\u003eMoran’s I scores of all methods across the three sections. We varied the clustering numbers (n=4, 6, 8, 10) and calculated the Moran’s I scores for each clustering result. In the box plots, the center lines indicate the median, boxes indicate the interquartile range, and whiskers indicate 1.5× interquartile range. \u003cstrong\u003eh \u003c/strong\u003eiLISI scores of all methods.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/495dd5235c300a13799ed8f1.png"},{"id":71218996,"identity":"a223571d-6d8f-4a57-9e58-10f45547dddf","added_by":"auto","created_at":"2024-12-12 09:04:30","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":61187710,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDownstream analysis results obtained with SpaMosaic and quantitative benchmarking on a human tonsil dataset. a \u003c/strong\u003eMosaic dataset configuration and clustering patterns obtained for all three sections. The RNA and protein clustering labels were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the batch-corrected low-dimension embeddings from Harmony. \u003cstrong\u003eb \u003c/strong\u003eSpatial plots of clustering labels from the methods for all three sections. All methods share the color map. Manual annotation was performed to identify SpaMosaic’s clusters. \u003cstrong\u003ec\u003c/strong\u003e RNA and protein expression of region-specific markers in SpaMosaic’s clusters. \u003cstrong\u003ed \u003c/strong\u003eARI scores on the whole dataset, section 1, section 2 and section 3 for all methods. The number of clusters was varied (n=4, 6, 8, 10) and we calculate the ARI for each clustering result. \u003cstrong\u003ee \u003c/strong\u003eMoran’s I scores for all methods across all sections. Similar to the ARI evaluation, the number of clusters was varied in the same fashion. \u003cstrong\u003ef \u003c/strong\u003eiLISI scores of all methods. \u003cstrong\u003eg \u003c/strong\u003eModality alignment scores of all methods (FOSCTTM and MS).\u003c/p\u003e","description":"","filename":"Fig5.png","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/b2fa32e2cb13751c34230c27.png"},{"id":71218984,"identity":"90ef9c0c-7a7a-49fe-a03e-a322119068ac","added_by":"auto","created_at":"2024-12-12 09:04:29","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":680065,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eApplication of SpaMosaic to five-modality mosaic integration.\u003c/strong\u003e \u003cstrong\u003ea\u003c/strong\u003e Mosaic dataset configuration and clustering patterns of the four sections. The RNA cluster labels were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the batch-corrected low-dimension embeddings from Harmony, and cluster labels for each epigenomic modality were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the principal components of the epigenomics profiles. The multi-modal clustering labels were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the concatenated embeddings of the RNA and epigenomics profiles. \u003cstrong\u003eb\u003c/strong\u003e Spatial plots of clusters obtained from the methods for all four sections. All methods share the color map. Manual annotation was performed to identify SpaMosaic’s clusters. \u003cstrong\u003ec\u003c/strong\u003e Allen Brain Atlas reference anatomic annotations for the postnatal mouse brain. \u003cstrong\u003ed\u003c/strong\u003e RNA expression of brain region markers in SpaMosaic’s clusters. \u003cstrong\u003ee\u003c/strong\u003e Moran’s I scores for the four sections. The number of clusters was varied (n=6, 8, 10, 12). \u003cstrong\u003ef\u003c/strong\u003e iLISI scores of all methods. \u003cstrong\u003eg\u003c/strong\u003e Relationships of the measured ATAC GAS, imputed H3K4me3 GAS and imputed H3K27ac GAS with the RNA in the CCG region of the fourth section. Only region-specific genes were kept for all modalities if their adjusted \u003cem\u003ep\u003c/em\u003e value was less than 1e-5 and the absolute value of log\u003csub\u003e2\u003c/sub\u003e fold change was larger than 0.5. \u003cstrong\u003eh\u003c/strong\u003e Upset plot of genes in quadrant I and quadrant III of all comparisons in \u003cstrong\u003eg\u003c/strong\u003e. \u003cstrong\u003ei\u003c/strong\u003e Enriched GO terms for genes shared across quadrant I of all comparisons, as well as for genes unique to quadrant I of the H3K27ac and H3K4me3 comparison. \u003cstrong\u003ej\u003c/strong\u003e Relationships of imputed H3K27me3 GAS, imputed H3K4me3 GAS and imputed H3K27ac GAS with measured ATAC GAS in the CCG region of the fourth section. Only region-specific genes were kept if their adjusted \u003cem\u003ep\u003c/em\u003e value was less than 1e-5 and the absolute value of log\u003csub\u003e2\u003c/sub\u003e fold change was larger than 0.5. \u003cstrong\u003ek\u003c/strong\u003e Enriched GO terms for the genes in quadrant IV of H3K27me3 vs ATAC, quadrant I of H3K4me3 vs ATAC and quadrant I of H3K27ac vs ATAC comparisons in\u003cstrong\u003e j\u003c/strong\u003e. \u003cstrong\u003el \u003c/strong\u003eRNA prediction error and Pearson correlation coefficient between measured RNA expression and predicted expression using different epigenomic modalities as input. For example, ‘H3K27me3 (section 1)’ denotes the result for predicting the RNA expression of the first section which also has a measured H3K27me3 modality. ‘Single modality’ indicates that the prediction model only used the measured H3K27me3 CSS while ‘four modalities’ indicates that the measured H3K27me3 CSS and all other imputed histone modification/ATAC GAS were used for prediction.\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/2665f6a5f906c4791d6e3425.png"},{"id":71218979,"identity":"55187517-f0d0-4484-93e1-eda7565326fe","added_by":"auto","created_at":"2024-12-12 09:04:29","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":169629,"visible":true,"origin":"","legend":"Supplementary material for the manuscript","description":"","filename":"SupplementaryInformation.docx","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/29ac9f24bc0f4fef23d75963.docx"},{"id":71218981,"identity":"400aaac7-c41d-4b38-997b-94e8c9fc6759","added_by":"auto","created_at":"2024-12-12 09:04:29","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":744457,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupp. Fig. S1 Graph neural network architecture and imputation strategy. a \u003c/strong\u003eAggregation and propagation processes of weighted LGCN. The weight for spatial neighbors is set to 1 while the weight for expression neighbors is set to 0.8. In each layer, each node aggregates the information from its neighbors using the weighted average. After two layers, the information from each node is concatenated across three output layers and then processed by a multiple layer perceptron. \u003cstrong\u003eb\u003c/strong\u003eImputation process. Taking section 3 as an example, which only has ATAC modality, we aim to impute its RNA and protein profiles. To impute the RNA profile of each spot in section 3, we find its RNA-specific neighbors from the other sections and then compute the average target expression profiles of these neighbors. The same approach is also taken for other modalities.\u003c/p\u003e","description":"","filename":"SuppFigS1.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/f3f8676ac99f556eaf6063a1.pdf"},{"id":71220626,"identity":"55318ea5-3b77-48f2-8013-1287a93ddf41","added_by":"auto","created_at":"2024-12-12 09:12:30","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":22138198,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupp. Fig. S2 Simulation settings and plots for competing methods. a\u003c/strong\u003e Five sets of simulation parameters used for data generation. \u003cstrong\u003eb\u003c/strong\u003e Mosaic configuration of each dataset. A dotted box denotes that this profile is missing. \u003cstrong\u003ec\u003c/strong\u003e ARI scores of clusters obtained by \u003cem\u003ek-\u003c/em\u003emeans and mclust clustering of the embeddings from the different methods and the raw data. \u003cstrong\u003ed\u003c/strong\u003e UMAP plots of the embeddings from all benchmarked methods. Spots are colored by section (first row), cluster (second row) and factor (third row). \u003cstrong\u003ee\u003c/strong\u003e Spatial plots of the clustering results of all methods. All methods and sections share the same color map.\u003c/p\u003e","description":"","filename":"SuppFigS2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/aa2406c0116f77329daf3388.pdf"},{"id":71220625,"identity":"07f546e6-c332-4dec-8e27-d9bc49d2af50","added_by":"auto","created_at":"2024-12-12 09:12:30","extension":"pdf","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":26082943,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupp. Fig. S3\u003c/strong\u003e \u003cstrong\u003eAdditional results from benchmarking with an embryonic mouse brain (RNA+ATAC) dataset.\u003c/strong\u003e \u003cstrong\u003ea\u003c/strong\u003e UMAP plots of the raw RNA and ATAC profiles. Spots are colored by section. \u003cstrong\u003eb\u003c/strong\u003e UMAP plots of the embeddings from all tested methods. Spots in the first row are colored by section, the second row by cluster and the third row by manual annotation. \u003cstrong\u003ec \u003c/strong\u003eMosaic dataset for evaluation with modality alignment metrics. The RNA and ATAC profiles of the E18.5 section were split into two separate sections to simulate originating from different experiments. Cluster labels were obtained from the original study. \u003cstrong\u003ed\u003c/strong\u003eUMAP plots of the raw RNA and ATAC profiles of the E15.5 and E18.5 sections. Spots are colored by section.\u003c/p\u003e","description":"","filename":"SuppFigS3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/d0b76f8ee181c983ba3fc0ea.pdf"},{"id":71219027,"identity":"dea61dfc-eecc-4b8c-a1e1-3ff30925ac9d","added_by":"auto","created_at":"2024-12-12 09:04:31","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":25875221,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupp. Fig. S4 Additional results of a postnatal mouse brain dataset (RNA+H3K4me3) benchmarking. a \u003c/strong\u003eUMAP plots of the raw RNA and H3K4me3 profiles. Spots are colored by section. \u003cstrong\u003eb \u003c/strong\u003eUMAP plots of the embeddings from all methods. Spots in the first row are colored by section, the second row by cluster.\u003c/p\u003e","description":"","filename":"SuppFigS4.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/280913d5e2a163fef2bc3f1a.pdf"},{"id":71220640,"identity":"596d98fc-e8f2-49cc-a580-480e28996800","added_by":"auto","created_at":"2024-12-12 09:12:31","extension":"pdf","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":65421498,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupp. Fig. S5\u003c/strong\u003e \u003cstrong\u003eDownstream analysis results obtained with SpaMosaic and quantitative benchmarking on a postnatal mouse brain (RNA+H3K27me3) dataset. a \u003c/strong\u003eMosaic dataset configuration and clustering patterns obtained for all three sections. The RNA and H3K27me3 cluster labels were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the batch-corrected low-dimensional embeddings from Harmony. The multi-modal cluster labels of section 1 were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the concatenated batch-corrected low-dimensional embeddings of the RNA and H3K27me3 profiles. \u003cstrong\u003eb \u003c/strong\u003eSpatial plots of the clusters from the methods for all three sections. All methods share the color map. Manual annotation was performed to identify SpaMosaic’s clusters. \u003cstrong\u003ec\u003c/strong\u003e RNA expression of marker genes of the six brain regions in SpaMosaic’s clusters. \u003cstrong\u003ed\u003c/strong\u003e Clusters overlaid on the image of the third section (top row) and spatial heatmap showing the gene activity scores of selected markers on the third section (middle and bottom rows). \u003cstrong\u003ee\u003c/strong\u003e Moran’s I scores of all methods across the three sections. We varied the clustering numbers (n=4, 6, 8, 10) and calculated the Moran’s I scores for each clustering result. \u003cstrong\u003ef \u003c/strong\u003eiLISI scores of all methods.\u003c/p\u003e","description":"","filename":"SuppFigS5.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/45629b87f0ac4f3e8480d937.pdf"},{"id":71220641,"identity":"c78fc330-8a82-4f11-af39-0e055b6e1edd","added_by":"auto","created_at":"2024-12-12 09:12:31","extension":"pdf","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":24042863,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupp. Fig. S6 Additional results of a postnatal mouse brain dataset (RNA+H3K27me3) benchmarking. a \u003c/strong\u003eUMAP plots of the raw RNA and H3K27me3 profiles. Spots are colored by section. \u003cstrong\u003eb \u003c/strong\u003eUMAP plots of the embeddings from all methods. Spots in the first row are colored by section, the second row by cluster.\u003c/p\u003e","description":"","filename":"SuppFigS6.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/5f2ffa841fdebda37212e280.pdf"},{"id":71219017,"identity":"064923c3-7d51-43bd-976a-7bd7fdc62c4f","added_by":"auto","created_at":"2024-12-12 09:04:31","extension":"pdf","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":45121674,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupp. Fig. S7\u003c/strong\u003e \u003cstrong\u003eDownstream analysis results obtained with SpaMosaic and quantitative benchmarking on a postnatal mouse brain dataset (RNA+ATAC). a \u003c/strong\u003eMosaic dataset configuration and clustering patterns obtained for all three sections. The RNA and ATAC cluster labels were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the batch-corrected low-dimensional embeddings from Harmony. The multi-modal cluster labels of section 1 were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the concatenated batch-corrected low-dimensional embeddings of the RNA and ATAC profiles. \u003cstrong\u003eb \u003c/strong\u003eSpatial plots of the clusters obtained from the methods for all three sections. All methods share the color map. Manual annotation was performed to identify SpaMosaic. \u003cstrong\u003ec\u003c/strong\u003e Heatmap of the DEGs of SpaMosaic’s clusters in section 1 and section 2. \u003cstrong\u003ed\u003c/strong\u003e Clusters overlaid on the image of the third section (left) and spatial heatmap depicting the gene activity scores of selected markers on the third section (middle and right). \u003cstrong\u003ee\u003c/strong\u003e Moran’s I scores of all methods across the three sections. We varied the clustering numbers (n=4, 6, 8, 10) and calculated the Moran’s I scores for each clustering result. \u003cstrong\u003ef \u003c/strong\u003eiLISI scores of all methods. \u003cstrong\u003eg \u003c/strong\u003eMosaic dataset for evaluation with modality alignment metrics. Raw RNA and ATAC profiles of the second section were split into two separate sections to simulate originating from different experiments. Cluster labels were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the batch-corrected low-dimensional embeddings from Harmony. \u003cstrong\u003eh\u003c/strong\u003eUMAP plots of the raw RNA profiles and ATAC profiles in \u003cstrong\u003eg\u003c/strong\u003e. Spots are colored by section. \u003cstrong\u003ei\u003c/strong\u003e Modality alignment scores (FOSCTTM and MS) of all methods.\u003c/p\u003e","description":"","filename":"SuppFigS7.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/b32e8d4ecdccb1c820d18653.pdf"},{"id":71218985,"identity":"d684d447-fbfd-4c82-9293-7bc989ef84f6","added_by":"auto","created_at":"2024-12-12 09:04:29","extension":"pdf","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":16885611,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupp. Fig. S8 Additional results of a postnatal mouse brain dataset (RNA+ATAC) benchmarking. a \u003c/strong\u003eUMAP plots of the raw RNA and ATAC profiles. Spots are colored by section. \u003cstrong\u003eb \u003c/strong\u003eUMAP plots of the embeddings from all methods. Spots in the first row are colored by section, the second row by cluster.\u003c/p\u003e","description":"","filename":"SuppFigS8.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/19f4debe8166eb593e5f1637.pdf"},{"id":71220639,"identity":"c62888c5-3842-43fa-b11a-0c81e4ee5ef8","added_by":"auto","created_at":"2024-12-12 09:12:31","extension":"pdf","order_by":10,"title":"","display":"","copyAsset":false,"role":"supplement","size":46271722,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupp. Fig. S9\u003c/strong\u003e \u003cstrong\u003eDownstream analysis results obtained with SpaMosaic and quantitative benchmarking on a postnatal mouse brain (RNA+H3K27ac) dataset. a \u003c/strong\u003eMosaic dataset configuration and clustering patterns obtained for all three sections. The RNA and H3K27ac cluster labels were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the batch-corrected low-dimensional embeddings from Harmony. The multi-modal cluster labels of section 1 were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the concatenated batch-corrected low-dimensional embeddings of the RNA and H3K27ac profiles. \u003cstrong\u003eb \u003c/strong\u003eSpatial plots of the clustering labels from the methods for all three sections. All methods share the color map. Manual annotation was performed to identify SpaMosaic’s clusters. \u003cstrong\u003ec\u003c/strong\u003e Heatmap of the DEG of SpaMosaic’s clusters in section 1 and section 2. \u003cstrong\u003ed\u003c/strong\u003e Clusters overlaid on the image of the third section (top left) and spatial heatmap showing the gene activity scores of selected markers on the third section. \u003cstrong\u003ee\u003c/strong\u003e Moran’s I scores of all methods across the three sections. We varied the clustering numbers (n=4, 6, 8, 10) and calculated the Moran’s I scores for each clustering result. \u003cstrong\u003ef \u003c/strong\u003eiLISI scores of all methods. \u003cstrong\u003eg \u003c/strong\u003eMosaic dataset for evaluation with modality alignment metrics. Raw RNA and H3K27ac profiles of the second section were split into two separate sections to simulate originating from different experiments. Cluster labels were obtained by performing \u003cem\u003ek\u003c/em\u003e-means clustering on the batch-corrected low-dimensional embeddings from Harmony.\u003cstrong\u003e h\u003c/strong\u003eUMAP plots of the raw RNA and H3K27ac profiles in \u003cstrong\u003eg\u003c/strong\u003e. Spots are colored by section. \u003cstrong\u003ei\u003c/strong\u003e Modality alignment scores of all methods (FOSCTTM and MS).\u003c/p\u003e","description":"","filename":"SuppFigS9.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5507983/v1/0f6a4a10ef25c2f93200e349.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Mosaic integration of spatial multi-omics with SpaMosaic","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":false,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"spatial multi-omics, mosaic integration, graph neural network, contrastive learning","lastPublishedDoi":"10.21203/rs.3.rs-5507983/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5507983/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"With the advent of spatial multi-omics, we can mosaic integrate diverse datasets with partially overlapping modalities to construct consensus multi-modal spatial atlases of the source tissue. SpaMosaic is a spatial multi-omics mosaic integration tool that employs contrastive learning and graph neural networks to construct a modality-agnostic and batch-corrected latent space suited for analyses like spatial domain identification and imputing missing omes. Using simulated and experimentally acquired datasets, we benchmarked SpaMosaic against single-cell multi-omics mosaic integration methods. The experimental data encompassed RNA and protein abundance, chromatin accessibility or histone modifications, acquired from brain, embryo, tonsil, and lymph node tissues. SpaMosaic achieved superior performance over existing methods in identifying known spatial domains with enhanced resolution and clarity while reducing noise and batch effects. It also ranked top for modality alignment, enabling seamless diagonal integration without the need for image registration. After integration, SpaMosaic can also impute missing modalities. With a mosaic set of mouse brain data with RNA and different epigenomic modalities, we integrated and imputed the missing omics. There we found the imputed gene activity scores of activating and silencing histone marks show the correct correlation with RNA. Moreover, we uncovered more region-specific genes and pathways showing the correct transcriptome-epigenome correlations in the imputed histone modification than in the measured chromatin accessibility modalities. Lastly, SpaMosaic’s imputation also allows the inference of relationships between different modalities without requiring co-profiling from the same section.","manuscriptTitle":"Mosaic integration of spatial multi-omics with SpaMosaic","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-12-12 09:04:24","doi":"10.21203/rs.3.rs-5507983/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"nature-genetics","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"ng","sideBox":"Learn more about [Nature Genetics](http://www.nature.com/ng/)","snPcode":"","submissionUrl":"","title":"Nature Genetics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Research","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"155f37af-a90e-491c-97e2-53477b4bcded","owner":[],"postedDate":"December 12th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":41436465,"name":"Biological sciences/Computational biology and bioinformatics/Data mining"},{"id":41436466,"name":"Biological sciences/Computational biology and bioinformatics/Data processing"}],"tags":[],"updatedAt":"2026-04-25T07:11:04+00:00","versionOfRecord":{"articleIdentity":"rs-5507983","link":"https://doi.org/10.1038/s41588-026-02573-3","journal":{"identity":"nature-genetics","isVorOnly":false,"title":"Nature Genetics"},"publishedOn":"2026-04-24 04:00:00","publishedOnDateReadable":"April 24th, 2026"},"versionCreatedAt":"2024-12-12 09:04:24","video":"","vorDoi":"10.1038/s41588-026-02573-3","vorDoiUrl":"https://doi.org/10.1038/s41588-026-02573-3","workflowStages":[]},"version":"v1","identity":"rs-5507983","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5507983","identity":"rs-5507983","version":["v1"]},"buildId":"Ocjqolx5-0U8gK8KHECj2","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

crossref
last seen: 2026-08-26T06:22:55.981938+00:00
europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0