Hundreds of cardiac MRI traits derived using 3D diffusion autoencoders share a common genetic architecture

preprint OA: closed

Abstract

Biobank-scale imaging provides an unprecedented opportunity to characterise thousands of organ phenotypes, how they vary in populations and how they relate to disease outcomes. However, deriving specific phenotypes from imaging data, such as Magnetic Resonance Imaging (MRI), requires time-consuming expert annotation, limiting scalability, and does not exploit how information-dense such image acquisitions are. In this study, we developed a 3D diffusion autoencoder to derive latent phenotypes from temporally resolved cardiac MRI data of 71,021 UK Biobank participants. These phenotypes were reproducible, heritable ( h 2 = [4 - 18%]), and significantly associated with cardiometabolic traits and outcomes, including atrial fibrillation ( P = 8.5 × 10 −29 ) and myocardial infarction ( P = 3.7 × 10 −12 ). By using latent space manipulation techniques, we were able to learn, directly interpret and visualise what specific latent phenotypes are capturing in a given MRI. To establish the genetic basis of such traits, we performed a genome-wide association study, identifying 89 significant common variants ( P 0.8) linked variants across phenotypic scales, from intermediate cardiac traits to cardiac disease endpoints. For example, rs142556838 that falls in CCDC141 colocalises with a latent imaging phenotype and a diastolic blood pressure locus. Using single-cell RNA-sequencing data we map CCDC141 expression specifically to a population of ventricular cardiomyocytes. Finally, Polygenic Risk Scores (PRS) derived from latent phenotypes demonstrated predictive power for a range of cardiometabolic diseases and enabled us to successfully stratify the individuals into different risk groups. In conclusion, this study showcases the use of diffusion autoencoding methods as powerful tools for unsupervised phenotyping, genetic discovery and disease risk prediction using cardiac MRI data.
Full text 3,715 characters · extracted from oa-doi-fallback · click to expand
Abstract Biobank-scale imaging provides an unprecedented opportunity to characterise thousands of organ phenotypes, how they vary in populations and how they relate to disease outcomes. However, deriving specific phenotypes from imaging data, such as Magnetic Resonance Imaging (MRI), requires time-consuming expert annotation, limiting scalability, and does not exploit how information-dense such image acquisitions are. In this study, we developed a 3D diffusion autoencoder to derive latent phenotypes from temporally resolved cardiac MRI data of 71,021 UK Biobank participants. These phenotypes were reproducible, heritable (h2 = [4 - 18%]), and significantly associated with cardiometabolic traits and outcomes, including atrial fibrillation (P = 8.5 × 10−29) and myocardial infarction (P = 3.7 × 10−12). By using latent space manipulation techniques, we were able to learn, directly interpret and visualise what specific latent phenotypes are capturing in a given MRI. To establish the genetic basis of such traits, we performed a genome-wide association study, identifying 89 significant common variants (P 0.8) linked variants across phenotypic scales, from intermediate cardiac traits to cardiac disease endpoints. For example, rs142556838 that falls in CCDC141 colocalises with a latent imaging phenotype and a diastolic blood pressure locus. Using single-cell RNA-sequencing data we map CCDC141 expression specifically to a population of ventricular cardiomyocytes. Finally, Polygenic Risk Scores (PRS) derived from latent phenotypes demonstrated predictive power for a range of cardiometabolic diseases and enabled us to successfully stratify the individuals into different risk groups. In conclusion, this study showcases the use of diffusion autoencoding methods as powerful tools for unsupervised phenotyping, genetic discovery and disease risk prediction using cardiac MRI data. Competing Interest Statement The authors have declared no competing interest. Funding Statement This study did not recieve any funding. Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This research has been conducted using the UK Biobank Resource under application number 82779. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Footnotes {sara.ometto{at}fht.org, soumick.chatterjee{at}fht.org, craig.glastonbury{at}fht.org} ↵5 Lead contact New title; Updated editorial changes (wording); Results for figure 2 now quantitative;

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00