Genetic underpinnings of predicted changes in cardiovascular function using self supervised learning

preprint OA: gold CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
⚙ AI-generated summary by qwen3.7-flash, 2026-09-10 ⓘ

This study links temporal cardiovascular state changes to genetics via Delta ECG, identifying genome-wide significant associations and gene expression patterns related to electron transport and immune pathways relevant to cardiovascular disease.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

⚙ AI-generated deep summary by qwen3.7-flash, 2026-09-24 · read from full text ⓘ

This study utilized self-supervised learning to analyze electrocardiogram signals from the Human Phenotype Project, creating a metric called Delta ECG that quantifies temporal shifts in cardiovascular state over two years. The researchers performed genome-wide association studies on this metric and predicted gene expression in peripheral blood mononuclear cells, identifying significant genetic links to the electron transport chain and immune pathways involving eosinophils and mast cells. The findings demonstrate that temporal changes in cardiovascular health share a genetic basis with cardiovascular disease risk factors and known correlates, validating the integration of deep learning embeddings with genetic analysis. Relevance to endometriosis: The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Background The genetic underpinnings of cardiovascular disease remain elusive. Contrastive learning algorithms have recently shown cutting-edge performance in extracting representations from electrocardiogram (ECG) signals that characterize cross-temporal cardiovascular state. However, there is currently no connection between these representations and genetics. Methods We designed a new metric, denoted as Delta ECG, which measures temporal shifts in patients’ cardiovascular state, and inherently adjusts for inter-patient differences at baseline. We extracted this measure for 4,782 patients in the Human Phenotype Project using a novel self-supervised learning model, and quantified the associated genetic signals with Genome-Wide-Association Studies (GWAS). We predicted the expression of thousands of genes extracted from Peripheral Blood Mononuclear Cells (PBMCs). Downstream, we ran enrichment and overrepresentation analysis of genes we identified as significantly predicted from ECG. Findings In a Genome-Wide Association Study (GWAS) of Delta ECG, we identified five associations that achieved genome-wide significance. From baseline embeddings, our models significantly predict the expression of 57 genes in men and 9 in women. Enrichment analysis showed that these genes were predominantly associated with the electron transport chain and the same immune pathways as identified in our GWAS. Conclusions We validate a novel method integrating self-supervised learning in the medical domain and simple linear models in genetics. Our results indicate that the processes underlying temporal changes in cardiovascular health share a genetic basis with CVD, its major risk factors, and its known correlates. Moreover, our functional analysis confirms the importance of leukocytes, specifically eosinophils and mast cells with respect to cardiac structure and function.
Full text 46,621 characters · extracted from preprint-html · click to expand
Genetic underpinnings of predicted changes in cardiovascular function using self supervised learning | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results Genetic underpinnings of predicted changes in cardiovascular function using self supervised learning View ORCID Profile Zachary Levine , Guy Lutsker , Anastasia Godneva , Adina Weinberger , Maya Pompan , Yeela Talmor-Barkan , View ORCID Profile Yotam Reisner , Hagai Rossman , View ORCID Profile Eran Segal doi: https://doi.org/10.1101/2024.08.15.608061 Zachary Levine 1 Department of Computer Science and Applied Mathematics, Weizmann Institute of Science , Rehovot, 76100, Israel Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Zachary Levine Guy Lutsker 1 Department of Computer Science and Applied Mathematics, Weizmann Institute of Science , Rehovot, 76100, Israel Find this author on Google Scholar Find this author on PubMed Search for this author on this site Anastasia Godneva 1 Department of Computer Science and Applied Mathematics, Weizmann Institute of Science , Rehovot, 76100, Israel Find this author on Google Scholar Find this author on PubMed Search for this author on this site Adina Weinberger 1 Department of Computer Science and Applied Mathematics, Weizmann Institute of Science , Rehovot, 76100, Israel 3 Pheno.AI , Tel-Aviv, Israel Find this author on Google Scholar Find this author on PubMed Search for this author on this site Maya Pompan 1 Department of Computer Science and Applied Mathematics, Weizmann Institute of Science , Rehovot, 76100, Israel Find this author on Google Scholar Find this author on PubMed Search for this author on this site Yeela Talmor-Barkan 1 Department of Computer Science and Applied Mathematics, Weizmann Institute of Science , Rehovot, 76100, Israel 4 Sackler Faculty of Medicine, Tel Aviv University , Tel-Aviv, 6997801, Israel 5 Department of Cardiology, Rabin Medical Center , Petah-Tikva, 49100, Israel Find this author on Google Scholar Find this author on PubMed Search for this author on this site Yotam Reisner 1 Department of Computer Science and Applied Mathematics, Weizmann Institute of Science , Rehovot, 76100, Israel 3 Pheno.AI , Tel-Aviv, Israel Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Yotam Reisner Hagai Rossman 3 Pheno.AI , Tel-Aviv, Israel Find this author on Google Scholar Find this author on PubMed Search for this author on this site Eran Segal 1 Department of Computer Science and Applied Mathematics, Weizmann Institute of Science , Rehovot, 76100, Israel Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Eran Segal For correspondence: eran.segal{at}weizmann.ac.il Abstract Full Text Info/History Metrics Preview PDF Abstract Background The genetic underpinnings of cardiovascular disease remain elusive. Contrastive learning algorithms have recently shown cutting-edge performance in extracting representations from electrocardiogram (ECG) signals that characterize cross-temporal cardiovascular state. However, there is currently no connection between these representations and genetics. Methods We designed a new metric, denoted as Delta ECG, which measures temporal shifts in patients’ cardiovascular state, and inherently adjusts for inter-patient differences at baseline. We extracted this measure for 4,782 patients in the Human Phenotype Project using a novel self-supervised learning model, and quantified the associated genetic signals with Genome-Wide-Association Studies (GWAS). We predicted the expression of thousands of genes extracted from Peripheral Blood Mononuclear Cells (PBMCs). Downstream, we ran enrichment and overrepresentation analysis of genes we identified as significantly predicted from ECG. Findings In a Genome-Wide Association Study (GWAS) of Delta ECG, we identified five associations that achieved genome-wide significance. From baseline embeddings, our models significantly predict the expression of 57 genes in men and 9 in women. Enrichment analysis showed that these genes were predominantly associated with the electron transport chain and the same immune pathways as identified in our GWAS. Conclusions We validate a novel method integrating self-supervised learning in the medical domain and simple linear models in genetics. Our results indicate that the processes underlying temporal changes in cardiovascular health share a genetic basis with CVD, its major risk factors, and its known correlates. Moreover, our functional analysis confirms the importance of leukocytes, specifically eosinophils and mast cells with respect to cardiac structure and function. Introduction Electrocardiography (ECG) is a critical tool in diagnostic cardiology, and exploring the genetic underpinnings of CVD though it remains an open but promising problem 1 . Existing works that apply deep learning to ECG either explore genetics but fix a supervisory signal in the form of target labels for age 2 and Atrial Fibrillation (AF) 3 , or learn embeddings using Self-Supervised-Learning (SSL), but do not consider genetic implications of changes in these features over time 4 , 5 . As well, these supervisory signals potentially bias learned features away from general health state, towards these individual target labels 3 . The branch of machine learning algorithms that use neural networks to learn discriminative features from data without target supervisory labels is known as SSL. These feature vectors, often denoted as embeddings, are traditionally learned through masked signal modeling or contrastive learning 6 . We used ECG recordings and genetics data (DNA, RNA expression) from the Human Phenotype Project (HPP) for this project. The HPP is a large-scale longitudinal study, centered on the deep phenotyping of people between 22 and 70 years of age in Israel. Over the course of several years, the study has collected a wide range of clinical data and biological data, and seeks to discover correlates for disease and targets for disease 7 . Recordings were collected at intake appointments, and repeated two years later for ~40% of the study cohort (see Methods). When neural networks are trained based on contrastive learning, coordinates in the latent space of the model are expected to correspond to various parameters of a patient’s health state 4 , 5 . We quantified each patient’s temporal shift in cardiovascular state through the distance between their baseline and follow up ECG embeddings. We defined our new metric, Delta ECG, as the cosine similarity between the embeddings of first and second appointments for each patient. This single number measures a patient’s changes in cardiovascular performance over a two year period, relative to their intake visit. As such, it functions as a valuable measure of cardiovascular aging over time, while inherently adjusting for baseline patient health differences across the HPP population. In our GWAS of Delta ECG, five SNPS reached genome-wide significance. Our pathway analyses and subsequent RNA expression predictions indicate that Delta ECG shares a genetic basis with CVD, its major risk factors, and its known correlates in the immune system. Our raw embeddings were directly able to predict the expression of genes along the pathways containing the GWAS hits for Delta ECG, validating both our findings and our method. Results Description of the Pooled HPP Cohort The present study used the cohort from the Human Phenotype Project, utilizing a pooled cohort consisting of 16,775 ECG recordings from 11,933 patients (diseased and healthy), subdivided into two non overlapping four second windows each. For a complete list of exclusion criteria, we direct readers to the original paper describing the cohort 7 . This distribution, along with descriptive summary statistics, can be found in Tables 1 and 2 . To understand the vast complexities of human health, including both function and dysfunction we pooled diseased and healthy individuals into one large dataset to train our model. View this table: View inline View popup Download powerpoint Table 1: Baseline features of included healthy cohorts, stratified by age. View this table: View inline View popup Download powerpoint Table 2: Baseline features of included disease cohorts. Self-Supervised Learning Objective In current literature, there are two leading methods of conducting self-supervised contrastive learning with ECG data: Contrastive Learning of Cardiac Signals (CLOCS/CMSC) 4 , and Patient Contrastive Learning (PCLR) 5 . To capitalize on the strengths of our dataset, a high number of repeat visits and high temporal resolution, we trained a model using a combination of both approaches. Overall, the training objective of the model, given two different views of an ECG signal, was to maximize the agreement between learned representations if those two windows are taken from the same patient, and to minimize them otherwise. Our complete training and validation curves can be found in Figure 2 , and architecture specifications/hyperparameters can be found in the Method section of this work. Download figure Open in new tab Figure 1: Our model correctly matches patients to their corresponding ECG recordings in the latent space, both across appointments and within them. We split ECG recordings in time, and passed them through our CNN/Transformer network. We trained our model based on a combination of the two leading contrastive learning algorithms for ECGs: Contrastive Multi-Segment Coding (bottom, left) and Patient Contrastive Learning (bottom, right). Our patient identification performance on these two tasks is shown in the associated barplot. Download figure Open in new tab Figure 2: The model converges based on contrastive loss after many iterations. a) Training loss (blue) and validation loss (purple) of the trained model. We initialized the model (Table S1, See Supplement) with random weights, and trained using the AdamW, with an initial learning rate of 1e-5. We reduced the learning rate by a factor of 5 for every 5 epochs that the validation loss did not decrease, and after 15 epochs of non-decreasing validation loss, we stopped training the model, and checkpointed the weights for later use. Model Validation: Patient Identification We trained our neural network (Table S1, See Supplement) until satisfaction of the stopping criteria (see Methods). To validate our model, we applied metrics over our embedding space that checked for how many patients, the closest embedding vector in the latent space to their first appointment was their second one (Top-1 across). We did the same for the matching of first and second halves of each recording (Top-1 within) 8 . Top-10 accuracy (in either case) was a relaxation so that the correct embedding need only be among the top 10 closest vectors. Our trained model achieved (test) accuracies of 57% (top-1) and 93% (top-10) across appointments. Within appointments, we achieved 90% (top-1) and 99%, (top-10). One might wonder how this performs relative to random assignment: 57% “across” patient identification is 150 times better than random. These results indicate that our learned representations effectively characterize patient health, and serve as validation for our model. The HPP Delta ECG distribution is as expected We extracted two distributions: the Delta ECG for each participant and the null distribution, which was the distance between all first and all second appointments, regardless of patient identity. These two distributions are shown in Figure 3 . We observed that the Delta ECG distribution is significantly to the right of the null one, which is centered around zero. This indicates that embeddings from a patient’s first appointment are much closer to embeddings from that patient’s second appointment than those from all other people. Download figure Open in new tab Figure 3: Patients are correctly matched to their repeat visits in the latent space Distribution (density) of cosine similarity between first and second visits of matched (green) and all (blue) appointments [the null distribution]. GWAS of Delta ECG: Five SNPs reach genome-wide significance Starting from the DNA level, we sought to understand the genetic associations for our embeddings. We adjusted for age, gender, and the top 10 Principal Components (PCs) of the variance-standardized relationship matrix, and pruned for first degree relations. Complete methods describing GWAS protocols and our thresholds for the HPP have been detailed previously 9 . There were no multi-trait hits in GWAS of the top 10 principal components from either first or second appointment embeddings. However, we identified five genome-wide significant variants associated with Delta ECG (see Introduction) ( Figure 4 and Table 3 ). We clumped the full unbiased GWAS results for Linkage Disequilibrium (LD), forming 629 clumps from the 1,806 top variants. Most importantly, all five of these SNPs were put in different clumps. Download figure Open in new tab Figure 4: Changes in ECG cardiovascular state share a genetic basis with CVD, its known correlates, and its major risk factors Manhattan plot of Delta ECG showing the top SNP associations for this trait, as assessed by our GWAS. Names of genes are labeled above each SNP they contain, and are highlighted in yellow. View this table: View inline View popup Download powerpoint Table 3: Delta ECG shares a genetic basis with CVD and its known risk factors and correlates. Significant hits for Delta ECG and their previous genetic associations. Delta ECG shares genetic underpinnings with CVD through ARID5B With respect to our GWAS results, no SNP had been previously reported in any GWAS. We thus performed our analyses on the gene level. Our strongest hit (rs139222531, P < 4e-9) is in ARID5B , a gene with previously established significant hits for CVD within the UK Biobank 10 . ARID5B encodes a protein that belongs to the AT-rich interaction domain (ARID) family. Previous studies have identified this gene as playing a role in the regulation of inflammation and immune responses with respect to atherosclerosis 11 . The identification of a previously established significant hit for CVD within our framework validates the usefulness of our model. Common genetic architecture between Delta ECG and glycemic CVD risk factors through SLC30A8 However, beyond CVD itself, we identified a significant hit in SLC30A8 , (rs117871919, P < 3e-8). SLC30A8 encodes a zinc transporter protein that is primarily expressed in pancreatic beta cells. SLC3A03 has known genetic associations for glycemic variance and control, namely HbA1c 12 , BMI 12 , and T2D 13 , which is a well-established risk factor for CVD 14 . These findings indicate that Delta ECG shares a genetic signal not just with CVD, but with its significant risk factors as well. Delta ECG also shares a genetic basis with known correlates of CVD through CDKL3 and MEGF6 We found a hit (rs113004948, P < 4e-8) in CDKL3 , a gene that encodes a protein belonging to the cyclin-dependent kinase (CDK) family. CDKL3 has previously established significant hits for Eosinophil counts 15 , which are known to have a strong relationship with frequency of cardiac complications (Bozkuş et al.). As well, our weakest but still genome-wide significant hit (rs74469838, P < 5e-8) is in MEGF6 , which has existing strong associations for Systemic Mastocytosis 16 , another correlate of CVD 17 . These results indicate that Delta ECG shares a genetic basis with correlates of CVD. We could predict the expression of 66 genes significantly from our model We wanted to understand the functional predictive power of our embeddings on a cellular level. In the HPP, RNA expression levels are collected from a sample containing a bulk of peripheral blood mononuclear cells (PBMCs) in a process known as RNA-Sequencing, or RNASeq (see Methods). After correction, in men, we were able to predict 57 genes significantly at a threshold of P<0.05, and 9 genes at the more stringent (0.01) level. In women, we could predict 9 genes at 0.05, and 0 at the more stringent testing threshold, after correction. These genes are displayed in Table S2. Among them are prominent mitochondrial genes such as MT-RNR1 / 2 and MT-TV . We aimed to explore which biological pathways were implicated amongst the genes whose expression values we could predict significantly in men and women. To this end, we ran overrepresentation analysis over three major gene sets: GO Biological Processes 2023, MSigDB_Hallmark_2020, and Kegg 2021 Human for the genes we could predict significantly (P* < 0.05) in both men and women. For women, there were no significant pathways after FDR correction, however for men we found 19 pathways at P* < 0.05 and 3 at the stricter 0.005 level. The complete pathway set that came up as significant in men (P* < 0.05) is displayed in Figure 5 . Download figure Open in new tab Figure 5: The electron transport chain and immune response are implicated in learned ECG baseline features Significantly over-represented pathways among genes with well-predicted expression (in men) from ECG embeddings The ETC is centrally implicated in Cardiovascular Disease The most dominant pathway set from our predictions of gene expression is that pertaining to the mitochondrial electron transport chain, in which we identified 9 pathways that were significantly over represented. These are, including some pathways duplicated across multiple gene sets: Cellular Respiration (GO:0045333, P*<0.005), Oxidative Phosphorylation (KEGG:hsa00190, P* < 0.005 & GO:0006119, P* < 0.05) & MSigDB Hallmark, P*<0.05) Proton Motive Force-Driven Force-Driven ATP Synthesis (GO:0015986, P* < 0.05), Mitochondrial ATP Synthesis (GO:0042776, P* < 0.05), Mitochondrial ATP Synthesis Coupled Electron Transport (GO:0042776, P* < 0.05) Aerobic Electron Transport Chain (GO:0019646, P* < 0.05), and Aerobic Respiration (GO:0009060, P* < 0.05). This replicates existing findings on the central importance of mitochondrial function in cardiovascular disease development 18 . Embeddings directly predict expression of genes along pathways containing GWAS hits In our GWAS results, we found a hit in (rs113004948, P < 4e-8) in CDKL3 . This gene also has a pre-existing known genetic association for Eosinophil count, a known correlate for cardiac complications. From this, we concluded that Delta ECG shares a genetic signal with CVD correlates. However, from the baseline embeddings of patients ECG we could also predict genes in which the pathway for Negative Regulation Of Leukocyte Mediated Cytotoxicity (GO:0001910, P* < 0.05) was over represented. The same can be said for our hit in MEGF6 , which has known associations for Systemic Mastocytosis 16 , while for RNA in men, the pathway for Mast Cell Degranulation (GO:0043304, P* < 0.05) was significantly over represented in our gene expression predictions. These results indicate that changes in cardiovascular state over time do not just share a genetic basis with eosinophil count and mast cell functioning, but that extracted features from ECG using our model are directly predictive of biomarkers pertaining to the functional role of the two at the cellular level, even in PBMCs. Purine Metabolism: Results from an application of enrichment analysis Lastly, we sought to leverage the RNA expression values directly in our pathway analyses. To do this, we applied gene set enrichment analysis with the case/control label set to a binary discretization of the distance between a patient’s baseline and follow-up appointment (see Methods). Here, women had no pathways, but in men, one pathway came up as significant: Human Purine Metabolism (KEGG:hsa00230, P* < 0.05) whose connection to CVD has been well established previously 19 , 20 . This once again can serve as validation of our method. Discussion To understand the genetic basis of individual-level changes in cardiovascular state over time, we conducted genetic analyses on the difference in learned patient cardiovascular state from a novel deep learning model. There are a few previous works that combine cardiac SSL and genetic studies, though existing studies with ECG are largely cross-sectional, and focus on the factors that cause an between-patient differences in cardiovascular state at a single time point 1 , 21 . With respect to repeated ECG recordings over time, only one previous work utilized data (from the UK Biobank) from repeat appointments, however similarity to baseline was used as a quality-control filter as opposed to a learning signal 22 . Aging is best observed within a single patient over time. Our work is the first to consider the genetic underpinnings of cross-temporal contrastive representations from ECG of the same patient. As such, our representations are perhaps more holistic views of cardiovascular health than those pertaining to a single disease outcome. More generally, our framework can be applied to any medical modality to quantify and understand the genetic underpinnings of temporal changes in phenotypic state as assessed by medical tests. Still, our work is not without limitations. We used generalized patient identification metrics for each task separately to verify that our model was correctly meaningful patient-specific representation from ECGs. Our “intra appointment” performance was lower than “inter “tasks. This is expected: periodicity within ECG signals across time means that predicting the first half from the second half of an ECG signal is much easier than predicting what a recording a few years down the line and with slight differences in electrode placement in different measurement occasions will look like. Still, our performance on both embedding metrics is indicative of the powerful representations captured by our model. As well, the sex-based differences we found in the over-representation and enrichment analyses are a result of the poorer gene expression predictions in women. This may be a result of the well-established lower CVD prevalence in women. However, we also cannot exclude the possibility that this has occurred as a result of batch effects or other problems with our ECG recordings, RNASequencing experiment, or both. The majority of the pooled HPP cohort studied here was relatively healthy. More significant cardiovascular deterioration, and therefore larger Delta ECG would be expected in a cohort comprised of more individuals with CVD. Still, the fact that among largely healthy individuals we were still able to detect disease signals is further proof of the strength of our method. Numerous works have shown the dependency of contrastive learning methods on using large datasets, and thus potential results are always a function of cohort size. Transfer effects across different datasets within physiological deep learning have been found to be highly significant across datasets 23 , 24 and thus we trained our model on our dataset directly, as opposed to training on a larger collection of ECGs such as the UK Biobank. Several future directions are suggested for our work, the first of which is training on a larger sample size. This would not only benefit our SSL: our sample size was below 5000, which is the minimum at which traditional heritability estimation methods, i.e (LD-based estimation and GREML) are sufficiently powered. Fitting the model on a larger cohort would allow us to explore the heritability of Delta ECG, and the genetic correlation between Delta ECG and other traditional cardiac phenotypes. As well, exploring more than one time point per patient could enable deeper exploration of patient trajectories. Methods Code and Model Weights Availability All model weights and associated code have been deposited on GitHub, accessible here. Architecture Specifications We aimed to build a novel architecture within which to evaluate our results. The natural starting place in the field of ECG analysis is the one dimensional convolutional neural network (1d-CNN). To improve expressibility as opposed to simply using CNNs, our architecture additionally utilizes transformer encoder layers. Combining CNNs with transformers allowed us to enhance ECG representation learning by using the CNN to generate a key set of temporal tokens, while the self-attention layers were able to model the deep global relationships between them. As we had two windows for each appointment, we averaged their embeddings to arrive at a single vector per visit, and applied cosine similarity to these vectors. Training Details We initialized the model described in Table S1 (See Supplement) with random weights, and trained using the AdamW optimizer 25 , with an initial learning rate of 1e-5, which we found empirically gave the best results on the validation set. We reduced the learning rate by a factor of 5 for every 5 epochs that the validation loss did not decrease, and after 15 epochs of non-decreasing validation loss, we stopped training the model, and checkpointed the weights for later use. As is standard in the contrastive learning literature, we used the linear projection head for training only, and dropped the projection during inference (generating embeddings per-person). We fit a residual convolutional tokenizer following a standard architecture scheme 26 , and followed by 12 transformer encoder layers 27 using the BERT Base hyperparameters 28 , including 12 attention heads, a hidden dimension of 3,024. In each forward pass, we pass each recording through the 1d-CNN encoder, yielding a sequence of 58 tokens aligned in time, over 768 channels. We set 768 to be the embedding dimension of the transformer network, and add positional embeddings. We then add a 59th additional token to the model whose value is instantiated randomly (Gaussian), similarly in essence to the CLS token of a standard Vision Transformer (ViT) 29 , which interacts with all other tokens and whose value is sent to the linear projection before becoming the final embedding. For each window per person we obtained a 768 dimensional embedding vector from our model. We used Python Version 3.11.3 for all analyses. We fit all models using PyTorch version ‘2.3.1+cu121’ 30 with the Nvidia Quadro RTX 8000 Graphics Processing Unit (GPU). ECG Dataset We began with a dataset comprising 16,775 ECGs, with 11,933 recordings from baseline visits, and the remainder from repeat appointments. Using 12 Lead NORAV ECG machine –PC-ECG 1200 31 with an integrated electrodes chest belt for precordial leads. The ECGs were all sampled at a rate of 1000 hz, for 10 seconds minimum, before applying a 50Hz AC noise filter and an EMG muscle noise filter at 35 Hz. The baseline filter was set to on, and we used 16 bit resolution. We split each recording into two non overlapping windows 4 seconds in length across all 12 channels. For the SSL pre-training task, we split the dataset randomly 90/5/5 into train/test/validation over patient identity. We excluded automatically-detected changes related to heart rate, artifacts (identified by focal changes in only part of the leads), and lead misplacement. Embedding Neighbourhood Metrics As we had two windows per recording, we averaged the embeddings from both windows to apply patient identification metrics 8 . We could have of course done this without averaging the embeddings from each window (treating each window separately), however we wanted to encourage all windows from both recordings to be close together, as opposed to just one of each. Within the sets of first and second appointments, we too can apply this metric without averaging the two windows: instead seeing whether the first half a patient’s ECG is closest to the second as compared to all other first or second windows. PBMCs Isolation Blood samples were collected with BD Vacutainer® CPT™ Cell Preparation Tube with Sodium Citrate and Ficoll (BD Ref# 362760), and PBCMs were purified according to the manufacturer’s instructions with minor modifications. Briefly, the PBMCs were collected from the CPT tubes after 20-minute centrifugation at 22oC followed by two washing steps with Washing media (0.9%RPMI 1640, Thermo Fisher Scientific, and 0.1%FBS, Sigma-Aldrich). The cells were resuspended with resuspension solution (0.5% RPMI 1640, Thermo Fisher Scientific and 0.5%FBS, Sigma-Aldrich) and frozen at −80oC until processing. All processes were performed on a Tecan Evo 100 automated platform. RNA extraction and library prep and sequencing RNA was extracted from frozen PBMCs with an All-prep DNA/RNA 96 (4) kit (QIAGEN, Cat# 20-80311) on a Tecan Evo 200 automated platform. Libraries for bulk mRNA sequencing were prepared with mcSCRB-seq methodology 32 . All steps were identical except for the final step of tagmentation. 96 amplified cDNA samples were pooled and 12ng from each pool was tagmented by adding 1uL of TDE1 and 2x Tagment DNA buffer (Illumina, Ref# 20034197) in 60uL followed by PCR amplification with Kapa HiFi HotStart ReadyMix (Kapa Biosystems) and 5 μM IDT for Illumina Nextera DNA Dual Indexes. The PCR conditions were: 3’ at 72°C, 30’’ at 95°C followed by 14 cycles of 10’’ at 95°C, 30’’ at 55°C, 1’ at 72°C and final elongation for 5’ at 72°C. PCR clean-up and size selection was performed using SPRI beads for each pool and eluted in 12uL. Libraries were sequenced to a minimum depth of 5M reads per sample on a NovaSeq 6000 instrument with a NovaSeq S1 v1.5 100-cycle kit (Cat# 20012865; Illumina). Deduplicated count data for the gene expression was computed using a previously validated pipeline 33 . These were then converted to counts per million mapped reads through normalization. RNA Expression Prediction For downstream task prediction, we stored embeddings from each person into a tabular dataset, and fit Ridge linear regression models over them to predict the output. We used cross-fold-validation over the dataset to predict the test prediction performance. Enrichment Analysis We discretized the Delta ECG distribution by assigning patients to the “case” category if their distance between their matched appointments was above the mean of the matched distance distribution (0.9), and “control” group otherwise. Using GSEAPY 34 , we ran over gene set enrichment analysis on all 3000 genes for both men and women, ranked by corrected prediction P value from ECG embeddings. We used 5000 permutations, and the log (base 2) ratio of classes. Phenotypes We set the data collection interval to be from the commencement of the HPP, in January of 2019 to June 2024. RNSeq: Ranked Multiple Hypothesis Testing Correction Our RNASequencing dataset includes expression values for over 10k genes per patient, though the data itself is sparse. The remaining genes which are expressed across people could be picked based on a combination of higher variance, higher expression, or randomly. Either way, a set number of genes needed to be subsampled for the experiment. As multiple hypothesis correction is a function of the number of tests performed, to exclude the possibility of hand-picking this number of genes so that our results were significant, we employed sequential multiple hypothesis correction in the following way: we first ordered all genes based on decreasing population-level variance. For the first (most-variable) gene, we left its P-value intact. For the nth-gene in the list, we adjusted its P value using Bonferonni adjustment for n tests. This is essentially an analogue of the Holm-Bonferonni method, without ordering the P values from lowest to highest first. Results from experiment consisting of the top n genes would therefore only be penalized for those genes. Genome Wide Association Study Our methods for GWAS and LD clumping have been described previously 9 . We use the same methodology for our eQTL analyses, replacing the clinical phenotypes with the normalized expression values for each of the 66 candidate genes across the population. Author Contributions Z.L. conceived the project, designed and trained the neural network, performed the genetic analyses, interpreted the results, and wrote the manuscript. G.L. contributed insights and assisted with the suggestion of ideas for the project. A.G. coordinated and provided the DNA data, coordinated, conducted, and designed the RNA Sequencing experiment, interpreted the results, and wrote the manuscript. A.W. coordinated, conducted, and designed the RNA Sequencing experiment, interpreted the results, and wrote the manuscript. M.P. coordinated, conducted, and designed the RNA Sequencing experiment and wrote the manuscript. Y.T.B and Y.R. interpreted the results, wrote the manuscript, and provided valuable clinical insights. H.R. provided data, interpreted the results, and wrote the manuscript. E.S. conceived, directed, and supervised the project and analyses Declaration of Interests H.R. and Y.R. and are employees of Pheno.AI, Ltd, a biomedical data science company from Tel-Aviv, Israel. A.W and, E.S. are paid consultants to Pheno.AI, Ltd. The rest of the authors declare no competing interests. Ethics The Weizmann Institute of Science review board (IRB) approved the study and its protocols. All identifying details of the participants were erased prior to statistical analysis, so informed consent was waived by the IRB. All participants had full knowledge of data handling, storage, and sharing methods. This information was given to all participants, and is in agreement with the data privacy and protection policy of the Weizmann Institute of Science ( https://www.weizmann.ac.il/pages/privacy-policy ). Acknowledgements We thank members of the Segal lab for useful discussions. E.S. is supported by the Crown Human Genome Center; Larson Charitable Foundation New Scientist Fund; Else Kroener Fresenius Foundation; White Rose International Foundation; Ben B. and Joyce E. Eisenberg Foundation; Nissenbaum Family; Marcos Pinheiro de Andrade and Vanessa Buchheim; Lady Michelle Michels; Aliza Moussaieff; and grants funded by the Minerva foundation with funding from the Federal German Ministry for Education and Research and by the European Research Council and the Israel Science Foundation. Footnotes ↵ 2 First author References 1. ↵ Radhakrishnan , A. et al. Cross-modal autoencoder framework learns holistic representations of cardiovascular state . Nat. Commun . 14 , 2436 ( 2023 ). OpenUrl 2. ↵ Libiseller-Egger , J. et al. Deep learning-derived cardiovascular age shares a genetic basis with other cardiac phenotypes . Sci. Rep . 12 , 22625 ( 2022 ). OpenUrl 3. ↵ Wang , X. et al. Genetic Susceptibility to Atrial Fibrillation Identified via Deep Learning of 12-Lead Electrocardiograms . Circ. Genomic Precis. Med . 16 , 340 – 349 ( 2023 ). OpenUrl 4. ↵ Kiyasseh , D. , Zhu , T. & Clifton , D. A. CLOCS: Contrastive Learning of Cardiac Signals Across Space, Time, and Patients . Preprint at http://arxiv.org/abs/2005.13249 ( 2021 ). 5. ↵ Diamant , N. et al. Patient Contrastive Learning: a Performant, Expressive, and Practical Approach to ECG Modeling . PLOS Comput. Biol . 18 , e1009862 ( 2022 ). OpenUrl CrossRef 6. ↵ Radford , A. , Narasimhan , K. , Salimans , T. & Sutskever , I. Improving Language Understanding by Generative Pre-Training . 7. ↵ Shilo , S. et al. 10 K: a large-scale prospective longitudinal study in Israel . Eur. J. Epidemiol . 36 , 1187 – 1194 ( 2021 ). OpenUrl CrossRef 8. ↵ Lutsker , G. , Rossman , H. , Godiva , N. & Segal , E. COMPRER: A Multimodal Multi-Objective Pretraining Framework for Enhanced Medical Image Representation . Preprint at doi: 10.48550/arXiv.2403.09672 ( 2024 ). 9. ↵ Levine , Z. et al. Genome-wide association studies and polygenic risk score phenome-wide association studies across complex phenotypes in the human phenotype project . Med 5 , 90 – 101.e4 ( 2024 ). OpenUrl 10. ↵ Kichaev , G. et al. Leveraging Polygenic Functional Enrichment to Improve GWAS Power . Am. J. Hum. Genet . 104 , 65 – 75 ( 2019 ). OpenUrl CrossRef PubMed 11. ↵ Liu , Y. et al. Blood monocyte transcriptome and epigenome analyses reveal loci associated with human atherosclerosis . Nat. Commun . 8 , 393 ( 2017 ). OpenUrl CrossRef 12. ↵ Kanai , M. et al. Genetic analysis of quantitative traits in the Japanese population links cell types to complex human diseases . Nat. Genet . 50 , 390 – 400 ( 2018 ). OpenUrl CrossRef PubMed 13. ↵ Diabetes Genetics Initiative of Broad Institute of Harvard and MIT, Lund University, and Novartis Institutes of BioMedical Research et al. Genome-wide association analysis identifies loci for type 2 diabetes and triglyceride levels . Science 316 , 1331 – 1336 ( 2007 ). OpenUrl Abstract / FREE Full Text 14. ↵ Pearson-Stuttard , J. et al. Trends in predominant causes of death in individuals with and without diabetes in England from 2001 to 2018: an epidemiological analysis of linked primary care records . Lancet Diabetes Endocrinol . 9 , 165 – 173 ( 2021 ). OpenUrl 15. ↵ Wj , A. et al. The Allelic Landscape of Human Blood Cell Trait Variation and Links to Common Complex Disease . Cell 167 , ( 2016 ). 16. ↵ B, N., et al. Results from a Genome-Wide Association Study (GWAS) in Mastocytosis Reveal New Gene Polymorphisms Associated with WHO Subgroups . Int. J. Mol. Sci . 21 , ( 2020 ). 17. ↵ Indhirajanti , S. et al. Systemic mastocytosis associates with cardiovascular events despite lower plasma lipid levels . Atherosclerosis 268 , 152 – 156 ( 2018 ). OpenUrl 18. ↵ Chistiakov , D. A. , Shkurat , T. P. , Melnichenko , A. A. , Grechko , A. V. & Orekhov , A. N. The role of mitochondrial dysfunction in cardiovascular disease: a brief review . Ann. Med . 50 , 121 – 127 ( 2018 ). OpenUrl CrossRef PubMed 19. ↵ Saito , Y. , Tanaka , A. , Node , K. & Kobayashi , Y. Uric acid and cardiovascular disease: A clinical review . J. Cardiol . 78 , 51 – 57 ( 2021 ). OpenUrl PubMed 20. ↵ Yu , W. & Cheng , J.-D. Uric Acid and Cardiovascular Disease: An Update From Molecular Mechanism to Clinical Perspective . Front. Pharmacol . 11 , 582680 ( 2020 ). OpenUrl 21. ↵ Yun , T. et al. Unsupervised representation learning on high-dimensional clinical data improves genomic discovery and prediction . Nat. Genet . 1 – 10 ( 2024 ) doi: 10.1038/s41588-024-01831-6 . OpenUrl CrossRef 22. ↵ Pirruccello , J. P. et al. Deep learning enables genetic analysis of the human thoracic aorta . Nat. Genet . 54 , 40 – 51 ( 2022 ). OpenUrl CrossRef 23. ↵ Ben-Moshe , N. et al. RawECGNet: Deep Learning Generalization for Atrial Fibrillation Detection From the Raw ECG . IEEE J. Biomed. Health Inform . 1 – 10 ( 2024 ) doi: 10.1109/JBHI.2024.3404877 . OpenUrl CrossRef 24. ↵ Vrudhula , A. et al. Impact of Case and Control Selection on Training Artificial Intelligence Screening of Cardiac Amyloidosis . JACC Adv . 0 . 25. ↵ Loshchilov , I. & Hutter , F. Decoupled Weight Decay Regularization . Preprint at doi: 10.48550/arXiv.1711.05101 ( 2019 ). OpenUrl CrossRef 26. ↵ Ribeiro , A. H. et al. Automatic diagnosis of the 12-lead ECG using a deep neural network . Nat. Commun . 11 , 1760 ( 2020 ). OpenUrl CrossRef 27. ↵ Vaswani , A. et al. Attention Is All You Need . Preprint at http://arxiv.org/abs/1706.03762 ( 2023 ). 28. ↵ Devlin , J. , Chang , M.-W. , Lee , K. & Toutanova , K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding . Preprint at doi: 10.48550/arXiv.1810.04805 ( 2019 ). OpenUrl CrossRef 29. ↵ Dosovitskiy , A. , et al. An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale . Preprint at doi: 10.48550/arXiv.2010.11929 ( 2021 ). OpenUrl CrossRef 30. ↵ Ansel , J. et al. PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation . in Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems , Volume 2 929 – 947 ( ACM , La Jolla CA USA , 2024 ). doi: 10.1145/3620665.3640366 . OpenUrl CrossRef 31. ↵ ECG Systems | ECG Excellence over 30 years - Norav Medical. https://www.noravmedical.com/ . 32. ↵ Bagnoli , J. W. et al. Sensitive and powerful single-cell RNA sequencing using mcSCRB-seq . Nat. Commun . 9 , 2937 ( 2018 ). OpenUrl CrossRef PubMed 33. ↵ Kohen , R. et al. UTAP: User-friendly Transcriptome Analysis Pipeline . BMC Bioinformatics 20 , 154 ( 2019 ). OpenUrl CrossRef 34. ↵ Fang , Z. , Liu , X. & Peltz , G. GSEApy: a comprehensive package for performing gene set enrichment analysis in Python . Bioinformatics 39 , btac757 ( 2023 ). OpenUrl Back to top Previous Next Posted August 21, 2024. Download PDF Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Genetic underpinnings of predicted changes in cardiovascular function using self supervised learning Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Genetic underpinnings of predicted changes in cardiovascular function using self supervised learning Zachary Levine , Guy Lutsker , Anastasia Godneva , Adina Weinberger , Maya Pompan , Yeela Talmor-Barkan , Yotam Reisner , Hagai Rossman , Eran Segal bioRxiv 2024.08.15.608061; doi: https://doi.org/10.1101/2024.08.15.608061 Share This Article: Copy Citation Tools Genetic underpinnings of predicted changes in cardiovascular function using self supervised learning Zachary Levine , Guy Lutsker , Anastasia Godneva , Adina Weinberger , Maya Pompan , Yeela Talmor-Barkan , Yotam Reisner , Hagai Rossman , Eran Segal bioRxiv 2024.08.15.608061; doi: https://doi.org/10.1101/2024.08.15.608061 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Systems Biology Subject Areas All Articles Animal Behavior and Cognition (8023) Biochemistry (18794) Bioengineering (14931) Bioinformatics (44508) Biophysics (22636) Cancer Biology (19779) Cell Biology (26956) Clinical Trials (138) Developmental Biology (14000) Ecology (21053) Epidemiology (2067) Evolutionary Biology (25483) Genetics (16187) Genomics (23537) Immunology (18737) Microbiology (42582) Molecular Biology (18107) Neuroscience (93652) Paleontology (701) Pathology (2989) Pharmacology and Toxicology (5105) Physiology (8133) Plant Biology (16030) Scientific Communication and Education (2098) Synthetic Biology (4574) Systems Biology (10255) Zoology (2393) window.__CF$cv$params={r:'a4022387ec8873e2',t:'MTc5MDI1NjU3NA==',u:'01a0d39b780270e1b571b48e8a8c4b65',ut:'t3Ae74C7AulVdyyvREmNtb4Wt5ow3_uPPqliSFHJbgg-1790256576-1.2.1.1-ZS5aWvBtx.zwV_GI46rMGZG7kh8h_kvAkkEOeyDKpBQfk2yMobZxYeIbu4pm_ckYsZnRhD9V95HUFzQ5oOzOuayikBXasPVafzai.BVeDow',i:60};(function(){if(!document.body)return;var s=document.createElement('script');s.src='/cdn-cgi/challenge-platform/scripts/precursor/main.js';document.head.appendChild(s);})();

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: preprint-html ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-10-07T06:25:12.689510+00:00
License: CC-BY-4.0 · commercial use OK · attribution required
Per Europe PMC