InfEHR: Resolving Clinical Uncertainty through Deep Geometric Learning on Electronic Health Records | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article InfEHR: Resolving Clinical Uncertainty through Deep Geometric Learning on Electronic Health Records Girish Nadkarni, Justin Kauffman, Emma Holmes, Akhil Vaid, Alexander Charney, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5953885/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 26 Sep, 2025 Read the published version in Nature Communications → Version 1 posted You are reading this latest preprint version Abstract Electronic health records (EHRs) contain multimodal data that can inform diagnostic and prognostic clinical decisions but are often unsuited for advanced machine learning (ML)–based patient-specific analyses. ML models and clinical heuristics learn generalizable relationships from predefined factors, yet many patients may not benefit if those factors are missing in the EHR or differ—however subtly—from typical training populations. Clinical heuristics are limited to low complexity, often linear, relationships and patterns between clinical variables. ML approaches in EHRs significantly expand pattern sophistication but require large, labeled datasets, which are often unattainable especially in low prevalence diseases and are limited by sources of random and non-random variation in EHRs. Deep learning (DL), in contrast with ML and clinical heuristics, learns features without predefinition but requires even greater label access for predictions. While DL can construct unsupervised EHR representations, the patterns and characteristics of less prevalent examples are poorly resolved, and downstream clinical applications still require labels. We present Inf-EHR, a framework to automatically compute clinical likelihoods from whole EHRs of patients from diverse clinical settings without need of large volumes of labeled training data. We apply deep geometric learning to EHRs through a novel procedure that converts whole EHRs to temporal graphs. These graphs naturally capture phenotypic temporal dynamics leading to unbiased representations. Using only a few labeled examples, InfEHR computes and automatically revises likelihoods leading to highly performant inferences especially in low prevalence diseases which are often the most clinically ambiguous. To demonstrate utility, we use EHRs from the Mount Sinai Health System and The University of California, Irvine Medical Center and test its performance compared to physician-provided clinical heuristics across two diseases with no clinical or epidemiological overlap: a rare disease (neonatal culture-negative sepsis) with prevalence of 2% in neonates, and a more common disease (adult post-operative acute kidney injury) with prevalence of 22%. We show that Inf-EHR is superior to existing clinical heuristics both for culture-negative sepsis (sensitivity: 0.65 vs .041, specificity: 0.99 vs.0.98) and post-operative acute kidney injury (sensitivity: 0.72 vs 0.20, specificity: 0.91 vs 0.97). We present the first application of geometric deep learning in EHRs that can be used in real world clinical settings at scale, for improving phenotype identification and resolving clinical uncertainty. Health sciences/Medical research/Outcomes research Health sciences/Health care/Diagnosis Full Text Additional Declarations There is NO Competing Interest. Supplementary Files InfEHRsuppnatmed.docx Supplemental Figures InfEHR Cite Share Download PDF Status: Published Journal Publication published 26 Sep, 2025 Read the published version in Nature Communications → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5953885","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":415771106,"identity":"ab88ef3f-fd17-47cd-af27-f06cca227ddb","order_by":0,"name":"Girish Nadkarni","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABL0lEQVRIie2RMUvDQBiGvxCIU3D9spi/kFCIFdz9G96SLOpSkLqUC4F0KfoP2l8g1CV0TDhIlwuucSyFzoqLkQzepQ7KteIoeM9wd7zHw3vcB6DR/EVMgNyg8hTKBY8+c+yufqX0ujT/SZF8UYDQrbIf98xeFc0Crg7HIXm9WfSj+/EkWD+3fZGAsR6+K4rPDjxmcxgg3zCn4niZcX7sFSkOkIHpV7eqkljAjFS8p46oE6dCqS8CLCgSysBy4slOpWiEMqujpBFKFEglb5HM9iiuaUFuC2Veh6VsOd8qFpK5VOibonhCYXaK5IFvwhOh+Bkvr7GSCTMSP6Zqy11pvjTpKZkuw95TnI7cYJlkOGxHZPrIihVt1Za823DXAOS8xM8oLWrxd9QWjUaj+Xd8AFXdc2ZOD81eAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0001-6319-4314","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":true,"prefix":"","firstName":"Girish","middleName":"","lastName":"Nadkarni","suffix":""},{"id":415771107,"identity":"9b61b0e2-f509-470a-99b8-64a6b55fa9f7","order_by":1,"name":"Justin Kauffman","email":"","orcid":"","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Justin","middleName":"","lastName":"Kauffman","suffix":""},{"id":415771108,"identity":"57a9f031-1d32-4635-990f-db02f2ccc5ae","order_by":2,"name":"Emma Holmes","email":"","orcid":"","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Emma","middleName":"","lastName":"Holmes","suffix":""},{"id":415771109,"identity":"c4102e9e-f570-4f82-8abf-b7ec45482aa2","order_by":3,"name":"Akhil Vaid","email":"","orcid":"https://orcid.org/0000-0002-3343-744X","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Akhil","middleName":"","lastName":"Vaid","suffix":""},{"id":415771110,"identity":"2a265fd5-fb15-4b2d-9ac3-99ec1bd1f8ba","order_by":4,"name":"Alexander Charney","email":"","orcid":"https://orcid.org/0000-0001-8135-6858","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Alexander","middleName":"","lastName":"Charney","suffix":""},{"id":415771111,"identity":"2db3c3e1-4f3d-4faf-b95e-1315b9afefdb","order_by":5,"name":"Patricia Kovatch","email":"","orcid":"https://orcid.org/0000-0001-8368-1742","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Patricia","middleName":"","lastName":"Kovatch","suffix":""},{"id":415771112,"identity":"92c23a33-554e-4c63-83ad-1dbb76265034","order_by":6,"name":"Joshua Lampert","email":"","orcid":"https://orcid.org/0000-0003-0081-0356","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Joshua","middleName":"","lastName":"Lampert","suffix":""},{"id":415771113,"identity":"f62183a3-9f04-4d76-ab41-48c9c10191fb","order_by":7,"name":"Ankit Sakhuja","email":"","orcid":"","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Ankit","middleName":"","lastName":"Sakhuja","suffix":""},{"id":415771114,"identity":"43683511-ac18-4a57-b827-b692c20fcfbc","order_by":8,"name":"Marinka Zitnik","email":"","orcid":"https://orcid.org/0000-0001-8530-7228","institution":"Harvard Medical School","correspondingAuthor":false,"prefix":"","firstName":"Marinka","middleName":"","lastName":"Zitnik","suffix":""},{"id":415771115,"identity":"0ec2d5bd-6bab-4084-b1e5-6e796ff93484","order_by":9,"name":"Benjamin Glicksberg","email":"","orcid":"https://orcid.org/0000-0003-4515-8090","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Benjamin","middleName":"","lastName":"Glicksberg","suffix":""},{"id":415771116,"identity":"71e47711-cabf-48e1-9a28-98657f3cc0ee","order_by":10,"name":"Ira Hofer","email":"","orcid":"","institution":"Icahn School of Medicine at Mount Sinai","correspondingAuthor":false,"prefix":"","firstName":"Ira","middleName":"","lastName":"Hofer","suffix":""}],"badges":[],"createdAt":"2025-02-03 22:15:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5953885/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5953885/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41467-025-63366-6","type":"published","date":"2025-09-26T04:00:00+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":92305122,"identity":"91445f31-9638-4df5-8f72-1284fc38a2be","added_by":"auto","created_at":"2025-09-27 07:11:22","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2002217,"visible":true,"origin":"","legend":"","description":"","filename":"InfEHRmainnatmed.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5953885/v1_covered_6837d16f-7205-49bf-b1dc-b9f6de688386.pdf"},{"id":76452537,"identity":"f2d59be4-347f-4d43-a0f2-617771153a7c","added_by":"auto","created_at":"2025-02-17 10:00:40","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":639117,"visible":true,"origin":"","legend":"Supplemental Figures InfEHR","description":"","filename":"InfEHRsuppnatmed.docx","url":"https://assets-eu.researchsquare.com/files/rs-5953885/v1/349d4beae7d742da5ab4a8a9.docx"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"InfEHR: Resolving Clinical Uncertainty through Deep Geometric Learning on Electronic Health Records","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-5953885/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5953885/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eElectronic health records (EHRs) contain multimodal data that can inform diagnostic and prognostic clinical decisions but are often unsuited for advanced machine learning (ML)\u0026ndash;based patient-specific analyses. ML models and clinical heuristics learn generalizable relationships from predefined factors, yet many patients may not benefit if those factors are missing in the EHR or differ\u0026mdash;however subtly\u0026mdash;from typical training populations. Clinical heuristics are limited to low complexity, often linear, relationships and patterns between clinical variables. ML approaches in EHRs significantly expand pattern sophistication but require large, labeled datasets, which are often unattainable especially in low prevalence diseases and are limited by sources of random and non-random variation in EHRs. Deep learning (DL), in contrast with ML and clinical heuristics, learns features without predefinition but requires even greater label access for predictions. While DL can construct unsupervised EHR representations, the patterns and characteristics of less prevalent examples are poorly resolved, and downstream clinical applications still require labels. We present Inf-EHR, a framework to automatically compute clinical likelihoods from whole EHRs of patients from diverse clinical settings without need of large volumes of labeled training data. We apply deep geometric learning to EHRs through a novel procedure that converts whole EHRs to temporal graphs. These graphs naturally capture phenotypic temporal dynamics leading to unbiased representations. Using only a few labeled examples, InfEHR computes and automatically revises likelihoods leading to highly performant inferences especially in low prevalence diseases which are often the most clinically ambiguous. To demonstrate utility, we use EHRs from the Mount Sinai Health System and The University of California, Irvine Medical Center and test its performance compared to physician-provided clinical heuristics across two diseases with no clinical or epidemiological overlap: a rare disease (neonatal culture-negative sepsis) with prevalence of 2% in neonates, and a more common disease (adult post-operative acute kidney injury) with prevalence of 22%. We show that Inf-EHR is superior to existing clinical heuristics both for culture-negative sepsis (sensitivity: 0.65 vs .041, specificity: 0.99 vs.0.98) and post-operative acute kidney injury (sensitivity: 0.72 vs 0.20, specificity: 0.91 vs 0.97). We present the first application of geometric deep learning in EHRs that can be used in real world clinical settings at scale, for improving phenotype identification and resolving clinical uncertainty.\u003c/p\u003e","manuscriptTitle":"InfEHR: Resolving Clinical Uncertainty through Deep Geometric Learning on Electronic Health Records","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-02-17 10:00:36","doi":"10.21203/rs.3.rs-5953885/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"nature-communications","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"NCOMMS","sideBox":"Learn more about [Nature Communications](http://www.nature.com/ncomms/)","snPcode":"","submissionUrl":"https://mts-ncomms.nature.com/","title":"Nature Communications","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Communications","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"b82453ca-d283-41ef-8d68-608b270b520c","owner":[],"postedDate":"February 17th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":44337300,"name":"Health sciences/Medical research/Outcomes research"},{"id":44337301,"name":"Health sciences/Health care/Diagnosis"}],"tags":[],"updatedAt":"2025-09-27T07:11:14+00:00","versionOfRecord":{"articleIdentity":"rs-5953885","link":"https://doi.org/10.1038/s41467-025-63366-6","journal":{"identity":"nature-communications","isVorOnly":false,"title":"Nature Communications"},"publishedOn":"2025-09-26 04:00:00","publishedOnDateReadable":"September 26th, 2025"},"versionCreatedAt":"2025-02-17 10:00:36","video":"","vorDoi":"10.1038/s41467-025-63366-6","vorDoiUrl":"https://doi.org/10.1038/s41467-025-63366-6","workflowStages":[]},"version":"v1","identity":"rs-5953885","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5953885","identity":"rs-5953885","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.