Symbolic Laws of Protein Folding: Interpretable Dynamics and Inverse Design via NEXUS

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract What are the symbolic laws that govern how proteins fold, evolve, and function? Despite revolutionary advances in deep learningbased structure prediction, modern biology lacks a compact, interpretable, and generative theory of protein folding rooted in first principles. In this work, we introduce a symbolic AI framework that discovers dynamic folding equations from amino acid sequences using a constrained operator search engine called NEXUS. Rather than predicting static structures, NEXUS uncovers interpretable dynamical laws—composed of memory-bearing derivatives, nonlocal residue coupling kernels, and energy gradients—that govern folding trajectories in symbolic form. We show that these equations accurately model folding for multiple fold classes, including α/β domains and β-barrels, and outperform black-box models in interpretability and biophysical insight. Crucially, the discovered equations are invertible: given a desired 3D fold, we solve symbolically for sequences that generate it, enabling inverse design. We validate our symbolic predictions with real PDB folds (e.g., 1IGY), contact maps, folding funnels, and mutation impact heatmaps derived from symbolic sensitivities (∆Ei). Our framework generalizes across domains—extending to RNA folding, intrinsically disordered proteins, and enzyme design—and produces experimentally actionable insights. By replacing statistical memorization with symbolic reasoning, we introduce not just a model, but a new paradigm: symbolic biology.
Full text 11,156 characters · extracted from preprint-html · click to expand
Symbolic Laws of Protein Folding: Interpretable Dynamics and Inverse Design via NEXUS | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Symbolic Laws of Protein Folding: Interpretable Dynamics and Inverse Design via NEXUS Mehardeep Singh This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6941169/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract What are the symbolic laws that govern how proteins fold, evolve, and function? Despite revolutionary advances in deep learningbased structure prediction, modern biology lacks a compact, interpretable, and generative theory of protein folding rooted in first principles. In this work, we introduce a symbolic AI framework that discovers dynamic folding equations from amino acid sequences using a constrained operator search engine called NEXUS. Rather than predicting static structures, NEXUS uncovers interpretable dynamical laws—composed of memory-bearing derivatives, nonlocal residue coupling kernels, and energy gradients—that govern folding trajectories in symbolic form. We show that these equations accurately model folding for multiple fold classes, including α/β domains and β-barrels, and outperform black-box models in interpretability and biophysical insight. Crucially, the discovered equations are invertible: given a desired 3D fold, we solve symbolically for sequences that generate it, enabling inverse design. We validate our symbolic predictions with real PDB folds (e.g., 1IGY), contact maps, folding funnels, and mutation impact heatmaps derived from symbolic sensitivities (∆Ei). Our framework generalizes across domains—extending to RNA folding, intrinsically disordered proteins, and enzyme design—and produces experimentally actionable insights. By replacing statistical memorization with symbolic reasoning, we introduce not just a model, but a new paradigm: symbolic biology. Biophysics Structural Biology Computational Biology Applied Biochemistry symbolic operator discovery protein folding dynamics fractional Langevin equation nonlocal interactions conformational evolution free energy landscape NEXUS framework symbolic regression biological physics fractional calculus molecular dynamics amino acid sequence encoding inverse problem stochastic modeling systems biophysics Full Text Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6941169","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":474260549,"identity":"ad18b9fb-5979-4303-a05c-188ecf9046b2","order_by":0,"name":"Mehardeep Singh","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA9klEQVRIiWNgGAWjYFACHiBmgzAPMFQASWbmBlK0nAFpYSRBCwNjG5jEr0W3vffg54oyGzlzBvaHh27Oq43mbwdq+VGxDacWszPnkiXPnEsztmzgMTicu+147ozDjA2MPWdu49ZyI8dAsrHtcOKGAzwMQC3HchuAWpgZ2/Bouf/G+Gdj2//6DQfYHxzOnXMsdz5BLTd4zIC2HEgwOMAAdFhDTe4GglrO5JhZNpxLNtzZDPRLzrEDuRuBWg7i9cvxM8Y3G8rs5M3Z2x9/zqmpy513/vDBBz8qcGuBAwNmMHUYTB4grB6kBULVEaV4FIyCUTAKRhYAAGvWYEsNMhMtAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0009-0002-0044-0236","institution":"Sri Chaitanaya","correspondingAuthor":true,"prefix":"","firstName":"Mehardeep","middleName":"","lastName":"Singh","suffix":""}],"badges":[],"createdAt":"2025-06-20 18:56:19","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":true,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":true},"doi":"10.21203/rs.3.rs-6941169/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6941169/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":85237417,"identity":"e7a4baea-b2b3-4cd6-8f72-eef5c34b4971","added_by":"auto","created_at":"2025-06-23 17:19:53","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":492747,"visible":true,"origin":"","legend":"","description":"","filename":"SymbolicLawsofProteinFolding.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6941169/v1_covered_8546b73f-b479-49e8-b1a1-e196ad3af5cb.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eSymbolic Laws of Protein Folding: Interpretable Dynamics and Inverse Design via NEXUS\u003c/p\u003e","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Sri Chaitnaya","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"symbolic operator discovery, protein folding dynamics, fractional Langevin equation, nonlocal interactions, conformational evolution, free energy landscape, NEXUS framework, symbolic regression, biological physics, fractional calculus, molecular dynamics, amino acid sequence encoding, inverse problem, stochastic modeling, systems biophysics","lastPublishedDoi":"10.21203/rs.3.rs-6941169/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6941169/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eWhat are the symbolic laws that govern how proteins fold, evolve, and function? Despite revolutionary advances in deep learningbased structure prediction, modern biology lacks a compact, interpretable, and generative theory of protein folding rooted in first principles. In this work, we introduce a symbolic AI framework that discovers dynamic folding equations from amino acid sequences using a constrained operator search engine called NEXUS. Rather than predicting static structures, NEXUS uncovers interpretable dynamical laws—composed of memory-bearing derivatives, nonlocal residue coupling kernels, and energy gradients—that govern folding trajectories in symbolic form. We show that these equations accurately model folding for multiple fold classes, including α/β domains and β-barrels, and outperform black-box models in interpretability and biophysical insight. Crucially, the discovered equations are invertible: given a desired 3D fold, we solve symbolically for sequences that generate it, enabling inverse design. We validate our symbolic predictions with real PDB folds (e.g., 1IGY), contact maps, folding funnels, and mutation impact heatmaps derived from symbolic sensitivities (∆Ei). Our framework generalizes across domains—extending to RNA folding, intrinsically disordered proteins, and enzyme design—and produces experimentally actionable insights. By replacing statistical memorization with symbolic reasoning, we introduce not just a model, but a new paradigm: symbolic biology.\u003c/p\u003e","manuscriptTitle":"Symbolic Laws of Protein Folding: Interpretable Dynamics and Inverse Design via NEXUS","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-23 17:11:44","doi":"10.21203/rs.3.rs-6941169/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0508ba54-d261-4ee1-8758-fd63ddfb64f7","owner":[],"postedDate":"June 23rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":50369795,"name":"Biophysics"},{"id":50369796,"name":"Structural Biology"},{"id":50369797,"name":"Computational Biology"},{"id":50369798,"name":"Applied Biochemistry"}],"tags":[],"updatedAt":"2025-06-23T17:11:44+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-23 17:11:44","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6941169","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6941169","identity":"rs-6941169","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-28T02:00:01.590549+00:00
License: CC-BY-4.0