Why world models fail under intervention: Ontological–causal separation as a necessary structure

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Learning world models that can understand, predict, and act within complex physical environments is a fundamental goal of embodied intelligence. However, many dominant approaches—such as latent video models, Dreamer-style imagination agents, and recent JEPA-based predictive architectures—operate on unstructured latent representations, where objects, relations, causality, and intention are implicitly entangled. While effective under observational evaluation, such representations provide limited support for intervention, planning, and counterfactual reasoning. This contrasts sharply with human cognition, which relies on structured reasoning over entities, attributes, and relations to understand and act in the world. In this work, we introduce LY-GWM (LingYang Graph World Model), a philosophy-inspired, graph-structured world model designed to examine a structural hypothesis: that explicit separation between ontological structure—what entities and relations constitute the world—and causal structure—how these entities evolve under actions—is a necessary condition for coherent world modeling. LY-GWM represents the environment as a dynamic graph composed of entities, attributes, and relations, and models causal dynamics as structured transformations over this graph. Building on this representation, LY-GWM incorporates three reasoning mechanisms inspired by long-standing philosophical traditions: causal dynamics for action-conditioned state transitions, teleological reasoning for goal-directed behavior, and dialectical novelty detection for identifying structural contradictions and unexplained changes. Through controlled diagnostic environments and high-fidelity humanoid simulation, we show that world models lacking explicit ontological–causal separation can perform well under observation yet systematically fail under intervention, whereas structured graph-based models maintain coherent reasoning across prediction, planning, and counterfactual analysis. Our results suggest that incorporating explicit ontological and causal structure is not merely a conceptual design choice, but a necessary condition for world models that support intervention-consistent reasoning in embodied intelligence.
Full text 11,743 characters · extracted from preprint-html · click to expand
Why world models fail under intervention: Ontological–causal separation as a necessary structure | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Why world models fail under intervention: Ontological–causal separation as a necessary structure Tian Liang Ma This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8377767/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Learning world models that can understand, predict, and act within complex physical environments is a fundamental goal of embodied intelligence. However, many dominant approaches—such as latent video models, Dreamer-style imagination agents, and recent JEPA-based predictive architectures—operate on unstructured latent representations, where objects, relations, causality, and intention are implicitly entangled. While effective under observational evaluation, such representations provide limited support for intervention, planning, and counterfactual reasoning. This contrasts sharply with human cognition, which relies on structured reasoning over entities, attributes, and relations to understand and act in the world. In this work, we introduce LY-GWM (LingYang Graph World Model), a philosophy-inspired, graph-structured world model designed to examine a structural hypothesis: that explicit separation between ontological structure—what entities and relations constitute the world—and causal structure—how these entities evolve under actions—is a necessary condition for coherent world modeling. LY-GWM represents the environment as a dynamic graph composed of entities, attributes, and relations, and models causal dynamics as structured transformations over this graph. Building on this representation, LY-GWM incorporates three reasoning mechanisms inspired by long-standing philosophical traditions: causal dynamics for action-conditioned state transitions, teleological reasoning for goal-directed behavior, and dialectical novelty detection for identifying structural contradictions and unexplained changes. Through controlled diagnostic environments and high-fidelity humanoid simulation, we show that world models lacking explicit ontological–causal separation can perform well under observation yet systematically fail under intervention, whereas structured graph-based models maintain coherent reasoning across prediction, planning, and counterfactual analysis. Our results suggest that incorporating explicit ontological and causal structure is not merely a conceptual design choice, but a necessary condition for world models that support intervention-consistent reasoning in embodied intelligence. Physical sciences/Mathematics and computing/Computational science Humanities/Language and linguistics Humanities/Complex networks Full Text Additional Declarations There is NO Competing Interest. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8377767","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":561714420,"identity":"e4d5c15f-bd85-4ae7-9fb7-9ba0a975bb7c","order_by":0,"name":"Tian Liang Ma","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+ElEQVRIie3PsUoDMRzH8YRAuvwl6x3Wgg8gBAKKUOyrXBCcHHQpHTr8QyC+Qjr5Cp0UtxyBcznqKnRRfIG7TRfxwFG9080h3/n3GX6EpFL/Mobx7X05mVmK8XUxBSFwQHBqIuGVkiNrXkh9Ns59GCS2I0yv4d4q6uJUYtEvDq4Nxgvg+jbTbvfSPYAkgTbt+c/ksOou+Gys7nxH/GYLRwxZvrrpJwEk38NH7TKYb+EYA2c7g6Rg9JPwDchQDJMIge2v69IqcOE3RGO5wkrlV8Y8+/oUcl/a/i8xxqbF5USw0VNoFiczIWzZtD3kuyj+bZ9KpVKpL30AdaJeeL9kf5IAAAAASUVORK5CYII=","orcid":"","institution":"Beijing Miwang Information Technology Co., Ltd.","correspondingAuthor":true,"prefix":"","firstName":"Tian","middleName":"Liang","lastName":"Ma","suffix":""}],"badges":[],"createdAt":"2025-12-16 15:27:38","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8377767/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8377767/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":98627131,"identity":"ebe2d830-76e5-4c5c-b04d-b43a05af3300","added_by":"auto","created_at":"2025-12-19 17:10:10","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":296923,"visible":true,"origin":"","legend":"Article File","description":"","filename":"NatureMachineIntelligence2025.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8377767/v1_covered_25f1343b-27f9-4935-8286-258af4f61ed9.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Why world models fail under intervention:\nOntological–causal separation as a necessary\nstructure","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-8377767/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8377767/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Learning world models that can understand, predict, and act within complex physical environments is a fundamental goal of embodied intelligence. However, many dominant approaches—such as latent video models, Dreamer-style imagination agents, and recent JEPA-based predictive architectures—operate on unstructured latent representations, where objects, relations, causality, and intention are implicitly entangled. While effective under observational evaluation, such representations provide limited support for intervention, planning, and counterfactual reasoning. This contrasts sharply with human cognition, which relies on structured reasoning over entities, attributes, and relations to understand and act in the world.\r\n\r\nIn this work, we introduce LY-GWM (LingYang Graph World Model), a philosophy-inspired, graph-structured world model designed to examine a structural hypothesis: that explicit separation between ontological structure—what entities and relations constitute the world—and causal structure—how these entities evolve under actions—is a necessary condition for coherent world modeling. LY-GWM represents the environment as a dynamic graph composed of entities, attributes, and relations, and models causal dynamics as structured transformations over this graph.\r\n\r\nBuilding on this representation, LY-GWM incorporates three reasoning mechanisms inspired by long-standing philosophical traditions: causal dynamics for action-conditioned state transitions, teleological reasoning for goal-directed behavior, and dialectical novelty detection for identifying structural contradictions and unexplained changes. Through controlled diagnostic environments and high-fidelity humanoid simulation, we show that world models lacking explicit ontological–causal separation can perform well under observation yet systematically fail under intervention, whereas structured graph-based models maintain coherent reasoning across prediction, planning, and counterfactual analysis.\r\n\r\nOur results suggest that incorporating explicit ontological and causal structure is not merely a conceptual design choice, but a necessary condition for world models that support intervention-consistent reasoning in embodied intelligence.","manuscriptTitle":"Why world models fail under intervention:\nOntological–causal separation as a necessary\nstructure","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-12-19 05:14:34","doi":"10.21203/rs.3.rs-8377767/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"64a6f035-c8b3-4add-90b8-c3821cb92800","owner":[],"postedDate":"December 19th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":59809866,"name":"Physical sciences/Mathematics and computing/Computational science"},{"id":59809867,"name":"Humanities/Language and linguistics"},{"id":59809868,"name":"Humanities/Complex networks"}],"tags":[],"updatedAt":"2025-12-19T05:14:34+00:00","versionOfRecord":[],"versionCreatedAt":"2025-12-19 05:14:34","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8377767","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8377767","identity":"rs-8377767","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00