PlanAgent: Embodied Visual-Language Model for Grounded Task planning with Environment Map

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Embodied Intelligence refers to the agent interacting with the environment, perceiving, planning, decision-making, and executing like humans, which is applicable in smart homes, drone inspections, and other domains. Embodied task planning is one of the main tasks of embodied intelligence, which generates detailed step-by-step plans while perceiving the surrounding environment and understanding language instruction. Visual-language models, with powerful multimodal representation capabilities, have been generalized to various tasks. When applied to embodied task planning, it still faces the following two challenges. Firstly, the intricate complexity of the environment leads to difficulties in global environment information modeling. Secondly, frequent turns in task paths result in the dependence on strong spatial reasoning ability. To overcome these challenges, we propose PlanAgent, the first embodied visual-language model for embodied task planning. Specifically, the environment map is employed to model the global environment information. Then we present the environment map encoder to extract task-related information from the environment. Further, to reduce task path planning's dependence on strong spatial reasoning, we introduce the self-posture-aware training strategy to break down long-term spatial reasoning into short-term. We build the EmbodiedPlan-20k dataset for grounded planning in embodied tasks. Our experiments on the dataset demonstrate that PlanAgent outperforms previous methods and all components are effective.
Full text 12,629 characters · extracted from preprint-html · click to expand
PlanAgent: Embodied Visual-Language Model for Grounded Task planning with Environment Map | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article PlanAgent: Embodied Visual-Language Model for Grounded Task planning with Environment Map Yuanchang Yue, Fanglong Yao, Youzhi Liu, Nayu Liu, Li Jin, Zequn Zhang, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4513731/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Embodied Intelligence refers to the agent interacting with the environment, perceiving, planning, decision-making, and executing like humans, which is applicable in smart homes, drone inspections, and other domains. Embodied task planning is one of the main tasks of embodied intelligence, which generates detailed step-by-step plans while perceiving the surrounding environment and understanding language instruction. Visual-language models, with powerful multimodal representation capabilities, have been generalized to various tasks. When applied to embodied task planning, it still faces the following two challenges. Firstly, the intricate complexity of the environment leads to difficulties in global environment information modeling. Secondly, frequent turns in task paths result in the dependence on strong spatial reasoning ability. To overcome these challenges, we propose PlanAgent, the first embodied visual-language model for embodied task planning. Specifically, the environment map is employed to model the global environment information. Then we present the environment map encoder to extract task-related information from the environment. Further, to reduce task path planning's dependence on strong spatial reasoning, we introduce the self-posture-aware training strategy to break down long-term spatial reasoning into short-term. We build the EmbodiedPlan-20k dataset for grounded planning in embodied tasks. Our experiments on the dataset demonstrate that PlanAgent outperforms previous methods and all components are effective. mbodied task planning embodied visual-language model grounded plan environment map Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4513731","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":314801423,"identity":"34396921-bc93-4c33-a512-c67774d46195","order_by":0,"name":"Yuanchang Yue","email":"","orcid":"","institution":"Aerospace Information Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Yuanchang","middleName":"","lastName":"Yue","suffix":""},{"id":314801424,"identity":"ad5d462c-2074-42b8-bb22-d0f5b720cd1a","order_by":1,"name":"Fanglong Yao","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABOUlEQVRIie2PvWrDMBRGrxHEixyvNv1JH0GeOrV+lRhDsnjoWKibygSUKXsG076FySgjcBZDOga8OBQyJVDoZGihsiEdakPo1kEHPt3LlQ6SABSKf8q7zHndiM97gJ5+2tAWcsF1w3EuFfQHBXGD1eXEebISKa2Wj9g8m71y+zkZ9JFeIhzeuqCLEqplW8lHXjTPV9iO8zvuJIXDECYIZ75H8Yho87yt8OC6NFiGySYYci8pNKkAMigaAgSANNZW1nsn+joqaVy4rH6YQZ9cMPfdyiZwpgYLpTLmaUQLjyEgUhEatbpvsTc7b3rBOLYXAQjICr/+SxpnK49ZO5LO20p/7afRgU0uTWv89gFhcfMym23Lffjgmqa/Lau2csWbImQw+Zk2w96x+cWANmUio5cd+wqFQqEA+AZG/G/ycBg2IwAAAABJRU5ErkJggg==","orcid":"","institution":"Aerospace Information Research Institute","correspondingAuthor":true,"prefix":"","firstName":"Fanglong","middleName":"","lastName":"Yao","suffix":""},{"id":314801426,"identity":"ea972d7b-2cde-4488-996e-15a87241f46b","order_by":2,"name":"Youzhi Liu","email":"","orcid":"","institution":"Aerospace Information Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Youzhi","middleName":"","lastName":"Liu","suffix":""},{"id":314801427,"identity":"cb2667eb-1959-4058-97f9-e0fd302b5874","order_by":3,"name":"Nayu Liu","email":"","orcid":"","institution":"Aerospace Information Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Nayu","middleName":"","lastName":"Liu","suffix":""},{"id":314801430,"identity":"00376d2d-aae0-49ec-a4ce-aea9ae213804","order_by":4,"name":"Li Jin","email":"","orcid":"","institution":"Aerospace Information Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Li","middleName":"","lastName":"Jin","suffix":""},{"id":314801431,"identity":"87bf913e-d110-47be-9a0b-fd50942e81fb","order_by":5,"name":"Zequn Zhang","email":"","orcid":"","institution":"Aerospace Information Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Zequn","middleName":"","lastName":"Zhang","suffix":""},{"id":314801432,"identity":"ab9a4157-3c90-45da-a0a7-81848c3a03e6","order_by":6,"name":"Daobing Zhang","email":"","orcid":"","institution":"Aerospace Information Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Daobing","middleName":"","lastName":"Zhang","suffix":""},{"id":314801433,"identity":"b71d64a7-b1c6-4c54-9b20-036dea127c74","order_by":7,"name":"Xian Sun","email":"","orcid":"","institution":"Aerospace Information Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Xian","middleName":"","lastName":"Sun","suffix":""},{"id":314801434,"identity":"809ff35d-4888-4011-9ef6-6980a331f34f","order_by":8,"name":"Kun Fu","email":"","orcid":"","institution":"Aerospace Information Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Kun","middleName":"","lastName":"Fu","suffix":""}],"badges":[],"createdAt":"2024-06-01 13:08:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4513731/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4513731/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":58973719,"identity":"7a0e3d15-d89c-4905-ab1a-84c08cc186da","added_by":"auto","created_at":"2024-06-24 22:22:29","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3083445,"visible":true,"origin":"","legend":"","description":"","filename":"PlanAgentEmbodiedVisualLanguageModelforGroundedTaskplanningwithEnvironmentMap.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4513731/v1_covered_bedab4ba-d7a2-4a43-90ff-f866a329d93d.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"PlanAgent: Embodied Visual-Language Model for Grounded Task planning with Environment Map","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"mbodied task planning, embodied visual-language model, grounded plan, environment map","lastPublishedDoi":"10.21203/rs.3.rs-4513731/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4513731/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Embodied Intelligence refers to the agent interacting with the environment, perceiving, planning, decision-making, and executing like humans, which is applicable in smart homes, drone inspections, and other domains. Embodied task planning is one of the main tasks of embodied intelligence, which generates detailed step-by-step plans while perceiving the surrounding environment and understanding language instruction. Visual-language models, with powerful multimodal representation capabilities, have been generalized to various tasks. When applied to embodied task planning, it still faces the following two challenges. Firstly, the intricate complexity of the environment leads to difficulties in global environment information modeling. Secondly, frequent turns in task paths result in the dependence on strong spatial reasoning ability. To overcome these challenges, we propose PlanAgent, the first embodied visual-language model for embodied task planning. Specifically, the environment map is employed to model the global environment information. Then we present the environment map encoder to extract task-related information from the environment. Further, to reduce task path planning's dependence on strong spatial reasoning, we introduce the self-posture-aware training strategy to break down long-term spatial reasoning into short-term. We build the EmbodiedPlan-20k dataset for grounded planning in embodied tasks. Our experiments on the dataset demonstrate that PlanAgent outperforms previous methods and all components are effective.","manuscriptTitle":"PlanAgent: Embodied Visual-Language Model for Grounded Task planning with Environment Map","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-06-19 04:41:00","doi":"10.21203/rs.3.rs-4513731/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7ad7b9c9-ab9e-4319-9da5-bc093194e82d","owner":[],"postedDate":"June 19th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-06-24T22:14:18+00:00","versionOfRecord":[],"versionCreatedAt":"2024-06-19 04:41:00","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4513731","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4513731","identity":"rs-4513731","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00