NavAI: A Generalizable LLM Framework for Navigation Tasks in Virtual Reality Environments

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Navigation is a fundamental capability for automated exploration in immersive virtual reality (VR) environments. While prior work has largely focused on path optimization in 360-degree image datasets or in pre-trained 3D simulators, such approaches are difficult to apply directly to general VR applications involving unseen scenes and dynamic interactions. To address this gap, we present NavAI, a general and extensible navigation framework that leverages large language models (LLMs) to support both basic action commands and multi-step goal-oriented navigation in arbitrary VR environments. NavAI decomposes navigation into perception, decision-making, and control execution, and explores multiple model variants that balance effectiveness, efficiency, and robustness. In particular, NavAI establishes an LLM-based navigation baseline and explores optimization strategies for virtual scene understanding and navigation goal decision-making, enabling systematic comparison between the baseline and optimized designs. Evaluation shows that the voting-based Full Mode achieves a 100% success rate across all tested environments and distance regimes, while Gemini-based single-model variants (with and without YOLO) match this effectiveness without ensemble voting. Moreover, YOLO-assisted variants reduce perception latency by up to ≈ 40% and lower token consumption by as much as ≈ 48% compared to LLM-only baselines, whereas Full Mode incurs up to 3.5× higher token cost due to redundant decision-making. Based on these findings, we discuss the limitations of LLM-driven navigation and outline directions for future optimizationand practical deployment in immersive VR systems.
Full text 13,141 characters · extracted from preprint-html · click to expand
NavAI: A Generalizable LLM Framework for Navigation Tasks in Virtual Reality Environments | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article NavAI: A Generalizable LLM Framework for Navigation Tasks in Virtual Reality Environments Jiajie Wang, Sumesh Surendran Letha, Matthew DiGiovanni, Kebin Peng, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9034054/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 6 You are reading this latest preprint version Abstract Navigation is a fundamental capability for automated exploration in immersive virtual reality (VR) environments. While prior work has largely focused on path optimization in 360-degree image datasets or in pre-trained 3D simulators, such approaches are difficult to apply directly to general VR applications involving unseen scenes and dynamic interactions. To address this gap, we present NavAI, a general and extensible navigation framework that leverages large language models (LLMs) to support both basic action commands and multi-step goal-oriented navigation in arbitrary VR environments. NavAI decomposes navigation into perception, decision-making, and control execution, and explores multiple model variants that balance effectiveness, efficiency, and robustness. In particular, NavAI establishes an LLM-based navigation baseline and explores optimization strategies for virtual scene understanding and navigation goal decision-making, enabling systematic comparison between the baseline and optimized designs. Evaluation shows that the voting-based Full Mode achieves a 100% success rate across all tested environments and distance regimes, while Gemini-based single-model variants (with and without YOLO) match this effectiveness without ensemble voting. Moreover, YOLO-assisted variants reduce perception latency by up to ≈ 40% and lower token consumption by as much as ≈ 48% compared to LLM-only baselines, whereas Full Mode incurs up to 3.5× higher token cost due to redundant decision-making. Based on these findings, we discuss the limitations of LLM-driven navigation and outline directions for future optimizationand practical deployment in immersive VR systems. Virtual Reality Large Language Model Vision Language Model Navigation Performance Evaluation Full Text Additional Declarations Competing interest reported. Institution COI with: The University of Arizona East Carolina University Villanova University Cite Share Download PDF Status: Under Review Version 1 posted Reviewers agreed at journal 14 May, 2026 Reviewers agreed at journal 20 Mar, 2026 Reviewers invited by journal 18 Mar, 2026 Editor assigned by journal 16 Mar, 2026 Submission checks completed at journal 06 Mar, 2026 First submitted to journal 04 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9034054","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":609722552,"identity":"884edea0-9c69-465c-9607-50124db36f67","order_by":0,"name":"Jiajie Wang","email":"","orcid":"","institution":"University of Arizona","correspondingAuthor":false,"prefix":"","firstName":"Jiajie","middleName":"","lastName":"Wang","suffix":""},{"id":609722553,"identity":"6b4db13c-1c13-4955-9329-88a0c8ae0516","order_by":1,"name":"Sumesh Surendran Letha","email":"","orcid":"","institution":"Villanova University","correspondingAuthor":false,"prefix":"","firstName":"Sumesh","middleName":"Surendran","lastName":"Letha","suffix":""},{"id":609722555,"identity":"960d6351-c0f5-42fa-b928-8c020780c6a3","order_by":2,"name":"Matthew DiGiovanni","email":"","orcid":"","institution":"Villanova University","correspondingAuthor":false,"prefix":"","firstName":"Matthew","middleName":"","lastName":"DiGiovanni","suffix":""},{"id":609722560,"identity":"303b45f0-bec7-49db-8031-cfd63f624183","order_by":3,"name":"Kebin Peng","email":"","orcid":"","institution":"East Carolina University","correspondingAuthor":false,"prefix":"","firstName":"Kebin","middleName":"","lastName":"Peng","suffix":""},{"id":609722562,"identity":"7987b9a1-b673-4a15-9d62-18c0ddd5efa3","order_by":4,"name":"Sen He","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAxElEQVRIiWNgGAWjYHACNgaGCjAJBowNxGk5Q7IWxjYEj7AW+fbeYw8+zrPL5+M/fOzBDwYb2Q0HCGgxOHMu3XDmtmTLNoZj6YY9DGnGhLVI5JhJ825jNmBj7DGT4GE4nEhQi/z8N2bSf+fUG7Ax85hJ/mH4T1gLww0eM2nGhsMGbGxABg/DAcJaDM7kmEn2HDtuwMbDliYtY5BsPJOgw9rPmEn8qKk2kO8/fEzyTYWdbB9Bh6FZSpryUTAKRsEoGAU4AADX2jptk9KMmQAAAABJRU5ErkJggg==","orcid":"","institution":"University of Arizona","correspondingAuthor":true,"prefix":"","firstName":"Sen","middleName":"","lastName":"He","suffix":""}],"badges":[],"createdAt":"2026-03-04 22:23:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9034054/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9034054/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105563959,"identity":"2a9e36ee-8c54-4501-bc81-4c99d628d26d","added_by":"auto","created_at":"2026-03-27 12:48:18","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":700929,"visible":true,"origin":"","legend":"","description":"","filename":"NavAI.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9034054/v1_covered_2cda4e3d-2f90-4deb-b0fd-a30ab335adae.pdf"}],"financialInterests":"Competing interest reported. Institution COI with:\nThe University of Arizona\nEast Carolina University\nVillanova University","formattedTitle":"NavAI: A Generalizable LLM Framework for Navigation Tasks in Virtual Reality Environments","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"automated-software-engineering","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ause","sideBox":"Learn more about [Automated Software Engineering](http://link.springer.com/journal/10515)","snPcode":"10515","submissionUrl":"https://submission.nature.com/new-submission/10515/3","title":"Automated Software Engineering","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Virtual Reality, Large Language Model, Vision Language Model, Navigation, Performance Evaluation","lastPublishedDoi":"10.21203/rs.3.rs-9034054/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9034054/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Navigation is a fundamental capability for automated exploration in immersive virtual reality (VR) environments. While prior work has largely focused on path optimization in 360-degree image datasets or in pre-trained 3D simulators, such approaches are difficult to apply directly to general VR applications involving unseen scenes and dynamic interactions. To address this gap, we present NavAI, a general and extensible navigation framework that leverages large language models (LLMs) to support both basic action commands and multi-step goal-oriented navigation in arbitrary VR environments. NavAI decomposes navigation into perception, decision-making, and control execution, and explores multiple model variants that balance effectiveness, efficiency, and robustness. In particular, NavAI establishes an LLM-based navigation baseline and explores optimization strategies for virtual scene understanding and navigation goal decision-making, enabling systematic comparison between the baseline and optimized designs. Evaluation shows that the voting-based Full Mode achieves a 100% success rate across all tested environments and distance regimes, while Gemini-based single-model variants (with and without YOLO) match this effectiveness without ensemble voting. Moreover, YOLO-assisted variants reduce perception latency by up to ≈ 40% and lower token consumption by as much as ≈ 48% compared to LLM-only baselines, whereas Full Mode incurs up to 3.5× higher token cost due to redundant decision-making. Based on these findings, we discuss the limitations of LLM-driven navigation and outline directions for future optimizationand practical deployment in immersive VR systems.","manuscriptTitle":"NavAI: A Generalizable LLM Framework for Navigation Tasks in Virtual Reality Environments","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-23 04:18:30","doi":"10.21203/rs.3.rs-9034054/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"121374155391890303564962031197455440042","date":"2026-05-14T22:08:36+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"335761471932660000762192668249806217656","date":"2026-03-21T03:20:04+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-19T01:18:28+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-16T13:58:36+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-06T05:01:45+00:00","index":"","fulltext":""},{"type":"submitted","content":"Automated Software Engineering","date":"2026-03-04T22:13:20+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"automated-software-engineering","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ause","sideBox":"Learn more about [Automated Software Engineering](http://link.springer.com/journal/10515)","snPcode":"10515","submissionUrl":"https://submission.nature.com/new-submission/10515/3","title":"Automated Software Engineering","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"57e40e0f-674d-4de0-9108-0a515ad414ab","owner":[],"postedDate":"March 23rd, 2026","published":true,"recentEditorialEvents":[{"type":"reviewerAgreed","content":"121374155391890303564962031197455440042","date":"2026-05-14T22:08:36+00:00","index":32,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-03-23T04:18:31+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-23 04:18:30","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9034054","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9034054","identity":"rs-9034054","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00