A Framework for Real-Time Stereoscopic Virtual Reality Conversion of Mainstream Game Output Using Synthetic Supervision and GPU-Resident Inference | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Framework for Real-Time Stereoscopic Virtual Reality Conversion of Mainstream Game Output Using Synthetic Supervision and GPU-Resident Inference Sahan Fernando, Dinesh Asanka This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9399054/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Mainstream digital games are overwhelmingly designed for monoscopic displays, whereas native virtual reality (VR) adaptation remains expensive, title-specific, and often unavailable for legacy or commercially closed content. This paper presents a render-stream-based framework for converting mainstream game output into stereoscopic VR from the externally rendered RGB stream, without relying on engine-internal geometry, depth, or pose buffers at inference time. The framework combines a Unity-based synthetic supervision pipeline, a two-stage neural architecture for monocular depth estimation and right-eye synthesis, and a GPU-resident deployment path built around mixed-precision TensorRT inference with DirectX-CUDA zero-copy integration. A task-specific dataset of 7,000 aligned left-right-depth samples was generated in Unity, and all experiments were conducted at a fixed 704 x 704 working resolution on consumer RTX-class hardware. Under joint quality, latency, and exportability screening, the final deployed pair consisted of DepthNet-EE and HybridStereoNet. The final runtime path reduced end-to-end latency from 142 ms in the initial HTTP baseline to 21.69 ms in the final zero-copy deployment. On the held-out test set, the final TensorRT engine achieved 37.09 dB peak signal-to-noise ratio, 0.9655 structural similarity, and 0.9311 on the reported BQA proxy, while integrated headset validation showed stable stereoscopic presentation at 45–50 FPS with Asynchronous Spacewarp disabled. Taken together, these results show that near-real-time stereoscopic VR conversion from rendered RGB streams is technically feasible under controlled conditions. Virtual reality stereoscopic rendering monocular-to-stereo conversion real-time inference Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9399054","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":627121221,"identity":"0b8999f3-6a90-4c39-bd72-a96647d8aef3","order_by":0,"name":"Sahan Fernando","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5ElEQVRIie3RsYrCMBzH8X8J/LsEb40U2ldo6CAFuWdJOWgXEcGloGDgIL5CBx8mJVCXPsCNHoKu3gucVt0cmo4O+Q4ZAh/4hQC4XG9YzLqjhGmIvvL048pHK/FkC3kyog08CRlGTBayfCCZBNvzQZdE4PgkzAI+o46IXpLuWi51i3MMcm0q+OKSoO4f9jPzvv8UXWJQSEOBCCC+tJDiV9b/LFPj/Z1shhDBZS3jTDHUHTEdsQxLqxmvdCMSpLkwNN5zZXv+hBWHi15fw2jb8CMtV9GHr+Je8roTwPIrLpfL5RrSDczWRBm8SWKsAAAAAElFTkSuQmCC","orcid":"","institution":"University of Kelaniya","correspondingAuthor":true,"prefix":"","firstName":"Sahan","middleName":"","lastName":"Fernando","suffix":""},{"id":627121222,"identity":"9ca73b22-a17a-48cd-9c57-8546e4a26104","order_by":1,"name":"Dinesh Asanka","email":"","orcid":"","institution":"University of Kelaniya","correspondingAuthor":false,"prefix":"","firstName":"Dinesh","middleName":"","lastName":"Asanka","suffix":""}],"badges":[],"createdAt":"2026-04-13 05:38:32","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9399054/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9399054/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":107756685,"identity":"bf93fa0c-6067-475d-80e0-2f52e07c62c6","added_by":"auto","created_at":"2026-04-24 19:21:52","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1096469,"visible":true,"origin":"","legend":"","description":"","filename":"AFrameworkforRealTimeStereoscopicVirtualRealityConversionofMainstreamGameOutputUsingSyntheticSupervisionandGPUResidentInference.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9399054/v1_covered_64c7f33e-4ef3-4162-9a20-14aeee1db0c3.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Framework for Real-Time Stereoscopic Virtual Reality Conversion of Mainstream Game Output Using Synthetic Supervision and GPU-Resident Inference","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Virtual reality, stereoscopic rendering, monocular-to-stereo conversion, real-time inference","lastPublishedDoi":"10.21203/rs.3.rs-9399054/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9399054/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eMainstream digital games are overwhelmingly designed for monoscopic displays, whereas native virtual reality (VR) adaptation remains expensive, title-specific, and often unavailable for legacy or commercially closed content. This paper presents a render-stream-based framework for converting mainstream game output into stereoscopic VR from the externally rendered RGB stream, without relying on engine-internal geometry, depth, or pose buffers at inference time. The framework combines a Unity-based synthetic supervision pipeline, a two-stage neural architecture for monocular depth estimation and right-eye synthesis, and a GPU-resident deployment path built around mixed-precision TensorRT inference with DirectX-CUDA zero-copy integration. A task-specific dataset of 7,000 aligned left-right-depth samples was generated in Unity, and all experiments were conducted at a fixed 704 x 704 working resolution on consumer RTX-class hardware. Under joint quality, latency, and exportability screening, the final deployed pair consisted of DepthNet-EE and HybridStereoNet. The final runtime path reduced end-to-end latency from 142 ms in the initial HTTP baseline to 21.69 ms in the final zero-copy deployment. On the held-out test set, the final TensorRT engine achieved 37.09 dB peak signal-to-noise ratio, 0.9655 structural similarity, and 0.9311 on the reported BQA proxy, while integrated headset validation showed stable stereoscopic presentation at 45\u0026ndash;50 FPS with Asynchronous Spacewarp disabled. Taken together, these results show that near-real-time stereoscopic VR conversion from rendered RGB streams is technically feasible under controlled conditions.\u003c/p\u003e","manuscriptTitle":"A Framework for Real-Time Stereoscopic Virtual Reality Conversion of Mainstream Game Output Using Synthetic Supervision and GPU-Resident Inference","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-24 19:21:32","doi":"10.21203/rs.3.rs-9399054/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"261ec6fe-a3d4-40de-948a-3901de2b7036","owner":[],"postedDate":"April 24th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-24T19:21:33+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-24 19:21:32","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9399054","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9399054","identity":"rs-9399054","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.