2D Image to 3D Box with Deep Learning Method | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article 2D Image to 3D Box with Deep Learning Method DONGJUN WU This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2960750/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract This article will show you a method for converting 2d images into 3d bounding boxes. The idea is to split the task into two steps. Firstly, we use neural network to output 3D object features displayed in 2D images. Then, we combine the boundary information of 2D objects in the images to generate geometric constraints on 3D objects, so as to get a complete 3D boundary box. In the first step, we defined an innovative loss function called combined_orientation_loss. This loss function took into account both the cosine distance and Euclidean distance, so that the length of the direction vector could also be reflected in the loss, which could enhance the sensitivity of the model to the prediction accuracy. In the second step, we return to the dimensions of 3D objects, which are relatively stable and therefore suitable for predicting a wide range of object types. Combining these two steps and the translation constraints of the object in the 2D image, we can accurately and stably restore its 3D border shape. We evaluated our approach on the KITTI target detection benchmark, including the official measure of 3D direction estimates and the accuracy of the resulting 3D bounding box. In the case of low calculation requirements achieved good results. 2D images 3D frame Deep learning Object features Geometric constraints Cosine distance Euclidean distance Dimension prediction Translation constraints Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2960750","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":202669988,"identity":"6cd7d6bc-112f-4cb8-81f9-a2e74c442043","order_by":0,"name":"DONGJUN WU","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7ElEQVRIiWNgGAWjYBACgxuJDQcSKg4n9sOFDhDUknzwwYczhwtnNsBUE9aSlmw4s+Vw5YYDxGvJMZPmbThcufl4A9vjj20Mcnw3Epg/fsGjxQysZcfh3G1nDrAbHGxjMJa8kcAmLUNQyxmgFqBKCaCWxA1ABrMEQS1th9M3z38A1lIP1ML8GZ8We7D32w4nb5BgAGtJMLiRwCD5AY8WyzOHQYGcljzjTGKbxJlzEoYzzzxsk8ajg8HgeCMoKm0S+9sPH5OoKLOR5zuefPjjD3x6EICxAUhIgBnMPMRpQdZNpC2jYBSMglEwMgAAuqFjEYyjflEAAAAASUVORK5CYII=","orcid":"","institution":"","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"DONGJUN","middleName":"","lastName":"WU","suffix":""}],"badges":[],"createdAt":"2023-05-20 17:29:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2960750/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2960750/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":37413868,"identity":"63126874-29d1-4d41-b430-ea51a0f90dcf","added_by":"auto","created_at":"2023-05-24 04:44:44","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":352021,"visible":true,"origin":"","legend":"","description":"","filename":"2DImageto3DBoxwithDeepLearningMethod.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2960750/v1_covered_702be7f6-ffb0-470d-bb0b-22f6c5a37ad0.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"2D Image to 3D Box with Deep Learning Method","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"2D images, 3D frame, Deep learning, Object features, Geometric constraints, Cosine distance, Euclidean distance, Dimension prediction, Translation constraints","lastPublishedDoi":"10.21203/rs.3.rs-2960750/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2960750/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"This article will show you a method for converting 2d images into 3d bounding boxes. The idea is to split the task into two steps. Firstly, we use neural network to output 3D object features displayed in 2D images. Then, we combine the boundary information of 2D objects in the images to generate geometric constraints on 3D objects, so as to get a complete 3D boundary box. In the first step, we defined an innovative loss function called combined_orientation_loss. This loss function took into account both the cosine distance and Euclidean distance, so that the length of the direction vector could also be reflected in the loss, which could enhance the sensitivity of the model to the prediction accuracy. In the second step, we return to the dimensions of 3D objects, which are relatively stable and therefore suitable for predicting a wide range of object types. Combining these two steps and the translation constraints of the object in the 2D image, we can accurately and stably restore its 3D border shape. We evaluated our approach on the KITTI target detection benchmark, including the official measure of 3D direction estimates and the accuracy of the resulting 3D bounding box. In the case of low calculation requirements achieved good results.","manuscriptTitle":"2D Image to 3D Box with Deep Learning Method","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-05-24 03:35:36","doi":"10.21203/rs.3.rs-2960750/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2196f093-a667-4ec6-8fcc-3815c8c3113e","owner":[],"postedDate":"May 24th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2023-05-24T04:44:29+00:00","versionOfRecord":[],"versionCreatedAt":"2023-05-24 03:35:36","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-2960750","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2960750","identity":"rs-2960750","version":["v1"]},"buildId":"cBFmMYwuxLRRLfASyISRj","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.