2D Image to 3D Box with Deep Learning Method

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-16

This paper presents a two-step deep learning method using a novel loss function to convert 2D images into 3D bounding boxes, achieving good results on the KITTI benchmark with low computational requirements.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-16 · read from full text

The paper proposes a two-step deep learning method for converting 2D images into 3D bounding boxes by first using a neural network to output 3D object features and then combining 2D object boundary information to generate geometric constraints that recover a complete 3D box. In the first step, the authors introduce a combined_orientation_loss that accounts for both cosine distance and Euclidean distance to improve sensitivity to direction and vector length in predictions. They then predict relatively stable 3D dimensions and incorporate translation constraints from the 2D image to produce the final 3D boundary shape. The approach is evaluated on the KITTI target detection benchmark using standard 3D direction and bounding-box accuracy metrics, and the authors report good performance with low computational requirements, with the caveat that it is a Research Square preprint not peer reviewed by a journal. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

This article will show you a method for converting 2d images into 3d bounding boxes. The idea is to split the task into two steps. Firstly, we use neural network to output 3D object features displayed in 2D images. Then, we combine the boundary information of 2D objects in the images to generate geometric constraints on 3D objects, so as to get a complete 3D boundary box. In the first step, we defined an innovative loss function called combined_orientation_loss. This loss function took into account both the cosine distance and Euclidean distance, so that the length of the direction vector could also be reflected in the loss, which could enhance the sensitivity of the model to the prediction accuracy. In the second step, we return to the dimensions of 3D objects, which are relatively stable and therefore suitable for predicting a wide range of object types. Combining these two steps and the translation constraints of the object in the 2D image, we can accurately and stably restore its 3D border shape. We evaluated our approach on the KITTI target detection benchmark, including the official measure of 3D direction estimates and the accuracy of the resulting 3D bounding box. In the case of low calculation requirements achieved good results.
Full text 9,501 characters · extracted from preprint-html · click to expand
2D Image to 3D Box with Deep Learning Method | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article 2D Image to 3D Box with Deep Learning Method DONGJUN WU This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2960750/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract This article will show you a method for converting 2d images into 3d bounding boxes. The idea is to split the task into two steps. Firstly, we use neural network to output 3D object features displayed in 2D images. Then, we combine the boundary information of 2D objects in the images to generate geometric constraints on 3D objects, so as to get a complete 3D boundary box. In the first step, we defined an innovative loss function called combined_orientation_loss. This loss function took into account both the cosine distance and Euclidean distance, so that the length of the direction vector could also be reflected in the loss, which could enhance the sensitivity of the model to the prediction accuracy. In the second step, we return to the dimensions of 3D objects, which are relatively stable and therefore suitable for predicting a wide range of object types. Combining these two steps and the translation constraints of the object in the 2D image, we can accurately and stably restore its 3D border shape. We evaluated our approach on the KITTI target detection benchmark, including the official measure of 3D direction estimates and the accuracy of the resulting 3D bounding box. In the case of low calculation requirements achieved good results. 2D images 3D frame Deep learning Object features Geometric constraints Cosine distance Euclidean distance Dimension prediction Translation constraints Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2960750","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":202669988,"identity":"6cd7d6bc-112f-4cb8-81f9-a2e74c442043","order_by":0,"name":"DONGJUN WU","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7ElEQVRIiWNgGAWjYBACgxuJDQcSKg4n9sOFDhDUknzwwYczhwtnNsBUE9aSlmw4s+Vw5YYDxGvJMZPmbThcufl4A9vjj20Mcnw3Epg/fsGjxQysZcfh3G1nDrAbHGxjMJa8kcAmLUNQyxmgFqBKCaCWxA1ABrMEQS1th9M3z38A1lIP1ML8GZ8We7D32w4nb5BgAGtJMLiRwCD5AY8WyzOHQYGcljzjTGKbxJlzEoYzzzxsk8ajg8HgeCMoKm0S+9sPH5OoKLOR5zuefPjjD3x6EICxAUhIgBnMPMRpQdZNpC2jYBSMglEwMgAAuqFjEYyjflEAAAAASUVORK5CYII=","orcid":"","institution":"","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"DONGJUN","middleName":"","lastName":"WU","suffix":""}],"badges":[],"createdAt":"2023-05-20 17:29:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2960750/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2960750/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":37413868,"identity":"63126874-29d1-4d41-b430-ea51a0f90dcf","added_by":"auto","created_at":"2023-05-24 04:44:44","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":352021,"visible":true,"origin":"","legend":"","description":"","filename":"2DImageto3DBoxwithDeepLearningMethod.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2960750/v1_covered_702be7f6-ffb0-470d-bb0b-22f6c5a37ad0.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"2D Image to 3D Box with Deep Learning Method","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"2D images, 3D frame, Deep learning, Object features, Geometric constraints, Cosine distance, Euclidean distance, Dimension prediction, Translation constraints","lastPublishedDoi":"10.21203/rs.3.rs-2960750/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2960750/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"This article will show you a method for converting 2d images into 3d bounding boxes. The idea is to split the task into two steps. Firstly, we use neural network to output 3D object features displayed in 2D images. Then, we combine the boundary information of 2D objects in the images to generate geometric constraints on 3D objects, so as to get a complete 3D boundary box. In the first step, we defined an innovative loss function called combined_orientation_loss. This loss function took into account both the cosine distance and Euclidean distance, so that the length of the direction vector could also be reflected in the loss, which could enhance the sensitivity of the model to the prediction accuracy. In the second step, we return to the dimensions of 3D objects, which are relatively stable and therefore suitable for predicting a wide range of object types. Combining these two steps and the translation constraints of the object in the 2D image, we can accurately and stably restore its 3D border shape. We evaluated our approach on the KITTI target detection benchmark, including the official measure of 3D direction estimates and the accuracy of the resulting 3D bounding box. In the case of low calculation requirements achieved good results.","manuscriptTitle":"2D Image to 3D Box with Deep Learning Method","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-05-24 03:35:36","doi":"10.21203/rs.3.rs-2960750/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2196f093-a667-4ec6-8fcc-3815c8c3113e","owner":[],"postedDate":"May 24th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2023-05-24T04:44:29+00:00","versionOfRecord":[],"versionCreatedAt":"2023-05-24 03:35:36","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-2960750","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2960750","identity":"rs-2960750","version":["v1"]},"buildId":"cBFmMYwuxLRRLfASyISRj","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-29T02:00:03.542394+00:00
License: CC-BY-4.0