Assessing the 3D position of a car with a single 2D camera using mainstream DCNN models

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Deep Convolutional Neural Networks (DCNNs) are regarded as one of the foundations of computer vision due to their unparalleled ability to process visual data. This work explores the use of DCNNs to estimate the orientation of vehicles from a single 2D image. The experiments included 48 training scenarios encompassing four dataset variations and 12 models, followed by an evaluation of each one based on four key metrics. Overall the best-performing architecture was EfficientNet-B2, achieving an accuracy of 97.22% consistently across all dataset variations, demonstrating its robustness to preprocessing techniques. Additionally, ResNet18 delivered competitive results, achieving the highest recorded accuracy of 98.61% on the original dataset, while MobileNetV2 also performed exceptionally well on the augmented and no-background datasets, reaching 98.61%. EfficientNet-B5 initially underperformed on the original dataset but significantly improved with augmentation, achieving 97.22% accuracy. The study revealed that dataset preprocessing played a crucial role in model performance, with augmentation and background removal significantly boosting accuracy. The classification mistakes were further analyzed using SHAP values, highlighting the importance of specific car features, such as the front and rear sections, in determining orientation. Overall, the results confirmed that vehicle orientation estimation can be effectively approached as a classification problem. The ResNet family proved highly robust, while EfficientNet-B2 emerged as a strong contender due to its consistency. MobileNetV2's efficiency and strong performance make it a viable option for real-time applications. Future work should explore transformer-based architectures and evaluate model performance on real-world datasets with varying environmental conditions.
Full text 11,932 characters · extracted from preprint-html · click to expand
Assessing the 3D position of a car with a single 2D camera using mainstream DCNN models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Assessing the 3D position of a car with a single 2D camera using mainstream DCNN models Youssef Yahia, Júlio Castro Lopes, Eduardo Bezerra, Pedro João Rodrigues, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5401929/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Deep Convolutional Neural Networks (DCNNs) are regarded as one of the foundations of computer vision due to their unparalleled ability to process visual data. This work explores the use of DCNNs to estimate the orientation of vehicles from a single 2D image. The experiments included 48 training scenarios encompassing four dataset variations and 12 models, followed by an evaluation of each one based on four key metrics. Overall the best-performing architecture was EfficientNet-B2, achieving an accuracy of 97.22% consistently across all dataset variations, demonstrating its robustness to preprocessing techniques. Additionally, ResNet18 delivered competitive results, achieving the highest recorded accuracy of 98.61% on the original dataset, while MobileNetV2 also performed exceptionally well on the augmented and no-background datasets, reaching 98.61%. EfficientNet-B5 initially underperformed on the original dataset but significantly improved with augmentation, achieving 97.22% accuracy. The study revealed that dataset preprocessing played a crucial role in model performance, with augmentation and background removal significantly boosting accuracy. The classification mistakes were further analyzed using SHAP values, highlighting the importance of specific car features, such as the front and rear sections, in determining orientation. Overall, the results confirmed that vehicle orientation estimation can be effectively approached as a classification problem. The ResNet family proved highly robust, while EfficientNet-B2 emerged as a strong contender due to its consistency. MobileNetV2's efficiency and strong performance make it a viable option for real-time applications. Future work should explore transformer-based architectures and evaluate model performance on real-world datasets with varying environmental conditions. Computer Vision Deep Learning Orientation Estimation SHAP EffcientNet ResNet MobileNet Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5401929","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":449216123,"identity":"11846bed-3bf4-4670-b38f-366ade122b45","order_by":0,"name":"Youssef Yahia","email":"","orcid":"","institution":"Polytechnic Institute of Bragança","correspondingAuthor":false,"prefix":"","firstName":"Youssef","middleName":"","lastName":"Yahia","suffix":""},{"id":449216124,"identity":"a280302e-b46c-4866-81ea-7ab271039a8c","order_by":1,"name":"Júlio Castro Lopes","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAxklEQVRIiWNgGAWjYBACxgbmhgMMBSAm84EDjA1EaWEEajEAMdkSwFp4iNHEANHCY8BAlBbmGYmNBz4Y2OTx85/5eLhwB0OePUE7ZiQ2HJxhkFYsOSN3w+GZZxiKCdoC0nKYx+Bw4oYbvBsO87YxJPYQrWX/+TMPSNSygSGHgUgtPQ8hfpG4kWYA9ItEMc8BAloM25MPf/hQAQyx/sOPPxfusMljbyCkBaogAUQwMzBIJBByF4M8A6oWBsJaRsEoGAWjYMQBAJl/SJPgbxYEAAAAAElFTkSuQmCC","orcid":"","institution":"Polytechnic Institute of Bragança","correspondingAuthor":true,"prefix":"","firstName":"Júlio","middleName":"Castro","lastName":"Lopes","suffix":""},{"id":449216125,"identity":"27c2de7d-5a66-4ab7-8856-e14844cf013f","order_by":2,"name":"Eduardo Bezerra","email":"","orcid":"","institution":"Federal Center for Technological Education (CEFET/RJ)","correspondingAuthor":false,"prefix":"","firstName":"Eduardo","middleName":"","lastName":"Bezerra","suffix":""},{"id":449216126,"identity":"41264846-e6ab-43db-9ea3-410909f016c1","order_by":3,"name":"Pedro João Rodrigues","email":"","orcid":"","institution":"Polytechnic Institute of Bragança","correspondingAuthor":false,"prefix":"","firstName":"Pedro","middleName":"João","lastName":"Rodrigues","suffix":""},{"id":449216127,"identity":"19bc6d75-0283-4850-919f-e46ea51db4de","order_by":4,"name":"Rui Pedro Lopes","email":"","orcid":"","institution":"Polytechnic Institute of Bragança","correspondingAuthor":false,"prefix":"","firstName":"Rui","middleName":"Pedro","lastName":"Lopes","suffix":""}],"badges":[],"createdAt":"2024-11-06 10:38:08","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5401929/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5401929/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":84860197,"identity":"7235c1e1-2e46-43a7-9249-9719bad91078","added_by":"auto","created_at":"2025-06-18 06:47:00","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1655852,"visible":true,"origin":"","legend":"","description":"","filename":"ArtixramProject2reviewed2.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5401929/v1_covered_ea7f4fad-595f-49b9-97a2-2855e776ce6e.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Assessing the 3D position of a car with a single 2D camera using mainstream DCNN models","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Computer Vision, Deep Learning, Orientation Estimation, SHAP, EffcientNet, ResNet, MobileNet","lastPublishedDoi":"10.21203/rs.3.rs-5401929/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5401929/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Deep Convolutional Neural Networks (DCNNs) are regarded as one of the foundations of computer vision due to their unparalleled ability to process visual data. This work explores the use of DCNNs to estimate the orientation of vehicles from a single 2D image. The experiments included 48 training scenarios encompassing four dataset variations and 12 models, followed by an evaluation of each one based on four key metrics. Overall the best-performing architecture was EfficientNet-B2, achieving an accuracy of 97.22% consistently across all dataset variations, demonstrating its robustness to preprocessing techniques. Additionally, ResNet18 delivered competitive results, achieving the highest recorded accuracy of 98.61% on the original dataset, while MobileNetV2 also performed exceptionally well on the augmented and no-background datasets, reaching 98.61%. EfficientNet-B5 initially underperformed on the original dataset but significantly improved with augmentation, achieving 97.22% accuracy. The study revealed that dataset preprocessing played a crucial role in model performance, with augmentation and background removal significantly boosting accuracy. The classification mistakes were further analyzed using SHAP values, highlighting the importance of specific car features, such as the front and rear sections, in determining orientation. Overall, the results confirmed that vehicle orientation estimation can be effectively approached as a classification problem. The ResNet family proved highly robust, while EfficientNet-B2 emerged as a strong contender due to its consistency. MobileNetV2's efficiency and strong performance make it a viable option for real-time applications. Future work should explore transformer-based architectures and evaluate model performance on real-world datasets with varying environmental conditions.","manuscriptTitle":"Assessing the 3D position of a car with a single 2D camera using mainstream DCNN models","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-29 12:57:42","doi":"10.21203/rs.3.rs-5401929/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0b1c8474-4f1d-4866-aa29-ed9eddc833b6","owner":[],"postedDate":"April 29th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-06-18T06:38:46+00:00","versionOfRecord":[],"versionCreatedAt":"2025-04-29 12:57:42","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5401929","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5401929","identity":"rs-5401929","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00