YOLO-ROC: A High-Precision and Ultra-Lightweight Model for Real-Time Road Damage Detection | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article YOLO-ROC: A High-Precision and Ultra-Lightweight Model for Real-Time Road Damage Detection Zicheng Lin, Weichao Pan This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7221917/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Road damage detection is a critical task for ensuring traffic safety and maintaining infrastructure integrity. While deep learning-based detection methods are now widely adopted, they still face two core challenges: first, the inadequate multi-scale feature extraction capabilities of existing networks for diverse targets like cracks and potholes, leading to high miss rates for small-scale damage; and second, the substantial parameter counts and computational demands of mainstream models, which hinder their deployment for efficient, real-time detection in practical applications. To address these issues, this paper proposes a high-precision and lightweight model, YOLO - Road Orthogonal Compact (YOLO-ROC). We designed a Bidirectional Multi-scale Spatial Pyramid Pooling Fast (BMS-SPPF) module to enhance multi-scale feature extraction and implemented a hierarchical channel compression strategy to reduce computational complexity. The BMS-SPPF module leverages a bidirectional spatialchannel attention mechanism to improve the detection of small targets. Concurrently, the channel compression strategy reduces the parameter count from 3.01M to 0.89M and GFLOPs from 8.1 to 2.6. Experiments on the RDD2022 China Drone dataset demonstrate that YOLO-ROC achieves a mAP50 of 67.6%, surpassing the baseline YOLOv8n by 2.11%. Notably, the mAP50 for the small-target D40 category improved by 16.8%, and the final model size is only 2.0 MB. Furthermore, the model exhibits excellent generalization performance on the RDD2022 China Motorbike dataset. Road Damage Detection YOLO Attention Mechanism Lightweight Model Multi-scale Feature Extraction Channel Compression Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7221917","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":492642203,"identity":"a7a31b30-fd9f-47e5-864c-7bf695a6530a","order_by":0,"name":"Zicheng Lin","email":"","orcid":"","institution":"Shandong Jianzhu University","correspondingAuthor":false,"prefix":"","firstName":"Zicheng","middleName":"","lastName":"Lin","suffix":""},{"id":492642204,"identity":"3a4f84ed-1582-4371-b05b-6b1eab15ca54","order_by":1,"name":"Weichao Pan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABM0lEQVRIie3RMWuDQBTA8ScHdjnieiK1X+GJEDIE8lVOAunSgF1Kp/ZK4Lq0dLXfIlNnxcFFOjta+gUsGWrAIc+0m6bNWKh/DnyoP7hTgKGhP5jFWFJKFHyWsqyqvm4i7i9xP7Hv9RzLcOJCphdGBPg7wTwf22V17QMNjH8T+IlAQbuiFdxFcvw+bZpAWen6cqvBHRXS2IRdYURSYktWQp77S42BEosQbQ2+XUjmRF3ChIxlSzQNzlK1hCN6GoJ1IU3aaidTBCpuyQMNzqQhYuWIgYbbQ4Tz1FBEfMFT0wGTCFwgJhokHiDiRDN6KlykwXvUvq/pLJ56Fd5z/rZyesgstT63dXPDkVkfZd24p09W+uLVV9OzUTZPNj2k53S06MeI9nOqY8A+Vh796tDQ0NB/aAdWVmh48+v+gwAAAABJRU5ErkJggg==","orcid":"","institution":"Shandong Jianzhu University","correspondingAuthor":true,"prefix":"","firstName":"Weichao","middleName":"","lastName":"Pan","suffix":""}],"badges":[],"createdAt":"2025-07-26 14:53:10","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7221917/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7221917/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":90737394,"identity":"ec118a8f-7b16-41c2-94cf-ddfc9f73a89e","added_by":"auto","created_at":"2025-09-06 22:46:22","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3022275,"visible":true,"origin":"","legend":"","description":"","filename":"YOLOROC.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7221917/v1_covered_2603f076-947d-496e-884d-737b8d845c12.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"YOLO-ROC: A High-Precision and Ultra-Lightweight Model for Real-Time Road Damage Detection","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Road Damage Detection, YOLO, Attention Mechanism, Lightweight Model, Multi-scale Feature Extraction, Channel Compression","lastPublishedDoi":"10.21203/rs.3.rs-7221917/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7221917/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eRoad damage detection is a critical task for ensuring traffic safety and maintaining infrastructure integrity. While deep learning-based detection methods are now widely adopted, they still face two core challenges: first, the inadequate multi-scale feature extraction capabilities of existing networks for diverse targets like cracks and potholes, leading to high miss rates for small-scale damage; and second, the substantial parameter counts and computational demands of mainstream models, which hinder their deployment for efficient, real-time detection in practical applications. To address these issues, this paper proposes a high-precision and lightweight model, YOLO - Road Orthogonal Compact (YOLO-ROC). We designed a Bidirectional Multi-scale Spatial Pyramid Pooling Fast (BMS-SPPF) module to enhance multi-scale feature extraction and implemented a hierarchical channel compression strategy to reduce computational complexity. The BMS-SPPF module leverages a bidirectional spatialchannel attention mechanism to improve the detection of small targets. Concurrently, the channel compression strategy reduces the parameter count from 3.01M to 0.89M and GFLOPs from 8.1 to 2.6. Experiments on the RDD2022 China Drone dataset demonstrate that YOLO-ROC achieves a mAP50 of 67.6%, surpassing the baseline YOLOv8n by 2.11%. Notably, the mAP50 for the small-target D40 category improved by 16.8%, and the final model size is only 2.0 MB. Furthermore, the model exhibits excellent generalization performance on the RDD2022 China Motorbike dataset.\u003c/p\u003e","manuscriptTitle":"YOLO-ROC: A High-Precision and Ultra-Lightweight Model for Real-Time Road Damage Detection","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-07-30 04:45:01","doi":"10.21203/rs.3.rs-7221917/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"bea4a469-f136-43ea-bb49-e404822ee029","owner":[],"postedDate":"July 30th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-09-06T22:38:07+00:00","versionOfRecord":[],"versionCreatedAt":"2025-07-30 04:45:01","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7221917","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7221917","identity":"rs-7221917","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.