LMATDet:Lightweight Multi-scale Aerial TargetDetection with attention and efficient loss | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article LMATDet:Lightweight Multi-scale Aerial TargetDetection with attention and efficient loss Yuqing Zhao This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5335228/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Aerial image-based target detection presents challenges such as low accuracy in multiscale target detection, slow detection speed, missed detections, and false positives. To address these issues, this paper proposes a novel aerial target detection algorithm, LMATDet (Lightweight Multi-scale Aerial Target Detection with attention and efficient loss), specifically designed for multiscale object detection in aerial images. First, the introduction of the Box-Attention mechanism enables the network to effectively focus on relevant target features while reducing interference from irrelevant information, thus improving detection accuracy in complex scenarios. Second, the proposed algorithm incorporates a lightweight network architecture, LMATNet, utilizing DO-Conv in place of conventional convolutional layers. This substitution accelerates convergence, reduces network complexity,and enhances detection speed, with the model size reduced to just 36.45 MB. Additionally, the Focal-EIOU loss function, tailored for multiscale object detection, is employed for boundary regression. This approach not only mitigates overfitting but also significantly reduces both false positives and missed detections.The proposed method is validated using three benchmark aerial datasets:RSOD, NWPU VHR-10, and VisDrone. The achieved mean average precision(mAP) values are 94.98%, 95.43%, and 42.71%, respectively. The results demonstrate that LMATDet not only offers robust performance for multiscale target detection but also reduces missed detections while improving both robustness and generalization capability. Compared to other state-of-the-art algorithms,LMATDet exhibits significantly enhanced detection accuracy, thereby proving its superiority and versatility in aerial target detection. Pattern recognition artificial intelligence multiscale feature attention mechanism Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5335228","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":373366276,"identity":"69826f16-aab1-4be8-8a80-01ecc4e9ca36","order_by":0,"name":"Yuqing Zhao","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA00lEQVRIiWNgGAWjYBACNmaGxMc/Kmrq+9mbDxCnhY+94bExw5ljjDN7jiUQp0WO5+AzYcY2ZsYNN3IMiHSYRHIac2Eb0H1nzny88YbBTk63gaCWtLTHM87JsDG29262nMOQbGx2gKCWnHQDnjI2Hmaes9ukeRgOJG4jrCX/mwQPG7MEUO8zIrXwHEiT5mljNuCRyGEjUgt7Q7LhjDPHEiR4jhlbzjEgwi/yzQyJDz5U1CTYH29+eONNhZ0cQS0oQIKHyKhB1kKqjlEwCkbBKBgRAAAXG0HDhvdTXAAAAABJRU5ErkJggg==","orcid":"","institution":"Guangdong University of Technology","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Yuqing","middleName":"","lastName":"Zhao","suffix":""}],"badges":[],"createdAt":"2024-10-26 03:08:03","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5335228/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5335228/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":69478371,"identity":"5c35ce0a-490b-4384-bf51-176020d76fd1","added_by":"auto","created_at":"2024-11-20 20:31:44","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1492212,"visible":true,"origin":"","legend":"","description":"","filename":"LMATDet.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5335228/v1_covered_e196d4a9-f7ec-4490-bb07-94bd68d5e589.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"LMATDet:Lightweight Multi-scale Aerial TargetDetection with attention and efficient loss","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Pattern recognition, artificial intelligence, multiscale feature, attention mechanism","lastPublishedDoi":"10.21203/rs.3.rs-5335228/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5335228/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eAerial image-based target detection presents challenges such as low accuracy in multiscale target detection, slow detection speed, missed detections, and false positives. To address these issues, this paper proposes a novel aerial target detection algorithm, LMATDet (Lightweight Multi-scale Aerial Target Detection with attention and efficient loss), specifically designed for multiscale object detection in aerial images. First, the introduction of the Box-Attention mechanism enables the network to effectively focus on relevant target features while reducing interference from irrelevant information, thus improving detection accuracy in complex scenarios. Second, the proposed algorithm incorporates a lightweight network architecture, LMATNet, utilizing DO-Conv in place of conventional convolutional layers. This substitution accelerates convergence, reduces network complexity,and enhances detection speed, with the model size reduced to just 36.45 MB. Additionally, the Focal-EIOU loss function, tailored for multiscale object detection, is employed for boundary regression. This approach not only mitigates overfitting but also significantly reduces both false positives and missed detections.The proposed method is validated using three benchmark aerial datasets:RSOD, NWPU VHR-10, and VisDrone. The achieved mean average precision(mAP) values are 94.98%, 95.43%, and 42.71%, respectively. The results demonstrate that LMATDet not only offers robust performance for multiscale target detection but also reduces missed detections while improving both robustness and generalization capability. Compared to other state-of-the-art algorithms,LMATDet exhibits significantly enhanced detection accuracy, thereby proving its superiority and versatility in aerial target detection.\u003c/p\u003e","manuscriptTitle":"LMATDet:Lightweight Multi-scale Aerial TargetDetection with attention and efficient loss","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-11-07 16:08:14","doi":"10.21203/rs.3.rs-5335228/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"5777f6df-68fa-43bb-be00-60aec5515727","owner":[],"postedDate":"November 7th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-12-27T12:53:24+00:00","versionOfRecord":[],"versionCreatedAt":"2024-11-07 16:08:14","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5335228","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5335228","identity":"rs-5335228","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.