QLDNet: A Lightweight and Efficient Network for High-Robustness Aerial Human Detection in UAV-Based Remote Sensing | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article QLDNet: A Lightweight and Efficient Network for High-Robustness Aerial Human Detection in UAV-Based Remote Sensing Mandun Zhang, Jiale Shi, Chenyue Hou, Yonghui Pang, Yiqi Zhang, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6036149/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 29 Aug, 2025 Read the published version in Journal of Real-Time Image Processing → Version 1 posted 10 You are reading this latest preprint version Abstract Accurate detection of extremely small targets in aerial high-resolution images is critical for military and civilian applications. However, challenges such as small target size, complex backgrounds, and target deformation hinder optimal performance. We propose QLDNet, a lightweight network that addresses these issues through a synergistic combination of non-strided convolution and decoupled large-kernel convolutional attention. Specifically, the non-strided convolution constructs a quadruple aggregate and connection feature extractor (QACFE) module, which maps features into the channel dimension while preserving spatial details. The decoupled large kernel convolutional attention then leverages these channel features to effectively extract structural and edge-related low-frequency information, while simultaneously reducing computational costs and model size caused by the increased channel dimensionality after QACFE. A learnable offset mechanism is introduced, transforming the detection head into a deformable detection head. Additionally, the network incorporates a PAFPN (Path Aggregation Feature Pyramid Network) structure to efficiently extract multi-scale features. During inference, a tailored approach performs pixel-level multi-scale detection. This enhances small target detection by merging features and pixels through dual-scale fusion, integrating multi-scale feature extraction with multilevel feature integration.Experimental results demonstrate that QLDNet achieves efficient and accurate small target detection, with an accuracy rate as high as 94.0\% and an inference time of 0.12 seconds per image, meeting the system's real-time requirements with fewer parameters and satisfying processing speeds.The code can be find at https://github.com/Sjl185721/QLDNet-v1.git. Mathematics Subject Classification (2020) MSC68T05 · MSC 86Axx · MSC 68T10 Object Detection Rescue Drone perspective Computer vision Remote Sensing Full Text Additional Declarations No competing interests reported. Supplementary Files LaTeXver2.0.zip Cite Share Download PDF Status: Published Journal Publication published 29 Aug, 2025 Read the published version in Journal of Real-Time Image Processing → Version 1 posted Editorial decision: Revision requested 28 Apr, 2025 Reviews received at journal 24 Apr, 2025 Reviews received at journal 22 Apr, 2025 Reviews received at journal 10 Apr, 2025 Reviewers agreed at journal 07 Apr, 2025 Reviewers agreed at journal 07 Apr, 2025 Reviewers agreed at journal 07 Apr, 2025 Reviewers invited by journal 07 Apr, 2025 Submission checks completed at journal 20 Mar, 2025 First submitted to journal 18 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6036149","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":447412186,"identity":"71bd6df7-5f24-403c-8502-ced8ff931fb1","order_by":0,"name":"Mandun Zhang","email":"","orcid":"","institution":"Hebei University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Mandun","middleName":"","lastName":"Zhang","suffix":""},{"id":447412187,"identity":"549d8da9-8182-4f8a-92c3-382416d0368c","order_by":1,"name":"Jiale Shi","email":"","orcid":"","institution":"Hebei University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Jiale","middleName":"","lastName":"Shi","suffix":""},{"id":447412188,"identity":"08e07b5b-b2b0-4b16-a1f5-c39629e806f0","order_by":2,"name":"Chenyue Hou","email":"","orcid":"","institution":"Hebei University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Chenyue","middleName":"","lastName":"Hou","suffix":""},{"id":447412189,"identity":"b96dd594-f046-409b-b9a4-5dd9dd52dccb","order_by":3,"name":"Yonghui Pang","email":"","orcid":"","institution":"Hebei University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Yonghui","middleName":"","lastName":"Pang","suffix":""},{"id":447412192,"identity":"84cdb1cd-7337-4cf8-b8a4-6940164305a6","order_by":4,"name":"Yiqi Zhang","email":"","orcid":"","institution":"Hebei University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Yiqi","middleName":"","lastName":"Zhang","suffix":""},{"id":447412195,"identity":"f9ed6655-82f1-4436-9207-41547c3f1026","order_by":5,"name":"Xiangsheng Huang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABE0lEQVRIiWNgGAWjYBACAwglASLYQIQcAzNjA2lajInVwgDXkohfPRCYSyQ/e/i1zULenH/5swcfd9SmbzjO3MDwo4ZB3hyHFssZaebGMmckDHfOeGNuOPPM8dyZzYwNjD3HGAx34rDP4EaCmbREhQTjhhtn2KR5247l9oP8wtvAkGBwAJeW9G/SEgYS9htuHH8m/bftWDobUAvjX7xacswkP1RIJG4432AmzdhWk8AP1MKM15Yzb8qkGc5IJG+4wWMm2dt2wBDkl8MyxyQMN+DScjx9m+TPtjrbDeePP5MAMuQNzh9/+PBNjY08LltAgJkHREokgMjDYJED0MjFCRh/gEh+sKF1eFWOglEwCkbByAQA9CJdoEUU3/kAAAAASUVORK5CYII=","orcid":"","institution":"Hebei Key Laboratory of Cognitive Intelligence,Xiong’an Institute of Innovation,Xiong’an New Area","correspondingAuthor":true,"prefix":"","firstName":"Xiangsheng","middleName":"","lastName":"Huang","suffix":""}],"badges":[],"createdAt":"2025-02-15 11:08:20","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6036149/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6036149/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s11554-025-01707-w","type":"published","date":"2025-08-29T15:57:47+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":90345001,"identity":"c3ea3d11-5521-4833-b765-be8a08a7c997","added_by":"auto","created_at":"2025-09-01 16:09:06","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2591019,"visible":true,"origin":"","legend":"","description":"","filename":"LaTeXver2.0.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6036149/v1_covered_1a36d17d-27a0-47e9-b180-cf48b3c460d1.pdf"},{"id":81531124,"identity":"51517880-bd2f-4ef4-97ad-78582abf78b9","added_by":"auto","created_at":"2025-04-28 09:29:04","extension":"zip","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":43306541,"visible":true,"origin":"","legend":"","description":"","filename":"LaTeXver2.0.zip","url":"https://assets-eu.researchsquare.com/files/rs-6036149/v1/a4dabbe60b736ca4ed3da04c.zip"}],"financialInterests":"No competing interests reported.","formattedTitle":"QLDNet: A Lightweight and Efficient Network for High-Robustness Aerial Human Detection in UAV-Based Remote Sensing","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"journal-of-real-time-image-processing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"rtip","sideBox":"Learn more about [Journal of Real-Time Image Processing](http://link.springer.com/journal/11554)","snPcode":"11554","submissionUrl":"https://submission.nature.com/new-submission/11554/3","title":"Journal of Real-Time Image Processing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Object Detection, Rescue, Drone perspective, Computer vision, Remote Sensing","lastPublishedDoi":"10.21203/rs.3.rs-6036149/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6036149/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eAccurate detection of extremely small targets in aerial high-resolution images is critical for military and civilian applications. However, challenges such as small target size, complex backgrounds, and target deformation hinder optimal performance. We propose QLDNet, a lightweight network that addresses these issues through a synergistic combination of non-strided convolution and decoupled large-kernel convolutional attention. Specifically, the non-strided convolution constructs a quadruple aggregate and connection feature extractor (QACFE) module, which maps features into the channel dimension while preserving spatial details. The decoupled large kernel convolutional attention then leverages these channel features to effectively extract structural and edge-related low-frequency information, while simultaneously reducing computational costs and model size caused by the increased channel dimensionality after QACFE. A learnable offset mechanism is introduced, transforming the detection head into a deformable detection head. Additionally, the network incorporates a PAFPN (Path Aggregation Feature Pyramid Network) structure to efficiently extract multi-scale features. During inference, a tailored approach performs pixel-level multi-scale detection. This enhances small target detection by merging features and pixels through dual-scale fusion, integrating multi-scale feature extraction with multilevel feature integration.Experimental results demonstrate that QLDNet achieves efficient and accurate small target detection, with an accuracy rate as high as 94.0\\% and an inference time of 0.12 seconds per image, meeting the system's real-time requirements with fewer parameters and satisfying processing speeds.The code can be find at https://github.com/Sjl185721/QLDNet-v1.git.\u003c/p\u003e\n\u003cp\u003eMathematics Subject Classification (2020) MSC68T05 · MSC 86Axx · MSC 68T10\u003c/p\u003e","manuscriptTitle":"QLDNet: A Lightweight and Efficient Network for High-Robustness Aerial Human Detection in UAV-Based Remote Sensing","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-28 09:28:55","doi":"10.21203/rs.3.rs-6036149/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-04-28T15:16:33+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-24T10:25:42+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-22T11:36:03+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-10T18:54:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"339616126763977004992643261427729266395","date":"2025-04-07T21:03:21+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"140922624758201485770258101696209160159","date":"2025-04-07T06:40:13+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"84227428429902111965043881622735209128","date":"2025-04-07T06:35:05+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-04-07T06:27:32+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-03-20T05:55:08+00:00","index":"","fulltext":""},{"type":"submitted","content":"Journal of Real-Time Image Processing","date":"2025-03-18T12:28:06+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"journal-of-real-time-image-processing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"rtip","sideBox":"Learn more about [Journal of Real-Time Image Processing](http://link.springer.com/journal/11554)","snPcode":"11554","submissionUrl":"https://submission.nature.com/new-submission/11554/3","title":"Journal of Real-Time Image Processing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"1c56d8f1-d8e2-45ea-adac-761be4b15e52","owner":[],"postedDate":"April 28th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-09-01T16:05:28+00:00","versionOfRecord":{"articleIdentity":"rs-6036149","link":"https://doi.org/10.1007/s11554-025-01707-w","journal":{"identity":"journal-of-real-time-image-processing","isVorOnly":false,"title":"Journal of Real-Time Image Processing"},"publishedOn":"2025-08-29 15:57:47","publishedOnDateReadable":"August 29th, 2025"},"versionCreatedAt":"2025-04-28 09:28:55","video":"","vorDoi":"10.1007/s11554-025-01707-w","vorDoiUrl":"https://doi.org/10.1007/s11554-025-01707-w","workflowStages":[]},"version":"v1","identity":"rs-6036149","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6036149","identity":"rs-6036149","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.