Adaptive Frequency Enhancement and Multi-Scale Semantic Interaction for Robust Cross-View Geo-Localization

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Cross-view geo-localization aims to match aerial and satellite images from dis- tinct viewpoints, facing challenges of geometric distortion and scale inconsistency. Existing methods focus on local features and lack robust modeling of con- text and frequency information. Here we present an adaptive framework for cross-view matching that integrates spectral enhancement and multi-scale seman- tic interaction. The method employs learnable frequency-domain attention to strengthen structural and detailed features, and uses multi-scale semantic guid- ance to improve feature alignment. We also propose an adaptive structured polyloss, which intelligently identifies hard examples and suppresses noise inter- ference. Experiments on University-1652 and SUES-200 datasets show that the proposed approach achieves 89.63% Recall@1 for drone-to-satellite matching and 92.87% Recall@1 for satellite-to-drone matching, outperforming state-of-the-art methods. This work provides a reliable solution for viewpoint-robust visual geo- localization and supports practical applications in drone navigation and remote sensing. Brief Explanation of Optimizations: The optimized abstract follows a clear logical flow, highlights quantitative results prominently, simplifies technical jargon, and emphasizes broader impact, which aligns with The Visual Computer’s preferred style and improves readability for a wide audience. Our code is publicly available at: https://github.com/WZJ952/AFMS-Net
Full text 12,410 characters · extracted from preprint-html · click to expand
Adaptive Frequency Enhancement and Multi-Scale Semantic Interaction for Robust Cross-View Geo-Localization | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Adaptive Frequency Enhancement and Multi-Scale Semantic Interaction for Robust Cross-View Geo-Localization He Xiao, Zhenju Wen, Qiuming Liu, Yu Du, Xing Le, Wenhao Zhong This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9537179/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Cross-view geo-localization aims to match aerial and satellite images from dis- tinct viewpoints, facing challenges of geometric distortion and scale inconsistency. Existing methods focus on local features and lack robust modeling of con- text and frequency information. Here we present an adaptive framework for cross-view matching that integrates spectral enhancement and multi-scale seman- tic interaction. The method employs learnable frequency-domain attention to strengthen structural and detailed features, and uses multi-scale semantic guid- ance to improve feature alignment. We also propose an adaptive structured polyloss, which intelligently identifies hard examples and suppresses noise inter- ference. Experiments on University-1652 and SUES-200 datasets show that the proposed approach achieves 89.63% Recall@1 for drone-to-satellite matching and 92.87% Recall@1 for satellite-to-drone matching, outperforming state-of-the-art methods. This work provides a reliable solution for viewpoint-robust visual geo- localization and supports practical applications in drone navigation and remote sensing. Brief Explanation of Optimizations: The optimized abstract follows a clear logical flow, highlights quantitative results prominently, simplifies technical jargon, and emphasizes broader impact, which aligns with The Visual Computer’s preferred style and improves readability for a wide audience. Our code is publicly available at: https://github.com/WZJ952/AFMS-Net Theoretical Computer Science Cross-View Geolocalization Drone Contextual Semantic Interaction Frequency Bands Full Text Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9537179","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":630059962,"identity":"1418e1c5-5961-4247-9c6f-cc732a92b1aa","order_by":0,"name":"He Xiao","email":"","orcid":"","institution":"Jiangxi University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"He","middleName":"","lastName":"Xiao","suffix":""},{"id":630059963,"identity":"2f4e33ea-6004-4fed-860f-12b9170eaa80","order_by":1,"name":"Zhenju Wen","email":"","orcid":"","institution":"Jiangxi University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Zhenju","middleName":"","lastName":"Wen","suffix":""},{"id":630059964,"identity":"59c2fee1-1ed1-4f8a-bae9-057b0b34f836","order_by":2,"name":"Qiuming Liu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA8ElEQVRIiWNgGAWjYBACPmYgkcDAwMPAwHzAACpogEcDAwMbQgtbAkipBGEtCCYPWCURWth5DG88bNsmY86/5kMxb1tdHQN78zYJhpo7eBzGY2yR2Habx3LG2w2GM9sOSzDwHCuTYDj2DJ8WMwmQFoMbZzcYfGw7IMEgkWMmwdhwmBgtZx4YJLbVSTDIvyFWy/keBqAtzEBbeAhpYSu2SDgHsoXNwHDGucOSbTxpQJFjuLXw8x/eePNH2W17g/OHnxnzlNXx87Mf3njjQw1uLSAgASET2MDxAY6pBLwaYFr4DzA/IKBwFIyCUTAKRigAABJES36LODBdAAAAAElFTkSuQmCC","orcid":"","institution":"Jiangxi University of Science and Technology","correspondingAuthor":true,"prefix":"","firstName":"Qiuming","middleName":"","lastName":"Liu","suffix":""},{"id":630059965,"identity":"5eacfe16-798e-4c10-86a4-eb56bc663e95","order_by":3,"name":"Yu Du","email":"","orcid":"","institution":"Jiangxi University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Yu","middleName":"","lastName":"Du","suffix":""},{"id":630059966,"identity":"da1ae392-29aa-45a7-b8bc-5e00e5344466","order_by":4,"name":"Xing Le","email":"","orcid":"","institution":"Jiangxi University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Xing","middleName":"","lastName":"Le","suffix":""},{"id":630059967,"identity":"dd4326c8-91c8-4231-8957-df6402d7a8df","order_by":5,"name":"Wenhao Zhong","email":"","orcid":"","institution":"Jiangxi University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Wenhao","middleName":"","lastName":"Zhong","suffix":""}],"badges":[],"createdAt":"2026-04-27 06:25:21","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-9537179/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9537179/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":108181529,"identity":"5bac3428-40b7-4bde-a499-131a63d0b942","added_by":"auto","created_at":"2026-04-30 08:58:44","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1206273,"visible":true,"origin":"","legend":"","description":"","filename":"snarticle.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9537179/v1_covered_c229900c-89f9-43a7-bb09-c3e08e82f82a.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eAdaptive Frequency Enhancement and Multi-Scale Semantic Interaction for Robust Cross-View Geo-Localization\u003c/p\u003e","fulltext":[],"fulltextSource":"","fullText":"","funders":[{"identity":"ad2d1dcd-4750-46eb-9ab6-84888e47260a","identifier":"10.13039/501100004479","name":"Natural Science Foundation of Jiangxi Province","awardNumber":"20242BAB25073","order_by":0}],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Jiangxi University of Science and Technology","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Cross-View Geolocalization, Drone, Contextual Semantic Interaction, Frequency Bands","lastPublishedDoi":"10.21203/rs.3.rs-9537179/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9537179/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eCross-view geo-localization aims to match aerial and satellite images from dis- tinct viewpoints, facing challenges of geometric distortion and scale inconsistency. Existing methods focus on local features and lack robust modeling of con- text and frequency information. Here we present an adaptive framework for cross-view matching that integrates spectral enhancement and multi-scale seman- tic interaction. The method employs learnable frequency-domain attention to strengthen structural and detailed features, and uses multi-scale semantic guid- ance to improve feature alignment. We also propose an adaptive structured polyloss, which intelligently identifies hard examples and suppresses noise inter- ference. Experiments on University-1652 and SUES-200 datasets show that the proposed approach achieves 89.63% Recall@1 for drone-to-satellite matching and 92.87% Recall@1 for satellite-to-drone matching, outperforming state-of-the-art methods. This work provides a reliable solution for viewpoint-robust visual geo- localization and supports practical applications in drone navigation and remote sensing. Brief Explanation of Optimizations: The optimized abstract follows a clear logical flow, highlights quantitative results prominently, simplifies technical jargon, and emphasizes broader impact, which aligns with The Visual Computer\u0026rsquo;s preferred style and improves readability for a wide audience. Our code is publicly available at: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/WZJ952/AFMS-Net\u003c/span\u003e\u003cspan address=\"https://github.com/WZJ952/AFMS-Net\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e","manuscriptTitle":"Adaptive Frequency Enhancement and Multi-Scale Semantic Interaction for Robust Cross-View Geo-Localization","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-28 10:01:47","doi":"10.21203/rs.3.rs-9537179/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"940142f7-2d45-4444-bb3a-f6a5fd7fd65d","owner":[],"postedDate":"April 28th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":67055388,"name":"Theoretical Computer Science"}],"tags":[],"updatedAt":"2026-04-28T10:01:48+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-28 10:01:47","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9537179","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9537179","identity":"rs-9537179","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00