UAPT: An Underwater Acoustic Target Recognition Method Based on Pre- trained Transformer | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article UAPT: An Underwater Acoustic Target Recognition Method Based on Pre- trained Transformer Jun Tang, Enxue Ma, Yang Qu, Wenbo Gao, Yuchen zhang, Lin Gan This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4253542/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 07 Jan, 2025 Read the published version in Multimedia Systems → Version 1 posted 11 You are reading this latest preprint version Abstract The Convolutional Neural Network (CNN) model in underwater acoustic target recognition (UATR) research reveals limitations arising from its inability to capture long-distance dependencies, impeding its capacity to focus on global information within the underwater acoustic signal. In contrast, the Transformer model has progressively emerged as the optimal choice in various studies, owing to its exclusive dependence on the attention mechanism for extracting global features from input data. Limited research utilizing the Transformer model in UATR has relied on an early ViT model, while in this paper, two refined Transformer models, namely Swin Transformer and Biformer, are adopted as the foundational networks, and a novel Swin Biformer model is proposed by harnessing the strengths of the two. Experimental results demonstrate the consistent superiority of the three models over CNN and ViT in UATR, and the Swin Biformer model remarkably attains the highest recognition accuracy of 94.3% evaluated on a dataset con-structed from the Deepship database. At the same time, this paper proposes an UATR method based on pre-trained Transformer, the effectiveness of which is underscored by experimental findings as a recognition accuracy of approximately 97% was achieved on a generalized dataset derived from the Shipsear database. Even with limited data samples and more stringent classification requirements, the method maintains a recognition accuracy of over 90%, all while significantly reducing the training duration. Underwater acoustic target recognition Transformer Transfer learning Deep learning Pre-trained Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 07 Jan, 2025 Read the published version in Multimedia Systems → Version 1 posted Editorial decision: Revision requested 14 Jul, 2024 Reviews received at journal 10 Jul, 2024 Reviewers agreed at journal 20 Jun, 2024 Reviewers agreed at journal 28 May, 2024 Reviews received at journal 26 May, 2024 Reviewers agreed at journal 03 May, 2024 Reviewers agreed at journal 03 May, 2024 Reviewers invited by journal 03 May, 2024 Editor assigned by journal 02 May, 2024 Submission checks completed at journal 13 Apr, 2024 First submitted to journal 11 Apr, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4253542","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":290801998,"identity":"1ea77305-d081-4e3c-a7ea-9cd3d860f28d","order_by":0,"name":"Jun Tang","email":"","orcid":"","institution":"Tianjin University","correspondingAuthor":false,"prefix":"","firstName":"Jun","middleName":"","lastName":"Tang","suffix":""},{"id":290801999,"identity":"63d25510-918b-4cec-804b-d1d795aca55b","order_by":1,"name":"Enxue Ma","email":"","orcid":"","institution":"Tianjin University","correspondingAuthor":false,"prefix":"","firstName":"Enxue","middleName":"","lastName":"Ma","suffix":""},{"id":290802000,"identity":"59dd1891-0f66-482f-91b3-54d9b8e05e4b","order_by":2,"name":"Yang Qu","email":"","orcid":"","institution":"Tianjin University","correspondingAuthor":false,"prefix":"","firstName":"Yang","middleName":"","lastName":"Qu","suffix":""},{"id":290802001,"identity":"84accd2e-d48a-49f6-885f-4ec3ab84ff50","order_by":3,"name":"Wenbo Gao","email":"","orcid":"","institution":"Tianjin University","correspondingAuthor":false,"prefix":"","firstName":"Wenbo","middleName":"","lastName":"Gao","suffix":""},{"id":290802002,"identity":"dd09b28a-7775-4c9d-97fb-5c68c09595dd","order_by":4,"name":"Yuchen zhang","email":"","orcid":"","institution":"Tianjin University","correspondingAuthor":false,"prefix":"","firstName":"Yuchen","middleName":"","lastName":"zhang","suffix":""},{"id":290802003,"identity":"87f493b8-1c5a-44ef-bf3b-17b96522ac6e","order_by":5,"name":"Lin Gan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAxklEQVRIiWNgGAWjYHACxgMJDDZQNhuReoBa0kjVwsBwmAQt8u5nDxx4uON8nsHtHgOGD2WHGfhnN+DXYngmL+FA4pnbxQZ3zhgwzjh3mEHizgECWhpyDA4ktt1O3HYjx4CZt+0wg4FEAgEt/W9AWs5BtPwlRou8BNiWAxAtjMRoMZAA25KcuP/OsYKDPefSeSRuELKlP8fw4c82u8SZs5s3PvhRZi3HP4OQLQdgLAlwBDHw4FcPsqUBScsoGAWjYBSMAqwAAL0fS1NUxzz9AAAAAElFTkSuQmCC","orcid":"","institution":"Northwestern Polytechnical University","correspondingAuthor":true,"prefix":"","firstName":"Lin","middleName":"","lastName":"Gan","suffix":""}],"badges":[],"createdAt":"2024-04-11 16:12:50","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4253542/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4253542/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s00530-024-01614-3","type":"published","date":"2025-01-07T15:57:45+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":73694242,"identity":"b9f70d53-b1df-46ba-b8cc-1a0a4357448e","added_by":"auto","created_at":"2025-01-13 16:12:47","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":755147,"visible":true,"origin":"","legend":"","description":"","filename":"Manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4253542/v1_covered_d7837582-08d7-4530-b0a7-f8742e6867df.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"UAPT: An Underwater Acoustic Target Recognition Method Based on Pre- trained Transformer","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"multimedia-systems","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mmsj","sideBox":"Learn more about [Multimedia Systems](http://link.springer.com/journal/530)","snPcode":"530","submissionUrl":"https://submission.nature.com/new-submission/530/3","title":"Multimedia Systems","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Underwater acoustic target recognition, Transformer, Transfer learning, Deep learning, Pre-trained","lastPublishedDoi":"10.21203/rs.3.rs-4253542/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4253542/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe Convolutional Neural Network (CNN) model in underwater acoustic target recognition (UATR) research reveals limitations arising from its inability to capture long-distance dependencies, impeding its capacity to focus on global information within the underwater acoustic signal. In contrast, the Transformer model has progressively emerged as the optimal choice in various studies, owing to its exclusive dependence on the attention mechanism for extracting global features from input data. Limited research utilizing the Transformer model in UATR has relied on an early ViT model, while in this paper, two refined Transformer models, namely Swin Transformer and Biformer, are adopted as the foundational networks, and a novel Swin Biformer model is proposed by harnessing the strengths of the two. Experimental results demonstrate the consistent superiority of the three models over CNN and ViT in UATR, and the Swin Biformer model remarkably attains the highest recognition accuracy of 94.3% evaluated on a dataset con-structed from the Deepship database. At the same time, this paper proposes an UATR method based on pre-trained Transformer, the effectiveness of which is underscored by experimental findings as a recognition accuracy of approximately 97% was achieved on a generalized dataset derived from the Shipsear database. Even with limited data samples and more stringent classification requirements, the method maintains a recognition accuracy of over 90%, all while significantly reducing the training duration.\u003c/p\u003e","manuscriptTitle":"UAPT: An Underwater Acoustic Target Recognition Method Based on Pre- trained Transformer","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-04-17 16:42:36","doi":"10.21203/rs.3.rs-4253542/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-07-14T05:34:10+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-07-10T13:10:30+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"315042089799599176235150124982823680260","date":"2024-06-20T06:56:17+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"23775593653427427604729473876180147678","date":"2024-05-28T14:53:20+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-05-26T10:26:26+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"312259341908292864868024439285316658034","date":"2024-05-04T01:15:38+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"155151773987778857661111049636550848906","date":"2024-05-03T20:09:38+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-05-03T15:53:40+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-05-02T04:47:45+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-04-13T12:03:15+00:00","index":"","fulltext":""},{"type":"submitted","content":"Multimedia Systems","date":"2024-04-11T16:11:35+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"multimedia-systems","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"mmsj","sideBox":"Learn more about [Multimedia Systems](http://link.springer.com/journal/530)","snPcode":"530","submissionUrl":"https://submission.nature.com/new-submission/530/3","title":"Multimedia Systems","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"eb45d85f-4b80-4ca5-b769-f6d78335eda3","owner":[],"postedDate":"April 17th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-01-13T16:07:14+00:00","versionOfRecord":{"articleIdentity":"rs-4253542","link":"https://doi.org/10.1007/s00530-024-01614-3","journal":{"identity":"multimedia-systems","isVorOnly":false,"title":"Multimedia Systems"},"publishedOn":"2025-01-07 15:57:45","publishedOnDateReadable":"January 7th, 2025"},"versionCreatedAt":"2024-04-17 16:42:36","video":"","vorDoi":"10.1007/s00530-024-01614-3","vorDoiUrl":"https://doi.org/10.1007/s00530-024-01614-3","workflowStages":[]},"version":"v1","identity":"rs-4253542","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4253542","identity":"rs-4253542","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.