Audio manipulation detection based on wavelet spectrogram and multidimensional feature fusion with dual-channel CNN | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Audio manipulation detection based on wavelet spectrogram and multidimensional feature fusion with dual-channel CNN Dongyu Wang, Canghong Shi, Xiaojie Li, Kai Peng, M. Abdullahi Sani This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6242466/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract With the prevalence of editing software, audio manipulation has become increasingly easy, posing a serious threat to the authenticity and integrity of audio data. Among various forms of manipulation, audio splicing and copy-move forgery are common. However, current audio forgery detection technologies typically can only detect one form of tampering, which presents significant limitations in practical applications. To address this issue, we propose a novel approach based on wavelet spectrograms and dual-channel convolutional neural network (DCCNN) for detecting audio splicing and copy-move forgery simultaneously. Initially, we transform the audio signals into wavelet spectrograms and then employ DCCNN for feature extraction and classification. Compared with the traditional convolutional neural network, this algorithm adopts a dual-channel convolutional structure combined with a multi-scale feature fusion technique, which is able to extract the local and global features of the wavelet spectrogram more effectively. By using convolution kernels of different sizes to perform convolution operations on the feature map and fusing these features, the feature expression ability of the image is enriched. Experiment result show that this method can simultaneously detect traces of splicing and copy-move forgery, and exhibits higher detection accuracy and robustness in audio forgery detection tasks. The algorithm was tested on dataset spliced and copy-move forged samples created from the Arabic speech corpus, TIMIT databases and ADD databsed, respectively. Subsequently, compared to existing state-of-the-art algorithms for audio splicing and copy-move forgery detection, the experimental results show that the proposed algorithm performs excellently in accuracy and significantly state-of-the-art methods. Audio copy-move Audio splcing The wavelet spectrogram Robustness Dual-channel convolutional neural network Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 22 Mar, 2025 Reviewers invited by journal 22 Mar, 2025 Editor assigned by journal 18 Mar, 2025 Submission checks completed at journal 18 Mar, 2025 First submitted to journal 17 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6242466","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":432580105,"identity":"e0fa17db-f045-414b-98bc-88cf7a2174a6","order_by":0,"name":"Dongyu Wang","email":"","orcid":"","institution":"Xihua University","correspondingAuthor":false,"prefix":"","firstName":"Dongyu","middleName":"","lastName":"Wang","suffix":""},{"id":432580106,"identity":"9fe7465e-6a12-46a6-ad39-ac4d2e30e0a3","order_by":1,"name":"Canghong Shi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA2klEQVRIiWNgGAWjYJACZjApwdj4IKGihjQtzQYPzhwjSQsDm+TDFmbCyvnbzx5+XVBxx27+7Oa2isQGNqBIdwJeLRJn8tKsZ5x5lrzhzsG2G4k7ZIAiZzfg1WLAkGNmzNt2ONlAIhGo5Qwbg4FELgEt/G8gWuRnJLYVJLYxE6FFIsf4MVCLHcONxDYGorRI3Hhjxsxz5nCCwY3EZomEM8d4CPqFvz/H+DNPxWF7+RnpDz/+qKiR42/vxa8FCNgkgERiA5THQ0g5CDB/ABL2xKgcBaNgFIyCEQoAjEtLcnz+eXYAAAAASUVORK5CYII=","orcid":"","institution":"Xihua University","correspondingAuthor":true,"prefix":"","firstName":"Canghong","middleName":"","lastName":"Shi","suffix":""},{"id":432580107,"identity":"b935b9ce-80ea-4ffa-b020-8449207ec61d","order_by":2,"name":"Xiaojie Li","email":"","orcid":"","institution":"Chengdu University of Information Technology","correspondingAuthor":false,"prefix":"","firstName":"Xiaojie","middleName":"","lastName":"Li","suffix":""},{"id":432580108,"identity":"a6381260-5441-4e28-9d4e-6695f7e0e6d0","order_by":3,"name":"Kai Peng","email":"","orcid":"","institution":"Xihua University","correspondingAuthor":false,"prefix":"","firstName":"Kai","middleName":"","lastName":"Peng","suffix":""},{"id":432580109,"identity":"b2a0e5e8-756a-43e9-9c9f-13f0a173da8c","order_by":4,"name":"M. Abdullahi Sani","email":"","orcid":"","institution":"University of Southern Denmark","correspondingAuthor":false,"prefix":"","firstName":"M.","middleName":"Abdullahi","lastName":"Sani","suffix":""}],"badges":[],"createdAt":"2025-03-17 08:23:22","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6242466/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6242466/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":80687804,"identity":"e8371da2-c204-4837-97fe-d7b744dd0368","added_by":"auto","created_at":"2025-04-16 04:17:26","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3509029,"visible":true,"origin":"","legend":"","description":"","filename":"ArticleTitle81.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6242466/v1_covered_f4ce2530-a6cb-4b52-b6e3-627bf2f7a6e6.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Audio manipulation detection based on wavelet spectrogram and multidimensional feature fusion with dual-channel CNN","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"signal-image-and-video-processing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"sivp","sideBox":"Learn more about [Signal, Image and Video Processing](http://link.springer.com/journal/11760)","snPcode":"11760","submissionUrl":"https://submission.nature.com/new-submission/11760/3","title":"Signal, Image and Video Processing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Audio copy-move, Audio splcing, The wavelet spectrogram, Robustness, Dual-channel convolutional neural network","lastPublishedDoi":"10.21203/rs.3.rs-6242466/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6242466/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"With the prevalence of editing software, audio manipulation has become increasingly easy, posing a serious threat to the authenticity and integrity of audio data. Among various forms of manipulation, audio splicing and copy-move forgery are common. However, current audio forgery detection technologies typically can only detect one form of tampering, which presents significant limitations in practical applications. To address this issue, we propose a novel approach based on wavelet spectrograms and dual-channel convolutional neural network (DCCNN) for detecting audio splicing and copy-move forgery simultaneously. Initially, we transform the audio signals into wavelet spectrograms and then employ DCCNN for feature extraction and classification. Compared with the traditional convolutional neural network, this algorithm adopts a dual-channel convolutional structure combined with a multi-scale feature fusion technique, which is able to extract the local and global features of the wavelet spectrogram more effectively. By using convolution kernels of different sizes to perform convolution operations on the feature map and fusing these features, the feature expression ability of the image is enriched. Experiment result show that this method can simultaneously detect traces of splicing and copy-move forgery, and exhibits higher detection accuracy and robustness in audio forgery detection tasks. The algorithm was tested on dataset spliced and copy-move forged samples created from the Arabic speech corpus, TIMIT databases and ADD databsed, respectively. Subsequently, compared to existing state-of-the-art algorithms for audio splicing and copy-move forgery detection, the experimental results show that the proposed algorithm performs excellently in accuracy and significantly state-of-the-art methods.","manuscriptTitle":"Audio manipulation detection based on wavelet spectrogram and multidimensional feature fusion with dual-channel CNN","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-16 04:09:13","doi":"10.21203/rs.3.rs-6242466/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-03-23T00:14:40+00:00","index":"","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-03-23T00:14:03+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-03-18T08:22:00+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-03-18T08:20:46+00:00","index":"","fulltext":""},{"type":"submitted","content":"Signal, Image and Video Processing","date":"2025-03-17T08:12:21+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"signal-image-and-video-processing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"sivp","sideBox":"Learn more about [Signal, Image and Video Processing](http://link.springer.com/journal/11760)","snPcode":"11760","submissionUrl":"https://submission.nature.com/new-submission/11760/3","title":"Signal, Image and Video Processing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"a0e6c912-003c-49e9-8bf6-0a134274f291","owner":[],"postedDate":"April 16th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2025-04-16T04:09:13+00:00","versionOfRecord":[],"versionCreatedAt":"2025-04-16 04:09:13","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6242466","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6242466","identity":"rs-6242466","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.