Deepfake audio detection and justication with Explainable Articial Intelligence 

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Deepfake audio refers to synthetically generated audio, often used as legal hoaxes to impersonate human voices. This paper generates fake audio from Fake or Real (FoR) dataset using Generative Adversarial Neural Networks (GANs). FoR dataset has the advantage of a diversity of speakers across 195,000 samples. The proposed work analyses the quality of the generated fake data using the Fréchet Audio Distance (FAD) score. FAD evaluation score of 23.814 indicates good quality fake has been produced by the generator. The study further enables glass box analysis of deepfake audio detection through Explainable Artificial Intelligence (XAI) models of LIME, SHAP and GradCAM. This research assists in understanding impact of frequency bands in audio classification based on the quantitative analysis of SHAPLey values and qualitative comparison of explainability masks of LIME and GradCAM. The use of FAD metric provides a quantitative evaluation of generator performance. XAI and FAD metrics help in the development of deepfake audio through GANs with minimal data input. The results of this research are applicable to detection of phishing audio calls and voice impersonation.
Full text 12,888 characters · extracted from preprint-html · click to expand
Deepfake audio detection and justication with Explainable Articial Intelligence | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Deepfake audio detection and justication with Explainable Articial Intelligence Aditi Govindu, Preeti Kale, Aamir Hullur, Atharva Gurav, Parth Godse This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6349440/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 10 You are reading this latest preprint version Abstract Deepfake audio refers to synthetically generated audio, often used as legal hoaxes to impersonate human voices. This paper generates fake audio from Fake or Real (FoR) dataset using Generative Adversarial Neural Networks (GANs). FoR dataset has the advantage of a diversity of speakers across 195,000 samples. The proposed work analyses the quality of the generated fake data using the Fréchet Audio Distance (FAD) score. FAD evaluation score of 23.814 indicates good quality fake has been produced by the generator. The study further enables glass box analysis of deepfake audio detection through Explainable Artificial Intelligence (XAI) models of LIME, SHAP and GradCAM. This research assists in understanding impact of frequency bands in audio classification based on the quantitative analysis of SHAPLey values and qualitative comparison of explainability masks of LIME and GradCAM. The use of FAD metric provides a quantitative evaluation of generator performance. XAI and FAD metrics help in the development of deepfake audio through GANs with minimal data input. The results of this research are applicable to detection of phishing audio calls and voice impersonation. Generative Adversarial Neural Networks (GANs) deepfake audio VGG16 Explainable Artificial Intelligence (XAI) Fréchet Audio Distance (FAD). Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 20 Dec, 2025 Reviews received at journal 26 Sep, 2025 Reviewers agreed at journal 08 Aug, 2025 Reviewers agreed at journal 07 Aug, 2025 Reviews received at journal 30 Jun, 2025 Reviewers agreed at journal 08 May, 2025 Reviewers invited by journal 06 May, 2025 Editor assigned by journal 23 Apr, 2025 Submission checks completed at journal 08 Apr, 2025 First submitted to journal 01 Apr, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6349440","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":453873832,"identity":"73cdeea8-d6ba-4954-bf2d-4ec394c79f57","order_by":0,"name":"Aditi Govindu","email":"","orcid":"","institution":"Dr. Vishwanath Karad MIT World Peace University","correspondingAuthor":false,"prefix":"","firstName":"Aditi","middleName":"","lastName":"Govindu","suffix":""},{"id":453873833,"identity":"4e0b6d72-cb8a-4455-8229-26c6e3e05c2b","order_by":1,"name":"Preeti Kale","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABAElEQVRIiWNgGAWjYNACAyBkZ2B8AKQggAeM8IADIC3MDMxAyoBYLQxgLWwSDAwILTiBufThg48/FDAY8zczH6u6UfAncbtEAuODt20MMuY4tFj2pSUbAB1mJnGYLe12joFB4s4ZCcyGc9sYeCwbsGsxOMNjJgHUYsNwmMcMpCV3w40ENmleoBaDA7i08H//AdIif5j/WzFUC/tv/Fp42EAhZmZwmIeNGWYLMz4tlj1sxhJnDCSMDQ+zGUvnGBjX7+x52Cw555wETi3mPMwPP1T8sTGcd7z54eecP3LG5uzJBz+8KbOxx+kwCCWBLMLYgCqCXQtekVEwCkbBKBjpAAATYlLUGQX+KQAAAABJRU5ErkJggg==","orcid":"","institution":"Dr. Vishwanath Karad MIT World Peace University","correspondingAuthor":true,"prefix":"","firstName":"Preeti","middleName":"","lastName":"Kale","suffix":""},{"id":453873834,"identity":"cf350fbc-f5c3-4c99-839d-c47b6867d44b","order_by":2,"name":"Aamir Hullur","email":"","orcid":"","institution":"Dr. Vishwanath Karad MIT World Peace University","correspondingAuthor":false,"prefix":"","firstName":"Aamir","middleName":"","lastName":"Hullur","suffix":""},{"id":453873835,"identity":"69e6e5da-e225-4654-b5e1-21d3a093f7ad","order_by":3,"name":"Atharva Gurav","email":"","orcid":"","institution":"Dr. Vishwanath Karad MIT World Peace University","correspondingAuthor":false,"prefix":"","firstName":"Atharva","middleName":"","lastName":"Gurav","suffix":""},{"id":453873836,"identity":"45a5ac1b-e450-4154-b1da-c04411a57240","order_by":4,"name":"Parth Godse","email":"","orcid":"","institution":"Dr. Vishwanath Karad MIT World Peace University","correspondingAuthor":false,"prefix":"","firstName":"Parth","middleName":"","lastName":"Godse","suffix":""}],"badges":[],"createdAt":"2025-04-01 04:53:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6349440/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6349440/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":82453395,"identity":"59b78fb2-644f-42a7-a80b-0a60e81462af","added_by":"auto","created_at":"2025-05-11 10:27:29","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":583004,"visible":true,"origin":"","legend":"","description":"","filename":"ManuscriptIJDS.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6349440/v1_covered_7f26b420-4fb3-47a9-b5cf-fefbb055df05.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Deepfake audio detection and justication with Explainable Articial Intelligence ","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"international-journal-of-data-science-and-analytics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jdsa","sideBox":"Learn more about [International Journal of Data Science and Analytics](http://link.springer.com/journal/41060)","snPcode":"41060","submissionUrl":"https://submission.nature.com/new-submission/41060/3","title":"International Journal of Data Science and Analytics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Generative Adversarial Neural Networks (GANs), deepfake audio, VGG16, Explainable Artificial Intelligence (XAI), Fréchet Audio Distance (FAD).","lastPublishedDoi":"10.21203/rs.3.rs-6349440/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6349440/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDeepfake audio refers to synthetically generated audio, often used as legal hoaxes to impersonate human voices. This paper generates fake audio from Fake or Real (FoR) dataset using Generative Adversarial Neural Networks (GANs). FoR dataset has the advantage of a diversity of speakers across 195,000 samples. The proposed work analyses the quality of the generated fake data using the Fr\u0026eacute;chet Audio Distance (FAD) score. FAD evaluation score of 23.814 indicates good quality fake has been produced by the generator. The study further enables glass box analysis of deepfake audio detection through Explainable Artificial Intelligence (XAI) models of LIME, SHAP and GradCAM. This research assists in understanding impact of frequency bands in audio classification based on the quantitative analysis of SHAPLey values and qualitative comparison of explainability masks of LIME and GradCAM. The use of FAD metric provides a quantitative evaluation of generator performance. XAI and FAD metrics help in the development of deepfake audio through GANs with minimal data input. The results of this research are applicable to detection of phishing audio calls and voice impersonation.\u003c/p\u003e","manuscriptTitle":"Deepfake audio detection and justication with Explainable Articial Intelligence ","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-05-11 09:55:24","doi":"10.21203/rs.3.rs-6349440/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-12-21T03:36:33+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-09-26T10:59:25+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"338512174081205354218150505887282053652","date":"2025-08-09T02:18:41+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"68346251356464885422439208901236562468","date":"2025-08-08T00:36:24+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-06-30T07:16:12+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"226245752043382194641378046521296065962","date":"2025-05-08T14:09:16+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-05-06T13:35:06+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-04-23T16:36:12+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-04-08T09:55:53+00:00","index":"","fulltext":""},{"type":"submitted","content":"International Journal of Data Science and Analytics","date":"2025-04-01T04:44:02+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"international-journal-of-data-science-and-analytics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jdsa","sideBox":"Learn more about [International Journal of Data Science and Analytics](http://link.springer.com/journal/41060)","snPcode":"41060","submissionUrl":"https://submission.nature.com/new-submission/41060/3","title":"International Journal of Data Science and Analytics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"d3563bc9-0a04-4942-a415-03eb83c856af","owner":[],"postedDate":"May 11th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-13T01:53:45+00:00","versionOfRecord":[],"versionCreatedAt":"2025-05-11 09:55:24","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6349440","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6349440","identity":"rs-6349440","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00