GReX-Bench: Benchmarking Generalization, Robustness, and Explainability in AI-Generated Image Detection | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article GReX-Bench: Benchmarking Generalization, Robustness, and Explainability in AI-Generated Image Detection Nusrat Tasnim, kutub Uddin, Khalid Malik This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8633550/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 12 You are reading this latest preprint version Abstract Generative AI has significantly transformed media creation and accessibility, enabling the rapid generation of fake content, particularly images. However, AI-generated images pose growing challenges for misinformation and biometric security, contributing to declining public trust, rising fraud, and increasing social engineering attacks.Despite progress in forensics, existing methods still face major limitations: (i) reliance on non-standardized benchmarks, (ii) inconsistent training protocols, (iii) limited evaluation metrics, (iv) weak interpretability, (v) lack of human-readable explanations, and (vi) insufficient attention to deployment and usability in real-world settings. These issues hinder fair comparison and obscure true reliability in security-critical applications.To address this, we introduce \textbf{GReX-Bench}, the first unified benchmarking framework for reproducible evaluation of forensic, anti-forensic (AF), and explainability. We benchmark sixteen prior detectors across eight public datasets (e.g., GAN, diffusion, and low-level vision) under six AF attacks. We analyze model behavior through confidence, ROC, and explainability techniques, including model-specific, model-agnostic, and generative LLM. Our findings reveal significant generalization gaps, with many methods performing well in-distribution but degrading across datasets, particularly under AF attacks. We also examine deployment factors such as efficiency, latency, and scalability to guide practical adoption. Generative Models GAN Diffusion AI-Generated Image Detection Generalization Robustness and Explainability Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 21 Mar, 2026 Reviews received at journal 21 Mar, 2026 Reviewers agreed at journal 01 Mar, 2026 Reviewers agreed at journal 28 Feb, 2026 Reviews received at journal 20 Feb, 2026 Reviews received at journal 20 Feb, 2026 Reviewers agreed at journal 10 Feb, 2026 Reviewers agreed at journal 09 Feb, 2026 Reviewers invited by journal 08 Feb, 2026 Editor assigned by journal 29 Jan, 2026 Submission checks completed at journal 20 Jan, 2026 First submitted to journal 18 Jan, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8633550","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":589465107,"identity":"e5b632cb-d29c-4be9-a7a8-2a5303793e27","order_by":0,"name":"Nusrat Tasnim","email":"","orcid":"","institution":"Korea Aerospace University","correspondingAuthor":false,"prefix":"","firstName":"Nusrat","middleName":"","lastName":"Tasnim","suffix":""},{"id":589465108,"identity":"22cb9ebe-54a7-4fc3-aeec-8bda4963c504","order_by":1,"name":"kutub Uddin","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAp0lEQVRIiWNgGAWjYBACCSBmZqhgYGBj4GEAkcRoYQbCMyRrYWwDMYnVIjkj/5h04bzD0XzsZw8wfCg7TFiLtEQym/TMbYdz23jyEhhnnCNCixxICy9IiwSPATNvG9Fa5kC1/CVGC9hhvA1QLYzEaJHseWxsPeNYOtAvOQYHe86lE9YicTzx4e2CGuvc+e1nDB/8KLMmrAUFHCBR/SgYBaNgFIwCXAAAPp8yI84uSUIAAAAASUVORK5CYII=","orcid":"","institution":"University of Michigan–Flint","correspondingAuthor":true,"prefix":"","firstName":"kutub","middleName":"","lastName":"Uddin","suffix":""},{"id":589465110,"identity":"30bf3360-2aa2-499e-95ec-3b8726632aae","order_by":2,"name":"Khalid Malik","email":"","orcid":"","institution":"University of Michigan–Flint","correspondingAuthor":false,"prefix":"","firstName":"Khalid","middleName":"","lastName":"Malik","suffix":""}],"badges":[],"createdAt":"2026-01-18 20:53:17","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8633550/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8633550/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":102521866,"identity":"f024b4dd-556e-4e87-bdc6-ee47193abb59","added_by":"auto","created_at":"2026-02-12 14:42:55","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":9463578,"visible":true,"origin":"","legend":"","description":"","filename":"air26manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8633550/v1_covered_fe1f76cd-8b6e-4705-a40b-11d43f1dbdaf.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"GReX-Bench: Benchmarking Generalization, Robustness, and Explainability in AI-Generated Image Detection","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"artificial-intelligence-review","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"aire","sideBox":"Learn more about [Artificial Intelligence Review](http://link.springer.com/journal/10462)","snPcode":"10462","submissionUrl":"https://submission.nature.com/new-submission/10462/3","title":"Artificial Intelligence Review","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Generative Models, GAN, Diffusion, AI-Generated Image Detection, Generalization, Robustness, and Explainability","lastPublishedDoi":"10.21203/rs.3.rs-8633550/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8633550/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Generative AI has significantly transformed media creation and accessibility, enabling the rapid generation of fake content, particularly images. However, AI-generated images pose growing challenges for misinformation and biometric security, contributing to declining public trust, rising fraud, and increasing social engineering attacks.Despite progress in forensics, existing methods still face major limitations: (i) reliance on non-standardized benchmarks, (ii) inconsistent training protocols, (iii) limited evaluation metrics, (iv) weak interpretability, (v) lack of human-readable explanations, and (vi) insufficient attention to deployment and usability in real-world settings. These issues hinder fair comparison and obscure true reliability in security-critical applications.To address this, we introduce \\textbf{GReX-Bench}, the first unified benchmarking framework for reproducible evaluation of forensic, anti-forensic (AF), and explainability. We benchmark sixteen prior detectors across eight public datasets (e.g., GAN, diffusion, and low-level vision) under six AF attacks. We analyze model behavior through confidence, ROC, and explainability techniques, including model-specific, model-agnostic, and generative LLM. Our findings reveal significant generalization gaps, with many methods performing well in-distribution but degrading across datasets, particularly under AF attacks. We also examine deployment factors such as efficiency, latency, and scalability to guide practical adoption.","manuscriptTitle":"GReX-Bench: Benchmarking Generalization, Robustness, and Explainability in AI-Generated Image Detection","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-12 14:40:58","doi":"10.21203/rs.3.rs-8633550/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-03-21T17:38:09+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-03-21T16:41:09+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"112254832619765323508628602015833995475","date":"2026-03-01T07:28:20+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"108925578480619149643392502621790455862","date":"2026-02-28T18:28:42+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-20T17:39:41+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-20T12:49:35+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"71212756765595773039501241154961811662","date":"2026-02-10T10:52:50+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"119827641910603195908343826815400602044","date":"2026-02-10T04:56:24+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-02-08T18:06:09+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-01-29T14:14:40+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-01-21T02:40:02+00:00","index":"","fulltext":""},{"type":"submitted","content":"Artificial Intelligence Review","date":"2026-01-18T20:43:28+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"artificial-intelligence-review","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"aire","sideBox":"Learn more about [Artificial Intelligence Review](http://link.springer.com/journal/10462)","snPcode":"10462","submissionUrl":"https://submission.nature.com/new-submission/10462/3","title":"Artificial Intelligence Review","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"ab3df8ca-f2f7-493d-ac8e-14cea5bbf454","owner":[],"postedDate":"February 12th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[],"tags":[],"updatedAt":"2026-03-21T17:53:20+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-12 14:40:58","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8633550","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8633550","identity":"rs-8633550","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.