A multi-model study of diagnostic faithfulness in AI-generated histopathology images | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article A multi-model study of diagnostic faithfulness in AI-generated histopathology images Hong-Yu Zhou, Qinxin Wang, Yuebin Jiang, Fan Yang, Yuqing Cao, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9432464/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract Translating pathological knowledge into images offers new ways to support medical education, illustrate rare disease patterns, and expand training data beyond the constraints of privacy regulation and annotation scarcity. Yet whether current generative AI can produce images diagnostically faithful to clinical descriptions remains unknown. We studied seven text-to-image models on an atlas-derived cohort spanning 6 organ systems and 19 tumor subcategories, combining tri-dimensional expert scoring with structured instruction-following verification across 9,612 pathologist ratings. Only the best-performing model approaches diagnostic interpretability; even so, it omits approximately one third of prompted features, with omission rates rising with case complexity. Automated metrics prove inadequate, as their scores do not reflect clinically meaningful quality differences. A model trained exclusively on breast histopathology is outperformed by general-purpose models within its own domain, suggesting that domain-specific training alone is insufficient without the underlying capacity to handle complex diagnostic descriptions. These findings reveal that visual plausibility is not a sufficient proxy for diagnostic utility, and that both generative models and evaluation standards must advance before synthetic pathology images can be responsibly adopted in oncology. Biological sciences/Cancer/Cancer imaging Biological sciences/Computational biology and bioinformatics/Computational models Full Text Additional Declarations There is NO Competing Interest. Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9432464","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":637034105,"identity":"e3a65ed0-7632-4f42-ba9f-362f29a4e9ec","order_by":0,"name":"Hong-Yu Zhou","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7UlEQVRIiWNgGAWjYDACCQYGZiCVAOFVADEzcwMpWs6AtDCSooWxDUzi12Jwu8f4c0HF4Tx+6fZr0rzzaqP524FaflRsw63lzhkD4xlnDhdLzjlTJs277XjujMOMDYw9Z27j1GJ2I8cgmbftcOKGGzlpQC3HchuAWpgZ2/BrOYzQMudY7nwitBg2Q7SkH5PmbajJ3UBIi/2NtGJmnjPpiTNn5DBbzjl2IHcjUMtBfH6RnJG8+TNPhXViv0T6wxtvaupy550/fPDBjwrcWpAAjwkwjg6DmQeIUQ8E7I8/MDDUEal4FIyCUTAKRhIAAAtjX64afOq2AAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0002-1256-7050","institution":"Tsinghua University","correspondingAuthor":true,"prefix":"","firstName":"Hong-Yu","middleName":"","lastName":"Zhou","suffix":""},{"id":637034106,"identity":"f250a251-51d3-4f88-9ab7-f7569b11eca3","order_by":1,"name":"Qinxin Wang","email":"","orcid":"","institution":"School of Biomedical Engineering, Tsinghua Medicine","correspondingAuthor":false,"prefix":"","firstName":"Qinxin","middleName":"","lastName":"Wang","suffix":""},{"id":637034107,"identity":"a3370707-ab59-4b9b-9e88-ee969e20c81e","order_by":2,"name":"Yuebin Jiang","email":"","orcid":"","institution":"Department of pathology, Zhangzhou Affiliated Hospital of Fujian Medical University","correspondingAuthor":false,"prefix":"","firstName":"Yuebin","middleName":"","lastName":"Jiang","suffix":""},{"id":637034108,"identity":"e59ddaf7-9d79-48ea-b652-577dfbe0962a","order_by":3,"name":"Fan Yang","email":"","orcid":"","institution":"Chinese Academy of Medical Sciences and Peking Union Medical College","correspondingAuthor":false,"prefix":"","firstName":"Fan","middleName":"","lastName":"Yang","suffix":""},{"id":637034109,"identity":"48bbcee2-5ad9-4940-8530-e03cf475cafe","order_by":4,"name":"Yuqing Cao","email":"","orcid":"","institution":"Chinese Academy of Medical Sciences and Peking Union Medical College","correspondingAuthor":false,"prefix":"","firstName":"Yuqing","middleName":"","lastName":"Cao","suffix":""},{"id":637034110,"identity":"63df4ecf-69c3-4488-861f-25deabcd657c","order_by":5,"name":"Jiang Chang","email":"","orcid":"","institution":"Chinese Academy of Medical Sciences and Peking Union Medical College","correspondingAuthor":false,"prefix":"","firstName":"Jiang","middleName":"","lastName":"Chang","suffix":""},{"id":637034111,"identity":"ca06686d-075d-48c8-afa4-07b923198efa","order_by":6,"name":"Peiqing Ma","email":"","orcid":"","institution":"National Cancer Center/National Clinical Research Center for Cancer/Cancer Hospital, Chinese Academy of Medical Sciences and Peking Union Medical College","correspondingAuthor":false,"prefix":"","firstName":"Peiqing","middleName":"","lastName":"Ma","suffix":""},{"id":637034112,"identity":"98099583-2dd8-419b-875a-dba1f4756805","order_by":7,"name":"Jiachen Ji","email":"","orcid":"","institution":"School of Biomedical Engineering, Tsinghua Medicine","correspondingAuthor":false,"prefix":"","firstName":"Jiachen","middleName":"","lastName":"Ji","suffix":""},{"id":637034113,"identity":"8c8fa3b9-9b1b-4c49-b37a-4488dd3bd314","order_by":8,"name":"Yi-Qun Che","email":"","orcid":"","institution":"Center for Clinical Laboratory, Beijing Friendship Hospital, Capital Medical University","correspondingAuthor":false,"prefix":"","firstName":"Yi-Qun","middleName":"","lastName":"Che","suffix":""},{"id":637034114,"identity":"3b253f06-9744-456c-aa9c-9d3d9ffa3aa3","order_by":9,"name":"Lin Yang","email":"","orcid":"","institution":"Chinese Academy of Medical Sciences and Peking Union Medical College","correspondingAuthor":false,"prefix":"","firstName":"Lin","middleName":"","lastName":"Yang","suffix":""},{"id":637034115,"identity":"9fafee71-98da-4149-a20b-712f5c275fcd","order_by":10,"name":"Xihai Zhao","email":"","orcid":"https://orcid.org/0000-0001-8953-0566","institution":"Department of Biomedical Engineering, School of Medicine, Tsinghua University","correspondingAuthor":false,"prefix":"","firstName":"Xihai","middleName":"","lastName":"Zhao","suffix":""}],"badges":[],"createdAt":"2026-04-16 03:15:52","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9432464/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9432464/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":109088213,"identity":"4272cc1c-0375-48ca-ad82-b25f143eabc6","added_by":"auto","created_at":"2026-05-12 13:20:06","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":13747603,"visible":true,"origin":"","legend":"Article File","description":"","filename":"maintext.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9432464/v1_covered_d3a8fde5-4926-49ee-8085-67a636cc5726.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"A multi-model study of diagnostic faithfulness in AI-generated histopathology images","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-9432464/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9432464/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Translating pathological knowledge into images offers new ways to support medical education, illustrate rare disease patterns, and expand training data beyond the constraints of privacy regulation and annotation scarcity. Yet whether current generative AI can produce images diagnostically faithful to clinical descriptions remains unknown. We studied seven text-to-image models on an atlas-derived cohort spanning 6 organ systems and 19 tumor subcategories, combining tri-dimensional expert scoring with structured instruction-following verification across 9,612 pathologist ratings. Only the best-performing model approaches diagnostic interpretability; even so, it omits approximately one third of prompted features, with omission rates rising with case complexity. Automated metrics prove inadequate, as their scores do not reflect clinically meaningful quality differences. A model trained exclusively on breast histopathology is outperformed by general-purpose models within its own domain, suggesting that domain-specific training alone is insufficient without the underlying capacity to handle complex diagnostic descriptions. These findings reveal that visual plausibility is not a sufficient proxy for diagnostic utility, and that both generative models and evaluation standards must advance before synthetic pathology images can be responsibly adopted in oncology.","manuscriptTitle":"A multi-model study of diagnostic faithfulness in AI-generated histopathology images","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-12 12:34:35","doi":"10.21203/rs.3.rs-9432464/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"nature-communications","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"NCOMMS","sideBox":"Learn more about [Nature Communications](http://www.nature.com/ncomms/)","snPcode":"","submissionUrl":"https://mts-ncomms.nature.com/","title":"Nature Communications","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Communications","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"e19c53e8-8bf3-4cc4-8f98-9b9351ca244b","owner":[],"postedDate":"May 12th, 2026","published":true,"recentEditorialEvents":[{"type":"reviewersInvited","content":"1","date":"2026-05-08T16:23:21+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":67808672,"name":"Biological sciences/Cancer/Cancer imaging"},{"id":67808673,"name":"Biological sciences/Computational biology and bioinformatics/Computational models"}],"tags":[],"updatedAt":"2026-05-12T12:34:36+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-12 12:34:35","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9432464","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9432464","identity":"rs-9432464","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.