When ten cases are not enough: Uncertainty and power in AI screening evaluations

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Real-world evaluations of artificial intelligence (AI) systems may yield highly variable performance estimates when based on limited sample sizes. In particular, a recent study compared two AI-based diabetic retinopathy screening systems based on only 10 positive cases, where a 20 percentage point difference in sensitivity corresponds to only two misclassified cases. Here, we show through simulation that, under such conditions, both the uncertainty of sensitivity estimates and the statistical power to detect differences between models are severely limited. Even under optimistic assumptions, sample sizes on the order of a few tens of positive cases lead to unstable estimates and insufficient power for meaningful comparison. These findings emphasize the importance of adequately powered studies, caution against overinterpretation of results derived from small cohorts, and support the systematic reporting of statistical power in comparison studies.
Full text 10,154 characters · extracted from preprint-html · click to expand
When ten cases are not enough: Uncertainty and power in AI screening evaluations | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article When ten cases are not enough: Uncertainty and power in AI screening evaluations Gwenolé Quellec This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9601448/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Real-world evaluations of artificial intelligence (AI) systems may yield highly variable performance estimates when based on limited sample sizes. In particular, a recent study compared two AI-based diabetic retinopathy screening systems based on only 10 positive cases, where a 20 percentage point difference in sensitivity corresponds to only two misclassified cases. Here, we show through simulation that, under such conditions, both the uncertainty of sensitivity estimates and the statistical power to detect differences between models are severely limited. Even under optimistic assumptions, sample sizes on the order of a few tens of positive cases lead to unstable estimates and insufficient power for meaningful comparison. These findings emphasize the importance of adequately powered studies, caution against overinterpretation of results derived from small cohorts, and support the systematic reporting of statistical power in comparison studies. Health sciences/Diseases Health sciences/Health care Physical sciences/Mathematics and computing Health sciences/Medical research artificial intelligence diabetic retinopathy screening statistical power uncertainty sample size Full Text Additional Declarations Competing interest reported. The author reports that a software license related to a diabetic retinopathy screening system (OphtAI) has been granted to Evolucare Technologies. The author has also served as a consultant for Evolucare Technologies. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9601448","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":634761301,"identity":"d535ba36-9a5c-4f6e-a5a4-9b7682bbed9c","order_by":0,"name":"Gwenolé Quellec","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABDklEQVRIie2QMUvEMBiGv1sylXNN6XH5CymFuhz9HY4pgWbJcY4nHNJb6iLOhfsfnT84OJe6d2wR6loXwUWNeh0ccjgK5hmSkHwP70sAHI4/y+XXOsEWYE6Od2enFf69oQCIRsXPf6uk46RVYRfy0A58AdNGIop1ou6Cm+7xSgM9tzhhk6mw5Bn4TSZQ1HJZzOoofKiAztCi7HQceHwPvNEc0wKXBdXE31ZwTS3Fwt3qxSjvRlkNmL6hIlQ9vRqF2hQWaGIU/EwBTHMUhIp4ckrhrI/9kkvPr3uO4iBDUywyxag95Xbf02GdzKf3smuHTcJYqbrnbbWwpxy/xQP8+WATTEo+ntA643A4HP+dD6xkWbOQxVfaAAAAAElFTkSuQmCC","orcid":"","institution":"Inserm","correspondingAuthor":true,"prefix":"","firstName":"Gwenolé","middleName":"","lastName":"Quellec","suffix":""}],"badges":[],"createdAt":"2026-05-03 16:38:29","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9601448/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9601448/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":108672719,"identity":"0c02678a-2c53-4d5e-b6fa-7be4d8e6932c","added_by":"auto","created_at":"2026-05-07 08:00:28","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":138318,"visible":true,"origin":"","legend":"","description":"","filename":"Manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9601448/v1_covered_3ab8d95a-f253-49b7-aeb4-86b9a0e11616.pdf"}],"financialInterests":"Competing interest reported. The author reports that a software license related to a diabetic retinopathy screening system (OphtAI) has been granted to Evolucare Technologies. The author has also served as a consultant for Evolucare Technologies.","formattedTitle":"When ten cases are not enough: Uncertainty and power in AI screening evaluations","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"artificial intelligence, diabetic retinopathy, screening, statistical power, uncertainty, sample size","lastPublishedDoi":"10.21203/rs.3.rs-9601448/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9601448/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e Real-world evaluations of artificial intelligence (AI) systems may yield highly variable performance estimates when based on limited sample sizes. In particular, a recent study compared two AI-based diabetic retinopathy screening systems based on only 10 positive cases, where a 20 percentage point difference in sensitivity corresponds to only two misclassified cases. Here, we show through simulation that, under such conditions, both the uncertainty of sensitivity estimates and the statistical power to detect differences between models are severely limited. Even under optimistic assumptions, sample sizes on the order of a few tens of positive cases lead to unstable estimates and insufficient power for meaningful comparison. These findings emphasize the importance of adequately powered studies, caution against overinterpretation of results derived from small cohorts, and support the systematic reporting of statistical power in comparison studies.\u003c/p\u003e","manuscriptTitle":"When ten cases are not enough: Uncertainty and power in AI screening evaluations","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-07 08:00:16","doi":"10.21203/rs.3.rs-9601448/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b3f072c2-f9c2-4874-9f43-de834645ef86","owner":[],"postedDate":"May 7th, 2026","published":true,"recentEditorialEvents":[{"type":"editorInvited","content":"","date":"2026-05-13T17:41:23+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-05-05T11:22:01+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-05-05T11:21:04+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2026-05-03T16:32:52+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":67552606,"name":"Health sciences/Diseases"},{"id":67552607,"name":"Health sciences/Health care"},{"id":67552608,"name":"Physical sciences/Mathematics and computing"},{"id":67552609,"name":"Health sciences/Medical research"}],"tags":[],"updatedAt":"2026-05-07T08:00:16+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-07 08:00:16","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9601448","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9601448","identity":"rs-9601448","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00