Small Language Models in Clinical Medicine: A Systematic Review of Performance, Safety, and Deployment Feasibility

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Large language models are increasingly used in clinical medicine, but their reliance on cloud servers conflicts with patient-privacy requirements and excludes resource-limited healthcare systems. Small language models (SLMs) of up to four billion parameters can run locally on a single commodity GPU, keeping data inside the institution while reaching performance comparable to much larger systems. Here we systematically review 14 studies that deploy SLMs for clinical prediction, information extraction, and medical question answering. Domain- adapted small models reached a median 91% of the best reported performance of larger baselines, and we found no significant correlation between parameter count and task accuracy. Only half of the studies evaluated hallucination rates, and none reported calibration or epistemic uncertainty. The computational case for on-premise clinical AI is therefore strong, but the safety engineering required for responsible deployment, particularly in agentic sub-agent pipelines, is largely absent.
Full text 13,268 characters · extracted from preprint-html · click to expand
Small Language Models in Clinical Medicine: A Systematic Review of Performance, Safety, and Deployment Feasibility | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Systematic Review Small Language Models in Clinical Medicine: A Systematic Review of Performance, Safety, and Deployment Feasibility Alon Gorenshtein, Mahmud Omar, Yiftach Barash, Jonathan B. Kruskal, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9488729/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Large language models are increasingly used in clinical medicine, but their reliance on cloud servers conflicts with patient-privacy requirements and excludes resource-limited healthcare systems. Small language models (SLMs) of up to four billion parameters can run locally on a single commodity GPU, keeping data inside the institution while reaching performance comparable to much larger systems. Here we systematically review 14 studies that deploy SLMs for clinical prediction, information extraction, and medical question answering. Domain- adapted small models reached a median 91% of the best reported performance of larger baselines, and we found no significant correlation between parameter count and task accuracy. Only half of the studies evaluated hallucination rates, and none reported calibration or epistemic uncertainty. The computational case for on-premise clinical AI is therefore strong, but the safety engineering required for responsible deployment, particularly in agentic sub-agent pipelines, is largely absent. Artificial Intelligence and Machine Learning Hospital Medicine Internal Medicine Small Language Models Agentic AI Large Language Models Systematic Review Clinical Natural Language Processing Edge Deployment PRISMA Full Text Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9488729","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Systematic Review","associatedPublications":[],"authors":[{"id":627307649,"identity":"e7126683-8583-49ab-8efe-fa3e00f8af6a","order_by":0,"name":"Alon Gorenshtein","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABC0lEQVRIiWNgGAWjYBCDBH4JxgYJEAPCZwPiAwS0SM4gWYvBDQYG4rTw9y8+9uHjjro849vNjbdu1DDk8c8+Y/i5ooxBju9GAlYtEjeeJc+ceeZwsdmdg83WOccYiiXO5RhLnjnHYCyJQwvDjTPGzLxtBxK33Uhsk85tYEhsOMOWINnYxpC4AYcWeZCWv211iZtnQLXMP8OW/BOopR6XFoPzPcbMjG3MiRskoFo2nGE+BrIFGCDYtRjeYEtm7G07nDgD4heJxI1ALZYN5yQMZ555gFWL3PnDhxl+Ah3WP7v94e2cGpvEeWcYm282lNnI8x3H4X0JVHEJDAYm4D+AW24UjIJRMApGARgAACqEahXT3MrWAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0009-0000-7542-8608","institution":"BRIDGE GenAI Lab, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA","correspondingAuthor":true,"prefix":"","firstName":"Alon","middleName":"","lastName":"Gorenshtein","suffix":""},{"id":627307650,"identity":"ac7ee54f-637e-4dc9-8819-39b75c75cafd","order_by":1,"name":"Mahmud Omar","email":"","orcid":"","institution":"BRIDGE GenAI Lab, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA","correspondingAuthor":false,"prefix":"","firstName":"Mahmud","middleName":"","lastName":"Omar","suffix":""},{"id":627307651,"identity":"cf9b311b-30ce-438d-a64c-f7b1ab808fd8","order_by":2,"name":"Yiftach Barash","email":"","orcid":"","institution":"BRIDGE GenAI Lab, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA","correspondingAuthor":false,"prefix":"","firstName":"Yiftach","middleName":"","lastName":"Barash","suffix":""},{"id":627307652,"identity":"da499e58-ed1e-4a24-b185-0f33f404eafc","order_by":3,"name":"Jonathan B. Kruskal","email":"","orcid":"","institution":"Department of Radiology, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA","correspondingAuthor":false,"prefix":"","firstName":"Jonathan","middleName":"B.","lastName":"Kruskal","suffix":""},{"id":627307653,"identity":"d58518ff-83b1-459f-8a15-2dcaf3042c28","order_by":4,"name":"Muneeb Ahmed","email":"","orcid":"","institution":"Department of Radiology, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA","correspondingAuthor":false,"prefix":"","firstName":"Muneeb","middleName":"","lastName":"Ahmed","suffix":""},{"id":627307654,"identity":"f5ec54ba-1e3f-4d5f-8c52-34a86d1f2862","order_by":5,"name":"Olga R. Brook","email":"","orcid":"","institution":"Department of Radiology, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA","correspondingAuthor":false,"prefix":"","firstName":"Olga","middleName":"R.","lastName":"Brook","suffix":""},{"id":627307655,"identity":"feebe8e3-b38c-4c68-9c8e-090699ef05af","order_by":6,"name":"Ben Illigens","email":"","orcid":"","institution":"Hasso Plattner Institute for Digital Health, Icahn School of Medicine at Mount Sinai, New York, NY, USA","correspondingAuthor":false,"prefix":"","firstName":"Ben","middleName":"","lastName":"Illigens","suffix":""},{"id":627307656,"identity":"3779a469-d75c-4d0d-9fc8-8403e3523f2d","order_by":7,"name":"Girish N. Nadkarni","email":"","orcid":"","institution":"The Windreich Department of Artificial Intelligence and Human Health, Mount Sinai Medical Center, New York, NY, USA","correspondingAuthor":false,"prefix":"","firstName":"Girish","middleName":"N.","lastName":"Nadkarni","suffix":""},{"id":627307657,"identity":"45be93b5-e6cf-4ff8-9793-045806a28a17","order_by":8,"name":"Eyal Klang","email":"","orcid":"","institution":"BRIDGE GenAI Lab, Beth Israel Deaconess Medical Center, Harvard Medical School, Boston, MA, USA","correspondingAuthor":false,"prefix":"","firstName":"Eyal","middleName":"","lastName":"Klang","suffix":""}],"badges":[],"createdAt":"2026-04-21 21:33:30","currentVersionCode":1,"declarations":{"humanSubjects":true,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":true,"humanSubjectConsent":true,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-9488729/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9488729/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":109203563,"identity":"8c395de9-194d-4fe3-9d94-612e946d6d46","added_by":"auto","created_at":"2026-05-13 14:39:28","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":975642,"visible":true,"origin":"","legend":"","description":"","filename":"manuscriptieee.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9488729/v1_covered_7946052d-d950-43ce-858b-0df849c5ac90.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eSmall Language Models in Clinical Medicine: A Systematic Review of Performance, Safety, and Deployment Feasibility\u003c/p\u003e","fulltext":[],"fulltextSource":"","fullText":"","funders":[{"identity":"2c3467e7-573a-49f1-a804-694d9d02a1be","identifier":"10.13039/100000002","name":"National Institutes of Health","awardNumber":"S10OD026880 ","order_by":0},{"identity":"80cdf3d1-061f-4117-91c8-c3fedf8210a9","identifier":"10.13039/100000002","name":"National Institutes of Health","awardNumber":"S10OD030463","order_by":1}],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Beth Israel Deaconess Medical Center","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Small Language Models, Agentic AI, Large Language Models, Systematic Review, Clinical Natural Language Processing, Edge Deployment, PRISMA","lastPublishedDoi":"10.21203/rs.3.rs-9488729/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9488729/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eLarge language models are increasingly used in clinical medicine, but their reliance on cloud servers conflicts with patient-privacy requirements and excludes resource-limited healthcare systems. Small language models (SLMs) of up to four billion parameters can run locally on a single commodity GPU, keeping data inside the institution while reaching performance comparable to much larger systems. Here we systematically review 14 studies that deploy SLMs for clinical prediction, information extraction, and medical question answering. Domain- adapted small models reached a median 91% of the best reported performance of larger baselines, and we found no significant correlation between parameter count and task accuracy. Only half of the studies evaluated hallucination rates, and none reported calibration or epistemic uncertainty. The computational case for on-premise clinical AI is therefore strong, but the safety engineering required for responsible deployment, particularly in agentic sub-agent pipelines, is largely absent.\u003c/p\u003e","manuscriptTitle":"Small Language Models in Clinical Medicine: A Systematic Review of Performance, Safety, and Deployment Feasibility","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-23 03:54:54","doi":"10.21203/rs.3.rs-9488729/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b737a566-2e5e-46db-807a-df6ebddb134e","owner":[],"postedDate":"April 23rd, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":66852691,"name":"Artificial Intelligence and Machine Learning"},{"id":66852692,"name":"Hospital Medicine"},{"id":66852693,"name":"Internal Medicine"}],"tags":[],"updatedAt":"2026-04-23T03:54:55+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-23 03:54:54","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9488729","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9488729","identity":"rs-9488729","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00