Assessing the Capability of Large Language Models in Answering Pediatric Critical Care Board-Style Questions

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background: The potential of Large Language Models (LLMs) in medicine is often linked to massive, resource-intensive models. However, their practical application in specialized fields like pediatric critical care requires exploring the capability of more efficient, locally-deployable open-source alternatives. In this study, we evaluated the accuracy and clinical reasoning of open-source LLMs of varying sizes, specifically assessing if smaller, efficient models can perform comparably to larger ones on pediatric critical care multiple-choice questions. Methods: A set of 100 pediatric critical care MCQs across six clinical domains, i.e. calculation, diagnosis, ethics, management, pharmacology, and physiology, was curated by two pediatric specialists to evaluate eight open-source LLMs, ranging from 2 to 70 billion parameters. The LLMs were assessed using the overall and category-specific accuracy and clinical reasoning quality score based on a 5-points Likert scale. Additionally, two pediatric critical care fellows completed the MCQs for comparison. Cochran’s Q test, McNemar’s test, the Friedman test, and Cohen’s kappa were used for the statistical analysis. Results: While the largest model (Llama-3.3-70B) achieved the highest accuracy (78%; 95% CI, 69%-86%), a key finding was the performance of the much smaller, 14.7-billion parameter Phi-4. This efficient model was strikingly comparable, with 75% accuracy (95% CI, 65%-83%) and a similar reasoning score (4.40 vs 4.49/5). Both models’ performance was on par with pediatric critical care fellows. The LLMs excelled in ethics but struggled with calculations. Inter-rater reliability was excellent for the clinical reasoning assessment (κ = 0.92). Conclusions: Our findings demonstrate that smaller, efficient LLMs can approach the performance of much larger models and pediatric critical care fellows for complex pediatric critical care reasoning. This suggests a viable pathway for developing secure, locally-deployable decision support tools without relying on massive, proprietary systems. At the same time, these models hold potential as complementary resources for trainee education in pediatric critical care. However, their identified weaknesses, especially in calculations, underscore that rigorous, domain-specific validation is an essential prerequisite to ensure safe use in both clinical and educational contexts. Trial registration: Not applicable.
Full text 27,095 characters · extracted from preprint-html · click to expand
Assessing the Capability of Large Language Models in Answering Pediatric Critical Care Board-Style Questions | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Assessing the Capability of Large Language Models in Answering Pediatric Critical Care Board-Style Questions Daniela Chanci, Ronald Moore, Henry P. Foote, Matthew A. Goldstein, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7714101/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 14 You are reading this latest preprint version Abstract Background: The potential of Large Language Models (LLMs) in medicine is often linked to massive, resource-intensive models. However, their practical application in specialized fields like pediatric critical care requires exploring the capability of more efficient, locally-deployable open-source alternatives. In this study, we evaluated the accuracy and clinical reasoning of open-source LLMs of varying sizes, specifically assessing if smaller, efficient models can perform comparably to larger ones on pediatric critical care multiple-choice questions. Methods: A set of 100 pediatric critical care MCQs across six clinical domains, i.e. calculation, diagnosis, ethics, management, pharmacology, and physiology, was curated by two pediatric specialists to evaluate eight open-source LLMs, ranging from 2 to 70 billion parameters. The LLMs were assessed using the overall and category-specific accuracy and clinical reasoning quality score based on a 5-points Likert scale. Additionally, two pediatric critical care fellows completed the MCQs for comparison. Cochran’s Q test, McNemar’s test, the Friedman test, and Cohen’s kappa were used for the statistical analysis. Results: While the largest model (Llama-3.3-70B) achieved the highest accuracy (78%; 95% CI, 69%-86%), a key finding was the performance of the much smaller, 14.7-billion parameter Phi-4. This efficient model was strikingly comparable, with 75% accuracy (95% CI, 65%-83%) and a similar reasoning score (4.40 vs 4.49/5). Both models’ performance was on par with pediatric critical care fellows. The LLMs excelled in ethics but struggled with calculations. Inter-rater reliability was excellent for the clinical reasoning assessment (κ = 0.92). Conclusions: Our findings demonstrate that smaller, efficient LLMs can approach the performance of much larger models and pediatric critical care fellows for complex pediatric critical care reasoning. This suggests a viable pathway for developing secure, locally-deployable decision support tools without relying on massive, proprietary systems. At the same time, these models hold potential as complementary resources for trainee education in pediatric critical care. However, their identified weaknesses, especially in calculations, underscore that rigorous, domain-specific validation is an essential prerequisite to ensure safe use in both clinical and educational contexts. Trial registration: Not applicable. Biological sciences/Computational biology and bioinformatics Health sciences/Health care Physical sciences/Mathematics and computing Health sciences/Medical research Large Language Models Pediatric critical care PICU Multiple-choice questions Clinical question-answering Benchmark Full Text Additional Declarations No competing interests reported. Supplementary Files llmpicusupplement.docx Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 27 Feb, 2026 Reviews received at journal 20 Feb, 2026 Reviews received at journal 11 Feb, 2026 Reviewers agreed at journal 09 Feb, 2026 Reviewers agreed at journal 09 Feb, 2026 Reviewers agreed at journal 19 Dec, 2025 Reviewers agreed at journal 18 Dec, 2025 Reviews received at journal 16 Dec, 2025 Reviewers agreed at journal 10 Dec, 2025 Reviewers invited by journal 22 Oct, 2025 Editor assigned by journal 22 Oct, 2025 Editor invited by journal 07 Oct, 2025 Submission checks completed at journal 03 Oct, 2025 First submitted to journal 03 Oct, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7714101","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":538515806,"identity":"2eab606a-81ec-4966-a0eb-c4bd24b280fa","order_by":0,"name":"Daniela Chanci","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABMklEQVRIie3RPWuDQBjA8RMhWc5mPUlfvsITAg0lkM+iZHAxbaYQSCFCIF3sbtDmYzifHMQlbVaHDmbpZItTsTSFnmmgFDVzh/tzyPngj4MTIZHoP0bzNeQbkr9B/sASRWP5Z4CrCPwlfLI+TtAvQQcizavJSfgY0BR6qOHOXuJs2LvpOANKP5ZdfXFnSfHrvEDU9bUWONBH5HnVadnQv/IiPrn3Dd3FVG49FAlQExgGGQHRLpsIZCDEBKr4TF8SrdZUSsgmAbaDKSfGOyfTPQm+PE4u4vpnGYn4KShfxMxPYXvCFIvpLkE1uYSoUQKBDSEmkTlS+QYIftPY6cpoL2x9pnpPxRvbmO00G0/OG47hk2w3AVIfBNvktnvmhCxIk1Hxlg+V/TIkWZXfi0Qikeho33lmbuOcF9jIAAAAAElFTkSuQmCC","orcid":"","institution":"Duke University","correspondingAuthor":true,"prefix":"","firstName":"Daniela","middleName":"","lastName":"Chanci","suffix":""},{"id":538515807,"identity":"d2c061c7-9eeb-4954-9438-e94449d094b6","order_by":1,"name":"Ronald Moore","email":"","orcid":"","institution":"Duke University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Ronald","middleName":"","lastName":"Moore","suffix":""},{"id":538515808,"identity":"0dba1154-1472-4233-a4db-db18d0e3388a","order_by":2,"name":"Henry P. Foote","email":"","orcid":"","institution":"Duke University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Henry","middleName":"P.","lastName":"Foote","suffix":""},{"id":538515809,"identity":"e3e6ba4c-9f6a-4abf-a65c-fe95cbdd0bf4","order_by":3,"name":"Matthew A. Goldstein","email":"","orcid":"","institution":"NYU Grossman School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Matthew","middleName":"A.","lastName":"Goldstein","suffix":""},{"id":538515810,"identity":"cf4bdec5-ed6d-4685-94e9-ca12072af4ed","order_by":4,"name":"Karan R. Kumar","email":"","orcid":"","institution":"Duke University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Karan","middleName":"R.","lastName":"Kumar","suffix":""},{"id":538515821,"identity":"d8ee4a40-236c-41db-ac37-e3e43a3b135a","order_by":5,"name":"Alexandre T. Rotta","email":"","orcid":"","institution":"Duke University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Alexandre","middleName":"T.","lastName":"Rotta","suffix":""},{"id":538515822,"identity":"ab28607e-9bf3-412f-b082-7bb56a4f9a0f","order_by":6,"name":"Christoph P. Hornik","email":"","orcid":"","institution":"Duke University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Christoph","middleName":"P.","lastName":"Hornik","suffix":""},{"id":538515823,"identity":"2c7073bd-ae46-4344-875c-9e8683509d1f","order_by":7,"name":"Marybeth Burriss-West","email":"","orcid":"","institution":"Duke University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Marybeth","middleName":"","lastName":"Burriss-West","suffix":""},{"id":538515824,"identity":"2073f9a2-c79d-47fe-9a9c-d4d2f3a8c1d8","order_by":8,"name":"Makenzie Hamilton","email":"","orcid":"","institution":"Duke University School of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Makenzie","middleName":"","lastName":"Hamilton","suffix":""},{"id":538515825,"identity":"4f59a3b9-3258-4a52-885e-36049dd6b350","order_by":9,"name":"Rishikesan Kamaleswaran","email":"","orcid":"","institution":"Duke University","correspondingAuthor":false,"prefix":"","firstName":"Rishikesan","middleName":"","lastName":"Kamaleswaran","suffix":""}],"badges":[],"createdAt":"2025-09-25 14:38:39","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7714101/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7714101/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":95037333,"identity":"498a202c-be1d-42b2-91de-844c5538b83f","added_by":"auto","created_at":"2025-11-03 15:34:13","extension":"tiff","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":33783190,"visible":true,"origin":"","legend":"","description":"","filename":"fig1.tiff","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/ca3a41d630c491570813012c.tiff"},{"id":95222137,"identity":"3a14f30f-75a7-48df-ba66-fd4905846281","added_by":"auto","created_at":"2025-11-05 16:20:10","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1486592,"visible":true,"origin":"","legend":"","description":"","filename":"llmpicuv3.docx","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/a77f9c5adf3b71f1d447d70d.docx"},{"id":95222294,"identity":"975375b4-5451-4d21-a232-894af56a7099","added_by":"auto","created_at":"2025-11-05 16:20:25","extension":"tiff","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":110892268,"visible":true,"origin":"","legend":"","description":"","filename":"fig2.tiff","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/9b4ce045cce3a76ccd376526.tiff"},{"id":95037334,"identity":"a98b686d-5b33-4cc4-ab19-b84975414784","added_by":"auto","created_at":"2025-11-03 15:34:13","extension":"tiff","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":26087900,"visible":true,"origin":"","legend":"","description":"","filename":"fig3.tiff","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/bbc2864bc739f3be63634442.tiff"},{"id":95037319,"identity":"5098570d-3496-476d-8466-62b7d11b644c","added_by":"auto","created_at":"2025-11-03 15:34:12","extension":"json","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":11806,"visible":true,"origin":"","legend":"","description":"","filename":"d7123253dc3c47bbb1285dae9a8d77a3.json","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/3e06edd8e40bce3036cb9c10.json"},{"id":95037318,"identity":"bf3f8dab-ab4c-4026-ac78-bcf93bfe2b72","added_by":"auto","created_at":"2025-11-03 15:34:12","extension":"docx","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":31811,"visible":true,"origin":"","legend":"","description":"","filename":"llmpicusupplement.docx","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/1bdc4fb72e947ab3127938a3.docx"},{"id":95037322,"identity":"a1033bab-0cda-4813-96c4-8c9b81c77d49","added_by":"auto","created_at":"2025-11-03 15:34:12","extension":"xml","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":103201,"visible":true,"origin":"","legend":"","description":"","filename":"d7123253dc3c47bbb1285dae9a8d77a31enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/90f9a1dc1c661e52f59ba5e6.xml"},{"id":95222211,"identity":"30d0b86b-30f6-4c24-8b83-6942f5347686","added_by":"auto","created_at":"2025-11-05 16:20:19","extension":"tiff","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":33783190,"visible":true,"origin":"","legend":"","description":"","filename":"fig1.tiff","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/d55dbb99a4a3ba274b02e8d4.tiff"},{"id":95037338,"identity":"3d34430d-73c1-4f74-8bfd-365c8049ae5c","added_by":"auto","created_at":"2025-11-03 15:34:14","extension":"tiff","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":110892268,"visible":true,"origin":"","legend":"","description":"","filename":"fig2.tiff","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/eb3795e3928c374e1fafb786.tiff"},{"id":95037336,"identity":"acd919a2-0dc3-4698-89e6-94eb4bf374b5","added_by":"auto","created_at":"2025-11-03 15:34:13","extension":"tiff","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":26087900,"visible":true,"origin":"","legend":"","description":"","filename":"fig3.tiff","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/bd00d4ff8e69cd258e0d3ed5.tiff"},{"id":95222303,"identity":"ed207f13-7b97-454f-a5df-88ba42cddfeb","added_by":"auto","created_at":"2025-11-05 16:20:26","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":792208,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/f0d8867b44e58f4bbd154e4a.png"},{"id":95221686,"identity":"13a21906-cc88-4c75-b0b9-5abe96ed0734","added_by":"auto","created_at":"2025-11-05 16:19:34","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":206213,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/0adc4233c0e222cef5764750.png"},{"id":95222343,"identity":"ed28db56-12c2-4f24-8a5a-cfa4cebe5960","added_by":"auto","created_at":"2025-11-05 16:20:30","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":371937,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/879dd0a7604a2945574a7e8a.png"},{"id":95222068,"identity":"1b01751c-bac7-4e85-b2a3-710e96f482d6","added_by":"auto","created_at":"2025-11-05 16:20:05","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":222024,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefig1.png","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/d9c67cac54be7455907f5a52.png"},{"id":95222329,"identity":"ea0fbb2b-b666-4371-8af9-cadba1229c18","added_by":"auto","created_at":"2025-11-05 16:20:27","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":244949,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefig2.png","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/786de69ad693be5e5ef82e44.png"},{"id":95037323,"identity":"b5df033d-bec7-4655-a692-b85acf0a8839","added_by":"auto","created_at":"2025-11-03 15:34:12","extension":"png","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":101544,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefig3.png","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/06c2cdd9c26c119d9e46702b.png"},{"id":95037329,"identity":"7646b5f5-1647-4007-aebf-5a1d45f812d8","added_by":"auto","created_at":"2025-11-03 15:34:12","extension":"png","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":221862,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/dfe372983dc606cc2ec7f565.png"},{"id":95037321,"identity":"62cb6dca-3a93-436e-9046-28b966984342","added_by":"auto","created_at":"2025-11-03 15:34:12","extension":"png","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":41004,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/6aee44ed2370565fdb544244.png"},{"id":95037326,"identity":"2fae9e1c-6a01-4c3a-9d6d-f9e12ae21dfc","added_by":"auto","created_at":"2025-11-03 15:34:12","extension":"png","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":101279,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/fe5e756cb7b8380728305254.png"},{"id":95037331,"identity":"e68163e7-6e36-4552-9576-86a060345e75","added_by":"auto","created_at":"2025-11-03 15:34:12","extension":"xml","order_by":19,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":100255,"visible":true,"origin":"","legend":"","description":"","filename":"d7123253dc3c47bbb1285dae9a8d77a31structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/ec2d4451edc17a3e8d66633d.xml"},{"id":95222299,"identity":"50265fcf-9c19-4675-8b38-94060c6d3716","added_by":"auto","created_at":"2025-11-05 16:20:25","extension":"html","order_by":20,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":115833,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/bfd69914f61231a26f1ad5d1.html"},{"id":95312138,"identity":"8092db61-48a7-4572-90e7-ee41640c20c1","added_by":"auto","created_at":"2025-11-06 15:47:28","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":866004,"visible":true,"origin":"","legend":"","description":"","filename":"llmpicuv3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1_covered_4b2f2467-19ae-45ca-9e32-9fc66728475a.pdf"},{"id":95037317,"identity":"3e1271ea-7151-48b6-831e-6af2d007e2b2","added_by":"auto","created_at":"2025-11-03 15:34:12","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":31811,"visible":true,"origin":"","legend":"","description":"","filename":"llmpicusupplement.docx","url":"https://assets-eu.researchsquare.com/files/rs-7714101/v1/b31425ebb32d6b7c4d9382e3.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Assessing the Capability of Large Language Models in Answering Pediatric Critical Care Board-Style Questions","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Large Language Models, Pediatric critical care, PICU, Multiple-choice questions, Clinical question-answering, Benchmark","lastPublishedDoi":"10.21203/rs.3.rs-7714101/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7714101/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground: \u003c/strong\u003eThe potential of Large Language Models (LLMs) in medicine is often linked to massive, resource-intensive models. However, their practical application in specialized fields like pediatric critical care requires exploring the capability of more efficient, locally-deployable open-source alternatives. In this study, we evaluated the accuracy and clinical reasoning of open-source LLMs of varying sizes, specifically assessing if smaller, efficient models can perform comparably to larger ones on pediatric critical care multiple-choice questions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods: \u003c/strong\u003eA set of 100 pediatric critical care MCQs across six clinical domains, i.e. calculation, diagnosis, ethics, management, pharmacology, and physiology, was curated by two pediatric specialists to evaluate eight open-source LLMs, ranging from 2 to 70 billion parameters. The LLMs were assessed using the overall and category-specific accuracy and clinical reasoning quality score based on a 5-points Likert scale. Additionally, two pediatric critical care fellows completed the MCQs for comparison. Cochran’s Q test, McNemar’s test, the Friedman test, and Cohen’s kappa were used for the statistical analysis.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults: \u003c/strong\u003eWhile the largest model (Llama-3.3-70B) achieved the highest accuracy (78%; 95% CI, 69%-86%), a key finding was the performance of the much smaller, 14.7-billion parameter Phi-4. This efficient model was strikingly comparable, with 75% accuracy (95% CI, 65%-83%) and a similar reasoning score (4.40 vs 4.49/5). Both models’ performance was on par with pediatric critical care fellows. The LLMs excelled in ethics but struggled with calculations. Inter-rater reliability was excellent for the clinical reasoning assessment (κ = 0.92).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions: \u003c/strong\u003eOur findings demonstrate that smaller, efficient LLMs can approach the performance of much larger models and pediatric critical care fellows for complex pediatric critical care reasoning. This suggests a viable pathway for developing secure, locally-deployable decision support tools without relying on massive, proprietary systems. At the same time, these models hold potential as complementary resources for trainee education in pediatric critical care. However, their identified weaknesses, especially in calculations, underscore that rigorous, domain-specific validation is an essential prerequisite to ensure safe use in both clinical and educational contexts.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTrial registration:\u003c/strong\u003e Not applicable.\u003c/p\u003e","manuscriptTitle":"Assessing the Capability of Large Language Models in Answering Pediatric Critical Care Board-Style Questions","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-03 15:34:07","doi":"10.21203/rs.3.rs-7714101/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-02-27T06:42:31+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-20T13:49:40+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-11T07:36:45+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"123518062132220104495047849053544524677","date":"2026-02-09T08:04:37+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"294230480888929761084667407468254443451","date":"2026-02-09T07:46:57+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"123322782493579523577358351838872426558","date":"2025-12-19T12:46:15+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"390331247903800433538143246902308719","date":"2025-12-18T16:18:27+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-12-17T02:34:27+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"235290951911801911733787741983945416690","date":"2025-12-10T07:05:31+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-10-22T17:07:57+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-10-22T17:02:05+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-10-07T05:09:39+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-10-03T14:28:30+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2025-10-03T14:24:45+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"dd76b59a-fdd4-4886-9bdb-f6dfb9dc2326","owner":[],"postedDate":"November 3rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":57277453,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":57277454,"name":"Health sciences/Health care"},{"id":57277455,"name":"Physical sciences/Mathematics and computing"},{"id":57277456,"name":"Health sciences/Medical research"}],"tags":[],"updatedAt":"2026-04-17T05:08:15+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-03 15:34:07","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7714101","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7714101","identity":"rs-7714101","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00