A Performance-Based Rubric for Generative AI Use in Medical Students’ Research Tasks: Development and Initial Psychometric Evaluation

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Background: As generative AI becomes embedded in medical training, patient safety depends on graduates’ ability to recognize AI limitations and bias, document AI involvement transparently, and verify AI-generated information rather than accept it uncritically. We developed a performance-based rubric to assess observable generative AI (LLM) literacy behaviors within authentic coursework. Methods: In a single-institution evaluation (Spring 2025), third-year medical students (n = 50 submissions) completed a structured research proposal and submitted the corresponding AI chat transcript and an AI-use disclosure. A four-domain rubric was developed through three pilot–revise cycles: AI Use Documentation, Prompt Generation, Verification, and Integration. Each domain was scored 0–3 (total 0–12). Three educators independently scored all submissions. Inter-rater reliability was assessed using ICC (average-measures, agreement). Construct-relevant patterns were examined via domain distributions (floor effects), performance bands (lower 25%, middle 50%, upper 25%), within-submission differences across domains (Friedman with Bonferroni-adjusted Wilcoxon tests), inter-domain associations (Spearman), and correlation with overall GPA (Spearman). Results: Mean (SD) domain scores were: AI Use Documentation 0.67 (1.08), Prompt Generation 1.33 (0.69), Verification 0.41 (0.71), and Integration 1.64 (0.67); total score 4.06 (1.80). Floor effects were substantial for AI Use Documentation (64% scored 0) and Verification (60% scored 0). Inter-rater reliability was high (ICC: Documentation 0.99, Prompt Generation 0.84, Verification 0.93, Integration 0.83). Verification was significantly lower than Prompt Generation and Integration (Bonferroni-adjusted p < 0.008). Inter-domain correlations were weak (ρ −0.206 to 0.310). Total scores showed no significant association with GPA (r = 0.194, p = 0.201). Conclusions: This rubric demonstrated strong scoring reliability and produced initial psychometric evidence consistent with measuring distinct, observable LLM-use competencies. Findings highlight prominent gaps in verification and transparent documentation, reinforcing competency guidance that emphasizes recognizing AI limitations and verifying AI output to protect patient safety. Further multi-site validation and implementation work is warranted.
Full text 19,409 characters · extracted from preprint-html · click to expand
A Performance-Based Rubric for Generative AI Use in Medical Students’ Research Tasks: Development and Initial Psychometric Evaluation | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Performance-Based Rubric for Generative AI Use in Medical Students’ Research Tasks: Development and Initial Psychometric Evaluation Nino Shiukashvili, Mariam Rochikashvili, Vasil Kupradze, Nana Gonjilashvili, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7600944/v2 This work is licensed under a CC BY 4.0 License Status: Posted Version 2 posted You are reading this latest preprint version Show more versions Abstract Background: As generative AI becomes embedded in medical training, patient safety depends on graduates’ ability to recognize AI limitations and bias, document AI involvement transparently, and verify AI-generated information rather than accept it uncritically. We developed a performance-based rubric to assess observable generative AI (LLM) literacy behaviors within authentic coursework. Methods: In a single-institution evaluation (Spring 2025), third-year medical students (n = 50 submissions) completed a structured research proposal and submitted the corresponding AI chat transcript and an AI-use disclosure. A four-domain rubric was developed through three pilot–revise cycles: AI Use Documentation, Prompt Generation, Verification, and Integration. Each domain was scored 0–3 (total 0–12). Three educators independently scored all submissions. Inter-rater reliability was assessed using ICC (average-measures, agreement). Construct-relevant patterns were examined via domain distributions (floor effects), performance bands (lower 25%, middle 50%, upper 25%), within-submission differences across domains (Friedman with Bonferroni-adjusted Wilcoxon tests), inter-domain associations (Spearman), and correlation with overall GPA (Spearman). Results: Mean (SD) domain scores were: AI Use Documentation 0.67 (1.08), Prompt Generation 1.33 (0.69), Verification 0.41 (0.71), and Integration 1.64 (0.67); total score 4.06 (1.80). Floor effects were substantial for AI Use Documentation (64% scored 0) and Verification (60% scored 0). Inter-rater reliability was high (ICC: Documentation 0.99, Prompt Generation 0.84, Verification 0.93, Integration 0.83). Verification was significantly lower than Prompt Generation and Integration (Bonferroni-adjusted p < 0.008). Inter-domain correlations were weak (ρ −0.206 to 0.310). Total scores showed no significant association with GPA (r = 0.194, p = 0.201). Conclusions: This rubric demonstrated strong scoring reliability and produced initial psychometric evidence consistent with measuring distinct, observable LLM-use competencies. Findings highlight prominent gaps in verification and transparent documentation, reinforcing competency guidance that emphasizes recognizing AI limitations and verifying AI output to protect patient safety. Further multi-site validation and implementation work is warranted. Artificial Intelligence AI Literacy Medical Students Performance-Based Assessment Medical Education Full Text Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 2 posted You are reading this latest preprint version Show more versions Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7600944","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":517246513,"identity":"9fa416c3-96fc-4882-ae47-331d4cc0c548","order_by":0,"name":"Nino Shiukashvili","email":"","orcid":"","institution":"Ken Walker International University","correspondingAuthor":false,"prefix":"","firstName":"Nino","middleName":"","lastName":"Shiukashvili","suffix":""},{"id":517246514,"identity":"18988e94-44a4-4762-bfb9-a7436c7bf2d5","order_by":1,"name":"Mariam Rochikashvili","email":"","orcid":"","institution":"Ken Walker International University","correspondingAuthor":false,"prefix":"","firstName":"Mariam","middleName":"","lastName":"Rochikashvili","suffix":""},{"id":517246515,"identity":"58ea6811-5030-4375-a0f2-672905aea7e5","order_by":2,"name":"Vasil Kupradze","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/UlEQVRIiWNgGAWjYJACxgYwlQBEFTYgfuMB4rU8OJMG5hOvhfFh22EwE68Wg+PHHz6c2XYnn789/eGDhDPn7da2HwbaUmMTjVPLmRxjw41tzyxnnHmQbJBQcTt525lEoJZjabkNOLRINuSwSQLdY8BwI+GYRMKZ28lmB4BaGBsO49bS//wZWIv8jcT2H4lt55LNzj/Er4VfIsFMciNQi8GNZDaGxLYDdmY3CNjCL/HG2HDGuWcGhmeeMQMdlpxgdgNoSwIev7Dxpz982FN2x0DuePrDjz8q7OzNzgOD7kONDU4tUHAAzkoEq0zArxxViz1hxaNgFIyCUTDSAADJRm95ilABLQAAAABJRU5ErkJggg==","orcid":"","institution":"Ken Walker International University","correspondingAuthor":true,"prefix":"","firstName":"Vasil","middleName":"","lastName":"Kupradze","suffix":""},{"id":517246516,"identity":"0d59a6aa-96ba-46ef-9467-d834fbd77c8a","order_by":3,"name":"Nana Gonjilashvili","email":"","orcid":"","institution":"Ken Walker International University","correspondingAuthor":false,"prefix":"","firstName":"Nana","middleName":"","lastName":"Gonjilashvili","suffix":""},{"id":517246517,"identity":"7234a5d1-ca15-4b5f-81f7-994862692199","order_by":4,"name":"Nino Gvajaia","email":"","orcid":"","institution":"Ken Walker International University","correspondingAuthor":false,"prefix":"","firstName":"Nino","middleName":"","lastName":"Gvajaia","suffix":""},{"id":517246518,"identity":"28fea24c-54f7-4726-bd69-6405959623ac","order_by":5,"name":"Luka Kutchava","email":"","orcid":"","institution":"Ken Walker International University","correspondingAuthor":false,"prefix":"","firstName":"Luka","middleName":"","lastName":"Kutchava","suffix":""},{"id":517246519,"identity":"95dd7eaf-7d38-4aac-bdbb-f6b18b6a5600","order_by":6,"name":"Nino Tevzadze","email":"","orcid":"","institution":"Ken Walker International University","correspondingAuthor":false,"prefix":"","firstName":"Nino","middleName":"","lastName":"Tevzadze","suffix":""},{"id":517246520,"identity":"effbd346-a4aa-490a-bf9a-3b66b280cc5f","order_by":7,"name":"Nona Janikashvili","email":"","orcid":"","institution":"Tbilisi State Medical University","correspondingAuthor":false,"prefix":"","firstName":"Nona","middleName":"","lastName":"Janikashvili","suffix":""},{"id":517246521,"identity":"b6325ffc-9068-4e41-a35d-d8b0b98bfd1d","order_by":8,"name":"Archil Undilashvili","email":"","orcid":"","institution":"Emory University","correspondingAuthor":false,"prefix":"","firstName":"Archil","middleName":"","lastName":"Undilashvili","suffix":""},{"id":517246522,"identity":"ad18ce8b-5b0f-4018-8378-309ddcd4e15c","order_by":9,"name":"Eka Ekaladze","email":"","orcid":"","institution":"Ken Walker International University","correspondingAuthor":false,"prefix":"","firstName":"Eka","middleName":"","lastName":"Ekaladze","suffix":""}],"badges":[],"createdAt":"2025-09-12 13:08:39","currentVersionCode":2,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-7600944/v2","doiUrl":"https://doi.org/10.21203/rs.3.rs-7600944/v2","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":100373070,"identity":"87000986-a51e-4f50-85eb-a83f38e02d88","added_by":"auto","created_at":"2026-01-16 08:13:34","extension":"png","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":122301,"visible":true,"origin":"","legend":"","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-7600944/v2/8f1f5e57c07d08a2cb6e6eb5.png"},{"id":100273518,"identity":"dabb5d88-6525-42f0-a753-66dffda21072","added_by":"auto","created_at":"2026-01-14 21:36:57","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":23326,"visible":true,"origin":"","legend":"","description":"","filename":"TheArtofAIDialogueMedicalStudentsAILiteracyintheAgeofConversationalTechnology.docx","url":"https://assets-eu.researchsquare.com/files/rs-7600944/v2/ef7ab3445905b91af2b1db06.docx"},{"id":100273511,"identity":"ea27110d-62a6-4dd6-97af-6c5ce8230334","added_by":"auto","created_at":"2026-01-14 21:36:57","extension":"json","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":12342,"visible":true,"origin":"","legend":"","description":"","filename":"4801c9ae88094aa7bb9cc37a4266a15b.json","url":"https://assets-eu.researchsquare.com/files/rs-7600944/v2/e3149765140229c89f8bf2b4.json"},{"id":100372574,"identity":"5b1aa25e-257c-48bf-bb9c-1754ba96ee56","added_by":"auto","created_at":"2026-01-16 08:12:41","extension":"xml","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":56647,"visible":true,"origin":"","legend":"","description":"","filename":"4801c9ae88094aa7bb9cc37a4266a15b1enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7600944/v2/87e0a409c0ba394a29518f31.xml"},{"id":100373078,"identity":"dfe81851-c7ec-408b-bb13-87e73c7a64a5","added_by":"auto","created_at":"2026-01-16 08:13:35","extension":"png","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":122301,"visible":true,"origin":"","legend":"","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-7600944/v2/0608d035c5d7e2f8e8bb2ba3.png"},{"id":100373254,"identity":"2f922ec9-7e96-4490-bffd-42d5f2a77211","added_by":"auto","created_at":"2026-01-16 08:13:55","extension":"png","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":48850,"visible":true,"origin":"","legend":"","description":"","filename":"OnlineFigure1.png","url":"https://assets-eu.researchsquare.com/files/rs-7600944/v2/099792f54fdaeb202c18bc3e.png"},{"id":100372209,"identity":"8c8d80aa-3bc0-47ec-9b3a-becd6fe713e1","added_by":"auto","created_at":"2026-01-16 08:11:51","extension":"xml","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":53162,"visible":true,"origin":"","legend":"","description":"","filename":"4801c9ae88094aa7bb9cc37a4266a15b1structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7600944/v2/160c621ece7f0e8b47f159c3.xml"},{"id":100273517,"identity":"d1420f26-5423-4351-9ff9-169f1cf4693c","added_by":"auto","created_at":"2026-01-14 21:36:57","extension":"html","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":63179,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7600944/v2/d307c69d79d36f89b77166aa.html"},{"id":100383670,"identity":"15243027-3f02-4384-a93f-cf18e752d6a5","added_by":"auto","created_at":"2026-01-16 10:48:00","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":388341,"visible":true,"origin":"","legend":"","description":"","filename":"APerformanceBasedRubricforGenerativeAIUseinMedicalStudentsResearchTasksDevelopmentandInitialPsychometricEvaluation.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7600944/v2_covered_3e40bbde-45b3-4d9b-9b94-9789f2edcde1.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eA Performance-Based Rubric for Generative AI Use in Medical Students’ Research Tasks: Development and Initial Psychometric Evaluation\u003c/strong\u003e\u003c/p\u003e","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Artificial Intelligence, AI Literacy, Medical Students, Performance-Based Assessment, Medical Education","lastPublishedDoi":"10.21203/rs.3.rs-7600944/v2","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7600944/v2","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground:\u003c/strong\u003e As generative AI becomes embedded in medical training, patient safety depends on graduates’ ability to recognize AI limitations and bias, document AI involvement transparently, and verify AI-generated information rather than accept it uncritically. We developed a performance-based rubric to assess observable generative AI (LLM) literacy behaviors within authentic coursework.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods:\u003c/strong\u003e In a single-institution evaluation (Spring 2025), third-year medical students (n = 50 submissions) completed a structured research proposal and submitted the corresponding AI chat transcript and an AI-use disclosure. A four-domain rubric was developed through three pilot–revise cycles: AI Use Documentation, Prompt Generation, Verification, and Integration. Each domain was scored 0–3 (total 0–12). Three educators independently scored all submissions. Inter-rater reliability was assessed using ICC (average-measures, agreement). Construct-relevant patterns were examined via domain distributions (floor effects), performance bands (lower 25%, middle 50%, upper 25%), within-submission differences across domains (Friedman with Bonferroni-adjusted Wilcoxon tests), inter-domain associations (Spearman), and correlation with overall GPA (Spearman).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults:\u003c/strong\u003e Mean (SD) domain scores were: AI Use Documentation 0.67 (1.08), Prompt Generation 1.33 (0.69), Verification 0.41 (0.71), and Integration 1.64 (0.67); total score 4.06 (1.80). Floor effects were substantial for AI Use Documentation (64% scored 0) and Verification (60% scored 0). Inter-rater reliability was high (ICC: Documentation 0.99, Prompt Generation 0.84, Verification 0.93, Integration 0.83). Verification was significantly lower than Prompt Generation and Integration (Bonferroni-adjusted p \u0026lt; 0.008). Inter-domain correlations were weak (ρ −0.206 to 0.310). Total scores showed no significant association with GPA (r = 0.194, p = 0.201).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions:\u003c/strong\u003e This rubric demonstrated strong scoring reliability and produced initial psychometric evidence consistent with measuring distinct, observable LLM-use competencies. Findings highlight prominent gaps in verification and transparent documentation, reinforcing competency guidance that emphasizes recognizing AI limitations and verifying AI output to protect patient safety. Further multi-site validation and implementation work is warranted.\u003c/p\u003e","manuscriptTitle":"A Performance-Based Rubric for Generative AI Use in Medical Students’ Research Tasks: Development and Initial Psychometric Evaluation","msid":"","msnumber":"","nonDraftVersions":[{"code":2,"date":"2026-01-14 21:36:52","doi":"10.21203/rs.3.rs-7600944/v2","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}},{"code":1,"date":"2025-09-18 18:02:55","doi":"10.21203/rs.3.rs-7600944/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"f39ec806-562d-4d83-929a-4a0613ffe1f5","owner":[],"postedDate":"January 14th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-11-13T06:23:55+00:00","versionOfRecord":[],"versionCreatedAt":"2026-01-14 21:36:52","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v2","identity":"rs-7600944","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7600944","identity":"rs-7600944","version":["v2"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-20T11:00:21.680559+00:00
License: CC-BY-4.0