Evaluating the Quality of AI-Generated Corrective Feedback for Medical Learners | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Evaluating the Quality of AI-Generated Corrective Feedback for Medical Learners Bonnie Desselle, Emma Simon, Amy Prudhomme, Christy Mumphrey, and 7 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9076837/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 13 You are reading this latest preprint version Abstract Background Faculty feedback is vital for the professional growth of learners. Barriers such as faculty’s confidence in their feedback skills and time limitations may inhibit the provision of quality feedback. Generative large language models (LLMs), such as ChatGPT, have the potential to support faculty in providing high-quality corrective feedback by producing effective narrative feedback scripts for faculty to reference. Objective This study explores the quality of Artificial Intelligence (AI)-generated text in the context of corrective feedback, based on standardized scenarios of learner deficiencies. Methods Six medical education leaders blindly rated two narrative feedback scripts for eighteen learner deficiency scenarios; one script for each scenario was authored by educators, the other script was generated by ChatGPT. The raters scored each script in four domains (behavior-centric, balanced between positive and negative, non-judgmental, and focused) using a 3-point scale (0–2), where 0 = No, 1 = Somewhat, 2 = Yes. The mean scores of the educator authored scripts and the ChatGPT-generated scripts were compared by domain. An overall score was also generated for the educator authored script and the ChatGPT-generated scripts using descriptive statistics. Results Across all four domains, ChatGPT-generated scripts received high mean scores, with all rated higher than the educator authored scripts. Also, ChatGPT-generated responses achieved higher overall scores than those authored by educators for 17 of 18 scenarios. Conclusion AI-generated responses can be of high quality and may provide a useful tool for faculty who are preparing to give corrective feedback. Further studies are needed to explore the utility of AI in supporting feedback delivery. feedback artificial intelligence large language models ChatGPT Figures Figure 1 Introduction Faculty feedback is essential to the development and professional growth of medical students, residents, and fellows. The Liaison Committee on Medical Education and the Accreditation Council for Graduate Medical Education require that faculty provide formative feedback to medical students, residents, and fellows.( 1 , 2 ) Common barriers, including limited time, insufficient observation, fear of damaging rapport, lack of training, and discomfort with delivering feedback, can impede the delivery of high-quality feedback.( 1 , 3 ) To address these challenges, many institutions invest in faculty development, offering workshops and digital resources to enhance feedback skills.( 4 – 12 ) Fromme et al developed the concept of feedback scripts, a tool that identifies the specific learner behavior that needs correction and guides the faculty in what to say to the learner.( 13 ) Feedback scripts were shown to be efficacious for faculty planning to deliver corrective feedback.( 13 ) Large language models (LLMs), a specific type of Artificial Intelligence, can produce fluent, conversational scripts, but do not possess clinical judgement of real world experience.( 14 , 15 ) Little is known about whether such models can generate corrective feedback scripts that meet established standards for effective educational practice. Prior studies have examined LLM communication quality in clinical or patient-facing settings, but less is known about the quality of AI generated formative feedback for learners in medical education.( 16 – 18 ) Medical educators increasingly face pressure to provide feedback, and existing faculty-development resources often require time investment and may not offer immediate, situation specific guidance, thus we sought to explore whether an LLM can assist by serving as a user-friendly tool that supports in-the-moment feedback delivery. Objective To address this gap, we conducted an exploratory study evaluating the quality of feedback produced by a commonly used LLM, ChatGPT, for a series of common learner performance issues. Methods Fromme et al published a series of 21 feedback scripts generated by pediatric educators attending a national conference.( 13 ) Feedback scripts addressed a variety of learner deficiency scenarios. Our study team used three scripts from this series for training purposes and 18 for the study. ChatGPT (GPT-4o; OpenAI, San Francisco, CA; accessed May–June 2025) was prompted to generate three to four sentences of written formative feedback to the learner for each scenario, responding as a pediatric faculty member. We compiled the ChatGPT-generated feedback alongside the original educator authored feedback scripts written by pediatric educators and published in the Fromme et al series.( 13 ) We recruited six faculty at our institution (who were not involved in writing the scenarios or scripts) to rate the educator authored feedback scripts and the feedback generated by ChatGPT. The raters were blinded to whether the scripts were authored by educators or generated by ChatGPT. Raters included six medical education leaders from a single academic institution, including two current or prior clerkship directors, two residency program directors, and two fellowship program directors. Scripts were ordered randomly and identified to the raters only as “Script 1” or “Script 2.” Each script was rated using a behaviorally anchored 4-domain assessment tool (supplemental materials), modified from the works of Halman et al and Schlair et al to align with the format of the feedback script being assessed.( 6 , 19 ) The revised tool assessed feedback scripts across domains established as important for effective feedback, specifically: ( 1 ) Behavior-centered, ( 2 ) Balanced between positive and negative feedback, ( 3 ) Non-judgmental, and ( 4 ) Focused. Each domain was scored using a 3-level rating scale (0–2), where 0 = No, 1 = Somewhat, 2 = Yes, yielding a minimum total score of 0 and a maximum total score of 2 for each domain and a minimum total score of 0 and a maximum total score of 8 for each script. Raters also indicated their overall script preference for each learner scenario Descriptive statistics were applied to compare the ratings using T-tests in Excel (Microsoft 365, Seattle, WA). We also compared the mean scores of each scenario and the four domains. The Louisiana State University Health Sciences Center Institutional Review Board granted full exempt status for this study. Results There was 100% completion by the raters. Across all four domains, ChatGPT-generated scripts received higher mean scores than educator authored scripts, as demonstrated in Fig. 1; behavior-centered domain (ChatGPT 1.90, SD 0.30 vs. educator 1.7, SD 0.53, P=.001 ), balanced domain (ChatGPT 1.87, SD 0.36 vs. educator 1.48 SD 0.73, P<.001), nonjudgmental domain (ChatGPT 1.87, SD 0.39 vs. educator 1.47, SD 0.39, P<.001), and focused domain (ChatGPT 1.91, SD 0.29 vs. educator 1.56, SD 0.29, P<.001). Mean scores of each domain are based on a 3-level rating scale (0 = No, 1 = Somewhat, 2 = Yes) with a maximum score of 2. As shown in Table 1 , analyzing each scenario individually, ChatGPT-generated scripts were rated higher than educator authored scripts in 17 of the 18 scenarios, half being of statistical significance. An analysis of combined average scores revealed that ChatGPT-generated scripts were more highly rated than educator authored scripts with a statistical significance (ChatGPT 7.55, SD 1.13 vs. educator authored 6.23, SD 1.60, P<.05). Table 1 Average Scores of ChatGPT and Educator Generated Narrative Scripts by Scenario Scenario Score of ChatGPT-Generated Script Score of Educator Authored Script P value Scenario 1 8.00 7.33 .073 Scenario 2 7.33 6.67 .243 Scenario 3 7.33 4.83 .046* Scenario 4 6.83 5.67 .164 Scenario 5 7.00 7.33 .651 Scenario 6 7.83 6.33 .016* Scenario 7 7.83 6.83 .076 Scenario 8 6.83 6.67 .871 Scenario 9 7.67 7.00 .144 Scenario 10 7.17 6.00 .034* Scenario 11 7.83 6.50 .046* Scenario 12 7.83 5.83 .033* Scenario 13 7.83 5.17 .026* Scenario 14 7.50 7.00 .438 Scenario 15 7.17 5.50 .0265* Scenario 16 8.00 5.50 .0303* Scenario 17 7.83 5.33 .0046* Scenario 18 8.00 6.67 .091 AVERAGE 7.55 6.23 < .001* Script scores are the sum of four domain ratings, where 0 = No, 1 = Somewhat, and 2 = Yes, yielding a maximum score of 8. The averages of the six raters' ChatGPT-generated and Educator authored script scores were compared. *Identifies statistically significant results. ChatGPT-generated feedback scripts were preferred over educator authored scripts in 84 of 108 ratings (78%). Only one educator authored script (number 5), addressing limited feedback receptivity of a learner was chosen by over half of the raters (83%). Discussion In this exploratory study, feedback generated by a LLM was highly rated by medical education leaders and the content favorably compared to educator authored scripts. ChatGPT’s strong performance in this communication-focused task reflects the intentional training on large-scale conversational text that optimizes conversational fluency.( 14 ) The overall preference of LLM-generated feedback suggests that generative AI may offer a viable, real-time tool to assist clinical educators in delivering high-quality feedback by offering support in communication, content, and tone consideration. AI-generated responses in this study were rated highly by medical education leaders, noted to be well-balanced between positive and negative feedback, were appropriately focused on modifiable behaviors, and were non-judgmental in nature. This aligns with Ayers et al.’s study, which found that chatbot responses were significantly preferred over physician responses, rated higher in quality, and viewed as more empathetic.( 17 ) Our findings suggest that AI may be useful for assisting faculty in quickly crafting sensitive, balanced feedback for learners across a variety of scenarios. While we inputted a single, short, straightforward prompt into ChatGPT, future studies may determine whether more detailed prompts result in even higher-quality feedback scripts. It is worth investigating whether adding additional details, such as the level of the learner, prior evaluations, and exam scores, meaningfully improves the quality of the feedback generated by the LLM. Given that AI-generated feedback has resulted in better performance by medical students in the context of simulated surgical procedures, further studies should focus on the impact of feedback quality on clinical performance.( 20 ) Limitations of this single-center, pediatrics focused study include considerations of generalizability. While 18 learner deficiency scenarios were evaluated, they are not comprehensive and cannot capture all the nuances experienced in clinical settings. This study used pre-written scenarios; results may differ if faculty vary the amount or order of details inputted into the LLM. Decision fatigue of raters could have impacted results of this study. Notably, the perspective of learners receiving the feedback is missing from this study. Future work is needed to explore the experience of those receiving feedback to understand their perspective on the quality of the AI-generated feedback and to compare it to educator authored scripts. Furthermore, additional studies are needed to understand in detail when and why AI-generated scripts are preferred. Conclusion In this exploratory study, ChatGPT-generated feedback was rated as high quality by medical education leaders. While AI cannot replace educator judgment or clinical experience, LLMs show promise as a useful tool to support faculty in crafting timely, balanced, behavior-focused formative feedback in medical education. Abbreviations AI Artificial intelligence LLM Large language model Declarations Ethics approval and consent to participate The Louisiana State University Health Sciences Center Institutional Review Board granted full exemption of this study. Consent for publication Not applicable. Availability of data and materials All data generated or analyzed during this study are included in this published article [and its supplementary information files]. Competing interests The authors declare that they have no competing interests Funding This study was conducted without any funding source. Authors' contributions B.D. conceived the study. All authors contributed to the development of methodology and interpretation of results. E.S. was responsible for data collection and analysis. All authors participated in the writing of the manuscript and approved the final manuscript. Acknowledgements Not applicable. References Functions and Structure of a Medical School Standards for Accreditation of Medical Education Programs Leading to the MD Degree [Internet]. 2025. Available from: https://lcme.org/publications/ ACGME Common Program Requirements [Internet]. 2025. Available from: https://www.acgme.org/globalassets/pfassets/programrequirements/2025-reformatted-requirements/cprresidency_2025_reformatted.pdf McCutcheon S, Duchemin AM. Overcoming barriers to effective feedback: a solution-focused faculty development approach. Int J Med Educ. 2020;11:230–2. 10.5116/ijme.5f7c.3157 . PubMed PMID: 33099519; PubMed Central PMCID: PMC7882126. Barak G, Foradori D, Fromme HB, Zuniga L, Dean A. Balancing Honest Assessment and Compassion for Learners Experiencing Burnout: A Workshop and Feedback Tool for Clinical Teachers. MedEdPORTAL. 2024;20:11449. 10.15766/mep_2374-8265 . .11449 PubMed PMID: 39410923; PubMed Central PMCID: PMC11473647. Broquet KE, Punwani M. Helping international medical graduates engage in effective feedback. Acad Psychiatry. 2012;36(4):282–7. 10.1176/appi.ap.11020031 . PubMed PMID: 22851024. Schlair S, Dyche L, Milan F. Longitudinal Faculty Development Program to Promote Effective Observation and Feedback Skills in Direct Clinical Observation. MedEdPORTAL. 2017;13:10648. 10. 15766/mep_2374-8265.10648 PubMed PMID: 30800849; PubMed Central PMCID: PMC6338150. Minehart RD, Rudolph J, Pian-Smith MCM, Raemer DB. Improving faculty feedback to resident trainees during a simulated case: a randomized, controlled trial of an educational intervention. Anesthesiology. 2014;120(1):160–71. 10. 1097/ALN.0000000000000058 PubMed PMID: 24398734. Giving Effective Feedback: Beyond Great Job [Internet]. 2007. Available from: https://www.youtube.com/watch?v=DbfISZjG9mU Gigante J, Dell M, Sharkey A. Getting beyond Good job: how to give effective feedback. Pediatrics. 2011;127(2):205–7. 10.1542/peds.2010-3351 . PubMed PMID: 21242222. Yarris LM, Fu R, LaMantia J, Linden JA, Gene Hern H, Lefebvre C, et al. Effect of an educational intervention on faculty and resident satisfaction with real-time feedback in the emergency department. Acad Emerg Med. 2011;18(5):504–12. 10.1111/j.1553-2712.2011.01055 . x PubMed PMID: 21569169; PubMed Central PMCID: PMC3095955. Lockyer J, Armson H, Könings KD, Lee-Krueger RCW, des Ordons AR, Ramani S, et al. In-the-Moment Feedback and Coaching: Improving R2C2 for a New Context. J Grad Med Educ. 2020;12(1):27–35. 10.4300/JGME-D-19-00508.1 . PubMed PMID: 32089791; PubMed Central PMCID: PMC7012514. Fromme HB, Ryan MS, Gray K, Black E, Paik S, Griego E, et al. A Script for What Ails Your Learners: Feedback Scripts to Promote Effective Learning. Acad Pediatr. 2020;20(5):721–3. 10.1016/j.acap.2020.02.005 . PubMed PMID: 32044468. Fromme HB, Ryan MS, Gray K, Black E, Paik S, Griego E, et al. A Script for What Ails Your Learners: Feedback Scripts to Promote Effective Learning. Acad Pediatr. 2020;20(5):721–3. 10.1016/j.acap.2020.02.005 . PubMed PMID: 32044468. Kalyan KS. A survey of GPT-3 family large language models including ChatGPT and GPT-4. Nat Lang Process J. 2024;6:100048. 10.1016/j.nlp.2023.100048 . Giray L. Enhancing Communication with ChatGPT: A Guide for Academic Writers, Teachers, and Professionals. J Pract Cardiovasc Sci. 2024;10(2):113–8. 10.4103/jpcs.jpcs_27_24 . Srivastava R, Srivastava S. Can Artificial Intelligence aid communication? Considering the possibilities of GPT-3 in Palliative care. Indian J Palliat Care. 2023;29(4):418–25. doi:10.25259/IJPC_155_2023 PubMed PMID: 38058478; PubMed Central PMCID: PMC10696352. Ayers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, et al. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Intern Med. 2023;183(6):589–96. 10.1001/jamainternmed.2023.1838 . PubMed PMID: 37115527; PubMed Central PMCID: PMC10148230. Dikici E, Bigelow M, Prevedello LM, White RD, Erdal BS. Integrating AI into radiology workflow: levels of research, production, and feedback maturity. J Med Imaging (Bellingham). 2020;7(1):016502. 10.1117/1.JMI.7.1.016502 . PubMed PMID: 32064302; PubMed Central PMCID: PMC7012173. Halman S, Dudek N, Wood T, Pugh D, Touchie C, McAleer S, et al. Direct Observation of Clinical Skills Feedback Scale: Development and Validity Evidence. Teach Learn Med. 2016;28(4):385–94. 10.1080/10401334.2016.1186552 . Fazlollahi AM, Bakhaidar M, Alsayegh A, Yilmaz R, Winkler-Schwartz A, Mirchi N, et al. Effect of Artificial Intelligence Tutoring vs Expert Instruction on Learning Simulated Surgical Skills Among Medical Students: A Randomized Clinical Trial. JAMA Netw Open. 2022;5(2):e2149008. 10.1001/jamanetworkopen.2021.49008 . PubMed PMID: 35191972; PubMed Central PMCID: PMC8864513. Additional Declarations No competing interests reported. Supplementary Files SupplementalmaterialsforBMCMedEdmanuscript.docx The prompt that was entered into ChatGPT. Eighteen learner deficiency scenarios, followed by the educator authored and ChatGPT-generated feedback scripts. The rubric used by raters to score each script. Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 19 May, 2026 Reviews received at journal 15 May, 2026 Reviewers agreed at journal 15 May, 2026 Reviews received at journal 14 May, 2026 Reviewers agreed at journal 14 May, 2026 Reviews received at journal 09 May, 2026 Reviewers agreed at journal 19 Apr, 2026 Reviewers agreed at journal 17 Apr, 2026 Reviewers invited by journal 09 Apr, 2026 Editor invited by journal 30 Mar, 2026 Editor assigned by journal 11 Mar, 2026 Submission checks completed at journal 11 Mar, 2026 First submitted to journal 09 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9076837","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":624716079,"identity":"3c6f3882-e300-4d86-bb0f-a409a7c8f9ea","order_by":0,"name":"Bonnie Desselle","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAwUlEQVRIiWNgGAWjYBACA4YEEGUjw8DAAxZgbCBSSxoPyVoOk6DFnD358McfFed55GfkHv7wg8FGdsMBAlose56lSfOcuc1jcCMvTbKHIc2YoBaDGzlmzIxtQC0SQAbQhYlEaMn//PHnv3NAh+UYf2Zg+E+MlhwGCd6GAzwMN3IMpBkYDhDWAvSLmTTPsWQegzNvzCR7DJKNZxLSAgyxxx9/1NjJybfnGH/4UWEn20dIC7o7SVM+CkbBKBgFowAHAABFCEJmMBZUTQAAAABJRU5ErkJggg==","orcid":"","institution":"Louisiana State University Health Sciences Center New Orleans","correspondingAuthor":true,"prefix":"","firstName":"Bonnie","middleName":"","lastName":"Desselle","suffix":""},{"id":624716080,"identity":"bf824fd9-9d35-454d-b0fd-a27a7afad5b6","order_by":1,"name":"Emma Simon","email":"","orcid":"","institution":"Manning Family Children’s Hospital","correspondingAuthor":false,"prefix":"","firstName":"Emma","middleName":"","lastName":"Simon","suffix":""},{"id":624716082,"identity":"d3f48429-9804-4669-a1fc-f605e4cd3ccb","order_by":2,"name":"Amy Prudhomme","email":"","orcid":"","institution":"Louisiana State University Health Sciences Center New Orleans","correspondingAuthor":false,"prefix":"","firstName":"Amy","middleName":"","lastName":"Prudhomme","suffix":""},{"id":624716083,"identity":"9054b301-0c00-4265-b817-4c9b40fe3c28","order_by":3,"name":"Christy Mumphrey","email":"","orcid":"","institution":"Louisiana State University Health Sciences Center New Orleans","correspondingAuthor":false,"prefix":"","firstName":"Christy","middleName":"","lastName":"Mumphrey","suffix":""},{"id":624716084,"identity":"7f9aac5b-616c-479a-9005-9710e4251cdd","order_by":4,"name":"Leslie Reilly","email":"","orcid":"","institution":"Louisiana State University Health Sciences Center New Orleans","correspondingAuthor":false,"prefix":"","firstName":"Leslie","middleName":"","lastName":"Reilly","suffix":""},{"id":624716085,"identity":"bd9c3da7-4717-42c8-bbca-3167a114cce3","order_by":5,"name":"Margaret Huntwork","email":"","orcid":"","institution":"Tulane University","correspondingAuthor":false,"prefix":"","firstName":"Margaret","middleName":"","lastName":"Huntwork","suffix":""},{"id":624716086,"identity":"e13a4d36-cd17-43cc-8f41-5b527485a804","order_by":6,"name":"Shubho Sarkar","email":"","orcid":"","institution":"Louisiana State University Health Sciences Center New Orleans","correspondingAuthor":false,"prefix":"","firstName":"Shubho","middleName":"","lastName":"Sarkar","suffix":""},{"id":624716087,"identity":"cc7084b0-e0c8-4e82-8954-4e30d50391d2","order_by":7,"name":"George Hescock","email":"","orcid":"","institution":"Louisiana State University Health Sciences Center New Orleans","correspondingAuthor":false,"prefix":"","firstName":"George","middleName":"","lastName":"Hescock","suffix":""},{"id":624716093,"identity":"abb39179-118e-49bb-b8c1-a1965f4bff9d","order_by":8,"name":"Jessica Patrick","email":"","orcid":"","institution":"Louisiana State University Health Sciences Center New Orleans","correspondingAuthor":false,"prefix":"","firstName":"Jessica","middleName":"","lastName":"Patrick","suffix":""},{"id":624716099,"identity":"37c149d0-3b46-460a-a971-1979a6f05870","order_by":9,"name":"Kelly Gajewski","email":"","orcid":"","institution":"Louisiana State University Health Sciences Center New Orleans","correspondingAuthor":false,"prefix":"","firstName":"Kelly","middleName":"","lastName":"Gajewski","suffix":""},{"id":624716101,"identity":"4a69ac5a-caf1-48f7-8250-767c4ff21954","order_by":10,"name":"Amy Creel","email":"","orcid":"","institution":"Louisiana State University Health Sciences Center New Orleans","correspondingAuthor":false,"prefix":"","firstName":"Amy","middleName":"","lastName":"Creel","suffix":""}],"badges":[],"createdAt":"2026-03-09 20:54:25","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9076837/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9076837/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":107191087,"identity":"445dd729-5f76-4609-b05d-c01870d26aac","added_by":"auto","created_at":"2026-04-17 21:15:18","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":32186,"visible":true,"origin":"","legend":"\u003cp\u003eChatGPT-generated scripts received higher mean scores than educator authored scripts,\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-9076837/v1/5315e43cb25655b858aaaccf.png"},{"id":107482635,"identity":"714f53dd-a9ae-420d-9c7e-3f89fc45912c","added_by":"auto","created_at":"2026-04-22 02:24:15","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":255777,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9076837/v1/877b32bf-a877-4765-b5ef-8e909bbfa073.pdf"},{"id":107191086,"identity":"bcd1ce5d-ae1c-4206-baac-df667ca80847","added_by":"auto","created_at":"2026-04-17 21:15:18","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":37693,"visible":true,"origin":"","legend":"\u003cp\u003e1.\tThe prompt that was entered into ChatGPT.\u003c/p\u003e\n\u003cp\u003e2.\tEighteen learner deficiency scenarios, followed by the educator authored and ChatGPT-generated feedback scripts.\u003c/p\u003e\n\u003cp\u003e3.\tThe rubric used by raters to score each script.\u003c/p\u003e","description":"","filename":"SupplementalmaterialsforBMCMedEdmanuscript.docx","url":"https://assets-eu.researchsquare.com/files/rs-9076837/v1/1719c58c5d124bc1fb4f2d7b.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Evaluating the Quality of AI-Generated Corrective Feedback for Medical Learners","fulltext":[{"header":"Introduction","content":"\u003cp\u003eFaculty feedback is essential to the development and professional growth of medical students, residents, and fellows. The Liaison Committee on Medical Education and the Accreditation Council for Graduate Medical Education require that faculty provide formative feedback to medical students, residents, and fellows.(\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) Common barriers, including limited time, insufficient observation, fear of damaging rapport, lack of training, and discomfort with delivering feedback, can impede the delivery of high-quality feedback.(\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e)\u003c/p\u003e \u003cp\u003eTo address these challenges, many institutions invest in faculty development, offering workshops and digital resources to enhance feedback skills.(\u003cspan additionalcitationids=\"CR5 CR6 CR7 CR8 CR9 CR10 CR11\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e) Fromme et al developed the concept of feedback scripts, a tool that identifies the specific learner behavior that needs correction and guides the faculty in what to say to the learner.(\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e) Feedback scripts were shown to be efficacious for faculty planning to deliver corrective feedback.(\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e) Large language models (LLMs), a specific type of Artificial Intelligence, can produce fluent, conversational scripts, but do not possess clinical judgement of real world experience.(\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e) Little is known about whether such models can generate corrective feedback scripts that meet established standards for effective educational practice. Prior studies have examined LLM communication quality in clinical or patient-facing settings, but less is known about the quality of AI generated formative feedback for learners in medical education.(\u003cspan additionalcitationids=\"CR17\" citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e)\u003c/p\u003e \u003cp\u003eMedical educators increasingly face pressure to provide feedback, and existing faculty-development resources often require time investment and may not offer immediate, situation specific guidance, thus we sought to explore whether an LLM can assist by serving as a user-friendly tool that supports in-the-moment feedback delivery.\u003c/p\u003e\u003ch2\u003eObjective\u003c/h2\u003e\u003cp\u003eTo address this gap, we conducted an exploratory study evaluating the quality of feedback produced by a commonly used LLM, ChatGPT, for a series of common learner performance issues.\u003c/p\u003e "},{"header":"Methods","content":"\u003cp\u003eFromme et al published a series of 21 feedback scripts generated by pediatric educators attending a national conference.(\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e) Feedback scripts addressed a variety of learner deficiency scenarios.\u003c/p\u003e \u003cp\u003eOur study team used three scripts from this series for training purposes and 18 for the study. ChatGPT (GPT-4o; OpenAI, San Francisco, CA; accessed May\u0026ndash;June 2025) was prompted to generate three to four sentences of written formative feedback to the learner for each scenario, responding as a pediatric faculty member. We compiled the ChatGPT-generated feedback alongside the original educator authored feedback scripts written by pediatric educators and published in the Fromme et al series.(\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e)\u003c/p\u003e \u003cp\u003eWe recruited six faculty at our institution (who were not involved in writing the scenarios or scripts) to rate the educator authored feedback scripts and the feedback generated by ChatGPT. The raters were blinded to whether the scripts were authored by educators or generated by ChatGPT. Raters included six medical education leaders from a single academic institution, including two current or prior clerkship directors, two residency program directors, and two fellowship program directors. Scripts were ordered randomly and identified to the raters only as \u0026ldquo;Script 1\u0026rdquo; or \u0026ldquo;Script 2.\u0026rdquo; Each script was rated using a behaviorally anchored 4-domain assessment tool (supplemental materials), modified from the works of Halman et al and Schlair et al to align with the format of the feedback script being assessed.(\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e) The revised tool assessed feedback scripts across domains established as important for effective feedback, specifically: (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) Behavior-centered, (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) Balanced between positive and negative feedback, (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e) Non-judgmental, and (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e) Focused. Each domain was scored using a 3-level rating scale (0\u0026ndash;2), where 0\u0026thinsp;=\u0026thinsp;No, 1\u0026thinsp;=\u0026thinsp;Somewhat, 2\u0026thinsp;=\u0026thinsp;Yes, yielding a minimum total score of 0 and a maximum total score of 2 for each domain and a minimum total score of 0 and a maximum total score of 8 for each script. Raters also indicated their overall script preference for each learner scenario\u003c/p\u003e \u003cp\u003eDescriptive statistics were applied to compare the ratings using T-tests in Excel (Microsoft 365, Seattle, WA). We also compared the mean scores of each scenario and the four domains.\u003c/p\u003e \u003cp\u003eThe Louisiana State University Health Sciences Center Institutional Review Board granted full exempt status for this study.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eThere was 100% completion by the raters. Across all four domains, ChatGPT-generated scripts received higher mean scores than educator authored scripts, as demonstrated in Fig.\u0026nbsp;1; behavior-centered domain (ChatGPT 1.90, SD 0.30 vs. educator 1.7, SD 0.53, P=.001 ), balanced domain (ChatGPT 1.87, SD 0.36 vs. educator 1.48 SD 0.73, P\u0026lt;.001), nonjudgmental domain (ChatGPT 1.87, SD 0.39 vs. educator 1.47, SD 0.39, P\u0026lt;.001), and focused domain (ChatGPT 1.91, SD 0.29 vs. educator 1.56, SD 0.29, P\u0026lt;.001). Mean scores of each domain are based on a 3-level rating scale (0\u0026thinsp;=\u0026thinsp;No, 1\u0026thinsp;=\u0026thinsp;Somewhat, 2\u0026thinsp;=\u0026thinsp;Yes) with a maximum score of 2.\u003c/p\u003e \u003cp\u003eAs shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, analyzing each scenario individually, ChatGPT-generated scripts were rated higher than educator authored scripts in 17 of the 18 scenarios, half being of statistical significance. An analysis of combined average scores revealed that ChatGPT-generated scripts were more highly rated than educator authored scripts with a statistical significance (ChatGPT 7.55, SD 1.13 vs. educator authored 6.23, SD 1.60, P\u0026lt;.05).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAverage Scores of ChatGPT and Educator Generated Narrative Scripts by Scenario\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eScore of ChatGPT-Generated Script\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eScore of Educator Authored Script\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eP value\u003c/em\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e7.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.073\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.243\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.046*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e6.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.164\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e7.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.651\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.016*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.076\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e6.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.871\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e7.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.144\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.034*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6.50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.046*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.033*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5.17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.026*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e7.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.438\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5.50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.0265*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5.50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.0303*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e7.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5.33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.0046*\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eScenario 18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e.091\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAVERAGE\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e7.55\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e6.23\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e\u0026lt;\u0026thinsp;.001*\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003eScript scores are the sum of four domain ratings, where 0\u0026thinsp;=\u0026thinsp;No, 1\u0026thinsp;=\u0026thinsp;Somewhat, and 2\u0026thinsp;=\u0026thinsp;Yes, yielding a maximum score of 8. The averages of the six raters' ChatGPT-generated and Educator authored script scores were compared.\u003c/p\u003e \u003cp\u003e*Identifies statistically significant results.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"1\" nameend=\"c5\" namest=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eChatGPT-generated feedback scripts were preferred over educator authored scripts in 84 of 108 ratings (78%). Only one educator authored script (number 5), addressing limited feedback receptivity of a learner was chosen by over half of the raters (83%).\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this exploratory study, feedback generated by a LLM was highly rated by medical education leaders and the content favorably compared to educator authored scripts. ChatGPT\u0026rsquo;s strong performance in this communication-focused task reflects the intentional training on large-scale conversational text that optimizes conversational fluency.(\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e) The overall preference of LLM-generated feedback suggests that generative AI may offer a viable, real-time tool to assist clinical educators in delivering high-quality feedback by offering support in communication, content, and tone consideration.\u003c/p\u003e \u003cp\u003eAI-generated responses in this study were rated highly by medical education leaders, noted to be well-balanced between positive and negative feedback, were appropriately focused on modifiable behaviors, and were non-judgmental in nature. This aligns with Ayers et al.\u0026rsquo;s study, which found that chatbot responses were significantly preferred over physician responses, rated higher in quality, and viewed as more empathetic.(\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e) Our findings suggest that AI may be useful for assisting faculty in quickly crafting sensitive, balanced feedback for learners across a variety of scenarios.\u003c/p\u003e \u003cp\u003eWhile we inputted a single, short, straightforward prompt into ChatGPT, future studies may determine whether more detailed prompts result in even higher-quality feedback scripts. It is worth investigating whether adding additional details, such as the level of the learner, prior evaluations, and exam scores, meaningfully improves the quality of the feedback generated by the LLM. Given that AI-generated feedback has resulted in better performance by medical students in the context of simulated surgical procedures, further studies should focus on the impact of feedback quality on clinical performance.(\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e)\u003c/p\u003e \u003cp\u003eLimitations of this single-center, pediatrics focused study include considerations of generalizability. While 18 learner deficiency scenarios were evaluated, they are not comprehensive and cannot capture all the nuances experienced in clinical settings. This study used pre-written scenarios; results may differ if faculty vary the amount or order of details inputted into the LLM. Decision fatigue of raters could have impacted results of this study. Notably, the perspective of learners receiving the feedback is missing from this study.\u003c/p\u003e \u003cp\u003eFuture work is needed to explore the experience of those receiving feedback to understand their perspective on the quality of the AI-generated feedback and to compare it to educator authored scripts. Furthermore, additional studies are needed to understand in detail when and why AI-generated scripts are preferred.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn this exploratory study, ChatGPT-generated feedback was rated as high quality by medical education leaders. While AI cannot replace educator judgment or clinical experience, LLMs show promise as a useful tool to support faculty in crafting timely, balanced, behavior-focused formative feedback in medical education.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cem\u003eAI\u003c/em\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eArtificial intelligence\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cem\u003eLLM\u003c/em\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eLarge language model\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cu\u003eEthics approval and consent to participate\u003c/u\u003e\u003c/p\u003e\n\u003cp\u003eThe Louisiana State University Health Sciences Center Institutional Review Board granted full exemption of this study.\u003c/p\u003e\n\u003cp\u003e\u003cu\u003eConsent for publication\u003c/u\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cu\u003eAvailability of data and materials\u003c/u\u003e\u003c/p\u003e\n\u003cp\u003eAll data generated or analyzed during this study are included in this published article [and its supplementary information files].\u003c/p\u003e\n\u003cp\u003e\u003cu\u003eCompeting interests\u003c/u\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests\u003c/p\u003e\n\u003cp\u003e\u003cu\u003eFunding\u003c/u\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThis study was conducted without any funding source.\u003c/p\u003e\n\u003cp\u003e\u003cu\u003eAuthors\u0026apos; contributions\u003c/u\u003e\u003c/p\u003e\n\u003cp\u003eB.D. conceived the study. \u0026nbsp;All authors contributed to the development of methodology and interpretation of results. E.S. was responsible for data collection and analysis. All authors participated in the writing of the manuscript and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cu\u003eAcknowledgements\u003c/u\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eFunctions and Structure of a Medical School Standards for Accreditation of Medical Education Programs Leading to the MD Degree [Internet]. 2025. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://lcme.org/publications/\u003c/span\u003e\u003cspan address=\"https://lcme.org/publications/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eACGME Common Program Requirements [Internet]. 2025. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.acgme.org/globalassets/pfassets/programrequirements/2025-reformatted-requirements/cprresidency_2025_reformatted.pdf\u003c/span\u003e\u003cspan address=\"https://www.acgme.org/globalassets/pfassets/programrequirements/2025-reformatted-requirements/cprresidency_2025_reformatted.pdf\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcCutcheon S, Duchemin AM. Overcoming barriers to effective feedback: a solution-focused faculty development approach. Int J Med Educ. 2020;11:230\u0026ndash;2. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.5116/ijme.5f7c.3157\u003c/span\u003e\u003cspan address=\"10.5116/ijme.5f7c.3157\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 33099519; PubMed Central PMCID: PMC7882126.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarak G, Foradori D, Fromme HB, Zuniga L, Dean A. Balancing Honest Assessment and Compassion for Learners Experiencing Burnout: A Workshop and Feedback Tool for Clinical Teachers. MedEdPORTAL. 2024;20:11449. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.15766/mep_2374-8265\u003c/span\u003e\u003cspan address=\"10.15766/mep_2374-8265\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. .11449 PubMed PMID: 39410923; PubMed Central PMCID: PMC11473647.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBroquet KE, Punwani M. Helping international medical graduates engage in effective feedback. Acad Psychiatry. 2012;36(4):282\u0026ndash;7. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1176/appi.ap.11020031\u003c/span\u003e\u003cspan address=\"10.1176/appi.ap.11020031\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 22851024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchlair S, Dyche L, Milan F. Longitudinal Faculty Development Program to Promote Effective Observation and Feedback Skills in Direct Clinical Observation. MedEdPORTAL. 2017;13:10648. 10. 15766/mep_2374-8265.10648 PubMed PMID: 30800849; PubMed Central PMCID: PMC6338150.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMinehart RD, Rudolph J, Pian-Smith MCM, Raemer DB. Improving faculty feedback to resident trainees during a simulated case: a randomized, controlled trial of an educational intervention. Anesthesiology. 2014;120(1):160\u0026ndash;71. 10. 1097/ALN.0000000000000058 PubMed PMID: 24398734.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGiving Effective Feedback: Beyond Great Job [Internet]. 2007. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.youtube.com/watch?v=DbfISZjG9mU\u003c/span\u003e\u003cspan address=\"https://www.youtube.com/watch?v=DbfISZjG9mU\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGigante J, Dell M, Sharkey A. Getting beyond Good job: how to give effective feedback. Pediatrics. 2011;127(2):205\u0026ndash;7. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1542/peds.2010-3351\u003c/span\u003e\u003cspan address=\"10.1542/peds.2010-3351\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 21242222.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYarris LM, Fu R, LaMantia J, Linden JA, Gene Hern H, Lefebvre C, et al. Effect of an educational intervention on faculty and resident satisfaction with real-time feedback in the emergency department. Acad Emerg Med. 2011;18(5):504\u0026ndash;12. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1111/j.1553-2712.2011.01055\u003c/span\u003e\u003cspan address=\"10.1111/j.1553-2712.2011.01055\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. x PubMed PMID: 21569169; PubMed Central PMCID: PMC3095955.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLockyer J, Armson H, K\u0026ouml;nings KD, Lee-Krueger RCW, des Ordons AR, Ramani S, et al. In-the-Moment Feedback and Coaching: Improving R2C2 for a New Context. J Grad Med Educ. 2020;12(1):27\u0026ndash;35. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.4300/JGME-D-19-00508.1\u003c/span\u003e\u003cspan address=\"10.4300/JGME-D-19-00508.1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 32089791; PubMed Central PMCID: PMC7012514.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFromme HB, Ryan MS, Gray K, Black E, Paik S, Griego E, et al. A Script for What Ails Your Learners: Feedback Scripts to Promote Effective Learning. Acad Pediatr. 2020;20(5):721\u0026ndash;3. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.acap.2020.02.005\u003c/span\u003e\u003cspan address=\"10.1016/j.acap.2020.02.005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 32044468.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFromme HB, Ryan MS, Gray K, Black E, Paik S, Griego E, et al. A Script for What Ails Your Learners: Feedback Scripts to Promote Effective Learning. Acad Pediatr. 2020;20(5):721\u0026ndash;3. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.acap.2020.02.005\u003c/span\u003e\u003cspan address=\"10.1016/j.acap.2020.02.005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 32044468.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKalyan KS. A survey of GPT-3 family large language models including ChatGPT and GPT-4. Nat Lang Process J. 2024;6:100048. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.nlp.2023.100048\u003c/span\u003e\u003cspan address=\"10.1016/j.nlp.2023.100048\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGiray L. Enhancing Communication with ChatGPT: A Guide for Academic Writers, Teachers, and Professionals. J Pract Cardiovasc Sci. 2024;10(2):113\u0026ndash;8. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.4103/jpcs.jpcs_27_24\u003c/span\u003e\u003cspan address=\"10.4103/jpcs.jpcs_27_24\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSrivastava R, Srivastava S. Can Artificial Intelligence aid communication? Considering the possibilities of GPT-3 in Palliative care. Indian J Palliat Care. 2023;29(4):418\u0026ndash;25. doi:10.25259/IJPC_155_2023 PubMed PMID: 38058478; PubMed Central PMCID: PMC10696352.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAyers JW, Poliak A, Dredze M, Leas EC, Zhu Z, Kelley JB, et al. Comparing Physician and Artificial Intelligence Chatbot Responses to Patient Questions Posted to a Public Social Media Forum. JAMA Intern Med. 2023;183(6):589\u0026ndash;96. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jamainternmed.2023.1838\u003c/span\u003e\u003cspan address=\"10.1001/jamainternmed.2023.1838\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 37115527; PubMed Central PMCID: PMC10148230.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDikici E, Bigelow M, Prevedello LM, White RD, Erdal BS. Integrating AI into radiology workflow: levels of research, production, and feedback maturity. J Med Imaging (Bellingham). 2020;7(1):016502. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1117/1.JMI.7.1.016502\u003c/span\u003e\u003cspan address=\"10.1117/1.JMI.7.1.016502\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 32064302; PubMed Central PMCID: PMC7012173.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHalman S, Dudek N, Wood T, Pugh D, Touchie C, McAleer S, et al. Direct Observation of Clinical Skills Feedback Scale: Development and Validity Evidence. Teach Learn Med. 2016;28(4):385\u0026ndash;94. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1080/10401334.2016.1186552\u003c/span\u003e\u003cspan address=\"10.1080/10401334.2016.1186552\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFazlollahi AM, Bakhaidar M, Alsayegh A, Yilmaz R, Winkler-Schwartz A, Mirchi N, et al. Effect of Artificial Intelligence Tutoring vs Expert Instruction on Learning Simulated Surgical Skills Among Medical Students: A Randomized Clinical Trial. JAMA Netw Open. 2022;5(2):e2149008. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jamanetworkopen.2021.49008\u003c/span\u003e\u003cspan address=\"10.1001/jamanetworkopen.2021.49008\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. PubMed PMID: 35191972; PubMed Central PMCID: PMC8864513.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-education","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"meed","sideBox":"Learn more about [BMC Medical Education](http://bmcmededuc.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/meed/default.aspx","title":"BMC Medical Education","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"feedback, artificial intelligence, large language models, ChatGPT","lastPublishedDoi":"10.21203/rs.3.rs-9076837/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9076837/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eFaculty feedback is vital for the professional growth of learners. Barriers such as faculty\u0026rsquo;s confidence in their feedback skills and time limitations may inhibit the provision of quality feedback. Generative large language models (LLMs), such as ChatGPT, have the potential to support faculty in providing high-quality corrective feedback by producing effective narrative feedback scripts for faculty to reference.\u003c/p\u003e\u003ch2\u003eObjective\u003c/h2\u003e \u003cp\u003eThis study explores the quality of Artificial Intelligence (AI)-generated text in the context of corrective feedback, based on standardized scenarios of learner deficiencies.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eSix medical education leaders blindly rated two narrative feedback scripts for eighteen learner deficiency scenarios; one script for each scenario was authored by educators, the other script was generated by ChatGPT. The raters scored each script in four domains (behavior-centric, balanced between positive and negative, non-judgmental, and focused) using a 3-point scale (0\u0026ndash;2), where 0\u0026thinsp;=\u0026thinsp;No, 1\u0026thinsp;=\u0026thinsp;Somewhat, 2\u0026thinsp;=\u0026thinsp;Yes. The mean scores of the educator authored scripts and the ChatGPT-generated scripts were compared by domain. An overall score was also generated for the educator authored script and the ChatGPT-generated scripts using descriptive statistics.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eAcross all four domains, ChatGPT-generated scripts received high mean scores, with all rated higher than the educator authored scripts. Also, ChatGPT-generated responses achieved higher overall scores than those authored by educators for 17 of 18 scenarios.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eAI-generated responses can be of high quality and may provide a useful tool for faculty who are preparing to give corrective feedback. Further studies are needed to explore the utility of AI in supporting feedback delivery.\u003c/p\u003e","manuscriptTitle":"Evaluating the Quality of AI-Generated Corrective Feedback for Medical Learners","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-17 21:15:14","doi":"10.21203/rs.3.rs-9076837/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-05-19T07:53:34+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-15T15:39:31+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"97828489935814465584816571926698163774","date":"2026-05-15T15:07:59+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-14T19:01:07+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"309417492838101721392339746445770851382","date":"2026-05-14T16:48:19+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-09T11:54:56+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"280015960562149467108557244653777916602","date":"2026-04-19T11:10:54+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"87529085876004613402471896183193216772","date":"2026-04-17T06:07:38+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-04-09T17:50:24+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-03-30T06:17:48+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-11T12:02:17+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-11T12:01:35+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Education","date":"2026-03-09T20:47:25+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-education","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"meed","sideBox":"Learn more about [BMC Medical Education](http://bmcmededuc.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/meed/default.aspx","title":"BMC Medical Education","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"401572b6-4924-4615-a55a-a52bd5026b6e","owner":[],"postedDate":"April 17th, 2026","published":true,"recentEditorialEvents":[{"type":"decision","content":"Revision requested","date":"2026-05-19T07:53:34+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-15T15:39:31+00:00","index":44,"fulltext":""},{"type":"reviewerAgreed","content":"97828489935814465584816571926698163774","date":"2026-05-15T15:07:59+00:00","index":43,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-14T19:01:07+00:00","index":41,"fulltext":""},{"type":"reviewerAgreed","content":"309417492838101721392339746445770851382","date":"2026-05-14T16:48:19+00:00","index":40,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-09T11:54:56+00:00","index":37,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[],"tags":[],"updatedAt":"2026-05-19T08:08:46+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-17 21:15:14","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9076837","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9076837","identity":"rs-9076837","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.