Exploring Generative Artificial Intelligence-Assisted Medical Education: Assessing Case-Based Learning for Medical Students.

OA: gold CC-BY-4.0

Abstract

The recent public release of generative artificial intelligence (GenAI) has brought fresh excitement by making access to GenAI for medical education easier than ever before. It is now incumbent upon both students and faculty to determine the optimal role of GenAI within the medical school curriculum. Given the promise and limitations of GenAI, this study aims to assess the current capabilities of a GenAI (Chat Generative Pre-trained Transformer, ChatGPT), specifically within the framework of a pre-clerkship case-based active learning curriculum. The role of GenAI is explored by evaluating its performance in generating educational materials, creating medical assessment questions, answering medical queries, and engaging in clinical reasoning by prompting it to respond to a problem-based learning scenario. Our results demonstrated that GenAI addressed epidemiology, diagnosis, and treatment questions well. However, there were still instances where it failed to provide comprehensive answers. Responses from GenAI might offer essential information, hint at the need for further inquiry, or sometimes omit critical details. GenAI struggled with generating information on complex topics, raising a significant concern when using it as a 'search engine' for medical student queries. This creates uncertainty for students regarding potentially missed critical information. With the increasing integration of GenAI into medical education, it is imperative for faculty to become well-versed in both its advantages and limitations. This awareness will enable them to educate students on using GenAI effectively in medical education.
Full text 21,237 characters · extracted from pmc-nxml · 5 sections · click to expand

Intro

Since the inception of artificial intelligence (AI), there has been interest in leveraging its potential within the field of medicine. Implementing the promise of AI has been pursued in nearly all areas of healthcare, with notable progress in many specialties, such as radiology and ophthalmology [ 1 , 2 ]. The field of medical education has similarly been interested in harnessing the power of AI to improve student learning and streamline administrative responsibilities [ 3 - 5 ]. The recent public release of generative AI (GenAI), Chat Generative Pre-trained Transformer (ChatGPT), has brought fresh excitement to this arena, making access to GenAI for medical education easier than ever. GenAI falls broadly in the category of machine learning and large language models but differs from other AI counterparts in that it produces new content in text, audio, images, simulations, and videos [ 4 , 6 ]. With these capabilities, it is understandable that there is both excitement and hesitation regarding applying GenAI in medical education. Focusing specifically on education within medical school, there are numerous ways in which GenAI can reshape current curriculum delivery and learning models and norms. While the rate of progress will not be homogenous across the following areas, GenAI is certain to impact educator content creation, student personalized learning, information gathering, and learner assessment [ 4 , 7 ]. For educator content creation, GenAI can help produce educational materials, exercises, and clinical cases based on provided prompts and information [ 8 ]. Students can utilize GenAI, albeit not originally designed for this purpose, as a sort of "search engine" to obtain answers to their specific questions with the additional benefit of analyzing their performance and adapting to their individual needs. Furthermore, GenAI can be used to develop and grade quizzes and other assessments, allowing educators to spend more time on other pursuits. While the potential benefits of GenAI are expansive, several limitations have arisen. One major area of concern is intellectual property rights [ 9 ]. The legal system is in the process of sorting out how the outcomes of GenAI fit within existing copyright law [ 10 ]. Additionally, school policies regarding plagiarism will be adapted due to the prevalence of GenAI [ 5 ]. Another noted limitation of GenAI is the bias of the algorithms and the bias of the information that feeds into the algorithm [ 11 , 12 ]. Ultimately, GenAI is currently still in its early phases and has documented instances of limited, misleading, and incorrect information. Hallucinations, for instance, are events in which GenAI extrapolates conclusions without sufficient supporting data [ 13 - 15 ]. Due to these limitations, many have advocated caution regarding using GenAI, and numerous institutions have made recommendations regarding its appropriate use. Given the promise and limitations of GenAI, this paper seeks to assess its current capabilities in the framework of a pre-clerkship case-based learning curriculum.

Results

Evaluation of GenAI in a medical school preclinic curriculum ChatGPT generated responses for all the prompts and selected diseases for each organ system. Regarding generating a study document, ChatGPT generated responses that typically involved seven sections: introduction, epidemiology, pathophysiology, clinical presentation, diagnosis, treatment, and conclusion. For generating test questions, ChatGPT was also able to generate questions for each selected disease. The test questions were formatted as multiple-choice questions with four possible answers labeled "a-d." In addition, the correct answer was listed below the question-and-answer choices. When prompted to answer a question on the epidemiology of the 11 selected diseases, ChatGPT was able to generate a response for each prompt. Although responses varied in their organization, they were generally divided into subsections about the following categories: introduction, prevalence, age and gender, age of onset, ethnicity, genetics, geography, risk factors, prognostic factors, comorbidities, mortality, complications, and conclusion. ChatGPT was able to generate a response for all questions asking about the diagnostic criteria for the 11 selected diseases. Responses did not follow a standard format but were tailored to the diagnostic criteria for each selected disease. A representative example of a response to the diagnostic criteria of melanoma is shown in Figure  1 . ChatGPT was also able to generate a response to the question asking about treatment for each of the 11 selected diseases. Similar to the responses for the diagnostic criteria prompt, responses for treatment were variable and tailored to the specific selected disease. A representative example of a response for treating endometriosis is shown in Figure  2 . Evaluation of GenAI for problem-based learning in the preclinic curriculum ChatGPT was able to generate a response to each question throughout the problem-based learning scenario. Additionally, it was able to generate responses that reflect the information that accumulates via the sequential reveal of the case. Prompts, questions, and responses generated for the problem-based learning scenario questions can be found in the Appendix (Tables 3 , 4 ). One example of a response to a question prompting discussion of a differential diagnosis based on initial symptoms is shown in Figure  3 . In addition to providing answers regarding differential diagnosis, ChatGPT was able to answer questions regarding gathering a patient's history, determining what to assess for on physical examination, selecting appropriate next steps in diagnosis and treatment, and describing the basic science behind disease and treatment. One limitation of ChatGPT in responding to this problem-based learning scenario is its inability to intake radiological images. This became of note when the scenario included a chest x-ray and a photograph of the culture results. Despite this limitation, the scenario included the chest X-ray, and culture descriptors that enabled ChatGPT to interpret the radiology and lab findings indirectly.

Discussion

Although still in its infancy, ChatGPT and other GenAIs are powerful tools with the potential to reshape medical education. Khan et al. outline seven roles for GenAI in medical education: automated scoring, teaching assistance, personalized learning, research assistance, quick access to information, generating case scenarios, and creating content to facilitate learning [ 17 ]. This provides an opportunity for both educators and students to streamline the educational process and gain back valuable time that can be spent on other academic, clinical, and research pursuits. Despite these advantages, the authors warn that GenAI, specifically referencing ChatGPT, is not without fault [ 17 ]. They note that ChatGPT can ignore context, leading to extraneous detail in the response. Additionally, it was noted that ChatGPT lacks more recent data, which is a concern given the rapid expansion of medical knowledge and clinically relevant changes that have occurred since [ 17 ]. The present study looked at the current ability of ChatGPT to act as a generator of study materials from which medical students can study, a "search engine" to answer questions from medical students, a generator of test questions for educators, and as an interactive tool to complement learning in case-based learning scenarios. Our study demonstrated that GenAI is currently not comprehensive enough to generate information to serve as a complete study guide for medical students. On the other hand, it did show tremendous promise in answering prompts regarding epidemiological, diagnostic, and treatment questions. GenAI also showed potential in developing test questions appropriate for medical student learning and preparation for board exams. Lastly, GenAI showed promise as an interactive tool to guide students through case-based learning scenarios. Given the findings of this study and the fact that GenAI is a new technology with the potential for significant improvement moving forward, students and faculty should explore ways to integrate GenAI into the learning process. Additionally, faculty should become well-versed in the advantages and shortcomings of GenAI to appropriately educate students on how it should be optimally used in medical education. A guideline set forth by the University of Southern California advises faculty members to educate themselves about GenAI proactively and to provide students with information about the technology's strengths and limitations [ 18 ]. In his commentary, Lee delves into both the potential benefits and shortcomings of ChatGPT in medical education [ 19 ]. He notes that GenAI can lead to improved efficiency, student engagement, and performance outcomes by providing opportunities for interactive simulations that educate and engage simultaneously. On the other hand, Lee notes concern over ethical concerns regarding mistakes produced by GenAI [ 19 ]. Generation of study materials Regarding being a generator of study materials, although ChatGPT provides accurate and relevant information, it is not as comprehensive as existing information stores commonly studied by medical students. Current and commonly used study resources offer more information for medical students to study. For instance, the study guide generated by ChatGPT for melanoma did not reference its association with the S-100 tumor marker, various forms of melanoma (superficial spreading, nodular, lentigo maligna, and acral lentiginous), and that patients with v-Raf murine sarcoma viral oncogene homolog B (BRAF) V600E mutations can stand to benefit from treatment with a BRAF kinase inhibitor. As such, students utilizing ChatGPT as a primary study resource may not be exposed to information critical to understanding the disease, disease diagnosis, and disease management. Furthermore, students would not be exposed to commonly tested information on board exams. Given the current limitations of GenAI in creating study guides, it is advisable not to rely on GenAI as the primary study resource. This is especially true when quality existing resources are available and contain more thorough and board-relevant information. In a study, Johnson et al. prompted ChatGPT to answer 180 questions submitted by 33 physicians representing 17 different specialties [ 20 ]. Among the answers, 53.5% were comprehensive, indicating that ChatGPT addressed all aspects of the questions and provided additional context beyond expectations [ 20 ]. However, just under half of the answers omitted significant portions or provided only the minimum required information to be considered complete. For a student seeking to grasp the nuances of clinical medicine, this level of missing content could present a reliability challenge and would discourage them from solely relying on GenAI responses. Clinical reasoning and answering medical queries Although ChatGPT did not generate comprehensive study guides for medical school, the GenAI addressed specific questions regarding epidemiology, diagnosis, and treatment better. For instance, taking the responses to multiple sclerosis as an example, GenAI correctly identified its association with age and gender, its complex genetics, and its significant impact on quality of life. Similarly, it outlined both appropriate diagnostic criteria and treatment options. Despite addressing these aspects, some key points were still not addressed. For instance, the GenAI did not list the exact treatment when referencing disease-modifying agents in managing multiple sclerosis. This leads to a key issue with using GenAI as a "search engine" to address questions in medical school: students do not know if they are missing out on critical pieces of information. This is especially noteworthy since different wordings of questions yield different answers. These answers may provide the necessary information, allude to the information requiring follow-up, or not allude to the information. The third category is dangerous since the student will not even know that he or she is missing out on the information. Another key issue is that GenAI does not adequately generate information on complex information, such as the genetic component of multiple sclerosis. These complex areas of science where a definitive answer is not yet known can result in AI delivering false information or result in AI hallucinations. Strong conducted an important experiment to assess ChatGPT's clinical reasoning abilities using clinical cases of varying difficulty levels [ 21 ]. He noted that its responses were comparable to those of first- or second-year medical students, but it missed critical details in complex cases involving multisystem conditions. There is some evidence that ChatGPT can clinically reason similarly to medical students and in-training residents on standardized examinations. Kung et al. evaluated the performance of ChatGPT on USMLE Step 1, Step 2, and Step 3 exams [ 22 ]. Through their analysis, the authors found that the GenAI could perform around the passing level for all three exams. As such, the authors concluded that there is potential for GenAI to teach medical students and assist them in passing these exams. Furthermore, the authors concluded that given the GenAI's performance on these exams, there is also potential for it to be used in a clinical setting. In a separate study, Li et al. utilized ChatGPT in virtual objective structured clinical examinations within the field of obstetrics and gynecology [ 23 ]. They observed that ChatGPT consistently generated factually accurate and contextually relevant responses to complex and dynamically evolving clinical inquiries. ChatGPT surpassed human candidates in various knowledge domains. In another study, Lum found that ChatGPT achieved scores in the 40th percentile when evaluated against first-year orthopedic surgery residents, in the eighth percentile for second-year residents, and the first percentile for third--, fourth--, and fifth-year residents [ 24 ]. These findings suggest that ChatGPT's testing performance is rather comparable to that of a first-year orthopedic surgery resident. The variation in results observed between these studies may be attributed to the specific fields under investigation, highlighting the necessity for further research in various domains within this context. Generating test questions In assessing the ability of GenAI to create test questions appropriate for medical students, it was found that ChatGPT had the ability to create high-quality multiple-choice exam questions. Many questions generated were brief clinical scenarios that required answerers to know not just one fact but several facts. This is similar to the questions that students encounter in national licensing exams and various course and clerkship shelf exams. As such, ChatGPT is a useful tool for faculty to create questions or for students to generate questions to test their knowledge. Although ChatGPT can generate high-quality, board-style questions, it also tends to create simpler questions relating to whether or not the answerer knows one specific fact. Overall, since test questions are central to the assessment of medical students throughout their preclinical and clinical curriculum, GenAI can potentially improve question generation and design. This is especially true for schools that rely entirely on faculty-generated questions or a hybrid of NBME and faculty questions. Others have employed ChatGPT to generate clinical case questions, noting its ability to produce plausible and well-structured questions within seconds; however, occasional inaccuracies in the content underscore the importance of subject matter expert review and verification [ 25 ]. Applicating to case-based learning The last aspect of this study involved assessing the potential impact of GenAI in assisting students in case-based learning. Over the last decade, more and more medical schools have adopted a case-based learning model that relies on faculty guiding students through clinical scenarios, a set of questions to be addressed, and defined learning objectives [ 26 ]. Given this model, it is conceivable that GenAI has the potential to serve as a personalized guide to students through these scenarios where it would adapt to a student's knowledge level. In assessing the role of GenAI in a problem-based learning scenario, it was noted that responses to questions and prompts included a significant amount of the key learning items outlined in the facilitator guide. For instance, when prompted to consider microorganisms that may cause pneumonia in a 12-year-old, GenAI correctly identified that mycoplasma pneumonia should be considered. Overall, this serves as a positive indicator that GenAI possesses the capability to correctly answer questions in a problem-based learning scenario. As a result, it is possible that it will soon have the ability to guide students through scenarios and identify knowledge gaps for students. However, much onus will be placed on the faculty and administration to properly integrate GenAI in case-based learning. In his study, Eysenbach highlights the potential for GenAI to create realistic patient scenarios and individualized learning experiences and to summarize studies for quickness and ease of understanding [ 27 ]. He explains how a conversation with someone with undiagnosed diabetes may proceed. From there, he continues to prompt the GenAI to provide additional information regarding a physician's conversations with a patient surrounding diabetes, including pathophysiology and management. He notes that ChatGPT is still limited and requires specific prompting that a medical student may not have the appropriate knowledge to ask. In another study, Buhr et al. evaluate ChatGPT's answers to otorhinolaryngology case-based questions compared to consultants' answers [ 28 ]. They noted that while ChatGPT provided longer answers, medical adequacy, and conciseness were significantly lower than consultants' answers. These limitations underscore the importance of recognizing that implementing ChatGPT in case-based learning depends on understanding its constraints and developing strategies to address them. Limitations and future directions This study assessed ChatGPT's performance in medical education from the authors' perspective. Future research should encompass medical students' perspectives and experiences while examining the short-term and long-term educational impacts of integrating GenAI into the curriculum through comparative study designs. As publicly available GenAI models become more prevalent, comparing ChatGPT against other GenAI models is imperative. Our study's difficulty level of questions and the clinical case were tailored to suit medical students. Caution should be exercised when extrapolating these findings to more complex teaching materials and cases with higher difficulty levels, such as those encountered by residents during training or practicing physicians.

Conclusions

Overall, GenAI has the potential to impact the medical school curriculum. While not yet comprehensive enough to serve as a repository of information, like a textbook, GenAI has the ability to answer questions, generate test questions, and appropriately respond to prompts in case-based learning scenarios. It performed well in addressing prompts related to epidemiology, diagnosis, and treatment but struggled to generate information on complex topics. It may fail to provide some key information, raising concerns about its suitability as a "search engine" for medical queries. This study offers insights into the current strengths of GenAI and its disadvantages. It can provide guidance to both students and faculty on how best to integrate GenAI, such as ChatGPT, into the current medical school paradigm.

Materials|Methods

A systematic approach was utilized to assess the capabilities of ChatGPT-4 in generating educational materials, creating medical assessment questions, answering medical questions, and reasoning through case-based learning scenarios. For educational materials, ChatGPT-4 was prompted to generate study guides for the medical conditions of each organ system (integumentary, skeletal, muscular, nervous, endocrine, cardiovascular, lymphatic, respiratory, digestive, urinary, and reproductive systems). For creating medical assessment questions, ChatGPT-4 was prompted to generate National Board of Medical Examiners (NBME)-)-style questions about the same medical conditions as for the summaries. For answering medical questions, specific questions were generated to assess Chat-GPT-4's ability to answer epidemiological, diagnostic, and treatment questions. The conditions selected and the precise language utilized are summarized in Tables  1 - 2 . ChatGPT was fed a representative case-based learning scenario from Edmunds et al. [ 16 ] and asked to address questions about the prompts to assess clinical reasoning within a case-based learning environment. A qualitative analysis of the performance of ChatGPT-4 with representative dialogue was performed.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-07-24T06:18:32.875334+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-4.0