Progression of an Artificial Intelligence Chatbot (ChatGPT) for Pediatric Cardiology Educational Knowledge Assessment: Considerable Gains in a Short Time

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Introduction: Artificial intelligence chatbots, like ChatGPT, have become powerful tools that are disrupting how humans interact with technology. The potential uses within medicine are vast. In medical education, these chatbots have shown improvements, in a short time span, in generalized medical examinations. We evaluated the overall performance and improvement between ChatGPT 3.5 and 4.0 in a test of pediatric cardiology knowledge. Methods: ChatGPT 3.5 and ChatGPT 4.0 were used to answer text based multiple choice questions derived from a Pediatric Cardiology Board Review textbook. Each chatbot, was given an 88 question test, subcategorized into 11 topics. We excluded questions with modalities other than text (sound clips or images). Statistical analysis was done using an unpaired two-tailed t-test. Results: Of the same 88 questions, ChatGPT 4.0 answered 66% of the questions correctly (n=58/88) which was significantly greater (p<0.0001) than ChatGPT 3.5, which only answered 38% (33/88). The ChatGPT 4.0 version also did better on each subspeciality topic as compared to ChatGPT 3.5. Conclusion: While acknowledging that ChatGPT does not yet offer subspecialty level knowledge in pediatric cardiology, the performance in pediatric cardiology educational assessments showed a considerable improvement in a short period of time between ChatGPT 3.5 to 4.0.
Full text 59,847 characters · extracted from preprint-html · click to expand
Progression of an Artificial Intelligence Chatbot (ChatGPT) for Pediatric Cardiology Educational Knowledge Assessment: Considerable Gains in a Short Time | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Progression of an Artificial Intelligence Chatbot (ChatGPT) for Pediatric Cardiology Educational Knowledge Assessment: Considerable Gains in a Short Time Michael N. Gritti, Hussain AlTurki, Pedrom Farid, Conall T. Morgan This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3360192/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 03 Jan, 2024 Read the published version in Pediatric Cardiology → Version 1 posted 10 You are reading this latest preprint version Abstract Introduction Artificial intelligence chatbots, like ChatGPT, have become powerful tools that are disrupting how humans interact with technology. The potential uses within medicine are vast. In medical education, these chatbots have shown improvements, in a short time span, in generalized medical examinations. We evaluated the overall performance and improvement between ChatGPT 3.5 and 4.0 in a test of pediatric cardiology knowledge. Methods ChatGPT 3.5 and ChatGPT 4.0 were used to answer text based multiple choice questions derived from a Pediatric Cardiology Board Review textbook. Each chatbot, was given an 88 question test, subcategorized into 11 topics. We excluded questions with modalities other than text (sound clips or images). Statistical analysis was done using an unpaired two-tailed t-test. Results Of the same 88 questions, ChatGPT 4.0 answered 66% of the questions correctly (n=58/88) which was significantly greater (p<0.0001) than ChatGPT 3.5, which only answered 38% (33/88). The ChatGPT 4.0 version also did better on each subspeciality topic as compared to ChatGPT 3.5. Conclusion While acknowledging that ChatGPT does not yet offer subspecialty level knowledge in pediatric cardiology, the performance in pediatric cardiology educational assessments showed a considerable improvement in a short period of time between ChatGPT 3.5 to 4.0. Artificial intelligence Natural language processing Pediatric Cardiology ChatGPT Education Figures Figure 1 Introduction Artificial intelligent (AI) chatbots have become a powerful technology changing how humans, in both a professional and personal sense, interact with digital interfaces. 1 Although the first AI chatbots were developed in the mid-20th century, only in the last decade, have the advancements in deep learning and neural networks allowed AI chatbots applicability to become more generalized to everyday life. 2 With the help of innovations like OpenAI's Generative Pre-Trained Transformer (GPT) architecture, the most advanced chatbots can now understand context, generate rational responses, and engage in lifelike conversations with humans. 3 Among these new chatbots, ChatGPT has recently garnered much attention due to its ability to improve user experiences, expedite processes, and spur innovation. 4 Now this technology has become accessible to the masses without specialized training. In healthcare, the potential uses for this software are vast and may revolutionize patient care, research, and medical education worldwide. ChatGPT and other natural language processing (NLP) models have already been used to help with various scientific research and medical endeavours, including writing medical abstracts 5 , conducting literature reviews 6 , helping with language barriers 7 , simplifying medical reports 8 , providing medical decision-making 9 , and helping with discharge summaries. 10,11 Recently, ChatGPT even passed the European Exam in Core Cardiology for specialty training in core adult cardiology. 12 In a short period of time, ChatGPT has already begun to upgrade its technology, with the application of ChatGPT 3.5 which was subsequently followed by ChatGPT 4.0 only a few months later. The newer model, ChatGPT 4.0 has shown improvement as compared to its predecessor in multiple domains, including the passing American Bar exam, 13 the USMLE 13,14 and more specialized medical disciplines like ophthalmology. 8,15 AI chatbots have the transformative opportunity to revolutionize patient care, medical diagnostics and improve healthcare dissemination but beforehand it must achieve a level of accuracy that can be relied upon consistently. Although the ChatGPT 4.0 version of the program has proven reliable for topics on generalized medical school training, to our knowledge, it’s performance has not yet been shown in specialized topics, such as pediatric cardiology. 16 Our study aims to assess the performance of ChatGPT in single best answer testing for pediatric cardiology, as well as compare the performance of ChatGPT 3.5 and ChatGPT 4.0 to assess the improvements over time. Methods We used a dataset of multiple choice questions derived from Pediatric Cardiology Board Review 17 by Eidem et al . with copyright permission from the publisher to test two AI chatbots. ChatGPT 3.5 and ChatGPT 4.0 were used to answer questions obtained from the question bank. The default modes for both ChatGPT 3.5 and ChatGPT 4.0 were used with ChatGPT Plus, and both chatbots were trained based on available data up to September 2021. As per the copyright permission agreement, only eight questions were allowed to be used per chapter. We arbitrarily used the first eight “text-only” questions from each of the 11 eligible chapters in the textbook for a total of 88 questions per test. We excluded questions that had associated images and other modalities such as sound, as ChatGPT 3.5 could only answer text based inputs. Each chapter contained eight questions regarding one of the following 11 pediatric cardiology topics: Cardiac Anatomy and Physiology , Congenital Cardiac Malformations , Diagnosis of Congenital Heart Disease , Cardiac Catheterization and Angiography , Non-invasive Cardiac Imaging , Electrophysiology Questions for Paediatrics , Exercise Physiology and Testing , Outpatient Cardiology , Cardiac Intensive Care and Heart Failure , Cardiac Pharmacology and Surgical Palliation and Repair of Congenital Heart Disease . Questions from other chapters were excluded if they were felt to be not specific to the specialization of pediatric cardiology, for example chapters on statistics. Each question was entered as a separate new prompt, and previous conversations were cleared to ensure no previous information affected the chatbot's answers. Each question was entered into ChatGPT 3.5 and ChatGPT 4.0 using the exact same wording. Each answer was then manually reviewed by members of our team (M.N.G and H.A.) to ensure that the chatbots had answered the question. If ChatGPT deemed that none, multiple or all the answers were correct, when this was not one of the multiple choice options, it was scored as incorrect. The answer key provided by the textbook was used to verify correct answers and both ChatGPT 3.5 and ChatGPT 4.0 were given a score out of 8 for each question they correctly answered in each chapter. The statistical analysis was done using an unpaired two-tailed t-test, with a p-value less than 0.05 being statistically significant. Due to the nature of this study, no research ethics approval from our institution was obtained. Results ChatGPT 3.5 and ChatGPT 4.0 were used to answer 88 text based multiple choice questions from the Pediatric Cardiology Board Review 17 textbook. ChatGPT 4.0 answered 58 questions correctly (65.9%), which was significantly greater than ChatGPT 3.5, which only answered 33 questions correctly (37.5%), p<0.0001. (Figure 1) ChatGPT 4.0 When broken down by each chapter and topic, ChatGPT 4.0 performed better or equal in each section compared to ChatGPT 3.5. ChatGPT 4.0 performed best on questions regarding Non-invasive Cardiac Imaging and Surgical Palliation and Repair of Congenital Heart Disease , correctly answering 82.5% (7/8) of the questions in each chapter. ChatGPT 4.0 scored lowest on Cardiac Catheterization and Angiography and Diagnosis of Congenital Heart Disease , scoring 37.5 % (3/8) and 50% (4/8) on these sections, respectively. (Table 1) ChatGPT 3.5 Based on the data, ChatGPT 3.5 highest grade on any topic was only 50%, which it scored across a range of topics, including Invasive Cardiac Imaging, Surgical Palliation and Repair of Congenital Heart Disease, Exercise Physiology and Testing, Cardiac Intensive Care, and Heart Failure . ChatGPT 3.5 scored lowest in Electrophysiology for Pediatrics , correctly answering only one question (1/8, 12.5%). This breakdown can be found in Table 1. Although ChatGPT 4.0 did better or equal in every section of the test compared to ChatGPT 3.5, ChatGPT 3.5 did answer three questions from different chapters correctly that ChatGPT 4.0 ultimately answered incorrectly. As this was out of keeping with the general trend of the results and to confirm that there was not a software malfunction, the three questions were re-entered twice more into both versions of ChatGPT and again showed the same results. Discussion AI chatbots such as ChatGPT 3.5 and ChatGPT 4.0 are novel NLP software’s with more advanced capabilities than previous iterations and are learning rapidly. 16 Within just one generation, released four months apart, ChatGPT 4.0 has already improved immensely compared to ChatGPT 3.5. As demonstrated in our study, when answering questions regarding various pediatric cardiology topics, ChatGPT 4.0 performed better or equal to ChatGPT3.5 on every topic and overall the score increased from a 38% to 66% score on our standardized test. This supports the findings by Michalek and colleagues that ChatGPT 4.0 had similar performance gains when dealing with a more specialized medical field. 8,15 Our study shows that, despite improvements, there are still substantial gains required in pediatric cardiology educational testing, as the newest AI chatbot still got 1/3 of the questions incorrect. This performance is substantially worse than what was seen by other researchers testing ChatGPT 4.0 on the USMLE Step 1 and Step 2, where it scored in the 90 th percentile, equivalent to an average score of 80-90%. 13 We believe this highlights the difficulties that the chatbots will have with decerning more complex medical topics than simpler general medical topics. One of the reasons for this difference in performance may be because of the goal of the tests. The USMLE is testing for a basic understanding of many topics whereas board exams in medical sub-specialties are assessing more nuanced understanding. The former gives ChatGPT an advantage because of its large dataset where it can quickly answer questions on generalized topics that undergraduate medical students try to memorize. Should one assume that the subject matter in pediatric cardiology is more nuanced, then one could postulate that if ChatGPT would have a more nuanced dataset, it may have faired much better. This the need for more software training in more specific fields of medicine. One noteworthy observation over the course of conducting the study, was that both versions of ChatGPT had poor insight into their ability to answer questions. For the questions that were answered incorrectly, the chatbot would still offer explanations that despite being incorrect would sound plausible. This phenomenon has been termed “hallucination,” and has been noted in other studies related to ChatGPT and other chatbots. 18 This creates a significant practical barrier to implantation as ChatGPT may not necessarily self-identify inaccuracies and would require that the content it generates be thoroughly reviewed. In the era of misinformation and falsehoods, this is the main drawback with using this type of technology for the general public when dealing with nuanced medical information. It is also interesting that there were 3 questions, all in different topics, which ChatGPT3.5 answered correctly while ChatGPT 4.0 did not, and this was confirmed on repeat testing for these questions. As such it is unlikely to be due to a software malfunction. One possibility for this, is that there is a difference between the two versions in how it processes the input, leading to the different answers, similar to how it has been shown that slight variations to a question may lead to different responses. 18 However, due to both the highly complex and proprietary nature of the program it is currently difficult to tell exactly why this phenomenon has occurred for these 3 questions. Over time, a trend may become apparent as more studies test its performance. This study had a few limitations that may have affected the generalizability of the study. One of the limitations is that the exam was created for this study. As a result, we need to determine how medical practitioners would have faired on the exam to grasp how well ChatGPT 4.0 indeed did on the examination. Therefore, our results can only be used to show the advancements made from ChatGPT 4.0 compared to ChatGPT 3.5 rather than how generalizable it is. There is also insufficient data on how well performance on these questions would translate to a more practical setting. Another limitation of the study is that we could not use any non-text modalities (like sound or images). A large part of forming a clinical picture in cardiology is being able to listen to sounds, interpret echocardiographic images and read electrocardiograms (ECGs). Trainees would be required and taught to have such understanding. While assessing the ability of ChatGPT 4.0 to interpret these images was outside the scope of our study, there is evidence that AI with a transformer based architecture, similar to ChatGPT, showed promising results in ECG interpretation. 19 There is also promising results with regards to use of AI in interpreting echocardiographic images. 20 However, these studies notably used AI models that were specifically trained with sample images. In theory, if ChatGPT or other transformer-based AI chatbots were similarly trained they may be able to show similar results if these modalities were added to standardized exams. Lastly, we only looked at multiple-choice answers, meaning the chatbot had a 20% chance of being correct. A long answer test format would have helped test the chatbots on a different aspect of knowledge and may be a source of future study. Conclusion In conclusion, ChatGPT 4.0 performed significantly better than ChatGPT 3.5 when tasked with answering specialized medical questions regarding pediatric cardiology. However, it still fell short of an acceptable grade. Although ChatGPT 4.0 is far more advanced than its predecessor, it still needs more training in pediatric cardiology and other specialized fields before we can trustily rely on it. Future research will be needed perpetually assess the accuracy of AI chatbots in this field of medicine, evaluate predictors of poor accuracy in the models, and both monitor the performance as well as utilization of AI chatbots within pediatric cardiology. Abbreviations AI: Artificial intelligence GPT: Generative Pre-Trained Transformer NLP: Natural language processing ECG: Electrocardiogram Declarations Financial Support: None. Source of Funding: None. Meeting Presentation: None. Conflict of Interest: No conflicting relationship exists for any author. Acknowledgements: None. References Liu PR, Lu L, Zhang JY, Huo TT, Liu SX, Ye ZW. Application of Artificial Intelligence in Medicine: An Overview. Curr Med Sci. 2021;41(6):1105-1115. Bart NK, Pepe S, Gregory AT, Denniss AR. Emerging Roles of Artificial Intelligence (AI) in Cardiology: Benefits and Barriers in a 'Brave New World'. Heart Lung Circ. 2023;32(8):883-888. Dave T, Athaluri SA, Singh S. ChatGPT in medicine: an overview of its applications, advantages, limitations, future prospects, and ethical considerations. Front Artif Intell. 2023;6:1169595. Sardana D, Fagan TR, Wright JT. ChatGPT: A disruptive innovation or disrupting innovation in academia? J Am Dent Assoc. 2023;154(5):361-364. Gao CA, Howard FM, Markov NS, et al. Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. NPJ Digit Med. 2023;6(1):75. Ruksakulpiwat S, Kumar A, Ajibade A. Using ChatGPT in Medical Research: Current Status and Future Directions. J Multidiscip Healthc. 2023;16:1513-1520. Teixeira da Silva JA. Can ChatGPT rescue or assist with language barriers in healthcare communication? Patient Educ Couns. 2023;115:107940. Mihalache A, Popovic MM, Muni RH. Performance of an Artificial Intelligence Chatbot in Ophthalmic Knowledge Assessment. JAMA Ophthalmol. 2023;141(6):589-597. Rao A, Kim J, Kamineni M, Pang M, Lie W, Succi MD. Evaluating ChatGPT as an Adjunct for Radiologic Decision-Making. medRxiv. 2023. Patel SB, Lam K. ChatGPT: the future of discharge summaries? Lancet Digit Health. 2023;5(3):e107-e108. Singh S, Djalilian A, Ali MJ. ChatGPT and Ophthalmology: Exploring Its Potential with Discharge Summaries and Operative Notes. Semin Ophthalmol. 2023;38(5):503-507. Skalidis I, Cagnina A, Luangphiphat W, et al. ChatGPT takes on the European Exam in Core Cardiology: an artificial intelligence success story? Eur Heart J Digit Health. 2023;4(3):279-281. Gibson A. How Does ChatGPT Perform on the Medical Licensing Exams? The Implications of Large Language Models for Medical Education and Knowledge Assessment. medRxiv. 2022. Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digit Health. 2023;2(2):e0000198. Mihalache A, Huang RS, Popovic MM, Muni RH. Performance of an Upgraded Artificial Intelligence Chatbot for Ophthalmic Knowledge Assessment. JAMA Ophthalmol. 2023;141(8):798-800. Waisberg E, Ong J, Masalkhi M, et al. GPT-4: a new era of artificial intelligence in medicine. Ir J Med Sci. 2023. Eidem B. Pediatric Cardiology Board Review. 2nd ed. Philadelphia, PA: Wolters Kluwer; 2023. Sallam M. ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns. Healthcare (Basel). 2023;11(6). Vaid A, Jiang J, Sawant A, et al. A foundational vision transformer improves diagnostic performance for electrocardiograms. NPJ Digit Med. 2023;6(1):108. He B, Kwan AC, Cho JH, et al. Blinded, randomized trial of sonographer versus AI cardiac function assessment. Nature. 2023;616(7957):520-524. Table Table 1. Number of correctly answered questions (out of 88) by ChatGPT 3.5 and ChatGPT 4.0 in each of the 11 chapters of the Pediatric Cardiology Board Review book. 17 ChatGPT 3.5 ChatGPT 4.0 Chapter 1: Cardiac Anatomy and Physiology 2 5 Chapter 2: Congenital Cardiac Malformations 2 5 Chapter 3: Diagnosis of Congenital Heart Disease 3 4 Chapter 4: Cardiac Catheterization and Angiography 3 3 Chapter 5: Non-invasive Cardiac Imaging 4 7 Chapter 6: Electrophysiology Questions for Paediatrics 1 5 Chapter 7: Exercise Physiology and Testing 4 6 Chapter 8: Outpatient Cardiology 3 5 Chapter 9: Cardiac Intensive Care and Heart Failure 4 5 Chapter 10: Cardiac Pharmacology 3 6 Chapter 11: Surgical Palliation and Repair of Congenital Heart Disease 4 7 Total 33 58 Score on entire examination 37.5% 65.9% Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 03 Jan, 2024 Read the published version in Pediatric Cardiology → Version 1 posted Editorial decision: Revision requested 26 Nov, 2023 Reviews received at journal 03 Oct, 2023 Reviewers agreed at journal 24 Sep, 2023 Reviewers agreed at journal 23 Sep, 2023 Reviewers agreed at journal 21 Sep, 2023 Reviewers agreed at journal 20 Sep, 2023 Reviewers invited by journal 20 Sep, 2023 Editor assigned by journal 16 Sep, 2023 Submission checks completed at journal 16 Sep, 2023 First submitted to journal 15 Sep, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3360192","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":233999830,"identity":"ed8c96e4-8164-4ebf-8c50-b19c52701a0c","order_by":0,"name":"Michael N. Gritti","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABEUlEQVRIie2PMUsDMRiGvxLILam3fkel9wuEFKE6FO+HuOQ46OxUOshxIuSWA1d/TiBw03FdhS4HQieHjh1UTFrtII11FMwDCeTje3jzAng8fxEC1NyX5vQKAD4xT0pA0KMKWuWuUHz6CwX2iolRoO2A7IYOzkqy6jZzhPhal8/rm0VyEtQZ72Y5hKU6qIw1vRhVDcKoTe3HlqlkUy1EqwEb4VAYxb40SrXtshQUg3uVSgUc3Er09r5X2uRTyYGHnVMZ9AvThW0V1ZNItUglAY6uFDoenNbIuFEeG57ZLhk3XRg+OVIWehW93E6GcRV06/nrVRKX9Xm0meXD8OFwyheMq++TH/ctcXF0xePxeP4rHynoV4j3XJ05AAAAAElFTkSuQmCC","orcid":"","institution":"The Hospital for Sick Children","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Michael","middleName":"N.","lastName":"Gritti","suffix":""},{"id":233999831,"identity":"258d3a6c-ba7a-4b32-b64d-442c51a565b9","order_by":1,"name":"Hussain AlTurki","email":"","orcid":"","institution":"The Hospital for Sick Children","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hussain","middleName":"","lastName":"AlTurki","suffix":""},{"id":233999832,"identity":"921c6a38-58f5-42ac-8cd8-e610ee95488c","order_by":2,"name":"Pedrom Farid","email":"","orcid":"","institution":"University of Western Ontario","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Pedrom","middleName":"","lastName":"Farid","suffix":""},{"id":233999833,"identity":"a8440d0c-0924-44aa-bc69-c2919cc49f1c","order_by":3,"name":"Conall T. Morgan","email":"","orcid":"","institution":"The Hospital for Sick Children","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Conall","middleName":"T.","lastName":"Morgan","suffix":""}],"badges":[],"createdAt":"2023-09-16 02:29:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3360192/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3360192/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s00246-023-03385-6","type":"published","date":"2024-01-03T15:00:37+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":43549937,"identity":"518a2d92-5355-4053-8590-3eccee0fa29a","added_by":"auto","created_at":"2023-09-22 18:36:00","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":95745,"visible":true,"origin":"","legend":"\u003cp\u003eChat GPT 3.5 and ChatGPT 4.0 were each asked the same eight text based multiple choice questions from each of 11 chapters in the \u003cem\u003ePediatric Cardiology Board Review\u003c/em\u003e.\u003csup\u003e17\u003c/sup\u003e Data shown is the mean percentage of correctly answered questions ± standard deviation by ChatGPT 3.5 (n=88) and ChatGPT 4.0 (n=88). Data were analyzed with an unpaired t-test and showed a significant difference between groups with a p-value \u0026lt;0.05.\u003c/p\u003e","description":"","filename":"Fig1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3360192/v1/5fb668acf9589b9202107e94.jpg"},{"id":49315844,"identity":"2e4676bb-060c-46ef-b0c7-affa3cf34cd4","added_by":"auto","created_at":"2024-01-08 15:10:50","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":263228,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3360192/v1/a2d05f4a-8e2a-46a6-8c27-ddd83a8d5f36.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Progression of an Artificial Intelligence Chatbot (ChatGPT) for Pediatric Cardiology Educational Knowledge Assessment: Considerable Gains in a Short Time","fulltext":[{"header":"Introduction","content":"\u003cp\u003eArtificial intelligent (AI) chatbots have become a powerful technology changing how humans, in both a professional and personal sense, interact with digital interfaces.\u003csup\u003e1\u003c/sup\u003e Although the first AI chatbots were developed in the mid-20th century, only in the last decade, have the advancements in deep learning and neural networks allowed AI chatbots applicability to become more generalized to everyday life.\u003csup\u003e2\u003c/sup\u003e With the help of innovations like OpenAI\u0026apos;s Generative Pre-Trained Transformer (GPT) architecture, the most advanced chatbots can now understand context, generate rational responses, and engage in lifelike conversations with humans.\u003csup\u003e3\u003c/sup\u003e Among these new chatbots, ChatGPT has recently garnered much attention due to its ability to improve user experiences, expedite processes, and spur innovation.\u003csup\u003e4\u003c/sup\u003e Now this technology has become accessible to the masses without specialized training.\u003c/p\u003e\n\u003cp\u003eIn healthcare, the potential uses for this software are vast and may revolutionize patient care, research, and medical education worldwide. ChatGPT and other natural\u0026nbsp;language processing (NLP) models have already been used to help with various scientific research and medical endeavours, including writing medical abstracts\u003csup\u003e5\u003c/sup\u003e, conducting literature reviews\u003csup\u003e6\u003c/sup\u003e, helping with language barriers\u003csup\u003e7\u003c/sup\u003e, simplifying medical reports\u003csup\u003e8\u003c/sup\u003e, providing medical decision-making\u003csup\u003e9\u003c/sup\u003e, and helping with discharge summaries.\u003csup\u003e10,11\u003c/sup\u003e Recently, ChatGPT even passed the European Exam in Core Cardiology for specialty training in core adult cardiology.\u003csup\u003e12\u003c/sup\u003e In a short period of time, ChatGPT has already begun to upgrade its technology, with the application of ChatGPT 3.5 which was subsequently followed by ChatGPT 4.0 only a few months later. The newer model, ChatGPT 4.0 has shown improvement as compared to its predecessor in multiple domains, including the passing American Bar exam,\u003csup\u003e13\u003c/sup\u003e the USMLE\u003csup\u003e13,14\u003c/sup\u003e and more specialized medical disciplines like ophthalmology.\u003csup\u003e8,15\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003eAI chatbots have the transformative opportunity to revolutionize patient care, medical diagnostics and improve healthcare dissemination but beforehand it must achieve a level of accuracy that can be relied upon consistently. Although the ChatGPT 4.0 version of the program has proven reliable for topics on generalized medical school training, to our knowledge, it\u0026rsquo;s performance has not yet been shown in specialized topics, such as pediatric cardiology.\u003csup\u003e16\u003c/sup\u003e Our study aims to assess the performance of ChatGPT in single best answer testing for pediatric cardiology, as well as compare the performance of ChatGPT 3.5 and ChatGPT 4.0 to assess the improvements over time.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eWe used a dataset of multiple choice questions derived from \u003cem\u003ePediatric Cardiology Board Review\u003c/em\u003e\u003csup\u003e17\u003c/sup\u003e by Eidem \u003cem\u003eet al\u003c/em\u003e. with copyright permission from the publisher to test two AI chatbots. ChatGPT 3.5 and ChatGPT 4.0 were used to answer questions obtained from the question bank. The default modes for both ChatGPT 3.5 and ChatGPT 4.0 were used with ChatGPT Plus, and both chatbots were trained based on available data up to September 2021. As per the copyright permission agreement, only eight questions were allowed to be used per chapter. We arbitrarily used the first eight \u0026ldquo;text-only\u0026rdquo; questions from each of the 11 eligible chapters in the textbook for a total of 88 questions per test. We excluded questions that had associated images and other modalities such as sound, as ChatGPT 3.5 could only answer text based inputs.\u003c/p\u003e\n\u003cp\u003eEach chapter contained eight questions regarding one of the following 11 pediatric cardiology topics: \u003cem\u003eCardiac Anatomy and Physiology\u003c/em\u003e, \u003cem\u003eCongenital Cardiac Malformations\u003c/em\u003e, \u003cem\u003eDiagnosis of Congenital Heart Disease\u003c/em\u003e, \u003cem\u003eCardiac Catheterization and Angiography\u003c/em\u003e, \u003cem\u003eNon-invasive Cardiac Imaging\u003c/em\u003e, \u003cem\u003eElectrophysiology Questions for Paediatrics\u003c/em\u003e, \u003cem\u003eExercise Physiology and Testing\u003c/em\u003e, \u003cem\u003eOutpatient Cardiology\u003c/em\u003e, \u003cem\u003eCardiac Intensive Care and Heart Failure\u003c/em\u003e, \u003cem\u003eCardiac Pharmacology\u003c/em\u003e and \u003cem\u003eSurgical Palliation and Repair of Congenital Heart Disease\u003c/em\u003e. Questions from other chapters were excluded if they were felt to be not specific to the specialization of pediatric cardiology, for example chapters on statistics.\u003c/p\u003e\n\u003cp\u003eEach question was entered as a separate new prompt, and previous conversations were cleared to ensure no previous information affected the chatbot\u0026apos;s answers. Each question was entered into ChatGPT 3.5 and ChatGPT 4.0 using the exact same wording. Each answer was then manually reviewed by members of our team (M.N.G and H.A.) to ensure that the chatbots had answered the question. If ChatGPT deemed that none, multiple or all the answers were correct, when this was not one of the multiple choice options, it was scored as incorrect. The answer key provided by the textbook was used to verify correct answers and both ChatGPT 3.5 and ChatGPT 4.0 were given a score out of 8 for each question they correctly answered in each chapter. The statistical analysis was done using an unpaired two-tailed t-test, with a p-value less than 0.05 being statistically significant. Due to the nature of this study, no research ethics approval from our institution was obtained.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eChatGPT 3.5 and ChatGPT 4.0 were used to answer 88 text based multiple choice questions from the \u003cem\u003ePediatric Cardiology Board Review\u003c/em\u003e\u003cem\u003e\u003csup\u003e17\u003c/sup\u003e\u003c/em\u003e\u003cem\u003e\u0026nbsp;\u003c/em\u003etextbook. ChatGPT 4.0 answered 58 questions correctly (65.9%), which was significantly greater than ChatGPT 3.5, which only answered 33 questions correctly (37.5%), p\u0026lt;0.0001. (Figure 1)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eChatGPT 4.0\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eWhen broken down by each chapter and topic, ChatGPT 4.0 performed better or equal in each section compared to ChatGPT 3.5. ChatGPT 4.0 performed best on questions regarding \u003cem\u003eNon-invasive Cardiac Imaging and Surgical Palliation\u003c/em\u003e and \u003cem\u003eRepair of Congenital Heart Disease\u003c/em\u003e, correctly answering 82.5% (7/8) of the questions in each chapter. ChatGPT 4.0 scored lowest on \u003cem\u003eCardiac Catheterization and Angiography\u003c/em\u003e and \u003cem\u003eDiagnosis of Congenital Heart Disease\u003c/em\u003e, scoring 37.5 % (3/8) and 50% (4/8) on these sections, respectively. (Table 1)\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eChatGPT 3.5\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003eBased on the data, ChatGPT 3.5 highest grade on any topic was only 50%, which it scored across a range of topics, including \u003cem\u003eInvasive Cardiac Imaging, Surgical Palliation and Repair of Congenital Heart Disease, Exercise Physiology and Testing, Cardiac Intensive Care, and Heart Failure\u003c/em\u003e. ChatGPT 3.5 scored lowest in \u003cem\u003eElectrophysiology for Pediatrics\u003c/em\u003e, correctly answering only one question (1/8, 12.5%). This breakdown can be found in Table 1. Although ChatGPT 4.0 did better or equal in every section of the test compared to ChatGPT 3.5, ChatGPT 3.5 did answer three questions from different chapters correctly that ChatGPT 4.0 ultimately answered incorrectly. As this was out of keeping with the general trend of the results and to confirm that there was not a software malfunction, the three questions were re-entered twice more into both versions of ChatGPT and again showed the same results.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eAI chatbots such as ChatGPT 3.5 and ChatGPT 4.0 are novel NLP software\u0026rsquo;s with more advanced capabilities than previous iterations and are learning rapidly.\u003csup\u003e16\u003c/sup\u003e Within just one generation, released four months apart, ChatGPT 4.0 has already improved immensely compared to ChatGPT 3.5. As demonstrated in our study, when answering questions regarding various pediatric cardiology topics, ChatGPT 4.0 performed better or equal to ChatGPT3.5 on every topic and overall the score increased from a 38% to 66% score on our standardized test. This supports the findings by Michalek and colleagues that ChatGPT 4.0 had similar performance gains when dealing with a more specialized medical field.\u003csup\u003e8,15\u003c/sup\u003e\u003c/p\u003e\n\u003cp\u003eOur study shows that, despite improvements, there are still substantial gains required in pediatric cardiology educational testing, as the newest AI chatbot still got 1/3 of the questions incorrect. This performance is substantially worse than what was seen by other researchers testing ChatGPT 4.0 on the USMLE Step 1 and Step 2, where it scored in the 90\u003csup\u003eth \u003c/sup\u003epercentile, equivalent to an average score of 80-90%.\u003csup\u003e13\u003c/sup\u003e We believe this highlights the difficulties that the chatbots will have with decerning more complex medical topics than simpler general medical topics. \u003c/p\u003e\n\u003cp\u003eOne of the reasons for this difference in performance may be because of the goal of the tests. The USMLE is testing for a basic understanding of many topics whereas board exams in medical sub-specialties are assessing more nuanced understanding. The former gives ChatGPT an advantage because of its large dataset where it can quickly answer questions on generalized topics that undergraduate medical students try to memorize. Should one assume that the subject matter in pediatric cardiology is more nuanced, then one could postulate that if ChatGPT would have a more nuanced dataset, it may have faired much better. This the need for more software training in more specific fields of medicine.\u003c/p\u003e\n\u003cp\u003eOne noteworthy observation over the course of conducting the study, was that both versions of ChatGPT had poor insight into their ability to answer questions. For the questions that were answered incorrectly, the chatbot would still offer explanations that despite being incorrect would sound plausible. This phenomenon has been termed \u0026ldquo;hallucination,\u0026rdquo; and has been noted in other studies related to ChatGPT and other chatbots.\u003csup\u003e18\u003c/sup\u003e This creates a significant practical barrier to implantation as ChatGPT may not necessarily self-identify inaccuracies and would require that the content it generates be thoroughly reviewed. In the era of misinformation and falsehoods, this is the main drawback with using this type of technology for the general public when dealing with nuanced medical information. \u003c/p\u003e\n\u003cp\u003eIt is also interesting that there were 3 questions, all in different topics, which ChatGPT3.5 answered correctly while ChatGPT 4.0 did not, and this was confirmed on repeat testing for these questions. As such it is unlikely to be due to a software malfunction. One possibility for this, is that there is a difference between the two versions in how it processes the input, leading to the different answers, similar to how it has been shown that slight variations to a question may lead to different responses.\u003csup\u003e18\u003c/sup\u003e However, due to both the highly complex and proprietary nature of the program it is currently difficult to tell exactly why this phenomenon has occurred for these 3 questions. Over time, a trend may become apparent as more studies test its performance.\u003c/p\u003e\n\u003cp\u003eThis study had a few limitations that may have affected the generalizability of the study. One of the limitations is that the exam was created for this study. As a result, we need to determine how medical practitioners would have faired on the exam to grasp how well ChatGPT 4.0 indeed did on the examination. Therefore, our results can only be used to show the advancements made from ChatGPT 4.0 compared to ChatGPT 3.5 rather than how generalizable it is. There is also insufficient data on how well performance on these questions would translate to a more practical setting. Another limitation of the study is that we could not use any non-text modalities (like sound or images). A large part of forming a clinical picture in cardiology is being able to listen to sounds, interpret echocardiographic images and read electrocardiograms (ECGs). Trainees would be required and taught to have such understanding. While assessing the ability of ChatGPT 4.0 to interpret these images was outside the scope of our study, there is evidence that AI with a transformer based architecture, similar to ChatGPT, showed promising results in ECG interpretation.\u003csup\u003e19\u003c/sup\u003e\u003csup\u003e \u003c/sup\u003eThere is also promising results with regards to use of AI in interpreting echocardiographic images.\u003csup\u003e20\u003c/sup\u003e\u003csup\u003e \u003c/sup\u003eHowever, these studies notably used AI models that were specifically trained with sample images. In theory, if ChatGPT or other transformer-based AI chatbots were similarly trained they may be able to show similar results if these modalities were added to standardized exams. Lastly, we only looked at multiple-choice answers, meaning the chatbot had a 20% chance of being correct. A long answer test format would have helped test the chatbots on a different aspect of knowledge and may be a source of future study.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn conclusion, ChatGPT 4.0 performed significantly better than ChatGPT 3.5 when tasked with answering specialized medical questions regarding pediatric cardiology. However, it still fell short of an acceptable grade. Although ChatGPT 4.0 is far more advanced than its predecessor, it still needs more training in pediatric cardiology and other specialized fields before we can trustily rely on it. Future research will be needed perpetually assess the accuracy of AI chatbots in this field of medicine, evaluate predictors of poor accuracy in the models, and both monitor the performance as well as utilization of AI chatbots within pediatric cardiology.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eAI: Artificial intelligence\u003c/p\u003e\n\u003cp\u003eGPT: Generative Pre-Trained Transformer\u003c/p\u003e\n\u003cp\u003eNLP: Natural\u0026nbsp;language processing\u003c/p\u003e\n\u003cp\u003eECG: Electrocardiogram\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFinancial Support:\u0026nbsp;\u003c/strong\u003eNone.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSource of Funding:\u003c/strong\u003e None.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMeeting Presentation:\u0026nbsp;\u003c/strong\u003eNone.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of Interest:\u0026nbsp;\u003c/strong\u003eNo conflicting relationship exists for any author.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements:\u003c/strong\u003e None.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eLiu PR, Lu L, Zhang JY, Huo TT, Liu SX, Ye ZW. Application of Artificial Intelligence in Medicine: An Overview. \u003cem\u003eCurr Med Sci. \u003c/em\u003e2021;41(6):1105-1115.\u003c/li\u003e\n\u003cli\u003eBart NK, Pepe S, Gregory AT, Denniss AR. Emerging Roles of Artificial Intelligence (AI) in Cardiology: Benefits and Barriers in a \u0026apos;Brave New World\u0026apos;. \u003cem\u003eHeart Lung Circ. \u003c/em\u003e2023;32(8):883-888.\u003c/li\u003e\n\u003cli\u003eDave T, Athaluri SA, Singh S. ChatGPT in medicine: an overview of its applications, advantages, limitations, future prospects, and ethical considerations. \u003cem\u003eFront Artif Intell. \u003c/em\u003e2023;6:1169595.\u003c/li\u003e\n\u003cli\u003eSardana D, Fagan TR, Wright JT. ChatGPT: A disruptive innovation or disrupting innovation in academia? \u003cem\u003eJ Am Dent Assoc. \u003c/em\u003e2023;154(5):361-364.\u003c/li\u003e\n\u003cli\u003eGao CA, Howard FM, Markov NS, et al. Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. \u003cem\u003eNPJ Digit Med. \u003c/em\u003e2023;6(1):75.\u003c/li\u003e\n\u003cli\u003eRuksakulpiwat S, Kumar A, Ajibade A. Using ChatGPT in Medical Research: Current Status and Future Directions. \u003cem\u003eJ Multidiscip Healthc. \u003c/em\u003e2023;16:1513-1520.\u003c/li\u003e\n\u003cli\u003eTeixeira da Silva JA. Can ChatGPT rescue or assist with language barriers in healthcare communication? \u003cem\u003ePatient Educ Couns. \u003c/em\u003e2023;115:107940.\u003c/li\u003e\n\u003cli\u003eMihalache A, Popovic MM, Muni RH. Performance of an Artificial Intelligence Chatbot in Ophthalmic Knowledge Assessment. \u003cem\u003eJAMA Ophthalmol. \u003c/em\u003e2023;141(6):589-597.\u003c/li\u003e\n\u003cli\u003eRao A, Kim J, Kamineni M, Pang M, Lie W, Succi MD. Evaluating ChatGPT as an Adjunct for Radiologic Decision-Making. \u003cem\u003emedRxiv. \u003c/em\u003e2023.\u003c/li\u003e\n\u003cli\u003ePatel SB, Lam K. ChatGPT: the future of discharge summaries? \u003cem\u003eLancet Digit Health. \u003c/em\u003e2023;5(3):e107-e108.\u003c/li\u003e\n\u003cli\u003eSingh S, Djalilian A, Ali MJ. ChatGPT and Ophthalmology: Exploring Its Potential with Discharge Summaries and Operative Notes. \u003cem\u003eSemin Ophthalmol. \u003c/em\u003e2023;38(5):503-507.\u003c/li\u003e\n\u003cli\u003eSkalidis I, Cagnina A, Luangphiphat W, et al. ChatGPT takes on the European Exam in Core Cardiology: an artificial intelligence success story? \u003cem\u003eEur Heart J Digit Health. \u003c/em\u003e2023;4(3):279-281.\u003c/li\u003e\n\u003cli\u003eGibson A. How Does ChatGPT Perform on the Medical Licensing Exams? The Implications of Large Language Models for Medical Education and Knowledge Assessment. \u003cem\u003emedRxiv. \u003c/em\u003e2022.\u003c/li\u003e\n\u003cli\u003eKung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. \u003cem\u003ePLOS Digit Health. \u003c/em\u003e2023;2(2):e0000198.\u003c/li\u003e\n\u003cli\u003eMihalache A, Huang RS, Popovic MM, Muni RH. Performance of an Upgraded Artificial Intelligence Chatbot for Ophthalmic Knowledge Assessment. \u003cem\u003eJAMA Ophthalmol. \u003c/em\u003e2023;141(8):798-800.\u003c/li\u003e\n\u003cli\u003eWaisberg E, Ong J, Masalkhi M, et al. GPT-4: a new era of artificial intelligence in medicine. \u003cem\u003eIr J Med Sci. \u003c/em\u003e2023.\u003c/li\u003e\n\u003cli\u003eEidem B. \u003cem\u003ePediatric Cardiology Board Review.\u003c/em\u003e 2nd ed. Philadelphia, PA: Wolters Kluwer; 2023.\u003c/li\u003e\n\u003cli\u003eSallam M. ChatGPT Utility in Healthcare Education, Research, and Practice: Systematic Review on the Promising Perspectives and Valid Concerns. \u003cem\u003eHealthcare (Basel). \u003c/em\u003e2023;11(6).\u003c/li\u003e\n\u003cli\u003eVaid A, Jiang J, Sawant A, et al. A foundational vision transformer improves diagnostic performance for electrocardiograms. \u003cem\u003eNPJ Digit Med. \u003c/em\u003e2023;6(1):108.\u003c/li\u003e\n\u003cli\u003eHe B, Kwan AC, Cho JH, et al. Blinded, randomized trial of sonographer versus AI cardiac function assessment. \u003cem\u003eNature. \u003c/em\u003e2023;616(7957):520-524.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Table","content":"\u003cp\u003e\u003cstrong\u003eTable 1.\u0026nbsp;\u003c/strong\u003eNumber of correctly answered questions (out of 88) by ChatGPT 3.5 and ChatGPT 4.0 in each of the 11 chapters of the \u003cem\u003ePediatric Cardiology Board Review\u0026nbsp;\u003c/em\u003ebook.\u003csup\u003e17\u003c/sup\u003e\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"633\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eChatGPT 3.5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eChatGPT \u0026nbsp;4.0\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eChapter 1: Cardiac Anatomy and Physiology\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eChapter 2: Congenital Cardiac Malformations\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e2\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eChapter 3: Diagnosis of Congenital Heart Disease\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e3\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e4\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eChapter 4: Cardiac Catheterization and Angiography\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e3\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e3\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eChapter 5: Non-invasive Cardiac Imaging\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e4\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e7\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eChapter 6: Electrophysiology Questions for Paediatrics\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e1\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eChapter 7: Exercise Physiology and Testing\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e4\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e6\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eChapter 8: Outpatient Cardiology\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e3\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eChapter 9: Cardiac Intensive Care and Heart Failure\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e4\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e5\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eChapter 10: Cardiac Pharmacology\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e3\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e6\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eChapter 11: Surgical Palliation and Repair of Congenital Heart Disease\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e4\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e7\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eTotal\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e33\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e58\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"64.24050632911393%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eScore on entire examination\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e37.5%\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.879746835443036%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003e65.9%\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"pediatric-cardiology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"pedc","sideBox":"Learn more about [Pediatric Cardiology](http://link.springer.com/journal/246)","snPcode":"246","submissionUrl":"https://submission.nature.com/new-submission/246/3","title":"Pediatric Cardiology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Artificial intelligence, Natural language processing, Pediatric Cardiology, ChatGPT, Education","lastPublishedDoi":"10.21203/rs.3.rs-3360192/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3360192/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eIntroduction\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eArtificial intelligence chatbots, like ChatGPT, have become powerful tools that are disrupting how humans interact with technology. The potential uses within medicine are vast. In medical education, these chatbots have shown improvements, in a short time span, in generalized medical examinations. We evaluated the overall performance and improvement between ChatGPT 3.5 and 4.0 in a test of pediatric cardiology knowledge.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eChatGPT 3.5 and ChatGPT 4.0 were used to answer text based multiple choice questions derived from a Pediatric Cardiology Board Review textbook. Each chatbot, was given an 88 question test, subcategorized into 11 topics. We excluded questions with modalities other than text (sound clips or images). Statistical analysis was done using an unpaired two-tailed t-test.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eOf the same 88 questions, ChatGPT 4.0 answered 66% of the questions correctly (n=58/88) which was significantly greater (p\u0026lt;0.0001) than ChatGPT 3.5, which only answered 38% (33/88). The ChatGPT 4.0 version also did better on each subspeciality topic as compared to ChatGPT 3.5.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWhile acknowledging that ChatGPT does not yet offer subspecialty level knowledge in pediatric cardiology, the performance in pediatric cardiology educational assessments showed a considerable improvement in a short period of time between ChatGPT 3.5 to 4.0.\u003c/p\u003e","manuscriptTitle":"Progression of an Artificial Intelligence Chatbot (ChatGPT) for Pediatric Cardiology Educational Knowledge Assessment: Considerable Gains in a Short Time","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-09-22 18:35:54","doi":"10.21203/rs.3.rs-3360192/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2023-11-26T20:42:22+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-10-03T18:04:34+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"3c27a74f-eab0-4595-8bc1-1afc20a51b69","date":"2023-09-25T01:39:01+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"21e1f8af-0404-47c6-92f0-f9062c003c3d","date":"2023-09-23T23:28:34+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"c79c6368-89bd-4056-9cd7-7cad015d3fdc","date":"2023-09-21T11:38:23+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"4f406b41-57f0-4dc8-bb75-cf4f91d8f071","date":"2023-09-20T23:24:41+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-09-20T23:11:53+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-09-16T07:00:30+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2023-09-16T07:00:29+00:00","index":"","fulltext":""},{"type":"submitted","content":"Pediatric Cardiology","date":"2023-09-16T02:27:16+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"pediatric-cardiology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"pedc","sideBox":"Learn more about [Pediatric Cardiology](http://link.springer.com/journal/246)","snPcode":"246","submissionUrl":"https://submission.nature.com/new-submission/246/3","title":"Pediatric Cardiology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"e7abfd99-e66c-48c5-aa68-881387cff437","owner":[],"postedDate":"September 22nd, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-01-08T15:07:22+00:00","versionOfRecord":{"articleIdentity":"rs-3360192","link":"https://doi.org/10.1007/s00246-023-03385-6","journal":{"identity":"pediatric-cardiology","isVorOnly":false,"title":"Pediatric Cardiology"},"publishedOn":"2024-01-03 15:00:37","publishedOnDateReadable":"January 3rd, 2024"},"versionCreatedAt":"2023-09-22 18:35:54","video":"","vorDoi":"10.1007/s00246-023-03385-6","vorDoiUrl":"https://doi.org/10.1007/s00246-023-03385-6","workflowStages":[]},"version":"v1","identity":"rs-3360192","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3360192","identity":"rs-3360192","version":["v1"]},"buildId":"ehx78VzkSd0WSzXnipQa-","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-06-06T02:00:05.402940+00:00
License: CC-BY-4.0