Assessing the Competence of ChatGPT-3.5 Artificial Intelligence System in Executing the ACLS Protocol of the AHA 2020

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Objectives: Artificial intelligence (AI) has become the focus of current studies, particularly due to its contribution in preventing human labor and time loss. The most important contribution of AI applications in the medical field will be to provide opportunities for increasing clinicians' gains, reducing costs, and improving public health. This study aims to assess the proficiency of ChatGPT-3.5, one of the most advanced AI applications available today, in its knowledge of current information based on the American Heart Association (AHA) 2020 guidelines. Methods: An 80-question quiz in a question-and-answer format, which includes the current AHA 2020 application steps, was prepared and applied to ChatGPT-3.5 in both English (ChatGPT-3.5 English) and native language (ChatGPT-3.5 Turkish) versions in March 2023. The questions were prepared only in the native language for emergency medicine specialists. Results: We found a similar success rate of over 80% in all questions asked to ChatGPT-3.5 and two independent emergency medicine specialists with at least 5 years of experience who did not know each other. ChatGPT-3.5 achieved a 100% success rate in all questions related to the General Overview for Current AHA Guideline, Airway Management, and Ventilation chapters in English. Conclusions: Our study indicates that ChatGPT-3.5 provides similar accurate and up-to-date responses as experienced emergency specialists in the AHA 2020 Advanced Cardiac Life Support Guidelines. This suggests that with future updated versions of ChatGPT, instant access to accurate and up-to-date information based on textbooks and guidelines will be possible.
Full text 67,734 characters · extracted from preprint-html · click to expand
Assessing the Competence of ChatGPT-3.5 Artificial Intelligence System in Executing the ACLS Protocol of the AHA 2020 | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Assessing the Competence of ChatGPT-3.5 Artificial Intelligence System in Executing the ACLS Protocol of the AHA 2020 İbrahim Altundağ, Sinem Doğruyol, Burcu Genç Yavuz, Kaan Yusufoğlu, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3035900/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Objectives: Artificial intelligence (AI) has become the focus of current studies, particularly due to its contribution in preventing human labor and time loss. The most important contribution of AI applications in the medical field will be to provide opportunities for increasing clinicians' gains, reducing costs, and improving public health. This study aims to assess the proficiency of ChatGPT-3.5, one of the most advanced AI applications available today, in its knowledge of current information based on the American Heart Association (AHA) 2020 guidelines. Methods: An 80-question quiz in a question-and-answer format, which includes the current AHA 2020 application steps, was prepared and applied to ChatGPT-3.5 in both English (ChatGPT-3.5 English) and native language (ChatGPT-3.5 Turkish) versions in March 2023. The questions were prepared only in the native language for emergency medicine specialists. Results: We found a similar success rate of over 80% in all questions asked to ChatGPT-3.5 and two independent emergency medicine specialists with at least 5 years of experience who did not know each other. ChatGPT-3.5 achieved a 100% success rate in all questions related to the General Overview for Current AHA Guideline, Airway Management, and Ventilation chapters in English. Conclusions: Our study indicates that ChatGPT-3.5 provides similar accurate and up-to-date responses as experienced emergency specialists in the AHA 2020 Advanced Cardiac Life Support Guidelines. This suggests that with future updated versions of ChatGPT, instant access to accurate and up-to-date information based on textbooks and guidelines will be possible. artificial intelligence AI chatbot generative pretrained transformer clinical decision support guidelines Figures Figure 1 Figure 2 1. Introduction One of the most important focal points of developing technology and computer systems is eliminating dependence on human labor and creating autonomous systems. Instead of systems that only carry out predetermined commands, the focus is on developing autonomous systems (artificial intelligence-AI) that can react appropriately to changing conditions. By utilizing intelligent algorithms and iterative processes, AI imitates human intelligence and has the capacity to enhance its performance by continuously updating the vast amount of information it gathers [ 1 ]. One of the most important features of AI is the natural language processing (NLP) function. NLP refers to the branch of AI that focuses on giving computers the ability to understand texts and spoken words in a way that humans can comprehend [ 2 ]. One of the most significant contributions of NLP technology in medical applications is the development of clinical decision support (CDS) systems. CDS is designed to assist healthcare professionals in making the most accurate decisions during the processes in healthcare sciences. The goal of CDS is to ensure that data in healthcare enterprises is presented to users completely, accurately, in a timely manner, and in the correct format [ 3 ]. NLP, which is closely related to CDS, is effective in directing CDS by using unstructured text information, representing clinical information and CDS interventions in standardized formats, and leveraging clinical narratives [ 4 ]. Today, one of the most notable examples of NLP and CDS models is the Generative Pretrained Transformer (GPT) model [ 1 ]. The popularity of AI has reached its peak lately in mainstream media and literature, with the emergence of Generative Pretrained Transformer-3 (GPT-3), a language model that can generate human-like texts [ 5 ]. Although not specifically designed for medical diagnoses, compared to previous CDS systems, the interactive version of GPT-3, ChatGPT-3, is a user-friendly AI model for general users with useful features such as a user-friendly interface, free text input, and sentence output [ 6 ]. The latest version of this system, ChatGPT-3.5, has been made available for free use since March 2023. Since the introduction of advanced AI applications like ChatGPT, the prominent feature that stands out in the CDS field is its use for case-based diagnosis and differential diagnosis purposes [ 7 , 8 ]. However, another important feature of ChatGPT that can be preferred is its ability to instantly provide users with desired information on any specific topic. Despite the abundance of information and resources in the health and medical field, the difficulty of accessing accurate information easily remains a challenge. Although the widespread use of the internet facilitates access to information, clinicians may still have difficulty accessing specific information today. Accessing accurate information quickly is the most basic need, especially for clinicians working in chaotic environments such as emergency departments. Querying conducted through ChatGPT, which can be referred to as consultation or advisory, will be one of the important areas of use of AI applications in emergency medicine, enabling clinicians to access up-to-date and accurate information within seconds. In our study, we aimed to test the proficiency of ChatGPT-3.5, one of the most advanced AI applications of today, in terms of its knowledge of current information based on the American Heart Association (AHA) 2020 Advanced Cardiac Life Support (ACLS) guidelines. For this purpose, a quiz prepared in both Turkish and English was presented to ChatGPT-3.5. The same quiz was also given to emergency medicine specialists to compare ChatGPT-3.5's success rates in Turkish and English. Additionally, considering that clinicians use ChatGPT-3.5 for queries in their native language, we aimed to evaluate the success levels in two different languages. 2. Methods Our study was conducted at the Emergency Medicine Department of Haydarpaşa Numune Training and Research Hospital in March 2023. 2.1. Study Design A quiz consisting of 80 questions in total, divided into 8 different chapters that aim to query current practices regarding the American Heart Association (AHA) 2020 Advanced Cardiac Life Support (ACLS) application steps, was prepared for use in our study. [ 9 ]. The prepared quiz was administered to both ChatGPT-3.5 and two emergency medicine specialists with a minimum of 5 years of experience who were not informed about the study protocol. The quiz results were evaluated by a blinded researcher. In the study, first, the accuracy of ChatGPT-3.5's answers was evaluated with reference to the AHA 2020 ACLS guidelines. In addition, the inter-rater agreement of the items included in the quiz was examined. Then, the accuracy rates of ChatGPT-3.5 and emergency medicine specialists' answers were compared (Fig. 1 ). The accuracy rates were evaluated both on individual question level and chapter level, and reported as 'success percentage'. An accuracy rate of 80% and above was considered successful in line with the literature [ 10 ]. 2.2 Quiz Components An 80-question quiz was created in a question-answer format that includes the current ACLS application steps. These steps were designed in a total of 8 chapters covering general information on current ACLS, basic-advanced airway management, high-quality cardiopulmonary resuscitation (CPR), ventilation, defibrillation, medications, vascular access and return of spontaneous circulation (ROSC), and finally questions related to class of recommendation. The quiz includes 5 questions testing general information on current ACLS guidelines, 7 questions on basic-advanced airway management, 11 questions on high-quality CPR, 8 questions on ventilation, 12 questions on defibrillation, 22 questions on medications, 6 questions on vascular access and ROSC, and finally 9 questions related to class of recommendation, in total 80 questions. 2.3 Obtaining Data The quiz was administered using the ChatGPT-3.5 March 2023 version. It was applied to the AI application in both English (ChatGPT-3.5 English) and native language (ChatGPT-3.5 Turkish), while for emergency medicine specialists, the questions were prepared only in the native language. The quiz questions were prepared to have a definitive, unchanging, and non-interpretable answer, and were transferred to ChatGPT-3.5 in plain text format, and the answers were evaluated for compliance with the AHA 2020 ACLS guidelines (Fig. 2 ). 2.4 Outcomes The primary aim of the study is to examine the accuracy of ChatGPT-3.5 in providing answers to questions related to current ACLS guidelines. The secondary aim is to compare the current knowledge level of experienced emergency medicine specialists with ChatGPT-3.5 regarding ACLS. The accuracy of the answers to questions related to ACLS was evaluated based on the 'Adult Basic and Advanced Life Support: 2020 AHA Guidelines for Cardiopulmonary Resuscitation and Emergency Cardiovascular Care' guide [ 9 ]. 3. Statistical analysis In our study, categorical data were expressed as numbers (n) and percentages (%). Chi-square test was used for comparing the data. The agreement between specialists and ChatGPT-3.5 for categorical data was evaluated with inter-rater agreement. Inter-rater agreement was expressed with Kappa value and 95% Confidence interval (CI), and the significance level was accepted as p < 0.05. Data analysis was performed using MedCalc Version 20.218 (MedCalc Software Ltd, Ostend, Belgium) and IBM SPSS version 20 (IBM Corp, Armonk, NY) programs. 4. Results The success rates for the entire quiz were as follows: for I.specialist, 81.3% (65/80); for II.specialist, 87.5% (70/80); for ChatGPT-3.5 in Turkish, 81.3% (65/80). The success rate for ChatGPT-3.5 in English was calculated as 86.3% (69/80). There was no statistically significant difference in success rates between I.specialist and II.specialist (p = 0.067). Similarly, there was no statistically significant difference in success rates between I.specialist and ChatGPT-3.5 in Turkish (p = 0.386), or between II.specialist and ChatGPT-3.5 in Turkish (p = 0.332). The distribution of correct answer rates and success rates by section is shown in Table 1 . In addition, in Table 1 , the success rates of specialists and ChatGPT-3.5 Turkish were compared chapter by chapter, and no statistically significant difference was found between the two groups. Table 1 Comparison of success percentages of ChatGPT-3.5 and specialists according to Quiz sections Chapters (questions) I. Specialist II. Specialist ChatGPT-3.5 Turkish ChatGPT-3.5 English P value* A – General Overview for Current AHA Guıdeline (n = 5) 5/5 (%100) 5/5 (%100) 5/5 (%100) 5/5 (%100) - B – Airway Management (n = 7) 5/7 (%71.4) 7/7 (%100) 7/7 (%100) 7/7 (%100) 0.110 C – High Quality CPR (n = 11) 8/11 (%72.7) 10/11 (%90.9) 11/11 (%100) 10/11 (%90.9) 0.137 D – Ventilation (n = 8) 7/8 (%87.5) 8/8 (%100) 5/8 (%62.5) 8/8 (%100) 0.122 E – Defibrillation (n = 12) 9/12 (%75) 10/12 (%83.3) 10/12 (%83.3) 9/12 (%75) 0.837 F – Medications (n = 22) 20/22 (%90.9) 18/22 (%81.8) 16/22 (%72.7) 19/22 (%86.4) 0.295 G – Vascular Access and ROSC (n = 6) 5/6 (%83.3) 4/6 (%66.6) 6/6 (%100) 5/6 (%83.3) 0.301 H – Class (Strength) of Recommendation (n = 9) 6/9 (%66.6) 8/9 (%88.8) 5/9 (%55.5) 6/9 (%66.6) 0.288 AHA: American Heart Association, ROSC: Return of spontaneous circulation, CPR: Cardiopulmonary resuscitation. The ratio of correct answers to all answers is given in parentheses as a percentage. *P values obtained by comparing the success rates of ChatGPT-3.5 Turkish with the success rates of specialists. The Kappa values obtained from the comparison of the correct/incorrect answers given to the questions in terms of interrater agreement are shown in Table 2 . Only the answers of ChatGPT-3.5 to the Turkish versions of the questions and the English versions of the questions showed fair agreement (Kappa = 0.27) (p = 0.015). Table 2 Inter-rater aggrement Kappa values Agreement Kappa Values 95% Confidence Interval I.specialist – II.specialist 0.20 -0.06-0.46 I.specialist – ChatGPT-3.5 Turkish 0.09 -0.14-0.33 II.specialist – ChatGPT-3.5 Turkish 0.10 -0.13-0.35 ChatGPT-3.5 Turkish- ChatGPT-3.5 English 0.27 0.00-0.53 The number of questions for which both emergency medicine specialists gave wrong answers was 4. It was observed that ChatGPT-3.5 Turkish also could not provide the correct answer to one of these questions (p = 0.742). This question was "What is the maximum initial dose for biphasic defibrillators according to the AHA 2020 ACLS algorithm?" However, it was seen that ChatGPT-3.5 gave the correct answer to this question in the English quiz. The number of questions on which both specialists answered correctly was 59, and when examined by chapters, it was found that the common correct answer rate of both specialists was 5 out of 5 (100%) in the 'A - General Overview for Current AHA Guideline' chapter. The chapter with the second highest common correct answer rate of both specialists (87.5%) was the 'D - Ventilation' chapter. It was found that ChatGPT-3.5 Turkish gave the correct answer to 50 out of the 59 questions on which both specialists answered correctly (p = 0.179). 5. Discussion As far as we know, our study is the first to evaluate AHA 2020 ACLS guidelines with ChatGPT-3.5. We observed a similar success rate of over 80% in all questions asked to ChatGPT-3.5 and two emergency medicine specialists working independently in different hospitals and who do not know each other. ChatGPT-3.5 answered all questions related to General Overview for Current AHA Guideline, Airway Management, and Ventilation chapters with 100% success rate in English, while it answered all questions related to General Overview for Current AHA Guideline, Airway Management, High Quality CPR, and Vascular Access and ROSC chapters with 100% success rate in Turkish. Based on the data we obtained, we found that ChatGPT-3.5 can provide highly accurate and up-to-date answers to questions about current ACLS practices. In addition, similar success rates were obtained when compared to emergency medicine specialists. The contribution of AI to preventing human labor and time loss is one of the most important features that makes current studies focus on it. The promise of AI applications in health and medical services is to provide opportunities for increasing gains for patients and the clinical team, reducing costs, and improving public health. The most important prerequisites for the successful use of AI applications in medicine are accessibility, standardization, quality, and encouraging data that represents the population [ 11 ]. There are many types of studies conducted with ChatGPT, including creating triage and differential diagnosis lists [ 6 ], querying drug interactions [ 12 ], AI-generated article writing [ 13 ], and medical problem and case-solving [ 14 ]. In general, the studies conducted through ChatGPT emphasize the CDS aspect, in which ChatGPT makes inferences about variable, fictional cases or situations and the accuracy of these inferences are evaluated. Studies that emphasize the evaluation of using ChatGPT as a reference source, rather than querying variable scenarios, like our study, will contribute to the literature in terms of using AI applications for a different purpose. The language selected for using ChatGPT as a reference source should be English, which is the main language of medical literature. In our study, the difference in accuracy rates between querying the same question in the native language and English (86.3% accuracy in English and 81.3% in the native language) supports the need for English as the selected language. ChatGPT's use of plain text sources may have contributed to the poor performance in the "Class of Recommendation" section, which was the least successful section in the Quiz for ChatGPT-3.5 (66.6% accuracy in English and 55% accuracy in Turkish) and led to a low rate of correct answers by specialists (66.6% and 88.8%, respectively). One possible reason for this may be that recommendations in guidelines are usually presented in tables, while ChatGPT prefers to use plain text instead of tables or figures as a source. Therefore, in medical literature queries performed through ChatGPT-3.5, it is essential to verify the accuracy of the information contained in tables, and it should not be forgotten that ChatGPT may provide incorrect answers. One of the other important points revealed by our study is the existence of incorrect answers alongside the correct answers provided by ChatGPT-3.5. This poses a risk of misleading the physician or user. In our study, ChatGPT-3.5 provided correct answers to 69 out of 80 questions asked in English and gave incorrect answers to 11 questions. This situation deviates from the idea that ChatGPT can be used as a definitive reference source. However, considering the similarity of the correct answer rates at the expert level, and with the development of algorithms and enriched databases, the reduction in incorrect answer rates in the coming years may bring back the possibility of using ChatGPT as a reference source. 6. Conclusion Our study showed that ChatGPT-3.5 has a similar level of accurate and up-to-date knowledge as an experienced emergency medicine specialist on the AHA 2020 Advanced Cardiac Life Support Guidelines. With the development of algorithms and new versions, ChatGPT's mastery of current information can be increased, and the number of incorrect answers can be reduced. Querying current guideline information through ChatGPT can serve as a consultant function easily accessible to emergency physicians. We believe that our study sheds light on the idea that ChatGPT can be used as a portal to instantly access accurate and up-to-date information based on textbooks and guidelines in the coming years. 7. Limitations One of the main limitations of our study is that ChatGPT-3.5 has not received clinical approval for obtaining healthcare information. Although our study has achieved successful results with ChatGPT-3.5, it should be kept in mind that AI applications, including ChatGPT, must be used with appropriate methods, and that ChatGPT is still being developed and its answers may be incorrect. One of the important limitations of ChatGPT is that its incorrect answers can lead to incorrect guidance and medical faults [ 15 ]. One of the other important limitations of our study is the language issue. Different results can be obtained when queries are performed in the native language compared to when they are performed in English. We believe that queries should be conducted in English due to the fact that medical literature is written in English, and medical terms should be searched using the forms used in the literature. In addition, another limitation of our study is that no time limit was set for both ChatGPT-3.5 and specialists to answer each question in the quiz. If time data is added to the study, it may be possible to compare the response times of AI and specialists. Declarations Competing interests Authors declare no conflict of interest. Acknowledgment funding The authors did not apply for a specific grant for this research from any funding agency in the public, commercial or non-profit sectors. References Amisha, Malik P, Pathania M, Rathaur VK. Overview of artificial intelligence in medicine. J Family Med Prim Care . 2019;8(7):2328-31. https://doi.org/10.4103/jfmpc.jfmpc_440_19. Reading Turchioe M, Volodarskiy A, Pathak J, Wright DN, Tcheng JE, Slotwiner D. Systematic review of current natural language processing methods and applications in cardiology. Heart . 2022;108(12):909-16. https://doi.org/10.1136/heartjnl-2021-319769. Gonçalves LS, Amaro MLM, Romero ALM, Schamne FK, Fressatto JL, Bezerra CW. Implementation of an Artificial Intelligence Algorithm for sepsis detection. Rev Bras Enferm . 2020;73(3):e20180421. https://doi.org/10.1590/0034-7167-2018-0421. Demner-Fushman D, Chapman WW, McDonald CJ. What can natural language processing do for clinical decision support? J Biomed Inform . 2009;42(5):760-72. https://doi.org/10.1016/j.jbi.2009.08.007. Nath S, Marie A, Ellershaw S, Korot E, Keane PA. New meaning for NLP: the trials and tribulations of natural language processing with GPT-3 in ophthalmology. Br J Ophthalmol . 2022;106(7):889-92. https://doi.org/10.1136/bjophthalmol-2022-321141. Hirosawa T, Harada Y, Yokose M, Sakamoto T, Kawamura R, Shimizu T. Diagnostic Accuracy of Differential-Diagnosis Lists Generated by Generative Pretrained Transformer 3 Chatbot for Clinical Vignettes with Common Chief Complaints: A Pilot Study. Int J Environ Res Public Health . 2023;20(4)https://doi.org/10.3390/ijerph20043378. Goodwin TR, Harabagiu SM. Medical Question Answering for Clinical Decision Support. Proc ACM Int Conf Inf Knowl Manag . 2016;2016:297-306. https://doi.org/10.1145/2983323.2983819. Xu G, Rong W, Wang Y, Ouyang Y, Xiong Z. External features enriched model for biomedical question answering. BMC Bioinformatics . 2021;22(1):272. https://doi.org/10.1186/s12859-021-04176-7. Panchal AR, Bartos JA, Cabañas JG, Donnino MW, Drennan IR, Hirsch KG, et al. Part 3: Adult Basic and Advanced Life Support: 2020 American Heart Association Guidelines for Cardiopulmonary Resuscitation and Emergency Cardiovascular Care. Circulation . 2020;142(16_suppl_2):S366-S468. https://doi.org/10.1161/CIR.0000000000000916. Banik R, Rahman M, Sikder MT, Rahman QM, Pranta MUR. Knowledge, attitudes, and practices related to the COVID-19 pandemic among Bangladeshi youth: a web-based cross-sectional analysis. Z Gesundh Wiss . 2023;31(1):9-19. https://doi.org/10.1007/s10389-020-01432-7. Matheny ME, Whicher D, Thadaney Israni S. Artificial Intelligence in Health Care: A Report From the National Academy of Medicine. JAMA . 2020;323(6):509-10. https://doi.org/10.1001/jama.2019.21579. Juhi A, Pipil N, Santra S, Mondal S, Behera JK, Mondal H. The Capability of ChatGPT in Predicting and Explaining Common Drug-Drug Interactions. Cureus . 2023;15(3):e36272. https://doi.org/10.7759/cureus.36272. Dergaa I, Chamari K, Zmijewski P, Ben Saad H. From human writing to artificial intelligence generated text: examining the prospects and potential threats of ChatGPT in academic writing. Biol Sport . 2023;40(2):615-22. https://doi.org/10.5114/biolsport.2023.125623. Sinha RK, Deb Roy A, Kumar N, Mondal H. Applicability of ChatGPT in Assisting to Solve Higher Order Problems in Pathology. Cureus . 2023;15(2):e35237. https://doi.org/10.7759/cureus.35237. King MR. The Future of AI in Medicine: A Perspective from a Chatbot. Ann Biomed Eng . 2023;51(2):291-5. https://doi.org/10.1007/s10439-022-03121-w. Additional Declarations No competing interests reported. Supplementary Files ChatGPTAITestResults.pdf ChatGPTAITestResults.xlsx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3035900","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":208035757,"identity":"4dc27fe7-512e-4291-99e1-e6600d525b46","order_by":0,"name":"İbrahim Altundağ","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABJUlEQVRIiWNgGAWjYDCCAwwMh4FUAhsDAyOQzcDYwMAMoiVkiNHCcOAAWAtbAkgLDz4tzCAtDAgtPAYgDk4tfLcPMB4u3GOXxyd2+MHhj213ZPv713x+daPGgoeB/fDRDVi0SJ5LYDg841lyMZt0msGBg23PjGfceLvNOucY0GE8aWk3sGgxOAP0C88B5sQ26QSQlsOJDTfObjPOYQNqkeAxw6OlHqgl/QNYy/wbZ54Z5/wjqOUwUEsOxJYN53uYH+e24dYieYax4fCMA8dBWgoOnDl32HjjDTYz5tw+CR42HH7hO8N8+HPBgerE+bPTNz6oKDssO+/84cefc77VyfGzHz6GTQs4IlCBRAKbBIhmw6ocK+A/wPyBeNWjYBSMglEwAgAAffF2crwveMQAAAAASUVORK5CYII=","orcid":"","institution":"University of Health Sciences, Başakşehir Çam and Sakura City Hospital","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"İbrahim","middleName":"","lastName":"Altundağ","suffix":""},{"id":208035758,"identity":"fe26c141-760e-4258-a220-47a2761502c2","order_by":1,"name":"Sinem Doğruyol","email":"","orcid":"","institution":"University of Health Sciences, Haydarpasa Numune Training and Research Hospital","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Sinem","middleName":"","lastName":"Doğruyol","suffix":""},{"id":208035759,"identity":"1d82af32-83c6-4638-afb8-fc1770508e74","order_by":2,"name":"Burcu Genç Yavuz","email":"","orcid":"","institution":"University of Health Sciences, Haydarpasa Numune Training and Research Hospital","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Burcu","middleName":"Genç","lastName":"Yavuz","suffix":""},{"id":208035760,"identity":"e98844f3-87b4-4eb8-92fb-f53d54c8cdf9","order_by":3,"name":"Kaan Yusufoğlu","email":"","orcid":"","institution":"University of Health Sciences, Haydarpasa Numune Training and Research Hospital","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Kaan","middleName":"","lastName":"Yusufoğlu","suffix":""},{"id":208035762,"identity":"2e659b22-0cdc-4f52-9361-b3686a09442f","order_by":4,"name":"Mustafa Ahmet Afacan","email":"","orcid":"","institution":"University of Health Sciences, Haydarpasa Numune Training and Research Hospital","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Mustafa","middleName":"Ahmet","lastName":"Afacan","suffix":""},{"id":208035764,"identity":"b67c284d-8aff-4401-8d6b-fc652ce81d7e","order_by":5,"name":"Şahin Çolak","email":"","orcid":"","institution":"University of Health Sciences, Haydarpasa Numune Training and Research Hospital","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Şahin","middleName":"","lastName":"Çolak","suffix":""}],"badges":[],"createdAt":"2023-06-07 19:59:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3035900/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3035900/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":38307550,"identity":"1035a147-aa99-47ed-b25a-1252cc67baf0","added_by":"auto","created_at":"2023-06-09 18:22:33","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":27279,"visible":true,"origin":"","legend":"\u003cp\u003eFlow diagram that illustrates the design of the study.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-3035900/v1/89eafc1eaedb66dcf9b81ddd.png"},{"id":38307551,"identity":"4ff33858-4ed2-4f0b-9171-dc6f7a8ff581","added_by":"auto","created_at":"2023-06-09 18:22:33","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":193598,"visible":true,"origin":"","legend":"\u003cp\u003eExample view of the use of ChatGPT-3.5. Sample questions with incorrect (above) and correct answers (below).\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-3035900/v1/c3606870a82d01f027137dc4.png"},{"id":39648653,"identity":"eaf26824-eb89-450c-862d-8ffc2adbbbce","added_by":"auto","created_at":"2023-07-06 16:29:39","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":486518,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3035900/v1/2a89a614-aed1-4136-a255-f4f9c0019f17.pdf"},{"id":38307553,"identity":"ef572df7-a804-47bb-adbc-83af474c85d8","added_by":"auto","created_at":"2023-06-09 18:22:33","extension":"pdf","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":185998,"visible":true,"origin":"","legend":"","description":"","filename":"ChatGPTAITestResults.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3035900/v1/7d2a9111fe8c61df0afb6c84.pdf"},{"id":38307552,"identity":"9520a370-c078-406f-8e27-3c48aeb43006","added_by":"auto","created_at":"2023-06-09 18:22:33","extension":"xlsx","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":17850,"visible":true,"origin":"","legend":"","description":"","filename":"ChatGPTAITestResults.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-3035900/v1/1d997aeed7dae61ddf0fc687.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Assessing the Competence of ChatGPT-3.5 Artificial Intelligence System in Executing the ACLS Protocol of the AHA 2020","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eOne of the most important focal points of developing technology and computer systems is eliminating dependence on human labor and creating autonomous systems. Instead of systems that only carry out predetermined commands, the focus is on developing autonomous systems (artificial intelligence-AI) that can react appropriately to changing conditions. By utilizing intelligent algorithms and iterative processes, AI imitates human intelligence and has the capacity to enhance its performance by continuously updating the vast amount of information it gathers [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. One of the most important features of AI is the natural language processing (NLP) function. NLP refers to the branch of AI that focuses on giving computers the ability to understand texts and spoken words in a way that humans can comprehend [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. One of the most significant contributions of NLP technology in medical applications is the development of clinical decision support (CDS) systems. CDS is designed to assist healthcare professionals in making the most accurate decisions during the processes in healthcare sciences. The goal of CDS is to ensure that data in healthcare enterprises is presented to users completely, accurately, in a timely manner, and in the correct format [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. NLP, which is closely related to CDS, is effective in directing CDS by using unstructured text information, representing clinical information and CDS interventions in standardized formats, and leveraging clinical narratives [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Today, one of the most notable examples of NLP and CDS models is the Generative Pretrained Transformer (GPT) model [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The popularity of AI has reached its peak lately in mainstream media and literature, with the emergence of Generative Pretrained Transformer-3 (GPT-3), a language model that can generate human-like texts [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAlthough not specifically designed for medical diagnoses, compared to previous CDS systems, the interactive version of GPT-3, ChatGPT-3, is a user-friendly AI model for general users with useful features such as a user-friendly interface, free text input, and sentence output [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. The latest version of this system, ChatGPT-3.5, has been made available for free use since March 2023.\u003c/p\u003e \u003cp\u003eSince the introduction of advanced AI applications like ChatGPT, the prominent feature that stands out in the CDS field is its use for case-based diagnosis and differential diagnosis purposes [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. However, another important feature of ChatGPT that can be preferred is its ability to instantly provide users with desired information on any specific topic. Despite the abundance of information and resources in the health and medical field, the difficulty of accessing accurate information easily remains a challenge. Although the widespread use of the internet facilitates access to information, clinicians may still have difficulty accessing specific information today. Accessing accurate information quickly is the most basic need, especially for clinicians working in chaotic environments such as emergency departments. Querying conducted through ChatGPT, which can be referred to as consultation or advisory, will be one of the important areas of use of AI applications in emergency medicine, enabling clinicians to access up-to-date and accurate information within seconds.\u003c/p\u003e \u003cp\u003e In our study, we aimed to test the proficiency of ChatGPT-3.5, one of the most advanced AI applications of today, in terms of its knowledge of current information based on the American Heart Association (AHA) 2020 Advanced Cardiac Life Support (ACLS) guidelines. For this purpose, a quiz prepared in both Turkish and English was presented to ChatGPT-3.5. The same quiz was also given to emergency medicine specialists to compare ChatGPT-3.5's success rates in Turkish and English. Additionally, considering that clinicians use ChatGPT-3.5 for queries in their native language, we aimed to evaluate the success levels in two different languages.\u003c/p\u003e"},{"header":"2. Methods","content":"\u003cp\u003eOur study was conducted at the Emergency Medicine Department of Haydarpaşa Numune Training and Research Hospital in March 2023.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. Study Design\u003c/h2\u003e \u003cp\u003eA quiz consisting of 80 questions in total, divided into 8 different chapters that aim to query current practices regarding the American Heart Association (AHA) 2020 Advanced Cardiac Life Support (ACLS) application steps, was prepared for use in our study. [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. The prepared quiz was administered to both ChatGPT-3.5 and two emergency medicine specialists with a minimum of 5 years of experience who were not informed about the study protocol. The quiz results were evaluated by a blinded researcher. In the study, first, the accuracy of ChatGPT-3.5's answers was evaluated with reference to the AHA 2020 ACLS guidelines. In addition, the inter-rater agreement of the items included in the quiz was examined. Then, the accuracy rates of ChatGPT-3.5 and emergency medicine specialists' answers were compared (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The accuracy rates were evaluated both on individual question level and chapter level, and reported as 'success percentage'. An accuracy rate of 80% and above was considered successful in line with the literature [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Quiz Components\u003c/h2\u003e \u003cp\u003eAn 80-question quiz was created in a question-answer format that includes the current ACLS application steps. These steps were designed in a total of 8 chapters covering general information on current ACLS, basic-advanced airway management, high-quality cardiopulmonary resuscitation (CPR), ventilation, defibrillation, medications, vascular access and return of spontaneous circulation (ROSC), and finally questions related to class of recommendation. The quiz includes 5 questions testing general information on current ACLS guidelines, 7 questions on basic-advanced airway management, 11 questions on high-quality CPR, 8 questions on ventilation, 12 questions on defibrillation, 22 questions on medications, 6 questions on vascular access and ROSC, and finally 9 questions related to class of recommendation, in total 80 questions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Obtaining Data\u003c/h2\u003e \u003cp\u003eThe quiz was administered using the ChatGPT-3.5 March 2023 version. It was applied to the AI application in both English (ChatGPT-3.5 English) and native language (ChatGPT-3.5 Turkish), while for emergency medicine specialists, the questions were prepared only in the native language. The quiz questions were prepared to have a definitive, unchanging, and non-interpretable answer, and were transferred to ChatGPT-3.5 in plain text format, and the answers were evaluated for compliance with the AHA 2020 ACLS guidelines (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Outcomes\u003c/h2\u003e \u003cp\u003e The primary aim of the study is to examine the accuracy of ChatGPT-3.5 in providing answers to questions related to current ACLS guidelines. The secondary aim is to compare the current knowledge level of experienced emergency medicine specialists with ChatGPT-3.5 regarding ACLS. The accuracy of the answers to questions related to ACLS was evaluated based on the 'Adult Basic and Advanced Life Support: 2020 AHA Guidelines for Cardiopulmonary Resuscitation and Emergency Cardiovascular Care' guide [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Statistical analysis","content":"\u003cp\u003eIn our study, categorical data were expressed as numbers (n) and percentages (%). Chi-square test was used for comparing the data. The agreement between specialists and ChatGPT-3.5 for categorical data was evaluated with inter-rater agreement. Inter-rater agreement was expressed with Kappa value and 95% Confidence interval (CI), and the significance level was accepted as p\u0026thinsp;\u0026lt;\u0026thinsp;0.05. Data analysis was performed using MedCalc Version 20.218 (MedCalc Software Ltd, Ostend, Belgium) and IBM SPSS version 20 (IBM Corp, Armonk, NY) programs.\u003c/p\u003e"},{"header":"4. Results","content":"\u003cp\u003eThe success rates for the entire quiz were as follows: for I.specialist, 81.3% (65/80); for II.specialist, 87.5% (70/80); for ChatGPT-3.5 in Turkish, 81.3% (65/80). The success rate for ChatGPT-3.5 in English was calculated as 86.3% (69/80). There was no statistically significant difference in success rates between I.specialist and II.specialist (p\u0026thinsp;=\u0026thinsp;0.067). Similarly, there was no statistically significant difference in success rates between I.specialist and ChatGPT-3.5 in Turkish (p\u0026thinsp;=\u0026thinsp;0.386), or between II.specialist and ChatGPT-3.5 in Turkish (p\u0026thinsp;=\u0026thinsp;0.332). The distribution of correct answer rates and success rates by section is shown in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e. In addition, in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e, the success rates of specialists and ChatGPT-3.5 Turkish were compared chapter by chapter, and no statistically significant difference was found between the two groups.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab1\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eComparison of success percentages of ChatGPT-3.5 and specialists according to Quiz sections\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eChapters\u003c/p\u003e\n\u003cp\u003e(questions)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eI. Specialist\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eII. Specialist\u003c/p\u003e\n\u003c/th\u003e\n\u003cth colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eChatGPT-3.5\u003c/p\u003e\n\u003cp\u003eTurkish\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eChatGPT-3.5 English\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eP value*\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eA \u0026ndash; General Overview for Current AHA Guıdeline (n\u0026thinsp;=\u0026thinsp;5)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5/5 (%100)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e5/5 (%100)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5/5\u003c/p\u003e\n\u003cp\u003e(%100)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5/5\u003c/p\u003e\n\u003cp\u003e(%100)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eB \u0026ndash; Airway Management (n\u0026thinsp;=\u0026thinsp;7)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5/7 (%71.4)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e7/7 (%100)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7/7\u003c/p\u003e\n\u003cp\u003e(%100)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7/7\u003c/p\u003e\n\u003cp\u003e(%100)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.110\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eC \u0026ndash; High Quality CPR (n\u0026thinsp;=\u0026thinsp;11)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8/11 (%72.7)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e10/11 (%90.9)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e11/11\u003c/p\u003e\n\u003cp\u003e(%100)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e10/11\u003c/p\u003e\n\u003cp\u003e(%90.9)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.137\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eD \u0026ndash; Ventilation (n\u0026thinsp;=\u0026thinsp;8)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7/8 (%87.5)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e8/8\u003c/p\u003e\n\u003cp\u003e(%100)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5/8\u003c/p\u003e\n\u003cp\u003e(%62.5)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8/8\u003c/p\u003e\n\u003cp\u003e(%100)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.122\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eE \u0026ndash; Defibrillation (n\u0026thinsp;=\u0026thinsp;12)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e9/12\u003c/p\u003e\n\u003cp\u003e(%75)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e10/12 (%83.3)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e10/12\u003c/p\u003e\n\u003cp\u003e(%83.3)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e9/12\u003c/p\u003e\n\u003cp\u003e(%75)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.837\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eF \u0026ndash; Medications (n\u0026thinsp;=\u0026thinsp;22)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e20/22 (%90.9)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e18/22 (%81.8)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e16/22\u003c/p\u003e\n\u003cp\u003e(%72.7)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e19/22\u003c/p\u003e\n\u003cp\u003e(%86.4)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.295\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eG \u0026ndash; Vascular Access and ROSC (n\u0026thinsp;=\u0026thinsp;6)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5/6 (%83.3)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e4/6 (%66.6)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6/6\u003c/p\u003e\n\u003cp\u003e(%100)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5/6\u003c/p\u003e\n\u003cp\u003e(%83.3)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.301\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eH \u0026ndash; Class (Strength) of Recommendation (n\u0026thinsp;=\u0026thinsp;9)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6/9 (%66.6)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003e8/9 (%88.8)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5/9\u003c/p\u003e\n\u003cp\u003e(%55.5)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6/9\u003c/p\u003e\n\u003cp\u003e(%66.6)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.288\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eAHA: American Heart Association, ROSC: Return of spontaneous circulation, CPR: Cardiopulmonary resuscitation. The ratio of correct answers to all answers is given in parentheses as a percentage. *P values obtained by comparing the success rates of ChatGPT-3.5 Turkish with the success rates of specialists.\u003c/p\u003e\n\u003cp\u003eThe Kappa values obtained from the comparison of the correct/incorrect answers given to the questions in terms of interrater agreement are shown in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e. Only the answers of ChatGPT-3.5 to the Turkish versions of the questions and the English versions of the questions showed fair agreement (Kappa\u0026thinsp;=\u0026thinsp;0.27) (p\u0026thinsp;=\u0026thinsp;0.015).\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab2\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eInter-rater aggrement Kappa values\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eAgreement\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eKappa Values\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e95% Confidence Interval\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eI.specialist \u0026ndash; II.specialist\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.20\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-0.06-0.46\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eI.specialist \u0026ndash; ChatGPT-3.5 Turkish\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.09\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-0.14-0.33\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eII.specialist \u0026ndash; ChatGPT-3.5 Turkish\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.10\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-0.13-0.35\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eChatGPT-3.5 Turkish- ChatGPT-3.5 English\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.27\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.00-0.53\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eThe number of questions for which both emergency medicine specialists gave wrong answers was 4. It was observed that ChatGPT-3.5 Turkish also could not provide the correct answer to one of these questions (p\u0026thinsp;=\u0026thinsp;0.742). This question was \"What is the maximum initial dose for biphasic defibrillators according to the AHA 2020 ACLS algorithm?\" However, it was seen that ChatGPT-3.5 gave the correct answer to this question in the English quiz.\u003c/p\u003e\n\u003cp\u003eThe number of questions on which both specialists answered correctly was 59, and when examined by chapters, it was found that the common correct answer rate of both specialists was 5 out of 5 (100%) in the 'A - General Overview for Current AHA Guideline' chapter. The chapter with the second highest common correct answer rate of both specialists (87.5%) was the 'D - Ventilation' chapter. It was found that ChatGPT-3.5 Turkish gave the correct answer to 50 out of the 59 questions on which both specialists answered correctly (p\u0026thinsp;=\u0026thinsp;0.179).\u003c/p\u003e"},{"header":"5. Discussion","content":"\u003cp\u003eAs far as we know, our study is the first to evaluate AHA 2020 ACLS guidelines with ChatGPT-3.5. We observed a similar success rate of over 80% in all questions asked to ChatGPT-3.5 and two emergency medicine specialists working independently in different hospitals and who do not know each other. ChatGPT-3.5 answered all questions related to General Overview for Current AHA Guideline, Airway Management, and Ventilation chapters with 100% success rate in English, while it answered all questions related to General Overview for Current AHA Guideline, Airway Management, High Quality CPR, and Vascular Access and ROSC chapters with 100% success rate in Turkish. Based on the data we obtained, we found that ChatGPT-3.5 can provide highly accurate and up-to-date answers to questions about current ACLS practices. In addition, similar success rates were obtained when compared to emergency medicine specialists.\u003c/p\u003e\n\u003cp\u003eThe contribution of AI to preventing human labor and time loss is one of the most important features that makes current studies focus on it. The promise of AI applications in health and medical services is to provide opportunities for increasing gains for patients and the clinical team, reducing costs, and improving public health. The most important prerequisites for the successful use of AI applications in medicine are accessibility, standardization, quality, and encouraging data that represents the population [\u003cspan class=\"CitationRef\"\u003e11\u003c/span\u003e]. There are many types of studies conducted with ChatGPT, including creating triage and differential diagnosis lists [\u003cspan class=\"CitationRef\"\u003e6\u003c/span\u003e], querying drug interactions [\u003cspan class=\"CitationRef\"\u003e12\u003c/span\u003e], AI-generated article writing [\u003cspan class=\"CitationRef\"\u003e13\u003c/span\u003e], and medical problem and case-solving [\u003cspan class=\"CitationRef\"\u003e14\u003c/span\u003e]. In general, the studies conducted through ChatGPT emphasize the CDS aspect, in which ChatGPT makes inferences about variable, fictional cases or situations and the accuracy of these inferences are evaluated. Studies that emphasize the evaluation of using ChatGPT as a reference source, rather than querying variable scenarios, like our study, will contribute to the literature in terms of using AI applications for a different purpose.\u003c/p\u003e\n\u003cp\u003eThe language selected for using ChatGPT as a reference source should be English, which is the main language of medical literature. In our study, the difference in accuracy rates between querying the same question in the native language and English (86.3% accuracy in English and 81.3% in the native language) supports the need for English as the selected language. ChatGPT\u0026apos;s use of plain text sources may have contributed to the poor performance in the \u0026quot;Class of Recommendation\u0026quot; section, which was the least successful section in the Quiz for ChatGPT-3.5 (66.6% accuracy in English and 55% accuracy in Turkish) and led to a low rate of correct answers by specialists (66.6% and 88.8%, respectively). One possible reason for this may be that recommendations in guidelines are usually presented in tables, while ChatGPT prefers to use plain text instead of tables or figures as a source. Therefore, in medical literature queries performed through ChatGPT-3.5, it is essential to verify the accuracy of the information contained in tables, and it should not be forgotten that ChatGPT may provide incorrect answers.\u003c/p\u003e\n\u003cp\u003eOne of the other important points revealed by our study is the existence of incorrect answers alongside the correct answers provided by ChatGPT-3.5. This poses a risk of misleading the physician or user. In our study, ChatGPT-3.5 provided correct answers to 69 out of 80 questions asked in English and gave incorrect answers to 11 questions. This situation deviates from the idea that ChatGPT can be used as a definitive reference source. However, considering the similarity of the correct answer rates at the expert level, and with the development of algorithms and enriched databases, the reduction in incorrect answer rates in the coming years may bring back the possibility of using ChatGPT as a reference source.\u003c/p\u003e"},{"header":"6. Conclusion","content":"\u003cp\u003e Our study showed that ChatGPT-3.5 has a similar level of accurate and up-to-date knowledge as an experienced emergency medicine specialist on the AHA 2020 Advanced Cardiac Life Support Guidelines. With the development of algorithms and new versions, ChatGPT's mastery of current information can be increased, and the number of incorrect answers can be reduced. Querying current guideline information through ChatGPT can serve as a consultant function easily accessible to emergency physicians. We believe that our study sheds light on the idea that ChatGPT can be used as a portal to instantly access accurate and up-to-date information based on textbooks and guidelines in the coming years.\u003c/p\u003e"},{"header":"7. Limitations","content":"\u003cp\u003eOne of the main limitations of our study is that ChatGPT-3.5 has not received clinical approval for obtaining healthcare information. Although our study has achieved successful results with ChatGPT-3.5, it should be kept in mind that AI applications, including ChatGPT, must be used with appropriate methods, and that ChatGPT is still being developed and its answers may be incorrect. One of the important limitations of ChatGPT is that its incorrect answers can lead to incorrect guidance and medical faults [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. One of the other important limitations of our study is the language issue. Different results can be obtained when queries are performed in the native language compared to when they are performed in English. We believe that queries should be conducted in English due to the fact that medical literature is written in English, and medical terms should be searched using the forms used in the literature. In addition, another limitation of our study is that no time limit was set for both ChatGPT-3.5 and specialists to answer each question in the quiz. If time data is added to the study, it may be possible to compare the response times of AI and specialists.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAuthors declare no conflict of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgment funding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors did not apply for a specific grant for this research from any funding agency in the public, commercial or non-profit sectors.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAmisha, Malik P, Pathania M, Rathaur VK. Overview of artificial intelligence in medicine. J Family Med Prim Care\u003cem\u003e.\u003c/em\u003e 2019;8(7):2328-31. https://doi.org/10.4103/jfmpc.jfmpc_440_19.\u003c/li\u003e\n\u003cli\u003eReading Turchioe M, Volodarskiy A, Pathak J, Wright DN, Tcheng JE, Slotwiner D. Systematic review of current natural language processing methods and applications in cardiology. Heart\u003cem\u003e.\u003c/em\u003e 2022;108(12):909-16. https://doi.org/10.1136/heartjnl-2021-319769.\u003c/li\u003e\n\u003cli\u003eGon\u0026ccedil;alves LS, Amaro MLM, Romero ALM, Schamne FK, Fressatto JL, Bezerra CW. Implementation of an Artificial Intelligence Algorithm for sepsis detection. Rev Bras Enferm\u003cem\u003e.\u003c/em\u003e 2020;73(3):e20180421. https://doi.org/10.1590/0034-7167-2018-0421.\u003c/li\u003e\n\u003cli\u003eDemner-Fushman D, Chapman WW, McDonald CJ. What can natural language processing do for clinical decision support? J Biomed Inform\u003cem\u003e.\u003c/em\u003e 2009;42(5):760-72. https://doi.org/10.1016/j.jbi.2009.08.007.\u003c/li\u003e\n\u003cli\u003eNath S, Marie A, Ellershaw S, Korot E, Keane PA. New meaning for NLP: the trials and tribulations of natural language processing with GPT-3 in ophthalmology. Br J Ophthalmol\u003cem\u003e.\u003c/em\u003e 2022;106(7):889-92. https://doi.org/10.1136/bjophthalmol-2022-321141.\u003c/li\u003e\n\u003cli\u003eHirosawa T, Harada Y, Yokose M, Sakamoto T, Kawamura R, Shimizu T. Diagnostic Accuracy of Differential-Diagnosis Lists Generated by Generative Pretrained Transformer 3 Chatbot for Clinical Vignettes with Common Chief Complaints: A Pilot Study. Int J Environ Res Public Health\u003cem\u003e.\u003c/em\u003e 2023;20(4)https://doi.org/10.3390/ijerph20043378.\u003c/li\u003e\n\u003cli\u003eGoodwin TR, Harabagiu SM. Medical Question Answering for Clinical Decision Support. Proc ACM Int Conf Inf Knowl Manag\u003cem\u003e.\u003c/em\u003e 2016;2016:297-306. https://doi.org/10.1145/2983323.2983819.\u003c/li\u003e\n\u003cli\u003eXu G, Rong W, Wang Y, Ouyang Y, Xiong Z. External features enriched model for biomedical question answering. BMC Bioinformatics\u003cem\u003e.\u003c/em\u003e 2021;22(1):272. https://doi.org/10.1186/s12859-021-04176-7.\u003c/li\u003e\n\u003cli\u003ePanchal AR, Bartos JA, Caba\u0026ntilde;as JG, Donnino MW, Drennan IR, Hirsch KG, et al. Part 3: Adult Basic and Advanced Life Support: 2020 American Heart Association Guidelines for Cardiopulmonary Resuscitation and Emergency Cardiovascular Care. Circulation\u003cem\u003e.\u003c/em\u003e 2020;142(16_suppl_2):S366-S468. https://doi.org/10.1161/CIR.0000000000000916.\u003c/li\u003e\n\u003cli\u003eBanik R, Rahman M, Sikder MT, Rahman QM, Pranta MUR. Knowledge, attitudes, and practices related to the COVID-19 pandemic among Bangladeshi youth: a web-based cross-sectional analysis. Z Gesundh Wiss\u003cem\u003e.\u003c/em\u003e 2023;31(1):9-19. https://doi.org/10.1007/s10389-020-01432-7.\u003c/li\u003e\n\u003cli\u003eMatheny ME, Whicher D, Thadaney Israni S. Artificial Intelligence in Health Care: A Report From the National Academy of Medicine. JAMA\u003cem\u003e.\u003c/em\u003e 2020;323(6):509-10. https://doi.org/10.1001/jama.2019.21579.\u003c/li\u003e\n\u003cli\u003eJuhi A, Pipil N, Santra S, Mondal S, Behera JK, Mondal H. The Capability of ChatGPT in Predicting and Explaining Common Drug-Drug Interactions. Cureus\u003cem\u003e.\u003c/em\u003e 2023;15(3):e36272. https://doi.org/10.7759/cureus.36272.\u003c/li\u003e\n\u003cli\u003eDergaa I, Chamari K, Zmijewski P, Ben Saad H. From human writing to artificial intelligence generated text: examining the prospects and potential threats of ChatGPT in academic writing. Biol Sport\u003cem\u003e.\u003c/em\u003e 2023;40(2):615-22. https://doi.org/10.5114/biolsport.2023.125623.\u003c/li\u003e\n\u003cli\u003eSinha RK, Deb Roy A, Kumar N, Mondal H. Applicability of ChatGPT in Assisting to Solve Higher Order Problems in Pathology. Cureus\u003cem\u003e.\u003c/em\u003e 2023;15(2):e35237. https://doi.org/10.7759/cureus.35237.\u003c/li\u003e\n\u003cli\u003eKing MR. The Future of AI in Medicine: A Perspective from a Chatbot. Ann Biomed Eng\u003cem\u003e.\u003c/em\u003e 2023;51(2):291-5. https://doi.org/10.1007/s10439-022-03121-w.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"artificial intelligence, AI chatbot, generative pretrained transformer, clinical decision support, guidelines","lastPublishedDoi":"10.21203/rs.3.rs-3035900/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3035900/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eObjectives: Artificial intelligence (AI) has become the focus of current studies, particularly due to its contribution in preventing human labor and time loss. The most important contribution of AI applications in the medical field will be to provide opportunities for increasing clinicians' gains, reducing costs, and improving public health. This study aims to assess the proficiency of ChatGPT-3.5, one of the most advanced AI applications available today, in its knowledge of current information based on the American Heart Association (AHA) 2020 guidelines.\u003c/p\u003e\n\u003cp\u003eMethods: An 80-question quiz in a question-and-answer format, which includes the current AHA 2020 application steps, was prepared and applied to ChatGPT-3.5 in both English (ChatGPT-3.5 English) and native language (ChatGPT-3.5 Turkish) versions in March 2023. The questions were prepared only in the native language for emergency medicine specialists.\u003c/p\u003e\n\u003cp\u003eResults: We found a similar success rate of over 80% in all questions asked to ChatGPT-3.5 and two independent emergency medicine specialists with at least 5 years of experience who did not know each other. ChatGPT-3.5 achieved a 100% success rate in all questions related to the General Overview for Current AHA Guideline, Airway Management, and Ventilation chapters in English.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eConclusions: Our study indicates that ChatGPT-3.5 provides similar accurate and up-to-date responses as experienced emergency specialists in the AHA 2020 Advanced Cardiac Life Support Guidelines. This suggests that with future updated versions of ChatGPT, instant access to accurate and up-to-date information based on textbooks and guidelines will be possible.\u003c/p\u003e","manuscriptTitle":"Assessing the Competence of ChatGPT-3.5 Artificial Intelligence System in Executing the ACLS Protocol of the AHA 2020","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-06-09 18:22:28","doi":"10.21203/rs.3.rs-3035900/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"d880a44c-f825-48bb-9200-633bfb1a8313","owner":[],"postedDate":"June 9th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2023-07-06T16:29:27+00:00","versionOfRecord":[],"versionCreatedAt":"2023-06-09 18:22:28","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3035900","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3035900","identity":"rs-3035900","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0