Credit
Natalie D. Cohen: Writing – review & editing, Writing – original draft, Methodology, Data curation. Milan Ho: Writing – review & editing, Writing – original draft. Donald McIntire: Validation, Formal analysis. Katherine Smith: Writing – original draft, Project administration, Conceptualization. Kimberly A. Kho: Writing – review & editing, Writing – original draft, Supervision, Methodology, Investigation, Data curation, Conceptualization.
Comment
The overall performance of LLMs was positive, with most giving accurate, although not comprehensive, answers to the questions. There were no questions the models could answer comprehensively or correctly across the reviewers. Different models displayed different levels of accuracy and comprehensiveness concerning patient questions about endometriosis. There were, however, numerous questions that led to inaccurate and not comprehensive answers by the LLMs. Only 1 of the 10 questions that the reviewers agreed with was answered completely correctly (“What is endometriosis?”) by all three chatbots with average scores above 4. The rest of the questions had average scores of below 4 for at least one of the chatbots. Only 3 out of 10 questions for Bard and Claude had average scores above 4. These results underscore the importance of continued evaluation of LLMs.
While most questions did not show differences between reviewer evaluation scores, two were found to have significant differences in the reviewer ratings. A possible explanation for this includes different reviewer interpretations of answers. For example, one of the questions with disagreement was, “Can endometriosis come back after a hysterectomy.” Some reviewers may interpret this question more nuanced than others, with chatbot answers not being comprehensive if there was no mention of age at the time of hysterectomy or oophorectomy at the time of hysterectomy. It is also possible that the disagreement reflects debate within the endometriosis community regarding the diagnosis of endometriosis or the risk of cancer. We attempted to mitigate eminence-based bias between reviewers by incorporating multiple guidelines.
Though these LLMs source information from the Internet, including medical textbooks and scientific publications, the information presented may not always be accurate or comprehensive. The cause of inadequacies and inaccuracies in the LLM responses is likely multifactorial. LLMs are trained from datasets that do not capture the most recent advances in clinical practice, leading to differing evidence bases for diagnosis and management guidelines. For example, ChatGPT is not continuously updated; rather, the data was most recent from 2021 when this study was conducted. 3 There are currently significant efforts by OpenAI to create more updated knowledge bases at faster intervals. 11 LLMs create generative responses in real-time. Thus, there is no fact-checking method within the program as they develop responses. There is also evidence of model “hallucinations” in which LLMs may fabricate data that is referenced. 12 Of note, no hallucinations were noted in the evaluation of responses. Currently, a warning on the ChatGPT interface states “ChatGPT can make mistakes. Consider checking important information.” The other two chatbots do not have such a warning. LLMs have been shown to invent terms and make incorrect assumptions to answer a user's questions. 3 Of note, chatbots do not list sources for their information.
Interestingly, there were more difficulties in the LLMs’ responses to questions about treatment rather than symptoms and disease process of endometriosis. The treatment of endometriosis is varied, nuanced, and can be complex. Debates on best practices related to medical and surgical management may be reflected in the dissimilar responses. Many chatbots do not have access to research articles that healthcare providers use to guide patients, such as professional society guidelines and committee opinions, which may be restricted to members only, and thus, the chatbots only use the training data provided. It is also possible that generative technology has more difficulty conveying nuanced answers to questions about endometriosis treatment than questions about pathophysiology and symptoms.
Since their release, LLMs have been questioned regarding the accuracy of the information presented due to known differences in training data. ChatGPT has been at the forefront of the literature on the evaluation of LLMs and has been regarded to be accurate but not necessarily comprehensive. 13 , 14 , 15 Other models, such as Claude by Anthropic and Bard by Google, remain under investigation both as standalone LLMs in comparison to other LLMs. Given the recent introduction of LLM chatbots, comparatively fewer studies compare LLM chatbot responses to physician responses in answering patient questions. 12
While this is the first study comparing multiple LLM answers to questions on endometriosis, previous literature has examined LLM responses to other conditions. 16 , 17 , 18 , 19 , 20 , 21 LLM (including ChatGPT and Bard) responses to patient questions have been evaluated for accuracy and comprehensiveness in topics such as colon cancer and colonoscopy, lung cancer, cirrhosis and hepatocellular carcinoma, head and neck cancer, breast reconstruction surgery, and total hip arthroplasty. 16 , 17 , 18 , 19 , 20 , 21 Most of the data currently available on LLM is with ChatGPT, and other models are not as well studied. 12 In our study, all the models exhibited some limitations. Our analysis reinforces prior findings that LLM responses are consistently accurate but require additional nuance and clarification when evaluated by physicians.
This study shows that much of the information presented by LLMs is accurate, which may help facilitate patient education. As most LLMs are currently free of charge, the ready availability of this information could help patients access care and information.
However, this study reveals gaps and inaccuracies in the medical information provided by leading LLMs. Most of the answers were not comprehensive and thus cannot replace patient-specific counseling by healthcare experts. In addition, clinicians should be aware of how patients’ utilization of LLMs may impact their knowledge of the disease.
LLMs are only as successful as the training data provided to them. This study underlines the idea that if the LLMs have greater access to current data and the expertise of specialists, they may be able to provide more accurate and comprehensive information. This includes continuous updating of the knowledge base LLMs utilize and greater access to recent literature and guidelines.
As chatbots become popularized and more widely used, gynecologists must work to improve both patients’ understanding of the limitations of responses they receive from LLMs and the accuracy of these responses. Specialist expertise and feedback should be integrated into fine-tuning LLM development and training data.
This study shows there is much to be learned about LLMs and how they may positively or negatively be incorporated into our patients’ understanding of endometriosis. While this study compared three leading LLMs, these models are continuously updated, and other models are being created. For example, the company Meta (Meta Platforms, Inc. Menlo Park, CA) recently released their AI chatbot, Llama, an open-source LLM. 22 Limited research has been done on customized language models that may be better tailored to specific topics as they are modeled. Proponents of open sourcing believe that increasing access to the underlying weightings and technologies improves access to and enhances the speed of innovation as users can use the models to serve their own purposes.
This study focused on commonly asked questions about endometriosis, but there are nuances to the questions that may be asked that could be evaluated in future research. Patients may utilize different language and syntax to ask the same question, which is left to interpretation by natural language processing and can yield different results. Further research could also elucidate how patient experiences obtaining information from healthcare providers can be compared to or complemented by LLMs.
This study has several strengths. This is the first study to comparatively evaluate the accuracy and comprehensiveness of LLM responses to patient questions on endometriosis. This study used three different LLMs to assess the technology's accuracy across various platforms and interfaces. Most research thus far has focused primarily on ChatGPT, whereas this study incorporated other LLMs. Additionally, using 10 depersonalized questions allowed for reproducibility and further study as new chatbots become available and the LLMs are updated with more current information. Lastly, all the evaluators were experienced academic gynecologists with sizeable clinical volume and expertise in endometriosis care.
This study does have limitations. First, though our reviewer base of gynecologists provides an expert opinion on the quality of chatbot responses, the determination of a response's quality is inevitably influenced by the individual perspectives of our reviewers. We aimed to reduce the effects of any individual reviewer's bias by including nine reviewers and incorporating current guidelines into the evaluation process. However, there were still two questions that showed significant disagreement between reviewers. Second, our 10 sample questions cannot capture the full scope of patients’ inquiries about endometriosis. In addition, the choice of questions was uniform and may not reflect patients’ semantics and word choices at various reading levels. Previous work has demonstrated that chatbot prompt differences can lead to varying outputs, and higher quality outputs can lead to higher quality responses. 23 At the time of writing this article, there is no known validated evaluation model for LLM responses.
Results
Results are presented in Table . There was no question that the models could provide comprehensive and correct answers to all the reviewers. Overall, there was consistency in the LLMs responding at least “mostly correct” for all but two questions. The model most associated with comprehensive and correct responses on average was ChatGPT. Table Comparison of responses Table Bard ChatGPT Claude Question Med [Q1, Q3] Mean ± std Med [Q1, Q3] Mean ± std Med [Q1, Q3] Mean ± std P W Friedman 1 What is endometriosis? 4 [4, 4] 4.0 ± 0.5 5 [4, 5] 4.3 ± 0.9 4 [4, 5] 4.2 ± 0.8 .779 0.676 2 How common is endometriosis? 4 [3, 4] 3.7 ± 0.9 4 [4, 4] 4.1 ± 0.6 5 [4, 5] 4.3 ± 1.1 .080 0.043 3 How do I know if I have endometriosis? 4 [4, 4] 3.9 ± 0.6 5 [5, 5] 4.8 ± 0.4 3 [3, 4] 3.2 ± 0.7 .017 0.002 4 What are the symptoms of endometriosis? 3 [3, 5] 3.8 ± 1.0 5 [4, 5] 4.6 ± 0.7 5 [4, 5] 4.6 ± 0.7 .264 0.041 5 How is endometriosis treated? 4 [4, 4] 3.9 ± 0.9 4 [4, 5] 4.2 ± 0.8 4 [4, 4] 3.7 ± 0.7 .256 0.119 6 How do I know what stage endometriosis I have? 3 [3, 4] 3.4 ± 1.0 4 [4, 5] 4.2 ± 0.7 4 [3, 4] 3.6 ± 0.9 .156 0.099 7 Can I get pregnant with endometriosis? 4 [4, 4] 4.0 ± 0.7 4 [4, 4] 4.0 ± 0.7 3 [3, 4] 3.4 ± 1.0 .264 0.102 8 Can endometriosis come back after hysterectomy? 3 [2, 3] 2.8 ± 0.8 3 [3, 4] 3.3 ± 0.9 3 [3, 4] 3.2 ± 0.8 .184 0.063 9 Does endometriosis increase my risk of cancer? 4 [2, 4] 3.0 ± 1.2 5 [4, 5] 4.6 ± 0.7 3 [2, 3] 2.9 ± 1.4 .017 0.007 10 Can endometriosis be cured? 5 [4, 5] 4.4 ± 1.0 5 [4, 5] 4.3 ± 0.9 4 [4, 5] 3.9 ± 1.3 .358 0.143 Med [Q1, Q3]: median [1st quartile, 3rd quartile]. Friedman, P value for the Friedman's test; Mean ± std : mean ± standard deviation; P W, P value for Kendall's W (also known as Kendall's coefficient of concordance). Cohen. A comparative analysis of generative artificial intelligence responses. AJOG Glob Rep 2024.
Comparison of responses
Med [Q1, Q3]: median [1st quartile, 3rd quartile].
Friedman, P value for the Friedman's test; Mean ± std : mean ± standard deviation; P W, P value for Kendall's W (also known as Kendall's coefficient of concordance).
There was some variability in the reviewers’ grading, with two of the questions noted to have significant differences. These questions were “How do I know if I have endometriosis?” and “Does endometriosis increase my risk of cancer?” For 8 of the 10 questions, the Friedman test did not indicate significant differences between reviewers. The Kendall's W test further substantiated these findings, indicating that there was evidence of significant disagreement between the graders for only those two questions previously noted.
Figure 2 presents the average scores for each model per question. The average scores for the 10 answers amongst Bard, Chat GPT, and Claude were 3.69, 4.24, and 3.7, respectively. The chatbots answered questions about the prevalence and symptoms of endometriosis more comprehensively; however, more discrepancy was noted with questions related to therapeutics and management of endometriosis. The question associated with the lowest scores across the reviewers was related to endometriosis recurrence after hysterectomy. Figure 2 Average score for each model for question. Figure 2 Cohen. A comparative analysis of generative artificial intelligence responses. AJOG Glob Rep 2024.
Average score for each model for question.
Materials
A set of 10 commonly asked questions about endometriosis were inputted to three LLMs in August of 2023. Questions were chosen by the research team. Models included Chat GPT-4 by the company Open AI (Open AI, Inc., San Francisco, CA), Claude by the company Anthropic (Anthropic, PBC, San Francisco, CA), and Bard by Google (Google, LLC, Mountain View, CA). The first generated response of each LLM was collected. An example response from each of the chatbots is shown in Figure 1 . Following the completed data collection, nine academic gynecologists with expertise in endometriosis evaluated each of the responses. The grading process was not blinded. Three guidelines on endometriosis treatment, including from the Cochrane database, the American Society of Reproductive Medicine, and the American College of Obstetrics and Gynecology, were provided to reviewers as examples of evidence-based expert-developed clinical guidelines for the evaluation process. 6 , 7 , 8 , 9 , 10 Figure 1 Example chatbot response to patient question “What is endometriosis?” Figure 1 Cohen. A comparative analysis of generative artificial intelligence responses. AJOG Glob Rep 2024.
Example chatbot response to patient question “What is endometriosis?”
The grading scale included the following: (1) Completely incorrect, (2) mostly incorrect and some correct, (3) mostly correct and some incorrect, (4) correct but inadequate, (5) correct and comprehensive. Institutional Review Board approval was not required for this study.
The individual 10 questions were examined independently to discern reviewer concordance across the three LLMs. Using Kendall's W for concordance, we evaluated the reviewer agreement in response across the three LLMs for each question. Kendall's W assessed inter-rater reliability with the significance of the related chi-square test indicating agreement among the raters. The P value was reported. Friedman's test was also utilized as a nonparametric one-way analysis of variance with repeated measures. The raters’ median and mean±standard deviation were presented for each LLM by each question. The P value of Friedman's test indicated whether there was a difference in the rating of the LLMs across the raters. The scale of the ratings was assumed ordinal in this case. Computations were conducted with SAS version 9.4, SAS Institute, Cary, NC.
Conclusions
Publicly accessible conversational chatbots provide an easy access point for patients seeking quick answers to questions about their health. The analysis of LLMs revealed that, on average, they offered mostly correct but inadequate responses to commonly asked patient questions on endometriosis. Of note, there were more difficulties in the LLMs’ responses to questions about treatment and recurrence of endometriosis rather than questions on symptoms and disease. There were differences in responses and subsequent scores for each model, revealing variability between the models.
While chatbot responses can serve as valuable supplements to information provided by licensed medical professionals, it is crucial to maintain a thorough ongoing evaluation process of outputs to provide the most comprehensive and accurate information to patients. Further research into this technology and its role in patient education and treatment is crucial as generative AI becomes more embedded in the medical field.
Introduction
The use of generative artificial intelligence (AI) chatbots came to the forefront in November 2022 with the release of ChatGPT by the company Open AI, a large language model (LLM) that quickly became one of the most rapidly expanding user bases within the first few months of its release. 1 Other companies have also released chatbots, such as Claude, by the company Anthropic and Bard, by Google, and they differ in architecture, natural language processing, and training data.
These LLMs offer a wide range of advantages in the medical field, including depth of knowledge, interaction, and presentation of complex medical information. In recent research, ChatGPT was found to pass the three exams of the USMLE, discuss public health topics with accurate definitions and examples of clinical studies, and demonstrate high diagnostic accuracy when presented with clinical vignettes. 2 These chatbots are expressive and interactive and have shown more empathy when comparing ChatGPT responses to physician-generated responses. 1 There have been efforts to incorporate chatbots into physician learning and clinical work because of their ability to translate complicated medical language into more accessible formats for patients. 3 , 4 However, with the advancement of this technology, there are concerns about its expandability, specifically that what patients expect chatbots to be able to answer differs from what they are capable of. 4
There are many applications of LLMs in Obstetrics and Gynecology, particularly in informing patients who may have undertreated and underdiagnosed conditions such as endometriosis. 3 Diagnostic delays in endometriosis from 7 to 9 years have been reported globally. 5 As a chronic disease that can be treated medically and surgically, as well as have a lasting impact throughout a patient's lifespan, accurate and comprehensive information is needed. While LLMs can offer an opportunity for patient education, what remains unclear is the current level of accuracy in the information presented to patients by LLMs related to endometriosis. This study aimed to evaluate three of the leading chatbots in terms of their ability to address questions pertaining to endometriosis. Specifically, we aimed to assess the accuracy and comprehensiveness of each LLM and determine the level of variability between the responses.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.