Accuracy and reproducibility of ChatGPT's free version answers about endometriosis

other OA: hybrid CC-BY-NC-ND-4.0
AI-generated summary by qwen3.7-flash, 2026-08-13

This study evaluated ChatGPT’s free version for endometriosis, finding high accuracy and reproducibility for general questions but lower performance on guideline-based inquiries regarding treatment and prevention.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-06, 2026-06-13 · read from full text

This study assessed the accuracy and reproducibility of ChatGPT (free version) answers to 81 endometriosis frequently asked questions gathered from internet sources and to 40 scientific questions derived from the ESHRE endometriosis guidelines. Answers were evaluated by an experienced gynecologist using a 1–4 grading scheme and reproducibility was tested by asking each question twice and checking whether the score category matched. ChatGPT answered 91.4% of FAQs completely accurately and sufficiently, with highest accuracy in symptoms/diagnosis and lowest in treatment, but only 67.5% of guideline-based questions were graded as completely correct; reproducibility was lowest for guideline-based questions (70.0%), and the authors note reliability concerns because much online content is not reviewed. This paper is centrally about endometriosis — evaluating how accurate and reproducible ChatGPT’s free-version responses are for endometriosis FAQs and ESHRE guideline-based questions.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

OBJECTIVE: To evaluate the accuracy and reproducibility of ChatGPT's free version answers about endometriosis for the first time. METHODS: Detailed internet searches to identify frequently asked questions (FAQs) about endometriosis have been performed. Scientific questions were prepared in accordance with the European Society of Human Reproduction and Embryology (ESHRE) endometriosis guidelines. An experienced gynecologist gave a score of 1-4 for each ChatGPT answer. The repeatability of ChatGPT answers about endometriosis was analyzed by asking each question twice, and the reproducibility of ChatGPT was accepted as scoring the answer to the same question in the same score category. RESULTS: A total of 91.4% (n = 71) of all FAQs were answered completely, accurately, and sufficiently. ChatGPT had the highest accuracy in the symptom and diagnosis category (94.1%, 16/17 questions) and the lowest accuracy in the treatment category (81.3%, 13/16 questions). Furthermore, of the 40 questions based on the ESHRE endometriosis guidelines, 27 (67.5%) were classified as grade 1, seven (17.5%) as grade 2, and six (15.0%) as grade 3. The reproducibility rate of FAQs in the prevention, symptoms, and diagnosis, and complications categories was the highest (100% for all categories). The reproducibility rate was the lowest for questions based on the ESHRE endometriosis guidelines (70.0%). CONCLUSION: ChatGPT accurately and satisfactorily responded to more than 90% of the questions about endometriosis, but to only 67.5% of questions based on the ESHRE endometriosis guidelines.
Full text 13,707 characters · extracted from oa-doi-fallback · 4 sections · click to expand

Abstract

Objective To evaluate the accuracy and reproducibility of ChatGPT's free version answers about endometriosis for the first time.

Methods

Detailed internet searches to identify frequently asked questions (FAQs) about endometriosis have been performed. Scientific questions were prepared in accordance with the European Society of Human Reproduction and Embryology (ESHRE) endometriosis guidelines. An experienced gynecologist gave a score of 1–4 for each ChatGPT answer. The repeatability of ChatGPT answers about endometriosis was analyzed by asking each question twice, and the reproducibility of ChatGPT was accepted as scoring the answer to the same question in the same score category.

Results

A total of 91.4% (n = 71) of all FAQs were answered completely, accurately, and sufficiently. ChatGPT had the highest accuracy in the symptom and diagnosis category (94.1%, 16/17 questions) and the lowest accuracy in the treatment category (81.3%, 13/16 questions). Furthermore, of the 40 questions based on the ESHRE endometriosis guidelines, 27 (67.5%) were classified as grade 1, seven (17.5%) as grade 2, and six (15.0%) as grade 3. The reproducibility rate of FAQs in the prevention, symptoms, and diagnosis, and complications categories was the highest (100% for all categories). The reproducibility rate was the lowest for questions based on the ESHRE endometriosis guidelines (70.0%).

Conclusion

ChatGPT accurately and satisfactorily responded to more than 90% of the questions about endometriosis, but to only 67.5% of questions based on the ESHRE endometriosis guidelines. 1 INTRODUCTION Endometriosis is a benign, chronic, inflammatory, estrogen-dependent disease wherein the endometrial glands and stroma are found outside the endometrial cavity. Endometriosis can involve distinct organs, for example, the brain, intestines, bladder, and eye, and cause anatomical or functional malfunctions depending on the organ involved.1 Furthermore, endometriosis leads to prolonged drug use, increased doctor visits, social isolation, and worsening of relationships. Estes et al. found that women with endometriosis lost more hours of employment productivity and had lower income than women without endometriosis.2 Nnoaham et al. stated that the educational life of one out of every four students was interrupted due to endometriosis-related symptoms.3 Ozgor et al. assessed the prevalence of endometriosis and found that one out of every six women had endometriosis-related symptoms. Moreover, the prevalence of endometriosis increases by up to 30% in cases with infertility and 45% in cases with chronic pelvic pain.4 Internet-based applications have recently become popular as the most frequently used resources by patients. ChatGPT (Open AI, San Francisco, CA, USA) is a relatively new artificial intelligence (AI) application with natural language processing; nonetheless, the potential advantages and limitations of ChatGPT in medical fields are still under investigation.5 Caglar et al. found that the answers provided by ChatGPT to patients' questions about pediatric urology and questions prepared based on pediatric urology guidelines were accurate and adequate.6 Moreover, Gilson et al. revealed that ChatGPT was successful in answering medical school examination questions.7 ChatGPT could also interpret radiological imaging findings with an acceptable error rate.8 Although few studies have evaluated the accuracy and reproducibility of ChatGPT answers in some diseases, to our knowledge, no study has assessed the quality of ChatGPT answers about endometriosis. Thus, we have analyzed the accuracy and reproducibility of ChatGPT's free version answers about endometriosis for the first time. 2 MATERIALS AND METHODS We performed a detailed internet search to identify frequently asked questions (FAQs) about endometriosis. The internet sources consulted included websites established by professional health institutions or workers, social media platforms such as Instagram and Facebook, and websites that patients and their relatives frequently visited. We then prepared a questionnaire based on all identified questions (Data S1). Scientific questions were prepared in accordance with the European Society of Human Reproduction and Embryology (ESHRE) endometriosis guidelines and were listed in a separate form (Data S1). Questions on personal information, questions for advertising purposes, repetitive questions, unrealistic questions, and questions with obvious grammatical errors were excluded. All FAQs (n = 81) were categorized as general information (n = 20), symptoms and diagnosis (n = 17), treatment (n = 16), prevention (n = 15), and complications (n = 13). A total of 40 scientific questions were created according to the ESHRE endometriosis guidelines. All ChatGPT answers were analyzed and scored by experienced endometriosis gynecologists. An experienced gynecologist gave a score of 1–4 for each ChatGPT answer (completely true answers scored as 1, accurate answers including insufficient data scored as 2, answers with correct information and incorrect information scored as 3, and completely incorrect answers scored as 4). The answer was accepted as completely correct when no further contribution was made by the experienced gynecologist to ChatGPT's answer. In addition, the ChatGPT answers for scientific questions were scored in accordance with the ESHRE endometriosis guidelines. The repeatability of ChatGPT answers about endometriosis was analyzed by asking each question twice, and the reproducibility of ChatGPT was accepted as scoring the answer to the same question in the same score category. Different score categories for two different ChatGPT answers for the same question resulted in negative ChatGPT reproducibility. As the present study did not include patient data, informed consent, and ethics committee approval were not required. 2.1 Statistical analysis Statistical analysis was performed using Excel version 16 (Microsoft Corporation, Redmond, WA, USA). The questions were separately evaluated as FAQs, and questions were prepared based on the ESHRE endometriosis guidelines. The scores given to the answers are shown as percentages. We did not perform any statistical hypothesis testing, as it was not applicable to our evaluation. 3 RESULTS A total of 144 FAQs were evaluated for suitability to the study, of which the following were excluded: 16 reposted questions, 21 grammatically incorrect questions, 12 questions without objective answers, and 14 questions related to personal information. Finally, 81 FAQs were answered by ChatGPT. In addition, 40 scientific questions were prepared based on the ESHRE endometriosis guidelines. The flowchart of the study of FAQs is summarized in Figure 1. A total of 91.4% (n = 71) of all FAQs were answered completely accurately and sufficiently. Missing information was detected in only one (1.2%) ChatGPT answer; none of ChatGPT answers was classified as grade 4. ChatGPT had the highest accuracy in the symptom and diagnosis category (94.1%, 16/17 questions) and the lowest accuracy in the treatment category (81.3%, 13/16 questions). Furthermore, of the 40 questions based on the ESHRE endometriosis guidelines, 27 (67.5%) were classified as grade 1, seven (17.5%) as grade 2, and six (15.0%) as grade 3. None of the ChatGPT answers to questions based on the ESHRE endometriosis guidelines were scored as grade 4. The grading of the ChatGPT answers is presented in Table 1. | Grade 1 | Grade 2 | Grade 3 | Grade 4 | | |---|---|---|---|---| | All questions (n = 81) | 74 (91.4%) | 7 (8.6%) | 1 (1.2%) | - | | General information (n = 20) | 18 (90.0%) | 2 (10.0%) | - | - | | Symptoms and diagnosis (n = 17) | 16 (94.1%) | 1 (5.9%) | - | - | | Treatment (n = 16) | 13 (81.3%) | 2 (12.5%) | 1 (6.2%) | - | | Prevention (n = 15) | 14 (93.3%) | 1 (6.7%) | - | - | | Complications (13) | 12 (92.3%) | 1 (7.7%) | - | - | | ESHRE endometriosis guideline recommendations (n = 40) | 27 (67.5%) | 7 (17.5%) | 6 (15.0%) | - | - Note: Grade 1, completely correct; Grade 2, correct but insufficient; grade 3, misleading information as well as correct information; grade 4, completely incorrect; ESHRE, European Society of Human Reproduction and Embryology. The reproducibility rate of FAQs in the prevention, symptoms and diagnosis, and complications categories was the highest (100% for all categories). The repeatability rate was at its lowest for the treatment category (81.3%) in the FAQs. Moreover, the reproducibility rate was lowest for questions based on the ESHRE endometriosis guidelines (70.0%). The reproducibility rates of ChatGPT answers to questions are shown in Figure 2. 4 DISCUSSION Artificial intelligence, including its pros and cons, has become one of the hottest topics in the medical field. The correct and efficient use of AI can enable better use of screening tests, help to categorize patients according to risk classifications, make faster diagnoses of diseases, and shorten the time of patients to get treatment. In addition to these possible advantages, there are still doubts and concerns about the use of AI in medical practice.9 Thus, we performed the present study to analyze ChatGPT's answers about endometriosis, which affects approximately one-sixth of the female population. The present study's results demonstrated that ChatGPT provided completely correct and satisfactory answers to 91.4% of the FAQs about endometriosis. However, ChatGPT satisfactorily answered only 67.5% of the questions based on the ESHRE endometriosis guidelines. The reproducibility rate was highest for FAQs about prevention, symptoms and diagnosis, and complications and lowest for questions based on the ESHRE endometriosis guidelines. Most of the information on internet resources is not reviewed, which is a significant limitation in terms of the reliability of these resources. Cetin et al. assessed the quality of YouTube videos about coronary artery bypass grafting and concluded that most of these videos contained misleading information.10 Alsyouf et al. assessed social media content (Instagram, Facebook, Pinterest, and Twitter) about prostate cancer and found that it included almost 30-fold more inaccurate information than accurate information.11 Bulck and Moons analyzed the performance of ChatGPT's answers about cardiologic disorders and found that ChatGPT gave correct and adequate answers to 17 of the 20 questions.12 The present study was the first to evaluate the accuracy of ChatGPT answers about endometriosis and to demonstrate that ChatGPT provided accurate and satisfactory answers to 91.4% of FAQs. While internet users can access and acquire information from unspecified content on social media platforms such as YouTube and Twitter, ChatGPT can access a large number of resources and retrieve extensive knowledge by analyzing these resources. We believe that analyzing data from numerous resources provided a higher accuracy rate for ChatGPT's answers. The guidelines are sources that have emerged from the results of numerous meta-analyses, reviews, and original studies and they manage clinical practice with specific recommendations. The success of AI may decrease in relation to answering questions based on scientific data. Antaki et al. answered ophthalmology resident exam questions using ChatGPT and found that ChatGPT scored 55.8 on the exam, which is similar to the mean of an ophthalmology resident's score.13 Furthermore, ChatGPT gave complete, accurate answers to 93.6% of pediatric urology questions based on the European Urology Association Pediatric Urology guidelines.6 In the present study, the accuracy rate of ChatGPT answers to the questions based on the ESHRE endometriosis guidelines remained at 67.5%, which was lower than that of the FAQs. Although our study is the first to analyze the knowledge of ChatGPT in respect of endometriosis, it has certain limitations. First, English was the only language used in the present study, and only ChatGPT answers in the English language were evaluated. However, using multiple languages during the study can be confusing and cause difficulty when trying to demonstrate the results. Additionally, English is the most widely used language in the world and in the scientific literature. Second, the study contains information from only a certain period, and the information on the internet is constantly increasing and being updated. The evaluation of ChatGPT's answers may have involved subjective decisions, as with all person-based evaluations. The scoring system, however, has been adapted from similar previous studies.6 Lastly, the public intelligibility of ChatGPT's answers was not evaluated, which is something that should be assessed in the future. 5 CONCLUSION In summary, ChatGPT accurately and satisfactorily responded to more than 90% of the questions about endometriosis, but to only 67.5% of questions based on the ESHRE endometriosis guidelines. The reproducibility rate was highest for FAQs about prevention, symptoms and diagnosis, and complications and lowest for questions based on the ESHRE endometriosis guidelines. We believe that the use of AI under the supervision of a gynecologist will be beneficial in the diagnosis, treatment, and follow-up of endometriosis. AUTHOR CONTRIBUTIONS Bahar Yuksel Ozgor: design, planning, conduct, data analysis, and manuscript writing. Melek Azade Simavi: design, planning, conduct, and data analysis. FUNDING INFORMATION None. CONFLICT OF INTEREST STATEMENT The authors have no conflicts of interest. DATA AVAILABILITY STATEMENT Research data are not shared.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Condition tags

endometriosis

MeSH descriptors

Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-09-08T06:17:11.951632+00:00
pubmed
last seen: 2026-09-08T06:15:13.928625+00:00
unpaywall
last seen: 2026-05-14T19:30:52.867331+00:00
License: CC-BY-NC-ND-4.0 · commercial use OK · attribution required
Courtesy of the U.S. National Library of Medicine