Can AI speak endo? A multi-platform evaluation of large language models against ESHRE endometriosis guidelines
other
public-domain-us
Abstract
STUDY QUESTION: How do the three most widely accessible large language models perform in terms of accuracy, consistency, and reliability when answering clinically relevant questions derived from the 2022 ESHRE guideline on endometriosis?
SUMMARY ANSWER: Model A achieved the highest accuracy scores, model C demonstrated significantly superior consistency across repeated queries, and all three models showed comparable but suboptimal reliability under free-tier access, while under official API access median accuracy converged across providers and reliability increased.
WHAT IS KNOWN ALREADY: Large language models are increasingly consulted by both clinicians and patients as readily accessible sources of medical information. In reproductive medicine, preliminary evidence suggests that individual platforms may retrieve endometriosis-related content with mixed fidelity. However, no study has simultaneously benchmarked multiple models against a single, internationally recognized endometriosis guideline. The extent to which these tools can be trusted to faithfully reproduce evidence-based recommendations on endometriosis diagnosis and management remains largely unexplored.
STUDY DESIGN, SIZE, DURATION: Cross-sectional, multi-platform comparative study. Fifty clinically relevant questions covering the diagnostic and therapeutic domains of the 2022 ESHRE endometriosis guideline were simultaneously submitted to all three models during December 2025. Each question was entered in duplicate using independent sessions to assess response consistency and reliability. Subgroup analyses were carried out by submitting the questions to the API version and to the free-tier version of the platforms available in May 2026.
PARTICIPANTS/MATERIALS, SETTING, METHODS: The three models were accessed through their respective free web interfaces using new accounts, without prompt engineering, prior training, or retrieval-augmented generation. A zero-shot prompting approach was adopted. Accuracy was evaluated by two independent experts using the Global Quality Score (GQS); consistency was defined as identical responses across the three iterations; reliability was defined as the alignment of each response with the ESHRE guideline. Discrepancies were settled by a third reviewer.
MAIN RESULTS AND THE ROLE OF CHANCE: Significant differences in accuracy were observed across models (Kruskal-Wallis H = 37.10, P < 0.001). Model A achieved the highest median GQS (5, interquartile range [IQR] 4-5), followed by model C (4, IQR 3-5) and model B (3, IQR 3-4). Post-hoc analysis confirmed that model A significantly outperformed model B (P < 0.001) but not model C (P = 0.123), while model C also scored significantly higher than model B (P < 0.001). For consistency, model C demonstrated a significantly higher rate of reproducible responses (92.0%) compared with model A (72.0%, P = 0.028) and model B (68.0%, P = 0.008). No significant between-model differences were found for reliability (χ2 = 1.029, P = 0.598), with rates of 76.0% for model C, 68.0% for model A, and 68.0% for model B. In API and free-tier May 2026 subgroup analyses, median GQS converged to 4 across all three providers, and reliability rose to at least 76% for each provider.
LIMITATIONS, REASONS FOR CAUTION: Model outputs may change with subsequent updates. GQS retains a degree of subjectivity despite expert adjudication. The study tested factual recall rather than complex clinical reasoning, limiting generalizability to real-world decision-making scenarios.
WIDER IMPLICATIONS OF THE FINDINGS: These findings provide the first multi-platform benchmark of large language models against the latest endometriosis guideline. While models A and C retrieved guideline-concordant information with acceptable fidelity, none of the models achieved a level of reliability required for unsupervised clinical use. The dissociation between accuracy and consistency at the free tier, and its attenuation under API and more recent accesses, indicates that both model capability and commercial tier shape user-facing outputs. The dissociation between accuracy and consistency underscores that a model producing high-quality answers does not necessarily do so in a reproducible manner. Expert oversight remains crucial when interpreting their output regarding endometriosis care. Future research should extend this framework to additional endometriosis guidelines and to longitudinal monitoring of models' performance with their upgrades.
FUNDING: No external funding was received for this study.
DISCLOSURES: The authors declare no competing interests.
TRIAL REGISTRATION NUMBER: N/A.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-07-22T06:15:39.293557+00:00
- pubmed
- last seen: 2026-07-22T06:09:22.412566+00:00
License: public-domain-us
· commercial use OK
· attribution required
Courtesy of the U.S. National Library of Medicine
Courtesy of the U.S. National Library of Medicine