Large Language Models Meet Gynecologic Ultrasound: Advancing the Characterization of ADNEXal Masses.

OA: gold CC-BY-4.0
⚙ AI-generated summary by qwen3.7-flash, 2026-10-04 ⓘ

This study evaluated ChatGPT’s ability to classify adnexal masses as benign or malignant and predict histology, finding that expert subjective assessment outperformed the large language model in diagnostic accuracy.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

⚙ AI-generated deep summary by qwen3.7-flash, 2026-10-04 · read from full text ⓘ

This study evaluated the diagnostic accuracy of the large language model ChatGPT (GPT-5) in distinguishing benign from malignant adnexal masses compared to established IOTA tools and expert sonographers. Researchers analyzed 300 surgically confirmed cases, inputting standardized ultrasound descriptors and CA125 levels into the model to assess its ability to predict histological outcomes. The results demonstrated that while ChatGPT could generate classifications, it exhibited significant limitations in reliability and applicability compared to traditional objective models and human expertise. Relevance to endometriosis: This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Ovarian cancer (OC) is the second most common gynecological malignancy and remains one of the leading causes of gynecological cancer-related mortality worldwide. A major clinical challenge is the lack of an accurate and widely applicable strategy for identifying patients at high risk of malignancy at an early stage. In this context, artificial intelligence (AI) has emerged as a promising tool to improve diagnostic performance. Among AI technologies, large language models (LLMs) have recently shown considerable potential in healthcare applications. In this study, we evaluated the diagnostic performance of ChatGPT (GPT-5) in classifying 300 adnexal masses as benign or malignant and compared its performance with that of the IOTA Simple Rules, the ADNEX model, and expert subjective assessment. We also assessed ChatGPT's ability to predict the most likely histological diagnosis for each lesion. All adnexal masses were described using the International Ovarian Tumor Analysis (IOTA) terminology, and histopathological examination served as the reference standard. Our findings showed that expert subjective assessment achieved the highest overall diagnostic performance for both benign/malignant classification (accuracy 87.3%; 95% CI, 83.0-90.9%) and prediction of the presumed histological diagnosis. ChatGPT A and ChatGPT B reached a sensitivity of 72.3% and 73.5%, a specificity of 74.5% and 75.9%, a positive predictive value of 75.2% and 76.5%, and a negative predictive value of 71.5% and 72.8%, respectively (inconclusive responses counted as misclassifications), with an overall accuracy of 73.3% and 74.7%. After adequate validation, large language models might complement existing decision-support tools for less experienced examiners, without replacing expert evaluation. Their ease of use and reliance on standardized ultrasound descriptors make them accessible to ultrasonographers with varying levels of expertise.
Full text 29,544 characters · extracted from pmc-nxml · 5 sections · click to expand

Section 2

All patients were enrolled at the Gynecology Oncology Unit of the “Giovanni Paolo II” I.R.C.C.S. Cancer Institute in Bari and of the Hospital Polyclinic in Bari between August 2022 and August 2025. All the ovarian masses were analyzed by ultrasound examination according to the IOTA criteria. Clinical and histopathological data were collected from clinical charts. Ultrasound examinations were performed by one expert operator at each participating center, each having at least 15 years of experience in gynecologic ultrasound. Transabdominal examination was performed with a 3.5–5.0 MHz convex transducer, and transvaginal examination with a 9.0–5.0 MHz broadband transducer (Voluson Expert 22, GE HealthCare, Chicago, IL, USA). The main stages of the study are briefly summarized in Figure 1 . The study was conducted according to the guidelines of the Declaration of Helsinki and the protocol was approved by the Institutional Review Board of “Giovanni Paolo II” I.R.C.C.S. Cancer Institute, Bari, Italy. All patients provided informed consent to use the data for scientific purposes. We collected 339 ultrasound examinations of adnexal masses, which were classified according to their morphology ( Figure 2 ), echogenicity ( Figure 3 ), and vascularization ( Figure 4 ). In the case of bilateral adnexal masses, the mass with the more complex ultrasound morphology was considered in the analysis. If both masses had similar ultrasound morphology, the larger mass or the one more easily accessible by ultrasound was included. Inclusion criteria were: - Suspected adnexal disease for which the patient was referred to centers adherent to the study; - Execution of a pre-surgical pelvic ultrasound examination; - Pre-surgical check of serum CA125; - Patients selected for surgery. Suspected adnexal disease for which the patient was referred to centers adherent to the study; Execution of a pre-surgical pelvic ultrasound examination; Pre-surgical check of serum CA125; Patients selected for surgery. We excluded 39 patients because of missing CA125. All the remaining 300 masses were evaluated by an expert ultrasound examiner who described them using IOTA terminology and assigned the presumptive diagnosis of benignity and malignancy using the Simple Rules, the ADNEX model, and subjective assessment. The approach to assess the pelvis was, at first, the transvaginal ultrasound, while the transabdominal ultrasound was used just for the masses which cannot be adequately evaluated for dimension, localization, or for patients who were virgo. Our reference was the histological examination which identified 48.33% benign masses and 51.67% malignant masses. The OpenAI ChatGPT system (GPT-5) was used to classify adnexal masses as benign or malignant based on standardized ultrasound descriptors. All cases were submitted through the ChatGPT web interface via a plus subscription, using ChatGPT (GPT-5; model introduced in ChatGPT on 7 August 2025 OpenAI, San Francisco, CA, USA), between 2 September 2025 and 13 October 2025. Web browsing, memory, and all additional tools were disabled. No system instructions other than those reported in Supplementary Material were provided. Sampling parameters (e.g., temperature, top_p) could not be customized in the web interface and were kept at default settings. Each case was submitted once in a newly opened, independent conversation, so that no information could carry over between cases. The prompt required a predefined response format, and responses were read and coded by G.S. and F.A.; a response was coded as inconclusive when it did not explicitly categorize the lesion as benign or malignant. The complete prompts, including the task instructions, the exact wording and order of the input variables, and the required output format, are reported verbatim in Supplementary Material . Only fully de-identified, structured ultrasound descriptors and serum CA125 values were entered into the platform; no images, personal identifiers, dates, or free-text clinical notes were transmitted. The information submitted could not, alone or in combination, allow patient re-identification, in compliance with the General Data Protection Regulation (GDPR). The main analysis focused on different combinations of parameters ( Table 3 ; Figure 5 ). At first, the accuracy of the ChatGPT prompting configurations in the diagnosis of benignity/malignancy was tested against the Simple Rules, the ADNEX model, and subjective assessment. Questions were asked in English describing the adnexal formations according to IOTA terminology and always using the same sequence: morphology, diameter, color score (ChatGPT A) and, then, the additional characterizing element for configurations ChatGPT B, C, D, and E ( Table 3 ). The statement “is a unilocular solid ovarian cyst of 14 mm with a color score 2 and a vascularized papilla more likely benign or malignant?” is an example of the question scheme we followed. Afterwards, the accuracy of ChatGPT in the determination of the presumed histological diagnosis was tested through standardized questions. The following information was provided: morphology, size, color score, and serum level of presurgical CA125. According to the IOTA Consensus Statement, we chose to add the information of the cystic content for unilocular and unilocular solid morphologies. For multilocular and multilocular solid morphologies, we also added the cystic content and the number of loci. The statement “which is the histological presumed diagnosis for a unilocular solid ovarian cyst with the following features: diameter 199 mm, color score 2, anechoic content, CA125 of 31?” is an example of the question scheme we followed. The result of the histological examination of the lesion after its surgical removal was used as the reference standard in all phases of the study. Agreement between the different diagnostic methods was evaluated using Cochran’s Q test for comparing more than two percentages from non-independent samples. If the null hypothesis of equality was rejected, Sheskin’s test for comparing two paired proportions was applied [ 25 ]. The accuracy assessment was carried out with particular reference to malignant cases, defining true positives (TP), false negatives (FN), true negatives (TN) and false positives (FP) ( Table 4 ). For each diagnostic method, sensitivity, specificity, positive and negative predictive values, and overall accuracy were determined, along with their respective confidence intervals using the Clopper–Pearson method. Inconclusive ChatGPT responses and cases in which the Simple Rules were not applicable were not treated as test-negative results: in the primary analysis, they were counted as misclassifications, i.e., as false negatives when the lesion was malignant and as false positives when it was benign at histology (intention-to-diagnose approach), and a sensitivity analysis restricted to conclusive classifications was also performed. Differences in sensitivity, specificity, and predictive values between methods were described by comparing point estimates and their 95% confidence intervals.

Intro

Ovarian cancer (OC) is the fifth most common oncological disease and the fourth leading cause of cancer death in the United States [ 1 ], with a 5-year survival rate of approximately 45% [ 2 ]. Because early detection significantly reduces mortality, accurately determining whether a pelvic mass is benign or malignant is critical. Expert sonographers are frequently able to determine the nature of an adnexal mass during the pre-surgical workup with an accuracy of 92%, compared with 82–87% for less experienced examiners [ 3 ]. In order to make this assessment objective, reproducible, and universally shared, the definition of a universal language has become necessary. To standardize this assessment universally, the IOTA (International Ovarian Tumor Analysis) group [ 4 ] introduced objective ultrasound models, such as the Simple Rules [ 5 ] and the ADNEX model (Assessment of Different NEoplasias in the adneXa) [ 6 ] ( Table 1 and Table 2 ). However, the subjective evaluation made by an experienced sonographer has consistently been confirmed as an accurate method to discriminate between benign and malignant lesions [ 3 ] with better performance compared with the other ultrasound methods [ 7 ]. Since expert sonographers are not always available, AI has emerged as a prominent source of healthcare innovation. Although gynecology relies heavily on imaging, AI implementation in this field is still in its early stages compared to other medical specialties [ 8 ]. Multiple studies have demonstrated its validity in the fields of screening [ 9 , 10 ], triage [ 11 , 12 ], diagnosis [ 13 , 14 ], drug development [ 15 , 16 ], treatment [ 17 , 18 ], monitoring [ 19 ], and image interpretation [ 20 , 21 ]. In 1999, Timmerman et al. showed that an artificial neural network can be trained to discriminate between a suspected benign or malignant nature of an adnexal mass using some simple ultrasound criteria and some clinical information (e.g., menopausal state and serum level of CA 125) [ 22 ]. Acharya et al. developed a machine learning model to automatically discriminate between the diagnosis of benignity and malignancy based on a dataset of 1000 benign and 1000 malignant ultrasound images. This model achieved an accuracy of 99.9%, a sensitivity of 100%, and a specificity of 99.8% [ 23 ]. A deep learning model based on a convolutional neural network (CNN) was proposed in the study by Aramendia-Vidaurreta et al. to obtain an automatic classification of adnexal masses, combining the ultrasound characteristics and the age of the patients. This model revealed an overall accuracy of 98.8% with a sensitivity of 98.5% and a specificity of 98.9% [ 24 ]. Currently, innovation is driven by large language models (LLMs). These deep learning systems are trained on massive datasets, allowing them to process realistic text and mimic human reasoning and communication. The most widespread variant today is ChatGPT (developed by OpenAI). Given the exponential growth of AI and the integration of ChatGPT into daily life, evaluating its reliability in healthcare has become an urgent priority. To date, however, no study has systematically evaluated a state-of-the-art LLM against the validated IOTA instruments (Simple Rules and ADNEX model) and expert subjective assessment on the same consecutive cohort of surgically confirmed adnexal masses described with standardized IOTA terminology. Prior work on artificial intelligence for adnexal masses has focused on image-based approaches, such as neural networks and radiomics [ 22 , 23 , 24 ], whereas the ability of LLMs to reason over standardized ultrasound descriptors remains essentially unexplored. This is the research gap addressed by the present study. The aim of our study is to evaluate what could be the prospects of the use of ChatGPT in ultrasound evaluation of adnexal masses, specifically exploring potential and limits in predicting the histological diagnosis. Therefore, this study aims to offer a point of reflection on the new application perspectives of AI such as ChatGPT in the healthcare field and on the possibility that its introduction could facilitate medical work or, conversely, hinder it.

Results

The study population consisted of 300 patients, with a median age of 52 years. At the time of the assessment, 162 patients were in menopause, 106 were of fertile age, and for the remaining 32 patients, these data were unavailable. Ascites was present in 41/300 cases. Table 5 summarizes the characteristics of the examined adnexal formations. The outcome of the histological examination reported a total of benign formations of 48.33% and a total of malignant formations of 51.67%. ChatGPT A and ChatGPT B were unable to determine whether the lesion was benign or malignant in 9% (27/300) and 4.33% (13/300) of cases, respectively; to reflect the clinical utility of the configuration, inconclusive responses and cases in which Simple Rules were not applicable were categorized as diagnostic failures (i.e., incorrect outputs), as the configuration failed to recognize the nature of the lesion. The configurations ChatGPT A and B which were assessed for the total study population agree with the histological examination ( Figure 6 ). The test to compare proportion of diagnosis among different methods showed a statistically significant difference for both benign ( Figure 6 , Q = 108.04, p < 0.001) and malignant ( Figure 6 , Q = 263.25, p < 0.001). The post hoc comparison between pairs was performed through the Sheskin method, which determined the minimum significant difference (MSD) between two paired proportions with a p -value lower than 0.05, if the Cochrane test was significant. For malignant, MSD was 10.1%, while for benign, the MSD was 9.5%. In a descriptive evaluation of the subpopulations (prompting configurations ChatGPT C, ChatGPT D, and ChatGPT E), performance was observed when additional parameters were considered beyond the predefined ones (morphology, size, and vascularization). Within its respective subpopulation, ChatGPT C aligned with the histological diagnosis for both malignant and benign lesions. Conversely, ChatGPT D aligned primarily with benign diagnoses, while ChatGPT E yielded a high proportion of inconclusive outputs ( Figure 7 for configurations C and D). Since configurations C, D, and E were evaluated in different eligible subpopulations, these comparisons are strictly descriptive and exploratory: a formal assessment of the incremental value of each additional variable would require comparing the baseline and expanded prompts within the same cases. Comparing the accuracy of the methods examined about the presumed diagnosis benignity vs. malignancy, the subjective assessment results were the most performing overall. In our study ( Table 6 ), the subjective assessment reported the highest accuracy (87.3%; 95% CI 83.0–90.9%), while the Simple Rules reported the lowest (66.7%; 95% CI 61.0–72.0%). Analyzing the agreement between the methods, it could be noticed that ADNEX had the highest percentage of malignant diagnosis (225/300 cases, 75%) ( Table 7 A), while the histological examination reported a percentage of 51.67% cases (155/300 cases). That result could be explained, in our examination, as referable to the nature of the method, which uses a cut-off of 10% to define an adnexal mass as a higher risk of malignancy, with subsequent surgical management. Therefore, in our study, we considered all the masses with an ADNEX score >10% as malignant, but, of course, not all of them were determined to be malignant at the histological diagnosis. This could be the reason why ADNEX overestimated malignant masses in our study population with considerable disagreement compared with other methods. Compared to ADNEX, ChatGPT A and B show a lower number of false positives but higher false negatives, resulting in lower sensitivity but greater specificity. On the other hand, subjective assessment seems to slightly overestimate the percentage of malignant masses while remaining homogeneous with the histological result. In this regard, the use of an additional method such as ChatGPT to support subjective diagnosis could be of clinical utility. After analyzing the agreement between methods, we calculated the positive predictive value (PPV), the negative predictive value (NPV), the sensitivity (SE), and the specificity (SP), with their 95% confidence intervals ( Table 7 B). Since the identification of tumor pathology is of greater clinical use, the focus of this analysis was on the malignant diagnosis. Subjective assessment showed a higher sensitivity than ChatGPT A (93.5% vs. 72.3%) and ChatGPT B (93.5% vs. 73.5%), with non-overlapping 95% confidence intervals ( Table 7 B). Configurations ChatGPT A and B showed similar sensitivity (72.3% vs. 73.5%) and specificity (74.5% vs. 75.9%), with largely overlapping confidence intervals. Regarding predictive values, the main difference between subjective assessment and ADNEX concerned the positive predictive value (83.8% vs. 67.6%, with non-overlapping 95% confidence intervals), which confirms our previous observation related to the ADNEX overestimation of malignancy. ChatGPT A and B were also similar in terms of positive predictive value (75.2% vs. 76.5%) and negative predictive value (71.5% vs. 72.8%). The similarity between sensitivity, specificity, and positive and negative predictive values between the two proposed configurations ChatGPT A and B could be due to the similarity between the two methods. In fact, ChatGPT B includes the same ultrasound parameters (morphology, size, and color score) plus the laboratory data (serum level of CA125). Finally, both predictive values of subjective assessment (PPV 83.8%, NPV 92.1%) were higher than those of ChatGPT A and B, the difference being most evident for the negative predictive value. The aim of the second phase of the study is to evaluate the feasibility of the use of LLM instead of an expert evaluation in the definition of the presumed histological diagnosis. The histological reports with the greatest statistical and clinical significance are cystadenoma, cystadenofibroma, teratoma, endometrioma, fibroid, borderline tumor, and invasive carcinoma. Compared with the histological report, the numbers of cases assigned by ChatGPT and by subjective assessment were similar for teratoma and endometrioma, whereas ChatGPT deviated more markedly than subjective assessment from the histological distribution for cystadenoma, cystadenofibroma, fibroid, and borderline tumor ( Table 8 and Figure 8 ). As can also be seen in Figure 8 , ChatGPT seems to overestimate some diagnoses (cystadenoma, borderline tumors, mature teratoma) and underestimate others (cystadenofibroma, invasive carcinomas or metastatic ones, fibroids). Looking at our results, it is possible to hypothesize that they are due to a limited ability of the system to interpret and contextualize some ultrasound features. For instance, teratomas and fibroids are both characterized by the ultrasound appearance of acoustic shadows. The discrepancy in the results obtained could be linked to the incorrect interpretation and contextualization of this ultrasound feature, resulting in incorrectly classifying many fibroids as teratomas. Likewise, cystadenoma and cystadenofibroma could be similar in their ultrasound appearance. Therefore, we could hypothesize that a large proportion of unidentified cystadenofibromas have been misclassified as cystadenomas. Furthermore, misdiagnosed invasive carcinomas may have been classified as borderline cancer. Lastly, ChatGPT does not express an opinion in 3.67% of cases.

Discussion

OC ranks as the seventh most diagnosed cancer among women worldwide and the second most common gynecological malignancy [ 26 ]. Its high lethality is primarily driven by a rapid peritoneal spread and the absence of effective screening programs capable of detecting the disease at an early stage [ 27 ]. Identifying accurate tools for early diagnosis and prognosis remains an unmet clinical need [ 28 ]. In this scenario, scientific research is increasingly exploring the applications of AI and LLMs, such as ChatGPT, within the gynecological field [ 29 , 30 , 31 , 32 , 33 , 34 ]. In 2021, the World Health Organization (WHO) published the “Ethics and Governance of Artificial Intelligence for Health” report, establishing a core principle: medical decisions should be made by humans, supported but never replaced by AI. Indeed, the misuse of AI entails significant ethical risks, profoundly impacting: Privacy and Cybersecurity: The vulnerability of sensitive patient data. Informed Consent and Transparency: The lack of clear protocols regarding how AI is utilized during the diagnostic–therapeutic process. Skill Degradation: The risk of a progressive decline in the clinical skills of healthcare professionals due to over-reliance. Patient-Guided Misuse: The danger of patients seeking care outside the health system without proper medical surveillance [ 35 ]. Privacy and Cybersecurity: The vulnerability of sensitive patient data. Informed Consent and Transparency: The lack of clear protocols regarding how AI is utilized during the diagnostic–therapeutic process. Skill Degradation: The risk of a progressive decline in the clinical skills of healthcare professionals due to over-reliance. Patient-Guided Misuse: The danger of patients seeking care outside the health system without proper medical surveillance [ 35 ]. LLMs have a unique ability to mimic human language. That may result in false, inaccurate, or incomplete statements with a potential negative consequence on patients’ health. Moreover, LLMs could be trained on poor quality or biased data [ 36 ]. Despite the growing literature on clinical LLMs, current reviews typically address wide-ranging clinical use cases at the expense of a dedicated focus on disease diagnosis. This has created a critical research gap, as the specific methodologies, strengths, and challenges of deploying LLMs for diagnostic tasks are rarely analyzed in depth. Existing specialty-specific reviews, focusing on areas such as digestive or infectious diseases [ 37 , 38 ], fail to provide a broader, cross-specialty analysis that incorporates diverse data types, LLM techniques, and clinical diagnostic workflows. Currently, there are still few studies concerning the application of LLM in the field of gynecology. A 2024 Italian study revealed that official national guidelines (AIOM) scored significantly higher in clarity, consistency, and completeness compared to answers provided by ChatGPT regarding OC [ 39 ]. On the other hand, the study by Cascella et al. highlighted that, despite LLMs’ impressive capabilities, they could obtain poor performance in real settings such as the medical field, where complex high-level thinking is needed [ 40 ]. Conversely, ChatGPT demonstrated high concordance with human evaluators when tested in triaging and recommending management interventions for fictitious obstetric and gynecologic emergencies [ 41 ]. The study by Ługowski et al. evaluated ChatGPT’s clinical knowledge by assessing its accuracy in answering questions from the Polish Specialty Certificate Examination (SCE) in Obstetrics and Gynecology [ 42 ]. They demonstrated that even though ChatGPT is capable of answering very difficult questions, its accuracy is better when solving less complex problems; therefore, the newer version should aid physicians in their daily practice. Moreover, it should be used in English to provide the best accuracy of answers. To mitigate the risks of misdiagnosis, it is essential to implement systematic processes, ensuring the accuracy of training datasets, and to establish clear procedures for handling situations where the AI’s output conflicts with clinical judgment. With this study, we aimed to expand the application of LLMs in the field of gynecology, placing a stronger emphasis on diagnostic models. Specifically, we evaluated the diagnostic potential of ChatGPT in the ultrasound assessment of 300 adnexal masses. Although the proposed AI configurations showed a lower overall accuracy than subjective expert assessment, the results are encouraging. ChatGPT offers clear advantages, including easy accessibility and educational utility due to its highly detailed explanations. However, the model has a critical structural limitation: it cannot interpret ultrasound images directly. It relies entirely on a human operator to provide correct data. In this study, the descriptions were inputted by an expert operator using standardized IOTA terminology. This demonstrates that the accuracy of the AI is strictly dependent on the expertise of the physician, confirming that AI should support, rather than replace, medical judgment. The lower diagnostic performance observed for ChatGPT is likely explained by its inability to discriminate and appropriately contextualize individual ultrasound features. Interestingly, incorporating additional input parameters, including serum CA-125 levels, vascularized papillary projections, acoustic shadows, and ascites, did not appreciably improve diagnostic accuracy. This finding is not unexpected, as none of these variables is pathognomonic for a specific diagnosis. For example, CA-125 may be elevated in benign conditions such as endometriosis; acoustic shadows may result from fibrous septa, calcifications, or solid tissue, which can be observed in both benign and malignant lesions; vascularized papillary projections are suggestive, but not exclusive, of borderline ovarian tumors; and ascites may occur in non-neoplastic conditions, whereas some patients with ovarian cancer present without ascites. Therefore, these variables require clinical interpretation rather than isolated pattern recognition. Our text-based approach should also be considered against state-of-the-art artificial intelligence systems that analyze ultrasound images directly. Deep learning models trained on pelvic ultrasound images have achieved expert-level discrimination of ovarian cancer in large multicenter series [ 43 ], and a recent systematic review confirmed the rapid growth and overall promising accuracy of image-based artificial intelligence in gynecologic oncology [ 44 ]. Radiomics-based machine learning has likewise shown good diagnostic performance for adnexal masses [ 45 ]. Compared with these systems, an LLM working on standardized IOTA descriptors currently offers lower discriminative performance, but it requires no imaging pipeline, is immediately available in clinical settings, and leverages the same standardized lexicon already used in routine reporting. The two approaches are therefore complementary, and multimodal systems combining direct image analysis with structured descriptors represent a natural next step. The application of LLMs in the clinical management of adnexal masses and ovarian cancer is supported by the recent literature across both pre-diagnostic screening and ultrasound stratification. LLM-based natural language processing has proven effective in extracting non-coded symptom signatures from unstructured electronic health records to support early cancer detection [ 46 ]. Additionally, LLMs have demonstrated high accuracy and reliability in standardizing O-RADS categorization directly from free-text ultrasound reports, highlighting their utility as clinical decision-support tools [ 47 ]. The principal limitation associated with the lower diagnostic accuracy of ChatGPT is the potential for inappropriate clinical management of adnexal masses. False positive results may lead to unnecessary surgery, whereas false negative classifications may delay referral and treatment of ovarian malignancies. Therefore, the probability of benignity or malignancy generated by ChatGPT should always be interpreted together with the patient’s clinical history, physical examination, laboratory findings, and comprehensive ultrasound assessment. Future studies should evaluate ChatGPT in real-world clinical settings, where more comprehensive clinical information is available than the limited set of ultrasound descriptors and additional variables included in the present study. Moreover, multimodal artificial intelligence models combining direct ultrasound image analysis with structured clinical data should be explored. Multicenter validation studies are also warranted to assess the generalizability of ChatGPT across different imaging platforms and healthcare settings. Finally, domain-specific fine-tuning of large language models using curated gynecological ultrasound reports may further improve diagnostic performance. Although the current evidence is encouraging, the integration of LLMs into routine clinical practice will require further improvements in model performance through training on larger, high-quality datasets, together with the development of robust ethical and regulatory frameworks addressing transparency, privacy, data security, and informed consent. Most importantly, LLMs should be regarded as decision-support tools that complement, rather than replace, clinical expertise. Importantly, we acknowledge that evaluating each case with a single prompt is a limitation of the present study as it does not capture output variability; accordingly, ongoing and future trials will incorporate repeated queries to formally assess response stability.

Conclusions

Expert subjective assessment remains the most accurate approach for the ultrasound characterization of adnexal masses. Although ChatGPT demonstrated lower diagnostic performance, it showed potential as an additional decision-support aid for less experienced examiners, to be used alongside validated IOTA instruments and never as a substitute for referral to an expert sonographer. Its performance depends on the use of standardized IOTA terminology and should always be interpreted in conjunction with clinical evaluation rather than as a stand-alone diagnostic method. Future research should focus on the paired evaluation of prompting strategies on identical case sets, the formal assessment of response stability across repeated queries, prospective multicenter validation, the comparison of different LLMs, including open-weight models deployable on premise for privacy reasons, and studies measuring the impact on clinician workload and diagnostic efficiency rather than diagnostic agreement alone.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-10-04T09:26:46.659050+00:00
License: CC-BY-4.0 · commercial use OK · attribution required
Per Europe PMC