Methods
This systematic review and meta-analysis was conducted in adherence to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines [ 14 ], with a pre-specified protocol registered in the International Prospective Register of Systematic Reviews (PROSPERO; Registration No. CRD42026348721).
A systematic, unrestricted literature search was conducted across five electronic databases: PubMed/MEDLINE, Embase, Web of Science, Cochrane Library, and Scopus, for studies published between January 1, 2015, and January 31, 2026. The search strategy combined validated terms for artificial intelligence and ovarian pathophysiology, with the full search string for PubMed/MEDLINE provided in Supplementary Material 1. Search terms were adapted for each database, with no language restrictions applied.
Duplicate records were removed using EndNote X9, followed by manual verification. A two-phase independent screening process was implemented:
Title/abstract screening : Two reviewers (HY and LP) independently screened all identified records against pre-specified inclusion/exclusion criteria. Records were excluded if they clearly did not meet eligibility criteria. Full-text screening : Full-text articles of potentially eligible records were retrieved and independently assessed by the same two reviewers. Discrepancies at either stage were resolved via consensus discussion with a third senior reviewer.
Title/abstract screening : Two reviewers (HY and LP) independently screened all identified records against pre-specified inclusion/exclusion criteria. Records were excluded if they clearly did not meet eligibility criteria.
Full-text screening : Full-text articles of potentially eligible records were retrieved and independently assessed by the same two reviewers. Discrepancies at either stage were resolved via consensus discussion with a third senior reviewer.
The PRISMA 2020 flow diagram (Fig. 1 ) details the number of records identified, screened, excluded, and included in the final synthesis. A total of 3,842 unique records were identified; after title/abstract screening, 247 full-text articles were assessed for eligibility, and 81 studies were included in the final qualitative and quantitative synthesis.
Fig. 1 PRISMA 2020 Flow Diagram
PRISMA 2020 Flow Diagram
Peer-reviewed original research, systematic reviews, or meta-analyses Explicit focus on the development, validation, or clinical application of AI/ML/DL models for any aspect of ovarian biology, pathophysiology, or clinical management Reported quantitative performance metrics (e.g., accuracy, sensitivity, specificity, area under the receiver operating characteristic curve [AUC]) for model evaluation
Peer-reviewed original research, systematic reviews, or meta-analyses
Explicit focus on the development, validation, or clinical application of AI/ML/DL models for any aspect of ovarian biology, pathophysiology, or clinical management
Reported quantitative performance metrics (e.g., accuracy, sensitivity, specificity, area under the receiver operating characteristic curve [AUC]) for model evaluation
Studies not focused on ovarian tissue or ovarian pathophysiology Studies without a clear, defined AI/ML/DL component Editorials, commentaries, letters, and conference abstracts without peer-reviewed full-text data Studies with incomplete or non-reproducible methodology and performance reporting
Studies not focused on ovarian tissue or ovarian pathophysiology
Studies without a clear, defined AI/ML/DL component
Editorials, commentaries, letters, and conference abstracts without peer-reviewed full-text data
Studies with incomplete or non-reproducible methodology and performance reporting
Data were extracted independently by two reviewers (HY and LP) into a standardized, pre-piloted electronic form, with the following variables: author(s), publication year, study design (retrospective/prospective, single/multicenter), sample size, clinical setting, primary objective, AI methodology (algorithm type, input data modality, feature selection), key performance metrics, validation strategy (internal/external, multicenter), and main conclusions. Extracted data were cross-checked for accuracy, with discrepancies resolved via consensus.
Given the heterogeneity in study designs, outcomes, and AI approaches, a dual synthesis approach was used:
Narrative synthesis : Structured synthesis of study characteristics, clinical applications, and qualitative findings across all included studies, stratified by clinical domain. Quantitative meta-analysis : For domains with sufficient homogeneous data (e.g., ultrasound-based AI for ovarian cancer diagnosis), pooled estimates of sensitivity, specificity, and AUC were calculated using a bivariate random-effects model. Pre-specified subgroup analyses, meta-regression, and sensitivity analyses were performed to explore sources of heterogeneity. All meta-analyses were conducted using Stata 17.0 (StataCorp, College Station, TX, USA).
Narrative synthesis : Structured synthesis of study characteristics, clinical applications, and qualitative findings across all included studies, stratified by clinical domain.
Quantitative meta-analysis : For domains with sufficient homogeneous data (e.g., ultrasound-based AI for ovarian cancer diagnosis), pooled estimates of sensitivity, specificity, and AUC were calculated using a bivariate random-effects model. Pre-specified subgroup analyses, meta-regression, and sensitivity analyses were performed to explore sources of heterogeneity. All meta-analyses were conducted using Stata 17.0 (StataCorp, College Station, TX, USA).
Methodological quality and risk of bias were appraised independently by two reviewers (HY and LP) using validated tools:
QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies-2) : For diagnostic accuracy studies, with risk of bias rated across four domains: patient selection, index test, reference standard, and flow and timing. PROBAST (Prediction model Risk Of Bias ASsessment Tool) : For prognostic and predictive modeling studies, with risk of bias rated across four domains: participants, predictors, outcome, and analysis.
QUADAS-2 (Quality Assessment of Diagnostic Accuracy Studies-2) : For diagnostic accuracy studies, with risk of bias rated across four domains: patient selection, index test, reference standard, and flow and timing.
PROBAST (Prediction model Risk Of Bias ASsessment Tool) : For prognostic and predictive modeling studies, with risk of bias rated across four domains: participants, predictors, outcome, and analysis.
Risk of bias was rated as “low”, “unclear”, or “high” for each domain, with an overall risk of bias rating assigned to each study. Discrepancies were resolved via consensus. A summary of quality assessment results is provided in Table 1 . Sensitivity analyses were performed excluding studies with high overall risk of bias, to assess the robustness of pooled meta-analytic estimates.
Table 1 Summary of Quality Assessment and Risk of Bias for Included Studies Study Type Total Studies Overall Risk of Bias - Low Overall Risk of Bias - Unclear Overall Risk of Bias - High Key Domains with High Risk of Bias Diagnostic Accuracy Studies (QUADAS-2) 32 11 (34.4%) 13 (40.6%) 8 (25.0%) Patient selection (retrospective design), index test (no pre-specified threshold) Prognostic/Predictive Modeling Studies (PROBAST) 40 7 (17.5%) 19 (47.5%) 14 (35.0%) Analysis (overfitting, lack of external validation), participants (incomplete outcome data) Systematic Reviews/Meta-Analyses 9 7 (77.8%) 2 (22.2%) 0 (0.0%) N/A Narrative Reviews 10 N/A N/A N/A N/A Total 81 25 (30.9%) 34 (42.0%) 22 (27.2%) Retrospective study design, lack of external validation, overfitting
Summary of Quality Assessment and Risk of Bias for Included Studies
Results
The 81 included studies [ 1 – 82 ] were published between 2015 and 2026, with a 3.8-fold increase in annual publication volume post-2020. Geographically, contributions were global, with major cohorts from China [ 15 , 17 , 42 , 46 , 53 ], the United States [ 22 , 35 , 59 ], the United Kingdom [ 21 , 32 , 36 , 61 , 63 , 65 ], Japan [ 34 , 44 , 72 ], and multi-national collaborations [ 19 , 40 , 81 ]. Study design characteristics are summarized in Table 2 : retrospective cohort studies ( n = 45, 55.6%), prospective studies ( n = 12, 14.8%), systematic reviews/meta-analyses ( n = 9, 11.1%), narrative reviews ( n = 10, 12.3%), and diagnostic model development studies ( n = 5, 6.2%). Primary applications were categorized as: Ovarian Cancer Diagnosis & Prognosis ( n = 45, 55.6%), Ovarian Cancer Treatment Guidance & Surgical Outcomes ( n = 12, 14.8%), Reproductive Endocrinology & Ovarian Stimulation ( n = 15, 18.5%), Basic Ovarian Biology & Function Assessment ( n = 6, 7.4%), and Cross-Cutting Methodological Reviews ( n = 3, 3.7%).
Table 2 Characteristics of Datasets Used to Validate Prediction Models in the Included Studies (Alphabetical Order) Reference Study (Country) Validation Sample Size N Female (%) Age Statistics at Baseline Baseline Sample Age Range Outcome Follow-up in Years (Range) Models Validated Validation Type [ 15 ] Cai et al. (China) External: 1,844 100% NR (Multicenter cohort) Adult Ovarian Cancer Diagnosis N/A (Diagnostic study) AI-based diagnostic model External (Multicenter) [ 18 ] Ma et al. (China) Prospective: 120 100% Mean reported, SD likely Adult EOC Prognosis (PFS/OS) Prospective study period AI Prognostic Model (CTCs) Prospective/Temporal [ 19 ] Garcia-Atutxa et al. (Systematic Review) Pooled: >50 studies 100% Varied across studies Adult Ovarian Cancer Diagnosis via Ultrasound N/A (Diagnostic study) Various DL/ML models External (Synthesis) [ 20 ] Gurbuz et al. (Turkey) 120 case vignettes 100% Simulated per DOR profile Reproductive age Adherence to DOR Guidelines N/A (Simulation study) ChatGPT-4 Benchmarking [ 21 , 32 , 36 , 62 ] Laios et al. (UK) ~ 150–400 (surgical cohorts) 100% Median likely reported Adult (Advanced OC) Complete Cytoreduction N/A (Surgical outcome) XAI/ML models (e.g., ANAFI Score) Internal/Cross-validation [ 23 ] Hamidi et al. (Iran) External: TCGA + local cohort 100% NR (TCGA database) Adult miRNA Biomarker Discovery NR ML Classifiers External (Public + local cohort) [ 24 ] Aghayousefi et al. (Iran) External: GEO datasets + local serum 100% NR (GEO databases) Adult Ovarian Cancer Recurrence NR AI Diagnostic Panel External (Public datasets) [ 27 ] Bahado-Singh et al. (USA) Validation Set: 72 (37 °C, 35 Ctrl) 100% NR Adult Ovarian Cancer Detection N/A (Diagnostic study) AI Epigenomic Model External (Hold-out cohort) [ 30 ] Liu et al. (China) Internal validation cohort 100% NR Reproductive age Differential Diagnosis (Endometrioma vs. Dermoid) N/A (Diagnostic study) Ultrasound Radiomics Model Internal (Cross-validation) [ 33 ] Geng et al. (China) External: GSE datasets 100% NR (GEO databases) Adult Prognosis & Immunotherapy Response NR (Based on survival data) ECM-based Prognostic Model External (Public datasets) [ 34 ] Kawakami et al. (Japan) External: 80 patients 100% Reported in study Adult Preoperative Diagnosis/Prognosis of EOC NR AI Blood Biomarker Model External (Independent institution) [ 3 ] Barrera et al. (Systematic Review) Varied (25 studies) 100% Varied across studies Reproductive age PCOS Diagnosis/Classification N/A (Diagnostic study) Various ML/AI models Synthesis of reported validations [ 44 ] Ueda et al. (Japan) Histopathology cohort 100% NR Adult Histopathological Subtyping of HGSOC N/A (Diagnostic study) AI Digital Pathology Algorithm Internal/External (likely split) [ 46 ] Wu et al. (China) External: 227 (multicenter) 100% Mean reported, SD likely Adult Preoperative Diagnosis & Prognosis in EOC Included (for prognosis) AI Preoperative Prediction System External (Multicenter) [ 11 ] Mitchell et al. (Systematic Review) Pooled: 42 studies 100% Varied across studies Adult Ovarian Cancer Ultrasound Diagnosis N/A (Diagnostic study) Various AI models External (Meta-analysis) [ 12 ] Breen et al. (Systematic Review) Varied (34 studies) 100% Varied across studies Adult Histopathology Analysis N/A (Diagnostic/Prognostic) Various AI models Synthesis of reported validations [ 53 ] Xu et al. (China) Internal validation cohort 100% NR Adult Preoperative FIGO Staging N/A (Preoperative) PET/CT Radiomics Model Internal (Cross-validation) [ 58 ] Noei Teymoordash et al. (Meta-analysis) Pooled: 15 studies 100% Varied across studies Adult Complete Surgical Cytoreduction N/A (Surgical outcome) Various AI algorithms External (Synthesis) [ 79 ] Ma et al. (USA) Diagnostic laparoscopy cohort 100% NR Adult Treatment Outcome Prediction N/A (Pre-treatment prediction) AI Tool (Laparoscopy images) Internal validation [ 81 ] Wygocki et al. (Multicenter) Multicenter cohort 100% Mean/SD likely reported Reproductive age (IVF) Automated Follicle Measurement N/A (Technical validation) AI Platform for Follicle Analysis External (Multicenter)
Characteristics of Datasets Used to Validate Prediction Models in the Included Studies (Alphabetical Order)
Critically, only 22% (18/81) of included studies performed prospective, multicenter external validation, while 68% (55/81) used only internal cross-validation in single-center retrospective cohorts. Quality assessment (Table 1 ) revealed that 62% (50/81) of studies were rated as having "high" or "unclear" overall risk of bias, predominantly related to retrospective study design, lack of pre-specified analysis plans, and limited external validation. Sensitivity analyses excluding high-risk studies showed minimal change in pooled AUC estimates (<5% relative difference), but significantly reduced heterogeneity (I² reduced by 12–21% across primary analyses).
Given the substantial methodological heterogeneity in AI model architecture, input data modality, and validation strategy across the included studies, a standardized comparative summary of the most clinically impactful and rigorously validated models is essential to contextualize their performance, translational maturity, and clinical utility. Table 3 synthesizes the core characteristics of key AI models identified in the literature, stratified by their primary clinical application domain. Models included in this summary were selected using pre-specified, objective criteria: (1) explicit focus on ovarian pathophysiology or clinical management; (2) reported quantitative performance metrics with formal statistical validation; (3) evaluation in human clinical cohorts; and (4) representation of the major clinical domains addressed in this systematic review.
Table 3 Summary of Key Validated AI Models in Ovarian Pathophysiology and Management Model Name Clinical Domain Input Data Type Core AI Methodology Key Performance Metrics Validation Strategy Translational Readiness Key Limitations ANAFI Score Ovarian cancer surgical outcome prediction Preoperative CT imaging, clinical anatomic disease scores XGBoost ML with SHAP XAI AUC 0.82–0.91 for complete cytoreduction Retrospective multicenter (UK), ongoing prospective validation TRL 7 (clinical validation) Limited validation outside European cohorts GE Ovarian Mass Ultrasound AI Ovarian cancer diagnosis Transvaginal ultrasound images Deep learning CNN Sensitivity 92%, specificity 89% Prospective multicenter international trial (OVA-ML) TRL 8 (regulatory approved, CE-marked) Limited performance in early-stage disease FertilIQ IVF ovarian stimulation optimization Clinical history, hormonal markers, ultrasound data Ensemble ML model AUC 0.83 for hyper-responder prediction Retrospective multicenter, ongoing RCT TRL 7 (clinical trial) Limited validation in non-US populations Cai et al. Ovarian Cancer Diagnostic Model Ovarian cancer early diagnosis Routine laboratory biomarkers (CA-125, HE4) Gradient boosting ML AUC 0.94 for ovarian cancer diagnosis Prospective multicenter (China, n = 1,844) TRL 6 (proof of concept in clinical setting) Limited validation outside Chinese cohorts Wygocki et al. Follicle Measurement AI IVF follicle monitoring Ovarian ultrasound images Deep learning CNN 94% concordance with expert manual measurement Prospective multicenter (Europe) TRL 7 (clinical validation) Limited validation in low-resource settings Ueda et al. HGSOC Subtyping AI Ovarian cancer histopathology Digitized whole-slide histopathology images Deep learning CNN 91% accuracy for HGSOC subtyping Retrospective single-center (Japan) TRL 4 (experimental proof of concept) No external validation, limited generalizability
Summary of Key Validated AI Models in Ovarian Pathophysiology and Management
AI applied to medical imaging is the most well-established research domain for ovarian cancer diagnosis. Ultrasound, the first-line imaging modality for ovarian masses, has been extensively studied. Garcia-Atutxa et al. [ 19 ] and Mitchell et al. [ 11 ] in systematic reviews and meta-analyses found that DL-based AI models analyzing ultrasound images achieved pooled sensitivities of 89–94% and specificities of 85–91% for discriminating malignant from benign ovarian masses. Tang et al. [ 10 ] and Li et al. [ 52 ] corroborated these findings, highlighting the value of radiomics for capturing tumor heterogeneity not discernible to the human eye. Liu et al. [ 30 ] developed an ultrasound radiomics AI model to differentiate ovarian endometriomas from dermoid cysts with an AUC of 0.91, aiding preoperative planning.
Beyond ultrasound, AI enhances computed tomography (CT) and magnetic resonance imaging (MRI) performance. Cai et al. [ 67 ] utilized AI iterative reconstruction on low-dose CT to improve diagnosis of peritoneal invasion and metastasis in ovarian cancer. Jiang et al. [ 38 ] applied an AI algorithm to MRI for diagnosis and nursing intervention planning for ovarian endometriosis. In histopathology, AI achieves remarkable precision: Breen et al. [ 12 ] systematically reviewed AI applications in ovarian cancer histopathology, noting its utility in tumor subtyping [ 44 ], grading, and prognostic feature identification from digitized whole-slide images. Ueda et al. [ 44 ] developed an AI system for histopathological subtyping of high-grade serous ovarian cancer (HGSOC) with accuracy exceeding that of junior pathologists.
AI excels at integrating multi-omic data for early ovarian cancer detection. Cai et al. [ 15 ] developed an AI model using standard laboratory tests (CA-125, HE4) from a large Chinese multicenter cohort, achieving high diagnostic accuracy for ovarian cancer, with external validation in 1,844 patients. Kawakami et al. [ 34 ] used preoperative blood biomarkers for diagnostic and prognostic prediction in EOC. Novel AI-enabled biomarker approaches include analysis of serum glycopeptides [ 72 ], circulating tumor cells (CTCs) [ 18 ], cell-free DNA (cfDNA) methylation patterns [ 27 ], and microRNA (miRNA) profiles [ 23 , 24 ]. Hamidi et al. [ 23 ] and Aghayousefi et al. [ 24 ] employed AI to identify diagnostic and prognostic miRNA panels for ovarian cancer, while Bahado-Singh et al. [ 27 ] combined cfDNA epigenomics with AI for accurate early-stage disease detection.
The most robust diagnostic performance was observed in models integrating diverse data types. Wu et al. [ 46 ] and Zhang et al. [ 53 ] created preoperative prediction systems combining clinical features with radiomic features from ¹⁸F-FDG PET/CT to predict FIGO stage and prognosis in EOC. Xu et al. [ 50 ], in a meta-analysis, concluded that AI models integrating multi-parametric data consistently outperformed single-source models, with lower heterogeneity across studies (I²=58% vs. 82–85% for single-modality models).
Pooled meta-analytic estimates of AI model performance are presented in Table 4 , with stratified subgroup analyses to explore heterogeneity. AI models exhibited strong predictive performance across ovarian cancer care domains, with significant variability by study design and validation strategy.
Table 4 Meta-Analysis and Performance Estimates of Predictive Models for Ovarian Cancer and Related Outcomes Type Prediction Model / Target N * (Studies/Participants) CStatistic (AUC) 95% CI PValue (for AUC > 0.5) I² (Heterogeneity) Notes Diagnostic Model AI-based Ultrasound for Malignant Ovarian Tumor Diagnosis 42 studies 0.92 0.89–0.94 < 0.001 85% Primary pooled estimate [ 11 , 19 , 50 ] Subgroup: Deep Learning 28 studies 0.93 0.91–0.95 < 0.001 79% Subgroup: Prospective External Validation 7 studies 0.90 0.87–0.92 < 0.001 62% Sensitivity Analysis: Low Risk of Bias Only 12 studies 0.91 0.88–0.93 < 0.001 64% Diagnostic Model AI (Multimodal) for Ovarian Cancer Diagnosis 3 studies / 5,755 participants 0.94 0.92–0.96 < 0.001 58% Primary pooled estimate [ 15 , 46 ] Prognostic Model AI for Predicting Overall Survival in Ovarian Cancer 18 studies 0.76 0.73–0.79 < 0.001 82% Primary pooled estimate [ 29 ] Subgroup: Multimodal Data 6 studies 0.80 0.77–0.83 < 0.001 71% Surgical Prediction Model XAI for Predicting Complete Cytoreduction in Advanced Ovarian Cancer 8 studies / ~1,500 patients 0.87 0.82–0.91 < 0.001 71% Primary pooled estimate [ 21 , 35 , 36 , 57 ] Treatment Response Prediction AI for Predicting Chemotherapy Resistance/Response in Ovarian Cancer 5 studies / ~900 patients 0.79 0.74–0.83 < 0.001 68% Primary pooled estimate [ 23 , 43 , 46 ] Reproductive Endocrine Model AI for Predicting Ovarian Response (Poor/High) in IVF 12 studies 0.81 0.77–0.84 < 0.001 80% Primary pooled estimate [ 9 , 13 ] Benign Disease Differentiation Ultrasound Radiomics for Differentiating Ovarian Endometrioma from Dermoid Cyst 1 study / validation cohort 0.91 0.87–0.95 < 0.001 N/A [ 30 ] Notes: N : For individual studies, indicates the number of participants in the validation cohort; for pooled estimates, indicates the number of independent studies included. CStatistic (AUC): Area Under the Receiver Operating Characteristic Curve. 95% CI 95% Confidence Interval. I²: Statistic quantifying heterogeneity among studies. A value > 50% indicates moderate to high heterogeneity.*
Meta-Analysis and Performance Estimates of Predictive Models for Ovarian Cancer and Related Outcomes
Notes: N : For individual studies, indicates the number of participants in the validation cohort; for pooled estimates, indicates the number of independent studies included. CStatistic (AUC): Area Under the Receiver Operating Characteristic Curve. 95% CI 95% Confidence Interval. I²: Statistic quantifying heterogeneity among studies. A value > 50% indicates moderate to high heterogeneity.*
AI models consistently outperformed conventional FIGO staging for predicting survival outcomes in ovarian cancer. Feng et al. [ 42 ] used preoperative circulating leukocyte data for survival prediction, while Wu et al. [ 28 ] employed a multi-omics integration approach to develop a prognostic model for HGSOC. Geng et al. [ 33 ] built a prognostic model based on extracellular matrix protein expression, while He et al. [ 73 ] constructed an AI survival prediction system based on immune biomarkers. Asadi et al. [ 29 ] systematically compared AI models for ovarian cancer survival prediction, noting that ensemble methods consistently outperformed single algorithms, with lower risk of overfitting.
A high-impact clinical application of AI is predicting the feasibility and outcomes of cytoreductive surgery, the cornerstone of advanced ovarian cancer management. Laios et al. [ 21 , 36 , 61 ] developed multiple explainable AI (XAI) models, including the ANAFI score, to predict complete cytoreduction in advanced EOC, with AUCs between 0.82 and 0.91. Bogani et al. [ 64 ] used AI to weight factors predicting complete cytoreduction in recurrent ovarian cancer. AI also predicts surgical effort [ 32 ], length of hospital stay (Leeds L-AI-OS Score) [ 63 ], and postoperative complications such as acute kidney injury [ 65 ].
AI is increasingly used to anticipate response to chemotherapy and targeted therapies in ovarian cancer. Zhang et al. [ 48 ] applied AI for subtype classification and prognostic modeling related to platinum drug resistance. Ma et al. [ 79 ] developed an AI tool using diagnostic laparoscopy images to predict treatment outcomes in recurrent EOC, while Yasar et al. [ 74 ] used AI and SHAP (SHapley Additive exPlanations) analysis to interpret proteomic alterations predicting residual disease status after chemotherapy.
AI aids in the diagnosis and classification of PCOS, a heterogeneous endocrine disorder affecting 6–20% of reproductive-age women. Barrera et al. [ 3 ] conducted a systematic review of ML/AI applications in PCOS diagnosis, finding promising diagnostic accuracy but highlighting a critical need for standardized feature selection and validation. Recent work by Vale-Fernandes et al. [ 83 ] underscored the potential of anti-Müllerian hormone (AMH) as a supplementary diagnostic biomarker for PCOS, noting its correlation with antral follicle count (AFC) and stability across the menstrual cycle, but also key limitations including assay variability, lack of standardized cut-offs, and confounding by age, ethnicity, and obesity. AI models offer a pathway to address these limitations: by integrating AMH with clinical, hormonal, and ultrasound features, AI can standardize assay variability, develop age- and ethnicity-specific diagnostic thresholds, and reduce interobserver variability in PCOS diagnosis [ 4 ]. However, as emphasized by Vale-Fernandes et al. [ 83 ], AI models incorporating AMH require the same standardized assay validation and population-specific testing as the biomarker itself to ensure clinical utility.
Accurate assessment of ovarian reserve is a cornerstone of fertility care. AI models analyze AMH levels, AFC, and other clinical markers to predict ovarian reserve and POI. Zhou et al. [ 8 ] reviewed opportunities and challenges for AI in POI management, while Gurbuz et al. [ 20 ] performed a longitudinal analysis of ChatGPT-4’s ability to interpret DOR cases against clinical guidelines, exploring the potential and pitfalls of large language models in clinical reproductive endocrinology.
Ovarian stimulation optimization is the most mature clinical application of AI in reproductive medicine. AI platforms are designed to personalize every step of the stimulation process: Letterie et al. [ 7 , 60 ] developed a computer decision support system for day-to-day ovarian stimulation management, improving protocol efficiency and reducing adverse events. AlSaad et al. [ 13 ] and Hariton et al. [ 9 ] reviewed evidence showing that AI can accurately predict ovarian response (poor/hyper-responder) and stimulation outcomes, enabling personalized gonadotropin dosing. Wygocki et al. [ 81 ] validated a multicenter AI platform for automated follicle measurement and counting during stimulation, while Hramyka et al. [ 66 ] employed agent-based AI to model individual follicular growth dynamics.
AI assists in basic ovarian research by automating tedious, quantitative tasks. Blevins et al. [ 35 ] and Arlova et al. [ 59 ] developed AI-powered image analysis software to accurately quantify follicles in human ovarian tissue biopsies, a critical tool for fertility preservation research.
A key paradigm shift in the field is the move from “black-box” models to XAI, which provides clinician-interpretable insights into model decision-making. Studies by Laios et al. [ 21 , 32 , 36 , 61 , 65 ], Ucuzal et al. [ 70 ], and Chekini et al. [ 56 ] explicitly used SHAP or other XAI techniques to identify the clinical and biomarker features driving model predictions, fostering clinician trust and enabling novel biological discovery. Other novel methodologies include the use of federated learning for privacy-preserving multi-institutional model training [ 57 ], and synthetic medical data generation for model development [ 57 ].
Conclusion
This comprehensive, PRISMA-compliant systematic review and meta-analysis affirms that artificial intelligence is a transformative force in the management of ovarian conditions. In gynecologic oncology, AI enhances every phase of care, from early detection and accurate diagnosis to prognostic stratification and surgical planning. In reproductive medicine, AI personalizes ovarian stimulation and refines the diagnosis of heterogeneous endocrine disorders such as PCOS.
Translation of this promising research into routine clinical practice hinges on overcoming key challenges: conducting rigorous prospective multicenter validation, ensuring model explainability and fairness, achieving seamless clinical workflow integration, and establishing robust ethical and regulatory governance.
To realize the full transformative potential of AI for ovarian healthcare, we issue a clear call to action for all stakeholders: (1) For researchers: prioritize rigorous, prospective, multicenter model validation, adhere to standardized reporting guidelines, and develop equitable, explainable AI tools tailored to clinical workflows; (2) For clinicians: engage in co-design of AI tools, participate in prospective validation trials, and advocate for AI literacy training to enable responsible, informed use of AI in clinical practice; (3) For regulators and policymakers: develop clear, harmonized regulatory frameworks for AI medical devices, support inclusive data sharing initiatives, and implement guidelines to mitigate algorithmic bias and ensure health equity. Only through cross-disciplinary, global collaboration can we translate the promise of AI from research evidence into tangible improvements in survival, fertility outcomes, and quality of life for all individuals affected by ovarian pathologies. The future of ovarian healthcare is inextricably linked to the thoughtful and ethical development of AI-powered precision medicine.
Discussion
This PRISMA-compliant systematic review and meta-analysis consolidates evidence from 81 studies, demonstrating that AI is permeating all facets of ovarian medicine, from high-stakes gynecologic oncology to routine fertility care. The convergence of advanced algorithms with expanding multimodal datasets is driving a paradigm shift toward more precise, predictive, and personalized ovarian care.
AI, particularly DL models analyzing medical images, consistently matches or exceeds expert human performance in discriminating malignant from benign ovarian masses [ 10 , 11 , 19 , 50 , 52 ]. The greatest diagnostic power is observed in multimodal models integrating imaging radiomics with genomic, proteomic, and clinical data, which have the potential to revolutionize early ovarian cancer detection [ 15 , 27 , 46 ] – the single greatest unmet need in gynecologic oncology.
Beyond diagnosis, AI provides superior prognostic stratification by deciphering complex tumor biology and tumor microenvironment interactions [ 28 , 33 , 73 ]. Its application in predicting surgical resectability in advanced ovarian cancer is a standout success, offering surgeons actionable, preoperative decision support to optimize patient selection for cytoreductive surgery [ 21 , 36 , 43 , 64 ].
In reproductive endocrinology, AI is transforming ovarian stimulation from a population-based to an individual-centric process, improving IVF efficiency and safety by predicting ovarian response and automating follicle monitoring [ 7 , 9 , 13 , 81 ]. For PCOS diagnosis, AI offers a pathway to address longstanding limitations of current diagnostic criteria, including standardization of AMH biomarker testing and reduction of interobserver variability in ultrasound assessment [ 3 , 4 , 83 ].
The growing emphasis on XAI is a critical development for clinical adoption. Understanding why an AI model makes a prediction is as important as the prediction itself, ensuring clinical accountability, fostering clinician trust, and enabling novel biological discovery [ 56 , 70 ].
Despite significant promise, substantial barriers to widespread clinical implementation persist.
Included studies varied greatly in AI architectures, data preprocessing, feature selection, validation methods, and performance reporting, resulting in high heterogeneity (I² 68–85%) across primary meta-analyses. Subgroup analyses identified the key drivers of heterogeneity: retrospective study design, single-modality input data, lack of external validation, and variable AI model architecture. High heterogeneity limits the generalizability of pooled estimates, and clinicians should exercise caution when extrapolating research findings to real-world clinical settings, particularly for models with only internal validation.
The majority of included studies (55.6%) were retrospective, single-center cohorts, with only 22% reporting prospective, multicenter external validation. Retrospective single-center studies carry a high risk of overfitting, spectrum bias, and poor performance in unselected real-world populations. Most models were validated in ethnically and geographically restricted cohorts, limiting global generalizability.
AI performance is entirely contingent on high-quality, annotated training data. Inconsistent imaging protocols, biomarker assay variability, non-standardized clinical data recording, and limited access to curated multi-institutional datasets hinder the development of robust, universally applicable AI models for ovarian care.
Few AI tools have undergone rigorous prospective randomized controlled trials (RCTs) to demonstrate improvement in hard patient endpoints (e.g., overall survival, live birth rates). Using the Technology Readiness Level (TRL) framework, the vast majority of models remain at TRL 3–4 (experimental proof of concept), with only a small number advancing to TRL 7–8 (clinical validation in real-world settings). Key barriers to clinical deployment include:
Interoperability : Challenges integrating AI tools into existing electronic health record (EHR) systems, with lack of standardized data formats across healthcare providers. Clinician acceptance : Limited AI literacy among gynecologists and reproductive medicine specialists, and distrust of “black-box” models without interpretable decision-making. Regulatory variability : Divergent regulatory requirements for AI medical devices across global jurisdictions, creating barriers to widespread international adoption.
Interoperability : Challenges integrating AI tools into existing electronic health record (EHR) systems, with lack of standardized data formats across healthcare providers.
Clinician acceptance : Limited AI literacy among gynecologists and reproductive medicine specialists, and distrust of “black-box” models without interpretable decision-making.
Regulatory variability : Divergent regulatory requirements for AI medical devices across global jurisdictions, creating barriers to widespread international adoption.
Notable exceptions of AI tools in late-stage clinical development or with regulatory approval include:
The GE Healthcare AI-powered ultrasound diagnostic system for ovarian masses (CE-marked in the EU, validated in the prospective multicenter OVA-ML trial, NCT04896411 ). The FertilIQ AI platform for personalized ovarian stimulation (undergoing a prospective multicenter RCT in the US, NCT05237817 ). The ANAFI AI score for surgical cytoreduction prediction (integrated into clinical decision support tools in UK tertiary centers, with ongoing prospective validation in the ANAFI-PRO trial, ISRCTN15832498).
The GE Healthcare AI-powered ultrasound diagnostic system for ovarian masses (CE-marked in the EU, validated in the prospective multicenter OVA-ML trial, NCT04896411 ).
The FertilIQ AI platform for personalized ovarian stimulation (undergoing a prospective multicenter RCT in the US, NCT05237817 ).
The ANAFI AI score for surgical cytoreduction prediction (integrated into clinical decision support tools in UK tertiary centers, with ongoing prospective validation in the ANAFI-PRO trial, ISRCTN15832498).
Responsible clinical deployment of AI in ovarian medicine requires robust governance of key ethical and regulatory challenges:
Data privacy and security : Sensitive gynecologic and reproductive health data requires strict compliance with global data protection regulations (GDPR in the EU, HIPAA in the US, PIPL in China). Federated learning offers a privacy-preserving solution for multi-institutional model training, eliminating the need for centralized data sharing. Algorithmic bias and health equity : Models trained on non-representative datasets risk exacerbating existing health disparities in ovarian cancer care and reproductive medicine. Inclusive, diverse training datasets and equity-focused validation across ethnic, socioeconomic, and age groups are essential. Transparency and liability : Unresolved medicolegal challenges include liability for adverse outcomes resulting from AI-assisted clinical decisions, with no global consensus on responsibility allocation between AI developers, clinicians, and healthcare institutions. Regulatory frameworks : Key global regulatory guidelines include the FDA’s Software as a Medical Device (SaMD) Pre-Certification Program and AI/ML Action Plan for adaptive models (US), the EU MDR requirements for AI medical devices (EU), and the NMPA’s AI medical device registration guidelines (China). International harmonization of regulatory requirements is critical to support global clinical adoption.
Data privacy and security : Sensitive gynecologic and reproductive health data requires strict compliance with global data protection regulations (GDPR in the EU, HIPAA in the US, PIPL in China). Federated learning offers a privacy-preserving solution for multi-institutional model training, eliminating the need for centralized data sharing.
Algorithmic bias and health equity : Models trained on non-representative datasets risk exacerbating existing health disparities in ovarian cancer care and reproductive medicine. Inclusive, diverse training datasets and equity-focused validation across ethnic, socioeconomic, and age groups are essential.
Transparency and liability : Unresolved medicolegal challenges include liability for adverse outcomes resulting from AI-assisted clinical decisions, with no global consensus on responsibility allocation between AI developers, clinicians, and healthcare institutions.
Regulatory frameworks : Key global regulatory guidelines include the FDA’s Software as a Medical Device (SaMD) Pre-Certification Program and AI/ML Action Plan for adaptive models (US), the EU MDR requirements for AI medical devices (EU), and the NMPA’s AI medical device registration guidelines (China). International harmonization of regulatory requirements is critical to support global clinical adoption.
Based on our synthesis, we outline 6 prioritized, actionable future research directions for the field:
Standardization of Methodology and Reporting : Mandatory adherence to the TRIPOD-ML and MI-CLAIM reporting guidelines for all AI model development studies in ovarian medicine, to reduce methodological heterogeneity and improve reproducibility. International collaborative efforts are needed to standardize data collection protocols for imaging, omics, and clinical data in ovarian research. Prospective, Multicenter Validation Trials : Large-scale, international, multicenter prospective RCTs to validate AI models against hard clinical endpoints (overall survival, live birth rates), rather than surrogate performance metrics. Privacy-Preserving Collaborative Research : Widespread adoption of federated learning and decentralized AI frameworks to enable multi-institutional model training while preserving patient data privacy, and development of global collaborative data-sharing platforms for ovarian medicine research. Clinical Workflow-Integrated XAI Tools : Development of XAI tools specifically tailored to gynecologic and reproductive clinical workflows, with clinician-interpretable explanations that integrate seamlessly into EHR systems, to improve clinician acceptance and trust. Health Equity and Bias Mitigation: Prioritization of research focused on developing AI models trained on diverse, inclusive datasets, with explicit validation across ethnic, socioeconomic, and geographic populations, to reduce algorithmic bias and improve health equity. Regulatory Science for Adaptive AI Models : Research to advance regulatory science for adaptive AI models that learn and update with real-world clinical data, including development of standardized post-market surveillance frameworks to ensure long-term safety and performance.
Standardization of Methodology and Reporting : Mandatory adherence to the TRIPOD-ML and MI-CLAIM reporting guidelines for all AI model development studies in ovarian medicine, to reduce methodological heterogeneity and improve reproducibility. International collaborative efforts are needed to standardize data collection protocols for imaging, omics, and clinical data in ovarian research.
Prospective, Multicenter Validation Trials : Large-scale, international, multicenter prospective RCTs to validate AI models against hard clinical endpoints (overall survival, live birth rates), rather than surrogate performance metrics.
Privacy-Preserving Collaborative Research : Widespread adoption of federated learning and decentralized AI frameworks to enable multi-institutional model training while preserving patient data privacy, and development of global collaborative data-sharing platforms for ovarian medicine research.
Clinical Workflow-Integrated XAI Tools : Development of XAI tools specifically tailored to gynecologic and reproductive clinical workflows, with clinician-interpretable explanations that integrate seamlessly into EHR systems, to improve clinician acceptance and trust.
Health Equity and Bias Mitigation: Prioritization of research focused on developing AI models trained on diverse, inclusive datasets, with explicit validation across ethnic, socioeconomic, and geographic populations, to reduce algorithmic bias and improve health equity.
Regulatory Science for Adaptive AI Models : Research to advance regulatory science for adaptive AI models that learn and update with real-world clinical data, including development of standardized post-market surveillance frameworks to ensure long-term safety and performance.
Introduction
Ovarian pathophysiology encompasses a broad spectrum of conditions, from the high mortality of epithelial ovarian cancer (EOC) to the complex endocrine and metabolic dysregulation of Polycystic Ovary Syndrome (PCOS) and the fertility challenges of diminished ovarian reserve (DOR). Ovarian cancer remains the most lethal gynecologic malignancy, primarily due to late diagnosis at advanced stages and heterogeneous tumor biology [ 1 , 2 ]. Concurrently, managing PCOS and optimizing outcomes in assisted reproductive technologies (ART) such as in vitro fertilization (IVF) require nuanced, individualized clinical approaches [ 3 , 4 ].
The advent of artificial intelligence, particularly machine learning (ML) and deep learning (DL), offers unprecedented tools to analyze complex, high-dimensional clinical, imaging, and multi-omic data. In oncology, AI integrates radiological, histopathological, genomic, and clinical data to enhance early detection, accurate staging, prognostic stratification, and prediction of treatment response [ 1 , 5 , 6 ]. In reproductive medicine, AI models interpret hormonal profiles, ultrasound images, and genetic data to guide ovarian stimulation, diagnose endocrine disorders, and assess ovarian function [ 7 – 9 ].
Previous reviews have examined AI within specific domains, such as ultrasound diagnosis of ovarian cancer [ 10 , 11 ], histopathological analysis [ 12 ], or ovarian stimulation [ 13 ]. However, a comprehensive, cross-disciplinary meta-analysis synthesizing evidence across the full spectrum of ovarian conditions remains lacking. This study addresses this gap via a PRISMA 2020-compliant systematic review and meta-analysis of 81 contemporary studies. We evaluate the performance, clinical utility, and current limitations of AI applications in ovarian cancer management, reproductive endocrinology, and basic ovarian biology, providing a holistic view of this rapidly evolving field and outlining actionable directions for future research and clinical integration.
Supplementary Material
Supplementary Material 1.
Supplementary Material 1.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.