Artificial intelligence for women's reproductive health: A scoping review of global diagnostic trends, methodological gaps, and a translational research agenda for low-resource settings

other public-domain-us
⚙ AI-generated deep summary by qwen3.7-flash, 2026-09-28 · read from full text ⓘ

This scoping review analyzed 116 studies on artificial intelligence applications for women’s reproductive and hormonal health published between 2021 and 2025. The authors found that while ensemble methods and neural networks demonstrated strong diagnostic performance, the majority of research originated from high-income countries and suffered from high risk of bias and a lack of external validation. Consequently, the current evidence base is insufficient to support population-level deployment in low-resource settings, highlighting significant geographic and methodological gaps. Relevance to endometriosis: listed as one condition where AI shows potential diagnostic performance, though the paper's main focus is the global distribution and methodological rigor of AI tools across multiple reproductive disorders.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Women's reproductive and endocrine disorders including Polycystic Ovary Syndrome (PCOS), endometriosis, thyroid disorders, infertility, and pregnancy-related complications remain a major global health burden. These conditions are especially difficult to manage in low- and middle-income countries (LMICs), where diagnostic facilities and specialist care are limited. We conducted a PRISMA-ScR-guided scoping review of artificial intelligence (AI) and machine learning (ML) applications in women's reproductive and hormonal health. Our search covered five databases PubMed, Scopus, IEEE Xplore, Web of Science, and Google Scholar and included peer-reviewed studies published between 2021 and 2025. A total of 116 eligible studies were identified and classified into six categories: clinical prediction and risk modeling, medical imaging and AI, biomarker discovery and multi-omics, reproductive and endocrine health applications, clinical decision support systems, and AI/ML methodology development. Ensemble methods such as Random Forest and Gradient Boosting showed strong and consistent diagnostic performance. Convolutional neural networks performed well in single-centre settings but showed reduced performance upon external validation, consistent with optimism bias. Overall, 83.6% of studies were classified as high risk of bias and 75.5% lacked external validation. After adjusting for population size, high-income countries produced 23 times more studies per million women of reproductive age than LMICs. Sub-Saharan Africa contributed fewer than 0.1 studies per million women, and no studies were identified from Nepal or Sri Lanka. Current evidence, which is heavily concentrated in high-income and East Asian settings, is insufficient to support population-level deployment. We introduce a four-tier LMIC feasibility classification system that links data modality, infrastructure requirements, and personnel needs. We also propose a phased research agenda that separates near-term priorities (0-2 years) from medium-term goals that depend on infrastructure investment.
Full text 86,495 characters · extracted from oa-doi-fallback · 2 sections · click to expand

Abstract

Women’s reproductive and endocrine disorders including Polycystic Ovary Syndrome (PCOS), endometriosis, thyroid disorders, infertility, and pregnancy-related complications remain a major global health burden. These conditions are especially difficult to manage in low- and middle-income countries (LMICs), where diagnostic facilities and specialist care are limited. We conducted a PRISMA-ScR–guided scoping review of artificial intelligence (AI) and machine learning (ML) applications in women’s reproductive and hormonal health. Our search covered five databases PubMed, Scopus, IEEE Xplore, Web of Science, and Google Scholar and included peer-reviewed studies published between 2021 and 2025. A total of 116 eligible studies were identified and classified into six categories: clinical prediction and risk modeling, medical imaging and AI, biomarker discovery and multi-omics, reproductive and endocrine health applications, clinical decision support systems, and AI/ML methodology development. Ensemble methods such as Random Forest and Gradient Boosting showed strong and consistent diagnostic performance. Convolutional neural networks performed well in single-centre settings but showed reduced performance upon external validation, consistent with optimism bias. Overall, 83.6% of studies were classified as high risk of bias and 75.5% lacked external validation. After adjusting for population size, high-income countries produced 23 times more studies per million women of reproductive age than LMICs. Sub-Saharan Africa contributed fewer than 0.1 studies per million women, and no studies were identified from Nepal or Sri Lanka. Current evidence, which is heavily concentrated in high-income and East Asian settings, is insufficient to support population-level deployment. We introduce a four-tier LMIC feasibility classification system that links data modality, infrastructure requirements, and personnel needs. We also propose a phased research agenda that separates near-term priorities (0–2 years) from medium-term goals that depend on infrastructure investment. Citation: Munna MMH, Islam T, Faruk O, Tulona RF, Sultana A, Saima F, et al. (2026) Artificial intelligence for women’s reproductive health: A scoping review of global diagnostic trends, methodological gaps, and a translational research agenda for low-resource settings. PLoS One 21(9): e0358557. https://doi.org/10.1371/journal.pone.0358557 Editor: Kwang-Sig Lee, Korea University - Seoul Campus: Korea University, KOREA, REPUBLIC OF Received: May 13, 2026; Accepted: September 2, 2026; Published: September 25, 2026 Copyright: © 2026 Munna et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Data Availability: Yes - all data are fully available without restriction; All data underlying this review are fully available in the Mendeley Data repository (https://doi.org/10.17632/ydx7v73t3k.2). The repository contains: (1) the full list of 116 included studies with citations; (2) the complete extracted study-level variables dataset; (3) coding and category definitions for the six taxonomy categories and LMIC feasibility tiers; (4) the complete, reproducible search strategies for all five databases; (5) the screening results with reasons for exclusion at each stage; and (6) the raw data used to generate all tables, figures, counts, and proportions. The completed PRISMA-ScR checklist (S1 File), risk-of-bias scoring rubric (S2 Table), and LMIC applicability categorical criteria (S2 Table) are also provided as supplementary files. Funding: The author(s) received no specific funding for this work. Competing interests: The authors have declared that no competing interests exist. 1. Introduction Women’s reproductive and hormonal disorders constitute a significant yet persistently underaddressed global health challenge, with profound implications for fertility, metabolic function, mental health, and quality of life [1–3]. Reproductive and hormonal disorders are complex conditions resulting from a mix of genetic, environmental and lifestyle factors [1,4,5]. Although these are very common, they are highly under-diagnosed and under-treated, resulting in considerable health burden, sub-fertility, disability-adjusted life years (DALYs), and negative health outcomes in later life, such as cardiovascular and metabolic diseases and psychological disturbances [2,5]. PCOS affects 5–18% of women of reproductive age, and is closely associated with a number of other important comorbidities, such as insulin resistance, obesity, dyslipidaemia and depression [1,4,6]. Endometriosis affects 5–10% of women, with diagnostic delays often exceeding 8–12 years [7,8]. Women’s health is a systemic domain where women’s symptoms are often normalised or dismissed [7–9], and hence inadequate care is a common issue affecting these disorders. A huge burden of reproductive and hormonal disorders also exists in low- and middle-income countries (LMICs) in South Asia, Sub-Saharan Africa, and South East Asia, where there is a coexistence of socio-economic burden, absence of specialist healthcare services, limited diagnostic capacity, and societal stigma [10]. In Bangladesh, the Maternal Mortality Ratio (MMR) has been reported to decrease from 323 deaths/100,000 live births in 2001–116 deaths/100,000 live births in 2017, but non-acute reproductive and hormonal disorders are not optimally managed. The prevalence of PCOS in Bangladesh has been estimated to be 7.8–20% as per the Rotterdam criteria, and diagnostic delays of over two years have been reported in rural Bangladesh [10,11]. Rapid expansion in mobile phone ownership and internet connectivity across LMICs provides a viable foundation for AI-enabled health solutions, yet pilot studies including work on early gestational diabetes prediction using wearable devices and explainable AI in South Africa highlight the gap between technological readiness and clinical integration [10,11]. Artificial intelligence (AI) and machine learning (ML) have shown strong potential in healthcare. They are particularly useful for pattern recognition, risk prediction, and clinical decision support. These methods have been applied to many types of data, including clinical records, hormonal profiles, medical imaging, and lifestyle indicators. Studies have reported high diagnostic performance for conditions such as PCOS, endometriosis, and gestational diabetes [12–16]. However, most existing models have important limitations. They were developed in high-income countries using small or single-centre datasets [12,14]. Many lack external validation or interpretability assessment. This raises serious concerns about whether these models can work reliably in low- and middle-income country (LMIC) settings [10]. No existing review has systematically examined methodological trends and translational readiness from an LMIC perspective. Consequently, the findings of this review reflect trends predominantly from East Asian and high-income settings and should not be interpreted as globally generalizable trends. Translational recommendations for LMICs remain theoretical until validated in local, resource-constrained populations. Some recent reviews have addressed related topics. Mengistu et al. [15] conducted a broad scoping review of AI in sexual and reproductive health. Ghaderzadeh et al. [12] systematically reviewed AI applications specifically in PCOS. However, neither review examined methodological rigor, resource-stratified deployment feasibility, and geographic equity together. These three factors are essential for assessing translational readiness in LMIC settings. The present review addresses this gap through three key contributions. First, it proposes a six-category methodological taxonomy covering all major reproductive and endocrine conditions. Second, it introduces a four-tier LMIC feasibility classification system that links data modality, infrastructure, and personnel requirements. Third, it presents a population-adjusted geographic equity analysis. This analysis reveals significant research voids in Nepal, Sri Lanka, and most of Sub-Saharan Africa [10,11]. This paper presents a PRISMA-ScR–guided scoping review of AI and ML applications in women’s reproductive and hormonal health (Fig 1). It covers studies published between 2021 and 2025. The review focuses explicitly on translational applicability in low-resource settings. Its goal is to support the development of robust, equitable, and context-aware AI solutions for LMICs [10,12]. 2. Materials and methods This scoping review followed the PRISMA-ScR guidelines (Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews) [17]. PRISMA-ScR includes 20 essential items and 2 optional items. The completed checklist is provided as S1 File. These guidelines were chosen to make the review transparent, reproducible, and complete. This was especially important because the review aimed to explore and map existing literature. A predefined protocol guided the entire review process. It covered all major stages: literature search, screening, eligibility assessment, data extraction, study classification, and synthesis. Since this is a scoping review rather than an intervention-based systematic review, formal registration in PROSPERO was not required. However, the protocol was developed before data extraction began and was deposited in the Mendeley Data repository [18]. The repository includes screening metadata for all 273 identified records. It contains the reasons for exclusion at both the first and second screening stages. It also includes final inclusion decisions for the 116 eligible studies and their taxonomy classifications. This was done in accordance with the PRISMA 2020 data sharing guidelines. 2.1 Research questions The research questions were formulated using the PICO (Population, Intervention, Comparison, Outcome) framework [19]. The population comprised women affected by reproductive and hormonal disorders; the intervention included AI and ML–based diagnostic or predictive approaches; comparisons involved conventional or non–AI-based strategies where applicable; and outcomes focused on diagnostic accuracy, risk prediction, clinical decision support, and translational feasibility in low-resource healthcare settings. The review was guided by the following research questions: - Q1: What are the clinical applications of AI and ML in women’s reproductive and hormonal health? - Q2: Which AI/ML tasks and models are most frequently used across different reproductive and endocrine disorders? - Q3: What data modalities are employed to develop AI-based diagnostic and predictive systems? - Q4: How are existing studies distributed across disease-specific and methodological taxonomies, and what translational gaps exist for low-resource healthcare settings? 2.2 Search strategy A comprehensive literature search was conducted across five major academic databases: PubMed, Scopus, IEEE Xplore, Web of Science, and Google Scholar. All searches were performed between December 2025 and February 2026, and results were date-stamped to ensure reproducibility. The search strategy was designed to identify peer-reviewed studies applying artificial intelligence, machine learning, or deep learning techniques to women’s reproductive and hormonal health. Search terms were organized into two conceptual domains and combined using Boolean operators (AND/OR). The complete database-specific search queries are summarized in Table 1. The 2021 start date was selected because it marks the inflection point where transformer-based architectures and the post-COVID acceleration of digital health research substantially reshaped the AI-in-healthcare literature, rendering pre-2021 methodological surveys insufficient for characterizing the current state of the field. Earlier work is well covered by previous reviews [12,15,16]. To ensure full reproducibility, we explicitly restricted the search to English-language publications. Grey literature, including preprints, unpublished theses, and conference abstracts without full peer-reviewed papers, were explicitly excluded to ensure the inclusion of only methodologically complete, validated studies. The complete, database-specific search strings, including all Boolean operators and field tags, are provided in Table 1. 2.3 Eligibility criteria Studies were considered eligible if they satisfied all of the following criteria: - Original peer-reviewed research articles; - Application of AI, ML, or deep learning techniques; - Focus on women’s reproductive or hormonal disorders; - Use of human clinical, imaging, biochemical, lifestyle, or multimodal data; - Publication in English between January 1, 2021 and December 31, 2025. Studies were excluded if they met any of the following conditions: review articles, meta-analyses, editorials, commentaries, or book chapters; conference abstracts without full papers; animal or laboratory-based studies; studies unrelated to reproductive or endocrine health; absence of an AI/ML methodological component; or inaccessible full text or insufficient methodological detail. 2.4 Study selection process All retrieved records were exported into Microsoft Excel, and duplicate entries were removed using automated deduplication followed by manual verification. Study selection was conducted using a predefined two-stage screening process managed through Rayyan AI [20,21] screening software. In the first stage, titles and abstracts were independently screened by two reviewers (MMHM and TI) to assess relevance to the research questions. Studies deemed irrelevant by both reviewers were excluded, while discrepancies were resolved through discussion with a third reviewer (TS). In the second stage, full-text articles of potentially eligible studies were independently assessed against the predefined eligibility criteria by two reviewers (AFS and FS). Reasons for exclusion at this stage were documented using a standardized classification system. Initial database searches identified 402 records. After removing 129 duplicates, 273 records underwent title/abstract screening. Of these, 167 proceeded to full-text assessment after excluding 73 records at title screening and 23 at abstract screening. Full-text review of 167 articles resulted in a further 51 exclusions, yielding 116 studies that satisfied all inclusion criteria for final synthesis (Fig 1). Of the 116 final studies, 94 (81.0%) were retrieved from PubMed, Scopus, IEEE Xplore, or Web of Science, and 22 (19.0%) were unique to Google Scholar. The Google Scholar 200-result limit therefore affected approximately one-fifth of the final corpus, and sensitivity to this cap was assessed by cross-checking citation overlap across databases. Illustrating the study identification, screening, eligibility assessment, and inclusion process. 2.5 Data extraction and charting Data extraction was performed using a standardized Excel-based data-charting form, which was iteratively refined during a pilot phase involving 10 randomly selected studies. For each included study, data were extracted across six specific domains: - Bibliographic information (database source, title, authors, publication year, journal/conference); - Study characteristics (research objective, target disease/condition, study design, sample size, geographic region, funding source); - Data modalities (clinical records, hormonal/biochemical markers, medical imaging, lifestyle/demographic data, multimodal integration); - AI/ML components (tasks, algorithms employed, software/implementation details); - Performance metrics (accuracy, sensitivity, specificity, AUC, F1-score, validation methods); and - Translational indicators (data accessibility, computational requirements, cost considerations, ethical considerations, deployment feasibility in LMICs). A complete data extraction dictionary detailing the operational definitions for all extracted variables is provided in the supplementary repository. To ensure high fidelity and minimize extraction errors, the data underwent a rigorous, multi-tier independent verification process. Initial data extraction was conducted by one reviewer (TI). In the second stage, the extracted data were independently verified by three reviewers (OF, AFS, and FS). Subsequently, a third verification was conducted by MMHM to ensure comprehensive accuracy across the entire dataset. A 20% random sample of the fully verified data was cross-verified, with inter-rater agreement calculated using Cohen’s (). Any remaining discrepancies at any stage of the verification process were resolved through discussion and consensus with the senior authors (RUI and TS). 2.6 Study classification and taxonomy The classification taxonomy was developed through an iterative three-step process: (1) initial framework based on prior AI-in-healthcare reviews [12,15], (2) pilot testing and refinement using 20 randomly selected studies, and (3) finalization through team consensus meetings involving all authors. Each study was assigned to a single primary taxonomy category based on its principal research objective, data modalities, and AI/ML tasks. In cases where studies addressed multiple domains, classification was based on the primary contribution of the work. The final taxonomy comprised six categories: (1) Clinical Prediction and Risk Modeling; (2) Medical Imaging and Artificial Intelligence; (3) Biomarker Discovery and Multi-Omics Analysis; (4) Reproductive and Endocrine Health Applications; (5) Clinical Decision Support and Digital Health Tools; and (6) AI/ML Methodology and Algorithmic Development. To resolve boundary cases between Category 2 and Category 4, the following decision rule was applied prospectively: Category 2 was assigned when the primary methodological contribution was an imaging architecture or segmentation method; Category 4 was assigned when the primary contribution was clinical integration of AI with reproductive health data, regardless of whether imaging was involved. Studies combining ultrasound features with clinical parameters without proposing a novel imaging methodology were classified as Category 4. 2.7 Data synthesis and reporting To ensure transparency in all descriptive summaries, the total number of included studies (N = 116) serves as the primary denominator for all reported counts, proportions, and percentages throughout the results, unless a specific subgroup denominator (e.g., studies specific to a single disease domain or algorithm) is explicitly stated in the corresponding table or text. Given the heterogeneity in datasets, disease domains, and evaluation metrics across included studies, quantitative meta-analysis was not performed. Three structural violations precluded valid pooling: first, substantial clinical heterogeneity across disease domains (PCOS, endometriosis, thyroid, GDM) means that a pooled AUC estimate would constitute an ecological fallacy; second, 31 of 116 studies report multiple algorithms on identical datasets, violating the independence assumption required for meta-analytic pooling; third, overlapping training cohorts from shared public repositories (Kaggle PCOS, UCI thyroid) further compound statistical dependence. Formal heterogeneity testing (I²) was therefore not performed, as the independence violations render such statistics uninterpretable. This decision is consistent with established scoping review methodology, which prioritises breadth of evidence mapping over quantitative synthesis. All performance summaries are presented as descriptive ranges to characterise the reviewed literature, in alignment with scoping review methodology. Although descriptive AUC ranges are reported to characterise central tendency within algorithm categories, several structural threats to validity must be acknowledged: (1) ecological fallacy, as study-level AUC estimates aggregate outcomes across different diseases, populations, and imaging modalities; (2) statistical dependence, as 31 of 116 studies report multiple algorithms on identical datasets, violating independence assumptions; and (3) overlapping training cohorts from shared public repositories (e.g., UCI thyroid, Kaggle PCOS datasets). Given these violations, no inferential statistical testing (p-values, t-tests, or significance claims) is reported, and all quantitative summaries in Tables 4 and 5 are presented as descriptive approximations; no causal or superiority claim should be inferred. Extracted data were summarised descriptively and thematically in relation to the four research questions: (1) clinical applications by disease domain, (2) AI/ML task and model distribution, (3) data modality trends, and (4) taxonomy-based patterns with translational implications. Descriptive AUC ranges were compiled from author-reported metrics for each algorithm category. 2.8 Risk-of-bias and quality assessment The included corpus spans four methodologically distinct study types — diagnostic imaging, clinical prediction modeling, biomarker discovery, and algorithm development — each governed by a different validated risk-of-bias instrument (QUADAS-AI for imaging, PROBAST-AI for prediction, and TRIPOD-AI reporting standards for algorithm development). Item-by-item application of any single instrument across this heterogeneous corpus would have produced systematic misclassification, because items relevant to one study type are inapplicable to another. To preserve cross-corpus comparability while retaining domain-specific validity, we adopted a harmonized five-domain framework synthesizing the shared core items of PROBAST-AI, QUADAS-AI, and TRIPOD-AI. Each study was evaluated across five domains: (1) participant selection and data representativeness; (2) predictor measurement; (3) outcome assessment and ground truth reliability; (4) model development and validation strategy; and (5) reporting transparency. Studies were categorized as having low, moderate, or high risk of bias based on cumulative assessment across domains. The complete scoring rubric, including explicit criteria for low- and high-risk classification in each domain, is provided in S1 Table. Studies meeting high-risk criteria in 1–2 domains were classified as moderate risk; studies meeting high-risk criteria in 3 or more domains were classified as high risk. The number of studies at each risk level per domain is reported in Table 10. 2.9 LMIC applicability assessment Each included study was evaluated for deployment feasibility in LMIC settings using a four-tier qualitative classification scheme, defined prospectively based on infrastructure requirements, personnel needs, cost per test, and appropriate care level. The four categories are: (1) High Feasibility — requires only basic clinical variables, no specialist personnel, cost < $10/test, deployable in primary care; (2) Moderate Feasibility — requires routine laboratory tests or basic ultrasound, nurse- or midwife-level personnel, cost $10–25/test, deployable in secondary care; (3) Low Feasibility — requires specialist imaging or advanced laboratory panels, specialist personnel, cost $25–50/test, requires district hospital; and (4) Very Low Feasibility — requires MRI or advanced imaging or multi-omics, highly specialized personnel, cost > $50/test, requires tertiary centre only. The full categorical criteria are provided in S2 Table. This qualitative scheme was adopted in preference to a continuous numeric score to avoid conveying unwarranted precision and to ensure reproducibility of classification across reviewers. 2.10 Limitations of the methodology - Language restriction: Inclusion of only English-language publications may exclude relevant studies from non-English speaking LMICs. - Database coverage: While five major databases were searched, some relevant studies in regional or specialized databases may have been missed. - Grey literature exclusion: Preprints, theses, and conference proceedings without full papers were excluded, potentially missing recent innovations. - Rapid field evolution: The AI/ML field evolves quickly, and some advances from late 2025 may not be fully captured. - Scoping nature: As a scoping review, detailed quality assessment of individual studies was not the primary focus, though general methodological trends were assessed. 2.11 Sensitivity analysis for overlapping public datasets Several included studies draw on identical publicly available repositories, including the Kaggle PCOS dataset and the UCI thyroid dataset. To assess sensitivity of descriptive AUC ranges to dataset duplication, a collapsed dataset-level analysis was performed. Studies sharing a common training source were treated as a single observation for descriptive pooling purposes. After collapsing, effective independent dataset counts were: - PCOS (Kaggle-derived): 20 studies collapsed to 11 independent datasets; - Thyroid (UCI-derived): 24 studies collapsed to 14 independent datasets. Collapsed descriptive AUC ranges shifted by less than 0.02 units across all algorithm families, suggesting descriptive patterns are not materially driven by dataset duplication. However, algorithm frequency rankings (Table 3) should be interpreted with caution, as Random Forest’s apparent dominance in PCOS studies partly reflects repeated application to the same Kaggle benchmark rather than independent clinical validation. 2.12 Use of artificial intelligence tools During the preparation of this manuscript, the authors used Grammarly and Quillbot for language editing and manuscript refinement. The authors reviewed and edited the AI-generated content and take full responsibility for the accuracy, validity, and originality of the final text. 3. Results 3.1 Overview of included studies Data from 116 studies were synthesised. An increase in the number of publications was observed from 2021 to 2022, and a larger surge was seen from 2023 onwards. Studies were reported from 27 countries across all regions (Table 2). Although the number of studies in LMICs was marginally higher in absolute terms (51.7%), when adjusted for the number of women of reproductive age per region, the number of studies per million women in HICs was 23 times that in LMICs, reflecting disparities in institutional and funding capacity [13,22–136]. 3.2 Findings by taxonomy category The distribution of included studies across the six taxonomy categories is reported in Table 3. Clinical Prediction & Risk Modeling (41 studies) included work on PCOS, gestational diabetes mellitus (GDM), and preeclampsia, with Random Forest and XGBoost as dominant algorithms. Medical Imaging + AI (22 studies) applied convolutional neural networks (CNNs) to endometriosis, uterine fibroid, and ovarian abnormality diagnosis, with accuracies in the range of 85–99%; almost all studies used single-centre datasets. Biomarker Discovery (13 studies) applied multi-omics approaches to identify potential novel biomarkers for endometriosis and PCOS, using models such as PLS-DA and LASSO regression. Reproductive & Endocrine Health (19 studies) predominantly addressed PCOS diagnosis, often combining ultrasound features with clinical parameters (accuracy range 89–100%). Clinical Decision Support (10 studies) reported mobile health and digital triage applications, though very few included validation in LMIC populations. AI/ML Methodology (11 studies) described novel hybrid CNN-ML approaches and interpretability approaches such as SHAP [137] (Fig 2). Illustrating major categories including menstrual dysfunction, pelvic structural disorders, androgen and metabolic abnormalities, endocrine gland dysfunction, fertility and pregnancy-related complications, and ovarian aging. 3.3 Disease-specific analysis As summarized in Table 4, endometriosis (n = 27, 23.3%) and PCOS (n = 20, 17.2%) represented the most extensively researched domains. PCOS studies reported accuracy in the range of 88–99% and AUC in the range of 0.89–0.98, supported by the integration of clinical (52%) and ultrasound (31%) data. Thyroid disorder studies (n = 24, 20.7%) similarly reported strong discrimination (AUC range 0.89–0.99), reflecting standardized biomarker thresholds in hormonal data. Endometriosis studies reported the widest range of performance metrics (AUC range 0.72–0.94) and the lowest proportion of externally validated models (11.1%). Infertility/IVF studies (n = 4, 3.4%) reported the narrowest range of performance (AUC range 0.74–0.89) and none had external validation. When performance is stratified by validation rigor (Table 5), a directional pattern of AUC reduction is observed as validation moves from single-centre to external testing. This pattern is directionally consistent with optimism bias across all algorithm categories and should not be interpreted as a formal statistical comparison, given the small matched subsets, heterogeneity in disease contexts, and violations of independence. Fig 3 illustrates the directional reduction in the reported AUC ranges for five algorithm families moving from single-centre to external validation settings. Calibration reporting was assessed as an additional aspect of study quality: calibration metrics (Brier score, calibration curve, or Hosmer–Lemeshow test) were reported in 8 of 116 studies (6.9%). Decision curve analysis or net benefit were reported in 3 studies (2.6%), all focused on GDM or preeclampsia. Clinical decision thresholds were clearly mentioned in the text of only 1 study (0.9%). A model may have high AUC but be poorly calibrated, resulting in a systematic bias toward over- or under-estimation of individual risk. Calibration metrics, in addition to discrimination metrics, should be included in minimum reporting standards. Bar chart shows mean AUC values reported across single-center, multi-center, and external validation settings for five algorithm families. Differences across validation tiers are directionally consistent with optimism bias; formal significance testing was not performed given violations of independence and small matched subsets (minimum n = 5 per algorithm). Values should not be interpreted as comparative efficacy estimates. 3.3.1 Sensitivity analysis of performance patterns. To assess robustness of the descriptive AUC decay patterns, three supplementary exploratory checks were performed. - Leave-one-out analysis: the largest single study contributing to each algorithm category was sequentially excluded; for Random Forest (largest study n = 12,000), exclusion changed the descriptive single-centre AUC range by less than 0.01, suggesting results are not driven by outlier studies. - Disease-stratified sub-analysis: decay patterns were re-examined separately for PCOS, endometriosis, and thyroid disorders. Thyroid models showed the smallest decay (mean across validation tiers), consistent with their well-standardised biomarker reference ranges. Endometriosis models showed the largest decay (), consistent with heterogeneous imaging protocols and the absence of consensus reference standards. - Modality stratification: imaging-based studies (ultrasound + MRI) showed larger decay than clinical-data-only models ( vs. 0.04), consistent with known scanner-level and protocol-level domain shift in medical imaging. These exploratory sensitivity checks reinforce the narrative interpretation that directional patterns are robust to outlier and stratification perturbations, and support prioritising clinical-data models for near-term LMIC deployment. These should not be interpreted as confirmatory analyses. 3.4 Data modality and resource requirements Basic clinical variables were used in 98 (84.5%) of the studies, and routine laboratory tests in 65 (56.0%), indicating reliance on low-cost and readily available data. Tier 1 modalities were used in all the studies concerning PCOS, GDM, and Preeclampsia, supporting feasibility in primary and secondary care settings. Refer to Table 6. Use of hormonal panels (36.2%) and ultrasound (32.8%) suggests a moderate level of resource dependency, implying a need for some level of laboratory and clinical skill. The low frequency of use of MRI/Advanced Imaging (19.0%) and multi-omics platforms (15.5%) suggests these modalities are primarily being used in tertiary care in HICs and in specialized centres, and are therefore not widely applicable in resource-constrained settings. As shown in Table 6, models requiring Tier 3 resources have the lowest LMIC feasibility. As shown in Table 7, studies that do not require specialist personnel (e.g., community hypertension screening, symptom-based endometriosis classifiers) fall into High Feasibility, while ultrasound-based fibroid and PCOS diagnostics fall into Low Feasibility. The resource-applicability patterns illustrated in Fig 4 highlight that low-resource models tend to fall within the primary-care-compatible quadrants, further emphasising the importance of simple infrastructure for feasibility and scalability. The scatter plot visualizes the trade-off between infrastructure requirements and specialist dependency, with the LMIC feasibility threshold indicated by the dashed line. 3.5 Geographic and equity analysis The geographic distribution showed notable discrepancies. As shown in Table 8, East Asia contributed the largest proportion of studies (39.6%), followed by North America (17.2%) and Europe (16.4%). More than one-third of the studies in each of these regions were multicentre (37–45%), with external validation rates ranging from 24% to 35%, and open-access rates reaching 60% in North America. The lower validation rates observed in South Asia (15.5%) and Sub-Saharan Africa (1.7%) were accompanied by smaller median sample sizes (309 and 104, respectively) and limited data-sharing practices (only 11.1% of South Asian studies shared code or data). The country-level analysis (Table 9) further highlights under representation of South Asian countries: there were only 4 studies from Bangladesh (3.4%), all from a single tertiary care hospital and using only internal cross-validation, and none from Nepal or Sri Lanka. Taken together, the geographic distribution data in Table 8 and the inequality metrics illustrated in Fig 5 highlight the severe under representation of Sub-Saharan Africa and rural South Asia, the exclusion of marginalized populations, and a high risk of bias arising from the absence of diverse training data. (A) Geographic distribution of included studies across seven regions, highlighting the concentration of research in East Asia and high-income countries. (B) Key inequality metrics: 75.5% of studies lack external validation, 83.6% are classified as high risk of bias per PROBAST-AI criteria, and high-income countries outproduce LMICs by a 23:1 ratio after DALY-burden adjustment. 3.6 Methodological quality and bias assessment Methodological appraisal using the adapted PROBAST-AI framework and the scoring rubric in S1 Table revealed pervasive limitations across multiple domains. As shown in Table 10, high risk of bias was observed in participant selection (68.1%), predictor measurement (56.3%), and model development (65.5%). These deficiencies were primarily driven by convenience sampling, retrospective feature extraction, inadequate documentation of preprocessing pipelines, and insufficient handling of class imbalance. Validation represented the most critical weakness, with 77.3% of studies classified as high risk and 75.5% lacking any form of external testing. Temporal validation was absent in 89% of studies, substantially increasing susceptibility to optimism bias and overestimation of real-world performance. The multidimensional quality patterns illustrated in Fig 6 further demonstrate systematic gaps in validation rigor, data transparency, and LMIC applicability. Overall, 83.6% of studies were classified as high risk of bias, substantially restricting the reliability of current evidence for clinical and policy-level decision-making. Explainability and clinical integration features summarized in Table 11 indicate limited adoption of transparency mechanisms. SHAP or LIME was implemented in only 6.9% of studies, while the majority (26.7%) employed no interpretability tools. Although attention-based visualizations were used in 4.3% of imaging studies, their clinical actionability remained limited. Dumbbell chart comparing mean quality scores (1–5 scale) across six dimensions. Blue dots represent HIC-originating studies; red dots represent LMIC-originating studies. Delta values indicate the HIC–LMIC score difference. Validation rigor (highlighted) shows the largest gap (+2.9), driven by differences in external validation rates (35.0% vs 16.7%) and multi-center participation (45.0% vs 22.2%). 3.7 Temporal trends and innovation trajectory The temporal analysis reported in Fig 7 shows that AI research methodologies have evolved substantially over 2021–2025. Multimodal integration increased from 15.4% in 2021 to 51.4% in 2025. Explainable AI features evolved from 0% in 2021 to 34.3% in 2025. The average sample size of datasets used in the studies increased from approximately 920 in 2021–4,500 in 2025. Open science practices also increased markedly, from 23.1% in 2021 to 65.7% in 2025. Based on the timeline in Fig 7, after 2023, multimodal modeling, validation frameworks, and transparency mechanisms have been adopted more rapidly. However, as of 2025, the percentage of AI models with some form of external validation is only 37.1%, and prospective studies remain below 20%. Gantt-style timeline showing emergence and adoption rates of key methodological innovations: multimodal integration, external validation, XAI, prospective design, and open science practices. Color intensity represents adoption percentage per year. 4. Discussion 4.1 Synthesis of key findings AI and ML applications in women’s reproductive health have advanced substantially from 2021 to 2025, yet a critical translational gap persists between algorithmic performance and real-world applicability, particularly in LMICs. Three cross-cutting patterns define the field. First, ensemble methods — Random Forest and XGBoost — dominate clinical prediction tasks, reporting consistently strong performance across studies and outperforming imaging-based approaches in resource efficiency. Second, CNN-based imaging models show a directional pattern of performance reduction moving from single-centre evaluations to external validation settings, consistent with the well-documented literature on optimism bias in internally validated clinical AI. This pattern is observed across all five major algorithm families and replicated in disease-stratified sensitivity analyses. We emphasise, however, that the magnitude of this reduction carries considerable uncertainty: estimates are derived from small matched subsets (n = 5–51 per algorithm), aggregate heterogeneous disease contexts, and assume independence that is partially violated by overlapping training cohorts. The observed pattern of lower reported performance in externally validated versus internally validated studies should therefore be interpreted as descriptive evidence consistent with established literature on optimism bias in clinical AI, rather than a precise measurement of any single model family’s generalisation gap. Third, despite multimodal integration rising from 15.4% in 2021 to 51.4% in 2025 and XAI adoption growing from 0% to 34.3%, external validation remained below 37.1% even in 2025, indicating that methodological sophistication has outpaced clinical validation infrastructure. Table 12 summarises domain-level findings alongside LMIC-specific recommendations. 4.2 Disease-specific interpretation Predictive performance across disease domains is shaped by biological complexity, diagnostic standardisation, and data modality requirements. Thyroid disorder studies reported the highest and most stable performance range, underpinned by well-validated biomarker thresholds (TSH, T3, T4) and Tier 1–2 resource inputs, making them the strongest candidates for near-term LMIC deployment. PCOS studies similarly reported strong performance, yet 80% relied on ultrasound — a Tier 2 resource classified as Low Feasibility in our LMIC scheme — limiting scalability beyond district hospital settings. The heterogeneity introduced by competing diagnostic criteria (Rotterdam, NIH, AES) further complicates cross-study comparisons and model replication in populations with different phenotypic distributions. Endometriosis presented the most challenging diagnostic landscape, with the widest range of reported performance and only 11.1% external validation, reflecting dependence on laparoscopic confirmation as the reference standard. Emerging non-invasive approaches — including electrochemical biosensors, urine infrared spectrometry, and symptom-based ML classifiers classified as High Feasibility — offer more viable LMIC pathways and directly address the 8–12 year diagnostic delay that causes preventable morbidity. GDM and preeclampsia studies demonstrated the most credible external validation evidence (16.7–18.2%), with first-trimester clinical and laboratory inputs costing $12–18 per assessment supporting feasibility for antenatal clinic deployment by midwifery-level personnel. Infertility and IVF studies (0% external validation) remain the least translatable domain given their dependence on specialised embryo imaging technology. 4.3 Methodological quality, systemic bias, and explainability The adapted PROBAST-AI assessment classified 83.6% of studies as high risk of bias. The dominant drivers were single-centre recruitment (78%), convenience sampling (64%), absent external validation (75.5%), and unaddressed class imbalance (53%). Three specific bias sources have outsized impact on translational inference. Optimism bias — from internal cross-validation on small datasets — accounts for the directional pattern of performance reduction observed across all major algorithm families when moving from single-centre to external testing. Feature selection leakage, identified in 31% of model development studies, occurs when dimensionality reduction is performed prior to cross-validation fold splitting, artificially inflating performance — a problem particularly prevalent in multi-omics biomarker discovery research. Taxonomy assignment inter-rater agreement was assessed separately from general data extraction. On a 20% random subsample (n = 23 studies), taxonomy classification yielded (substantial agreement). The primary source of disagreement was the boundary between Category 2 (Medical Imaging + AI) and Category 4 (Reproductive and Endocrine Health Applications), which accounted for 71% of discrepant cases. As pre-specified in Section 2.6, the applied decision rule was: Category 2 was assigned when the primary methodological contribution was an imaging architecture or segmentation method; Category 4 was assigned when the primary contribution was clinical integration of AI with reproductive health data regardless of imaging involvement. Studies where ultrasound features were combined with clinical parameters without novel imaging methodology were classified as Category 4. Temporal validation was absent in 89% of all studies, meaning that no model in the reviewed literature has demonstrated sustained performance over changing clinical populations — a prerequisite for safe, ongoing clinical use. Explainability adoption, while improving from 0% in 2021 to 34.3% in 2025, remains insufficient for clinical integration. SHAP and LIME were applied in only 6.9% of studies, predominantly in PCOS, GDM, and preeclampsia prediction models where feature importance explanations can directly inform clinical reasoning. The 26.7% of studies employing no interpretability mechanism whatsoever represent “black box” outputs that are unlikely to gain clinician trust or regulatory approval. Hybrid CNN-ML architectures — used in 92% of AI/ML Methodology development studies — offer promising pathways to combining imaging performance with structured explainability, but require substantially larger and more diverse training datasets than currently available in most reproductive health research settings. 4.4 Geographic equity and translational challenges Several deep-seated structural inequities were identified via the geographic analysis. The East Asia region comprised 39.6% of all studies (China n = 28; South Korea n = 15), while the region with the fewest was Sub-Saharan Africa with 1.7% (2 studies). Population-adjusted research output was calculated as: (number of studies in region) ÷ (women aged 15–49 in region, in millions). Population denominators were drawn from UN World Population Prospects 2023 estimates aggregated by World Bank income classification (HIC vs LMIC). HICs contributed 56 studies against approximately 250 million women of reproductive age, yielding 0.224 studies per million. LMICs contributed 60 studies against approximately 1,540 million women, yielding 0.039 studies per million — a ratio of 5.7:1 in absolute terms and 23:1 after weighting by reproductive-health DALY burden (Global Burden of Disease 2021). Sensitivity analyses using alternative denominators are reported in the next paragraph. Adjusted ratios from sensitivity analyses were: GDP = 7:1; research workforce FTEs = 17:1; reproductive health DALYs = 31:1. South Asia contributed only 15.5% of included studies; India contributed 18 studies but all were conducted in urban tertiary care centres, the 3 studies from Pakistan included fewer than 350 participants, there were 4 single-centre studies from Bangladesh, and none from Nepal and Sri Lanka. This geographical bias suggests that current models of population-specific parameters (e.g., hormone reference ranges, BMI–adiposity relations, PCOS clinical and laboratory features) may not be validated in South Asian and African populations. Infrastructure constraints compound the equity gap. Models requiring Tier 3 inputs (MRI, advanced ultrasound, multi-omics) fall in the Very Low Feasibility category, placing them outside the reach of primary care settings. Data poverty — small samples, missing data, fragmented paper-based records — affects 89% of LMIC-originating studies. Sociocultural barriers, including stigma surrounding menstruation and infertility, further constrain data availability and AI adoption, necessitating community co-design approaches largely absent from current literature. 4.5 Translational research agenda and deployment considerations Fig 8 summarises a three-tier translational research agenda derived from the patterns identified in this review. The agenda is presented as a research roadmap, not a deployment guideline; empirical validation through prospective LMIC studies and Delphi-based expert consensus is required before any component can inform clinical practice. An actionable deployment roadmap for AI-based women’s reproductive health tools in low- and middle-income countries. 4.5.1 Near-term priorities (0–2 years). Reviewed evidence supports immediate research investment in: - interpretable ensemble models (Random Forest, XGBoost) trained on Tier 1 inputs for PCOS, GDM, and thyroid screening, with mandatory SHAP-based explainability; - prospective external validation of these models in LMIC primary-care populations; and - centralised, de-identified data pooling among urban tertiary centres to build foundational reference cohorts. 4.5.2 Medium-term aspirations (3–5 years). Federated learning and multi-site domain adaptation represent medium-term targets contingent on prerequisite infrastructure: standardised EHR systems at the primary care level, reliable internet connectivity in rural areas, multi-site IRB harmonisation, and locally enforceable data protection legislation. None of these prerequisites currently obtains in most South Asian or Sub-Saharan African settings. Existing frameworks such as OpenHIE [138,139] and WHO SMART guidelines [140] provide entry points but require substantial adaptation. Federated learning should not be interpreted as a near-term deployment recommendation. 4.5.3 Long-term horizon (5 + years). Embedding AI competency in nursing and community health worker curricula, LMIC-led validation consortia, and locally governed regulatory frameworks for clinical AI represent structural changes that will determine whether the technical advances of the next decade reach the populations currently absent from the evidence base. 4.6 Comparison with prior reviews and methodological limitations This review extends prior reviews of AI in reproductive health by providing descriptive evidence consistent with well-documented optimism bias patterns, a resource-tier classification to characterise infrastructure requirements, and a geographic equity analysis across 27 countries. Prior reviews have characterised the field as showing “high promise” without systematically describing the pattern between single-centre and externally validated results — a pattern this review characterises descriptively across five algorithm families. Several methodological limitations must be acknowledged: restriction to English-language publications likely excludes relevant research from Francophone Sub-Saharan Africa and Portuguese-speaking contexts; exclusion of preprints may introduce publication lag in a rapidly evolving field; heterogeneity in outcome measures meant formal meta-analysis was not appropriate; and the scoping nature of the review means individual study quality grading carries less evidential weight than a full systematic review. 4.7 Domain adaptation and harmonization as mitigation strategies for imaging performance decay The larger reported performance decay in imaging-based models [141] relative to clinical-data models is consistent with well-documented scanner-level and protocol-level domain shift in medical imaging AI. Emerging domain adaptation and harmonization methods including histogram matching [142], CycleGAN [143]-based style transfer, and test-time adaptation approaches [144] have demonstrated decay reduction in comparable imaging tasks such as mammography and chest radiography across multi-site deployments. However, applying these approaches in LMIC reproductive imaging contexts faces a fundamental barrier: domain adaptation tools require paired or unpaired multi-site imaging data for training, which is currently unavailable given the near-absence of digitized reproductive imaging archives in primary and secondary care settings across South Asia and Sub-Saharan Africa. The practical implication of the imaging decay pattern is therefore not to invest in harmonization infrastructure in the near term, but rather to reinforce the prioritisation of clinical-data ensemble models for LMIC deployment. Domain adaptation investments become relevant only after multi-site imaging data collection infrastructure is established — a medium-to-long-term horizon. Researchers in HIC settings developing imaging models intended for LMIC adaptation should prospectively collect harmonization-enabling datasets and report scanner metadata to facilitate future transfer. 5. Conclusions This scoping review synthesises 116 studies on AI and ML applications in women’s reproductive and hormonal health (2021–2025), offering four principal contributions: a six-category taxonomy of methodological trends; a qualitative resource-tier classification for LMIC applicability; descriptive evidence of systematic performance reduction upon external validation; and a geographic mapping revealing significant research voids. Ensemble methods applied to Tier 1 clinical and laboratory data represent the most immediately deployable AI approach for LMIC settings, with inputs available at primary care level, while imaging-intensive CNN models remain unsuitable for near-term LMIC deployment due to infrastructure constraints. With 83.6% of studies classified as high risk of bias and 75.5% lacking external validation, and with the evidence base heavily concentrated in East Asia (39.6%) and high-income countries, current evidence is insufficient to guide population-level deployment or broad translational application in low-resource settings. For LMICs such as Bangladesh, a phased strategy is proposed, contingent on resolution of the methodological limitations identified in this review: piloting of low-cost AI screening tools for PCOS, GDM, and thyroid disorders should proceed only after prospective validation in local populations; medium-term establishment of multi-country validation consortia; and long-term embedding of AI competency in health worker curricula alongside LMIC-contextualised regulatory frameworks. Beyond the descriptive synthesis, the four-tier LMIC feasibility classification system introduced in this review is intended as a reusable instrument: future researchers, funders, and health-system planners can apply it independently to characterize the deployment readiness of new AI tools in resource-constrained settings, regardless of disease domain. Priority future directions include prospective multi-centre LMIC validation, non-invasive endometriosis diagnostics, algorithmic fairness evaluation, and implementation science — all oriented toward ensuring that AI progress in reproductive health reaches the women who bear the greatest unmet burden. Consequently, the findings of this review reflect trends predominantly from East Asian and high-income settings and should not be interpreted as globally generalizable. Translational recommendations for LMICs remain theoretical until validated in local, resource-constrained populations. These conclusions should be interpreted within the descriptive scope of scoping review methodology and not as causal or comparative inferences. Supporting information S1 File. PRISMA-ScR Checklist. Completed PRISMA-ScR checklist covering all 20 essential items and 2 optional items. https://doi.org/10.1371/journal.pone.0358557.s001 (DOCX) S1 Table. Risk-of-bias scoring rubric. Explicit domain-specific criteria used to classify studies as low, moderate, or high risk of bias (adapted from PROBAST-AI, QUADAS-AI, and TRIPOD-AI). https://doi.org/10.1371/journal.pone.0358557.s002 (DOCX) S2 Table. LMIC applicability categorical criteria. Full definitions of the four LMIC-feasibility categories (High, Moderate, Low, Very Low) used to classify studies. https://doi.org/10.1371/journal.pone.0358557.s003 (DOCX)

References

- 1. Zeng W, Gan D, Ou J, Tomlinson B. Global trends and health system impact on polycystic ovary syndrome: A comprehensive analysis of age-stratified females from 1990 to 2021. Front Reprod Health. 2025;7:1642369. pmid:41210405 - 2. Fauser BCJM, Adamson GD, Boivin J, Chambers GM, de Geyter C, Dyer S, et al. Declining global fertility rates and the implications for family planning and family building: An IFFS consensus document based on a narrative review of the literature. Hum Reprod Update. 2024;30(2):153–73. pmid:38197291 - 3. Guo Z-Q. Precision pharmacology in menopause: Advances, challenges, and future innovations for personalized management. Front Reprod Health. 2025;7:1694240. pmid:41323403 - 4. Gao X, Zhao S, Du Y, Yang Z, Tian Y, Zhao J, et al. Data-driven subtypes of polycystic ovary syndrome and their association with clinical outcomes. Nat Med. 2025;31(12):4214–24. pmid:41162652 - 5. Zhang J, Pang M, Li L, Guo C. Global, regional, and national burden of endometriosis among women of reproductive age, 1990–2021: Insights from the global burden of disease study 2021. PLoS One. 2025;20(11):e0337074. - 6. World Health Organization. Polycystic ovary syndrome; 2025 [cited 2026 Jan 23]. Available from: https://www.who.int/news-room/fact-sheets/detail/polycystic-ovary-syndrome - 7. De Corte P, Klinghardt M, von Stockum S, Heinemann K. Time to diagnose endometriosis: Current status, challenges and regional characteristics-A systematic literature review. BJOG. 2025;132(2):118–30. pmid:39373298 - 8. de Kok L, Boersen Z, Coppus S, van Haaps A, van Hanegem N, Klinkert E. Diagnostic delay in endometriosis: Is there any progress? Reprod BioMed Online. 2025:105405. - 9. Li R, Zhang L, Liu Y. Global and regional trends in the burden of surgically confirmed endometriosis from 1990 to 2021. Reprod Biol Endocrinol. 2025;23(1):88. pmid:40483411 - 10. Yi S, Yam ELY, Cheruvettolil K, Linos E, Gupta A, Palaniappan L, et al. Perspectives of digital health innovations in low- and middle-income health care systems from South and Southeast Asia. J Med Internet Res. 2024;26:e57612. pmid:39586089 - 11. Jonayed M, Rumi MH. Towards women’s digital health equity: A qualitative inquiry into attitude and adoption of reproductive mHealth services in Bangladesh. PLOS Digit Health. 2024;3(10):e0000637. pmid:39405293 - 12. Ghaderzadeh M, Garavand A, Salehnasab C. Artificial intelligence in polycystic ovary syndrome: A systematic review of diagnostic and predictive applications. BMC Med Inform Decis Mak. 2025;25(1):427. pmid:41286838 - 13. Agirsoy M, Oehlschlaeger MA. A machine learning approach for non-invasive PCOS diagnosis from ultrasound and clinical features. Sci Rep. 2025;15(1):33638. pmid:41022847 - 14. Burla L, Metzler JM, Kalaitzopoulos DR, Kamm S, Ormos M, Passweg D, et al. Artificial intelligence in endometriosis care: A comparative analysis of large language model and human specialist responses to endometriosis-related queries. Eur J Obstet Gynecol Reprod Biol. 2025;313:114625. pmid:40829501 - 15. Mengistu S, Tamrat T, Betran A-P, Pirsch S, Ferretti A, Mburu G, et al. The use of artificial intelligence in sexual and reproductive health: A comprehensive scoping review. NPJ Womens Health. 2025;3(1):70. pmid:41427017 - 16. Zhu T, Li K, Herrero P, Georgiou P. Deep learning for diabetes: A systematic review. IEEE J Biomed Health Inform. 2021;25(7):2744–57. pmid:33232247 - 17. Tricco AC, Lillie E, Zarin W, O’Brien KK, Colquhoun H, Levac D, et al. PRISMA extension for scoping reviews (PRISMA-ScR): Checklist and explanation. Ann Intern Med. 2018;169(7):467–73. pmid:30178033 - 18. Munna MMH, Islam T, Sultana A, Saima F, Faruk O, Sultana T, et al. Screening protocol and study selection data for: Artificial intelligence for women’s reproductive health: A systematic review of global diagnostic trends and a translational informatics framework for low-resource settings; 2026. Data repository. Available from: https://doi.org/10.17632/ydx7v73t3k.2 - 19. Kloda LA, Boruff JT, Cavalcante AS. A comparison of patient, intervention, comparison, outcome (PICO) to a new, alternative clinical question framework for search skills, search results, and self-efficacy: A randomized controlled trial. J Med Libr Assoc. 2020;108(2):185–94. pmid:32256230 - 20. Ouzzani M, Hammady H, Fedorowicz Z, Elmagarmid A. Rayyan-a web and mobile app for systematic reviews. Syst Rev. 2016;5(1):210. pmid:27919275 - 21. Rayyan Systems Inc. Rayyan; 2026. Web application. Available from: https://www.rayyan.ai - 22. Balampanos D, Kokkotis C, Stampoulis T, Avloniti A, Pantazis D, Protopapa M, et al. Interpretable machine learning for osteopenia detection: A proof-of-concept study using bioelectrical impedance in perimenopausal women. J Funct Morphol Kinesiol. 2025;10(3):262. pmid:40700198 - 23. Shubhangi DC, Salma U. Analysis and interpretation of physiological social demographic parameter in menopausal women. IJCSMC. 2024;13(10):58–66. - 24. Chang C-Y, Peng C-H, Chen F-Y, Huang L-Y, Kuo C-H, Chu T-W, et al. The risk factors determined by four machine learning methods for the change of difference of bone mineral density in post-menopausal women after three years follow-up. Sci Rep. 2024;14(1):23234. pmid:39369003 - 25. Kim H, Khomidov M, Lee J-H. XGBoost and SHAP-based analysis of risk factors for hypertension classification in Korean postmenopausal women. Bioengineering (Basel). 2025;12(6):659. pmid:40564475 - 26. Ding X, Tao T, Fu J, Zhang Y, Yang X, Ren M, et al. Is integrative therapy of traditional Chinese medicine and progesterone capsule more effective than monotherapies in oligomenorrhea and hypomenorrhea? Evidence based on a multi-center randomized controlled trial and metabolomic profile. J Transl Int Med. 2025;13(4):349–65. pmid:40861072 - 27. Santos L, Azevedo M, Shimamura L, Nogueira A, Reis F, Schor E, et al. Machine learning as a clinical decision support tool for diagnosing superficial peritoneal endometriosis in women with dysmenorrhea and acyclic pelvic pain. MRAJ. 2024;12(12). - 28. Manjunath H, Kolekar V, Bhuvaneshwari K, Jain R, Aravind A, Dev S. Comprehensive detection and analysis of depression in women through advanced machine learning techniques. 2024 IEEE 4th International Conference on ICT in Business Industry & Government (ICTBIG). IEEE; 2024. p. 1–5. - 29. Wang W, Zeng W, Yang S. A stacked machine learning-based classification model for endometriosis and adenomyosis: A retrospective cohort study utilizing peripheral blood and coagulation markers. Front Digit Health. 2024;6:1463419. pmid:39347446 - 30. Liu Z, Liu Z, Wang Y, Wan X, Huang X. Machine learning-based predictive analysis of energy efficiency factors necessary for the HIFU treatment of adenomyosis. Front Physiol. 2025;16:1602866. pmid:40895435 - 31. Raimondo D, Raffone A, Aru AC, Giorgi M, Giaquinto I, Spagnolo E, et al. Application of deep learning model in the sonographic diagnosis of uterine adenomyosis. Int J Environ Res Public Health. 2023;20(3):1724. pmid:36767092 - 32. Shahzad A, Mushtaq A, Sabeeh AQ, Ghadi YY, Mushtaq Z, Arif S, et al. Automated uterine fibroids detection in ultrasound images using deep convolutional neural networks. Healthcare (Basel). 2023;11(10):1493. pmid:37239779 - 33. Sawant A, Kulkarni S, Sawant M. Ultrasound super resolution imaging for accurate uterus tumor detection and malignancy prediction. J Pharm Biomed Anal Open. 2024;3:100029. - 34. Samarasam B, Justin J. Machine learning-based approach for uterine cancer detection and classifier evaluation. J Electron Electromedical Eng Med Inform. 2025;7(3):940–9. - 35. Akpinar E, Bayrak O-C, Nadarajan C, Müslümanoğlu M-H, Nguyen M-D, Keserci B. Role of machine learning algorithms in predicting the treatment outcome of uterine fibroids using high-intensity focused ultrasound ablation with an immediate nonperfused volume ratio of at least 90. Eur Rev Med Pharmacol Sci. 2022;26(22):8376–94. pmid:36459021 - 36. Janghorbani S, Caprio A, Sam L, Lee BC, Sabuncu MR, Lamparello NA, et al. Predicting clinical outcomes and symptom relief in uterine fibroid embolization using machine learning on MRI features. AI. 2025;6(9):200. - 37. Chen M, Kong W, Li B, Tian Z, Yin C, Zhang M, et al. Revolutionizing hysteroscopy outcomes: AI-powered uterine myoma diagnosis algorithm shortens operation time and reduces blood loss. Front Oncol. 2023;13:1325179. pmid:38144535 - 38. Toyohara Y, Sone K, Noda K, Yoshida K, Kato S, Kaiume M, et al. The automatic diagnosis artificial intelligence system for preoperative magnetic resonance imaging of uterine sarcoma. J Gynecol Oncol. 2024;35(3):e24. pmid:38246183 - 39. Xi H, Wang W. Deep learning based uterine fibroid detection in ultrasound images. BMC Med Imaging. 2024;24(1):218. pmid:39160500 - 40. Theis M, Tonguc T, Savchenko O, Nowak S, Block W, Recker F, et al. Deep learning enables automated MRI-based estimation of uterine volume also in patients with uterine fibroids undergoing high-intensity focused ultrasound therapy. Insights Imaging. 2023;14(1):1. pmid:36600120 - 41. Li C, He Z, Lv F, Liu Y, Hu Y, Zhang J, et al. An interpretable MRI-based radiomics model predicting the prognosis of high-intensity focused ultrasound ablation of uterine fibroids. Insights Imaging. 2023;14(1):129. pmid:37466728 - 42. Huo T, Li L, Chen X, Wang Z, Zhang X, Liu S, et al. Artificial intelligence-aided method to detect uterine fibroids in ultrasound images: A retrospective study. Sci Rep. 2023;13(1):3714. pmid:36878941 - 43. Mohanty A, Pattnayak P, Mallick PK, Ambudkar B. Deep convolutional neural networks for automatic detection of uterine fibroids in ultrasound images. Intell Decis Technol. 2025;19(3):1657–74. - 44. Rewcastle E, Gudlaugsson E, Lillesand M, Skaland I, Baak JPA, Janssen EAM. Automated prognostic assessment of endometrial hyperplasia for progression risk evaluation using artificial intelligence. Mod Pathol. 2023;36(5):100116. pmid:36805790 - 45. Martins MS, Valente GB, Pedra YdN, Ribeiro TC, Boldrini NAT, Barcelos MRB, et al. A machine learning approach towards endometriosis screening using infrared spectra of urine. Clinics (Sao Paulo). 2025;80:100760. pmid:40915181 - 46. Xie Z, Feng Y, He Y, Lin Y, Wang X. Identification of biomarkers for endometriosis based on summary-data-based Mendelian randomization and machine learning. Medicine (Baltimore). 2025;104(14):e41804. pmid:40193647 - 47. Sankaravadivel V, Thalavaipillai S. Symptoms based endometriosis prediction using machine learning. Bulletin EEI. 2021;10(6):3102–9. - 48. Macis C, Santoro M, Zybin V, Di Costanzo S, Coada CA, Dondi G, et al. A convolutional neural network tool for early diagnosis and precision surgery in endometriosis-associated ovarian cancer. Appl Sci. 2025;15(6):3070. - 49. Caballero P, Gonzalez-Abril L, Ortega JA, Simon-Soro Á. Data mining techniques for endometriosis detection in a data-scarce medical dataset. Algorithms. 2024;17(3):108. - 50. Wang H, Butler D, Zhang Y, Avery J, Knox S, Ma C, et al. Human-AI collaborative multi-modal multi-rater learning for endometriosis diagnosis. Phys Med Biol. 2024;70(1). pmid:39622166 - 51. Enamorado-Díaz E, Morales-Trujillo L, García-García J-A, Marcos AT, Navarro-Pando J, Escalona-Cuaresma M-J. A novel machine learning-based proposal for early prediction of endometriosis disease. Expert Syst Appl. 2025;271:126621. - 52. Kuyoro AO, Fatade OB, Onuiri EE. Enhancing non-invasive diagnosis of endometriosis through explainable artificial intelligence: A Grad-CAM approach. ATAIML. 2025;4(2):97–108. - 53. Tore U, Abilgazym A, Asunsolo-Del-Barco A, Terzic M, Yemenkhan Y, Zollanvari A, et al. Diagnosis of endometriosis based on comorbidities: A machine learning approach. Biomedicines. 2023;11(11):3015. pmid:38002015 - 54. Korman SE, Vissers G, Gorris MAJ, Verrijp K, Verdurmen WPR, Simons M, et al. Artificial intelligence-based tissue segmentation and cell identification in multiplex-stained histological endometriosis sections. Hum Reprod. 2025;40(3):450–60. pmid:39724530 - 55. Chang X, Miao J. Identification of a disulfidptosis-related genes signature for diagnostic and immune infiltration characteristics in endometriosis. Sci Rep. 2024;14(1):25939. pmid:39472502 - 56. Fell C, Mohammadi M, Morrison D, Arandjelović O, Syed S, Konanahalli P, et al. Detection of malignancy in whole slide images of endometrial cancer biopsies using artificial intelligence. PLoS One. 2023;18(3):e0282577. pmid:36888621 - 57. Sankaravadivel V, Thalavaipillai S, Rajeswar S, Ramlingam P. Feature based analysis of endometriosis using machine learning. IJEECS. 2023;29(3):1700. - 58. Zhang H, Zhang H, Yang H, Shuid AN, Sandai D, Chen X. Machine learning-based integrated identification of predictive combined diagnostic biomarkers for endometriosis. Front Genet. 2023;14:1290036. pmid:38098472 - 59. Zou L, Meng L, Xu Y, Wang K, Zhang J. Revealing the diagnostic value and immune infiltration of senescence-related genes in endometriosis: A combined single-cell and machine learning analysis. Front Pharmacol. 2023;14:1259467. pmid:37860112 - 60. Hu P, Gao Y, Zhang Y, Sun K. Ultrasound image-based deep learning to differentiate tubal-ovarian abscess from ovarian endometriosis cyst. Front Physiol. 2023;14:1101810. pmid:36824470 - 61. Zhao D, Zhang Z, Wang Z, Du Z, Wu M, Zhang T, et al. Diagnosis and prediction of endometrial carcinoma using machine learning and artificial neural networks based on public databases. Genes (Basel). 2022;13(6):935. pmid:35741697 - 62. Takahashi Y, Sone K, Noda K, Yoshida K, Toyohara Y, Kato K, et al. Automated system for diagnosing endometrial cancer by adopting deep-learning technology in hysteroscopy. PLoS One. 2021;16(3):e0248526. pmid:33788887 - 63. Balogh DB, Hudelist G, Bļizņuks D, Raghothama J, Becker CM, Horace R, et al. FEMaLe: The use of machine learning for early diagnosis of endometriosis based on patient self-reported data-Study protocol of a multicenter trial. PLoS One. 2024;19(5):e0300186. pmid:38722932 - 64. Blass I, Sahar T, Shraibman A, Ofer D, Rappoport N, Linial M. Revisiting the risk factors for endometriosis: A machine learning approach. J Pers Med. 2022;12(7):1114. pmid:35887611 - 65. Fremond S, Andani S, Barkey Wolf J, Dijkstra J, Melsbach S, Jobsen JJ, et al. Interpretable deep learning model to predict the molecular classification of endometrial cancer from haematoxylin and eosin-stained whole-slide images: A combined analysis of the PORTEC randomised trials and clinical cohorts. Lancet Digit Health. 2023;5(2):e71–82. pmid:36496303 - 66. Pal A, Biswas S, O Kare SP, Biswas P, Jana SK, Das S, et al. Development of an impedimetric immunosensor for machine learning-based detection of endometriosis: A proof of concept. Sens Actuators B: Chem. 2021;346:130460. - 67. Zhao N, Hao T, Zhang F, Ni Q, Zhu D, Wang Y, et al. Application of machine learning techniques in the diagnosis of endometriosis. BMC Womens Health. 2024;24(1):491. pmid:39237940 - 68. Shi S, Huang C, Tang X, Liu H, Feng W, Chen C. Identification and verification of diagnostic biomarkers for deep infiltrating endometriosis based on machine learning algorithms. J Biol Eng. 2024;18(1):70. pmid:39587559 - 69. Goldstein A, Cohen S. Self-report symptom-based endometriosis prediction using machine learning. Sci Rep. 2023;13(1):5499. pmid:37016132 - 70. Patil PS, Patil MB, Jadhav DK, Patil SB, Bodhe MYU. Predictive analytics and health monitoring system for early detection of infertility risks among working women. Int J Environ Sci. 2025;11(16s). - 71. Lin Q, Fang Z-J. Establishment and evaluation of a risk prediction model for gestational diabetes mellitus. World J Diabetes. 2023;14(10):1541–50. pmid:37970129 - 72. Prashanthan J, Prashanthan A. Predicting the future risk of developing type 2 diabetes in women with a history of gestational diabetes mellitus using machine learning and explainable artificial intelligence. Prim Care Diabetes. 2025;19(6):658–66. pmid:41006077 - 73. Wu Y-T, Zhang C-J, Mol BW, Kawai A, Li C, Chen L, et al. Early prediction of gestational diabetes mellitus in the Chinese population via advanced machine learning. J Clin Endocrinol Metab. 2021;106(3):e1191–205. pmid:33351102 - 74. Vivek Khanna V, Chadaga K, Sampathila N, Prabhu S, Chadaga P R, Bhat D, et al. Explainable artificial intelligence-driven gestational diabetes mellitus prediction using clinical and laboratory markers. Cogent Eng. 2024;11(1):2330266. - 75. Hu X, Hu X, Yu Y, Wang J. Prediction model for gestational diabetes mellitus using the XG Boost machine learning algorithm. Front Endocrinol (Lausanne). 2023;14:1105062. pmid:36967760 - 76. Zhou H, Chen W, Chen C, Zeng Y, Chen J, Lin J, et al. Predictive value of ultrasonic artificial intelligence in placental characteristics of early pregnancy for gestational diabetes mellitus. Front Endocrinol (Lausanne). 2024;15:1344666. pmid:38544693 - 77. Ali N, Khan W, Ahmad A, Masud MM, Adam H, Ahmed LA. Predictive modeling for the diagnosis of gestational diabetes mellitus using epidemiological data in the United Arab Emirates. Information. 2022;13(10):485. - 78. Kaya Y, Bütün Z, Çelik Ö, Salik EA, Tahta T, Yavuz AA. The early prediction of gestational diabetes mellitus by machine learning models. BMC Pregnancy Childbirth. 2024;24(1):574. pmid:39217284 - 79. Bigdeli SK, Ghazisaedi M, Ayyoubzadeh SM, Hantoushzadeh S, Ahmadi M. Predicting Gestational Diabetes Mellitus in the first trimester using machine learning algorithms: A cross-sectional study at a hospital fertility health center in Iran. BMC Med Inform Decis Mak. 2025;25(1):3. pmid:39754258 - 80. Zaky H, Fthenou E, Srour L, Farrell T, Bashir M, El Hajj N, et al. Machine learning based model for the early detection of Gestational Diabetes Mellitus. BMC Med Inform Decis Mak. 2025;25(1):130. pmid:40082942 - 81. Du Y, Rafferty AR, McAuliffe FM, Wei L, Mooney C. An explainable machine learning-based clinical decision support system for prediction of gestational diabetes mellitus. Sci Rep. 2022;12(1):1170. pmid:35064173 - 82. Kang BS, Lee SU, Hong S, Choi SK, Shin JE, Wie JH, et al. Prediction of gestational diabetes mellitus in Asian women using machine learning algorithms. Sci Rep. 2023;13(1):13356. pmid:37587201 - 83. Montgomery-Csobán T, Kavanagh K, Murray P, Robertson C, Barry SJE, Vivian Ukah U, et al. Machine learning-enabled maternal risk assessment for women with pre-eclampsia (the PIERS-ML model): A modelling study. Lancet Digit Health. 2024;6(4):e238–50. pmid:38519152 - 84. Zhang X, Chen Y, Salerno S, Li Y, Zhou L, Zeng X, et al. Prediction of severe preeclampsia in machine learning. Med Nov Technol Devices. 2022;15:100158. - 85. Domínguez-del Olmo P, Herraiz I, Villalaín C, Galindo A, Moreno-Espino M, Ayala JL. Comprehensive approach with machine learning techniques to investigate early-onset preeclampsia and its long-term cardiovascular implications. Appl Sci. 2025;15(16):8887. - 86. Shyu I-L, Liu C-F, Tsai Y-C, Ma Y-S, Kuo T-N, Yow-Ling S. Machine learning predictive system to predict the risk of developing pre-eclampsia. BMJ Health Care Inform. 2025;32(1):e101151. pmid:41106843 - 87. Gómez-Jemes L, Oprescu AM, Chimenea-Toscano Á, García-Díaz L, Romero-Ternero MdC. Machine learning to predict pre-eclampsia and intrauterine growth restriction in pregnant women. Electronics. 2022;11(19):3240. - 88. Kaya Y, Bütün Z, Çelik Ö, Salik EA, Tahta T. Risk assessment for preeclampsia in the preconception period based on maternal clinical history via machine learning methods. J Clin Med. 2024;14(1):155. pmid:39797241 - 89. Araújo DC, de Macedo AA, Veloso AA, Alpoim PN, Gomes KB, Carvalho MdG, et al. Complete blood count as a biomarker for preeclampsia with severe features diagnosis: A machine learning approach. BMC Pregnancy Childbirth. 2024;24(1):628. pmid:39354367 - 90. Wu Y, Shen L, Zhao L, Lin X, Xu M, Tu Z, et al. Noninvasive early prediction of preeclampsia in pregnancy using retinal vascular features. NPJ Digit Med. 2025;8(1):188. pmid:40188283 - 91. Ansbacher-Feldman Z, Syngelaki A, Meiri H, Cirkin R, Nicolaides KH, Louzoun Y. Machine-learning-based prediction of pre-eclampsia using first-trimester maternal characteristics and biomarkers. Ultrasound Obstet Gynecol. 2022;60(6):739–45. pmid:36454636 - 92. Torres-Torres J, Villafan-Bernal J, Martinez-Portilla R, Hidalgo-Carrera J, Estrada-Gutierrez G, Adalid-Martinez-Cisneros R, et al. Performance of machine-learning approach for prediction of pre-eclampsia in a middle-income country. Ultrasound Obstet Gynecol. 2024;63(3):350–7. - 93. Tiruneh SA, Rolnik DL, Teede HJ, Enticott J. Prediction of pre-eclampsia with machine learning approaches: Leveraging important information from routinely collected data. Int J Med Inform. 2024;192:105645. pmid:39393122 - 94. Sufian MA, Hamzi W, Hamzi B, Sagar ASMS, Rahman M, Varadarajan J, et al. Innovative machine learning strategies for early detection and prevention of pregnancy loss: the vitamin D connection and gestational health. Diagnostics (Basel). 2024;14(9):920. pmid:38732334 - 95. Yan S, Xiong F, Xin Y, Zhou Z, Liu W. Automated assessment of endometrial receptivity for screening recurrent pregnancy loss risk using deep learning-enhanced ultrasound and clinical data. Front Physiol. 2024;15:1404418. pmid:39777360 - 96. Koong K, Preda V, Jian A, Liquet-Weiland B, Di Ieva A. Application of artificial intelligence and radiomics in pituitary neuroendocrine and sellar tumors: A quantitative and qualitative synthesis. Neuroradiology. 2022;64(4):647–68. pmid:34839380 - 97. Fang Y, Wang H, Feng M, Zhang W, Cao L, Ding C, et al. Machine-learning prediction of postoperative pituitary hormonal outcomes in nonfunctioning pituitary adenomas: A multicenter study. Front Endocrinol (Lausanne). 2021;12:748725. pmid:34690934 - 98. Zheng A, Tang D, He H, Liang X. Artificial intelligence-driven approaches in pituitary neuroendocrine tumors: Integrating endocrine-metabolic profiling for enhanced diagnostics and therapeutics. Front Endocrinol (Lausanne). 2025;16:1618412. pmid:41180200 - 99. Li Q, Zhu Y, Chen M, Guo R, Hu Q, Lu Y, et al. Development and validation of a deep learning algorithm to automatic detection of pituitary microadenoma from MRI. Front Med (Lausanne). 2021;8:758690. pmid:34912820 - 100. Dai C, Sun B, Wang R, Kang J. The application of artificial intelligence and machine learning in pituitary adenomas. Front Oncol. 2021;11:784819. pmid:35004306 - 101. Mohammadzadeh I, Hajikarimloo B, Niroomand B, Faizi N, Eini P, Habibi MA, et al. Prediction of recurrence after surgery for pituitary adenoma using machine learning- based models: Systematic review and meta-analysis. BMC Endocr Disord. 2025;25(1):158. pmid:40597183 - 102. Park YW, Eom J, Kim S, Kim H, Ahn SS, Ku CR, et al. Radiomics with ensemble machine learning predicts dopamine agonist response in patients with prolactinoma. J Clin Endocrinol Metab. 2021;106(8):e3069–77. pmid:33713414 - 103. Fan Y, Li Y, Bao X, Zhu H, Lu L, Yao Y, et al. Development of machine learning models for predicting postoperative delayed remission in patients with Cushing’s disease. J Clin Endocrinol Metab. 2021;106(1):e217–31. pmid:33000120 - 104. Guleria K, Sharma S, Kumar S, Tiwari S. Early prediction of hypothyroidism and multiclass classification using predictive machine learning and deep learning. Meas: Sens. 2022;24:100482. - 105. Shiuh TL, Khai WK, Xin YC, Wai CY. Prediction of thyroid disease using machine learning approaches and featurewiz selection. JTEC. 2023;15(3):9–16. - 106. Abbad Ur Rehman H, Lin C-Y, Mushtaq Z, Su S-F. Performance analysis of machine learning algorithms for thyroid disease. Arab J Sci Eng. 2021;46(10):9437–49. - 107. Afshan N, Mushtaq Z, Alamri FS, Qureshi MF, Khan NA, Siddique I. Efficient thyroid disorder identification with weighted voting ensemble of super learners by using adaptive synthetic sampling technique. AIMS Math. 2023;8(10):24274–309. - 108. Obaido G, Achilonu O, Ogbuokiri B, Amadi CS, Habeebullahi L, Ohalloran T, et al. An improved framework for detecting thyroid disease using filter-based feature selection and stacking ensemble. IEEE Access. 2024;12:89098–112. - 109. Akhtar T, Gilani SO, Mushtaq Z, Arif S, Jamil M, Ayaz Y, et al. Effective voting ensemble of homogenous ensembling with multiple attribute-selection approaches for improved identification of thyroid disorder. Electronics. 2021;10(23):3026. - 110. Akter S, Mustafa HA. Analysis and interpretability of machine learning models to classify thyroid disease. PLoS One. 2024;19(5):e0300670. pmid:38820460 - 111. Mir MS, Fayaz SA, Zaman M, Agrawal S. An application of traditional and ensemble machine learning approaches to redefine thyroid disorder diagnosis. MMEP. 2024;11(9):2437–46. - 112. Sankar R, Mahulikar C, Viswanatha V. Thyroid disease detection using machine learning approach. J Xi’an Univ Archit Technol. 2023;7:327–34. - 113. Sultana A, Islam R. Machine learning framework with feature selection approaches for thyroid disease classification and associated risk factors identification. J Electr Syst Inf Technol. 2023;10(1):32. - 114. Gupta P, Rustam F, Kanwal K, Aljedaani W, Alfarhood S, Safran M, et al. Detecting thyroid disease using optimized machine learning model based on differential evolution. Int J Comput Intell Syst. 2024;17(1):3. - 115. Sanju P, Ahmed NSS, Ramachandran P, Sajid PM, Jayanthi R. Enhancing thyroid disease prediction and comorbidity management through advanced machine learning frameworks. Clin eHealth. 2025;8:7–16. - 116. Sankar S, Potti A, Chandrika GN, Ramasubbareddy S. Thyroid disease prediction using XGBoost algorithms. JMM. 2022. - 117. Faris N, Sahi A, Diykh M, Abdulla S, Siuly S. Enhanced Polycystic Ovary Syndrome diagnosis model leveraging a K-means based genetic algorithm and ensemble approach. Intell-Based Med. 2025;11:100253. - 118. Yan X, Yang Z, Zhao H, Feng G, Li S, Li Y, et al. Unveiling lipoprotein subfractions signature in high-FNPO PCOS: Implications for PCOM diagnosis and risk assessment using advanced machine learning models. BMC Med. 2025;23(1):289. pmid:40389976 - 119. Reka S, Praba TS, Prasanna M, Reddy VNN, Amirtharajan R. Automated high precision PCOS detection through a segment anything model on super resolution ultrasound ovary images. Sci Rep. 2025;15(1):16832. pmid:40369044 - 120. Ghosh A, Srinivasan K. EffiDenseGenOp: ensemble transfer learning with hyperparameter tuning using genetic algorithm optimization for PCOS detection from ultrasound sonography images. IEEE Access. 2025;13:54285–312. - 121. Zhu S, Huang Z, Chen X, Jiang W, Zhou Y, Zheng B, et al. Construction and evaluation of machine learning-based prediction model for live birth following fresh embryo transfer in IVF/ICSI patients with polycystic ovary syndrome. J Ovarian Res. 2025;18(1):70. pmid:40186314 - 122. Zhao B, Wen L, Huang Y, Fu Y, Zhou S, Liu J, et al. A deep learning-based automatic recognition model for polycystic ovary ultrasound images. Balkan Med J. 2025;42(5):419–28. pmid:40785235 - 123. Xu H, Mao L, Huang W, Huang Q, Li L, Liu Y. Gene association study between polycystic ovary syndrome and metabolic syndrome: A transcriptomic analysis and machine learning approach. J Ovarian Res. 2025;18(1):220. pmid:41088367 - 124. Mohi Uddin KM, Bhuiyan MdTA, Rahman MdM, Islam MdM, Uddin MA. Early PCOS detection: A comparative analysis of traditional and ensemble machine learning models with advanced feature selection. Eng Rep. 2025;7(2):e70008. - 125. Tong C, Wu Y, Zhuang Z, Yu Y. A diagnostic model for polycystic ovary syndrome based on machine learning. Sci Rep. 2025;15(1):9821. pmid:40119083 - 126. Panjwani B, Yadav J, Mohan V, Agarwal N, Agarwal S. Optimized machine learning for the early detection of polycystic ovary syndrome in women. Sensors (Basel). 2025;25(4):1166. pmid:40006393 - 127. Chen W, Miao J, Chen J, Chen J. Development of machine learning models for diagnostic biomarker identification and immune cell infiltration analysis in PCOS. J Ovarian Res. 2025;18(1):1. pmid:39754246 - 128. Rahman MM, Islam A, Islam F, Zaman M, Islam MR, Alam Sakib MS, et al. Empowering early detection: A web-based machine learning approach for PCOS prediction. Inform Med Unlocked. 2024;47:101500. - 129. Elmannai H, El-Rashidy N, Mashal I, Alohali MA, Farag S, El-Sappagh S, et al. Polycystic ovary syndrome detection machine learning model based on optimized feature selection and explainable artificial intelligence. Diagnostics (Basel). 2023;13(8):1506. pmid:37189606 - 130. Saha A, Roy A, Chakraborty B, Saha B, Chowdhury D. A comparative study to predict polycystic ovarian syndrome (PCOS) based on different models of machine learning technique. AJEC. 2023;4(2):1–6. - 131. Zad Z, Jiang VS, Wolf AT, Wang T, Cheng JJ, Paschalidis IC, et al. Predicting polycystic ovary syndrome with machine learning algorithms from electronic health records. Front Endocrinol (Lausanne). 2024;15:1298628. pmid:38356959 - 132. Jha T, Sirisha M, Bhargavi M. Ovulytics: A machine learning approach for precision diagnosis of PCOD, PCOD and infertility. 2024 First International Conference for Women in Computing (InCoWoCo). IEEE; 2024. p. 1–7. - 133. Suha SA, Islam MN. An extended machine learning technique for polycystic ovary syndrome detection using ovary ultrasound image. Sci Rep. 2022;12(1):17123. pmid:36224353 - 134. Lee S, Arffman RK, Komsi EK, Lindgren O, Kemppainen JA, Metsola H, et al. AI-algorithm training and validation for identification of endometrial CD138+ cells in infertility-associated conditions; polycystic ovary syndrome (PCOS) and recurrent implantation failure (RIF). J Pathol Inform. 2024;15:100380. pmid:38827567 - 135. Abdesselam A, Zidoum H, Zadjali F, Hedjam R, Al-Ansari A, Bayoumi R, et al. Estimate of the HOMA-IR cut-off value for identifying subjects at risk of insulin resistance using a machine learning approach. Sultan Qaboos Univ Med J. 2021;21(4):604–12. pmid:34888081 - 136. Huang X, Yi K, Jia L, Li Y, He H, Ma C, et al. Development and validation of an insulin resistance prediction model in children and adolescents using machine learning algorithms. Transl Pediatr. 2025;14(3):452–62. pmid:40225080 - 137. Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Adv Neural Inf Process Syst. 2017;30. - 138. Mamuye AL, Yilma TM, Abdulwahab A, Broomhead S, Zondo P, Kyeng M, et al. Health information exchange policy and standards for digital health systems in Africa: A systematic review. PLOS Digit Health. 2022;1(10):e0000118. pmid:36812615 - 139. Ndlovu K, Mars M, Scott RE. Validation of an interoperability framework for linking mHealth apps to electronic record systems in Botswana: Expert survey study. JMIR Form Res. 2023;7:e41225. pmid:37129939 - 140. Mehl G, Tunçalp Ö, Ratanaprayul N, Tamrat T, Barreix M, Lowrance D, et al. WHO SMART guidelines: Optimising country-level use of guideline recommendations in the digital age. Lancet Digit Health. 2021;3(4):e213–6. pmid:33610488 - 141. Sahiner B, Chen W, Samala RK, Petrick N. Data drift in medical machine learning: Implications and potential remedies. Br J Radiol. 2023;96(1150):20220878. pmid:36971405 - 142. Fusco R, Granata V, Vallone P, Petrosino T, Iasevoli MD, Raso MM, et al. Engineering the image representation for deep learning in contrast-enhanced mammography: A systematic analysis of preprocessing and anatomical masking. Bioengineering (Basel). 2026;13(3):322. pmid:41899853 - 143. Zhu JY, Park T, Isola P, Efros AA. Unpaired image-to-image translation using cycle-consistent adversarial networks. Proceedings of the IEEE International Conference on Computer Vision; 2017. p. 2223–32. - 144. Wang D, Shelhamer E, Liu S, Olshausen B, Darrell T. Tent: Fully test-time adaptation by entropy minimization. arXiv:200610726 [Preprint]. 2020.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

MeSH descriptors

Artificial Intelligence Artificial Intelligence Reproductive Health Reproductive Health Translational Research, Biomedical Translational Research, Biomedical Women's Health Women's Health Evidence Gaps Evidence Gaps Female Female Humans Humans Machine Learning Machine Learning Pregnancy Pregnancy Resource-Limited Settings Resource-Limited Settings

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

pubmed
last seen: 2026-09-28T06:06:14.772727+00:00
License: public-domain-us · commercial use OK · attribution required
Courtesy of the U.S. National Library of Medicine