An explainable machine learning model for identifying the risk of suboptimal oocyte maturation in patients with polycystic ovary syndrome undergoing GnRH-antagonist protocol.

OA: gold CC-BY-NC-ND-4.0
AI-generated summary by gemini-2.5-flash-lite, 2026-07-29

An XGBoost machine learning model accurately identified suboptimal oocyte maturation risk in PCOS patients, with timing of Gn administration, basal E2, age, AMH, and infertility duration identified as key predictors.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by qwen3.7-flash, 2026-08-18 · read from full text

This retrospective cohort study developed and validated explainable machine learning models to predict suboptimal oocyte maturation in women with polycystic ovary syndrome undergoing GnRH-antagonist controlled ovarian stimulation. The researchers compared predictive features between PCOS patients and a non-PCOS control group, utilizing clinical, laboratory, and ultrasound data to identify condition-specific determinants of oocyte maturity using Shapley Additive Explanations. A key limitation noted was the absence of a universally accepted cutoff for defining optimal versus suboptimal maturation rates, necessitating the selection of a conservative 75% threshold based on existing literature. Relevance to endometriosis: The paper explicitly excludes patients with endometriosis or adenomyosis from its inclusion criteria, focusing solely on PCOS pathophysiology and IVF outcomes.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

BACKGROUND: Polycystic ovary syndrome (PCOS) is a prevalent and heterogeneous endocrine disorder and a leading cause of female infertility. Impaired oocyte maturation in PCOS patients undergoing in vitro fertilization (IVF) adversely affects fertilization, embryo quality, and pregnancy outcomes. Identifying the underlying factors of suboptimal oocyte maturation in patients with PCOS is challenging due to heterogeneous ovarian responses and complex clinical profiles. Machine learning (ML) provides a promising approach to capture complex interactions among clinical parameters, enabling individualized risk assessment of oocyte maturation and the identification of PCOS-specific factors. METHODS: A total of 759 patients with PCOS and 809 non-PCOS controls undergoing the GnRH-antagonist protocol were included and randomly divided into training and validation sets. Feature selection was performed using the least absolute shrinkage and selection operator (LASSO) regression and the Boruta algorithm. Based on 13 selected features, ten machine learning models were developed. Model performance was evaluated using multiple metrics, including the area under the receiver operating characteristic curve (AUC), accuracy, precision, sensitivity, specificity, negative predictive value (NPV), and F1 score. SHapley Additive exPlanations (SHAP) were used to interpret feature importance and to distinguish PCOS-specific factors. RESULTS: Among ten ML models, the XGBoost model showed the best performance (AUC = 0.868). SHAP analysis identified initial timing of gonadotropin (Gn) administration, basal estradiol (E2) levels, age, anti-Müllerian hormone (AMH), and infertility duration as key contributors to oocyte maturation rate prediction in PCOS patients. Individual-level SHAP analysis further revealed substantial inter-patient heterogeneity in feature contributions, enabling personalized interpretation. Comparative modeling and SHAP interpretation in a matched non-PCOS control cohort identified distinct PCOS-specific predictors, underscoring the differential clinical relevance of key variables in PCOS versus non-PCOS populations. CONCLUSION: Our ML model effectively identified the risk of suboptimal oocyte maturation in PCOS patients undergoing GnRH-antagonist protocol. The identified key features provide clinically interpretable insights that may inform individualized ovarian stimulation strategies.
Full text 52,170 characters · extracted from pmc-nxml · 6 sections · click to expand

Results

Between December 2014 and December 2024, a total of 759 patients diagnosed with PCOS were included in this study. Among them, 606 patients achieved an oocyte maturation rate ≥ 75%, while 153 patients had a maturation rate < 75%. Baseline characteristics of the PCOS cohort are summarized in Table  1 . Significant differences ( P  < 0.05) were observed for age, type of infertility, initial timing of Gn, and serum LH levels. Notably, patients with higher maturation rates were older, higher LH levels, and exhibited a greater proportion of secondary infertility. The cohort was randomly divided into training and validation sets for model development and validation. There were no statistically significant differences in demographic data, clinical characteristics, or maturation rates between the two sets, indicating good comparability (Table  2 ). Table 1 Demographic and clinical characteristics of high maturation rate group and low maturation rate group in PCOS cohorts Characteristics High maturation rate ( n  = 606) Low maturation rate ( n  = 153) P value Clinical characteristics  Age (years) 30 (27–32) 28 (27–30)  < 0.001  BMI (kg/m 2 ) 22.22 (20.24–24.73) 22.19 (20.13–24.71) 0.823  Infertility duration (years) 3 (2–4) 3 (2–5) 0.052  Type of infertility, n (%) 0.011 Primary infertility 390 (64.4) 115 (75.2) Secondary infertility 216 (35.6) 38 (24.8) Cycle-specific information  Initial dose of Gn (IU) 150 (150–200) 150 (150–175) 0.264  Initial timing of Gn, n (%)  < 0.001   2 46 (7.6) 21 (13.7)   3 276 (45.5) 91 (59.5)   4 274 (45.2) 39 (25.5)   5 10 (1.7) 2 (1.3) Trigger protocol 0.118  Single trigger 350 (57.8) 99 (64.7)  Dual trigger 256 (42.2) 54 (35.3) LH supplementation 0.361  Yes 202(33.3) 57(37.3)  No 404(66.7) 96(62.7) Laboratory parameters  AMH (ng/mL) 9.03 (5.83–12.53) 9.86 (6.96–11.81) 0.287  Basal FSH (IU/L) 4.98 (4.34–5.81) 4.92 (4.21–5.61) 0.266  Basal LH (IU/L) 5.38 (3.54–8.37) 4.54 (3.29–6.89) 0.016  Basal E 2 (pg/mL) 32.00 (24.00–42.00) 32.00 (25.88–42.75) 0.302  Basal T (ng/mL) 0.42 (0.33–0.53) 0.45 (0.36–0.53) 0.090  PRL (ng/mL) 13.58 (10.17–19.41) 13.20 (10.44–17.83) 0.515  FSH on Gn initiation day (IU/L) 5.00 (4.24–5.79) 4.76 (4.20–5.75) 0.368  LH on Gn initiation day (IU/L) 5.00 (3.47–7.32) 4.83 (3.30–7.05) 0.329  E 2 on Gn initiation day (pg/mL) 30.00 (24.00–40.00) 30.00 (21.00–38.50) 0.556 Ultrasound index  AFC 20 (16–25) 21 (17–25) 0.329 Continuous values are presented as medians (IQR), and categorical values were reported as numbers (%) BMI body mass index, Gn gonadotropin, AMH anti-Müllerian hormone, FSH follicle-stimulating hormone, LH luteinizing hormone, E 2 Estradiol, T testosterone, PRL prolactin, AFC antral follicle count Table 2 Demographic and clinical characteristics of training and validation cohorts in PCOS group Characteristics Training set ( n  = 607) Validation set ( n  = 152) P value Clinical characteristics  Age (years) 29 (27–32) 30 (28–32) 0.037  BMI (kg/m2) 22.22 (20.20–24.80) 22.22 (20.23–24.32) 0.552  Infertility duration (years) 3 (2–5) 3 (2–4) 0.083 Type of infertility, n (%) 0.359  Primary infertility 415 (68.4) 98 (64.5)  Secondary infertility 192 (31.6) 54 (35.5) Cycle-specific information  Initial dose of Gn (IU) 150 (150–187.5) 150 (150–200) 0.741  Initial timing of Gn, n (%) 0.722   2 52 (8.6) 15 (9.9)   3 295 (48.6) 72 (47.4)   4 249 (41.0) 64 (42.1)   5 11 (1.8) 1 (0.7) Trigger Protocol, n (%) 0.723  Single trigger 361 (59.5) 88 (57.9)  Dual trigger 246 (40.5) 64 (42.1) Laboratory parameters  AMH (ng/mL) 9.25 (6.06–12.41) 9.31 (5.94–12.4) 0.848  Basal FSH (IU/L) 4.97 (4.27–5.70) 4.98 (4.38–5.87) 0.443  Basal LH (IU/L) 5.14 (3.50–8.03) 5.59 (3.59–8.10) 0.749  Basal E 2 (pg/mL) 32.00 (25.00–42.00) 33.00 (25.25–40.75) 0.991  Basal T (ng/mL) 0.42 (0.34–0.53) 0.42 (0.32–0.52) 0.500  PRL (ng/mL) 13.74 (10.46–19.10) 12.42 (8.78–18.86) 0.011  FSH on Gn initiation day (IU/L) 4.93 (4.25–5.76) 4.91 (3.91–5.87) 0.676  LH on Gn initiation day (IU/L) 4.90 (3.48–7.21) 5.09 (3.35–7.34) 0.536  E 2 on Gn initiation day (pg/mL) 30.00 (23.00–39.00) 31.00 (25.00–41.00) 0.179 Ultrasound index  AFC 20 (16–25) 20 (15–25) 0.255 Outcome  Maturation rate, n (%) 0.133 High 478 (78.7) 128 (84.2) Low 129 (21.3) 24 (15.8) Continuous values are presented as medians (IQR), and categorical values were reported as numbers (%) BMI body mass index, Gn gonadotropin, AMH anti-Müllerian hormone, FSH follicle-stimulating hormone, LH luteinizing hormone, E 2 Estradiol, T testosterone, PRL prolactin, AFC antral follicle count Demographic and clinical characteristics of high maturation rate group and low maturation rate group in PCOS cohorts Continuous values are presented as medians (IQR), and categorical values were reported as numbers (%) BMI body mass index, Gn gonadotropin, AMH anti-Müllerian hormone, FSH follicle-stimulating hormone, LH luteinizing hormone, E 2 Estradiol, T testosterone, PRL prolactin, AFC antral follicle count Demographic and clinical characteristics of training and validation cohorts in PCOS group Continuous values are presented as medians (IQR), and categorical values were reported as numbers (%) BMI body mass index, Gn gonadotropin, AMH anti-Müllerian hormone, FSH follicle-stimulating hormone, LH luteinizing hormone, E 2 Estradiol, T testosterone, PRL prolactin, AFC antral follicle count Feature selection was performed using both LASSO regression and the Boruta algorithm (Fig.  2 ). Figure  2 A–C illustrate the results of LASSO regression with five-fold cross-validation, which identified 11 features based on the optimal lambda value (lambda.min). The Boruta algorithm identified 5 key features, ranked in descending order of importance (Fig.  2 D). The final feature sets of two methods are presented in Fig.  2 E. To improve model robustness and capture a broader spectrum of predictive information, we combined the two sets and used all 13 features to develop the machine learning models. Spearman correlation analysis revealed generally low pairwise correlations among the selected features, indicating no substantial multicollinearity (all |r|< 0.5; Fig.  2 F). Fig. 2 Feature selection by LASSO regression and Boruta algorithm. A The Lasso regression coefficient path diagram. B LASSO regression with tenfold cross-validation. Misclassification errors was plotted against log(λ). The vertical dotted lines correspond to the optimal λ with the minimum cross-validated error (lambda.min) and one standard error of the minimum (lambda.1se). C Features with non-zero coefficients selected by LASSO regression. D Features selected by Boruta algorithm. Only features classified as important were retained. E Venn diagram of selected features from LASSO regression and the Boruta algorithm. The union of two feature sets yield 11 predictors. F Heatmap of Spearman correlation analysis between selected features. E2, estradiol; T, testosterone; FSH, follicle stimulating hormone; Gn, gonadotropin; PRL, prolactin; LH, luteinizing hormone; AMH, anti-Müllerian hormone; LASSO, least absolute shrinkage and selection operator Feature selection by LASSO regression and Boruta algorithm. A The Lasso regression coefficient path diagram. B LASSO regression with tenfold cross-validation. Misclassification errors was plotted against log(λ). The vertical dotted lines correspond to the optimal λ with the minimum cross-validated error (lambda.min) and one standard error of the minimum (lambda.1se). C Features with non-zero coefficients selected by LASSO regression. D Features selected by Boruta algorithm. Only features classified as important were retained. E Venn diagram of selected features from LASSO regression and the Boruta algorithm. The union of two feature sets yield 11 predictors. F Heatmap of Spearman correlation analysis between selected features. E2, estradiol; T, testosterone; FSH, follicle stimulating hormone; Gn, gonadotropin; PRL, prolactin; LH, luteinizing hormone; AMH, anti-Müllerian hormone; LASSO, least absolute shrinkage and selection operator Ten machine learning algorithms were applied to construct predictive models based on the selected features. Hyperparameters were optimized using the Optuna framework with five-fold cross-validation to prevent overfitting and improve generalizability. To ensure robust model selection, we evaluated performance metrics on both training and validation sets. While multiple algorithms achieved near-perfect scores on the training set (Table  3 ), the XGBoost model was selected as the optimal model due to its superior generalizability, yielding an AUC of 0.868 in the validation set (95% CI: 0.762–0.958) (Fig.  3 ). This consistent performance on unseen data confirms that the XGBoost model effectively captured underlying clinical patterns rather than overfitting to training noise. Additional performance metrics further confirmed the model’s superiority: accuracy 0.914, precision 0.762, sensitivity 0.667, specificity 0.961, NPV 0.939, and F1 score 0.711 (Table  3 ). Given its consistent outperformance, downstream analysis was based on the XGBoost model. Table 3 Predictive performance of ten machine learning models in PCOS cohort Model Dataset AUC Accuracy Precision Sensitivity Specificity NPV F1 Score AdaBoost Training 0.919 0.839 0.819 0.870 0.808 0.862 0.844 Validation 0.761 0.789 0.389 0.583 0.828 0.914 0.467 ANN Training 0.951 0.857 0.927 0.774 0.939 0.806 0.844 Validation 0.783 0.862 0.565 0.542 0.922 0.915 0.553 Decision Tree Training 0.792 0.742 0.735 0.755 0.728 0.748 0.745 Validation 0.707 0.770 0.366 0.625 0.797 0.919 0.462 Extra Trees Training 1.000 1.000 1.000 1.000 1.000 1.000 1.000 Validation 0.805 0.882 0.636 0.583 0.938 0.923 0.609 Gradient Boosting Training 1.000 1.000 1.000 1.000 1.000 1.000 1.000 Validation 0.854 0.895 0.682 0.625 0.945 0.931 0.652 KNN Training 0.963 0.879 0.830 0.952 0.805 0.944 0.887 Validation 0.743 0.658 0.274 0.708 0.648 0.922 0.395 Logistic Regression Training 0.813 0.743 0.732 0.766 0.720 0.754 0.748 Validation 0.715 0.770 0.378 0.708 0.781 0.935 0.493 Random Forest Training 1.000 1.000 1.000 1.000 1.000 1.000 1.000 Validation 0.807 0.901 0.737 0.583 0.961 0.925 0.651 SVM Training 0.812 0.749 0.732 0.785 0.713 0.768 0.758 Validation 0.723 0.737 0.340 0.708 0.742 0.931 0.459 XGBoost Training 1.000 1.000 1.000 1.000 1.000 1.000 1.000 Validation 0.868 0.914 0.762 0.667 0.961 0.939 0.711 Adaboost adaptive boosting, ANN artificial neural network, Extra Trees extremely randomized trees, GBDT gradient boosting decision trees, KNN K-nearest neighbors, SVM support vector machine, XGBoost eXtreme gradient boosting, AUC area under the receiver-operating-characteristic curve, NPV negative predictive value Fig. 3 Receiver operating characteristic (ROC) curves of the machine learning models in PCOS group. A ROC curve in the training set. B ROC curve in the validation set. PCOS, polycystic ovary syndrome; Adaboost, adaptive boosting; ANN, artificial neural network; Extra Trees, extremely randomized trees; Gradient Boosting, gradient boosting decision trees; KNN, K-nearest neighbors; SVM, support vector machine; XGBoost, eXtreme gradient boosting; AUC, area under the receiver-operating-characteristic curve Predictive performance of ten machine learning models in PCOS cohort Adaboost adaptive boosting, ANN artificial neural network, Extra Trees extremely randomized trees, GBDT gradient boosting decision trees, KNN K-nearest neighbors, SVM support vector machine, XGBoost eXtreme gradient boosting, AUC area under the receiver-operating-characteristic curve, NPV negative predictive value Receiver operating characteristic (ROC) curves of the machine learning models in PCOS group. A ROC curve in the training set. B ROC curve in the validation set. PCOS, polycystic ovary syndrome; Adaboost, adaptive boosting; ANN, artificial neural network; Extra Trees, extremely randomized trees; Gradient Boosting, gradient boosting decision trees; KNN, K-nearest neighbors; SVM, support vector machine; XGBoost, eXtreme gradient boosting; AUC, area under the receiver-operating-characteristic curve To enhance clinical interpretability, SHAP was used to quantify the contribution of each feature to the model’s predictions. SHAP provides both global and local explanations: the former reveals feature importance across the entire population, whereas the latter offers insights into individual-level predictions. As shown in Fig.  4 A, features were ranked by their mean absolute SHAP values in descending order. The top five influential features were the initial timing of Gn, basal serum E 2 levels, age, AMH, and infertility duration. SHAP summary plots (Fig.  4 B) illustrated how each feature influenced the model’s output. In the plot, the y-axis lists the features, the x-axis reflects their impact on outcome, and the color represents the actual feature value. For example, delayed initiation of Gn was associated with higher maturation rates, while longer infertility duration decreased the likelihood of achieving a high maturation rate. SHAP dependence plots (Figure S1A) further elucidated the relationship between individual feature values and their corresponding SHAP values. SHAP values greater than zero corresponded to predictions of the low maturation rate. For instance, initiating ovarian stimulation on or before menstrual cycle day 3, younger age (≤ 30 years), intermediate AMH levels (7.5–13.0 ng/mL), and either very low (≤ 23 pg/mL) or high (≥ 48 pg/mL) serum E 2 levels on the day of Gn initiation were associated with positive SHAP values, suggesting a higher probability of being classified into the low maturation rate group. Fig. 4 Global model explanation in PCOS group. A SHAP summary bar plot shows feature importance for each predictor of the XGBoost model in descending order. B SHAP summary dot plot. Each dot represents a SHAP value for a given feature in a patient. The x axis indicates the SHAP value, reflecting the feature’s impact on the model output. The color gradient from red to blue corresponds to high to low feature values, respectively. SHAP, SHapley Additive exPlanations; PCOS, polycystic ovary syndrome; Gn, gonadotropin; E2, estradiol; AMH, anti-Müllerian hormone; LH, luteinizing hormone; T, testosterone; PRL, prolactin; FSH, follicle stimulating hormone Global model explanation in PCOS group. A SHAP summary bar plot shows feature importance for each predictor of the XGBoost model in descending order. B SHAP summary dot plot. Each dot represents a SHAP value for a given feature in a patient. The x axis indicates the SHAP value, reflecting the feature’s impact on the model output. The color gradient from red to blue corresponds to high to low feature values, respectively. SHAP, SHapley Additive exPlanations; PCOS, polycystic ovary syndrome; Gn, gonadotropin; E2, estradiol; AMH, anti-Müllerian hormone; LH, luteinizing hormone; T, testosterone; PRL, prolactin; FSH, follicle stimulating hormone SHAP force plots provide insights into how individual predictions are generated by incorporating personalized input data (Fig.  5 ). The base value represents the average model output, while each feature contributes either positively (red) or negatively (blue) to the final prediction. The length of each color bar reflects the magnitude of its contribution. Figure  5 A shows a true positive case, where the model predicted a low maturation rate with 97.8% probability. Figure  5 B shows a true negative case, with a high maturation rate predicted at 89.6%. In patient A (Fig.  5 A), shorter infertility duration, elevated serum T levels, and younger age shifted the prediction toward a low maturation rate. Conversely, the use of a dual trigger protocol, higher PRL levels, and basal FSH levels contributed in the opposite direction, favoring a high maturation rate. The final prediction reflects the net effect of these opposing influences. SHAP force plots thus provide an intuitive and interpretable visualization of individualized predictions. Fig. 5 Local model explanation in PCOS group. SHAP force plots illustrate how individual features contributed to the model’s prediction for A patient A (true positive) and B Patient B (true negative). The color represents the contributions of each feature, with red being positive (pushing the prediction towards the “low maturation rate” class) and blue being negative (pushing the prediction towards the “high maturation rate” class). The length of the color bar represents the contribution strength. f(x) denotes the model output in log-odds. For instance, f(x) = 3.76 corresponds to a predicted probability of 97.8%. SHAP, SHapley Additive exPlanations; PCOS, polycystic ovary syndrome; T, testosterone; Gn, gonadotropin; LH, luteinizing hormone; AMH, anti-Müllerian hormone; E2, estradiol; PRL, prolactin; FSH, follicle stimulating hormone Local model explanation in PCOS group. SHAP force plots illustrate how individual features contributed to the model’s prediction for A patient A (true positive) and B Patient B (true negative). The color represents the contributions of each feature, with red being positive (pushing the prediction towards the “low maturation rate” class) and blue being negative (pushing the prediction towards the “high maturation rate” class). The length of the color bar represents the contribution strength. f(x) denotes the model output in log-odds. For instance, f(x) = 3.76 corresponds to a predicted probability of 97.8%. SHAP, SHapley Additive exPlanations; PCOS, polycystic ovary syndrome; T, testosterone; Gn, gonadotropin; LH, luteinizing hormone; AMH, anti-Müllerian hormone; E2, estradiol; PRL, prolactin; FSH, follicle stimulating hormone To evaluate the specificity of the identified predictive patterns to PCOS, a control group of 809 non-PCOS patients was established at the same center during the same period, based on the inclusion criteria. Baseline characteristics are detailed in Table  4 . Significant differences ( P  < 0.05) were observed for type of infertility, serum PRL levels, and basal E 2 levels. Similar to the PCOS cohort, the control group was split into training and validation sets with no significant differences (Table  5 ). Table 4 Demographic and clinical characteristics of high maturation rate group and low maturation rate group in control cohort Characteristics High maturation rate ( n  = 676) Low maturation rate ( n  = 133) P value Clinical characteristics  Age (years) 31 (28–35) 31 (28–35) 0.971  Infertility duration (years) 3 (1–5) 3 (1–4) 0.154 Type of infertility, n (%)  < 0.001  Primary infertility 222 (32.8) 66 (49.6)  Secondary infertility 454 (67.2) 67 (50.4) Cycle-specific information  Initial dose of Gn (IU) 225 (150–300) 225 (150–300) 0.472  Initial timing of Gn, n (%) 0.262   2 27 (4.0) 5 (3.8)   3 191 (28.3) 50 (37.6)   4 414 (61.2) 72 (54.1)   5 41 (6.1) 6 (4.5)   6 3 (0.4) 0 (0.0) Trigger Protocol 0.978  Single trigger 589 (87.1) 116 (87.2)  Dual trigger 87 (12.9) 17 (12.8) LH supplementation 0.751  Yes 371(54.9) 71(53.4)  No 305(45.1) 62(46.6)  Basal FSH (IU/L) 5.70 (4.73–6.72) 5.44 (4.66–6.50) 0.154  Basal LH (IU/L) 3.18 (2.36–4.28) 3.49 (2.47–4.53) 0.142  Basal E 2 (pg/mL) 30.00 (23.00–40.00) 28.00 (22.00–38.00) 0.233  Basal T (ng/mL) 0.30 (0.24–0.38) 0.32 (0.26–0.38) 0.207  PRL (ng/mL) 15.48 (11.10–20.83) 17.28 (12.57–23.89) 0.027  E 2 on Gn initiation day (pg/mL) 28.00 (22.00–37.00) 26.00 (19.00–34.75) 0.023 Continuous values are presented as medians (IQR), and categorical values were reported as numbers (%). Abbreviations: Gn, gonadotropin; AMH, anti-Müllerian hormone; FSH, follicle-stimulating hormone; LH, luteinizing hormone; E 2 , Estradiol; T, testosterone; PRL, prolactin Table 5 Demographic and clinical characteristics of training and validation cohorts in control group Characteristics Training set ( n  = 647) Validation set ( n  = 162) P value Clinical characteristics  Age (years) 31 (28–35) 31 (28–35) 0.556  Infertility duration (years) 3 (1–5) 3 (1–4) 0.946 Type of infertility, n (%) 0.391  Primary infertility 235 (36.3) 53 (32.7)  Secondary infertility 412 (63.7) 109 (67.3) Cycle-specific information  Initial dose of Gn (IU) 225 (150–300) 225 (150–300) 0.214  Initial timing of Gn, n (%) 0.440   2 23 (3.6) 9 (5.6)   3 194 (30.0) 47 (29.0)   4 394 (60.9) 92 (56.8)   5 34 (5.3) 13 (8.0)   6 2 (0.3) 1 (0.6) Trigger Protocol 0.964  Single trigger 564 (87.2) 141 (87.0)  Dual trigger 83 (12.8) 21 (13.0) Laboratory parameters  AMH (ng/mL) 3.98 (2.04–6.31) 4.57 (1.99–6.89) 0.212  Basal FSH (IU/L) 5.70 (4.72–6.74) 5.41 (4.68–6.40) 0.093  Basal LH (IU/L) 3.27 (2.38–4.33) 3.22 (2.40–4.28) 0.922  Basal E 2 (pg/mL) 30.00 (23.00–40.00) 31.00 (22.00–37.00) 0.826  Basal T (ng/mL) 0.30 (0.24–0.38) 0.30 (0.24–0.37) 0.598  PRL (ng/mL) 15.76 (11.27–21.18) 14.81 (11.13–23.07) 0.884  E 2 on Gn initiation day (pg/mL) 28.00 (22.00–37.00) 28.00 (21.00–36.00) 0.632 Outcome  Maturation rate, n (%) 0.881 High 540 (83.5) 136 (84.0) Low 107 (16.5) 26 (16.0) Continuous values are presented as medians (IQR), and categorical values were reported as numbers (%) Gn gonadotropin, AMH anti-Müllerian hormone, FSH follicle-stimulating hormone, LH luteinizing hormone, E 2 Estradiol, T testosterone, PRL prolactin Demographic and clinical characteristics of high maturation rate group and low maturation rate group in control cohort Continuous values are presented as medians (IQR), and categorical values were reported as numbers (%). Abbreviations: Gn, gonadotropin; AMH, anti-Müllerian hormone; FSH, follicle-stimulating hormone; LH, luteinizing hormone; E 2 , Estradiol; T, testosterone; PRL, prolactin Demographic and clinical characteristics of training and validation cohorts in control group Continuous values are presented as medians (IQR), and categorical values were reported as numbers (%) Gn gonadotropin, AMH anti-Müllerian hormone, FSH follicle-stimulating hormone, LH luteinizing hormone, E 2 Estradiol, T testosterone, PRL prolactin Applying the same robust model selection criteria, the XGBoost model was identified as the optimal model for the control cohort, achieving an AUC of 0.847 (95% CI: 0.746–0.935), with accuracy 0.889, precision 0.700, sensitivity 0.538, specificity 0.956, NPV 0.915, and F1 score 0.609 in the validation set. (Fig.  6 , Table  6 ). SHAP analysis in the control group (Fig.  7 A) identified the top five contributors: type of infertility, infertility duration, AMH, serum E 2 levels on the day of Gn initiation, and serum PRL levels. SHAP summary and dependence plots provided further insights into their predictive roles, enabling comparison with the PCOS cohort. (Fig.  7 B , Figure S1B). Fig. 6 Receiver operating characteristic (ROC) curves of the machine learning models in control group. A ROC curve in the training set. B ROC curve in the validation set. Adaboost, adaptive boosting; ANN, artificial neural network; Extra Trees, extremely randomized trees; Gradient Boosting, gradient boosting decision trees; KNN, K-nearest neighbors; SVM, support vector machine; XGBoost, eXtreme gradient boosting; AUC, area under the receiver-operating-characteristic curve Table 6 Predictive performance of ten machine learning models in control cohort Model Dataset AUC Accuracy Precision Sensitivity Specificity NPV F1 Score AdaBoost Training 0.976 0.910 0.916 0.904 0.917 0.905 0.910 Validation 0.768 0.833 0.484 0.577 0.882 0.916 0.526 ANN Training 0.967 0.905 0.900 0.909 0.900 0.908 0.905 Validation 0.828 0.864 0.571 0.615 0.912 0.925 0.593 Decision Tree Training 1.000 1.000 1.000 1.000 1.000 1.000 1.000 Validation 0.707 0.821 0.452 0.538 0.875 0.908 0.491 Extra Trees Training 1.000 1.000 1.000 1.000 1.000 1.000 1.000 Validation 0.750 0.870 0.632 0.462 0.949 0.902 0.533 Gradient Boosting Training 1.000 1.000 1.000 1.000 1.000 1.000 1.000 Validation 0.834 0.877 0.688 0.423 0.963 0.897 0.524 KNN Training 0.970 0.847 0.774 0.981 0.713 0.975 0.865 Validation 0.729 0.642 0.271 0.731 0.625 0.924 0.396 Logistic Regression Training 0.733 0.688 0.682 0.706 0.670 0.695 0.693 Validation 0.675 0.667 0.259 0.577 0.684 0.894 0.357 Random Forest Training 1.000 1.000 1.000 1.000 1.000 1.000 1.000 Validation 0.778 0.877 0.667 0.462 0.956 0.903 0.545 SVM Training 1.000 1.000 0.998 1.000 0.998 1.000 0.999 Validation 0.729 0.846 0.533 0.308 0.949 0.878 0.390 XGBoost Training 0.999 0.982 0.992 0.972 0.993 0.973 0.982 Validation 0.847 0.889 0.700 0.538 0.956 0.915 0.609 Adaboost adaptive boosting, ANN artificial neural network, Extra Trees extremely randomized trees, GBDT gradient boosting decision trees, KNN K-nearest neighbors, SVM support vector machine, XGBoost eXtreme gradient boosting, AUC area under the receiver-operating-characteristic curve, NPV negative predictive value Fig. 7 Model explanation by the SHAP method in control group. A SHAP summary bar plot shows feature importance for each predictor of the XGBoost model in descending order. B SHAP summary dot plot. Each dot represents a SHAP value for a given feature in a patient. The x axis indicates the SHAP value, reflecting the feature’s impact on the model output. The color gradient from red to blue corresponds to high to low feature values, respectively. SHAP, SHapley Additive exPlanations; Gn, gonadotropin; E2, Estradiol; AMH, anti-Müllerian hormone; LH, luteinizing hormone; T, testosterone; PRL, prolactin; FSH, follicle stimulating hormone Receiver operating characteristic (ROC) curves of the machine learning models in control group. A ROC curve in the training set. B ROC curve in the validation set. Adaboost, adaptive boosting; ANN, artificial neural network; Extra Trees, extremely randomized trees; Gradient Boosting, gradient boosting decision trees; KNN, K-nearest neighbors; SVM, support vector machine; XGBoost, eXtreme gradient boosting; AUC, area under the receiver-operating-characteristic curve Predictive performance of ten machine learning models in control cohort Adaboost adaptive boosting, ANN artificial neural network, Extra Trees extremely randomized trees, GBDT gradient boosting decision trees, KNN K-nearest neighbors, SVM support vector machine, XGBoost eXtreme gradient boosting, AUC area under the receiver-operating-characteristic curve, NPV negative predictive value Model explanation by the SHAP method in control group. A SHAP summary bar plot shows feature importance for each predictor of the XGBoost model in descending order. B SHAP summary dot plot. Each dot represents a SHAP value for a given feature in a patient. The x axis indicates the SHAP value, reflecting the feature’s impact on the model output. The color gradient from red to blue corresponds to high to low feature values, respectively. SHAP, SHapley Additive exPlanations; Gn, gonadotropin; E2, Estradiol; AMH, anti-Müllerian hormone; LH, luteinizing hormone; T, testosterone; PRL, prolactin; FSH, follicle stimulating hormone

Materials

This retrospective cohort study included patients who met the eligibility criteria at the Reproductive Medicine Center of the First Affiliated Hospital of Sun Yat-sen University (Guangzhou, China) between December 2014 and December 2024. This study was approved by the institutional ethics committee (ethical approval no. [2025]421). Informed consent was waived due to the retrospective nature of the study. All procedures were conducted in accordance with relevant clinical guidelines and institutional regulations. Inclusion criteria for the PCOS group were: (1) diagnosis of PCOS based on the Rotterdam criteria [ 21 ]; (2) undergoing a first intracytoplasmic sperm injection (ICSI) cycle with complete clinical and laboratory records; and (3) use of a GnRH antagonist protocol. Inclusion criteria for the control group were: (1) not meeting the Rotterdam criteria; (2) undergoing a first ICSI cycle with clinical and laboratory complete records; and (3) use of a GnRH antagonist protocol. Exclusion criteria included: (1) ovarian abnormalities or a history of ovarian surgery; (2) malignancies; (3) endometriosis or adenomyosis; (4) chromosomal abnormalities or polymorphisms; (5) metabolic disorders including diabetes, hyperprolactinemia, hyperthyroidism, hypothyroidism, and autoimmune diseases; and (6) cycles with > 20% missing data or incomplete outcome records. Basal sex hormones were measured on days 2–4 of the menstrual cycle. Anti-Mullerian hormone (AMH) and body mass index (BMI) were recorded prior to COS. All assays of serum sex hormones and AMH levels were performed in the specialized endocrinology laboratory within the medicine center. Follicular development was consistently monitored by transvaginal ultrasound, performed by a relatively fixed team of experienced reproductive physicians. Antral follicle count (AFC) was defined as the total number of follicles measuring 2–9 mm in both ovaries. All patients underwent COS according to standard GnRH antagonist protocol, as described previously [ 22 , 23 ]. Recombinant follicle-stimulating hormone (FSH; Gonal-F, Merck-Serono) or a combination of recombinant FSH and human menopausal gonadotropins (hMG; Menopur, Ferring) was administered for ovarian stimulation. Gonadotropin (Gn) dosage was individualized according to age, body weight, and ovarian reserve, and further adjusted based on serum estradiol (E 2 ) levels and follicular diameter. Daily administration of 0.25 mg GnRH antagonist (Ganirelix, Organon; or Cetrotide, Merck Serono) was initiated on stimulation day 6–7, upon detection of a dominant follicle ≥ 12 mm. Ovulation was triggered with either 250 μg of recombinant human chorionic gonadotropin (hCG; Ovidrel, Merck-Serono), 10,000 IU hCG or a dual trigger with 0.2 mg triptorelin acetate (Decapeptyl, Ferring) plus 2,000 IU hCG when two follicles reached 18 mm or three follicles reached 17 mm in diameter. Transvaginal ultrasound-guided oocyte retrieval was performed 35–38 h after trigger by experienced and well-trained physicians. Cumulus–oocyte complexes (COCs) were collected and incubated in IVF fertilization medium (VitroLife, Gothenburg) at 37 °C in a humidified atmosphere containing 6% CO₂ for 2–4 h. Cumulus and corona radiata cells were subsequently removed using 80 U/mL hyaluronidase (VitroLife, Gothenburg) and gentle mechanical pipetting, after which the maturity of oocytes was carefully assessed for subsequent use [ 22 , 23 ]. Clinical characteristics, cycle-specific parameters, laboratory and ultrasound results from eligible patients were collected. To minimize potential bias introduced by missing data, variables and cycles with > 20% missing data were excluded. Remaining missing values were imputed using multiple imputation by chained equations (MICE) [ 24 ], to preserve data integrity for model development. Finally, clinical characteristics including age, BMI, infertility duration, type of infertility, cycle-specific information including initial dose of Gn, initial timing of Gn, trigger protocol, laboratory results including AMH, basal FSH, LH, E 2 and testosterone (T), Prolactin (PRL), FSH, luteinizing hormone (LH) and E 2 on Gn initiation day, ultrasound parameters including AFC were utilized for further analysis. Among these predictors, Type of Infertility, Initial timing of Gn, and Trigger Protocol were processed as categorical variables, while the remaining features were treated as continuous variables. Mature oocytes were defined as those at the metaphase II (MII) stage with a visible first polar body, in accordance with the International Glossary on Infertility and Fertility Care [ 25 ]. The oocyte maturation rate was calculated as the number of mature (MII) oocytes divided by the total number of oocytes denudated for micromanipulation. The primary outcome variable for analysis was the categorization of oocyte maturation competence. Consequently, Patients were stratified into high (≥ 75%) and low (< 75%) maturation groups based on a 75% cut-off, as described in the Introduction. Normality of continuous variables was assessed using the Kolmogorov–Smirnov test. Data with normal distributions were analyzed using independent-samples t-tests and presented as means ± standard deviations (SD). Non-normally distributed data were evaluated using the Mann–Whitney U test and presented as medians and interquartile ranges (IQR). Categorical variables were compared using the Chi-squared test and reported as frequencies and percentages. All analyses were performed using IBM SPSS Statistics 29.0, with a significance threshold of P  < 0.05 (two-sided). The dataset was randomly split into training (80%) and validation (20%) sets. Feature selection was performed exclusively on the training set using a combination of least absolute shrinkage and selection operator (LASSO) regression and the Boruta algorithm to enhance model interpretability and predictive accuracy. Selected features were used to develop 10 machine learning (ML) models: adaptive boosting (Adaboost), artificial neural network (ANN), decision tree (DT), extremely randomized trees (Extra Trees), gradient boosting decision trees (GBDT), K-nearest neighbors (KNN), logistic regression (LR), random forest (RF), support vector machine (SVM), and eXtreme gradient boosting (XGBoost). Hyperparameters were optimized using the Optuna framework. Five-fold cross-validation was implemented exclusively on the training set to prevent overfitting and ensure model generalizability, while the validation set was reserved for final model evaluation. Model performance was evaluated using several metrics, including the area under the receiver operating characteristic curve (AUC), accuracy, precision, sensitivity, specificity, negative predictive value (NPV), and F1 score. Model interpretation was performed using SHapley Additive exPlanations (SHAP, version 0.40.0), which quantified both global and individual-level contributions of each feature. Mean absolute SHAP values were calculated to identify the most influential predictors across the dataset. An overview of the model development and analysis workflow is presented in Fig.  1 . All data preprocessing, modeling, and statistical analyses were conducted using Python 3.11 and R 4.3.2. Fig. 1 Flow chart of the study design. PCOS, polycystic ovary syndrome; IVF, in vitro fertilization; ICSI, intracytoplasmic sperm injection; LASSO, least absolute shrinkage and selection operator; Adaboost, adaptive boosting; ANN, artificial neural network; Extra Trees, extremely randomized trees; GBDT, gradient boosting decision trees; KNN, K-nearest neighbors; SVM, support vector machine; XGBoost, eXtreme gradient boosting; AUC, area under the receiver-operating-characteristic curve; NPV, negative predictive value Flow chart of the study design. PCOS, polycystic ovary syndrome; IVF, in vitro fertilization; ICSI, intracytoplasmic sperm injection; LASSO, least absolute shrinkage and selection operator; Adaboost, adaptive boosting; ANN, artificial neural network; Extra Trees, extremely randomized trees; GBDT, gradient boosting decision trees; KNN, K-nearest neighbors; SVM, support vector machine; XGBoost, eXtreme gradient boosting; AUC, area under the receiver-operating-characteristic curve; NPV, negative predictive value

Conclusion

In conclusion, we developed and validated a machine learning-based model that identifies the risk of suboptimal oocyte maturation in PCOS patients undergoing the GnRH-antagonist protocol with high accuracy. This model holds substantial promise as a clinical tool for identifying individuals at risk of suboptimal oocyte maturation, facilitating timely monitoring during IVF cycles, guiding personalized preventive and therapeutic strategies, and ultimately improving reproductive outcomes.

Discussion

To our knowledge, this is the first study to develop ML models specifically aimed at assessing the risk of suboptimal oocyte maturation in patients with PCOS. To better identify PCOS-specific influencing factors, we additionally constructed and analyzed prediction models using data from non-PCOS patients treated at the same medical center and during the same period. Ten ML algorithms were evaluated and compared, among which the XGBoost model exhibited the best performance in both PCOS and non-PCOS cohorts. Furthermore, SHAP analysis was applied to interpret the contribution of individual predictors, offering both a foundation for model refinement and valuable insights to support personalized treatment planning and optimization of COS protocols. Compared with traditional statistical approaches, ML offers significant advantages in predicting clinical outcomes. It can capture complex, non-linear relationships within high-dimensional datasets and identifying subtle patterns and interactions that might be overlooked using conventional methods [ 17 ]. In our study, the XGBoost model demonstrated superior performance in both PCOS and non-PCOS groups, achieving AUC values of 0.868 and 0.847, respectively. XGBoost is a robust ensemble learning algorithm that sequentially builds decision trees, with each subsequent tree correcting the errors of the previous ones [ 26 ]. It has been widely applied in medical research with impressive predictive accuracy [ 27 , 28 ]. In the present work, we utilized XGBoost to develop a final model incorporating clinical and demographic features available prior to IVF cycle initiation, making it a promising tool for early identification of patients at risk of low oocyte maturation rate and allowing for timely clinical intervention. To further elucidate the factors influencing oocyte maturation, SHAP analysis was performed on the XGBoost models for both PCOS and control groups, allowing direct comparison of feature importance and identification of PCOS-specific patterns. The initial timing of Gn administration emerged as the most influential predictor in the PCOS group but ranked only eighth in the control group. Delayed initiation beyond menstrual cycle day 3 was associated with higher oocyte maturation rates in both populations. Currently, no standardized protocol exists for determining the timing of Gn initiation, and clinical decisions are often guided by experience and patient presentation. Few studies have explored this variable, but one investigation among PCOS patients reported that starting stimulation on day 2 resulted in a higher total oocyte yield, whereas initiation on day 5 improved fertilization outcomes [ 29 ]. To our knowledge, our study is the first to identify an association between Gn initiation timing and oocyte maturation rate in PCOS patients undergoing IVF, offering a new perspective for clinical practice. Nevertheless, additional well-designed, large-scale studies are needed to explore and confirm these findings. Age was also identified as a key influencing factor in both models, ranking third in the PCOS group and sixth in the control group. SHAP dependence plots indicated that in PCOS patients, age over 31 was associated with higher predicted oocyte maturation, while in non-PCOS patients, the predicted oocyte maturation rate declined after age 32. While numerous studies have demonstrated that advanced maternal age is generally associated with lower oocyte maturation rate [ 20 , 29 ]—consistent with our findings in the control group—this trend is not observed across all populations. Some studies have reported no significant age-related differences in maturation rates [ 6 ], while others observed a U-shaped distribution, with the lowest maturation rates in women aged ≥ 41 years, followed by those under 30 [ 30 ]. Such inconsistencies may result from variations in age profiles and clinical characteristics among study cohorts, emphasizing the need to investigate oocyte maturation predictors specifically in patients with PCOS. Serum hormone levels are widely recognized as critical indicators in the IVF process. Among them, AMH is a key marker of ovarian reserve and was among the top predictors of oocyte maturation in both groups (ranked 4th in the PCOS model and 3rd in the control model based on SHAP analysis). In the control group, higher AMH levels were positively associated with oocyte maturation, in line with the findings of previous studies [ 6 , 20 , 29 , 31 ]. However, this pattern was not observed in the PCOS group. Notably, within the AMH range of 7.5 to 13.0 ng/mL, SHAP values were predominantly positive, suggesting an association with reduced oocyte maturation. This paradox may reflect the qualitative differences in ovarian reserve and follicular dynamics in PCOS patients, where excessively high AMH levels may reflect follicular dysregulation rather than competence. These findings underscore the importance of cautious interpretation of AMH values in PCOS populations. Our findings further indicated a positive association between basal LH levels and oocyte maturation rate. Consistently, previous research has suggested that elevated LH is a favorable predictor for oocyte maturation [ 32 ]. Animal model studies support this association, showing that suppression of basal LH levels can downregulate genes essential for oocyte development and competence [ 33 ]. Notably, our analysis revealed no significant difference in the proportion of LH supplementation between different outcome groups. Consequently, this variable was not prioritized in our model construction. Nevertheless, further exploration of the interaction between high basal LH and exogenous supplementation remains a valuable direction for future research. In addition, basal T levels emerged as a relatively more important predictor in the PCOS models. Elevated serum T levels were associated with lower oocyte maturation rates. Hyperandrogenism is a hallmark of PCOS [ 21 ] and is linked not only to reduced oocyte maturation but also to higher miscarriage rates, lower live birth rates, and increased risks of gestational diabetes and preeclampsia [ 34 – 36 ]. These results suggest that managing androgen levels may improve both IVF success and pregnancy outcomes in PCOS patients. The consistency of our model with these established clinical observations enhances its credibility. Obesity is well known to impair metabolic and cardiovascular function and adversely affect reproductive outcomes, including oocyte maturation, IVF success and pregnancy outcomes [ 37 ]. However, BMI did not emerge as a strong predictor in our models. This finding suggests a more complex relationship between BMI and reproductive potential, underscoring the need for more refined and individualized management strategies in IVF treatment. Importantly, lifestyle interventions remain essential for restoring menstrual regularity and improving reproductive outcomes in overweight or obese patients with PCOS [ 38 ]. This study has several limitations. First, constraints regarding data source and validation methodology limit the external validity of our findings. Our data were obtained from a single medical center, and we employed a random split approach to partition the dataset into training and validation sets. While random splitting is a common practice, we acknowledge its methodological limitations [ 39 ]. This approach reduces the effective sample size available for model development, potentially increasing the risk of overfitting, while simultaneously yielding a smaller validation set that may result in less precise performance estimates. Furthermore, since the validation set originates from the same source population, it represents only internal validation, limiting the assessment of the model’s generalizability to other clinical settings. Consequently, future studies utilizing multicenter, independent cohorts are essential. Second, there is currently no universally accepted classification standard for oocyte maturation rate. In this study, we selected a cut-off value informed by previous literature and our clinical experience, which we believe offers meaningful insights. Nevertheless, further research is warranted to establish a more precise and widely applicable classification system. Third, due to the retrospective design and the extended duration of data collection, certain clinical records and laboratory parameters necessary for sub-phenotyping were not consistently available across the entire cohort. As a result, phenotype-specific subgroup analyses could not be conducted. systematically collected and comprehensive phenotypic data are required to further delineate predictive factors across distinct PCOS subtypes. Finally, our dataset lacked information on glucose metabolism. Insulin resistance (IR) is a common metabolic disturbance in patients with PCOS and has been associated with poor IVF outcomes [ 1 ]. Previous studies have demonstrated that PCOS patients with IR tend to have lower oocyte maturation rates than those without IR [ 38 ]. The absence of such metabolic data may have limited both the predictive performance and clinical interpretability of our model.

Introduction

Polycystic ovary syndrome (PCOS) is a common and heterogeneous endocrine disorder affecting approximately 6–15% of women of reproductive age. It is characterized by anovulation, hyperandrogenism (HA), and polycystic ovarian morphology. The clinical manifestations of PCOS are diverse and closely associated with various comorbidities, including metabolic dysfunction, cardiovascular disease, psychological disorders, and pregnancy-related complications, all of which significantly impact women’s health. Notably, PCOS is also one of the leading causes of female infertility [ 1 ]. With the increasing prevalence of infertility, a growing number of women with PCOS are turning to in vitro fertilization (IVF) to achieve pregnancy [ 2 ]. This complex process involves the retrieval of mature oocytes, fertilization in vitro, and subsequent embryo transfer. Among these steps, oocyte maturation is a critical determinant of IVF success [ 3 ]. Immature oocytes impair fertilization, embryo quality, and implantation potential [ 4 ]. A lower oocyte maturation rate has been linked to reduced rates of normal fertilization, fewer high-quality embryos, lower clinical pregnancy and live birth rates, thereby imposing substantial physical, emotional, and financial burdens on patients [ 5 – 7 ]. Oocyte maturation arrest is a common complication in PCOS patients [ 8 , 9 ]. Previous studies have reported significantly lower oocyte maturation rate in PCOS patients during IVF cycles, with considerable inter-individual and inter-subtype variability [ 10 – 13 ]. Gonadotropin-releasing hormone (GnRH) antagonist protocol has gained popularity for controlled ovarian stimulation (COS) in recent years due to several advantages, such as shorter treatment duration, reduced gonadotropin requirement, and lower risk of ovarian hyperstimulation syndrome (OHSS) [ 14 ]. This protocol is particularly favored for patients with PCOS. This is due to their increased susceptibility to OHSS, irregular menstrual cycles, and the need for safer and more flexible stimulation strategies [ 15 ]. However, a notable drawback of this protocol is a potentially reduced oocyte maturation rate, which may be attributed to suboptimal follicular synchronization [ 16 ]. These challenges underscore the importance of accurately assessing the risk of suboptimal oocyte maturation and identifying clinical predictors in PCOS patients to inform personalized interventions and improve IVF outcomes. However, current evidence regarding the factors associated with impaired oocyte maturation remains limited and inconclusive, especially in women with PCOS. Moreover, IVF process generates large volumes of complex and heterogeneous data with intricate interdependencies, making it difficult for clinicians to interpret and apply this information in real-time clinical decision-making. Machine learning (ML) has emerged as a promising approach to address these challenges. ML algorithms are capable of detecting non-linear patterns, managing high-dimensional data, and modeling complex interactions among variables [ 17 ]. These capabilities make ML suitable for developing accurate and individualized prediction models. ML has demonstrated utility in early diagnosis of PCOS [ 18 ] and prediction of live birth outcomes [ 19 ], providing clinicians with valuable support for early intervention and personalized treatment planning. Nevertheless, ML-based models specifically targeting oocyte maturation rate in PCOS patients remain scarce, and most existing studies lack a parallel non-PCOS control cohort. Without a comparative control model, it is difficult to distinguish universal IVF predictors from those specific to PCOS pathophysiology. Therefore, distinct modeling is essential to elucidate group-specific feature contribution patterns and uncover PCOS-specific determinants that might otherwise be obscured in a unified model. In this study, we developed interpretable and personalized ML models to assess the risk of suboptimal oocyte maturation and identify clinical predictors in patients with and without PCOS undergoing the GnRH antagonist protocol. A total of ten ML models were constructed using a combination of clinical characteristics, cycle-specific variables, laboratory measurements, and ultrasound parameters. Their predictive performances were systematically evaluated and compared. Although oocyte maturation rate is a critical factor influencing IVF success, there is currently no universally accepted cut-off to define “high” versus “low” maturation rate. Previous studies have identified a threshold of 75% as optimal for improved reproductive outcomes and patients with maturation rates below this level have been reported to exhibit significantly reduced pregnancy rates [ 6 ]. Considering that average oocyte maturation rates in IVF cycles typically range between 77 and 85% [ 20 ], a threshold of 75% was selected as a clinically conservative cutoff to distinguish cycles with relatively suboptimal maturation performance. This cutoff lies slightly below the reported average range, allowing identification of patients who may be at increased risk of impaired oocyte maturation while maintaining clinical interpretability. Based on this rationale, 75% was used as the cut-off to stratify patients for the development of predictive classification models. To overcome the black-box nature of ML models, we employed Shapley Additive Explanations (SHAP) to quantify and visualize the contribution of each predictor, thereby enhancing model transparency and interpretability. By comparing SHAP value patterns of predictors between PCOS and non-PCOS groups, we aimed to identify PCOS-specific predictors that may explain differences in oocyte maturation rates. These insights could offer valuable clinical guidance for individualized treatment planning and decision-making, ultimately optimizing treatment strategies and improving reproductive outcomes in this population.

Supplementary Material

Supplementary Material 1. Figure S1. SHAP dependence plots for the prediction models. (A) SHAP dependence plot for the prediction model in PCOS group. (B) SHAP dependence plot for the prediction model in control group. Each dependence plot shows how a single feature affects the output of the prediction model, and each dot represents a single patient. SHAP, SHapley Additive exPlanations; Gn, gonadotropin; E2, Estradiol; AMH, anti-Müllerian hormone; LH, luteinizing hormone; T, testosterone; PRL, prolactin. Supplementary Material 2. Supplementary Material 1. Figure S1. SHAP dependence plots for the prediction models. (A) SHAP dependence plot for the prediction model in PCOS group. (B) SHAP dependence plot for the prediction model in control group. Each dependence plot shows how a single feature affects the output of the prediction model, and each dot represents a single patient. SHAP, SHapley Additive exPlanations; Gn, gonadotropin; E2, Estradiol; AMH, anti-Müllerian hormone; LH, luteinizing hormone; T, testosterone; PRL, prolactin. Supplementary Material 2.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-23T09:30:01.253652+00:00
License: CC-BY-NC-ND-4.0