Development of a single-center predictive model for conventional in vitro fertilization outcomes excluding total fertilization failure: implications for protocol selection.

OA: gold

Abstract

ObjectivesTo develop a multidimensional clinical indicator-based prediction model for identifying high-risk patients with fertilization failure conventional in vitro fertilization (c-IVF) cycles, thereby optimizing therapeutic decision-making.MethodsThis retrospective single-center study analyzed 691 cycles (594 c-IVF, 97 rescue ICSI) from January 2019 to August 2024. Key parameters included female age, BMI, male semen parameters (sperm concentration, total progressive motile sperm count [TPMC], DNA fragmentation index [DFI]), and infertility duration. Three machine learning models (logistic regression, random forest, XGBoost) were developed and validated using a nested cross-validation framework with SMOTE oversampling.ResultsThe logistic regression model demonstrated superior predictive performance (mean AUC = 0.734 ± 0.049), significantly outperforming random forest (0.714 ± 0.034) and XGBoost (0.697 ± 0.038). Significant predictors included protective factors-male age (OR = 0.642, 95%CI:0.598-0.689) and TPMC (OR = 0.428, 95%CI:0.392-0.466), and risk factors-female BMI (OR = 1.268, 95%CI:1.191-1.351) and DFI (OR = 1.362, 95%CI:1.274-1.455). The nomogram showed moderate-to-high discriminative power (C-index = 0.722, 95%CI:0.667-0.773) upon internal validation. Decision curve analysis confirmed clinical utility at threshold probabilities between 0.05 and 0.60.ConclusionsThe logistic regression-based prediction model exhibits robust performance in assessing c-IVF fertilization failure risk. While optimized for our center's specific clinical context, external multicenter validation is required to confirm broader clinical applicability.
Full text 40,007 characters · extracted from pmc-nxml · 5 sections · click to expand

Results

Comparison of baseline characteristics between the two groups revealed statistically significant differences in female age, male age, female BMI index, type of infertility, and years of infertility ( P  < 0.05), as shown in Table  1 . Table 1 Comparison of baseline characteristics between the two groups Variable c-IVF Group Rescue ICSI Group χ 2 /U Value P -value n 594 97 Female age [years, Median (P25, P75)] 34.0 (30.0, 38.0) 32.0 (28.0, 35.0) 34017.0 0.004 * Male age [years, Median (P25, P75)] 35.0 (31.0, 40.0) 34.0 (31.0, 37.0) 33667.0 0.008 * Female BMI [kg/m², Median (P25, P75)] 22.2 (20.0, 24.2) 22.9 (20.8, 25.4) 24518.0 0.019 * Type of infertility (%) 5.802 0.016 * Primary infertility 31.31% (186/594) 44.33% (43/97) Secondary infertility 68.69% (408/594) 55.67% (54/97) Infertility factors (%) 7.613 0.107 Ovulatory factors 17.85% (106/594) 23.71% (23/97) Male factors 8.59% (51/594) 15.46% (15/97) Pelvic factors 58.42% (347/594) 47.42% (46/97) Combined factors 4.55% (27/594) 4.12% (4/97) Other factor 10.61% (63/594) 9.28% (9/97) HBsAg (%) 0.898 0.343 Positive 6.60% (40/594) 4.12% (4/97) Negative 93.40% (566/594) 95.88% (93/97) Years of infertility [years] 2 (1, 3) 2 (2, 3) 25162.0 0.037 * Occupation (%) 1.019 0.961 White-collar 16.33% (97/594) 15.46% (15/97) Service and commercial 4.04% (24/594) 3.09% (3/97) Blue-collar and manufacturing 8.92% (53/594) 11.34% (11/97) Self-employed/entrepreneur 52.19% (310/594) 52.58% (51/97) Professional/technical 13.13% (78/594) 13.40% (13/97) Unemployed/other 5.39% (32/594) 4.12% (4/97) Education level (%) 2.615 0.455 Middle school or below 46.30% (275/594) 42.27% (41/97) College/university 35.86% (213/594) 32.99% (32/97) Vocational school/high school 17.00% (101/594) 23.71% (23/97) Postgraduate or above 0.84% (5/594) 1.03% (1/97) AMH [ng/mL, Median (P25,P75)] 2.69 (1.3425, 4.38) 2.9 (1.82, 4.67) 26309.5 0.170 Basal FSH [mIU/mL, Median (P25, P75)] 6.87 (5.56, 8.44) 6.59 (5.58, 8.23) 29402.5 0.745 Basal LH [mIU/mL, Median (P25, P75)] 4.555 (3.21, 6.79) 4.86 (3.75, 6.38) 26272.0 0.164 Basal E2 [pmol/L, Median (P25, P75)] 132.25 (93.1, 199.0) 134.2 (102.3, 191.7) 28003.0 0.658 Note: * indicates P  < 0.05 Comparison of baseline characteristics between the two groups Note: * indicates P  < 0.05 Comparison of ovarian stimulation parameters between the two groups showed no statistically significant differences in the adopted protocols, which included agonist protocol, antagonist protocol, follicular phase high progesterone ovarian stimulation protocol, minimal stimulation protocol, luteal phase stimulation protocol, and natural cycle protocol ( P  > 0.05). Similarly, no significant intergroup differences were observed in gonadotropin (Gn) dosage, Gn administration duration, or trigger-day estradiol (E2) levels ( P  > 0.05), as detailed in Table  2 . Table 2 Comparison of ovarian stimulation parameters between the two groups Variable c-IVF Group Rescue ICSI Group χ 2 /U Value P -value N 594 97 Ovarian stimulation protocols (%) 2.474 0.780 Agonist protocols (ultra-long protocol, short-acting long protocol, modified follicular-phase long protocol, luteal-phase long protocol, follicular-phase long protocol) 55.56% (330/594) 58.76% (57/97) Antagonist protocol 23.23% (138/594) 22.68% (22/97) Follicular-phase high progesterone ovarian stimulation protocol 10.77% (64/594) 12.37% (12/97) Minimal stimulation protocol 9.09% (54/594) 6.19% (6/97) Luteal-phase stimulation protocol 1.01% (6/594) 0.00% (0/97) Natural cycle protocol 0.34% (2/594) 0.00% (0/97) Gonadotropin (Gn) dosage (U) 2025.0 (1575.0, 2584.0) 2175.0 (1775.0, 2850.0) 25276.5 0.053 Gn administration duration (days) 11.0 (9.0, 12.0) 11.0 (9.0, 13.0) 27231.5 0.384 Trigger-day E2 [pmol/L, M(P25, P75)] 7952.7 (4753.75, 15436.3) 8508.0 (5420.0, 15231.0) 27643.0 0.523 Comparison of ovarian stimulation parameters between the two groups Comparison of male semen parameters between the two groups revealed no statistically significant differences in semen volume or non-progressive motility (NP) ( P  > 0.05). However, significant intergroup differences were observed in sperm concentration, total progressive motile sperm count, progressive motility (PR), immotile sperm rate (IM), sperm DNA fragmentation index (DFI), and sperm acrosin activity ( P  < 0.05), as detailed in Table  3 . Table 3 Male semen analysis results Variable c-IVF Group Rescue ICSI Group U Value P -value Semen volume [ml, M (P25, P75)] 2.5 (2.0, 3.0) 2.0 (2.0, 2.5) 31257.0 0.174 Sperm concentration [×10⁶/ml, M (P25, P75)] 70.0 (50.0, 97.5) 40.0 (30.0, 70.0) 40360.0 < 0.001 * Total progressive motile sperm count [×10⁶, M (P25, P75)] 60.0 (35.25, 100.0) 35.0 (18.0, 60.0) 40635.5 < 0.001 * Progressive motility (PR) [%, M (P25, P75)] 40.0 (30.0, 50.0) 30.0 (30.0, 40.0) 36311.0 < 0.001 * Non-progressive motility (NP) [%, M (P25, P75)] 20.0 (10.0, 20.0) 20.0 (10.0, 20.0) 30345.5 0.338 Immotile sperm (IM) [%, M (P25, P75)] 40.0 (40.0, 50.0) 50.0 (40.0, 60.0) 20503.5 < 0.001 * Sperm DNA fragmentation index (DFI) [%, M (P25, P75)] 10.5 (6.3, 16.4475) 14.3 (9.49, 20.58) 21946.5 < 0.001 * Acrosin activity [uIU/10⁶ sperm, M (P25, P75)] 76.1 (53.4, 106.9) 73.3 (40.4, 94.3) 32766.5 0.030 * Note: * indicates P  < 0.05 Male semen analysis results Note: * indicates P   0.05). Before rescue ICSI, significant intergroup differences were observed in fertilization outcomes (normal fertilization, abnormal fertilization, unfertilized oocytes) ( P  < 0.05). Following rescue ICSI intervention, the normal fertilization rate increased from 17.16% (150/874) to 61.78% (540/874), the abnormal fertilization rate rose from 2.97% (26/874) to 9.95% (87/874), and the unfertilized rate decreased from 79.86% (698/874) to 28.26% (247/874), with persistent statistical significance between groups ( P  < 0.05), as detailed in Table  4 . Table 4 Oocyte retrieval and fertilization outcomes Variable c-IVF Group Rescue ICSI Group χ 2 Value P -value Number of oocytes retrieved (%) 1.714 0.190 Mature oocytes 88.46% (5517/6237) 86.97% (874/1005) Immature oocytes 11.54% (720/6237) 13.03% (131/1005) Fertilization outcomes before rescue ICSI (%) 1882.518 < 0.001 * Normal fertilization 67.34% (3715/5517) 17.16% (150/874) Abnormal fertilization 18.98% (1047/5517) 2.97% (26/874) Unfertilized oocytes 13.68% (755/5517) 79.86% (698/874) Fertilization outcomes after rescue ICSI (%) 140.355 < 0.001 * Normal fertilization 67.34% (3715/5517) 61.78% (540/874) Abnormal fertilization 18.97% (1047/5517) 9.95% (87/874) Unfertilized oocytes 13.69% (755/5517) 28.26% (247/874) Note: * indicates P  < 0.05 Oocyte retrieval and fertilization outcomes Note: * indicates P  < 0.05 After aggregating all P  < 0.05 variables from Tables  1 and 2 , and 3 (Female age, Male age, Female BMI, Sperm acrosin activity, Sperm concentration, PR, TPMC, IM, DFI, Type of infertility, Years of infertility) and performing collinearity tests by checking the Variance Inflation Factor (VIF), backward stepwise regression analysis identified Male age, Female BMI, TPMC, and DFI as predictors for the logistic regression model (Fig.  3 ). Subsequently, nine variables with cumulative Gini importance > 90%—Female age, Male age, Sperm acrosin activity, Sperm concentration, PR, TPMC, IM, DFI, and Years of infertility—were selected for the random forest model (Fig.  4 ). Finally, XGBoost algorithm-based feature importance assessment (with a cumulative Gain value threshold > 90%) identified nine core predictors: Male age, Female BMI, Sperm concentration, PR, TPMC, IM, DFI, Type of infertility, and Years of infertility (Fig.  5 ). Fig. 3 This heatmap illustrates the variable screening process based on backward stepwise regression analysis. The x-axis represents selection steps (Step 1–9), and the y-axis denotes candidate variables. The color intensity of each cell reflects the p -value at the corresponding step (darker blue indicates lower p -values and higher statistical significance). The final predictors selected for the logistic regression model include Male age, Female BMI, TPMC, and DFI ( P  < 0.05 at Step 9). The color bar on the right indicates the p -value range (0-0.7), with yellow representing high p -values (low significance) and blue representing low p -values (high significance) This heatmap illustrates the variable screening process based on backward stepwise regression analysis. The x-axis represents selection steps (Step 1–9), and the y-axis denotes candidate variables. The color intensity of each cell reflects the p -value at the corresponding step (darker blue indicates lower p -values and higher statistical significance). The final predictors selected for the logistic regression model include Male age, Female BMI, TPMC, and DFI ( P  < 0.05 at Step 9). The color bar on the right indicates the p -value range (0-0.7), with yellow representing high p -values (low significance) and blue representing low p -values (high significance) Fig. 4 This figure illustrates the feature importance ranking based on Gini criteria from a random forest model. The x-axis represents standardized Gini importance scores, and the y-axis lists clinical variables sorted in descending order of importance. To maximize model performance, a cumulative Gini importance threshold of > 90% was applied, resulting in the selection of nine core predictors: Years of infertility, TPMC, Sperm concentration, IM, DFI, Male age, PR, Female age, and Sperm acrosin activity. Notably, years of infertility and total progressive motile sperm count (TPMC) exhibited the highest contributions (Gini importance > 0.14), while female BMI and type of infertility were excluded due to lower importance scores This figure illustrates the feature importance ranking based on Gini criteria from a random forest model. The x-axis represents standardized Gini importance scores, and the y-axis lists clinical variables sorted in descending order of importance. To maximize model performance, a cumulative Gini importance threshold of > 90% was applied, resulting in the selection of nine core predictors: Years of infertility, TPMC, Sperm concentration, IM, DFI, Male age, PR, Female age, and Sperm acrosin activity. Notably, years of infertility and total progressive motile sperm count (TPMC) exhibited the highest contributions (Gini importance > 0.14), while female BMI and type of infertility were excluded due to lower importance scores Fig. 5 This figure presents feature importance scores (Gain values) from an XGBoost model. The x-axis represents standardized Gain values, and the y-axis lists clinical variables sorted in descending order of importance. By applying a cumulative Gain threshold of > 90%, nine core predictors were identified: Sperm concentration, IM, TPMC, PR, Years of infertility, Type of infertility, DFI, Male age, and Female BMI. Notably, sperm concentration exhibited the highest contribution, while female age was excluded due to lower importance. This selection strategy balances model complexity with predictive power, ensuring included features significantly contribute to predicting fertilization failure risk This figure presents feature importance scores (Gain values) from an XGBoost model. The x-axis represents standardized Gain values, and the y-axis lists clinical variables sorted in descending order of importance. By applying a cumulative Gain threshold of > 90%, nine core predictors were identified: Sperm concentration, IM, TPMC, PR, Years of infertility, Type of infertility, DFI, Male age, and Female BMI. Notably, sperm concentration exhibited the highest contribution, while female age was excluded due to lower importance. This selection strategy balances model complexity with predictive power, ensuring included features significantly contribute to predicting fertilization failure risk During model development and validation, all three models (logistic regression, random forest, and XGBoost) demonstrated robust predictive performance through nested 5-fold cross-validation, with outer-loop average AUCs of 0.734 (95% CI 0.668–0.802), 0.714 (95% CI 0.667–0.762), and 0.697 (95% CI 0.645–0.748), respectively (Fig.  6 ). Due to limited test set sample size, all models exhibited lower precision and recall in the rescue ICSI group validation (Table  5 ). The logistic regression model achieved optimal performance in the c-IVF group (precision 0.91, recall 0.87, F1-score 0.88, overall accuracy 81%). While the random forest model showed marginally higher recall (0.88) and F1-score (0.89) in the c-IVF group compared to logistic regression, its precision (0.89) and overall accuracy (81%) remained comparable without significant differences. The XGBoost model demonstrated slight fluctuations in the c-IVF group (precision 0.88, recall 0.85, F1-score 0.87, accuracy 78%), though performance remained within acceptable ranges. In the rescue ICSI group, logistic regression maintained the highest recall (0.44) among the three models, though absolute values remained low. Random forest exhibited better recall (0.34 vs. 0.32) and significantly higher precision (0.34 vs. 0.26) than XGBoost in this group, suggesting higher false-positive risk for XGBoost predictions. Overall, all models showed significantly reduced classification performance in the rescue ICSI group compared to the c-IVF group (e.g., logistic regression F1-score decreased from 0.88 to 0.38; XGBoost from 0.87 to 0.29), indicating that feature distribution or sample size in the rescue ICSI group may pose substantial challenges for model learning. Targeted feature engineering optimization or algorithm adaptation strategies for small-sample/high-noise scenarios are recommended for future improvements. Fig. 6 This figure presents receiver operating characteristic (ROC) curves for three models during nested 5-fold cross-validation: Logistic Regression: Mean AUC = 0.734 (SD = 0.049), with a narrow confidence interval and significant separation from the random classifier (diagonal line). Random Forest: Mean AUC = 0.714 (SD = 0.034), exhibiting similar curve shape to logistic regression but slightly wider confidence intervals. XGBoost: Mean AUC = 0.697 (SD = 0.038), with the lowest overall curve elevation and broadest confidence interval, indicating weaker predictive performance. Notably, inter-model differences are most pronounced in high-specificity regions (FPR < 0.2), where logistic regression maintains higher TPR at low FPR compared to XGBoost This figure presents receiver operating characteristic (ROC) curves for three models during nested 5-fold cross-validation: Logistic Regression: Mean AUC = 0.734 (SD = 0.049), with a narrow confidence interval and significant separation from the random classifier (diagonal line). Random Forest: Mean AUC = 0.714 (SD = 0.034), exhibiting similar curve shape to logistic regression but slightly wider confidence intervals. XGBoost: Mean AUC = 0.697 (SD = 0.038), with the lowest overall curve elevation and broadest confidence interval, indicating weaker predictive performance. Notably, inter-model differences are most pronounced in high-specificity regions (FPR < 0.2), where logistic regression maintains higher TPR at low FPR compared to XGBoost Table 5 Summary of model classification performance Models Groups Precision Recall F1-Score Accuracy Logistic Regression c-IVF Group 0.91 0.87 0.88 0.81 RICSI Group 0.35 0.44 0.38 Random Forest c-IVF Group 0.89 0.88 0.89 0.81 RICSI Group 0.34 0.34 0.33 XGBoost c-IVF Group 0.88 0.85 0.87 0.78 RICSI Group 0.26 0.32 0.29 Note: The data in this table represent the average statistics from five outer-loop cross-validation runs. Precision indicates the proportion of samples predicted as positive that are actually positive; Recall represents the proportion of actual positive samples that are correctly predicted; F1-Score is the harmonic mean of Precision and Recall; Accuracy reflects the proportion of all samples that are correctly predicted Summary of model classification performance Note: The data in this table represent the average statistics from five outer-loop cross-validation runs. Precision indicates the proportion of samples predicted as positive that are actually positive; Recall represents the proportion of actual positive samples that are correctly predicted; F1-Score is the harmonic mean of Precision and Recall; Accuracy reflects the proportion of all samples that are correctly predicted The optimal logistic regression model parameters (Table  6 ) revealed that TPMC (OR = 0.428; 95% CI: 0.392–0.466; P  < 0.001) and male age (OR = 0.642; 95% CI: 0.598–0.689; P  < 0.001) were protective factors against fertilization failure, with higher values linked to reduced risk. In contrast, female BMI (OR = 1.268; 95% CI: 1.191–1.351; P  = 0.005) and DFI (OR = 1.362; 95% CI: 1.274–1.455; P  = 0.005) acted as independent risk factors, where elevated values significantly increased fertilization failure risk. All variables achieved statistical significance ( P  < 0.05), confirming their predictive validity. The nomogram (Fig.  7 ) constructed based on male age (22–57 years), female BMI (14.6–35.2 kg/m²), TPMC (2–250 × 10⁶), and DFI (0.39–62.07%) demonstrated a C-index of 0.722 (95% CI: 0.667–0.773) through bootstrap validation (1000 resamples), indicating moderate-to-strong discrimination for c-IVF fertilization failure risk. This performance was significantly superior to random guessing (C-index = 0.5, P  < 0.001). To evaluate the clinical utility of our predictive model, decision curve analysis (DCA) was performed (Fig.  8 ). The DCA demonstrated that the logistic regression model provided a positive net benefit across threshold probabilities from 0.05 to approximately 0.60, indicating clinical usefulness within this range. The model outperformed both the ‘treat all’ and ‘treat none’ strategies within these thresholds, suggesting potential clinical value in guiding early intervention decisions for patients at risk of c-IVF fertilization failure. However, at higher threshold probabilities (> 0.60), the net benefit decreased, suggesting cautious application of the model when extremely high certainty is required before intervention. External validation is recommended to enhance clinical applicability. Table 6 Features selected in the conventional logistic regression B S.E. OR P 95% CI Lower (OR) 95% CI Upper (OR) Intercept -0.353 0.036 0.702 < 0.001 0.655 0.754 Male age -0.444 0.036 0.642 < 0.001 0.598 0.689 Female BMI 0.238 0.032 1.268 0.005 1.191 1.351 TPMC -0.850 0.044 0.428 < 0.001 0.392 0.466 DFI 0.309 0.034 1.362 0.005 1.274 1.455 Notes: (a) Results show averaged statistics over 5 outer folds with pooled standard errors; (b) ORs represent relative risk per 1-SD increase in standardized features, 95% CIs calculated via pooled SEs Features selected in the conventional logistic regression Notes: (a) Results show averaged statistics over 5 outer folds with pooled standard errors; (b) ORs represent relative risk per 1-SD increase in standardized features, 95% CIs calculated via pooled SEs Fig. 7 This nomogram, constructed using logistic regression coefficients, quantifies the contribution of each variable (male age, female BMI, TPMC, DFI) to fertilization failure risk. Contribution scores (Points) are calculated via weighted model coefficients, with an overall score range of 100.0–150.0 (Overall point axis). Higher scores correlate with increased risk. The Positive Risk axis at the bottom visualizes the predicted probability (0–1 scale), where a green dashed line (threshold = 0.2) denotes the risk threshold. When the overall score exceeds 112 points, the predicted risk surpasses this threshold and rises sharply, indicating the need for early intervention This nomogram, constructed using logistic regression coefficients, quantifies the contribution of each variable (male age, female BMI, TPMC, DFI) to fertilization failure risk. Contribution scores (Points) are calculated via weighted model coefficients, with an overall score range of 100.0–150.0 (Overall point axis). Higher scores correlate with increased risk. The Positive Risk axis at the bottom visualizes the predicted probability (0–1 scale), where a green dashed line (threshold = 0.2) denotes the risk threshold. When the overall score exceeds 112 points, the predicted risk surpasses this threshold and rises sharply, indicating the need for early intervention Fig. 8 Decision curve analysis (DCA) of the logistic regression model for predicting c-IVF fertilization failure. The x-axis represents threshold probability, and the y-axis represents net benefit. The blue line represents the mean model performance across all five cross-validation folds, with the gray shaded area indicating ± 1 standard deviation. The model demonstrates positive net benefit compared to “treat all” (dashed line) and “treat none” (solid horizontal line) strategies within threshold probabilities ranging from 0.05 to approximately 0.60 Decision curve analysis (DCA) of the logistic regression model for predicting c-IVF fertilization failure. The x-axis represents threshold probability, and the y-axis represents net benefit. The blue line represents the mean model performance across all five cross-validation folds, with the gray shaded area indicating ± 1 standard deviation. The model demonstrates positive net benefit compared to “treat all” (dashed line) and “treat none” (solid horizontal line) strategies within threshold probabilities ranging from 0.05 to approximately 0.60

Materials

This study retrospectively analyzed clinical data from the Reproductive Medicine Center of our hospital between January 2019 and August 2024. The initial inclusion criterion was all cycles undergoing ART treatment, with exclusion criteria including cycles resulting in TFF, cycles treated with ICSI, and cycles treated with half-volume intracytoplasmic sperm injection (Half-ICSI). A total of 691 cycles were included and divided into two groups based on a threshold of 2 PB extrusion rate at 6 h post-IVF insemination: 594 cycles with 2 PB extrusion rate ≥ 30% were categorized as the c-IVF group, and 97 cycles with 2 PB extrusion rate < 30% were assigned to the rescue ICSI group. The process is illustrated in Fig.  1 . Fig. 1 Visual diagram of the detailed process for data collection Visual diagram of the detailed process for data collection The following data were retrieved through the ART Workflow Management System (Wuhan Huchuang Technology Co., Ltd., version v9.2.28.30) at our center: The baseline characteristics of enrolled patients included: female age, male age, female BMI, occupation type, educational level, years of infertility, infertility-related factors (e.g., tubal factors, ovulatory dysfunction), type of infertility (primary or secondary), controlled ovarian stimulation protocol, total gonadotropin (Gn) dosage, and Gn administration duration. Hormonal assessments included basal sex hormones [estradiol (E2), luteinizing hormone (LH), follicle-stimulating hormone (FSH)], trigger-day E2 levels, and anti-Müllerian hormone (AMH). All measurements were performed using the Beckman Coulter DxI 800 Immunoassay Analyzer with chemiluminescence immunoassay technology. Hepatitis B surface antigen (HBsAg) testing was conducted using the Shenzhen Yhlo Biotech iFlash 3000 Chemiluminescence Immunoassay Analyzer. Semen samples were collected after an abstinence period of 3–5 days on the day of oocyte retrieval, strictly following the standardized protocols outlined in the WHO Laboratory Manual for the Examination and Processing of Human Semen (5th edition). Semen volume was measured using the gravimetric method, and key parameters—including sperm concentration, progressive motility (PR), non-progressive motility (NP), and immotile sperm (IM)—were quantitatively analyzed using a Makler Counting Chamber (Sefi Medical Instruments, Israel). The total progressively motile sperm count (TPMC) was calculated using the formula: semen volume ( mL )× sperm concentration (×10⁶/ mL )× PR . Sperm selection was performed using the Isolate ® double-density gradient centrifugation method (with Sperm Separation Medium Upper Layer and Sperm Separation Medium Lower Layer), strictly adhering to standard operating procedures (SOPs). Finally, the selected sperm were suspended in G-IVF PLUS medium, adjusted to a PR concentration of 3 × 10⁶/mL with a total volume of 0.5 mL, for subsequent c-IVF insemination. The most recent semen analysis results (within 3 months) were obtained using the modified Kennedy enzymatic method. The procedure was as follows: Freshly liquefied semen was thoroughly mixed, and sperm density was determined through preliminary testing. A precisely measured volume of semen containing 7.5 × 10⁶ spermatozoa was added to both experimental (test) tubes and control tubes. After centrifugation at 3000×g for 10 min to remove supernatant, the pellet was processed according to the reagent instructions to establish the enzymatic reaction system. Absorbance values were measured at 405 nm using a microplate reader, and experimental data were quantitatively analyzed based on the absorbance difference between test and control tubes. The most recent semen analysis results (within 3 months) were obtained using the Sperm Chromatin Structure Assay (SCSA). The principle of this method is based on standard acid denaturation, followed by metachromatic staining with acridine orange dye, which selectively binds to different DNA structures: normal double-stranded DNA in sperm emits green fluorescence (excitation wavelength 530 nm), while abnormal single-stranded DNA emits red fluorescence (excitation wavelength 640 nm). A flow cytometry system was used to analyze 5,000 sperm events per sample, and the DNA fragmentation index (DFI) was quantitatively determined by calculating the ratio of green fluorescence intensity (double-stranded DNA) to red fluorescence intensity (single-stranded DNA). Number of Oocytes Retrieved: Total number of cumulus-oocyte complexes (COCs) obtained via transvaginal ultrasound-guided follicular aspiration, counted based on the presence of intact oocyte structures under microscopic observation. Number of Mature Oocytes: Cumulus-oocyte complexes were mechanically denuded 4 h post-IVF insemination, and oocyte morphology was assessed under an inverted microscope (×200 magnification). Mature oocytes were defined as those with uniform cytoplasm, intact zona pellucida, and completion of the first meiotic division (Metaphase II, MII stage). The specific criterion for maturity was the visibility of the first polar body (First Polar Body, 1 PB), confirming completion of the first meiotic division (MII stage). 2 PB was not required for maturity assessment: this structure is only extruded post-fertilization and serves to evaluate fertilization status rather than oocyte maturation. Number of Fertilized Oocytes: Evaluated 18 ± 1 h post-IVF insemination, with fertilization confirmed by the presence of distinct pronuclear (PN) structures. Normal Fertilization Count: Number of oocytes exhibiting two pronuclei (2PN) of equal size with clearly discernible nucleoli, consistent with normal human fertilization characteristics. Abnormal Fertilization Count: Includes oocytes with single pronucleus (1PN), three pronuclei (3PN), or multiple pronuclei (≥ 4PN). All fertilization evaluations described above were independently reviewed and confirmed by two senior embryologists. Normality and homogeneity of variance tests were performed for all data. Normally distributed data are presented as mean ± standard deviation and analyzed using independent-sample t-tests, while non-normally distributed data are expressed as median (interquartile range) M ( P 25, P 75) and compared between groups using the Mann-Whitney U test. Categorical variables are presented as percentages (%) and analyzed using the chi-square (χ²) test. During model development, three predictive models were constructed: a conventional logistic regression model, a random forest model, and an XGBoost algorithm model. A nested cross-validation framework was employed to evaluate model performance. The outer validation used stratified 5-fold cross-validation to partition the dataset into training (80%) and testing sets (20%), while the inner validation applied 5-fold stratified cross-validation to optimize hyperparameters for each outer fold. To address sample imbalance, the Synthetic Minority Over-sampling Technique (SMOTE) was integrated into the modeling pipeline, ensuring that synthetic sampling was applied exclusively to training folds during inner validation, with testing folds retaining the original distribution. The outer validation phase utilized untreated independent test sets for final performance evaluation. This dual validation mechanism effectively controlled information leakage risks and reduced random bias from data partitioning through repeated sampling (as shown in Fig.  2 ). Statistically significant clinical indicators between the two groups were incorporated into model analysis: (1) A conventional logistic regression model was built using variables selected via backward stepwise regression with chi-square test criteria; (2) The random forest model was constructed based on cumulative Gini importance for feature selection; (3) The XGBoost model was used to analyze variable contributions to fertilization failure prediction through Gain values. All models underwent hyperparameter tuning to maximize performance. Model comparisons were conducted using receiver operating characteristic (ROC) curve area under the curve (AUC), test set prediction accuracy, regression coefficients, and F1 scores. All analyses were performed using Python 3.8 with standardized data processing, and statistical significance was set at P  < 0.05. Fig. 2 Model algorithm flowchart Model algorithm flowchart This retrospective study utilized anonymized data generated from routine clinical practice at our reproductive medicine center. The research protocol was approved by the Institutional Review Board of Wenzhou People’s Hospital (Approval No.KY-202501-010), and all procedures complied with the Ethical Guidelines for Human Assisted Reproductive Technology and the Declaration of Helsinki. We confirm that written informed consent encompassing potential research use was obtained from all participants during their initial c-IVF treatment, and all analyzed data were fully de-identified.

Conclusion

This study demonstrates the clinical feasibility of machine learning-based predictive tools for early identification of fertilization failure in c-IVF cycles. The multidimensional model integrating sperm DFI, TPMC, and metabolic parameters significantly enhances predictive accuracy for fertilization outcomes. Although developed using single-center data, which provides a center-specific tool with high local relevance, our model lays the groundwork for precision and personalized reproductive care. Future investigations should focus on refining predictive models through multi-omics integration, multi-center validation, and advanced bioinformatics to elucidate underlying biological mechanisms.

Discussion

Factors influencing oocyte fertilization​​ include oocyte quality (maturity, genetic integrity), sperm parameters (motility, DFI, acrosin activity), and in vitro culture systems (medium composition, culture conditions, standardized procedures). To optimize outcomes for c-IVF patients, our center implemented the following quality control strategies: individualized ovarian stimulation protocols based on follicular development [ 10 ] sperm selection optimization [ 11 ] and strict standardization of embryology laboratory protocols [ 12 ]. Nevertheless, 14.04% (97/691) of cycles exhibited low c-IVF fertilization rates. Following rescue ICSI, the normal fertilization rate significantly increased from 17.16 to 61.78%. Although overall fertilization outcomes (abnormal fertilization and total fertilization failure rates) remained statistically different compared to c-IVF (χ²=140.355, P  < 0.001), the intergroup difference in the key clinical parameter—normal fertilization rate—was only 5.56% (67.34% vs. 61.78%), with a 95% confidence interval (2.11-9.01%) whose lower bound approached the predefined clinically meaningful difference threshold (5%). This demonstrates that rescue ICSI can effectively maintain normal fertilization competence comparable to c-IVF. Comparative analysis of predictive models​​ identified statistically significant predictors through systematic evaluation of patient demographics, ovarian stimulation parameters, and semen characteristics. Key factors included female parameters (age, BMI), male parameters (age, sperm concentration, TPMC, sperm acrosin activity, PR, IM, DFI), and infertility-related factors (infertility type, duration). Existing literature has extensively documented the regulatory roles of these factors in gamete and embryo quality [ 13 – 22 ]. We developed three predictive models—logistic regression, random forest, and XGBoost—and evaluated their performance using a nested five-fold cross-validation framework. The logistic regression model demonstrated superior predictive performance (mean AUC = 0.734 ± 0.049) compared to random forest (0.714 ± 0.034) and XGBoost (0.697 ± 0.038) ( p  < 0.05). Classification metrics (precision, recall, F1-score, accuracy) further confirmed the dominance of logistic regression. Given the linear separability of data and low feature complexity, logistic regression was selected as the optimal predictive model. ​Multivariate regression analysis​​ identified male age, female BMI, TPMC, and DFI as significant predictors of IVF fertilization outcomes, with TPMC demonstrating particularly strong predictive power for normal fertilization rates. While prior studies have established that advanced male age correlates with increased sperm DNA damage, impaired chromatin maturity, and reduced motility—all contributing to diminished fertility [ 23 – 25 ] —our data revealed younger parental ages in the rescue ICSI group compared to the c-IVF group. This discrepancy may stem from: (1) single-center study limitations with restricted sample size; (2) preferential use of ICSI in older males with poor semen quality [ 26 ] (due to anticipated low c-IVF success), creating a “healthy survivor effect” in the c-IVF group (i.e., only healthier older males opting for c-IVF) [ 27 , 28 ]. Furthermore, female BMI > 25 kg/m² is associated with diminished ovarian function and reserve, potentially compromising ART outcomes [ 29 , 30 ]. Elevated BMI also correlates with increased gonadotropin requirements [ 31 ] reduced embryo quality, higher miscarriage rates, and lower live birth rates [ 32 ]. TPMC has been widely validated as a predictor of fertilization failure [ 18 , 33 – 35 ] with a sensitivity of 80% when using TPMC < 1.5 × 10⁶ as the threshold [ 36 ]. Although sperm DFI primarily impacts post-fertilization outcomes—including transferable embryo yield, pregnancy rates, and early miscarriage risk [ 37 – 39 ] —our logistic regression analysis detected its independent influence on fertilization rates using data from semen analyses within 3 months (despite no DFI assessment on oocyte retrieval day), suggesting potential predictive value warranting further investigation. The exploration of predictive indicators for c-IVF fertilization failure has evolved progressively. Early models primarily focused on male age, c-IVF cycle parameters, and semen characteristics (e.g., Repping et al. [ 40 ] achieved AUC = 0.75 using total motile sperm count [TMC]), yet overlooked sperm morphological and DNA integrity metrics. Although TMC has been shown to reflect sperm functional capacity, it does not directly assess DNA integrity or morphological quality. Subsequent studies incorporated ovarian responsiveness indicators (e.g., Wang et al. [ 41 ] improved AUC to 0.899 by integrating basal LH, progesterone levels, and progressive motility [PR%]) and laboratory processing parameters (e.g., Li et al. [ 42 ] reached AUC = 0.815 with swim-up sperm concentration and normal morphology rate), but omitted metabolic factors. While recent advances by Xingnan et al. [ 43 ] included ovarian reserve markers (infertility duration, mature oocyte rate), body composition parameters like BMI remained unaddressed. These ovarian-related factors primarily reflect female contributions and do not fully capture the male-side risk profile. Discrepancies in our predictive factors may stem from three considerations: the innovative inclusion of sperm DNA fragmentation index (DFI) enhances sensitivity to genetic material integrity; regional variations in population characteristics and laboratory protocols might alter variable weighting; furthermore, optimized feature selection through random forest and XGBoost algorithms enables precise detection of interaction effects between metabolic indicators (e.g., BMI) and male age, a dimension under captured by previous single-model approaches. It is also crucial to note significant heterogeneity in the definition of ‘fertilization failure’ among the referenced prediction models. For example, Repping et al. [ 40 ] and Li et al. [ 42 ] predicted TFF assessed at the traditional 16–18 h mark (long protocol). Wang et al. [ 41 ] developed models for both TFF and LFR. Conversely, Xingnan et al. [ 43 ] predicted LFR based on an early 4-hour assessment (short protocol). ​​These differences stem fundamentally from variations in laboratory workflows regarding fertilization assessment timing (short vs. long protocol), which directly dictates the clinical window for rescue ICSI intervention.​​ Identifying potential LFR early via a short protocol allows for more timely rescue ICSI, potentially improving outcomes compared to the later intervention dictated by TFF diagnosis in a long protocol. ​​Consequently, the choice of predictive target (TFF vs. LFR) inherently aligns with specific laboratory practices and the clinical problem they aim to solve (late vs. early intervention). This fundamental variation in outcome definition contributes to differences in reported predictors and model performance.​. Study limitations should be carefully considered. Despite utilizing five-year data, the single-center design with limited sample size may constrain the generalizability of conclusions. Our study exclusively utilized clinical data from our reproductive center, deliberately choosing this approach due to challenges in accessing multi-center data. While this provides high internal consistency and a tool optimized for our specific clinical practices, it limits broader applicability [ 44 ]. Technical variations during the study period (e.g., personnel turnover, methodological updates) could introduce confounding biases [ 45 ]. The significant sample size imbalance between groups necessitated SMOTE application [ 46 ] though synthetic samples may not fully capture complex biological mechanisms underlying fertilization failure. Furthermore, the study omitted potential determinants such as sperm morphology, lifestyle factors (smoking/alcohol/dietary patterns) [ 47 ] and genetic/endocrine profiles, which may interact with identified variables. The clinical utility of our model requires validation through multi-center prospective studies [ 48 ] as thresholds for parameters like DFI and BMI may require adjustment for different populations [ 49 ]. Future multi-center studies with expanded cohorts and multi-omics integration could systematically elucidate the multidimensional regulatory networks underlying c-IVF fertilization rates.

Supplementary Material

Below is the link to the electronic supplementary material. Supplementary Material 1 Supplementary Material 1

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-16T09:21:09.727480+00:00