Methods
We enrolled the following infertile women who planned to receive IVF-ET from 2018 to 2021, and the following inclusion criteria were used: (1) aged 22 to 43 years; (2) with the normal karyotype; (3) BMI: 16 ~ 33 kg/m 2 ; (4) the endometrial thickness was not less than 6 mm when a dominant follicle ≥ 14 mm in a natural cycle (or on day 10 of the HRT cycle). The exclusion criteria were as follows: endocrine abnormalities (AFC < 5), uterine leiomyomas (with intermural leiomyomas of more than 3 cm or submucosal leiomyomas regardless of size), endometrial polyps (≥ 1.5 cm), and uterine adhesions. Clinical characteristics of the enrolled women were recorded, including number of failed embryo transfers, number of failed embryos (failed in pregnancy), and whether they had recurrent implantation failure. Recurrent implantation failure (RIF) was defined as the failure to achieve a clinical pregnancy after two or more embryo-transfer cycles with good-quality embryos, consistent with previous literature (Polanski et al. 2014 , Coughlan et al. 2014 ). Other variables recorded included the antral follicle count (AFC, including left AFC and right AFC), endometrial thickness, endometrial morphology under ultrasound imaging, and hormone test indicators (FSH, LH, PRL, and AMH). The IVF-ET treatment was performed as routine. The quality of inner cell mass (ICM) and trophectoderm cells were documented. For transferred embryos, the following features were recorded: fresh embryo or frozen embryo, natural cycle or artificial cycle, type of embryo transferred (cleavage or blastocyst), and the rating scores of the cleavage embryos or the blastocyst embryos (the expansion state of blastocele, ICM and trophectoderm cells). For each patient, if only one embryo was transferred, the current embryo was regarded as the optimal embryo; if two embryos were transferred, the better rated embryo was the optimal embryo, irrespective of whether at the cleavage or blastocyst stage. And the later outcomes (pregnancy / live birth) were documented.
In natural cycles, patients performed ovulation tests to determine the luteinizing hormone (LH) surge, and endometrial biopsies were collected within the same cycle on day 7–9 after the LH surge (LH + 7–9). Urinary LH surge was detected using commercially available ovulation predictor kits (Clearblue ® , Swiss Precision Diagnostics, UK). The first day of a positive test result was defined as LH + 0, and endometrial biopsy was subsequently scheduled on LH + 7–9. In hormone replacement therapy (HRT) cycles, biopsies were performed 5–7 days after progesterone initiation ( P + 5–7). This timing corresponds to the window of implantation and was synchronized with the day of embryo transfer for each patient, ensuring that all samples represented receptive-phase endometrium. Endometrial tissues were obtained using disposable uterine-cavity tissue-suction catheters and immediately stored in liquid nitrogen for RNA extraction. Total RNA was isolated using the TRIzol reagent, and RNA quantity and purity were assessed with a NanoDrop spectrophotometer (NanoDrop Technologies, USA). mRNA sequencing was performed on the Illumina HiSeq platform (Illumina, USA).
Differentially expressed genes (DEGs) between pregnancy and no-pregnancy groups were first identified using the DESeq2 R package, with FDR 1. This primary analysis yielded 180 DEGs (131 with lower expression in the pregnancy group and 49 with higher expression), which are shown in the volcano plot and heatmap (Fig. 1 A, B) and were used for the initial GO/KEGG enrichment analyses (Fig. 1 C, D).
Fig. 1 Differentially expressed genes (DEGs) associated with pregnancy in the transcriptomic analysis and related enrichments. A Volcano plot of DEGs between pregnancy and no-pregnancy groups in the primary DESeq2 analysis (FDR 1), including 131 pregnancy-lower genes and 49 pregnancy-higher genes. B Heatmap of 180 DEGs. C Bubble plot of the top enriched GO biological processes based on the 180-gene discovery panel. D Bubble plot of the top enriched KEGG pathways based on the 180-gene discovery panel
Differentially expressed genes (DEGs) associated with pregnancy in the transcriptomic analysis and related enrichments. A Volcano plot of DEGs between pregnancy and no-pregnancy groups in the primary DESeq2 analysis (FDR 1), including 131 pregnancy-lower genes and 49 pregnancy-higher genes. B Heatmap of 180 DEGs. C Bubble plot of the top enriched GO biological processes based on the 180-gene discovery panel. D Bubble plot of the top enriched KEGG pathways based on the 180-gene discovery panel
To evaluate whether baseline clinical variables could confound these transcriptomic differences, a covariate-adjusted sensitivity analysis was then conducted in DESeq2 using a multivariable design formula (age + RIF_status + rightAFC + outcome). In this adjusted analysis, 168 of the 180 DEGs (93.3%) remained significant with consistent effect direction. We further included cycle type (natural vs. HRT) as an additional covariate (cycle_type + age + RIF_status + rightAFC + outcome) and performed cycle-type–stratified analyses (natural-only and HRT-only). The vast majority of pregnancy-associated signals were preserved across analyses, supporting robustness of the transcriptomic findings.
For all downstream predictive modelling (ExtraTrees and logistic regression) and drug-repositioning analyses, we restricted the gene inputs to this covariate-validated 168-gene panel in order to minimise residual confounding.
Terminology for directionality: To avoid causal implications, we refer to DEGs with lower expression in the pregnancy group as “pregnancy-lower (PL)” genes and those with higher expression as “pregnancy-higher (PH)” genes. Labels such as “risk” or “protective” are not used to denote expression direction alone; any inference regarding adverse or beneficial association is instead based on model coefficients, pathway context, and supporting literature, and should be considered hypothesis-generating rather than causal.
As detailed above, we additionally included cycle type (natural vs. HRT) as a covariate and performed cycle-type–stratified DESeq2 analyses (natural-only and HRT-only), which confirmed that the core pregnancy-associated transcriptomic signals were largely preserved across cycle types (Table 1 ).
Table 1 Characteristics of patients with endometrium samples Variable All N = 147 Pregnancy N = 92 No Pregnancy N = 55 P -value Age (year) 34.61 ± 3.89 34.11 ± 3.94 35.45 ± 3.70 0.042 Body mass index (kg/m 2 ) 22.80 ± 3.39 22.64 ± 3.40 23.63 ± 3.40 0.450 Duration of infertility (years) 3.62 ± 2.16 3.91 ± 2.42 3.46 ± 2.09 0.234 Primary infertility 88 (59.9%) 55 (59.8%) 33 (60.0%) 0.86
Indications for IVF
Hysteromyoma 9 (6.1%) 5 (5.4%) 4 (7.3%) 0.653 Endometriosis 17 (11.6%) 9 (9.8%) 8 (14.5) 0.382 Tubal abnormality 62 (42.2%) 41 (44.6%) 21 (38.2%) 0.445 Male factor 114 (77.6%) 75 (81.5%) 39 (70.9%) 0.136 Ovulation disorders 24 (16.3%) 19 (20.7%) 5 (9.1%) 0.066
Failure history
Number of failed transfer cycles 1.60 ± 1.95 1.28 ± 1.82 2.13 ± 2.07 0.011 Number of failed embryos 2.66 ± 3.41 2.32 ± 3.41 3.22 ± 3.36 0.124 Recurrent implantation failure 37 (25.2%) 16 (17.4%) 21 (38.2%) 0.006 Hormone features FSH (ng/ml) 8.23 ± 3.97 7.89 ± 2.57 8.80 ± 5.57 0.19 LH (ng/ml) 4.20 ± 2.25 4.28 ± 2.33 4.09 ± 2.13 0.63 PRL (ng/ml) 13.17 ± 6.96 12.52 ± 5.78 14.33 ± 8.64 0.19 AMH (ng/ml) 4.0 ± 3.63 4.0 ± 2.92 4.0 ± 4.62 0.997 Left antral follicle count 5.78 ± 3.34 5.93 ± 3.46 5.53 ± 3.14 0.476 Right antral follicle count 5.96 ± 3.32 6.48 ± 3.60 5.09 ± 2.60 0.008 Endometrial thickness (mm) 9.69 ± 2.29 9.81 ± 2.27 9.47 ± 2.33 0.391
Type of embryo transferred
0.186 Cleavage 61 (41.5%) 42 (45.7%) 19 (34.5%) Blastocyst 86 (58.5%) 50 (54.3%) 36 (65.5%)
Cycle type (n, %)
1
Natural cycle 57 (47.9%) 38 (41.3%) 19 (34.5%) 0.49 HRT 62 (52.1%) 54 (58.7%) 36 (65.5%)
Embryo characteristics
Fresh embryo transfer, n (%) 28 (19.0%) 22 (23.9%) 6 (10.9%) 0.08 Frozen embryo transfer, n (%) 119 (81.0%) 70 (76.1%) 49 (89.1%) — Morphologically good embryo (MGE), n (%) 2 117 (82.4%) 77 (84.6%) 40 (78.4%) 0.26 1 Cycle type information was available for 119 of 147 patients due to incomplete documentation in early cases 2 MGE indicates embryos with high morphological quality (ICM ≥ B and trophectoderm ≥ B according to Gardner’s grading system)
Characteristics of patients with endometrium samples
1 Cycle type information was available for 119 of 147 patients due to incomplete documentation in early cases
2 MGE indicates embryos with high morphological quality (ICM ≥ B and trophectoderm ≥ B according to Gardner’s grading system)
SPSS software (version 25.0) was used for descriptive statistics and for constructing the final pregnancy-prediction and live-birth-prediction models, whereas the PyCaret framework was used as an auxiliary tool for model comparison and gene-importance exploration. Categorical variables were summarized as counts and percentages, and continuous variables were expressed as mean ± standard deviation (SD). A P value < 0.05 was considered statistically significant. The primary outcome was clinical pregnancy and the secondary outcome was live birth. Candidate gene predictors were drawn from the covariate-validated panel of 168 pregnancy-associated DEGs derived from the 180-gene discovery set, together with clinically relevant baseline variables. In SPSS, multivariable logistic regression models were fitted using all available cases to derive parsimonious prediction models for pregnancy and live birth (Tables 2 , 3 , 4 and 5 ). For pregnancy prediction, the final model contained 17 genes and a single clinical variable (history of recurrent implantation failure). For live-birth prediction, the final model included 20 genes and three clinical variables (ovulation disorder, right antral follicle count, and primary infertility). Model coefficients, odds ratios (ORs) with 95% confidence intervals (CIs), Wald statistics and two-sided P values were reported. Classification performance (including overall accuracy and area under the receiver operating characteristic curve [AUC]) represents apparent, within-sample performance and should not be interpreted as external validation.
Table 2 The logistic regression model for clinical pregnancy Variables P -value (Wald test) P -value (Wald test) BPIFB1 404.772 0.014 GALNT5 3309.964 0.027 PKP1 -449.534 0.008 MMP10 7.288 0.007 KRT16 -1536.459 0.019 FUT6 -834.601 0.041 CEACAM7 3013.749 0.001 TMPRSS11B 151122.926 0.056 HS3ST3A1 -930.356 0.007 CH25H 131.958 0.006 SPOCD1 -761.01 0.01 IGLV4-69 64.035 0.003 SLC12A3 -104.461 0.016 CALB2 -7.429 0.005 MUCL1 -453.997 0.045 CEACAM6 -1160.622 0.014 REG1A -733.69 0.009 Recurrent Implantation Failure -352.383 0.011 Constant 546.185 0.012 The model was estimated in n = 125 patients with complete gene-expression and clinical data B denotes the logistic regression coefficient; P values are from Wald tests This internally validated model reflects apparent within-sample performance and has not undergone external validation
The logistic regression model for clinical pregnancy
The model was estimated in n = 125 patients with complete gene-expression and clinical data
B denotes the logistic regression coefficient; P values are from Wald tests
This internally validated model reflects apparent within-sample performance and has not undergone external validation
Table 3 Internal confusion matrix of the logistic regression model for clinical pregnancy ( n = 125) Prediction no-pregnancy Prediction pregnancy Correct% Actual no-pregnancy 47 0 100% Actual pregnancy 0 78 100% Overall 100% The confusion matrix summarizes apparent internal performance in the subset of patients with complete data No external validation was performed; therefore, the 100% accuracy reflects within-sample performance and should be interpreted cautiously
Internal confusion matrix of the logistic regression model for clinical pregnancy ( n = 125)
The confusion matrix summarizes apparent internal performance in the subset of patients with complete data
No external validation was performed; therefore, the 100% accuracy reflects within-sample performance and should be interpreted cautiously
Table 4 The logistic regression model for live birth Variables B (logit coefficient) B (logit coefficient) MAL -59.573 0.058 FAT2 -824.852 0.062 GPR87 -1550.72 0.091 TMPRSS11E 1706.511 0.076 SERPINB2 -7545.215 0.087 CHAC1 -635.93 0.098 FUT6 -3454.399 0.084 MTND1P23 -0.536 0.088 CPA4 -217.474 0.011 ADRB2 -890.797 0.086 DAZ4 -37048.875 0.025 LHX1 21793.482 0.067 RPPH1 3.699 0.083 IGLV4-69 2739.932 0.088 DAZ3 93092.499 0.027 CALB2 -27.992 0.07 FCER2 -4229.768 0.08 PITX1 -195.014 0.101 SFRP5 306.88 0.083 POTEH 9739.826 0.085 Ovulation disorder (Yes vs. No) 985.427 0.101 Right antral follicles 190.685 0.087 Primary infertility (Yes vs. No) 887.768 0.089 Constant 505.034 0.09 Clinical predictors were coded as binary variables (Yes vs. No) except for right antral follicles, which was modeled as a continuous variable The model was fitted in n = 124 patients with complete outcome and predictor data P values are from Wald tests and the model has not been externally validated
The logistic regression model for live birth
Clinical predictors were coded as binary variables (Yes vs. No) except for right antral follicles, which was modeled as a continuous variable
The model was fitted in n = 124 patients with complete outcome and predictor data
P values are from Wald tests and the model has not been externally validated
Table 5 Internal confusion matrix of the logistic regression model for live birth ( n = 124) Prediction no live birth Prediction live birth Correct% Actual no live birth 56 0 100% Actual live birth 0 68 100% Overall 100% The confusion matrix summarizes apparent internal performance in the subset of patients with complete data The 100% accuracy reflects within-sample performance and may overestimate performance in external populations
Internal confusion matrix of the logistic regression model for live birth ( n = 124)
The confusion matrix summarizes apparent internal performance in the subset of patients with complete data
The 100% accuracy reflects within-sample performance and may overestimate performance in external populations
In parallel, exploratory machine-learning analyses were performed using PyCaret to compare algorithms based solely on the covariate-validated 168-gene panel, without clinical features, and to obtain model-agnostic rankings of gene importance. For this purpose, the dataset was randomly split into a training set (70%) and a testing set (30%) by stratified sampling on the pregnancy outcome to preserve class balance. The random seed was fixed to ensure reproducibility. Multiple classification algorithms (e.g., logistic regression, random forest, extra trees, gradient boosting) were trained on the 70% training subset, and their performance on the independent 30% testing subset was evaluated using accuracy and AUC. The extra trees classifier achieved the best apparent performance and was therefore used to derive variable-importance scores for pregnancy and live-birth prediction (Fig. 2 C–D). These PyCaret-based analyses were exploratory and served to support feature prioritization and subsequent drug-repositioning analyses, rather than to define the primary clinical prediction models.
Fig. 2 Interaction networks of DEGs and important features in ExtraTrees classifier models. A Genetic interaction network of pregnancy-associated DEGs generated using GeneMANIA. B Physical interaction network of pregnancy-associated DEGs generated using GeneMANIA. C For clinical pregnancy, an ExtraTrees classifier using only the covariate-validated 168-gene panel (without clinical features) outperformed other algorithms. The top model-contributing genes included PTPN5, FAT2, PKP1, SLC12A3, KRT5, IGHD, FUT6, IL1RN, MAL, and CHAC1. The lower panel shows the corresponding confusion matrix in the 30% test set. D For live birth, the ExtraTrees classifier using the same 168-gene panel showed the best internal performance. The most important features included MTND1P23, FLRT3, SLC12A3, S100A2, TFF3, MAL, IL8, TACR1, CHL1 and CHAC1 with the corresponding test-set confusion matrix shown below
Interaction networks of DEGs and important features in ExtraTrees classifier models. A Genetic interaction network of pregnancy-associated DEGs generated using GeneMANIA. B Physical interaction network of pregnancy-associated DEGs generated using GeneMANIA. C For clinical pregnancy, an ExtraTrees classifier using only the covariate-validated 168-gene panel (without clinical features) outperformed other algorithms. The top model-contributing genes included PTPN5, FAT2, PKP1, SLC12A3, KRT5, IGHD, FUT6, IL1RN, MAL, and CHAC1. The lower panel shows the corresponding confusion matrix in the 30% test set. D For live birth, the ExtraTrees classifier using the same 168-gene panel showed the best internal performance. The most important features included MTND1P23, FLRT3, SLC12A3, S100A2, TFF3, MAL, IL8, TACR1, CHL1 and CHAC1 with the corresponding test-set confusion matrix shown below
Drug-repositioning analysis was performed as an exploratory, hypothesis-generating step using the Comparative Toxicogenomics Database (CTD, https://ctdbase.org/ ). We constructed gene query sets from the covariate-validated 168-gene panel as follows: (i) the top 10 pregnancy-lower (PL) and top 10 pregnancy-higher (PH) DEGs (pregnancy vs. no-pregnancy; FDR 1); (ii) model-contributing genes identified by the best-performing machine-learning model (PyCaret); and (iii) hub genes from genetic/physical interaction networks. For each set, we retrieved compounds reported in CTD to downregulate PL candidates or upregulate PH candidates. Compounds were then aggregated by multi-target coverage. We excluded agents that were clearly toxic, not for human use, or known teratogens. The resulting list was considered hypothesis-generating and reported in the Results.
Results
The clinical characteristics of 147 patients with endometrial samples are summarized in Table 1 . Age was significantly lower in the pregnancy group compared with the no-pregnancy group ( P = 0.042), whereas BMI and other baseline characteristics showed no significant differences between groups. Cycle type distribution (natural vs. artificial) was comparable between pregnancy and no-pregnancy groups ( P = 0.49), suggesting that endometrial preparation regimen was unlikely to confound the observed transcriptomic differences. As shown in Table 1 , there were 147 patients who received IVF-ET, including 92 cases with a pregnancy outcome and 55 without pregnancy. The average age was 34.61 ± 3.89 years. The infertility duration was 3.62 ± 2.16 years. The main indications for IVF included male factor infertility (77.6%), tubal abnormality (42.2%), ovulation disorders (16.3%), endometriosis (11.6%), and uterine myoma (6.1%). The mean number of previous failed transfer cycles was 1.60 ± 1.95, and the mean number of previously failed embryos was 2.66 ± 3.41. There were 37 (25.2%) cases that had repeated implantation failure. In the ultrasound imaging, the antral follicle count on the left side was 5.78 ± 3.34, and that on the right side was 5.96 ± 3.32. The average endometrial thickness was 9.69 ± 2.29 mm. Besides, there were 27 (18.4%) cases receiving fresh ET and 120 (81.6%) cases receiving frozen ET. Finally, 61 (41.5%) cases received embryos at the cleavage stage, and the other 86 cases received the blastocyst embryos. Embryo stage (cleavage vs. blastocyst) and embryo morphological quality (MGE) were comparable between pregnancy and no-pregnancy groups ( p = 0.186 and P = 0.26, respectively; Table 1 ), indicating that embryo development stage and quality were unlikely to confound the observed endometrial transcriptomic differences. Between the pregnancy and no-pregnancy groups, there were significant differences in age ( P < 0.05), number of failed transfer cycles ( P < 0.05), presence of recurrent implantation failure ( P < 0.01), and the antral follicle count on the right side ( P < 0.01).
Based on the criteria of FDR 1 in the primary (unadjusted) DESeq2 analysis, a total of 180 pregnancy-associated DEGs were identified between the pregnancy and no-pregnancy groups, including 131 pregnancy-lower genes and 49 pregnancy-higher genes. These DEGs are displayed in the volcano plot and heatmap in Fig. 1 A and B. GO and KEGG enrichment analyses based on this 180-gene discovery panel identified the top significantly enriched biological processes and pathways (Fig. 1 C, D).
To assess whether baseline clinical differences could explain these transcriptomic changes, we subsequently performed a covariate-adjusted DESeq2 analysis including age, RIF status, and right-side AFC as covariates (age + RIF_status + rightAFC + outcome). After adjustment, 168 of the 180 DEGs (93.3%) remained significant with consistent direction, indicating that the transcriptomic differences were not solely driven by these clinical variations. Further adjustment for cycle type and cycle-type–stratified analyses (natural-only and HRT-only) showed similar patterns, supporting the robustness of the core pregnancy-associated signals. Unless otherwise specified, all downstream predictive modelling and drug-repositioning analyses were restricted to this 168-gene panel.
Next, we explored potential hub genes within the interaction networks of these DEGs using the GeneMANIA online tool. In the genetic interaction network (Fig. 2 A), hub genes included TMEM108, DSC3, TRIM29, PRDM16, CHL1, DDIT4L, ADRB2, and EYA4. In the physical interaction network (Fig. 2 B), the key nodes were CXCL6, CXCL5, IL8, KRT17, CXCL3, MMP7, CXCL1, and KRT5.
Using the covariate-validated panel of 168 pregnancy-associated differentially expressed genes (DEGs) derived from the 180-gene discovery set, we developed internally validated logistic regression models for predicting pregnancy and live-birth outcomes by combining gene expression and clinical features. Candidate predictors were drawn from the 168-gene panel together with baseline variables that differed between the pregnancy and no-pregnancy groups. For pregnancy prediction, stepwise selection yielded a parsimonious final model including the expression of 17 genes (BPIFB1, GALNT5, PKP1, MMP10, KRT16, FUT6, CEACAM7, TMPRSS11B, HS3ST3A1, CH25H, SPOCD1, IGLV4-69, SLC12A3, CALB2, MUCL1, CEACAM6 and REG1A) and a single clinical variable, history of recurrent implantation failure (RIF) (Table 2 ). RIF was defined as the failure to achieve a clinical pregnancy after two or more embryo-transfer cycles with good-quality embryos, consistent with the definition used in the Methods section. As shown in Table 3 , this model achieved an apparent 100% within-sample classification accuracy in 125 cases with complete data (47 failures and 78 successes), which likely reflects overfitting within the small dataset and therefore should be interpreted cautiously.
Similarly, starting from the same 168-gene panel and clinically relevant variables, another logistic regression model was developed for live-birth prediction. The final model incorporated the expression of 20 genes together with three clinical features—ovulation disorder, right antral follicle count, and primary infertility (Table 4 ). This live-birth model also achieved an apparent 100% within-sample classification accuracy across 124 available cases (Table 5 ). This perfect performance represents within-sample accuracy rather than external validation and may again reflect overfitting in a limited dataset.
In parallel, to explore gene-level contributions independent of clinical variables and to support downstream drug-repositioning analyses, we used the PyCaret tool to compare multiple machine-learning algorithms based solely on the covariate-validated 168-gene panel (without clinical features). Model performance was evaluated by accuracy and area under the receiver operating characteristic curve (AUC). For pregnancy, the ExtraTrees classifier achieved the highest apparent accuracy (Fig. 2 C), with top-ranked genes including PTPN5, FAT2, PKP1, SLC12A3, KRT5, IGHD, FUT6, IL1RN, MAL, and CHAC1. For live birth, the same algorithm showed the best internal performance, and the most informative genes included MTND1P23, FLRT3, SLC12A3, S100A2, TFF3, MAL, IL8, TACR1, CHL1 and CHAC1(Fig. 2 D).
Given the small cohort size and single-center design, these internally validated models should be regarded as proof-of-concept tools rather than ready-for-clinical-use predictors. External, multi-center validation will be necessary to confirm their robustness and generalizability.
Lastly, the important genes associated with IVF-ET outcomes (clinical pregnancy and live birth) were categorized as pregnancy-lower (PL) and pregnancy-higher (PH) candidates according to their expression direction between the pregnancy and no-pregnancy groups, based on the covariate-validated 168-gene panel (Fig. 3 A). The top pregnancy-lower genes (lower expression in the pregnancy group) were considered candidates potentially associated with unfavorable implantation environments. In addition, important genes from the extra-trees classifier models that were significantly pregnancy-lower, as well as pregnancy-lower hub genes identified in the interaction networks, were included in this candidate list. These candidate genes were queried in the Comparative Toxicogenomics Database (CTD) to identify compounds reported to downregulate their expression. The compounds with the greatest number of gene targets are presented in Fig. 3 B. After excluding clearly toxic agents, non-human drugs, and known teratogens, four compounds—cyclosporine, acetaminophen, tretinoin, and estradiol—emerged as potentially beneficial candidates for improving IVF-ET outcomes. We next examined pregnancy-higher genes (Fig. 3 A, lower), including the top increased genes, those identified as significant contributors in the extra-trees classifier models (none), and hub genes with higher expression in the pregnancy group (none). Using the CTD database, we searched for compounds reported to upregulate these pregnancy-higher genes. However, all identified compounds affected only single targets, and no multi-target compounds were found. Therefore, these four compounds may potentially contribute to improved IVF-ET outcomes by downregulating pregnancy-lower candidate genes associated with unfavorable implantation environments. These findings are associative and hypothesis-generating only and require experimental validation before any clinical consideration.
Fig. 3 Potential drugs for improving IVF-ET outcomes. A Important genes associated with IVF-ET outcomes (clinical pregnancy and live birth) were categorized as pregnancy-lower (PL) and pregnancy-higher (PH) candidates based on the covariate-validated 168-gene panel. The upper panel summarises PL candidates compiled from: (i) the top pregnancy-lower DEGs (pregnancy vs. no-pregnancy), (ii) genes that were both pregnancy-lower and highly ranked in the ExtraTrees models, and (iii) pregnancy-lower hub genes in the interaction networks. The lower panel summarises PH candidates from the top pregnancy-higher DEGs; no genes from the ExtraTrees models or hub analyses met the PH criteria. B PL and PH candidates were queried in the CTD database to identify compounds reported to downregulate PL genes or upregulate PH genes. After excluding clearly toxic compounds, non-human agents, and known teratogens, four compounds—cyclosporine, acetaminophen, tretinoin, and estradiol—emerged as multi-target candidates suppressing multiple PL genes. All compounds affecting PH genes had only single targets, and no multi-target PH-enhancing drugs were identified. These findings are associative and hypothesis-generating and do not imply causality
Potential drugs for improving IVF-ET outcomes. A Important genes associated with IVF-ET outcomes (clinical pregnancy and live birth) were categorized as pregnancy-lower (PL) and pregnancy-higher (PH) candidates based on the covariate-validated 168-gene panel. The upper panel summarises PL candidates compiled from: (i) the top pregnancy-lower DEGs (pregnancy vs. no-pregnancy), (ii) genes that were both pregnancy-lower and highly ranked in the ExtraTrees models, and (iii) pregnancy-lower hub genes in the interaction networks. The lower panel summarises PH candidates from the top pregnancy-higher DEGs; no genes from the ExtraTrees models or hub analyses met the PH criteria. B PL and PH candidates were queried in the CTD database to identify compounds reported to downregulate PL genes or upregulate PH genes. After excluding clearly toxic compounds, non-human agents, and known teratogens, four compounds—cyclosporine, acetaminophen, tretinoin, and estradiol—emerged as multi-target candidates suppressing multiple PL genes. All compounds affecting PH genes had only single targets, and no multi-target PH-enhancing drugs were identified. These findings are associative and hypothesis-generating and do not imply causality
Conclusion
In summary, our study identified an initial discovery panel of 180 differentially expressed genes (DEGs) associated with clinical pregnancy outcomes in IVF-ET patients; 168 of these remained significant after covariate-adjusted analyses and formed the covariate-validated gene panel that we used to construct two internally validated predictive models that achieved satisfactory apparent accuracy for pregnancy and live-birth prediction. Moreover, using this 168-gene panel, our study further identified several novel candidate drugs (cyclosporine, acetaminophen, tretinoin, and estradiol) that may have potential to improve IVF-ET outcomes, although these findings are hypothesis-generating and require functional validation. These findings highlight the potential of integrating transcriptomic and clinical features to refine reproductive outcome prediction and guide personalized therapeutic strategies, while emphasising that external validation in larger multi-center cohorts remains essential before any clinical application.
Discussion
The major findings of this study are: (1) we discovered a panel of 180 genes associated with the clinical pregnancy outcome, 168 of which remained significant after covariate-adjusted analyses and thus formed a robust 168-gene panel used for all downstream modelling and drug-repositioning analyses; and (2) we developed two internally validated models that achieved apparent 100% accuracy in internal validation for predicting pregnancy and live birth by combining key genes and clinical features. These internally validated models are intended as proof-of-concept tools, illustrating the potential predictive value of endometrial gene expression when integrated with clinical parameters. External validation in independent cohorts will be necessary to confirm generalizability and any clinical applicability.
We found for the first time that, based on the covariate-validated 168-gene panel, only one historical indicator—recurrent implantation failure (RIF)—was required to predict whether the current ET would result in a successful pregnancy. RIF refers to the failure to achieve a clinical pregnancy after two or more embryo-transfer cycles with good-quality embryos (Zeng et al. 2022 ). The mechanisms underlying RIF remain incompletely understood (Jiang et al. 2017 , Shen et al. 2022 , Fu et al. 2020 , Dong et al. 2023 ), and a history of RIF is indeed a risk factor that increases the likelihood of implantation failure or pregnancy loss (Ozer et al. 2023 , Pagidas et al. 2008 ). Similarly, for live birth prediction, three clinical features (ovulation disorder, right antral follicle count, and primary infertility) were sufficient to achieve apparent 100% accuracy in internal validation when combined with the covariate-validated 168-gene expression panel. A noteworthy finding is that patients with ovulation disorder appeared to benefit more from IVF-ET than those with other infertility causes. This result is consistent with the univariate analyses (Chi-square analysis in Table 1 ). Therefore, our results suggest that IVF-ET may represent a preferred assisted-reproduction option for infertile patients with ovulatory disorders. Ovulation disorders are associated with higher risks of pregnancy and neonatal complications (Wang et al. 2021 , Palomba et al. 2015 , Grigorescu et al. 2014 , Stern et al. 2015 ). Our findings, however, suggest that ovulation disorders may be more amenable to resolution by IVF-ET than by improving the endometrial environment.
The antral follicle count (AFC) is a well-known indicator of ovarian reserve, and for patients receiving IVF-ET, it is a significant predictor of reproductive outcomes, including successful pregnancy and live birth (Chen et al. 2020 , Maseelall et al. 2009 , Bonilla-Musoles et al. 2012 , Chang et al. 1998 , Tadros et al. 2016 , Tsakos et al. 2014 , Zhou et al. 2020 ). Our study reconfirmed the importance of AFC as a predictor. Interestingly, we found that unilateral AFC (right side) rather than bilateral AFC was more informative. This asymmetry has been described previously, but differences in prognosis have rarely been reported. A cross-sectional study investigating side differences in AFC and ovarian volume found that the right ovary contained 8.1% more antral follicles and had 10.7% larger volume compared with the left (Korsholm et al. 2017 ). Similarly, a right-left comparison in PCOS diagnosis revealed mean AFC differences of 0.24 in the overall population, 0.14 in PCOS patients, and 0.42 in controls (Köninger et al. 2014 ). Another study including 6,617 ultrasound records also found that the number of antral follicles in the right ovary was significantly higher than in the left, in both PCOS and normal populations (AlSerri et al. 2014 ). Differently, we discovered that focusing solely on the right AFC could successfully predict live-birth rates in IVF-ET. Moreover, we observed that primary rather than secondary infertility appeared more likely to benefit from IVF-ET. However, this observation requires further confirmation because of the modest statistical significance and limited sample size. It is also inconsistent with some published studies—for instance, a Chinese study reported that primary infertility is associated with a lower fertilization proportion in IVF-ET (Xia et al. 2020 ). To date, few studies have directly compared IVF-ET outcomes between primary and secondary infertility types. Recently, a study investigating factors affecting cumulative live-birth rates among PCOS patients receiving IVF-ET found no significant difference between primary and secondary infertility (Li et al. 2023 ).
Another novelty of this study is that we developed a discovery panel of 180 DEGs by transcriptomic analysis of endometrial samples. Covariate-adjusted analyses confirmed that 168 of these remained significant and were therefore used as a robust gene panel for subsequent modelling and drug-repositioning analyses. This panel not only assists in predicting pregnancy outcomes but also provides numerous potential targets for exploring the molecular mechanisms by which the endometrial environment influences pregnancy. Because several baseline characteristics (age, RIF status, and right-side AFC) differed between the pregnancy and non-pregnancy groups, we performed a covariate-adjusted differential expression analysis in DESeq2. Most DEGs remained significant and showed consistent expression direction after adjustment, suggesting that the observed endometrial transcriptomic differences were not solely attributable to these clinical variations but likely reflect intrinsic biological differences in endometrial receptivity. Furthermore, given the well-known hormonal divergence between natural and HRT cycles, we also examined cycle type as a potential confounder. Covariate-adjusted and cycle-type–stratified analyses confirmed that the core transcriptomic signals were not solely driven by cycle-type composition, supporting an intrinsic endometrial component to receptivity. For example, MAL was significantly decreased in the pregnancy group, identified as a risk factor in both the logistic and extra-trees models. The role of MAL in endometrial receptivity remains poorly understood. DNA methylation of MAL has been implicated in endometrial cancer development (Strooper et al. 2014 ), but its influence on implantation tolerance has not been reported. Similarly, CXCL5, CXCL6, and KRT17 were not only hub genes in the interaction networks but also among the top decreased DEGs, suggesting they may be key risk genes impairing implantation. CXCL5 and CXCL6, both CXC chemokines, could theoretically exert adverse effects by increasing inflammatory levels within the endometrial immune microenvironment. CXCL5-CXCR2 signaling has been recognized as a senescence-associated secretory phenotype in preimplantation embryos; CXCL5 treatment suppressed trophoblast outgrowth in young mouse blastocysts, whereas suppression of CXCL5-CXCR2 signaling improved implantation in aging embryos (Kawagoe et al. 2020 ). However, direct evidence remains limited. Some studies have found no significant role of CXCL5 or CXCL6 in spontaneous abortion (Spathakis et al. 2023 ), while others reported lower CXCL5 levels in villous tissue from recurrent spontaneous abortion patients compared with controls (Zhang et al. 2021 ). CXCL6 acts as a neutrophil chemoattractant, detectable in amniotic fluid, with concentrations increasing with gestational age but unrelated to term parturition (Mittal et al. 2008 ). KRT17 (Keratin 17) has been identified as a negative prognostic biomarker in high-grade endometrial carcinoma, with high expression implying poor survival (Bai et al. 2019 ). To date, however, the effects of CXCL5, CXCL6, and KRT17 on endometrial physiology during IVF-ET remain unexplored and warrant further investigation. Importantly, expression direction alone does not define risk. In this study, we therefore interpret pregnancy-lower and pregnancy-higher (PL/PH) signals in conjunction with model coefficients, pathway context, and prior literature, and we present them as hypothesis-generating associations rather than causal determinants. Moreover, FCER2, SLC12A3, CEACAM6, and GPR87 were repeatedly included in prognostic models (in both logistic regression and extra-trees classifier analyses). Their associations with endometrial biology or pregnancy outcomes have not yet been reported. Together, we have identified several novel endometrial gene targets that may contribute to improved IVF-ET outcomes.
Finally, by prioritizing pregnancy-lower (PL) candidates alongside pathway and literature context, we identified four candidate drugs—cyclosporine, acetaminophen, tretinoin, and estradiol—as potentially beneficial for IVF-ET patients. The available evidence supports this hypothesis. Cyclosporine is a potent immunosuppressant with wide clinical applications. Prior studies have shown that cyclosporine A improves pregnancy outcomes in women with recurrent pregnancy loss and elevated Th1/Th2 ratios (Azizi et al. 2019 ), and a 2022 retrospective study in China confirmed its benefit in patients with repeated implantation failure without increasing obstetric or pediatric complications (Cheng et al. 2022 ). Acetaminophen is generally safe in pregnancy, although its impact on endometrial function remains underexplored (Scialli et al. 2010 , Patel et al. 2022 , Castro et al. 2022 ). Tretinoin (retinoic acid) metabolism plays an essential role in implantation and decidualization of endometrial stromal cells (Han et al. 2010 , Ozaki et al. 2017 , Hohn and Denker 2002 , Osteen et al. 2002 , Rajakumar et al. 2020 , Sidell et al. 2010 , Wu et al. 2011 , Zheng et al. 2000 ). Additionally, several studies have demonstrated that increased serum estradiol levels are associated with higher pregnancy rates, especially in younger women (Wei et al. 2023 ), and estrogen supplementation may further improve IVF-ET outcomes (Jung and Roh 2000 , Lewin et al. 1994 , Propst et al. 2001 , Unfer et al. 2004 ). However, large-scale data on estrogen or cyclosporine use in IVF-ET remain limited, and no study has assessed the potential of acetaminophen or tretinoin in this context. Collectively, we propose these four agents as promising candidates that warrant further clinical validation. These drug suggestions are exploratory, guided by PL/PH directionality and existing literature, and should not be construed as clinical recommendations.
Still, this study has several limitations. Firstly, although the two logistic regression models achieved apparent 100% accuracy, this represents internal performance only and lacks external validation. Given the modest sample size and a high predictor-to-sample ratio, overfitting is likely; thus, reported accuracy should be interpreted as within-sample performance rather than true predictive ability on independent data. Despite stratified 70/30 train–test splits, overall efficacy remains constrained by sample size and the single-center design, which may limit generalizability. In addition, although key baseline variables (age, RIF status, and right-side AFC) were adjusted for in sensitivity analyses, residual confounding by unmeasured factors and by cycle-type–related aspects (e.g., specific HRT regimens or dosing) cannot be fully excluded. Although embryo stage and morphological quality were comparable between groups, these metrics were not explicitly modeled as covariates due to sample-size limitations; future multi-center studies should incorporate detailed embryo-quality metrics (ICM and trophectoderm grading). Furthermore, while we identified potential drugs from transcriptomic signals, these are hypothesis-generating only; functional validation (in vitro/in vivo) and prospective clinical evaluation are required before any clinical consideration. Larger, multi-center cohorts are needed for external validation and to confirm model robustness across diverse settings.
Introduction
In vitro fertilization (IVF) with embryo transfer (ET) is increasingly the standard approach. Nonetheless, the current outcome and live birth are not fully satisfactory (Cheng et al. 2023 , Gao et al. 2019 , Liang et al. 2023 , Tan et al. 2023 , Wei et al. 2023 , Zhang et al. 2023 ). Previous studies have explored various risk and protective factors that may influence IVF-ET outcomes and developed prediction models. However, most of these models rely on demographic and clinical factors, such as obstetrical history, BMI, physical examination, and infertility work-up (Cheng et al. 2023 , Tan et al. 2023 ). Recently, new clinical features like sperm DNA fragmentation, chromosome polymorphism, and sleep quality have gained attention (Zhang et al. 2023 , Zhu et al. 2022 , Liu et al. 2023 ). Despite these advancements, the overall predictive value of these factors is limited, and more indicative information, in particular at the genetic level, is urgently needed. A successful pregnancy depends not only on good embryo quality but also on satisfactory endometrial receptivity. In particular, the latter is crucial for the embryo attachment and growth during implantation. Poor endometrial receptivity can account for a majority of ET failures (Foster et al. 2009 ). Previously, the endometrial receptivity array has been applied in the clinical practice for individualized determination, but its efficacy still needs to be improved (Rubin et al. 2023 , Robert 2020 , Bassil et al. 2018 , Tan et al. 2018 ). Thus, to date, there remains a need to explore more genes associated with IVF-ET outcomes through higher-throughput assays in comparison with the endometrial receptivity array (e.g., transcriptome sequencing to detect RNA expression levels). Additionally, more novel targets based on endometrial gene expression need to be explored, which help to understand the mechanisms of implantation failure and improve the outcomes. Theoretically, enough gene expression data, combined with clinical profiles, would significantly improve model performance, and shed light on those more critical features. This study aimed to discover more differential genes linked with the IVF‑ET outcomes and develop new models in combination of a panel of genes and clinical profiles. By conducting a comprehensive analysis of gene expression data and clinical profiles, we hope to improve model performance and identify critical factors that could significantly impact IVF-ET outcomes. Our study may contribute to more effective and personalized treatments for IVF-ET patients.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.