Results
In this study, 18,750 patients were finally enrolled based on the inclusion and exclusion criteria. The data were randomly divided into a training set ( n = 13,125) and a validation set ( n = 5,625) in a 7:3 ratio. Comparisons of risk warning indicators between the two groups showed no significant differences in the study variables ( P > 0.05), indicating good homogeneity between the two datasets (Table 1 ; Fig. 1 ). Basic characteristics and differences of the training set are provided in Supplementary Table 2. Fig. 1 The flow diagram of developing and validating the prediction model
The flow diagram of developing and validating the prediction model
Table 1 Comparison of characteristics between the training and validation groups Variables Total ( n = 18750) Validation ( n = 5625) Training ( n = 13125)
P
Age (years) 62.00 (47.00, 75.00) 62.00 (47.00, 75.00) 63.00 (48.00, 75.00) 0.752 Height (cm) 163.00 (157.00, 165.00) 163.00 (157.00, 168.00) 163.00 (157.00, 165.00) 0.244 Weight (kg) 72.00 (60.00, 87.23) 72.00 (60.19, 88.22) 71.85 (60.00, 86.91) 0.085 BMI 27.68 (23.48, 33.42) 27.77 (23.46, 33.76) 27.64 (23.48, 33.30) 0.134 Na (mEq/L) 139.00 (137.00, 141.00) 139.00 (137.00, 141.00) 139.00 (137.00, 141.00) 0.341 K (mEq/L) 4.10 (3.80, 4.40) 4.10 (3.80, 4.40) 4.10 (3.80, 4.40) 0.926 Glu (mg/dl) 108.00 (92.00, 139.00) 108.00 (92.00, 139.00) 108.00 (92.00, 138.00) 0.630 HbA1c (%) 5.70 (5.40, 6.20) 5.70 (5.40, 6.20) 5.70 (5.40, 6.20) 0.418 Rbc (m/uL) 4.03 (3.54, 4.43) 4.04 (3.55, 4.43) 4.03 (3.53, 4.43) 0.706 Wbc (K/uL) 8.40 (6.33, 11.40) 8.30 (6.30, 11.40) 8.40 (6.40, 11.40) 0.581 Plt (K/uL) 250.00 (194.00, 314.00) 249.00 (194.00, 315.00) 251.00 (193.00, 314.00) 0.901 Hb (g/dl) 11.90 (10.30, 13.10) 11.90 (10.30, 13.10) 11.90 (10.30, 13.10) 0.748 Ast (IU/L) 24.00 (18.00, 39.00) 24.00 (19.00, 38.00) 24.00 (18.00, 40.00) 0.303 Alt (IU/L) 20.00 (14.00, 32.00) 19.00 (14.00, 32.00) 20.00 (14.00, 32.00) 0.231 Alb (g/dl) 3.90 (3.30, 4.30) 3.90 (3.30, 4.30) 3.90 (3.30, 4.30) 0.914 Ldl (mg/dl) 93.00 (68.00, 121.00) 95.00 (69.00, 121.00) 93.00 (67.00, 121.00) 0.121 Hdl (mg/dl) 51.00 (39.00, 65.00) 51.00 (39.00, 65.00) 51.00 (39.00, 65.00) 0.685 Tc (mg/dl) 175.00 (141.00, 209.00) 176.00 (142.00, 209.00) 174.00 (141.00, 209.00) 0.107 Tg (mg/dl) 117.00 (82.00, 174.00) 118.00 (82.00, 174.00) 117.00 (81.00, 174.00) 0.516 TyG index 8.80 (8.36, 9.32) 8.82 (8.37, 9.31) 8.80 (8.36, 9.32) 0.274 Bun (mg/dl) 15.00 (11.00, 23.00) 15.00 (11.00, 23.00) 15.00 (11.00, 22.00) 0.587 Scr (mg/dl) 0.80 (0.70, 1.00) 0.80 (0.70, 1.00) 0.80 (0.70, 1.00) 0.301 Inr 1.20 (1.10, 1.40) 1.20 (1.10, 1.30) 1.20 (1.10, 1.40) 0.096 Pt (s) 12.90 (11.80, 14.90) 12.90 (11.70, 14.80) 12.90 (11.80, 14.90) 0.144 Ptt (s) 30.00 (26.80, 36.10) 30.00 (26.80, 35.80) 30.10 (26.80, 36.20) 0.425 Race, n(%) 0.200 Asian 381 (2.03) 102 (1.81) 279 (2.13) White 9631 (51.37) 2852 (50.70) 6779 (51.65) Black 1620 (8.64) 481 (8.55) 1139 (8.68) Other 7118 (37.96) 2190 (38.93) 4928 (37.55) Copd, n(%) 0.990 No 15,301 (81.61) 4590 (81.60) 10,711 (81.61) Yes 3449 (18.39) 1035 (18.40) 2414 (18.39) Coronary, n(%) 0.863 No 13,746 (73.31) 4119 (73.23) 9627 (73.35) Yes 5004 (26.69) 1506 (26.77) 3498 (26.65) Diabetes, n(%) 0.870 No 13,219 (70.50) 3961 (70.42) 9258 (70.54) Yes 5531 (29.50) 1664 (29.58) 3867 (29.46) Hypertension, n(%) 0.562 No 12,843 (68.50) 3836 (68.20) 9007 (68.62) Yes 5907 (31.50) 1789 (31.80) 4118 (31.38) Leiomyoma, n(%) 0.114 No 15,124 (80.66) 4498 (79.96) 10,626 (80.96) Yes 3626 (19.34) 1127 (20.04) 2499 (19.04)
Comparison of characteristics between the training and validation groups
LASSO regression with 10-fold cross-validation was used to screen predictive variables from the 30 candidate variables. The analysis was performed using the glmnet package in R. The initial lambda (λ) was set to iterate 100 times. In the LASSO regression model, the optimal penalty coefficient λ was determined by plotting the coefficient changes (Fig. 2 A), where the left dashed line corresponds to the minimum λ (λmin), and the right dashed line corresponds to λmin plus one standard error. This value was selected as the optimal λ for the model. As the penalty coefficient λ increased, the coefficients of the initial variables were gradually compressed, with some coefficients shrinking to zero (Fig. 2 B). This process minimized overfitting and optimized variable selection. To balance model performance and parsimony, the optimal λ was chosen based on the 10-fold cross-validation error at λmin plus one standard error. This resulted in the selection of 20 variables as predictive factors: Race, Coronary, Hypertension, Age, Height, Weight, Na, Rbc, Wbc, Plt, Hb, Glu, Ast, Alb, Hdl, Tc, TyG, Bun, INR, and Ptt. Fig. 2 Using the LASSO regression model for variable selection. A Illustrates the selection process for the optimal penalty coefficient λ in the LASSO model. The optimal λ was determined using 10-fold cross-validation, with the left dashed line indicating the minimum λ (λmin) and the right dashed line indicating λmin plus one standard error, corresponding to the optimal number of selected variables. B Presents the LASSO coefficient curves for the 30 variables, showing how coefficients are penalized and reduced to zero as λ increases, aiding in variable selection
Using the LASSO regression model for variable selection. A Illustrates the selection process for the optimal penalty coefficient λ in the LASSO model. The optimal λ was determined using 10-fold cross-validation, with the left dashed line indicating the minimum λ (λmin) and the right dashed line indicating λmin plus one standard error, corresponding to the optimal number of selected variables. B Presents the LASSO coefficient curves for the 30 variables, showing how coefficients are penalized and reduced to zero as λ increases, aiding in variable selection
Multivariate logistic regression was performed on the selected feature variables (using forward and backward stepwise regression with a variable criterion of P < 0.05) to establish a clinical prediction model. A total of 16 predictors were ultimately identified: Race, Coronary, Hypertension, Age, Height, Weight, Rbc, Wbc, Plt, Hb, Ast, Alb, Tc, TyG, Bun, and INR were identified as risk factors for uterine fibroids. The results showed that compared to Caucasians, individuals of Asian, African, and other races had a higher risk of developing uterine fibroids, which was statistically significant ( P < 0.05). The presence of cardiovascular disease was associated with a reduced risk of uterine fibroids (OR = 0.0484, 95% CI 0.383–0.613, P < 0.001). Hypertension was associated with an increased risk of uterine fibroids (OR = 1.496, 95% CI 1.260–1.776, P < 0.001). Younger age was associated with an increased risk of uterine fibroids (OR = 0.951, 95% CI 0.947–0.955, P < 0.001). Increased height was associated with an increased risk of uterine fibroids (OR = 1.030, 95% CI 1.021–1.040, P < 0.001). Increased weight was associated with an increased risk of uterine fibroids (OR = 1.011, 95% CI 1.008–1.015, P < 0.001). Elevated Rbc values were associated with an increased risk of uterine fibroids (OR = 1.609, 95% CI 1.343–1.927, P < 0.001). Decreased Wbc values were associated with an increased risk of uterine fibroids (OR = 0.944, 95% CI 0.927–0.961, P < 0.001). Elevated Plt values were associated with an increased risk of uterine fibroids (OR = 1.002, 95% CI 1.002–1.003, P < 0.001). Decreased Hb values were associated with an increased risk of uterine fibroids (OR = 0.809, 95% CI 0.763–0.858, P < 0.001). Decreased Ast values were associated with an increased risk of uterine fibroids (OR = 0.998, 95% CI 0.997–0.999, P < 0.001). Elevated Alb values were associated with an increased risk of uterine fibroids (OR = 2.138, 95% CI 1.894–2.415, P < 0.001). Elevated Tc values were associated with an increased risk of uterine fibroids (OR = 1.003, 95% CI 1.001–1.004, P = 0.002). Decreased TyG values were associated with an increased risk of uterine fibroids (OR = 0.534, 95% CI 0.472–0.605, P < 0.001). Decreased Bun values were associated with an increased risk of uterine fibroids (OR = 0.988, 95% CI 0.979–0.996, P = 0.003). Decreased INR values were associated with an increased risk of uterine fibroids (OR = 0.563, 95% CI 0.463–0.684, P < 0.001). A nomogram model for predicting the probability of uterine fibroid occurrence was developed based on the above risk factors (Table 2 ; Fig. 3 ). The OR values were converted to natural logarithms (log10), and a forest plot of risk factors was generated (Fig. 4 ).
Table 2 Univariate and multivariate logistic regression results Variables Univariate Multivariate β
P
OR (95%CI) β
P
OR (95%CI) Race Asian 1.00 (Reference) 1.00 (Reference) White −1.27
< 0.001
0.28 (0.14 ~ 0.57) −1.07
0.006
0.34 (0.16 ~ 0.73) Black 0.69 0.055 1.99 (0.98 ~ 4.04) 0.30 0.447 1.35 (0.63 ~ 2.90) Other 3.31
< 0.001
27.48 (14.11 ~ 53.51) 3.09
< 0.001
21.88 (10.62 ~ 45.06) Coronary No 1.00 (Reference) 1.00 (Reference) Yes −2.17
< 0.001
0.11 (0.10 ~ 0.14) −0.72
< 0.001
0.48 (0.38 ~ 0.61) Hypertension No 1.00 (Reference) 1.00 (Reference) Yes −0.68
< 0.001
0.51 (0.46 ~ 0.56) 0.40
< 0.001
1.50 (1.26 ~ 1.78) Age −0.07
< 0.001
0.93 (0.93 ~ 0.93) −0.05
< 0.001
0.95 (0.95 ~ 0.96) Height 0.05
< 0.001
1.05 (1.05 ~ 1.06) 0.03
< 0.001
1.03 (1.02 ~ 1.04) Weight 0.01
< 0.001
1.01 (1.01 ~ 1.02) 0.01
< 0.001
1.01 (1.01 ~ 1.01) Na 0.01 0.386 1.01 (0.99 ~ 1.02) Rbc 0.61
< 0.001
1.84 (1.71 ~ 1.97) 0.48
< 0.001
1.61 (1.34 ~ 1.93) Wbc −0.08
< 0.001
0.92 (0.91 ~ 0.93) −0.06
< 0.001
0.94 (0.93 ~ 0.96) Plt 0.01
< 0.001
1.01 (1.01 ~ 1.01) 0.01
< 0.001
1.01 (1.01 ~ 1.01) Hb 0.06
< 0.001
1.06 (1.03 ~ 1.08) −0.21
< 0.001
0.81 (0.76 ~ 0.86) Glu −0.02
< 0.001
0.98 (0.98 ~ 0.99) Ast −0.01
< 0.001
0.99 (0.99 ~ 0.99) −0.01
< 0.001
0.99 (0.99 ~ 0.99) Alb 1.22
< 0.001
3.38 (3.12 ~ 3.67) 0.76
< 0.001
2.14 (1.89 ~ 2.42) Hdl 0.01
< 0.001
1.01 (1.01 ~ 1.02) 0.00 0.074 1.00 (1.00 ~ 1.01) Tc 0.01
< 0.001
1.01 (1.01 ~ 1.01) 0.01
0.002
1.01 (1.01 ~ 1.01) TyG −0.82
< 0.001
0.44 (0.41 ~ 0.47) −0.63
< 0.001
0.53 (0.47 ~ 0.61) Bun −0.10
< 0.001
0.90 (0.89 ~ 0.91) −0.01
0.003
0.99 (0.98 ~ 0.99) Inr −1.81
< 0.001
0.16 (0.13 ~ 0.20) −0.57
< 0.001
0.56 (0.46 ~ 0.68) Ptt −0.02
< 0.001
0.98 (0.97 ~ 0.98)
Univariate and multivariate logistic regression results
Fig. 3 Nomogram predicting the risk of uterine fibroids in female patients
Nomogram predicting the risk of uterine fibroids in female patients
Fig. 4 Forest plot of the association between risk factors and uterine fibroids in female patients
Forest plot of the association between risk factors and uterine fibroids in female patients
Calibration curves of the nomogram in the training set showed good consistency between predicted and observed outcomes (Fig. 5 A), with the Hosmer-Lemeshow test indicating no significant differences ( P = 1), suggesting a good fit. The prediction performance of the nomogram was assessed using the receiver operating characteristic (ROC) curve, which yielded an area under the curve (AUC) of 0.95 (95% CI, 0.95–0.96), with a sensitivity of 0.88, specificity of 0.92, prediction accuracy of 0.89 (95% CI, 0.88–0.89), positive predictive value (PPV) of 0.98, and negative predictive value (NPV) of 0.65 (Table 3 ; Fig. 5 C). In the validation group, the calibration curves also demonstrated good consistency, and the Hosmer-Lemeshow test showed no significant differences ( P > 0.05), confirming high calibration. The AUC was 0.95 (95% CI, 0.94–0.96), with a sensitivity of 0.88, specificity of 0.93, prediction accuracy of 0.89 (95% CI, 0.88–0.90), PPV of 0.98, and NPV of 0.66 (Table 3 ; Fig. 5 C). There was no statistically significant difference in AUC between the training and validation sets ( P > 0.05) (Fig. 5 B). Fig. 5 Discrimination and calibration of the nomogram prediction model in the training and validation groups. A Calibration Curve in the Training Group; ( B) Calibration Curve in the Validation Group; ( C) ROC Curves in the Training and Validation Groups
Discrimination and calibration of the nomogram prediction model in the training and validation groups. A Calibration Curve in the Training Group; ( B) Calibration Curve in the Validation Group; ( C) ROC Curves in the Training and Validation Groups
Table 3 Diagnostic performance of the nomogram model for uterine fibroids in female patients in the training and validation groups Data AUC (95%CI) Accuracy (95%CI) Sensitivity (95%CI) Specificity (95%CI) PPV (95%CI) NPV (95%CI) cut off Training 0.95 (0.95–0.96) 0.89 (0.88–0.89) 0.88 (0.88–0.89) 0.92 (0.91–0.93) 0.98 (0.98–0.98) 0.65 (0.63–0.66) 0.217 Validation 0.95 (0.94–0.96) 0.89 (0.88–0.90) 0.88 (0.87–0.89) 0.93 (0.91–0.94) 0.98 (0.98–0.98) 0.66 (0.63–0.68) 0.217
Diagnostic performance of the nomogram model for uterine fibroids in female patients in the training and validation groups
Decision curve analysis (DCA) is a method used to evaluate the clinical utility of a diagnostic test by quantifying the net benefit across different threshold probabilities. In this study, DCA was applied to assess the clinical utility of the nomogram model. Compared to the thresholds of “no intervention” and “intervention for all,” the nomogram demonstrated higher clinical net benefits in both the training and validation groups (Fig. 6 A and D). Clinical impact curves further indicated that the number of high-risk patients predicted to experience uterine fibroids and the actual number of uterine fibroid events tended to align closely with the risk threshold. This suggests that the model has significant predictive power and good clinical applications (Fig. 6 B and C). Fig. 6 Clinical applications assessment of the nomogram prediction model in the training and validation groups. A DCA in the Training Group; ( B ) DCA in the Validation Group; ( C ): Clinical Impact Curve in the Training Group; ( D) : Clinical Impact Curve in the Validation Group
Clinical applications assessment of the nomogram prediction model in the training and validation groups. A DCA in the Training Group; ( B ) DCA in the Validation Group; ( C ): Clinical Impact Curve in the Training Group; ( D) : Clinical Impact Curve in the Validation Group
Materials
This study employed a retrospective cohort design, utilizing data from the Medical Information Mart for Intensive Care (MIMIC-IV3.0) database for female patients admitted between 2008 and 2022. The MIMIC-IV3.0 database is a publicly accessible resource approved by the Institutional Review Boards of Beth Israel Deaconess Medical Center and the Massachusetts Institute of Technology. Data extraction was performed by trained personnel who obtained certification for database access and download (ID: 49746917). This study adheres to the principles of the Declaration of Helsinki. Inclusion Criteria: Female patients diagnosed with uterine fibroids based on the International Classification of Diseases, Ninth and Tenth Revisions (ICD-9 and ICD-10) codes. Exclusion Criteria: (1) Age < 18 years; (2) Hospitalization duration < 24 h; (3) Missing data on weight, height, triglycerides (TG), or fasting blood glucose (FBG); (4) Multiple admissions, with only the first admission data included; (5) Patients with a history of hysterectomy or prior fibroid removal procedures. Variables Collected: Demographic Information: Age, weight, height, and race. Comorbidities: Hypertension, diabetes, cardiovascular disease, and uterine fibroids. Laboratory Variables: White blood cell count, red blood cell count, hemoglobin, platelet count, alanine aminotransferase, aspartate aminotransferase, albumin, high-density lipoprotein, low-density lipoprotein, cholesterol, triglycerides, serum creatinine, blood urea nitrogen, glucose, potassium, sodium, activated partial thromboplastin time (APTT), international normalized ratio (INR), and prothrombin time (PT). The value of the TyG index was calculated as ln [fasting TG (mg/dl) × FBG (mg/dl)/2].
Data extraction was performed using PostgreSQL (version 16) and Navicate Premium (version 17), and statistical analysis and visualization were conducted using R (version 4.4.0, www.R-project.org/ ). The R packages used in our study were displayed in Supplementary Table 1. Continuous Variables: Non-normally distributed continuous variables are presented as medians (interquartile ranges, M [Q1, Q3]) and were compared between groups using the Mann-Whitney U test. Categorical Variables: Categorical variables are presented as frequencies and percentages, and group comparisons were performed using the chi-square test. Missing Data Handling: Variables with missing values exceeding 10% were excluded. Variables with missing values between 5% and 10% were imputed using multiple imputation. Variables with missing values less than 5% were replaced with the mean value of the variable. Outliers were managed using the winsorization method, with cutoff points set at the 1 st and 99th percentiles. To address multicollinearity among variables, LASSO regression with 10-fold cross-validation was used to select predictive variables [ 12 ]. Specified the number of iterations (100) for lambda selection in LASSO. A multivariable logistic regression model (using stepwise selection with a significance threshold of P < 0.05) was then constructed based on the selected variables. A nomogram was created to visualize the prediction model. Model Evaluation: Discrimination: Assessed using the receiver operating characteristic (ROC) curve. Calibration: Evaluated using calibration curves. Clinical Utility: Assessed using decision curve analysis (DCA) and clinical impact curves. A two-sided P value of less than 0.05 was considered statistically significant. Patients were randomly divided into a training set (70%) and a validation set (30%) using a fixed random seed (set.seed(123)) to ensure reproducibility. The prediction model for uterine fibroid risk was developed and validated using the training set.
Conclusion
This study demonstrates that Race, Coronary, Hypertension, Age, Height, Weight, Rbc, Wbc, Plt, Hb, Ast, Alb, Tc, TyG, Bun, and INR are significantly associated with the occurrence of uterine fibroids. A nomogram model was developed and validated to predict the risk of uterine fibroids, enabling individualized patient prediction and offering the potential for early risk prediction and disease prevention.
Discussion
This study established a prediction model for the risk of uterine fibroids in women using the MIMIC database. The model identified Race, Coronary, Hypertension, Age, Height, Weight, Rbc, Wbc, Plt, Hb, Ast, Alb, Tc, TyG, Bun, and INR as risk factors. Validation confirmed the model’s reliability, enhancing its practicality for predicting uterine fibroid risk in women. Compared with the PRS-based model by Piekos et al.[ 13 ]. (AUC = 0.60), The AUC of this study to create a predictive model using only clinical indicators can reach 0.95 and does not require genotype data.
Numerous studies have identified hypertension and cardiovascular diseases as significant risk factors for uterine fibroids [ 14 – 17 ]. In a single-center case-control study conducted in Japan [ 18 ], women with hypertension (defined as systolic blood pressure ≥ 140 mmHg, diastolic blood pressure ≥ 90 mmHg, or current use of antihypertensive medications) had nearly five times the risk of developing uterine fibroids compared to women with normal blood pressure. Therefore, the incidence of uterine fibroids is often higher in women with hypertension and cardiovascular diseases. There is a significant difference in the incidence of uterine fibroids among different races, confirming that uterine fibroids are more common in black women than in white women, and the incidence rate in black women is higher than that in other races [ 2 , 19 ], which suggests that there may be a genetic basis for this condition. Some studies have also indicated that age, height, and weight can influence the occurrence of uterine fibroids [ 20 – 22 ]. The incidence of uterine fibroids increases with age. A prospective study on the incidence and growth of uterine fibroids in young African American women showed that the incidence of fibroids increased from 6% in women aged 23–25 years to 13% in those aged 32–35 years as they grew older [ 23 ].
The inclusion of metabolic markers such as the TyG index, total cholesterol (TC), and hemoglobin (Hb) in the model suggests that metabolic abnormalities may promote fibroid growth by influencing estrogen levels or the local microenvironment. This is consistent with previous studies linking obesity and insulin resistance to an increased risk of uterine fibroids. Additionally, the significance of white blood cell (WBC) and platelet (PLT) counts reflects the role of chronic inflammation in fibroid development, supporting the hypothesis that uterine fibroids are associated with the activation of local inflammatory factors such as interleukin-6 (IL-6) and tumor necrosis factor-alpha (TNF-α)[ 24 ]. This study, for the first time, integrates blood biochemical indicators (e.g.: Ast, Alb, INR) with clinical characteristics (e.g.: height, weight) to construct a prediction model, overcoming the limitations of traditional studies that have predominantly focused on obstetric history or imaging features. The TyG (Triglyceride-Glucose) index, a marker of insulin resistance (IR) and metabolic syndrome is closely associated with metabolic disorders such as obesity, dyslipidemia, and hyperglycemia. Metabolic abnormalities, in turn, are linked to obesity, estrogen levels, and IR, indirectly promoting the proliferation of uterine fibroid cells. This study establishes a direct relationship between the TyG index and uterine fibroids. By incorporating TyG as a marker of IR, the model enhances its ability to identify high-risk populations for uterine fibroids. This approach is analogous to the use of metabolic indicators in risk models for cervical cancer [ 25 ].
Similar machine learning approaches have been successfully applied in other gynecological conditions, such as endometriosis and adenomyosis [ 9 ], and could be adapted for risk prediction in those contexts. Integration of this model into AI-based clinical decision support systems could enhance its utility in routine clinical practice, providing clinicians with a tool for early identification of high-risk patients. However, external validation in diverse populations and healthcare settings is necessary to ensure generalizability. Future studies could employ transfer learning techniques to adapt the model to local data sources, improving its performance across different regions.
(1) This study is retrospective and may have residual or unmeasured confounding factors. However, rigorous statistical methods were used to strengthen the validity of the results. (2) The MIMIC database primarily includes critically ill patients, which may introduce selection bias. Additionally, the absence of certain indicators in the database limited the study. (3) Patient data were collected from a single center, leading to significant heterogeneity. Incomplete imaging data were not included in the analysis. Future research will involve collecting comprehensive demographic, laboratory, and imaging data for prospective, multicenter studies.
Introduction
Uterine fibroids (also known as leiomyomas or myomas) are common clonal neoplasms of the uterus, composed of smooth muscle cells and fibroblasts, and are rich in extracellular matrix (ECM). Their development and gene expression are regulated by sex steroid hormones, primarily estrogen and progesterone, in a menstrual cycle-dependent manner [ 1 , 2 ]. According to ultrasound screening studies and pathological data, uterine fibroids are highly prevalent, with up to 80% of women having detectable fibroids by the age of 50. The majority of patients with uterine fibroids are asymptomatic, and the lesions are predominantly benign [ 3 ]. Uterine fibroids exhibit an incidence associated with age and race. Compared to Caucasian women, Black women have a higher risk of developing fibroids at an earlier age and often experience greater disease severity [ 4 – 6 ]. The diagnosis of uterine fibroids is influenced by several factors, as the size, location, and number of fibroids vary among patients, and the symptoms caused by fibroids also differ [ 7 ]. The diagnosis of uterine fibroids is influenced by several factors, including the variability in size, location, and number of fibroids among patients, as well as the diversity of symptoms they cause. Many women do not associate their symptoms with fibroids, leading to potential delays in diagnosis. Additionally, some fibroids are asymptomatic and may remain undetected. To date, identified serum biomarkers lack sufficient predictive value to enhance diagnostic accuracy beyond that provided by physical examination and imaging. If biomarkers could accurately identify women at high risk of uterine fibroids or distinguish those with stable disease from those at higher risk, they might have clinical utility. Furthermore, biomarkers could be useful in differentiating benign uterine fibroids from other pathologies, thereby improving diagnostic precision and clinical decision-making [ 8 ].
In recent years, machine learning-based modeling has been widely applied in obstetrics and gynecology for diagnostic and risk prediction purposes. For instance, studies analyzing elastography values in endometrioma patients alongside clinical findings have demonstrated enhanced predictive performance when using machine learning algorithms [ 9 ]. Moreover, artificial intelligence techniques have shown considerable advantages in improving the accuracy and speed of diagnosis and treatment in uterine fibroid management [ 10 ]. In obstetrics, early-pregnancy serum biomarkers such as PAPP-A and β-hCG have been successfully used in machine learning models to predict the risk of gestational diabetes [ 11 ]. In this context, the model developed in the present study contributes meaningfully and innovatively to the existing literature by enabling early prediction of uterine fibroid risk.
Currently, no prior studies have focused on developing a prediction model for uterine fibroids. To address this gap, we constructed and validated a prediction model for uterine fibroid risk using data from the Medical Information Mart for Intensive Care (MIMIC-IV3.0) database.