Construction and validation of risk prediction model for uterine fibroids: a retrospective cohort study based on MIMIC database.

OA: gold CC-BY-NC-ND-4.0
AI-generated summary by claude@2026-07, 2026-07-13

This study developed and validated a uterine fibroid risk prediction model using 16 predictors identified from the MIMIC database, demonstrating high accuracy and clinical utility.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-13 · read from full text

This retrospective cohort study constructed and validated a uterine fibroid risk prediction model using female patient data from the MIMIC-IV3.0 database (admissions 2008–2022), selecting adults with uterine fibroids defined by ICD-9/10 codes and excluding those with short stays, missing key variables, multiple admissions (keeping the first), or prior hysterectomy/fibroid removal. From 30 candidate demographic, comorbidity, and laboratory variables (including derived TyG index), the authors used LASSO with 10-fold cross-validation for variable selection followed by multivariable logistic regression, then assessed discrimination (ROC), calibration, and clinical utility (decision curve analysis and clinical impact curves) in a 70/30 training/validation split (n=18,750 total). They report selecting 20 predictive factors and finding no significant differences in baseline characteristics between training and validation sets. A major limitation is that the model is built from retrospective hospital admission data without an explicit external validation cohort. Relevance to endometriosis: the paper is included in the corpus via keyword match to machine-learning risk prediction in gynecologic conditions, but it does not explicitly discuss endometriosis or adenomyosis.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

OBJECTIVE: To develop and validate a prediction model for uterine fibroid risk based on clinical characteristics and blood biochemical indicators, and to explore its risk factors and clinical application value. METHODS: This retrospective cohort study utilized data from the Medical Information Mart for Intensive Care (MIMIC-IV3.0) database, including records of female patients admitted between 2008 and 2022. Patient data, including demographic information, vital signs, clinical symptoms, and laboratory test results, were collected. Patients were divided into training and validation groups in a 7:3 ratio. To avoid multicollinearity, the Least Absolute Shrinkage and Selection Operator (LASSO) regression was used to select variables in the training group, and a prediction model for uterine fibroid risk was constructed using logistic regression analysis. The performance and clinical utility of the model were evaluated using the area under the receiver operating characteristic curve (AUC), calibration curves, and decision curve analysis (DCA). RESULTS: A total of 18,750 patients were included. LASSO regression and logistic regression identified 16 predictors closely associated with uterine fibroid risk. The calibration curve demonstrated good consistency between predicted and observed results, with no significant difference in the Hosmer-Lemeshow test (P = 1). The AUC of the model was 0.95 (95% CI, 0.95–0.96), with a sensitivity of 0.88, specificity of 0.92, and prediction accuracy of 0.89 (95% CI, 0.88–0.89). In the validation group, the AUC was 0.95 (95% CI, 0.94–0.96), with a sensitivity of 0.88, specificity of 0.93, and prediction accuracy of 0.89 (95% CI, 0.88–0.90). The model exhibited excellent predictive performance and clinical utility. CONCLUSION: This study developed a uterine fibroid risk prediction model using the MIMIC database. Through rigorous statistical methods, 16 predictors closely associated with uterine fibroid risk were identified, and the reliability and clinical applicability of the model were validated. This model enables individualized risk prediction for patients and provides a powerful tool for early prediction and risk reduction of uterine fibroids.
Full text 26,751 characters · extracted from pmc-nxml · 6 sections · click to expand

Results

In this study, 18,750 patients were finally enrolled based on the inclusion and exclusion criteria. The data were randomly divided into a training set ( n  = 13,125) and a validation set ( n  = 5,625) in a 7:3 ratio. Comparisons of risk warning indicators between the two groups showed no significant differences in the study variables ( P  > 0.05), indicating good homogeneity between the two datasets (Table 1 ; Fig. 1 ). Basic characteristics and differences of the training set are provided in Supplementary Table 2. Fig. 1 The flow diagram of developing and validating the prediction model The flow diagram of developing and validating the prediction model Table 1 Comparison of characteristics between the training and validation groups Variables Total ( n  = 18750) Validation ( n  = 5625) Training ( n  = 13125) P Age (years) 62.00 (47.00, 75.00) 62.00 (47.00, 75.00) 63.00 (48.00, 75.00) 0.752 Height (cm) 163.00 (157.00, 165.00) 163.00 (157.00, 168.00) 163.00 (157.00, 165.00) 0.244 Weight (kg) 72.00 (60.00, 87.23) 72.00 (60.19, 88.22) 71.85 (60.00, 86.91) 0.085 BMI 27.68 (23.48, 33.42) 27.77 (23.46, 33.76) 27.64 (23.48, 33.30) 0.134 Na (mEq/L) 139.00 (137.00, 141.00) 139.00 (137.00, 141.00) 139.00 (137.00, 141.00) 0.341 K (mEq/L) 4.10 (3.80, 4.40) 4.10 (3.80, 4.40) 4.10 (3.80, 4.40) 0.926 Glu (mg/dl) 108.00 (92.00, 139.00) 108.00 (92.00, 139.00) 108.00 (92.00, 138.00) 0.630 HbA1c (%) 5.70 (5.40, 6.20) 5.70 (5.40, 6.20) 5.70 (5.40, 6.20) 0.418 Rbc (m/uL) 4.03 (3.54, 4.43) 4.04 (3.55, 4.43) 4.03 (3.53, 4.43) 0.706 Wbc (K/uL) 8.40 (6.33, 11.40) 8.30 (6.30, 11.40) 8.40 (6.40, 11.40) 0.581 Plt (K/uL) 250.00 (194.00, 314.00) 249.00 (194.00, 315.00) 251.00 (193.00, 314.00) 0.901 Hb (g/dl) 11.90 (10.30, 13.10) 11.90 (10.30, 13.10) 11.90 (10.30, 13.10) 0.748 Ast (IU/L) 24.00 (18.00, 39.00) 24.00 (19.00, 38.00) 24.00 (18.00, 40.00) 0.303 Alt (IU/L) 20.00 (14.00, 32.00) 19.00 (14.00, 32.00) 20.00 (14.00, 32.00) 0.231 Alb (g/dl) 3.90 (3.30, 4.30) 3.90 (3.30, 4.30) 3.90 (3.30, 4.30) 0.914 Ldl (mg/dl) 93.00 (68.00, 121.00) 95.00 (69.00, 121.00) 93.00 (67.00, 121.00) 0.121 Hdl (mg/dl) 51.00 (39.00, 65.00) 51.00 (39.00, 65.00) 51.00 (39.00, 65.00) 0.685 Tc (mg/dl) 175.00 (141.00, 209.00) 176.00 (142.00, 209.00) 174.00 (141.00, 209.00) 0.107 Tg (mg/dl) 117.00 (82.00, 174.00) 118.00 (82.00, 174.00) 117.00 (81.00, 174.00) 0.516 TyG index 8.80 (8.36, 9.32) 8.82 (8.37, 9.31) 8.80 (8.36, 9.32) 0.274 Bun (mg/dl) 15.00 (11.00, 23.00) 15.00 (11.00, 23.00) 15.00 (11.00, 22.00) 0.587 Scr (mg/dl) 0.80 (0.70, 1.00) 0.80 (0.70, 1.00) 0.80 (0.70, 1.00) 0.301 Inr 1.20 (1.10, 1.40) 1.20 (1.10, 1.30) 1.20 (1.10, 1.40) 0.096 Pt (s) 12.90 (11.80, 14.90) 12.90 (11.70, 14.80) 12.90 (11.80, 14.90) 0.144 Ptt (s) 30.00 (26.80, 36.10) 30.00 (26.80, 35.80) 30.10 (26.80, 36.20) 0.425 Race, n(%) 0.200  Asian 381 (2.03) 102 (1.81) 279 (2.13)  White 9631 (51.37) 2852 (50.70) 6779 (51.65)  Black 1620 (8.64) 481 (8.55) 1139 (8.68)  Other 7118 (37.96) 2190 (38.93) 4928 (37.55) Copd, n(%) 0.990  No 15,301 (81.61) 4590 (81.60) 10,711 (81.61)  Yes 3449 (18.39) 1035 (18.40) 2414 (18.39) Coronary, n(%) 0.863  No 13,746 (73.31) 4119 (73.23) 9627 (73.35)  Yes 5004 (26.69) 1506 (26.77) 3498 (26.65) Diabetes, n(%) 0.870  No 13,219 (70.50) 3961 (70.42) 9258 (70.54)  Yes 5531 (29.50) 1664 (29.58) 3867 (29.46) Hypertension, n(%) 0.562  No 12,843 (68.50) 3836 (68.20) 9007 (68.62)  Yes 5907 (31.50) 1789 (31.80) 4118 (31.38) Leiomyoma, n(%) 0.114  No 15,124 (80.66) 4498 (79.96) 10,626 (80.96)  Yes 3626 (19.34) 1127 (20.04) 2499 (19.04) Comparison of characteristics between the training and validation groups LASSO regression with 10-fold cross-validation was used to screen predictive variables from the 30 candidate variables. The analysis was performed using the glmnet package in R. The initial lambda (λ) was set to iterate 100 times. In the LASSO regression model, the optimal penalty coefficient λ was determined by plotting the coefficient changes (Fig. 2 A), where the left dashed line corresponds to the minimum λ (λmin), and the right dashed line corresponds to λmin plus one standard error. This value was selected as the optimal λ for the model. As the penalty coefficient λ increased, the coefficients of the initial variables were gradually compressed, with some coefficients shrinking to zero (Fig. 2 B). This process minimized overfitting and optimized variable selection. To balance model performance and parsimony, the optimal λ was chosen based on the 10-fold cross-validation error at λmin plus one standard error. This resulted in the selection of 20 variables as predictive factors: Race, Coronary, Hypertension, Age, Height, Weight, Na, Rbc, Wbc, Plt, Hb, Glu, Ast, Alb, Hdl, Tc, TyG, Bun, INR, and Ptt. Fig. 2 Using the LASSO regression model for variable selection. A Illustrates the selection process for the optimal penalty coefficient λ in the LASSO model. The optimal λ was determined using 10-fold cross-validation, with the left dashed line indicating the minimum λ (λmin) and the right dashed line indicating λmin plus one standard error, corresponding to the optimal number of selected variables. B Presents the LASSO coefficient curves for the 30 variables, showing how coefficients are penalized and reduced to zero as λ increases, aiding in variable selection Using the LASSO regression model for variable selection. A Illustrates the selection process for the optimal penalty coefficient λ in the LASSO model. The optimal λ was determined using 10-fold cross-validation, with the left dashed line indicating the minimum λ (λmin) and the right dashed line indicating λmin plus one standard error, corresponding to the optimal number of selected variables. B Presents the LASSO coefficient curves for the 30 variables, showing how coefficients are penalized and reduced to zero as λ increases, aiding in variable selection Multivariate logistic regression was performed on the selected feature variables (using forward and backward stepwise regression with a variable criterion of P  < 0.05) to establish a clinical prediction model. A total of 16 predictors were ultimately identified: Race, Coronary, Hypertension, Age, Height, Weight, Rbc, Wbc, Plt, Hb, Ast, Alb, Tc, TyG, Bun, and INR were identified as risk factors for uterine fibroids. The results showed that compared to Caucasians, individuals of Asian, African, and other races had a higher risk of developing uterine fibroids, which was statistically significant ( P  < 0.05). The presence of cardiovascular disease was associated with a reduced risk of uterine fibroids (OR = 0.0484, 95% CI 0.383–0.613, P  < 0.001). Hypertension was associated with an increased risk of uterine fibroids (OR = 1.496, 95% CI 1.260–1.776, P  < 0.001). Younger age was associated with an increased risk of uterine fibroids (OR = 0.951, 95% CI 0.947–0.955, P  < 0.001). Increased height was associated with an increased risk of uterine fibroids (OR = 1.030, 95% CI 1.021–1.040, P  < 0.001). Increased weight was associated with an increased risk of uterine fibroids (OR = 1.011, 95% CI 1.008–1.015, P  < 0.001). Elevated Rbc values were associated with an increased risk of uterine fibroids (OR = 1.609, 95% CI 1.343–1.927, P  < 0.001). Decreased Wbc values were associated with an increased risk of uterine fibroids (OR = 0.944, 95% CI 0.927–0.961, P  < 0.001). Elevated Plt values were associated with an increased risk of uterine fibroids (OR = 1.002, 95% CI 1.002–1.003, P  < 0.001). Decreased Hb values were associated with an increased risk of uterine fibroids (OR = 0.809, 95% CI 0.763–0.858, P  < 0.001). Decreased Ast values were associated with an increased risk of uterine fibroids (OR = 0.998, 95% CI 0.997–0.999, P  < 0.001). Elevated Alb values were associated with an increased risk of uterine fibroids (OR = 2.138, 95% CI 1.894–2.415, P  < 0.001). Elevated Tc values were associated with an increased risk of uterine fibroids (OR = 1.003, 95% CI 1.001–1.004, P  = 0.002). Decreased TyG values were associated with an increased risk of uterine fibroids (OR = 0.534, 95% CI 0.472–0.605, P  < 0.001). Decreased Bun values were associated with an increased risk of uterine fibroids (OR = 0.988, 95% CI 0.979–0.996, P  = 0.003). Decreased INR values were associated with an increased risk of uterine fibroids (OR = 0.563, 95% CI 0.463–0.684, P  < 0.001). A nomogram model for predicting the probability of uterine fibroid occurrence was developed based on the above risk factors (Table 2 ; Fig. 3 ). The OR values were converted to natural logarithms (log10), and a forest plot of risk factors was generated (Fig. 4 ). Table 2 Univariate and multivariate logistic regression results Variables Univariate Multivariate β P OR (95%CI) β P OR (95%CI) Race  Asian 1.00 (Reference) 1.00 (Reference)  White −1.27 < 0.001 0.28 (0.14 ~ 0.57) −1.07 0.006 0.34 (0.16 ~ 0.73)  Black 0.69 0.055 1.99 (0.98 ~ 4.04) 0.30 0.447 1.35 (0.63 ~ 2.90)  Other 3.31 < 0.001 27.48 (14.11 ~ 53.51) 3.09 < 0.001 21.88 (10.62 ~ 45.06) Coronary  No 1.00 (Reference) 1.00 (Reference)  Yes −2.17 < 0.001 0.11 (0.10 ~ 0.14) −0.72 < 0.001 0.48 (0.38 ~ 0.61) Hypertension  No 1.00 (Reference) 1.00 (Reference)  Yes −0.68 < 0.001 0.51 (0.46 ~ 0.56) 0.40 < 0.001 1.50 (1.26 ~ 1.78) Age −0.07 < 0.001 0.93 (0.93 ~ 0.93) −0.05 < 0.001 0.95 (0.95 ~ 0.96) Height 0.05 < 0.001 1.05 (1.05 ~ 1.06) 0.03 < 0.001 1.03 (1.02 ~ 1.04) Weight 0.01 < 0.001 1.01 (1.01 ~ 1.02) 0.01 < 0.001 1.01 (1.01 ~ 1.01) Na 0.01 0.386 1.01 (0.99 ~ 1.02) Rbc 0.61 < 0.001 1.84 (1.71 ~ 1.97) 0.48 < 0.001 1.61 (1.34 ~ 1.93) Wbc −0.08 < 0.001 0.92 (0.91 ~ 0.93) −0.06 < 0.001 0.94 (0.93 ~ 0.96) Plt 0.01 < 0.001 1.01 (1.01 ~ 1.01) 0.01 < 0.001 1.01 (1.01 ~ 1.01) Hb 0.06 < 0.001 1.06 (1.03 ~ 1.08) −0.21 < 0.001 0.81 (0.76 ~ 0.86) Glu −0.02 < 0.001 0.98 (0.98 ~ 0.99) Ast −0.01 < 0.001 0.99 (0.99 ~ 0.99) −0.01 < 0.001 0.99 (0.99 ~ 0.99) Alb 1.22 < 0.001 3.38 (3.12 ~ 3.67) 0.76 < 0.001 2.14 (1.89 ~ 2.42) Hdl 0.01 < 0.001 1.01 (1.01 ~ 1.02) 0.00 0.074 1.00 (1.00 ~ 1.01) Tc 0.01 < 0.001 1.01 (1.01 ~ 1.01) 0.01 0.002 1.01 (1.01 ~ 1.01) TyG −0.82 < 0.001 0.44 (0.41 ~ 0.47) −0.63 < 0.001 0.53 (0.47 ~ 0.61) Bun −0.10 < 0.001 0.90 (0.89 ~ 0.91) −0.01 0.003 0.99 (0.98 ~ 0.99) Inr −1.81 < 0.001 0.16 (0.13 ~ 0.20) −0.57 < 0.001 0.56 (0.46 ~ 0.68) Ptt −0.02 < 0.001 0.98 (0.97 ~ 0.98) Univariate and multivariate logistic regression results Fig. 3 Nomogram predicting the risk of uterine fibroids in female patients Nomogram predicting the risk of uterine fibroids in female patients Fig. 4 Forest plot of the association between risk factors and uterine fibroids in female patients Forest plot of the association between risk factors and uterine fibroids in female patients Calibration curves of the nomogram in the training set showed good consistency between predicted and observed outcomes (Fig. 5 A), with the Hosmer-Lemeshow test indicating no significant differences ( P  = 1), suggesting a good fit. The prediction performance of the nomogram was assessed using the receiver operating characteristic (ROC) curve, which yielded an area under the curve (AUC) of 0.95 (95% CI, 0.95–0.96), with a sensitivity of 0.88, specificity of 0.92, prediction accuracy of 0.89 (95% CI, 0.88–0.89), positive predictive value (PPV) of 0.98, and negative predictive value (NPV) of 0.65 (Table 3 ; Fig. 5 C). In the validation group, the calibration curves also demonstrated good consistency, and the Hosmer-Lemeshow test showed no significant differences ( P  > 0.05), confirming high calibration. The AUC was 0.95 (95% CI, 0.94–0.96), with a sensitivity of 0.88, specificity of 0.93, prediction accuracy of 0.89 (95% CI, 0.88–0.90), PPV of 0.98, and NPV of 0.66 (Table 3 ; Fig. 5 C). There was no statistically significant difference in AUC between the training and validation sets ( P  > 0.05) (Fig. 5 B). Fig. 5 Discrimination and calibration of the nomogram prediction model in the training and validation groups. A Calibration Curve in the Training Group; ( B) Calibration Curve in the Validation Group; ( C) ROC Curves in the Training and Validation Groups Discrimination and calibration of the nomogram prediction model in the training and validation groups. A Calibration Curve in the Training Group; ( B) Calibration Curve in the Validation Group; ( C) ROC Curves in the Training and Validation Groups Table 3 Diagnostic performance of the nomogram model for uterine fibroids in female patients in the training and validation groups Data AUC (95%CI) Accuracy (95%CI) Sensitivity (95%CI) Specificity (95%CI) PPV (95%CI) NPV (95%CI) cut off Training 0.95 (0.95–0.96) 0.89 (0.88–0.89) 0.88 (0.88–0.89) 0.92 (0.91–0.93) 0.98 (0.98–0.98) 0.65 (0.63–0.66) 0.217 Validation 0.95 (0.94–0.96) 0.89 (0.88–0.90) 0.88 (0.87–0.89) 0.93 (0.91–0.94) 0.98 (0.98–0.98) 0.66 (0.63–0.68) 0.217 Diagnostic performance of the nomogram model for uterine fibroids in female patients in the training and validation groups Decision curve analysis (DCA) is a method used to evaluate the clinical utility of a diagnostic test by quantifying the net benefit across different threshold probabilities. In this study, DCA was applied to assess the clinical utility of the nomogram model. Compared to the thresholds of “no intervention” and “intervention for all,” the nomogram demonstrated higher clinical net benefits in both the training and validation groups (Fig. 6 A and D). Clinical impact curves further indicated that the number of high-risk patients predicted to experience uterine fibroids and the actual number of uterine fibroid events tended to align closely with the risk threshold. This suggests that the model has significant predictive power and good clinical applications (Fig. 6 B and C). Fig. 6 Clinical applications assessment of the nomogram prediction model in the training and validation groups. A DCA in the Training Group; ( B ) DCA in the Validation Group; ( C ): Clinical Impact Curve in the Training Group; ( D) : Clinical Impact Curve in the Validation Group Clinical applications assessment of the nomogram prediction model in the training and validation groups. A DCA in the Training Group; ( B ) DCA in the Validation Group; ( C ): Clinical Impact Curve in the Training Group; ( D) : Clinical Impact Curve in the Validation Group

Materials

This study employed a retrospective cohort design, utilizing data from the Medical Information Mart for Intensive Care (MIMIC-IV3.0) database for female patients admitted between 2008 and 2022. The MIMIC-IV3.0 database is a publicly accessible resource approved by the Institutional Review Boards of Beth Israel Deaconess Medical Center and the Massachusetts Institute of Technology. Data extraction was performed by trained personnel who obtained certification for database access and download (ID: 49746917). This study adheres to the principles of the Declaration of Helsinki. Inclusion Criteria: Female patients diagnosed with uterine fibroids based on the International Classification of Diseases, Ninth and Tenth Revisions (ICD-9 and ICD-10) codes. Exclusion Criteria: (1) Age < 18 years; (2) Hospitalization duration < 24 h; (3) Missing data on weight, height, triglycerides (TG), or fasting blood glucose (FBG); (4) Multiple admissions, with only the first admission data included; (5) Patients with a history of hysterectomy or prior fibroid removal procedures. Variables Collected: Demographic Information: Age, weight, height, and race. Comorbidities: Hypertension, diabetes, cardiovascular disease, and uterine fibroids. Laboratory Variables: White blood cell count, red blood cell count, hemoglobin, platelet count, alanine aminotransferase, aspartate aminotransferase, albumin, high-density lipoprotein, low-density lipoprotein, cholesterol, triglycerides, serum creatinine, blood urea nitrogen, glucose, potassium, sodium, activated partial thromboplastin time (APTT), international normalized ratio (INR), and prothrombin time (PT). The value of the TyG index was calculated as ln [fasting TG (mg/dl) × FBG (mg/dl)/2]. Data extraction was performed using PostgreSQL (version 16) and Navicate Premium (version 17), and statistical analysis and visualization were conducted using R (version 4.4.0, www.R-project.org/ ). The R packages used in our study were displayed in Supplementary Table 1. Continuous Variables: Non-normally distributed continuous variables are presented as medians (interquartile ranges, M [Q1, Q3]) and were compared between groups using the Mann-Whitney U test. Categorical Variables: Categorical variables are presented as frequencies and percentages, and group comparisons were performed using the chi-square test. Missing Data Handling: Variables with missing values exceeding 10% were excluded. Variables with missing values between 5% and 10% were imputed using multiple imputation. Variables with missing values less than 5% were replaced with the mean value of the variable. Outliers were managed using the winsorization method, with cutoff points set at the 1 st and 99th percentiles. To address multicollinearity among variables, LASSO regression with 10-fold cross-validation was used to select predictive variables [ 12 ]. Specified the number of iterations (100) for lambda selection in LASSO. A multivariable logistic regression model (using stepwise selection with a significance threshold of P < 0.05) was then constructed based on the selected variables. A nomogram was created to visualize the prediction model. Model Evaluation: Discrimination: Assessed using the receiver operating characteristic (ROC) curve. Calibration: Evaluated using calibration curves. Clinical Utility: Assessed using decision curve analysis (DCA) and clinical impact curves. A two-sided P value of less than 0.05 was considered statistically significant. Patients were randomly divided into a training set (70%) and a validation set (30%) using a fixed random seed (set.seed(123)) to ensure reproducibility. The prediction model for uterine fibroid risk was developed and validated using the training set.

Conclusion

This study demonstrates that Race, Coronary, Hypertension, Age, Height, Weight, Rbc, Wbc, Plt, Hb, Ast, Alb, Tc, TyG, Bun, and INR are significantly associated with the occurrence of uterine fibroids. A nomogram model was developed and validated to predict the risk of uterine fibroids, enabling individualized patient prediction and offering the potential for early risk prediction and disease prevention.

Discussion

This study established a prediction model for the risk of uterine fibroids in women using the MIMIC database. The model identified Race, Coronary, Hypertension, Age, Height, Weight, Rbc, Wbc, Plt, Hb, Ast, Alb, Tc, TyG, Bun, and INR as risk factors. Validation confirmed the model’s reliability, enhancing its practicality for predicting uterine fibroid risk in women. Compared with the PRS-based model by Piekos et al.[ 13 ]. (AUC = 0.60), The AUC of this study to create a predictive model using only clinical indicators can reach 0.95 and does not require genotype data. Numerous studies have identified hypertension and cardiovascular diseases as significant risk factors for uterine fibroids [ 14 – 17 ]. In a single-center case-control study conducted in Japan [ 18 ], women with hypertension (defined as systolic blood pressure ≥ 140 mmHg, diastolic blood pressure ≥ 90 mmHg, or current use of antihypertensive medications) had nearly five times the risk of developing uterine fibroids compared to women with normal blood pressure. Therefore, the incidence of uterine fibroids is often higher in women with hypertension and cardiovascular diseases. There is a significant difference in the incidence of uterine fibroids among different races, confirming that uterine fibroids are more common in black women than in white women, and the incidence rate in black women is higher than that in other races [ 2 , 19 ], which suggests that there may be a genetic basis for this condition. Some studies have also indicated that age, height, and weight can influence the occurrence of uterine fibroids [ 20 – 22 ]. The incidence of uterine fibroids increases with age. A prospective study on the incidence and growth of uterine fibroids in young African American women showed that the incidence of fibroids increased from 6% in women aged 23–25 years to 13% in those aged 32–35 years as they grew older [ 23 ]. The inclusion of metabolic markers such as the TyG index, total cholesterol (TC), and hemoglobin (Hb) in the model suggests that metabolic abnormalities may promote fibroid growth by influencing estrogen levels or the local microenvironment. This is consistent with previous studies linking obesity and insulin resistance to an increased risk of uterine fibroids. Additionally, the significance of white blood cell (WBC) and platelet (PLT) counts reflects the role of chronic inflammation in fibroid development, supporting the hypothesis that uterine fibroids are associated with the activation of local inflammatory factors such as interleukin-6 (IL-6) and tumor necrosis factor-alpha (TNF-α)[ 24 ]. This study, for the first time, integrates blood biochemical indicators (e.g.: Ast, Alb, INR) with clinical characteristics (e.g.: height, weight) to construct a prediction model, overcoming the limitations of traditional studies that have predominantly focused on obstetric history or imaging features. The TyG (Triglyceride-Glucose) index, a marker of insulin resistance (IR) and metabolic syndrome is closely associated with metabolic disorders such as obesity, dyslipidemia, and hyperglycemia. Metabolic abnormalities, in turn, are linked to obesity, estrogen levels, and IR, indirectly promoting the proliferation of uterine fibroid cells. This study establishes a direct relationship between the TyG index and uterine fibroids. By incorporating TyG as a marker of IR, the model enhances its ability to identify high-risk populations for uterine fibroids. This approach is analogous to the use of metabolic indicators in risk models for cervical cancer [ 25 ]. Similar machine learning approaches have been successfully applied in other gynecological conditions, such as endometriosis and adenomyosis [ 9 ], and could be adapted for risk prediction in those contexts. Integration of this model into AI-based clinical decision support systems could enhance its utility in routine clinical practice, providing clinicians with a tool for early identification of high-risk patients. However, external validation in diverse populations and healthcare settings is necessary to ensure generalizability. Future studies could employ transfer learning techniques to adapt the model to local data sources, improving its performance across different regions. (1) This study is retrospective and may have residual or unmeasured confounding factors. However, rigorous statistical methods were used to strengthen the validity of the results. (2) The MIMIC database primarily includes critically ill patients, which may introduce selection bias. Additionally, the absence of certain indicators in the database limited the study. (3) Patient data were collected from a single center, leading to significant heterogeneity. Incomplete imaging data were not included in the analysis. Future research will involve collecting comprehensive demographic, laboratory, and imaging data for prospective, multicenter studies.

Introduction

Uterine fibroids (also known as leiomyomas or myomas) are common clonal neoplasms of the uterus, composed of smooth muscle cells and fibroblasts, and are rich in extracellular matrix (ECM). Their development and gene expression are regulated by sex steroid hormones, primarily estrogen and progesterone, in a menstrual cycle-dependent manner [ 1 , 2 ]. According to ultrasound screening studies and pathological data, uterine fibroids are highly prevalent, with up to 80% of women having detectable fibroids by the age of 50. The majority of patients with uterine fibroids are asymptomatic, and the lesions are predominantly benign [ 3 ]. Uterine fibroids exhibit an incidence associated with age and race. Compared to Caucasian women, Black women have a higher risk of developing fibroids at an earlier age and often experience greater disease severity [ 4 – 6 ]. The diagnosis of uterine fibroids is influenced by several factors, as the size, location, and number of fibroids vary among patients, and the symptoms caused by fibroids also differ [ 7 ]. The diagnosis of uterine fibroids is influenced by several factors, including the variability in size, location, and number of fibroids among patients, as well as the diversity of symptoms they cause. Many women do not associate their symptoms with fibroids, leading to potential delays in diagnosis. Additionally, some fibroids are asymptomatic and may remain undetected. To date, identified serum biomarkers lack sufficient predictive value to enhance diagnostic accuracy beyond that provided by physical examination and imaging. If biomarkers could accurately identify women at high risk of uterine fibroids or distinguish those with stable disease from those at higher risk, they might have clinical utility. Furthermore, biomarkers could be useful in differentiating benign uterine fibroids from other pathologies, thereby improving diagnostic precision and clinical decision-making [ 8 ]. In recent years, machine learning-based modeling has been widely applied in obstetrics and gynecology for diagnostic and risk prediction purposes. For instance, studies analyzing elastography values in endometrioma patients alongside clinical findings have demonstrated enhanced predictive performance when using machine learning algorithms [ 9 ]. Moreover, artificial intelligence techniques have shown considerable advantages in improving the accuracy and speed of diagnosis and treatment in uterine fibroid management [ 10 ]. In obstetrics, early-pregnancy serum biomarkers such as PAPP-A and β-hCG have been successfully used in machine learning models to predict the risk of gestational diabetes [ 11 ]. In this context, the model developed in the present study contributes meaningfully and innovatively to the existing literature by enabling early prediction of uterine fibroid risk. Currently, no prior studies have focused on developing a prediction model for uterine fibroids. To address this gap, we constructed and validated a prediction model for uterine fibroid risk using data from the Medical Information Mart for Intensive Care (MIMIC-IV3.0) database.

Supplementary Material

Supplementary Material 1. Supplementary Material 1. Supplementary Material 2. Supplementary Material 2.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-01T06:07:04.264727+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-NC-ND-4.0