A Machine Learning–Based Prediction Model for Deep Infiltrating Endometriosis

In: International Journal of Women's Health · 2026 · vol. Volume 18 , pp. 1–15 · doi:10.2147/ijwh.s627925 · W7214984031
article OA: gold CC0
⚙ AI-generated summary by qwen3.7-flash, 2026-10-04 ⓘ

This retrospective study developed a machine learning-based predictive model using clinical and imaging data to identify deep infiltrating endometriosis among patients with confirmed or suspected disease, finding the random forest algorithm achieved moderate internal validation performance.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

⚙ AI-generated deep summary by qwen3.7-flash, 2026-10-04 · read from full text ⓘ

This retrospective study developed and internally validated a machine learning-based prediction model for deep infiltrating endometriosis using clinical, imaging, and laboratory data from 250 surgically confirmed patients. Researchers compared Random Forest, Gradient Boosting Machine, Support Vector Machine, and conventional logistic regression algorithms, finding that the Random Forest model achieved the highest discriminative ability with an area under the curve of 0.787. The authors note that while the Random Forest model demonstrated moderate predictive performance and favorable calibration, it currently lacks clinical readiness and requires external multicenter validation to confirm generalizability. This paper is centrally about endometriosis — specifically focused on developing diagnostic tools for the deep infiltrating phenotype.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Objective: Deep infiltrating endometriosis (DIE) is a severe endometriosis phenotype. This study aimed to develop and internally validate a machine learning-based predictive model for DIE using retrospective clinical data from a single center to improve diagnostic accuracy. Materials and Methods: Clinical, imaging, and laboratory data were retrospectively collected from 250 surgically confirmed DIE patients and non-DIE endometriosis controls (2020– 2025). Samples were randomly split 7:3 into training and internal validation sets. Three machine learning algorithms-Random Forest (RF), Gradient Boosting Machine (GBM), and Support Vector Machine (SVM)-were developed and compared with conventional logistic regression. Model performance was assessed using AUC, sensitivity, specificity, calibration, and decision curves. The model was designed to identify DIE among patients with confirmed or suspected endometriosis. Results: Univariate analysis identified seven independent predictors (age, BMI, CA125, CA199, neutrophil-to-lymphocyte ratio, albumin, ovarian endometrioma), confirmed by multivariate logistic regression. In the validation set, AUCs were 0.787 (RF), 0.740 (GBM), 0.685 (SVM), and 0.709 (logistic regression). RF achieved 0.714 accuracy and 0.821 specificity, alongside favorable calibration and a positive net clinical benefit. Conclusion: In this retrospective single-center exploratory study, the RF model demonstrated moderate non-invasive predictive performance for DIE. However, this performance level does not yet support clinical readiness, and external multicenter validation is needed to confirm generalizability. Keywords: deep infiltrating endometriosis, machine learning, predictive model, internal validation
Full text 44,412 characters · extracted from oa-html · 11 sections · click to expand

Objective

Deep infiltrating endometriosis (DIE) is a severe endometriosis phenotype. This study aimed to develop and internally validate a machine learning-based predictive model for DIE using retrospective clinical data from a single center to improve diagnostic accuracy.

Materials and methods

Clinical, imaging, and laboratory data were retrospectively collected from 250 surgically confirmed DIE patients and non-DIE endometriosis controls (2020– 2025). Samples were randomly split 7:3 into training and internal validation sets. Three machine learning algorithms-Random Forest (RF), Gradient Boosting Machine (GBM), and Support Vector Machine (SVM)-were developed and compared with conventional logistic regression. Model performance was assessed using AUC, sensitivity, specificity, calibration, and decision curves. The model was designed to identify DIE among patients with confirmed or suspected endometriosis.

Results

Univariate analysis identified seven independent predictors (age, BMI, CA125, CA199, neutrophil-to-lymphocyte ratio, albumin, ovarian endometrioma), confirmed by multivariate logistic regression. In the validation set, AUCs were 0.787 (RF), 0.740 (GBM), 0.685 (SVM), and 0.709 (logistic regression). RF achieved 0.714 accuracy and 0.821 specificity, alongside favorable calibration and a positive net clinical benefit.

Conclusion

In this retrospective single-center exploratory study, the RF model demonstrated moderate non-invasive predictive performance for DIE. However, this performance level does not yet support clinical readiness, and external multicenter validation is needed to confirm generalizability.

Keywords

deep infiltrating endometriosis, machine learning, predictive model, internal validation

Introduction

Deep infiltrating endometriosis (DIE) is the most aggressive phenotype of endometriosis, histopathologically defined as endometrial-like glands and stroma infiltrating to a depth of ≥5 mm beneath the peritoneal surface. The Enzian classification is currently the most widely used system for describing DIE lesion locations, categorizing involvement of the posterior and anterior pelvic compartments.1 Clinically, DIE frequently involves pelvic organs such as the rectum, ureters, and bladder, leading to debilitating pelvic pain, dyspareunia, infertility, and intestinal or urinary tract symptoms.2 Because the clinical presentation of DIE lacks specificity and overlaps with that of ovarian endometrioma or superficial peritoneal endometriosis, early identification based solely on symptoms is challenging.3 Currently, transvaginal ultrasound and pelvic magnetic resonance imaging (MRI) are commonly used non-invasive imaging examinations. Expert-performed ultrasound and MRI have demonstrated high diagnostic accuracy for DIE, with reported sensitivity and specificity exceeding 90% for certain anatomical sites. However, these examinations are highly operator-dependent, require specialized expertise, and may not be readily available in all clinical settings. Consequently, diagnostic accuracy in routine practice varies considerably, and definitive diagnosis still requires laparoscopic surgery and pathological biopsy.4,5 Consequently, developing a non-invasive predictive model based on routine clinical, imaging, and laboratory indicators holds significant clinical value for the early screening, stratified management, and avoidance of unnecessary invasive procedures for DIE. In existing studies on predictive models for DIE, early investigations primarily utilized conventional logistic regression. For instance, Perelló et al (2017) developed a DIE prediction model using clinical indicators from patients with ovarian endometrioma, achieving an area under the curve (AUC) of 0.91 with internal validation.6 Arfi et al (2019) developed a nomogram (AUC 0.81) to predict live birth rates after DIE surgery. Although these studies confirmed the predictive value of clinical data, they did not incorporate machine learning methods.7 With the advancement of artificial intelligence, machine learning has gradually been applied to imaging-assisted diagnosis of DIE. Xu et al (2025) used a YOLOv8 deep learning model to detect DIE, significantly improving the diagnostic performance of junior sonographers (AUC increased from 0.748 to 0.878).8 Peluso et al (2026) combined MRI radiomics, graph computing, and machine learning to distinguish active from fibrotic DIE lesions (balanced accuracy 72.7%), demonstrating the potential of integrating radiomics and machine learning, although their objective was not to predict the presence of DIE based on multimodal clinical data.9 Simultaneously, numerous machine learning diagnostic models based on gene expression data have emerged. Shi et al (2024) identified the biomarker USP14 (AUC 0.786) using Least Absolute Shrinkage and Selection Operator (LASSO), Random Forest (RF), and Support Vector Machine (SVM).10 She et al (2022) combined RF and an artificial neural network to construct an endometriosis diagnostic model.11 Zhang et al (2023) used 11 algorithms, including RF, SVM, Gradient Boosting Machine (GBM), and neural networks, to screen five key genes, achieving an AUC of 0.836.12 Similarly, Zou et al (2023) applied RF, SVM, and LASSO to identify aging-related genes and constructed a neural network classifier.13 Recent systematic reviews have confirmed these trends. A meta-analysis by Zhang et al (2026) included 45 studies and found that machine learning models based on clinical features achieved a validation set AUC of only approximately 0.80, whereas models based on imaging and genetic data achieved AUCs of up to 0.98 and 0.87, respectively.14 A literature review by Shrestha et al (2025) also indicated promising prospects for machine learning in imaging diagnosis of endometriosis, but noted a persistent lack of predictive models specifically designed for DIE that are based on routine clinical/imaging/laboratory indicators and employ multiple mainstream machine learning algorithms with internal validation.15 Most existing machine learning models rely on costly or less accessible genetic data, focus on lesion detection rather than comprehensive clinical diagnosis of DIE, and are often limited by small sample sizes. Although individual markers such as CA125, CA199, and NLR have been associated with DIE, no single indicator offers sufficient diagnostic accuracy. A model integrating multiple readily available clinical, imaging, and laboratory parameters could provide standardized risk assessment when expert imaging is unavailable or inconclusive. Therefore, this study retrospectively collects clinical, imaging, and laboratory data from 250 patients over the past five years. Three machine learning algorithms will be employed to construct and internally validate DIE prediction models. By comparing their discriminative ability, calibration, and clinical utility, we will select the optimal model. Intended for use in patients with suspected or surgically confirmed endometriosis—rather than for asymptomatic screening—this tool can assist outpatient physicians in timely risk stratification and referral decisions, especially when expert ultrasound or MRI is unavailable or inconclusive.

Materials and methods

Sample Size Calculation According to the principles of sample size estimation for machine learning prediction model development, the Events Per Variable (EPV) method was used. This study planned to include 5–7 predictor variables. Referring to the study by Perelló et al (2017), the incidence of DIE was assumed to be 50%–60%.6 To ensure model stability and avoid overfitting, an EPV of 15 was adopted. With 7 predictor variables, the required number of DIE events was 7×15 = 105, and the minimum total sample size was 105/0.5 = 210 cases. With 5 predictor variables and EPV = 20, the minimum total sample size was 5×20/0.5 = 200 cases. Consequently, this study ultimately plans to enroll 250 cases, meeting the sample size requirements for machine learning prediction models. Inclusion Criteria (1) All patients underwent laparoscopic or open surgical resection of lesions, with postoperative pathological examination definitively diagnosing endometriosis. (2) Pelvic MRI or transvaginal ultrasound was performed within 2 weeks before surgery, with complete and interpretable imaging data. (3) Peripheral venous blood was collected within 1 month before surgery to complete laboratory tests. (4) Complete clinical medical records were available. Exclusion Criteria (1) Preoperative hormonal therapy for more than 1 month. (2) Co-existing other malignant tumors or severe systemic diseases. (3) Prior pelvic radiotherapy or radical surgery for pelvic malignancy. (4) More than 20% missing key clinical, imaging, or laboratory data. (5) Pregnancy or lactation. (6) Postoperative pathological diagnosis of diseases other than endometriosis. Pathological Assessment All surgical specimens were fixed in 10% neutral-buffered formalin, routinely processed, embedded in paraffin, and sectioned at 4-μm thickness. Sections were stained with hematoxylin and eosin (H&E) for histopathological evaluation. DIE was defined as the presence of endometrial glands and stroma infiltrating to a depth of ≥5 mm beneath the peritoneal surface, as measured by an ocular micrometer on the H&E-stained sections. The depth of infiltration was independently measured by two pathologists. In cases of disagreement, a third senior pathologist was consulted to reach a consensus. The presence of fibrosis, smooth muscle metaplasia, and surrounding tissue involvement was also recorded as ancillary diagnostic features. Pathologists were blinded to the patients’ clinical and imaging data during the histopathological evaluation. Data Collection All patients were consecutively enrolled from the Department of Obstetrics and Gynecology of Hangzhou Red Cross Hospital between January 2020 and December 2025. No specific sampling strategy was applied; all eligible patients who met the inclusion criteria during the study period were included to minimize selection bias. General Baseline and Past Medical History Data General demographic and physiological baseline data were collected, including age in years at admission, fasting height and weight measured by nursing staff according to standard protocols at admission (used to calculate Body Mass Index, BMI), self-reported age in years at menarche, regularity of menstrual cycles during the 6 months before admission, all clinically confirmed previous pregnancies, and the number of previous pregnancy terminations before 28 weeks of gestation. All baseline data were verified against original medical records to ensure accuracy. Patients’ histories of prior pelvic surgeries and diseases were collected, including a history of prior laparoscopic or open surgery for endometriosis (confirmed by previous surgical records and postoperative pathology reports). Additionally, a history of other pelvic surgeries (excluding those for endometriosis) was collected, including cesarean section, uterine myomectomy, benign ovarian tumor surgery, fallopian tube-related surgery, and pelvic adhesiolysis. Clinical Symptom Data Clinical symptoms related to endometriosis during the 6 months before admission were collected. These included lower abdominal distending pain during menstruation clearly related to the menstrual cycle, persistent or intermittent pelvic pain lasting more than 6 months, deep pelvic pain in the lower abdomen during or after sexual intercourse, and defecation-related discomfort (eg, anal tenesmus, difficulty defecating, or sensation of incomplete evacuation) occurring during or just before menstruation. All symptoms were documented based on the admission history of present illness and gynecological physical examination records, while excluding symptoms caused by other definite organic diseases. Laboratory Test Indicators Fasting peripheral venous blood test results collected within 24 hours after admission and before surgery were obtained. All blood samples were collected, centrifuged, and tested according to the standardized procedures of our hospital’s clinical laboratory, using fixed instruments and reagents, with qualified internal quality control and external quality assessment. Specific indicators included serum carbohydrate antigen 125 (CA125), carbohydrate antigen 199 (CA199), carcinoembryonic antigen level (CEA), the neutrophil-to-lymphocyte ratio calculated from complete blood count, and serum albumin level. All test results were directly extracted from the official reports issued by the clinical laboratory, ensuring complete consistency with the original test reports. Imaging Lesion Data Imaging examination results completed at our hospital within 1 month before admission were collected. All images were independently interpreted in a double-blinded manner by two senior specialist physicians. In cases of disagreement, a chief physician with higher seniority reviewed the images to reach a final diagnosis. The collected data specifically included the presence of ovarian endometrioma assessed by transvaginal gynecological ultrasound and the presence of ureteral dilation assessed by urinary ultrasound or whole-abdominal computed tomography. Diagnoses followed standard imaging criteria for the respective diseases, and imaging abnormalities caused by factors unrelated to endometriosis were excluded. Outcome Definition The primary outcome of this study was the presence or absence of DIE, using postoperative pathological examination as the gold standard. According to the Enzian classification and pathological diagnostic criteria, DIE was defined as lesions infiltrating to a depth of ≥ 5 mm, involving pelvic structures such as the uterosacral ligaments, rectum, vagina, bladder, or ureters.16 All patients underwent laparoscopic or open surgical resection, and postoperative pathological reports were independently issued by two pathologists. The pathologists were blinded to the patients’ clinical, imaging, and laboratory data during histopathological evaluation. Inter-observer agreement was assessed, and disagreements were resolved by consultation with a third senior pathologist. Statistical Analysis Statistical analyses were performed using Statistical Product and Service Solutions 26.0 (SPSS 26.0) and R Programming Language 4.5.1 (R 4.5.1). Graphs were generated using GraphPad Prism 9.0. The significance level was set at α = 0.05, and a two-tailed P < 0.05 was considered statistically significant. Continuous data were first tested for normality. Data following a normal distribution were presented as mean ± standard deviation (X±S), and comparisons between groups were performed using the t-test. Data not following a normal distribution were presented as median (interquartile range), and comparisons between groups were performed using the Mann–Whitney U-test. Categorical data were presented as counts (percentages), and comparisons between groups were performed using the χ2 test, or Fisher’s exact test when the expected frequency was < 5. Patients were randomly divided into a training set and a validation set at a ratio of 7:3 using a random number table. Before grouping, the balance of baseline clinical characteristics between the two sets was verified using the aforementioned tests. Missing data, which accounted for less than 5% of all included variables, were handled by median imputation for continuous variables and mode imputation for categorical variables using training set statistics. The dataset showed no significant class imbalance (DIE: 50.9% in training, 49.3% in validation), thus no oversampling or undersampling techniques were applied. In the training set, univariate analysis was performed to screen for potential predictors with P < 0.05. LASSO regression was used for variable compression and redundancy elimination. The optimal parameters were determined using 10-fold cross-validation to obtain the core predictor variables, which were then included in a multivariable logistic regression model using a stepwise backward elimination method to select predictive factors. The multicollinearity diagnostic test was performed using variance inflation factor (VIF), with VIF < 10 indicating no significant multicollinearity. Based on the predictors screened by the above process, four prediction models were developed on the training set: Random Forest (RF), Gradient Boosting Machine (GBM), Support Vector Machine (SVM), and conventional logistic regression. For each machine learning algorithm, hyperparameters were tuned using grid search with 10-fold cross-validation within the training set. The final hyperparameters for RF were ntree=500 and mtry=3; for GBM, n.trees=100, interaction.depth=3, and shrinkage=0.1; and for SVM, cost=1 and gamma=0.1 with a radial basis kernel. Logistic regression was developed using the same set of predictors without hyperparameter tuning. The classification threshold was set at 0.5 by default, as no alternative threshold was optimized. Continuous variables were not separately standardized, as tree-based algorithms (RF and GBM) are scale-invariant, and SVM performance was not superior in our preliminary analysis. This sequential procedure—univariate screening, LASSO compression, and multivariable logistic regression—was applied solely to identify the final set of predictors for inclusion in the machine learning models. The final machine learning models were then developed independently using this same set of predictors, without further feature selection, to avoid circularity and ensure comparability across models. Receiver Operating Characteristic (ROC) curves were plotted, and the AUC with 95% Confidence Interval (CI) was calculated to evaluate model discrimination. Model discrimination was further evaluated using sensitivity, positive predictive value (PPV), negative predictive value (NPV), and F1-score. 95% confidence intervals for AUC, accuracy, sensitivity, and specificity were calculated using the bootstrap method with 1000 resamples. Pairwise comparisons of AUCs between models were performed using DeLong’s test. Calibration curves were plotted, and the Hosmer-Lemeshow test was used to assess the agreement between predicted probabilities and actual outcomes. In addition, calibration was assessed by calibration slope, intercept, and Brier score. Decision curve analysis (DCA) was performed to evaluate the clinical net benefit of each model across a range of threshold probabilities. For the RF model, the out-of-bag (OOB) error rate was monitored during training to assess generalization performance. To enhance model interpretability, SHapley Additive exPlanations (SHAP) analysis was conducted to quantify the contribution of each predictor Internal validation was performed using 1000 bootstrap samples to correct for overfitting bias. For machine learning model development, all preprocessing and feature selection steps were performed exclusively on the training set, with the validation set held out and not used until the final model evaluation to prevent information leakage.

Results

Comparison of General Data Between Training and Validation Sets Among the 175 patients in the training set, 89 (50.86%) had DIE. Among the 75 patients in the validation set, 37 (49.33%) had DIE. No statistically significant differences were found in the general data between the training and validation sets (P>0.05) (Table 1). | Table 1 Comparison of General Data Between the Training and Validation Sets | Univariate Analysis for the Presence of DIE Univariate analysis revealed significant differences between the DIE and non-DIE groups in the training set regarding age, BMI, CA125, CA199, NLR, albumin, and the presence of ovarian endometrioma (P<0.05) (Supplemental Table 1). Multivariate Logistic Regression Analysis for the Presence of DIE Using postoperatively confirmed DIE as the binary dependent variable (DIE = 1, non-DIE = 0), variables with P < 0.05 in the univariate analysis were incorporated into the LASSO regression for variable selection. The lambda.1se criterion was applied to select the final variables (Figure 1). The multicollinearity diagnostic test yielded variance inflation factor (VIF) values of 1.230, 1.109, 1.150, 1.191, 1.118, and 1.115 for the selected indicators, all of which met the inclusion criteria for multivariate regression analysis. Consequently, all these variables were included in the subsequent variable selection and model development. A multivariate binary logistic regression model was then constructed. The results demonstrated that age, BMI, serum CA125, CA199, NLR, albumin level, and concomitant ovarian endometrioma were independent influencing factors for the occurrence of DIE (P<0.05) (Table 2). | Figure 1 Lasso coefficient screening plot (A) and waterfall plot (B). | | Table 2 Multivariate Logistic Regression Analysis for the Presence of DIE | Predictive Performance of Machine Learning Models in the Training and Validation Sets The Hosmer-Lemeshow test results indicated that all P-values for the models in both the training and validation sets exceeded 0.05, suggesting good agreement between the models’ predicted probabilities and the observed probabilities. The calibration performance met the requirements for clinical prediction models. In the training set, the χ2 values for the four models were 11.658, 4.603, 9.648, and 16.658, with corresponding P-values of 0.167, 0.799, 0.291, and 0.053, respectively. In the validation set, the χ2 values for the four models were 6.532, 8.575, 3.349, and 6.273, with corresponding P of 0.588, 0.379, 0.911, and 0.617, respectively. The calibration slopes for RF, GBM, SVM, and logistic regression were 1.07, 1.21, 0.88, and 0.94, respectively, with intercepts of 0.06, −0.22, 0.28, and 0.10. The Brier scores were 0.187, 0.203, 0.214, and 0.198, indicating good calibration for all models. Calibration metrics should be interpreted with caution given the relatively small validation sample size. (Figure 2 and Table 3). | Figure 2 Calibration curve analysis of machine learning-based DIE prediction models. | | Table 3 Diagnostic Performance and Calibration Test Results of Each Prediction Model in the Training and Validation Sets | ROC curve analysis revealed that, in the training set, the AUC values for the RF, GBM, SVM, and conventional Logistic regression models were 0.794, 0.757, 0.756, and 0.728, respectively. Their corresponding accuracy values were 0.731, 0.714, 0.714, and 0.731, and specificity values were 0.879, 0.793, 0.741, and 0.638, respectively. In the internal validation set, the AUC values for the four aforementioned models were 0.787, 0.740, 0.685, and 0.709, with accuracy values of 0.714, 0.679, 0.643, and 0.714, and specificity values of 0.821, 0.821, 0.750, and 0.607, respectively. Notably, no major performance degradation was observed after internal validation, suggesting reasonable stability within this single-center dataset. However, given the limited size of the validation set and its single-center origin, this finding should be interpreted cautiously (Figure 3 and Table 3). The performance metrics for all four models in both the training and validation sets were summarized in Table 3. DeLong’s test for pairwise AUC comparisons was performed in both training and validation sets. In the internal validation set, RF demonstrated significantly higher AUC than SVM (P = 0.018) and logistic regression (P = 0.036), but not significantly different from GBM (P = 0.107). In the training set, RF was significantly superior to all other three models, with P = 0.002 vs SVM, P = 0.005 vs GBM, and P = 0.001 vs logistic regression. | Figure 3 ROC curve analysis of machine learning-based DIE prediction models. | Decision Curve Analysis of Machine Learning Models In the validation set, the RF model demonstrated a higher net benefit than the other models across threshold probabilities between 0.2 and 0.6. Within this range, the net benefit of the RF model ranged from 0.05 to 0.15, suggesting that using the model to guide referral decisions could reduce unnecessary procedures while maintaining acceptable detection rates. However, the clinical value of the model is currently limited to the specific scenario of differentiating DIE among patients with suspected endometriosis, and it should be interpreted as a supplement to rather than a replacement for expert imaging assessment. (Figure 4). | Figure 4 Decision curve analysis of machine learning-based DIE prediction models. | The error rate plot of the RF model showed a rapid decline in the out-of-bag (OOB) error rate, which eventually stabilized at a very low level. This observation reflects the strong generalization ability of the RF model; even when evaluated using OOB data, the error rate remained low and stable (Figure 5). The SHAP variable importance plot illustrates the contribution of each variable to the model (Figure 6). The SHAP values reflect the impact of each variable on the model’s predicted probability of DIE, with the direction and magnitude of contribution varying across individuals. Among the seven predictors, BMI and CA125 showed relatively larger SHAP magnitudes, suggesting stronger influence on model predictions, while age and NLR demonstrated comparatively smaller contributions. This analysis is intended to enhance model interpretability rather than to establish causal relationships, and the clinical utility of SHAP-derived feature importance requires further investigation in prospective studies. | Figure 5 Error rate plot of the RF model. | | Figure 6 Shap variable importance plot. |

Discussion

The results demonstrated that the RF model outperformed the other models in terms of discrimination, calibration, and clinical net benefit. Independent predictors of DIE included age, BMI, serum levels of CA125 and CA199, NLR, albumin level, and the presence of ovarian endometrioma. While these associations are biologically plausible, the current study design does not permit mechanistic conclusions, and these interpretations should be viewed as hypothesis-generating. Although these predictors are individually known to be associated with endometriosis severity, their integration into a composite model that provides a standardized, quantitative probability estimate represents the main added value of this study. Elevated serum CA125 level is a classic biomarker for DIE.17 CA125 is primarily secreted by coelomic epithelial cells upon inflammatory or mechanical stimulation. DIE lesions, which infiltrate to a depth exceeding 5 mm, often involve the pelvic peritoneum and serosal layers, inducing a robust local inflammatory response and peritoneal irritation. Consequently, CA125 is released into the bloodstream.18 In the present study, serum CA125 levels were significantly higher in the DIE group than in the non-DIE group. However, CA125 is not a specific marker for DIE; elevated levels can also occur in patients with ovarian endometrioma or pelvic inflammatory disease. Therefore, the diagnostic accuracy of CA125 alone for DIE remains limited.19 Serum CA199 levels were also significantly elevated in the DIE group. CA199 is a common biomarker for gastrointestinal malignancies. Nevertheless, in the context of endometriosis, when lesions invade pelvic organs such as the rectum, sigmoid colon, or ureters, local tissue injury and repair processes can induce CA199 expression in epithelial cells.19 Patients with DIE involving the bowel or urinary tract exhibit markedly higher serum CA199 levels than those without such involvement.20 Although the present study did not directly confirm a correlation between CA199 elevation and specific sites of organ involvement, the findings suggest that elevated CA199 may serve as an indirect indicator of DIE affecting deep pelvic organs. The NLR is a composite marker reflecting systemic inflammatory status. DIE is recognized as a chronic inflammatory disease. Ectopic endometrial tissue undergoes repeated bleeding and necrosis within the pelvic cavity, activating macrophages and neutrophils and releasing multiple pro-inflammatory cytokines.21 Elevated neutrophil counts accompanied by relative lymphopenia result in an increased NLR. This ratio remains relatively stable and is less influenced by acute factors such as hydration status.22 Reduced albumin level represents another noteworthy predictive factor. Albumin is synthesized by the liver, and its level is influenced by both nutritional status and inflammatory state.23 Patients with DIE often experience chronic pain, gastrointestinal symptoms, and potential intestinal malabsorption, leading to inadequate protein intake or increased protein catabolism. Furthermore, chronic inflammation suppresses albumin gene transcription through inflammatory cytokines such as tumor necrosis factor-alpha and interleukin-6, further decreasing serum albumin levels.24 In this study, patients in the DIE group were younger and had lower BMI compared with those in the non-DIE group. Younger women have a relatively compact pelvic anatomy and higher estrogen levels, which may create favorable conditions for aggressive DIE growth. Low BMI may be related to two factors: (1) weight loss resulting from DIE-related gastrointestinal symptoms and chronic pain, and (2) potential effects of hormones secreted by adipose tissue, such as leptin, on the development of endometriosis.25 In contrast, other studies have reported opposite conclusions. These discrepancies may be attributable to differences in the composition of control groups across studies.26 The control group in the present study included patients with ovarian endometrioma, a condition often associated with higher BMI. Consequently, the intergroup comparisons warrant cautious interpretation. Ovarian endometrioma is one of the most common comorbidities of DIE. In the present study, more than half of the patients with DIE also had ovarian endometrioma. From an anatomical and pathogenetic perspective, endometrial tissue released upon rupture of an ovarian endometrioma can implant onto the posterior pelvic compartment, promoting the formation of DIE at sites such as the uterosacral ligaments and the rectovaginal septum.27 Therefore, when ovarian endometrioma is clinically identified, a high index of suspicion for concurrent DIE is warranted, and further imaging evaluation should be considered when necessary. The strength of this study lies in its systematic comparison of three mainstream machine learning algorithms using routine clinical, imaging, and laboratory indicators, coupled with internal validation. We acknowledge that several recent machine-learning studies have applied similar analytical pipelines to endometriosis prediction. Our model distinguishes itself by three aspects. First, it is specifically designed for DIE—the most severe phenotypic subtype—rather than general endometriosis or ovarian endometrioma. Second, it relies exclusively on routine clinical, laboratory, and basic ultrasound variables that are widely accessible across different healthcare settings. Third, it provides a standardized, quantitative probability estimate for DIE that can assist physicians in outpatient or primary care settings when expert imaging is unavailable or inconclusive. This study has several limitations. First, the data were retrospectively collected from a single center. Although the sample size met the basic requirements for machine learning modeling, the events-per-variable approach used for sample size estimation is more suited to conventional regression than to comparing and tuning multiple machine learning algorithms, and the internal validation set contained only 75 patients, including 37 DIE cases, which may limit the precision of performance estimates. Selection bias and information bias remain possible. Second, and most critically, only internal validation was performed; an independent external validation dataset was lacking. Consequently, the generalizability of the model across different populations, healthcare centers, and geographic regions remains uncertain. Third, imaging variables were confined to transvaginal ultrasound findings of ovarian endometrioma and ureteral dilation, whereas a broader set of detailed imaging features—including lesion location, uterosacral ligament involvement, rectovaginal septum involvement, bowel or bladder infiltration, nodule size, and MRI-specific signs—was not incorporated, even though these parameters are well established as highly informative for DIE diagnosis in expert ultrasound or MRI settings. This constrains the model’s scope considerably, given that comprehensive imaging assessment remains the cornerstone of current clinical DIE evaluation. Fourth, the control group comprised patients with endometriosis but without DIE during the same period, rather than healthy individuals. Thus, the model is primarily applicable for further identifying the DIE subtype among patients already highly suspected of having endometriosis, rather than serving as a screening tool in the general population. Additionally, the study timeframe (2020–2025) supports the development and internal validation of a diagnostic prediction model based on surgical and pathological outcomes, but is insufficient to assess long-term clinical utility, impact on clinical decision-making, or patient outcomes. Therefore, the model should be interpreted as exploratory at this stage. If validated externally, such noninvasive prediction models could eventually support risk stratification and referral decisions in resource-limited settings where expert imaging is unavailable. However, these potential applications remain speculative at this stage. Future research should focus on external validation, comparison with expert imaging-based diagnostic pathways, and prospective assessment of clinical impact before any implementation can be considered.

Conclusions

In this retrospective single-center exploratory study, the RF model demonstrated moderate predictive performance for noninvasive DIE identification using routine clinical and laboratory variables. The model was developed for differentiating DIE from non-DIE endometriosis among patients with surgically confirmed endometriosis, rather than for general population screening or initial diagnosis. The intended clinical scenario is to assist physicians in outpatient or primary care settings where expert ultrasound or MRI is unavailable or inconclusive, by providing a standardized probability estimate to support referral decisions. However, the model has only been internally validated, and its performance, generalizability, and added clinical value require further assessment. External validation in independent cohorts, comparison with expert imaging-based diagnostic pathways, and prospective evaluation are prerequisites before any clinical implementation can be considered. Data Sharing Statement The data used and/or analysed during the current study available from the corresponding author on reasonable request. Ethics Approval and Consent to Participate The study was approved by the Ethics Committee of Hangzhou Red Cross Hospital (No. HZ-审-03-65). The approval fully covered retrospective data extraction from electronic medical records. Informed consent was obtained from all patients. This study was conducted in accordance with the Declaration of Helsinki. Author Contributions All authors made a significant contribution to the work reported, whether that is in the conception, study design, execution, acquisition of data, analysis and interpretation, or in all these areas; took part in drafting, revising or critically reviewing the article; gave final approval of the version to be published; have agreed on the journal to which the article has been submitted; and agree to be accountable for all aspects of the work. Funding There is no funding to report. Disclosure The authors declare that they have no competing interests.

References

1. O’leary M, Neary C, Lawrence E. The diagnostic accuracy of magnetic resonance imaging versus transvaginal ultrasound in deep infiltrating endometriosis and their impact on surgical Decision-Making: a systematic review. Diagnostics. 2025;15(22). doi:10.3390/diagnostics15222856 2. Kim HJ, Lee EJ, Hong SS, et al. Comprehensive review of endometriosis based on imaging findings: what radiologists need to know. J Korean Soc Radiol. 2025;86(5):720–15. doi:10.3348/jksr.2024.0104 3. Ottolina J, Villanacci R, D’alessandro S, et al. Endometriosis and adenomyosis: modern concepts of their clinical outcomes, Treatment, and management. J Clin Med. 2024;13(14):3996. doi:10.3390/jcm13143996 4. Guerriero S, Saba L, Pascual M A, et al. Transvaginal ultrasound vs magnetic resonance imaging for diagnosing deep infiltrating endometriosis: systematic review and meta-analysis. Ultrasound Obstet Gynecol. 2018;51(5):586–595. 5. Chen Q, Jia L, Wang S, et al. Douglas pouch fluid improves the accuracy of transvaginal ultrasound in the diagnosis of uterosacral ligaments deep infiltration endometriosis: a prospective study. J Ultrasound Med. 2025;44(1):111–117. 6. Perelló M, Martínez-Zamora M A, Torres X, et al. Markers of deep infiltrating endometriosis in patients with ovarian endometrioma: a predictive model. Eur J Obstet, Gynecol, Reprod Biol. 2017;209:55–60. 7. Arfi A, Bendifallah S, D’argent EM, et al. Nomogram predicting the likelihood of live-birth rate after surgery for deep infiltrating endometriosis without bowel involvement in women who wish to conceive: a retrospective study. Eur J Obstet, Gynecol, Reprod Biol. 2019;235:81–87. 8. Xu J, Zhang A, Zheng Z, Cao J, Zhang X. Development and validation an AI model to improve the diagnosis of deep infiltrating endometriosis for Junior Sonologists. Ultrasound Med Biol. 2025;51(7):1143–1147. doi:10.1016/j.ultrasmedbio.2025.03.012 9. Peluso S, Lucidi V, Di Giovanni M C, et al. Artificial intelligence-enhanced MRI radiomics for discriminating active and fibrotic lesions to Support therapeutic Decision Making in deep endometriosis. J Minim Invasive Gynecol. 2026;33:778–786. doi:10.1016/j.jmig.2025.12.037 10. Shi S, Huang C, Tang X, et al. Identification and verification of diagnostic biomarkers for deep infiltrating endometriosis based on machine learning algorithms. J Biol Eng. 2024;18(1):70. doi:10.1186/s13036-024-00466-9 11. She J, Su D, Diao R, et al. A joint model of Random Forest and artificial neural network for the diagnosis of endometriosis. Front Genet. 2022;13:848116. 12. Zhang H, Zhang H, Yang H, et al. Machine learning-based integrated identification of predictive combined diagnostic biomarkers for endometriosis. Front Genet. 2023;14:1290036. 13. Zou L, Meng L, Xu Y, et al. Revealing the diagnostic value and immune infiltration of senescence-related genes in endometriosis: a combined single-cell and machine learning analysis. Front Pharmacol. 2023;14:1259467. 14. Zhang B, Lv X, Li D, et al. Diagnostic accuracy of machine learning for endometriosis: a systematic review and meta-analysis. Front Endocrinol. 2025;16:1735567. 15. Shrestha P, Shrestha B, Sherestha J, et al. Current status and future potential of machine learning in diagnostic imaging of endometriosis: a literature review. JNMA J Nepal Med Assoc. 2025;63(283):205–211. 16. Becker CM, Bokor A, Heikinheimo O, et al. ESHRE guideline: endometriosis. Hum Reprod Open. 2022;2022(2):hoac009. 17. Abdul Karim A K, Abd Aziz N H, Md Zin R R, et al. The effect of surgical intervention of endometriosis to CA-125 and pain. Malays J Med Sci. 2020;27(6):7–14. 18. Shokrnejad-Namin T, Farzizadeh N, Najmi Z, et al. Evaluating the diagnostic potential of CA-125 and miRNA levels in endometriosis: a narrative review. Int J Gynaecol Obstet. 2026;172(3):1392–1429. 19. Chen T, Wei J L, Leng T, et al. The diagnostic value of the combination of hemoglobin, CA199, CA125, and HE4 in endometriosis. J Clin Lab Anal. 2021;35(9):e23947. 20. Dong Y, Wang L, Chen Y, et al. Malignant endometriosis of the rectovaginal septum: a case report. Asian J Surg. 2023;46(6):2626–2627. 21. Dominoni M, Pasquali M F, Musacchi V, et al. Neutrophil to lymphocytes ratio in deep infiltrating endometriosis as a new toll for clinical management. Sci Rep. 2024;14(1):7575. doi:10.1038/s41598-024-58115-6 22. Islam MM, Satici MO, Eroglu S E. Unraveling the clinical significance and prognostic value of the neutrophil-to-lymphocyte ratio, platelet-to-lymphocyte ratio, systemic immune-inflammation index, systemic inflammation response index, and delta neutrophil index: an extensive literature review. Turk J Emerg Med. 2024;24(1):8–19. doi:10.4103/tjem.tjem_198_23 23. Don BR, Kaysen G. Poor nutritional status and inflammation: serum albumin: relationship to inflammation and nutrition. Semin Dial. 2004;17(6):432–437. doi:10.1111/j.0894-0959.2004.17603.x 24. Kim Y, Molnar M Z, Rattanasompattikul M, et al. Relative contributions of inflammation and inadequate protein intake to hypoalbuminemia in patients on maintenance hemodialysis. Int Urol Nephrol. 2013;45(1):215–227. doi:10.1007/s11255-012-0170-8 25. Martire FG, Giorgi M, D’abate C, et al. Deep infiltrating endometriosis in adolescence: early diagnosis and possible prevention of disease progression. J Clin Med. 2024;13(2):550. doi:10.3390/jcm13020550 26. Rowlands IJ, Hockey R, Abbott J A, et al. Body mass index and the diagnosis of endometriosis: findings from a national data linkage cohort study. Obes Res Clin Pract. 2022;16(3):235–241. 27. Koninckx PR, Fernandes R, Ussia A, et al. Pathogenesis based diagnosis and Treatment of endometriosis. Front Endocrinol. 2021;12:745548. © 2026 The Author(s). This work is published and licensed by Dove Medical Press Limited. The full terms of this license are available at https://www.dovepress.com/terms and incorporate the Creative Commons Attribution - Non Commercial (unported, 4.0) License. By accessing the work you hereby accept the Terms. Non-commercial uses of the work are permitted without any further permission from Dove Medical Press Limited, provided the work is properly attributed. For permission for commercial use of this work, please see paragraphs 4.2 and 5 of our Terms. Recommended articles Prediction of Acute Kidney Injury in Intracerebral Hemorrhage Patients Using Machine Learning She S, Shen Y, Luo K, Zhang X, Luo C Neuropsychiatric Disease and Treatment 2023, 19:2765-2773 Published Date: 11 December 2023 A Machine-Learning Model Based on Clinical Features for the Prediction of Severe Dysphagia After Ischemic Stroke Ye F, Cheng LL, Li WM, Guo Y, Fan XF International Journal of General Medicine 2024, 17:5623-5631 Published Date: 28 November 2024 Development and Validation of a Neonatal Hypothermia Prediction Model for In-Hospital Transport Using Machine Learning Algorithms: A Single-Center Retrospective Study Zhang W, Gu X, Gu C, Yao L, Zhang Y, Wang K Journal of Multidisciplinary Healthcare 2025, 18:3205-3217 Published Date: 4 June 2025 Using Machine Learning and the HAMD-24 Scale to Predict Suicide Ideation in Depressed Patients Chen Y, Jiang ZY, Dong GZ, Zhang WY, Wang K, Yang HY Psychology Research and Behavior Management 2025, 18:2153-2165 Published Date: 12 October 2025 Development and Validation of a Disability Risk Prediction Model for Older Adults Based on Machine Learning: A Multi-Algorithm Comparison with SHAP Interpretation Yang S, Liang H, Peng Y, Zhang Y, Mao Q, Wen H, Tang J, Yuan X Clinical Interventions in Aging 2026, 21:613798 Published Date: 24 August 2026

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: oa-html ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

openalex
last seen: 2026-10-07T06:01:52.021920+00:00
License: CC0 · commercial use OK