Developing symptom-based predictive models of endometriosis as a clinical screening tool: results from a multicenter study

article OA: gold CC0 ⤵ 120 in-corpus citations
AI-generated summary by gemini-2.5-flash-lite, 2026-06-07

Symptom-based models accurately predicted stage III/IV endometriosis in symptomatic women, but only predicted any-stage endometriosis with moderate accuracy, aiding surgical investigation prioritization.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-06, 2026-06-12 · read from full text

This multicenter, two-phase clinic-based study (19 hospitals in 13 countries) prospectively recruited 1,396 pre-menopausal women undergoing diagnostic laparoscopy for symptoms such as dysmenorrhoea, dyspareunia, nonmenstrual pelvic pain, menstrual dyschezia, and/or infertility, and excluded women with recent hormonal medication, pregnancy, amenorrhoea, or prior endometriosis diagnosis. Participants completed a 25-item, language-adapted preoperative questionnaire covering histories and pelvic pain intensity/frequency; experienced gynecologists recorded laparoscopic findings, and endometriosis cases were defined by visual diagnosis and staged using revised American Fertility Society criteria. The authors developed symptom-based logistic regression models to predict any-stage and stage III/IV endometriosis, with performance measured by AUC and calibrated cutoffs, and found that models incorporating ultrasound evidence generally explained more variability and discriminated better in validation than models without ultrasound, though external validation reduced performance in some full models. This paper is centrally about endometriosis — developing symptom-based predictive models for clinical screening of endometriosis.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

OBJECTIVE: To generate and validate symptom-based models to predict endometriosis among symptomatic women prior to undergoing their first laparoscopy. DESIGN: Prospective, observational, two-phase study, in which women completed a 25-item questionnaire prior to surgery. SETTING: Nineteen hospitals in 13 countries. PATIENT(S): Symptomatic women (n = 1,396) scheduled for laparoscopy without a previous surgical diagnosis of endometriosis. INTERVENTION(S): None. MAIN OUTCOME MEASURE(S): Sensitivity and specificity of endometriosis diagnosis predicted by symptoms and patient characteristics from optimal models developed using multiple logistic regression analyses in one data set (phase I), and independently validated in a second data set (phase II) by receiver operating characteristic (ROC) curve analysis. RESULT(S): Three hundred sixty (46.7%) women in phase I and 364 (58.2%) in phase II were diagnosed with endometriosis at laparoscopy. Menstrual dyschezia (pain on opening bowels) and a history of benign ovarian cysts most strongly predicted both any and stage III and IV endometriosis in both phases. Prediction of any-stage endometriosis, although improved by ultrasound scan evidence of cyst/nodules, was relatively poor (area under the curve [AUC] = 68.3). Stage III and IV disease was predicted with good accuracy (AUC = 84.9, sensitivity of 82.3% and specificity 75.8% at an optimal cut-off of 0.24). CONCLUSION(S): Our symptom-based models predict any-stage endometriosis relatively poorly and stage III and IV disease with good accuracy. Predictive tools based on such models could help to prioritize women for surgical investigation in clinical practice and thus contribute to reducing time to diagnosis. We invite other researchers to validate the key models in additional populations.
Full text 21,933 characters · extracted from pmc-nxml · 3 sections · click to expand

Results

As shown in Supplemental Figure 1 , 771 (phase I) and 625 (phase II) of the women recruited met the inclusion criteria and had complete surgical information at the close of the study. Among participants, the proportions of cases at centers varied from 35% in Ibadan to 97% in Guangzhou. Supplemental Table 2 shows the average age of case and control women in both phases and the frequency of pathology found at surgery. Endometriosis was diagnosed in 360 (46.7%) women in phase I and 364 (58.2%) women in phase II. The results of the best-fitting final models for any-stage and stage III and IV endometriosis, respectively, are shown in Supplemental Table 3 and Table 1 , respectively. The results of the reduced models (only retaining variables with consistent evidence in phases I and II [see Methods]) are not shown, but these variables are highlighted in bold. The full “any-stage no ultrasound” model 1 (see Supplemental Table 3) had 26 variables but 12 of these (46%) were dropped in the reduced model 2 (see Methods). Similarly, 9 of 23 variables (39%) in the full “any-stage ultrasound” model 3 were dropped in the corresponding reduced model 4. Notably, menstrual dyschezia (pain on opening bowels during periods) and a medical history of benign ovarian cysts were most strongly associated with any-stage endometriosis in models with ultrasound (Phase II OR = 3.12, 95% CI = 1.07–9.10, P =.037, and OR = 2.92, 95% CI = 1.35–6.30, P =.006, respectively) and without ultrasound (Phase II OR = 3.47, 95% CI = 1.40–8.57, P =.007, and OR = 4.15, 95% CI = 2.19–7.86, P <.001, respectively). Rectal bleeding during menstruation, IBS (Rome III), unspecified functional bowel disorder, duration of smoking, subfertility due to blocked tubes, and ethnicity (Asian/Oriental and other/mixed) were inconsistently associated with endometriosis in both any-stage endometriosis models. The any-stage ultrasound model 3 explained substantially more variability in endometriosis than the any-stage no ultrasound model 1 (Nagelkerke's R 2  = 0.54 vs. 0.44 in phase I). Although the any-stage model including ultrasound (model 3) retained its R 2 value of 0.54 in phase II, its value dropped to 0.30 for model 1, indicating a better predictive performance of models, including ultrasound evidence. The full stage III and IV no ultrasound model 5 ( Table 1 ), had 23 variables, but 6 of these were dropped in the reduced model 6. In comparison, 3 of 18 variables in the full stage III and IV ultrasound model 7 were dropped in the corresponding reduced model 8. Menstrual dyschezia, a medical history of benign ovarian cysts, and Black ethnicity were most strongly associated with stage III and IV endometriosis, whereas other/mixed ethnicity was the only variable inconsistently associated with endometriosis in both any-stage endometriosis models in phases I and II. Although model 7 including ultrasound explained greater variability in diagnosis of stage III and IV endometriosis than model 5 without ultrasound (Nagelkerke's R 2  = 0.57 vs. 0.47 in phase I), the drop in R 2 was comparable for both models (to R 2  = 0.51 vs. 0.42). All four full predictive models fitted the data well as shown by the Hosmer and Lemeshow test P values (all  P >.05). All four full models were evaluated for their predictive performance in the phase II data. As expected, ultrasound evidence alone showed high sensitivity but very low specificity in the prediction of any or stage III or IV endometriosis ( Table 2 ), that is, positive ultrasound evidence was a good predictor of the presence of endometriosis, but negative evidence was a very poor predictor of absence of disease. As shown in Table 2 and Figure 2 , the full any-stage no ultrasound model 1 had good discrimination in phase I data (AUC = 84.2, 95% CI = 81.1–87.0, P <.0001), but its performance was reduced substantially when applied to the phase II validation data set (AUC = 68.3, 95% CI = 63.9–72.4, P <.0001). Predictive ability in phase II data was somewhat improved in the reduced model 2 to AUC = 72.2 (95% CI = 68.1–76.1, P <.001), although the optimal model cut-off of 0.59 would provide a low sensitivity of 54%. Similarly, the full any-stage ultrasound model 3 had good discrimination in phase I data (AUC = 87.3, 95% CI = 84.2–90.0, P <.0001), and although some reduction in its performance was evident when applied to phase II data (AUC = 80.0, 95% CI = 75.6–83.3, P <.0001), this reduction was not as substantial as for the any-stage no ultrasound model 1. The corresponding reduced model 4 improved predictive ability in phase II data to AUC = 85.1 (95% CI = 81.5–88.2, P <.001), its optimal model cut-off of 0.51 providing a sensitivity of 80% and a specificity of 77% (cf. respective values of 0.80, 58% and 89% for the full model). The full stage III or IV no ultrasound model 5 had good discrimination in phase I data (AUC = 87.3, 95% CI = 84.3–89.8, P <.0001) and retained good predictive power when applied to the phase II data set (AUC = 83.3, 95% CI = 79.6–86.6, P <.0001). Notably, the reduced model 6 did not offer improvement in model performance (AUC = 81.1, 95% CI = 77.2–84.6, P <.0001). The reduced model optimal cut-off was 0.24, with a sensitivity of 74% and specificity of 78% (cf. respective values of 0.45, 71%, and 85% for the full model). Similarly, the stage III or IV ultrasound model 7 had good discrimination in both phase I (AUC = 90.8, 95% CI = 88.1–93.0, P <.0001) and phase II data (AUC = 84.9, 95% CI = 081.4–88.0, P <.0001). The corresponding reduced model 8 did not improve model performance in phase II data (AUC = 85.5, 95% CI = 82.0–88.5, P <.0001). The reduced model optimal cut-off was 0.29, with a sensitivity of 80% and a specificity of 80% (cf. respective values of 0.24, 82%, and 76% for the full model).

Discussion

In this multicenter study of symptomatic premenopausal women presenting with symptoms that were potentially indicative of endometriosis, we show that a combination of symptom characteristics and variables in the medical history, with or without ultrasound evidence of cysts/nodules, can predict the finding of stage III and IV endometriosis at laparoscopy with reasonably good accuracy. The best-fitting predictive model included, along with ultrasound evidence, menstrual dyschezia, ethnicity, and a history of benign ovarian cysts as the variables with the strongest predictive performance. These variables are mostly disease risk factors reported in previous studies. Specifically, menstrual dyschezia is strongly associated with deep infiltrating endometriosis, a severe form of the disease (27) , which had a relatively low prevalence in our clinical population (7.0% in phase I, and 7.8% in phase II). The positive association of stage III and IV endometriosis with Black ethnicity, however, conflicts with previous reports which suggest that White and Asian women have a greater risk of disease (28) , although these reports generally relate to any-stage, rather than stage III and IV, endometriosis. In this study, we report the development in one sample population, and validation in another drawn from the same source population, of models predicting [1] any-stage and [2] stage III and IV, endometriosis, with or without ultrasound scan evidence. The ability to predict any-stage endometriosis in the model excluding ultrasound evidence was generally poor (AUC = 68.3). Some improvement (AUC = 72.2) could be gained by removing from the model variables with inconsistent association with endometriosis across phase I and II data, but the external validity of the reduced model would need to be evaluated in an independent data set. When including ultrasound scan evidence, the prediction of any-stage endometriosis was improved (AUC = 80.0), but the optimal model cut-off results in a relatively low sensitivity (58%), which reduces the utility of the model as a potential clinical screening tool. The poor predictability of any-stage endometriosis is not surprising given that similar findings are reported for models based on serum markers (29) , and stage I (minimal) endometriosis is considered pathogenetically to be different to stage III and IV disease (30–33) . In contrast to our findings, both any-stage and stage III and IV endometriosis were reported to be predictable from the medical history of 1,079 prospectively recruited subfertile women in Portugal (34) . However, as the predictive models were not validated in an independent data set, their findings should be interpreted with some caution. The models predicting stage III and IV endometriosis (±ultrasound evidence) showed much better performance. However, the model that included ultrasound evidence showed better performance, which is not surprising given the value of ultrasound to diagnose ovarian endometriomas (35) and deep infiltrating endometriosis affecting the bowel (36) . Optimal model cut-offs resulted in sensitivities of 70.9% and 82.3%, and specificities of 84.7% and 75.8%, for models excluding and including ultrasound evidence, respectively. In contrast to any-stage models, only marginal improvement could potentially be gained by excluding variables with inconsistent association with stage III and IV endometriosis across the data from phases I and II; this relative consistency in association across phases provided further evidence of the superior predictability of stage III and IV endometriosis over “any stage” disease. The stage III and IV endometriosis models could, in addition to ultrasound and physical examination, be used to prioritize women presenting with symptoms for laparoscopy in clinical practice, that is, mirroring the setting in which the present study was conducted, or to initiate medical therapies sometimes reserved until a surgical diagnosis of endometriosis has been made. To what extent the models have predictive power in other settings (e.g., self-selected women with pelvic pain symptoms in the general population) is unknown, and therefore the utility of the tool should not be advocated for this purpose. Indeed, although identifying a noninvasive diagnostic test for endometriosis is an explicit priority in endometriosis research, other authors have cautioned that such a tool could be misappropriated as a population screening tool for a disease that may not fit a population screening model (37) . A potential argument against the use of the models to prioritize symptomatic women for surgery is that a high prevalence of other pathologies was found among controls in this study (72%), which could warrant surgical intervention. However, whether surgical intervention would be deemed appropriate for these pathologies, or indeed whether they were likely to be the underlying reason for the symptoms or a coincidental finding, is a matter for debate. We believe that the stage III and IV prediction models in particular are potentially useful clinically, as the likely presence of moderate/severe disease would be a good basis for prioritization of surgical exploration and intervention. As far as we know, this is the first study to use robust modeling techniques for model generation, followed by external validation to generate and validate symptom-based predictive models of endometriosis in a large prospectively recruited cohort of women across different countries and ethnicities. Previous attempts, focused on subtypes of endometriosis (13, 14) , have been hampered by small sample sizes (11) , and failed to validate models in populations independent to those from which model parameters were generated (34) . The enrollment of women from diverse backgrounds according to a uniform set of criteria potentially addresses issues with the global utility of clinical prediction of endometriosis arising from inconsistencies in disease definition across studies and population. We invite other research groups to validate the key models in this paper in additional populations, as well as in subgroups that may be of specific clinical or population-based interest (e.g., those women who had infertility as the only surgical indication, who had biopsy-proven disease, or who were of a particular ethnicity). Although the WHSS was designed to improve on the limitations of earlier studies, it had itself potential limitations. First, endometriosis was diagnosed visually, without histologic confirmation, although this followed the European Society of Human Reproduction and Embryology guideline (23) , based on the premise that negative histology does not exclude the presence of disease. Consequently, disease status may have been inappropriately assigned; however, participating hospitals were experienced in diagnosing endometriosis. In a separate diagnostic validation study, 29 surgeons from the participating centers viewed nine standardized videos to allow, in a blinded manner, the assessment of consistency in diagnosis and staging of disease. Preliminary analysis suggested substantial inter-rater agreement in disease identification and staging (both Fleiss k  > 0.60; C. Becker and K. May, unpublished data). Second, the generation of the models with reduced numbers of variables was based on both phase I and phase II data. More parsimonious models are always preferable (38) ; however, because they are partly based on phase II evidence, they would require additional external validation. Third, although a strength of the study was that results were generated using data from a wide variety of clinical centers worldwide covering a range of patient profiles, this meant that the results may have been affected to some extent by selection bias possibly arising from [1] differential frequency of concomitant pathologies, in particular the higher proportion of women with nonendometriotic adhesions amongst controls compared to cases (32.4% vs. 14.7% in phase I), and [2] the significant variations across centers in proportions of cases in the sample populations. This is another reason why we call for further independent validation of the models in additional clinical populations, which are likely to each have their own unique patient population. In conclusion, the diagnostic delay, high investigation costs, and personal suffering associated with endometriosis might be reduced by access to a screening tool that predicts endometriosis with good accuracy in women presenting in a clinical setting. Although prediction of any-stage endometriosis is relatively poor, the symptom-based models developed and validated in this study predict stage III and IV endometriosis with a good degree of accuracy. They suggest that such a tool might help to prioritize women for surgical investigation in gynecologic practice.

Materials|Methods

The WHSS was a two-phase (model development/validation), clinic-based study in 19 hospitals in 13 countries. Between September 2008 and January 2010, we prospectively recruited 1,396 consecutive pre-menopausal women, aged 18–45, undergoing diagnostic laparoscopy because of at least one of the following symptoms: dysmenorrhoea (34.0% in phase I vs. 31.8% in phase II), dyspareunia (12.3% in phase I vs. 14.5% in phase II), nonmenstrual pelvic pain (36.1% in phase I vs. 37.0% in phase II), menstrual dyschezia (6.9% in phase I vs. 8.0% in phase II), or infertility (56.5% in phase I vs. 47.5% in phase II). Women with a previous surgical diagnosis of endometriosis, amenorrhoea or current pregnancy, or who had taken hormonal medication (including combined oral contraception) within the previous 3 months, were excluded. During the model generating phase I (September 2008 to June 2009), consenting women who met the inclusion criteria completed, prior to their scheduled surgery, a 25-item self-administered questionnaire in their own language ( www.endometriosisfoundation.org/WERF-WHSS-Questionnaire-English.pdf ). During the model validating phase II (July 2009 to January 2010), premenopausal women, aged 18–45 years, were recruited using the same inclusion and exclusion criteria as in phase I; they completed the same 25-item questionnaire before their surgery. The questionnaire incorporated items to elicit women's past medical, obstetric, and family histories, as well as items to evaluate the intensity and frequency of pelvic pain. Pelvic pain intensity was assessed on 11-point numerical pain rating scales (20) ranging in possible values from 0 (no pain) to 10 (worst possible pain). The questionnaire also included standardized questions previously validated in women with pelvic pain or other symptom groups. These instruments included [1] the IBS Rome III questionnaire to identify women with pelvic pain symptoms due to irritable bowel syndrome (21) and [2] standardized pelvic pain symptom assessment used in earlier studies in Oxford (22, 23) . The questionnaire also asked for sociodemographic, lifestyle, and physical attributes. Experienced gynecol-ogists recorded the laparoscopic findings in a standard manner ( http://www.endometriosisfoundation.org/WERF-GSWH-WHSS-surgical-sheet.pdf ). For those women who had preoperative pelvic ultrasound (84.7% and 92.2% of women in phases I and II, respectively), imaging findings were recorded on the surgical sheet. Cases were defined as women in the study populations who, at laparoscopy, were found to have endometriosis—diagnosed on visual evidence alone according to the European Society of Human Reproduction and Embryology guideline (1) and staged using the revised American Fertility Society classification: I (minimal), II (mild), III (moderate), or IV (severe) (24) . Controls were women in the study populations without endometriosis (with or without other diagnoses) at laparoscopy. The Mid- and South Buckinghamshire Research Ethics Committee in the United Kingdom approved the study, followed by approval from all the local ethics committees. Women who had [1] any stage of endometriosis and [2] stage III or IV endometriosis were compared to controls. To compare cases and controls on categorical variables, Pearson's χ 2 tests or the Fisher's exact test were used where appropriate. Continuous parametric variables were assessed using the Student's t -test and nonparametric variables with the Mann-Whitney U test. As the outcome of interest (presence of any-stage and stage III or IV endometriosis) was binary in nature, logistic regression was used for the predictive modeling. The WHSS questionnaire contained more than 200 variables, which, if entered into one logistic regression model, would have resulted in over fitting of models to the data. As a guide, in the final model, the number of degrees of freedom (df) should not exceed approximately 10% of the number of observations in the smaller outcome category (25) . We therefore employed a tiered approach to building the predictive model in phase I ( Fig. 1 ). First, groups of clinically related variables were assessed in a multivariate logistic regression framework, to assess which variables within each group showed little or no association with endometriosis, and could therefore be excluded. For each group, this was done iteratively by first excluding variables for which tests of association with endometriosis resulted in significance levels of P >.5, progressing to dropping variables with P >.2, resulting in a submodel for each of the groups. In each group, the goodness of fit of the submodel to the data was tested using the Hosmer and Lemeshow test (26) , whereas the drop in Nagelkerke's R 2 (assessing the disease variance explained by the variables in the model) when removing variables was not to exceed 10%. Variables from each of the submodels were then included in the complete models, which were again reduced iteratively by considering [1] the Hosmer and Lemeshow test for goodness of fit, [2] the drop in Nagelkerke's R 2  when removing variables, and [3] the significance of the association of each variable in the model with endometriosis (dropping variables from P >.5 down to P >.2). This resulted in best-fitting final models for the prediction of any-stage and stage III and IV endometriosis based on the phase I data. For each of the two diagnostic outcomes, a model including and excluding preoperative ultrasound evidence of cysts/nodules was generated ( Supplemental Table 1 ). Final models were subsequently fitted to phase II data to assess their predictive performance (see model validation below). In addition, reduced models were generated, which excluded variables in the best-fitting final models that had opposite directions of effect in phase I and II data sets. These reduced models were also assessed for performance; thus, a total of eight model types were generated and assessed ( Supplemental Table 1 ). To illustrate the performance of the derived predictive models in the phase I population data, and the relative drop in performance when externally validating in phase II data, the accuracy of the models in predicting any-stage and stage III and IV endometriosis was assessed by analyzing the receiver operating characteristic (ROC) curve in both phase I and II. The ROC curve displays the relationship between sensitivity and 1-specificity and the area under the ROC curve (AUC) depicts how well the model distinguishes women with and without endometriosis; a model with a greater AUC has a better-performing risk function. Model sensitivity, specificity, and positive and negative likelihood ratios were also calculated, and the best model cut-off points were considered to be those that corresponded to the highest sum of specificity and sensitivity. All univariate and logistic regression analyses were done using the Statistical Package for the Social Sciences 16.0 (SPSS, Inc.); prediction analyses in both phase I and II populations were conducted within the binary logistic module in SPLUS 6.0 (TIBCO Software, Inc.); and ROC analysis using MedCalc 11.6 (MedCalc Software).

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Condition tags

endometriosis

MeSH descriptors

Endometriosis Adolescent Adult Endometriosis Female Humans Logistic Models Middle Aged Models, Theoretical Prospective Studies

Citation neighborhood

Papers in the corpus that this work cites (lower rings, blue) and that cite this one (upper rings, green). Dot size scales with the paper's in-corpus citation count — bigger dot = more influential within the endo/adeno field. Click a dot to open that paper. [ expand to 2 hops ] — adds papers reached through this work's immediate citers/citees. Heavier; up to 60 extra dots.

References (43)

Cited by (50)

Source provenance

europepmc
last seen: 2026-09-21T06:08:07.822426+00:00
openalex
last seen: 2026-06-10T17:14:06.276822+00:00
pubmed
last seen: 2026-05-13T22:16:11.197438+00:00
License: CC0 · commercial use OK