Feature-Limited Performance in Machine Learning Prediction of Endometriosis from Clinical Symptoms

In: Medical Research Archives · 2025 · vol. 14(2) · doi:10.18103/mra.v14i2.7221 · W7133320939
OA: diamond CC0
⚙ AI-generated summary by qwen3.7-flash, 2026-10-06 ⓘ

Machine learning models for endometriosis prediction using clinical symptoms plateaued at an AUC of 0.65–0.67, indicating that performance is limited by the information content of symptoms rather than model architecture or data size.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

⚙ AI-generated deep summary by qwen3.7-flash, 2026-08-27 · read from full text ⓘ

This study evaluated five machine learning architectures to predict endometriosis using only six base clinical symptoms from 10,000 patient records. The models, including logistic regression and gradient-boosted decision trees, converged on a similar test AUC of approximately 0.67, with feature engineering providing negligible performance gains. Learning curves plateaued early, indicating that prediction limits are constrained by the intrinsic information content of symptom-based features rather than model complexity or sample size. This paper is centrally about endometriosis — specifically the diagnostic challenges and limitations of noninvasive, symptom-based machine learning screening for the condition.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Background: Endometriosis affects approximately 10% of women of reproductive age worldwide, yet diagnosis remains challenging due to nonspecific symptoms and reliance on invasive laparoscopic confirmation, resulting in diagnostic delays averaging seven to ten years. Machine learning approaches have shown promise for noninvasive screening, but fundamental questions remain regarding whether performance limitations arise from model architecture constraints, insufficient training data, or intrinsic information ceilings imposed by symptom-based clinical features. Methods: Five machine learning architectures (logistic regression with L1 regularization, support vector machines with radial basis function kernels, gradient-boosted decision trees, random forests, and deep neural networks) were systematically compared for endometriosis prediction using six base clinical variables from 10,000 patient records. The pipeline incorporated stratified data splitting (80%/10%/10% train/validation/test), label noise mitigation through ambiguity-based instance weighting, cost-sensitive learning prioritizing false negative reduction, and cross-validated threshold optimization. Feature engineering expanded the base features to 21 variables through interaction terms, polynomial transformations, and discretized bins. Learning curve analysis assessed whether performance was constrained by training set size or feature informativeness. Results: All five model architectures converged to similar test performance (AUC range: 0.653–0.674), with the selected logistic regression model achieving test AUC of 0.674, recall of 0.566, precision of 0.562, and specificity of 0.696 at the Youden-optimized threshold. Feature engineering yielded negligible improvements, with mean test AUC changing by only 0.002 between baseline (6 features) and engineered (21 features) configurations. Learning curves plateaued beyond 60% of training data, with training and validation AUC converging to 0.667 and 0.641 respectively, and the gap narrowing from 0.058 to 0.026. Conclusions: The convergence of multiple model families to similar performance limits, minimal gains from feature engineering, and plateaued learning curves provide empirical evidence that model performance is constrained by the information content of symptom-based clinical features rather than by model architecture, sample size, or feature representation sophistication. The observed AUC ceiling of approximately 0.65–0.67 aligns with published literature on symptom-based endometriosis screening and indicates that clinically actionable discrimination performance requires data enrichment through incorporation of imaging findings, biomarkers, or genomic risk factors rather than algorithmic innovation.
Full text 9,704 characters · extracted from oa-doi-fallback · 5 sections · click to expand

Abstract

Background: Endometriosis affects approximately 10% of women of reproductive age worldwide, yet diagnosis remains challenging due to nonspecific symptoms and reliance on invasive laparoscopic confirmation, resulting in diagnostic delays averaging seven to ten years. Machine learning approaches have shown promise for noninvasive screening, but fundamental questions remain regarding whether performance limitations arise from model architecture constraints, insufficient training data, or intrinsic information ceilings imposed by symptom-based clinical features.

Methods

Five machine learning architectures (logistic regression with L1 regularization, support vector machines with radial basis function kernels, gradient-boosted decision trees, random forests, and deep neural networks) were systematically compared for endometriosis prediction using six base clinical variables from 10,000 patient records. The pipeline incorporated stratified data splitting (80%/10%/10% train/validation/test), label noise mitigation through ambiguity-based instance weighting, cost-sensitive learning prioritizing false negative reduction, and cross-validated threshold optimization. Feature engineering expanded the base features to 21 variables through interaction terms, polynomial transformations, and discretized bins. Learning curve analysis assessed whether performance was constrained by training set size or feature informativeness.

Results

All five model architectures converged to similar test performance (AUC range: 0.653–0.674), with the selected logistic regression model achieving test AUC of 0.674, recall of 0.566, precision of 0.562, and specificity of 0.696 at the Youden-optimized threshold. Feature engineering yielded negligible improvements, with mean test AUC changing by only 0.002 between baseline (6 features) and engineered (21 features) configurations. Learning curves plateaued beyond 60% of training data, with training and validation AUC converging to 0.667 and 0.641 respectively, and the gap narrowing from 0.058 to 0.026.

Conclusions

The convergence of multiple model families to similar performance limits, minimal gains from feature engineering, and plateaued learning curves provide empirical evidence that model performance is constrained by the information content of symptom-based clinical features rather than by model architecture, sample size, or feature representation sophistication. The observed AUC ceiling of approximately 0.65–0.67 aligns with published literature on symptom-based endometriosis screening and indicates that clinically actionable discrimination performance requires data enrichment through incorporation of imaging findings, biomarkers, or genomic risk factors rather than algorithmic innovation. Article Details The Medical Research Archives grants authors the right to publish and reproduce the unrevised contribution in whole or in part at any time and in any form for any scholarly non-commercial purpose with the condition that all publications of the contribution include a full citation to the journal as published by the Medical Research Archives.

References

2. Yan H, Li X, Dai Y, et al. Global, regional, and national burdens of endometriosis from 1990 to 2021: a trend analysis. Front Med (Lausanne). 2025; 12:1562196. doi:10.3389/fmed.2025.1562196 3. Parasar P, Ozcan P, Terry KL. Endometriosis: epidemiology, diagnosis and clinical management. Curr Obstet Gynecol Rep. 2017;6(1):34-41. doi:10. 1007/s13669-017-0187-1 4. Giudice LC, Kao LC. Endometriosis. Lancet. 2004 ;364(9447):1789-1799. doi:10.1016/s0140-6736(0 4)17403-5 5. Li W, Feng H, Ye Q. Factors contributing to the delayed diagnosis of endometriosis—a systematic review and meta-analysis. Front Med (Lausanne). 2025;12:1576490. doi:10.3389/fmed.2025.1576490 6. De Corte P, Klinghardt M, von Stockum S, Heinemann K. Time to diagnose endometriosis: current status, challenges and regional characteristics—a systematic literature review. BJOG. 2024;132(2):118-130. doi: 10.1111/1471-0528.17973 7. Dantkale KS, Agrawal M. A comprehensive review of the diagnostic landscape of endometriosis: assessing tools, uncovering strengths, and acknowledging limitations. Cureus. 2024;16:e56978. doi:10.7759/ cureus.56978 8. Davenport S, Smith D, Green DJ. Barriers to a timely diagnosis of endometriosis: a qualitative systematic review. Obstet Gynecol. 2023;142(3):57 1-583. doi:10.1097/aog.0000000000005255 9. Pascoal E, Wessels JM, Aas-Eng MK, et al. Strengths and limitations of diagnostic tools for endometriosis and relevance in diagnostic test accuracy research. Ultrasound Obstet Gynecol. 2022;60(3):309-327. doi:10.1002/uog.24892 10. Harzif AK, Nurbaeti P, Putri AS, et al. Factors associated with delayed diagnosis of endometriosis: a systematic review. J Endometr Pelvic Pain Disord. 2024;12:1291120. doi:10.1177/22840265241291120 11. Nnoaham KE, Hummelshoj L, Kennedy SH, Jenkinson C, Zondervan KT. Developing symptom-based predictive models of endometriosis as a clinical screening tool: results from a multicenter study. Fertil Steril. 2012;98(3):692-701.e5. doi:10.1 016/j.fertnstert.2012.04.022 12. Stegmann BJ, Funk MJ, Sinaii N, et al. A logistic model for the prediction of endometriosis. Fertil Steril. 2009;91(1):51-55. doi:10.1016/j.fertnst ert.2007.11.038 13. Tore U, Abilgazym A, Asunsolo-del Barco A, et al. Diagnosis of endometriosis based on comorbidities: a machine learning approach. Biomedicines. 2023; 11(11):3015. doi:10.3390/biomedicines11113015 14. Nouri B, Hashemi SH, Ghadimi DJ, Roshandel S, Akhlaghdoust M. Machine learning-based detection of endometriosis: a retrospective study in a population of Iranian female patients. Int J Fertil Steril. 2024;18(4):1519. doi:10.22074/ijfs.202 4.2009338.1519 15. Goldstein A, Cohen S. Self-report symptom-based endometriosis prediction using machine learning. Sci Rep. 2023;13(1):32761. doi:10.1038/s 41598-023-32761-8 16. Zhao N, Hao T, Zhang F, et al. Application of machine learning techniques in the diagnosis of endometriosis. BMC Womens Health. 2024;24(1): 33342. doi:10.1186/s12905-024-03334-2 17. Kucukakcali Z, Akbulut S, Colak C. Prediction of genomic biomarkers for endometriosis using the transcriptomic dataset. World J Clin Cases. 2025; 13(20):104556. doi:10.12998/wjcc.v13.i20.104556 18. Zhang H, Zhang H, Yang H, Shuid AN, Sandai D, Chen X. Machine learning-based integrated identification of predictive combined diagnostic biomarkers for endometriosis. Front Genet. 2023; 14:1290036. doi:10.3389/fgene.2023.1290036 19. Blass I, Sahar T, Shraibman A, Ofer D, Rappoport N, Linial M. Revisiting the risk factors for endometriosis: a machine learning approach. J Pers Med. 2022;12(7):1114. doi:10.3390/jpm12071114 20. Shrestha P, Shrestha B, Shrestha J, Chen J. Current status and future potential of machine learning in diagnostic imaging of endometriosis: a literature review. J Nepal Med Assoc. 2025;63(283) :205-211. doi:10.31729/jnma.8897 21. Ramadan ZM, Mouazen SM, Khan SS, Khan S, Farag NS. Machine learning in the early detection of endometriosis: a literature review on symptom clustering and imaging integration. Precis Future Med. 2025;9(3):117-128. doi:10.23838/pfm.2025.00177 22. Mary JJ, Shanthi V. Early detection of endometriosis: integrating medical imaging and machine learning algorithms for non-invasive diagnosis. Int Res J Adv Eng Manag. 2025;3(3):591-595. doi:10.47392/irjaem.2025.0095 23. Cao S, Li X, Zheng X, Zhang J, Ji Z, Liu Y. Identification and validation of a novel machine learning model for predicting severe pelvic endometriosis: a retrospective study. Sci Rep. 2025;15(1):96093. doi:10.1038/s41598-025-96093-5 24. Balogh DB, Hudelist G, Blizņuks D, et al. The use of machine learning for early diagnosis of endometriosis based on patient self-reported data—study protocol of a multicenter trial. PLoS One. 2024;19(5):e0300186. doi:10.1371/journal. pone.0300186 25. Chrysa N, Lamprini C, Maria-Konstantina C, Constantinos K. Deep learning improves accuracy of laparoscopic imaging classification for endometriosis diagnosis. J Clin Med Surg. 2024;4(1):1137. doi:10.52768/2833-5465/1137 26. Chetcuti K, Chilungulo C. Case-based review of low-field MRI in resource-constrained settings: a clinical perspective from Malawi. BJR Open. 2024;7(1):tzaf028. doi:10.1093/bjro/tzaf028 27. Lyimo BM, Popkin-Hall ZR, Giesbrecht DJ, et al. Potential opportunities and challenges of deploying next generation sequencing and CRISPR-Cas systems to support diagnostics and surveillance towards malaria control and elimination in Africa. Front Cell Infect Microbiol. 2022;12:757844. doi:10.3389/fcimb.2022.757844 28. Helmy M, Awad M, Mosa KA. Limited resources of genome sequencing in developing countries: challenges and solutions. Appl Transl Genom. 2016;9:15-19. doi:10.1016/j.atg.2016.03.003 29. Tekola-Ayele F, Rotimi CN. Translational genomics in low- and middle-income countries: opportunities and challenges. Public Health Genomics. 2015;18(4):242-247. doi:10.1159/000433518 30. Malkin RA. Barriers for medical devices for the developing world. Expert Rev Med Devices. 2007;4(6):759-763. doi:10.1586/17434440.4.6.759 31. Husain G, Nasef D, Jose R, et al. SMOTE vs. SMOTEENN: a study on the performance of resampling algorithms for addressing class imbalance in regression models. Algorithms. 2025; 18(1):37. doi:10.3390/a18010037 32. Toma M. AI-Assisted Medical Diagnostics: A Clinical Guide to Next-Generation Diagnostics. New York, NY: Dawning Research Press; 2025. https://openlibrary.org/works/OL44048041W/. Accessed January 14, 2026.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Condition tags

endometriosis

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

openalex
last seen: 2026-06-10T17:14:06.276822+00:00
License: CC0 · commercial use OK