Comorbidity analysis and clustering of endometriosis patients using electronic health records

other OA: gold CC-BY-4.0
AI-generated summary by gemini-2.5-flash-lite, 2026-06-07

This study analyzed electronic health records from over forty thousand endometriosis patients and identified hundreds of associated conditions, revealing distinct patient subpopulations with varying comorbidity patterns.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-06, 2026-06-07 · read from full text

This paper used electronic health record data to analyze comorbidities and patient heterogeneity in 19,059 individuals with endometriosis at UCSF and 24,453 across five additional UC medical centers, comparing them with matched non-endometriosis controls via propensity score matching and estimating enriched diagnoses using odds ratios. Using unsupervised clustering (nearest-neighbors graph plus Leiden clustering) on diagnosis patterns, the authors identified endometriosis patient subgroups and assessed concordance of enriched conditions across data sources. They found 661 significantly enriched comorbidities at UCSF, including strong associations with uterine adenomyosis, pelvic peritoneal adhesions, pelvic pain, ovarian cysts, infertility, and autoimmune disease, with additional enrichment of conditions such as migraines, GERD, asthma, and vitamin D deficiency that were less commonly reported in smaller studies; 302 overlapping enriched conditions were statistically concordant in the UC-wide analysis. The paper’s main stated caveat is potential diagnostic overlap in coding (e.g., uterine adenomyosis vs uterine endometriosis), which could affect interpreted associations. This paper is centrally about endometriosis—comorbidity profiling and clustering of endometriosis patients using multicenter EHR data.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Endometriosis is a prevalent, complex, inflammatory condition associated with a diverse range of symptoms and comorbidities. Despite its substantial burden on patients, population-level studies that explore its comorbid patterns and heterogeneity are limited. In this retrospective case-control study, we analyze comorbidities from over forty thousand endometriosis patients across six University of California medical centers using de-identified electronic health record (EHR) data. We find hundreds of conditions significantly associated with endometriosis, including genitourinary disorders, neoplasms, and autoimmune diseases, with strong replication across datasets. Clustering analyses identify patient subpopulations with distinct comorbidity patterns, including psychiatric and autoimmune conditions. This study provides a comprehensive analysis of endometriosis comorbidities and highlights the heterogeneity within the patient population. Our findings demonstrate the utility of EHR data in uncovering clinically meaningful patterns and suggest pathways for personalized disease management and future research on biological mechanisms underlying endometriosis.
Full text 40,540 characters · extracted from pmc · 9 sections · click to expand

Author

Conceptualization, U.K., T.T.O., K.G., L.C.G., and M.S.; methodology, U.K., J.R., and M.S.; software, U.K.; investigation, U.K. and T.T.O.; writing – original draft, U.K., T.T.O., and B.D.Y.; writing – review and editing, all authors; supervision, J.C.I., J.O.-A., N.E., L.C.G., and M.S.; project administration, M.S.; funding acquisition, L.C.G. and M.S.

Results

We identified a total of 19,059 endometriosis patients at UCSF and an additional 24,453 endometriosis patients across University of California’s Health Data Warehouse (UCHDW) ( Figures 1 C; Table 1 ; Figure S1 ). These patients had a mean age of 52.6 years at UCSF and 46.5 years in the UCHDW. In both data sources, endometriosis patients were predominantly white (51.7% at UCSF and 53.9% in the UCHDW) with “Asian” and “other” being the next two most common classifications. Similarly, most patients (76.1% at UCSF and 71.8% in the UCHDW) were identified as not Hispanic or Latino. Most of our UCHDW patients come from records at two large centers (44.6% and 24.8%, respectively) with the remaining being split across the three other centers. Table 1 Patient characteristics at UCSF and across the other UC medical centers analyzed UCSF UC-wide Total ( n ) 19,059 24,453 Age (mean, SD) 52.6 (SD 14.3) 46.5 (SD 11.9) Gender  Female 18,995 (99.6%) 24,411 (99.8%)  Male 36 (0.2%) 34 (0.1%)  Unknown 21 (0.1%) 6 (0.0%)  Other 7 (0.0%) 2 (0.0%) Race  White 9,859 (51.7%) 13,168 (53.9%)  Asian 2,980 (15.6%) 2,767 (11.3%)  Other 2,327 (12.2%) 3,170 (13.0%)  Unknown 2,150 (11.3%) 2,628 (10.7%)  Black or African American 1,105 (5.8%) 1,404 (5.7%)  Native Hawaiian or other Pacific Islander 491 (2.6%) 121 (0.5%)  American Indian or Alaska Native 147 (0.8%) 112 (0.4%)  Multirace – 1,083 (4.4%) Ethnicity  Not Hispanic or Latino 14,502 (76.1%) 17,560 (71.8%)  Unknown 2,570 (13.5%) 2,378 (9.7%)  Hispanic or Latino 1,987 (10.4%) 4,515 (18.5%) Location  UC center #1 – 10,901 (44.6%)  UC center #2 – 6,069 (24.8%)  UC center #3 – 3,638 (14.9%)  UC center #4 – 2,986 (12.2%)  UC center #5 – 859 (3.5%) Patient characteristics at UCSF and across the other UC medical centers analyzed When considering the full condition set, our analysis revealed 661 significantly enriched (aggregated adjusted p value < 0.05) comorbidities at UCSF spanning across nearly all ICD chapters ( Figure 2 A), reflecting the diverse clinical presentations associated with endometriosis ( Table S1 ; Figure S2 ). Among the most significantly enriched conditions were uterine adenomyosis (odds ratio [OR] = 181, noting potential overlap in diagnostic coding with uterine endometriosis), pelvic peritoneal adhesions (OR = 51.1), non-inflammatory disorders of the female genital organs (OR = 30.2), pain in female pelvis (OR = 26.3), and cyst of ovary (OR = 16), all of which were also significant in the UCHDW ( Figure 2 B). We also see significant enrichments for female infertility (OR = 5) and general autoimmune disease (OR = 4.3), consistent with prior literature on endometriosis-associated infertility 18 and systemic inflammation. 19 Interestingly, we identified several conditions less commonly reported in smaller studies, including migraines 20 , 21 , 22 (OR = 4), gastroesophageal reflux disease 23 (OR = 3.6), asthma 24 (OR = 2.5), and vitamin D deficiency 25 (OR = 3.8). We found that 302 conditions of these conditions (45% of the complete set) were significantly enriched in the UCHDW data as well ( Table S2 ), with statistically significant correlation of the log-odds ratios (Pearson r = 0.864 and p = 2.38 × 10 −91 ) ( Figure 2 C). Figure 2 Results of comorbidity analysis (A) Manhattan plots of comorbidities observed at UCSF, organized by ICD chapter. (B) Strip plots of top twenty comorbidities (by aggregate p value) observed at UCSF, showing computed odds ratios across thirty replicate control groups. (C) Comparison to UC-wide analysis, showing number of significant overlapping comorbidities and concordance of odds ratios for overlapping conditions. See also Figures S2 and S3 and Tables S1 , S2 , S3 , and S4 . Results of comorbidity analysis (A) Manhattan plots of comorbidities observed at UCSF, organized by ICD chapter. (B) Strip plots of top twenty comorbidities (by aggregate p value) observed at UCSF, showing computed odds ratios across thirty replicate control groups. (C) Comparison to UC-wide analysis, showing number of significant overlapping comorbidities and concordance of odds ratios for overlapping conditions. See also Figures S2 and S3 and Tables S1 , S2 , S3 , and S4 . When considering the pre-endometriosis condition set, our analysis revealed 106 significantly enriched (aggregated adjusted p value < 0.05) comorbidities at UCSF ( Table S3 ; Figure S2 ). These conditions span a narrower range of ICD chapters, with most enriched conditions falling into those spanning the genitourinary system, related symptoms, and neoplasms ( Figure 2 A). Among the most significantly enriched conditions were cyst of ovary (OR = 6.6), dysmenorrhea (OR = 8.3), female genital organ symptoms (OR = 4.9), disorder of female genital organs (OR = 4.3), and pain in female pelvis (OR = 15.2) ( Figure 2 B). We also see a significant enrichment for increased cancer antigen 125 (OR = 17.9), consistent with prior literature describing the relationship between this biomarker and endometriosis. 26 Interestingly, the association between migraines and endometriosis remained significant even prior to the diagnosis of endometriosis (OR = 2). We found that 34 of these conditions (32% of the complete set) were significantly enriched in the UCHDW data as well ( Table S4 ), with statistically significant correlation of the log-odds ratios (Pearson r = 0.865 and p = 4.00 × 10 −11 ) ( Figure 2 C). The pre-endometriosis analysis in the UCHDW dataset found several notable negative associations, with odds ratios less than one, including hyperlipidemia (OR = 0.67) and mixed hyperlipidemia (OR = 0.67). Our unsupervised clustering analysis identified distinct subpopulations of endometriosis patients, characterized by shared diagnosis patterns. Specifically, we identified 21 clusters in the UCSF cohort and 26 clusters in the UCHDW cohort when analyzing the full condition set ( Figure 3 ; Tables S5 and S6 ). Similarly, the pre-endometriosis condition set yielded 31 clusters at UCSF and 41 clusters in the UCHDW cohort ( Figure 4 ; Tables S7 and S8 ). These clusters revealed diverse patterns of comorbidities, with some clusters being dominated by autoimmune disorders, others by pregnancy complications, and still others by psychiatric conditions. Annotation of clusters suggested biologically and clinically meaningful subgroupings, which may reflect underlying heterogeneity in the disease’s pathophysiology or healthcare utilization patterns. Comparing cluster concordance across our two independent datasets, we saw UC-wide clusters highlighted by pregnancy and cancer-related conditions when considering all diagnoses ( Figure 3 ) and pregnancy and urinary tract conditions when considering pre-endometriosis diagnoses ( Figure 4 ). Clusters at UCSF also highlighted a number of other disease categories, including skin disorders, renal disorders, and mental health conditions. Figure 3 Clustering results using all conditions At the top, a bipartite graph showing clusters from the UCSF and UC-wide datasets with overlapping enriched conditions. The legend below the graph shows how node size maps to the number of patients in each cluster, how edge size maps to the percentage of UCSF conditions that are found in the matched UC-wide cluster, and how the size of the node border relates to the number of uniquely enriched conditions found in each cluster. Below the legend, visualizations of clusters computed using all diagnoses overlaid on two-dimensional uniform manifold approximation and projection (UMAP) representations of endometriosis patients, at UCSF and UC-wide. Note that only clusters with overlapping conditions between the two datasets are labeled, with the rest colored in gray. To the right, cluster annotations for clusters represented in the bipartite graph. See also Tables S5 , S6 , and S10 . Figure 4 Clustering results using pre-endometriosis conditions At the top, a bipartite graph showing clusters from the UCSF and UC-wide datasets with overlapping enriched conditions. The legend below the graph shows how node size maps to the number of patients in each cluster, how edge size maps to the percentage of UCSF conditions that are found in the matched UC-wide cluster, and how the size of the node border relates to the number of uniquely enriched conditions found in each cluster. Below the legend, visualizations of clusters computed using pre-endometriosis diagnoses overlaid on two-dimensional UMAP representations of endometriosis patients, at UCSF and UC-wide. Note that only clusters with overlapping conditions between the two datasets are labeled, with the rest colored in gray. To the right, cluster annotations for clusters represented in the bipartite graph. See also Tables S7 , S8 , and S10 . Clustering results using all conditions At the top, a bipartite graph showing clusters from the UCSF and UC-wide datasets with overlapping enriched conditions. The legend below the graph shows how node size maps to the number of patients in each cluster, how edge size maps to the percentage of UCSF conditions that are found in the matched UC-wide cluster, and how the size of the node border relates to the number of uniquely enriched conditions found in each cluster. Below the legend, visualizations of clusters computed using all diagnoses overlaid on two-dimensional uniform manifold approximation and projection (UMAP) representations of endometriosis patients, at UCSF and UC-wide. Note that only clusters with overlapping conditions between the two datasets are labeled, with the rest colored in gray. To the right, cluster annotations for clusters represented in the bipartite graph. See also Tables S5 , S6 , and S10 . Clustering results using pre-endometriosis conditions At the top, a bipartite graph showing clusters from the UCSF and UC-wide datasets with overlapping enriched conditions. The legend below the graph shows how node size maps to the number of patients in each cluster, how edge size maps to the percentage of UCSF conditions that are found in the matched UC-wide cluster, and how the size of the node border relates to the number of uniquely enriched conditions found in each cluster. Below the legend, visualizations of clusters computed using pre-endometriosis diagnoses overlaid on two-dimensional UMAP representations of endometriosis patients, at UCSF and UC-wide. Note that only clusters with overlapping conditions between the two datasets are labeled, with the rest colored in gray. To the right, cluster annotations for clusters represented in the bipartite graph. See also Tables S7 , S8 , and S10 . In addition to a strong concordance of associations when comparing pre-endometriosis diagnoses to all diagnoses ( Figure S3 ), analysis of the patient clusters at UCSF reveals that certain groups of patients remain consistently clustered across both condition sets ( Figure 5 ). This suggests that these groups of patients, initially assigned to pre-endometriosis clusters associated with hyperlipidemia, mental health conditions, pregnancy, and anemia, might experience similar clinical trajectories after their endometriosis diagnosis. Figure 5 Changes in cluster membership before and after endometriosis diagnosis Alluvial plot showing change in cluster membership for patients at UCSF when considering pre-endometriosis diagnoses versus all diagnoses. Each cluster on the right (all diagnoses) is composed of subsets of clusters on the left (pre-endometriosis diagnosis), representing patient subsets. All clusters are labeled with their annotated themes. Changes in cluster membership before and after endometriosis diagnosis Alluvial plot showing change in cluster membership for patients at UCSF when considering pre-endometriosis diagnoses versus all diagnoses. Each cluster on the right (all diagnoses) is composed of subsets of clusters on the left (pre-endometriosis diagnosis), representing patient subsets. All clusters are labeled with their annotated themes.

Resource

Requests for further information and resources should be directed to and will be fulfilled by the lead contact, Marina Sirota ( [email protected] ). This study did not generate any new materials. The data that support the findings of this study are not openly available to individuals unaffiliated with UCSF due to the sensitivity of medical records, with the exception of collaborators. Individuals not affiliated with UCSF may set up an official collaboration with a UCSF-affiliated investigator by reaching out to the lead contact , Marina Sirota ( [email protected] ). UCSF-affiliated individuals may contact UCSF’s Clinical and Translational Science Institute ( [email protected] ) or UCSF’s Information Commons team for more information ( [email protected] ). UC-wide data are only available to UC researchers who have completed analyses in their respective UC first and have provided justification for scaling their analyses across UC health centers. Censored code for the analysis and visualizations in this study can be found at https://github.com/khanu263/comorbidities-clustering-endo . Any additional information required to reanalyze the data reported in this work paper is available from the lead contact upon request.

Discussion

Our analysis of comorbidities in endometriosis patients highlights findings that align closely with previous investigations. Many of the observed associations between endometriosis and its comorbidities, such as chronic pain conditions 27 and gastrointestinal diseases, 28 are consistent with the established literature. Several of the comorbidities identified in our analysis have also been implicated in large-scale genetic studies of endometriosis, reinforcing the biological plausibility of these associations. For example, a recent genome-wide association study identified shared genetic loci between endometriosis and migraines, suggesting overlapping neuroinflammatory pathways. 29 Similarly, recent Mendelian randomization studies support relationships between endometriosis and gastrointestinal conditions, including gastroesophageal reflux disease and irritable bowel syndrome. 23 , 30 A genetic overlap between endometriosis and asthma has also been reported, with enrichment in hormonal and immune-related pathways. 24 Our study complements these findings by providing phenotypic validation across large and diverse healthcare systems. These parallels across genomic and clinical domains underscore the multifactorial nature of endometriosis and its shared etiology with other systemic conditions. One of the strongest comorbid associations we observed was between endometriosis and adenomyosis. While this likely reflects a true relationship, given the shared pathophysiology, 31 it is also important to acknowledge potential ambiguity in diagnostic coding. In many EHR systems, adenomyosis and endometriosis are sometimes conflated, especially in ICD-based coding hierarchies. Although our analysis uses Systematized Nomenclature of Medicine (SNOMED) conditions, these were mapped from underlying diagnostic codes in the source EHRs after de-identification. As a result, subtle distinctions between conditions such as “endometriosis of uterus” and “uterine adenomyosis” may not be reliably preserved. We therefore interpret this association with caution, as it may represent a mixture of true co-occurrence and heterogeneity in coding practices. Future analyses incorporating clinical notes may be better suited to disentangle this relationship. In addition, our analysis identified several less-frequently reported associations, demonstrating the power of this data-driven approach. For instance, the protective associations with hyperlipidemia and mixed hyperlipidemia in the UCHDW dataset are particularly interesting in the context of a body of literature identifying statins as potential therapeutic avenues for endometriosis, 32 , 33 , 34 , 35 since patients with these conditions are likely taking statins as treatment. Similarly, the repeated appearance of migraines as a significant association, both before and after endometriosis diagnosis, underscores the emerging concept of repurposing drugs from comorbidities to treat endometriosis pain; in a recent study, Fattori et al. showed that a migraine drug decreased disease burden and pain in a mouse model of endometriosis. 36 These results provide further evidence of the broad and multifaceted clinical impact of endometriosis, underscoring the importance of a comprehensive approach to its management. Notably, many of these associations replicated across two separate data sources and time periods, reinforcing the robustness of our findings. This replication suggests that the observed patterns are not artifacts of a single dataset or limited to a specific population but instead represent generalizable trends across diverse healthcare environments. Such consistency enhances the credibility of our results and their relevance to broader patient populations. Through unsupervised clustering, we identified distinct subpopulations of endometriosis patients characterized by unique patterns of comorbidities. These clusters reveal insights into the heterogeneity of endometriosis, suggesting that patients may fall into clinically meaningful subgroups based on their associated diagnoses. While our clustering analyses were exploratory in nature, we observed substantial consistency across different data sources and diagnostic time frames. Notably, clusters associated with themes like pregnancy complications and cancers appeared across both UCSF and UC-wide cohorts and when analyzing all diagnoses or only pre-endometriosis conditions. This replication across settings provides evidence of cluster robustness. Though our clustering results were generally consistent at a thematic level, we did not attempt to track individual patient membership across clusterings derived from different condition sets, since the input data and resulting cluster structures vary substantially across these settings. As such, we focus our interpretation on group-level concordance rather than patient-level retention. These findings may serve as a foundation for future studies aiming to tailor treatments or management strategies to specific patient subgroups. For instance, these subgroups may be associated with differing patterns of drug utilization, providing a starting point for examining patient outcomes. Our study has several strengths worth highlighting. First, the use of data from multiple medical centers and time periods allows for a more comprehensive and generalizable analysis of endometriosis and its comorbidities. Second, the robust control replication method employed in our analyses mitigates the risk of false-positive associations, which is often a concern with studies of this nature. Finally, our use of EHR data provides a valuable perspective on real-world population patterns, which can inform clinical practice and policy. As with any research leveraging EHRs, our analysis is subject to inherent data issues, including missingness, patients moving between healthcare systems, and coding differences across institutions. Additionally, our selection criteria were deliberately permissive, as we defined endometriosis cases based on EHR-documented diagnoses rather than surgically confirmed cases. While this approach maximized the size and diversity of our sample, it may have introduced misclassification bias. Moreover, the case-control design of our study precludes any conclusions about causality or temporality. Another key limitation is the geographic context of our data. All participating institutions are part of the University of California health system, which primarily serves populations within a single state. Although this system includes both academic and community-based medical centers, patients accessing UC care may have higher socioeconomic status or better access to specialized care. As such, our findings may not fully capture the experiences of individuals in under-resourced settings, those without insurance, or those outside the United States healthcare system. Broader validation in other populations and healthcare contexts is warranted to assess the generalizability of these patterns. Despite these limitations, our study provides valuable insights into the comorbidities and heterogeneity of endometriosis. By leveraging data from multiple institutions and time periods, employing robust statistical analysis methods, and considering the impact of healthcare utilization patterns on EHR data, we address several of the challenges inherent to EHR-based research. Our findings contribute to a growing body of evidence that underscores the complexity of endometriosis and highlights the potential for EHR data to advance our understanding of this condition on a population level. In summary, our study provides a comprehensive analysis of endometriosis comorbidities using large-scale EHR data from multiple medical centers. By comparing endometriosis patients to matched controls, we captured both the broad range of associated comorbidities and the specific patterns preceding diagnosis. Importantly, our clustering analysis revealed meaningful subgroups within the endometriosis patient population, highlighting the heterogeneity of the condition and suggesting potential pathways for personalized management strategies. The replication of key findings across two independent datasets reinforces the robustness and generalizability of our results. This consistency underscores the value of leveraging EHR data to study complex, heterogeneous conditions like endometriosis at scale. However, our findings also emphasize the challenges inherent in analyzing real-world data, such as biases introduced by healthcare utilization patterns and limitations in diagnostic coding. Looking ahead, our work lays a foundation for future studies to explore further relationships between endometriosis and its comorbidities, as well as the biological mechanisms underlying the observed heterogeneity. Integrating genomic, clinical, and patient-reported data with EHR-based findings may further enhance our understanding and ultimately support the development of targeted diagnostic tools and treatment strategies, including with novel machine learning approaches. By advancing knowledge of endometriosis and its comorbidities, this research contributes to ongoing efforts to improve patient care, reduce diagnostic delays, and address the significant burden of this disease.

Declaration

During the preparation of this work, the authors used ChatGPT-4o (released May 13, 2024, by OpenAI) in order to turn section outlines into complete initial drafts and for refining language and improving readability. After using this tool, the authors reviewed and edited the content as needed, and we take full responsibility for the content of the publication.

Introduction

Endometriosis is a chronic and often debilitating inflammatory and systemic condition that affects millions of individuals worldwide. 1 It is characterized by the presence of endometrial-like tissue outside the uterus, leading to inflammation, scarring, and adhesions. 2 Endometriosis is estimated to affect approximately 10% of reproductive-aged individuals with a uterus, making it a significant public health concern. The condition is associated with a wide range of symptoms, including chronic pelvic pain, infertility, dysmenorrhea, and gastrointestinal disorders, all of which contribute to a substantial burden on patients’ quality of life. 3 Despite its prevalence, diagnosing and managing endometriosis remain challenging due in part to wide variability in manifestation. Patients often take years to reach a diagnosis, during which symptoms are frequently misattributed to other conditions and the disease can progress. 4 Even though many clinicians use clinical judgment to diagnose and initiate treatment, the diagnostic delay is often further influenced by the reliance on surgery for definitive diagnosis. Furthermore, treatment options are complex, and response rates are variable. While hormonal therapies and surgical interventions can provide symptom relief, they are associated with side effects and a high likelihood of symptom recurrence. 5 Beyond these clinical hurdles, endometriosis imposes a substantial psychosocial burden on patients, who may face stigma, invalidation of their pain, and diminished quality of life. 6 Although many smaller studies have provided valuable insights into endometriosis’s heterogeneity, 7 , 8 , 9 , 10 they often cannot capture broad population-level features that characterize the disease. Electronic health records (EHRs) offer an opportunity to study large patient populations and uncover patterns that may not be apparent in smaller-scale studies. 11 By leveraging these data, researchers can identify trends and associations that might otherwise go unnoticed. The application of EHRs to study endometriosis is particularly promising, as it allows for a more comprehensive examination of the disease and its comorbidities in diverse and extensive patient populations. 12 Prior studies investigating endometriosis using EHRs have highlighted their potential for providing insights into the condition. For instance, Choi et al. analyzed patient information from South Korea’s Health Insurance Review and Assessment data and found 44 International Classification of Diseases (10th revision) (ICD-10) codes that were significantly associated with endometriosis. 13 Also, Estes et al. examined administrative health claims data from Optum’s Clinformatics DataMart to compare the incidence of mental health outcomes in patients in the United States with and without endometriosis, finding an increased risk for anxiety, depression, and self-directed violence. 14 In a similar vein, Melgaard et al. explored pre-diagnosis healthcare utilization in a Danish set of patients with and without endometriosis, reporting a higher number of hospital contacts among women with endometriosis. 15 However, these studies focus on specific aspects of the disease, such as particular classes of comorbidities or certain patient subpopulations, and do not attempt to validate their findings across independent data sources. In the clustering domain, Sarria-Santamera et al. performed hierarchical subtyping of endometriosis patients in the Spanish National Health System, revealing clusters with higher comorbidity levels and particular diagnoses like anemia and infertility but limited to a narrow time window (2013–2017) and without an independent data source for validation. 16 Along similar lines, Urteaga et al. used patient-reported data from a smartphone app to identify subtypes of endometriosis. 17 Though not built around structured EHR data, this work demonstrates the utility of data-driven methods in characterizing this complex disease. In this work, we aim to address these gaps by analyzing the comorbidities of endometriosis patients across multiple medical centers. Specifically, we examine data from the University of California, San Francisco (UCSF), and five other University of California (UC) medical centers, utilizing odds ratio analysis and unsupervised clustering techniques to identify and characterize patterns of diagnoses associated with endometriosis ( Figure 1 A). Our analysis compares endometriosis patients with matched controls, considering both pre-endometriosis conditions and diagnoses across the full patient timeline. Clustering is then performed on the endometriosis patients to identify subgroups that may reflect different disease trajectories or phenotypes ( Figure 1 B). This dual focus enables us to capture the overall landscape of endometriosis comorbidities compared to controls, while also examining the heterogeneity within the endometriosis patient population. By uncovering these patterns, we aim to contribute to a more comprehensive understanding of endometriosis, its clinical features, and its impact on patient health. Figure 1 Graphical overview of the study and analysis (A) Overview of the study design. Endometriosis patients were selected from electronic medical records by diagnoses. We then (1) compute odds ratios for comorbidities by comparing against matched controls and (2) identify groupings of endometriosis patients using unsupervised clustering. (B) Overview of the clustering process. From the initial matrix of diagnoses, a nearest-neighbors graph is constructed based on patient-to-patient distances. The Leiden algorithm is then applied to this graph. Each cluster from UCSF is matched to the UC-wide cluster with the greatest proportion of overlapping significantly enriched conditions. (C) Overview of patient selection. We select cases using the “endometriosis” concepts from the Observational Medical Outcomes Partnership (OMOP) concept hierarchy and match against controls from the non-endometriosis population using propensity score matching on both demographic and healthcare utilization covariates. See also Figure S1 . Graphical overview of the study and analysis (A) Overview of the study design. Endometriosis patients were selected from electronic medical records by diagnoses. We then (1) compute odds ratios for comorbidities by comparing against matched controls and (2) identify groupings of endometriosis patients using unsupervised clustering. (B) Overview of the clustering process. From the initial matrix of diagnoses, a nearest-neighbors graph is constructed based on patient-to-patient distances. The Leiden algorithm is then applied to this graph. Each cluster from UCSF is matched to the UC-wide cluster with the greatest proportion of overlapping significantly enriched conditions. (C) Overview of patient selection. We select cases using the “endometriosis” concepts from the Observational Medical Outcomes Partnership (OMOP) concept hierarchy and match against controls from the non-endometriosis population using propensity score matching on both demographic and healthcare utilization covariates. See also Figure S1 .

Coi Statement

L.C.G. is a consultant to Myovant Sciences, Gesynta Pharma, Celmatix, NextGen Jane, and Chugai Pharmaceutical Co.

Star★Methods

REAGENT or RESOURCE SOURCE IDENTIFIER Deposited data original code this paper https://github.com/khanu263/comorbidities-clustering-endo Software and algorithms Python (3.11) Python Software Foundation https://www.python.org/ R (4.1.3) R Development Core Team https://www.r-project.org/ MatchIt (4.5.5) Comprehensive R Archive Network https://cran.r-project.org/web/packages/MatchIt/index.html scanpy (1.11.1) Python Package Index https://pypi.org/project/scanpy/ scipy (1.14.1) Python Package Index https://pypi.org/project/scipy/ matplotlib (3.9.1) Python Package Index https://pypi.org/project/matplotlib/ Cytoscape Cytoscape Consortium https://cytoscape.org/ original code this paper https://github.com/khanu263/comorbidities-clustering-endo Patients were selected from two separate de-identified EHR databases, both of which conform to the Observational Medical Outcomes Partnership 37 (OMOP) Common Data Model 38 (CDM) schema. We initially performed our analysis using UCSF’s records, which run from 1988 and comprise over 5.5 million patients. We then expanded our scope to the University of California’s Health Data Warehouse (UCHDW), which aggregates medical records from over 8 million patients seen at six University of California medical centers, spanning 2012 to the present day. We explicitly excluded UCSF patients from our UCHDW queries to avoid patient overlap ( Figure 1 C). We defined endometriosis cases as patients who at any point in their medical history were assigned at least one of the 49 standard SNOMED condition IDs descended from “endometriosis” in OMOP’s concept hierarchy ( Table S9 ). Conversely, we identified a background population by selecting patients who have at least one condition and were never diagnosed with endometriosis. We selected controls via propensity score matching against cases on age, gender, race, and ethnicity (and location for UCHDW patients) using the MatchIt 39 package in R (version 4.5.5, nearest neighbor method) ( Figure S1 ). For each endometriosis case, we select thirty control patients, allowing us to repeat the association analysis and build a distribution of odds ratios for each condition, as described in the next section. To consider the effect of healthcare utilization on our observed associations, we also selected utilization-matched controls by adding the number of recorded visits and record duration as additional covariates for propensity score matching. These additional control patients were used to filter significant associations as explained in the next section. Our initial analysis considers all diagnoses across any given patient’s record. To examine the pattern of comorbidities observed prior to endometriosis, we repeat the following methods using only patients with conditions assigned prior to their first endometriosis diagnosis, and their matched controls. We restrict the condition set to those that appear in the record before the first endometriosis condition, or conditions assigned at any time to the controls. All analysis of UCSF and UC-wide EHR data was performed under the approval of the Institutional Review Boards. All clinical data were de-identified and written informed consent was waived by the institutions. Patients were selected using the UCSF de-identified and UC-wide HIPAA-compliant limited dataset OMOP EHR databases. We iterated over the control patients in thirty separate groups (corresponding with the propensity score matching results from the previous section), and for each of the SNOMED conditions ever assigned to a patient, we calculated the odds ratio between cases and selected controls and computed a p -value by the hypergeometric test. We applied a Bonferroni correction to the p -value by multiplying it with the number of conditions being tested in a given iteration. This resulted in a distribution of thirty odds ratios and corresponding Bonferroni-corrected p -values for each condition. To aggregate across these iterations, ensuring our associations are robust to variation in the control patients, we compute the mean odds ratio and an aggregated p -value (defined as twice the mean corrected p -value 40 ) for each condition. We then accounted for healthcare utilization by repeating the previous procedure using the utilization-matched control patients from the previous section. We considered a condition to be significant in our final analysis if it was first identified as significant in the non-utilization-matched comparison and remained significant after adjusting for healthcare utilization. This approach eliminates associations that appear significant solely due to increased interactions with the healthcare system stemming from endometriosis and its symptoms, while minimizing falsely protective associations that arise from comparing to control patients with higher overall healthcare needs. To identify clusters of endometriosis patients, we first construct a matrix where each row represents an endometriosis patient, and each column represents a condition. An entry is set to 1 if the corresponding patient has the given condition, and 0 otherwise. To identify clusters, we applied principal component analysis to this matrix to reduce each patient’s representation down to 1,000 dimensions, explaining ∼80% of the variance. We then constructed a neighborhood graph and applied the Leiden community detection algorithm 41 using scanpy. 42 We used a fixed Leiden resolution parameter of 1.0 across all clustering analyses. This resolution value was selected based on standard practice in exploratory clustering and provided interpretable granularity. We tested a small range of resolution values (0.5–1.0) and found that while the exact number of clusters varied, the key subgroupings remained consistent. The number of clusters reported in each setting (UCSF vs. UC-wide, all vs. pre-endometriosis diagnoses) reflects the output of the community detection process, rather than a predefined input. To examine each of the resulting clusters, we perform the odds ratio analysis procedure described previously, with the “cases” being endometriosis patients in the cluster under consideration and the “controls” being patients in all other clusters. We then identify significantly enriched conditions that are exclusive to each cluster. We compared both the association and clustering analyses across the two data sources. For the association analysis, we first found the intersection of significantly enriched conditions across the two data sources and then computed the Pearson correlation coefficient and associated p -value between the log-odds ratios of those conditions. For the clustering analysis, we matched each cluster found at UCSF to the UCHDW cluster with the most overlapping significant conditions ( Figure 1 B), if there was one, and visualized the results as a bipartite graph. Note that only clusters that have overlapping conditions between the two data sources are visualized in the graphs. To further quantify similarity between matched clusters across datasets, we computed the Jaccard similarity index between the sets of significantly enriched conditions in each cluster pair. These values are reported in Table S10 . We applied the Uniform Manifold Approximation and Projection (UMAP) algorithm 43 to the diagnosis matrix like the one used for the clustering analysis but including all endometriosis patients and their closest matched controls, to visualize the comorbidity structure. We used Cytoscape 44 to construct the bipartite graph visualizations and associated legends for the clustering comparison between the two data sources. All other illustrations were created with BioRender, and all other data visualizations were generated using matplotlib. 45 To annotate SNOMED concepts with ICD chapter labels for visualization, we used the SNOMED CT to ICD-10-CM Map provided in the Unified Medical Language System created by the National Library of Medicine. 46 Matching to construct study groups was performed using the matchit function of the MatchIt package 39 (version 4.5.5) in R. We used the nearest neighbor method, and matched on a combination of covariates including age, gender, race, ethnicity, location (for UCHDW patients), number of recorded visits, and record duration, as described in the method details. All statistical tests were performed using the hypergeom function from the scipy package (version 1.14.1) in Python. To account for multiple hypothesis testing, we applied a Bonferroni correction to the p -values from our hypergeometric tests by multiplying them with the number of conditions being tested in a given control group iteration. To aggregate p -values across the thirty control groups, we computed the mean odds ratio and an aggregated p -value (defined as twice the mean corrected p -value 40 ) for each condition.

Acknowledgments

The authors would like to thank Parker Grosjean, Duncan Muir, Chimno Nnadi, Michael Keiser, Stacey Missmer, Christina Theodoris, Tony Capra, and all members of the Sirota Lab for insightful conversations and guidance throughout the preparation of this work. The authors acknowledge the use of the UCSF Information Commons computational research platform, developed and supported by the UCSF Bakar Computational Health Sciences Institute in collaboration with the IT Academic Research Services, the Center for Intelligent Imaging Computational Core, and the CTSI Research Technology Program. This manuscript was supported by the 10.13039/100009633 Eunice Kennedy Shriver National Institute of Child Health and Human Development , P01HD106414 (U.K., T.T.O., J.C.I., J.O.-A., L.C.G., and M.S.), and the 10.13039/100000057 National Institute of General Medical Sciences , T32GM067547 (U.K.) and T32GM142516 (K.G.). This project was supported by the 10.13039/100006108 National Center for Advancing Translational Sciences , National Institutes of Health, through UCSF-CTSI grant number UL1TR001872 .

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Condition tags

endometriosis

MeSH descriptors

Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis Endometriosis

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-09-13T09:25:22.628771+00:00
pmc
last seen: 2026-05-13T20:22:03.195721+00:00
pubmed
last seen: 2026-09-13T06:07:03.576298+00:00
unpaywall
last seen: 2026-05-11T08:34:28.763810+00:00
License: CC-BY-4.0 · commercial use OK · attribution required
Courtesy of the U.S. National Library of Medicine