The Problem of Pain in Rheumatology: Variations in Case Definitions Derived From Chronic Pain Phenotyping Algorithms Using Electronic Health Records.

OA: closed
AI-generated deep summary by qwen3.7-flash, 2026-08-22 · read from full text

This study developed and validated computational phenotyping algorithms to estimate the prevalence of chronic pain in patients with autoimmune rheumatic diseases using electronic health records from Stanford Health Care. By analyzing data from over 3,000 patients with conditions such as systemic lupus erythematosus and Sjögren’s syndrome, the researchers compared four unimodal criteria—pain scores, diagnostic codes, analgesic prescriptions, and pain interventions—to create more precise case definitions than previous methods. The findings revealed significant variability in pain prevalence across different rheumatic diseases and highlighted that females consistently exhibited higher rates of chronic pain indicators across all algorithmic measures. Relevance to endometriosis: endometriosis is listed among the ICD-10 codes used to define chronic pain diagnoses within the study's phenotyping algorithm, though the paper's primary focus remains on autoimmune rheumatic diseases.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

ObjectiveThe aim of this study was to investigate and compare different case definitions for chronic pain to provide estimates of possible misclassification when researchers are limited by available electronic health record and administrative claims data, allowing for greater precision in case definitions.MethodsWe compared the prevalence of different case definitions for chronic pain (N = 3042) in patients with autoimmune rheumatic diseases. We estimated the prevalence of chronic pain based on 15 unique combinations of pain scores, diagnostic codes, analgesic medications, and pain interventions.ResultsChronic pain prevalence was lowest in unimodal pain phenotyping algorithms: 15% using analgesic medications, 18% using pain scores, 21% using pain diagnostic codes, and 22% using pain interventions. In comparison, the prevalence using a well-validated phenotyping algorithm was 37%. The prevalence of chronic pain also increased with the increasing number (bimodal to quadrimodal) of phenotyping algorithms that comprised the multimodal phenotyping algorithms. The highest estimated chronic pain prevalence (47%) was the multimodal phenotyping algorithm that combined pain scores, diagnostic codes, analgesic medications, and pain interventions. However, this quadrimodal phenotyping algorithm yielded a 10% overestimation of chronic pain compared to the well-validated algorithm.ConclusionThis is the first empirical study to our knowledge that shows that established common modes of phenotyping chronic pain can lead to substantially varying estimates of the number of patients with chronic pain. These findings can be a reference for biases in case definitions for chronic pain and could be used to estimate the extent of possible misclassifications or corrections in using datasets that cannot include specific data elements.
Full text 25,181 characters · extracted from pmc-nxml · 3 sections · click to expand

Results

The number of patients attending rheumatology clinics in 2019 ranged from 242 diagnosed with ankylosing spondylitis to 842 diagnosed with Sjogren’s syndrome ( Table 1 ). There was a preponderance of female patients (55%–93%) in all conditions, except for ankylosing spondylitis. The proportion of non-White patients ranged from 50% in systemic lupus to 29% in psoriatic arthritis and Hispanics comprised 19% of patients with systemic lupus and 8% of patients with ankylosing spondylitis. The average age at the first visit ranged from 42 years in systemic lupus to 56 years in Sjogren’s syndrome. The proportion of publicly/privately insured did not vary considerably by disease status: overall, 93%. Approximately 18% of patients attending the rheumatology clinics met the criteria for chronic pain using pain scores based on having two or more pain scores (≥4 points out of 10) at least three months apart. Pain scores were available for >95% of patient records. The prevalence of the pain score phenotyping algorithm varied considerably by sex: 20% in females and 12% in males. Females had a higher prevalence of each unimodal pain phenotyping algorithm ( Figure 2 ). The prevalence of pain ICD phenotyping algorithm (i.e., having two or more pain ICD codes at least three months apart) was higher in females (23% compared to 15% in males). However, the prevalence of the pain prescriptions phenotyping algorithm (i.e., having two or more pain medications at least three months apart) was slightly higher in females (16% compared to 13% in males). Similarly, the prevalence of the pain interventions phenotyping algorithm was slightly higher in females compared to males (23% versus 19%). There was also considerable variability in the prevalence of each unimodal pain phenotyping algorithm by rheumatologic disease ( Figure 3 ). The prevalence of the pain score phenotyping algorithm was highest among patients with systemic lupus erythematosus (22%) and Sjogren’s syndrome (19%). The prevalence of the ICD code pain phenotyping algorithm was highest among patients with Sjogren’s syndrome (~24%). The prevalence of the pain prescriptions phenotyping algorithm was similarly highest among patients with Sjogren’s syndrome and lowest among patients with systemic sclerosis. However, 17% of patients with ankylosing spondylitis had pain interventions compared with 26% of patients with systemic sclerosis. The correlation coefficients between the unimodal pain phenotyping algorithms were generally low to medium in strength (Supplementary Table 4). The lowest correlation was between pain ICD codes and pain interventions (~0.23), while the highest correlation was between pain scores and pain ICD codes (~0.46). There were also differences in the correlations by sex, with males tending to have lower correlation coefficients. However, male patients had the highest correlation between pain scores and analgesic prescriptions (0.50) and almost no correlation between analgesic prescriptions and pain interventions. Table 2 shows the prevalence of chronic pain based on bimodal (combination of two phenotyping algorithms), trimodal (combination of three phenotyping algorithms), and quadrimodal phenotyping algorithms (combination of four phenotyping algorithms). Among the eleven possible combinations of the unimodal pain phenotyping algorithms, the highest estimated prevalence of chronic pain was the combination of all four unimodal algorithms into a multimodal phenotyping algorithm, #11 (47%). The lowest estimated prevalence of chronic pain was the combination comprised of two unimodal phenotyping algorithms (pain scores and pain prescriptions), #2 (28%). The Tian et al chronic pain algorithm combined pain scores, ICD codes and prescription and the prevalence of chronic pain based on this algorithm (#7) was 37%. Using the Tian et al algorithm, the prevalence of chronic pain was slightly higher among females (39% compared to 29%). The prevalence of chronic pain was also highest among black patients (~42%) compared with Asian patients (27%). The prevalence of chronic pain was also highest among the publicly insured across multiple combinations. The prevalence of chronic pain using these multimodal definitions was again highest in patients with Sjogren’s syndrome and lowest in systemic sclerosis. When estimating the extent of the over-estimation or under-estimation of chronic pain by the multimodal pain phenotyping algorithms compared with the Tian et al algorithm, we found substantial variation overall and within key population subgroups ( Figure 4 ). The combination that was closest in prevalence to the Tian et al algorithm was ICD codes and pain interventions (overall difference was −0.4%), with the differences in estimates of chronic pain varying from −11% to +5% in various subgroups. The combination that was furthest in prevalence to the Tian et al algorithm was pain scores, ICD codes, prescriptions, and interventions (overall difference was +10%), with over-estimation as high as 13% in subgroups. We quantified the strength of association between sex, race and disease status, and chronic pain using different case definitions according to the multimodal pain phenotyping algorithms (Supplementary Figures 1 and 2). When estimating odds ratios comparing females to males, we found that female patients had >50% higher odds of chronic pain using the Tian et al algorithm (highlighted in red). 7 While this point estimate varied by the multimodal pain phenotyping algorithm, the confidence intervals overlapped for the eleven multimodal phenotyping algorithms. Similarly, across multiple case definitions, Asian patients had 50% lower odds of chronic pain compared with white patients based on the Tian et al pain algorithm; this point estimate did not vary considerably using other case definitions. However, there were more variations in point estimates by disease status, although most confidence intervals overlapped considerably.

Discussion

This study used records from ~3,000 patients seen at rheumatology clinics to create unimodal and multimodal chronic pain phenotyping algorithms. The aim was to investigate and compare different case definitions for chronic pain for use to provide estimates of possible misclassification when researchers are limited by available EHR data. We did not determine the accuracy of each pain phenotyping algorithm – this is the subject of future research. As expected, the prevalences of chronic pain using unimodal phenotyping algorithms were lower than those using multimodal pain phenotyping algorithms. The prevalences of chronic pain also increased with increasing number (bimodal to quadrimodal) of phenotyping algorithms that comprised the multimodal phenotyping algorithms. We applied the 2012 Tian et al multimodal phenotyping algorithm that combines pain scores, ICD codes and prescription to our current data to yield a chronic pain prevalence estimate of 37%, which exceeded the 19% chronic pain prevalence estimate that the algorithm authors reported in their 2012 study. 7 Moreover, we improved on the Tian et al general chronic pain phenotyping algorithm by better aligning with rheumatology clinics and patients in the following ways. First, the Tian study setting was in a multisite community center whereas we used academic rheumatology clinics in Northern California and a patient population with significant burden of chronic pain. Second, the Tian et al phenotyping algorithm restricted their examination of analgesic medications to opioids only. In contrast, we added other analgesic medications that better align with contemporary changes in the management of chronic pain, and are commonly prescribed for inflammatory pain conditions. When we limited our case definition to opioids only the prevalence of chronic pain was ~31%. The Tian et al multimodal phenotyping algorithm showed good to excellent accuracy during derivation and validation. The positive predictive value of the Tian et al multimodal phenotyping algorithm was 98% in the derivation dataset and 91% upon validation, the sensitivity was 85% and the specificity was 98%. 7 These accuracy metrics were superior to those of unimodal phenotyping algorithms (pain scores, opioids, and ICD codes). The Tian et al multimodal phenotyping algorithm (or variations of it) has been adapted for myriad purposes including: characterizing the demographics of chronic pain patients in the state of Maine, 8 estimating the costs and gender differences in the use of complementary and integrative health intervention for chronic pain in veterans, 13 , 14 predicting pain intensity improvements, 15 and evaluating the burden of neuropathic pain in Canada. 16 The choice of which unimodal or multimodal pain phenotyping algorithm to use depends on the availability of the data elements in EHR, claims or registry data and the research question or hypothesis. For example, most administrative claims databases such as Truven Health (IBM) MarketScan, Optum Insight, Medicare, and Medicaid may only have ICD codes, medications, and interventions, but may not have pain scores. However, a trimodal phenotyping algorithm with that combination would yield a chronic pain estimate that is 5% higher than the one estimated by the Tian et al phenotyping algorithm. An alternative would be the bimodal phenotyping algorithm that combines ICD codes and interventions where the chronic pain estimate was only 0.4% less than the Tian et al phenotyping algorithm. One caveat to bear in mind: one cannot study the effects of elements that comprise a particular phenotyping algorithm. If a chronic pain phenotyping algorithm includes prescriptions, researchers cannot study the effect or trends of prescriptions using that phenotyping algorithm. One study that adapted the Tian et al algorithm for estimating the cost of complementary and integrative health approaches for chronic musculoskeletal pain in younger US Veterans did not include the use of medications in the case definition for pain because they were evaluating the use of pain interventions. 13 Researchers in that study elected to use ICD codes and pain scores, a phenotyping algorithm that yielded a 6% underestimate of chronic pain prevalence in the present study. 13 This phenotyping algorithm has been shown to have a predictive value of 82%. 7 In another study, improvements in pain intensity were evaluated and the chronic pain definition used only ICD codes. 15 In this study, the unimodal phenotyping algorithm consisting of ICD codes yielded a chronic pain prevalence of 21%, a potential 18% underestimate. This phenotyping algorithm has been shown to have a predictive value of 89–95%, but the sensitivity ranged between 20–71%. 7 As expected, the prevalence of chronic pain was higher among female patients in each case definition. Female patients had a higher prevalence of high pain scores and pain ICD codes. A female preponderance in the burden and intensity of chronic pain is well documented in the literature. 17 , 18 For example, several studies have reported higher prevalence of temporomandibular disorders, neuropathic pain in diabetes, and postoperative pain in female patients. 19 – 22 In a study of 72,000 patients using the EHR from Stanford Hospital and Clinics, Ruau et al found that, on average, women reported higher pain scores in 72% of diagnoses, including rheumatoid arthritis, osteoarthritis and diabetes. 23 Other studies have also shown the higher use of pain medications in females compared to males, including opioid use and complementary/integrative health. 24 – 26 To our knowledge, this study is the first to highlight sex differences across four chronic pain phenotyping algorithms. Consistent with previous research, we found higher prescriptions of pain medications and intervention in females than males, but these differences were not significant. This study has a few limitations. We did not include patients with rheumatoid arthritis (RA) in this study because patients with this condition are selectively given the RAPID3 questionnaire at every visit in our clinics, which includes the visual analog pain scale (VAS). Thus, the inclusion of patients with RA have an artificially higher prevalence of pain relative to the other groups because of the systematic collection of VAS pain in RA (Supplementary Figure 3). We included only medications that are prescribed nationally to ensure generalizability, thus excluding pain treatment modalities such as cannabis. We were unable to capture social determinants of health such as educational attainment, income, and employment status – factors known to impact chronic pain and disability. We also did not include comorbidities such as diabetes, heart disease, and mood disorders. We did not include information on healthcare utilization metrics such as emergency department visits and surgeries. There are numerous other lifestyle factors that could potentially influence pain levels and the accuracy of pain classifications, including diet, physical activity, alcohol consumption, and sleep quality. However, we were only able to extract smoking status and BMI from the EHR system. We also did not include information on the prescribers of the pain medications and did not include status of specific conditions such as duration of disease, severity, and rheumatologic medications such as disease-modifying antirheumatic drugs. We found that effect estimates of associations between key demographic variables and chronic pain based on the different case definitions did not vary considerably. It is impossible to tell which directions the effect sizes will go during multivariable regression modeling. In addition, these findings may not generalize beyond rheumatology clinics and these clinics’ geographical/sociodemographic contexts. However, chronic pain remains the most salient source of disability in rheumatology and the sociodemographic composition of the study participants is racially diverse (~40% non-white). We expect this phenomenon is true in other contexts, however, future studies should include validation of these findings in other specialties. We speculate that error would be on the side of underestimating. In addition, misclassification bias in case ascertainment may have profound implications for estimation of disease burden, thus, impacting risk estimations, patient care, cost estimations, and health policy decisions. However, perfect case definitions do not exist and there needs to be careful consideration of the costs of over- or under-estimation of chronic pain cases, e.g., over-screening and missed cases. We did not estimate the accuracy of the case definitions in ascertaining chronic pain. This is the subject of a future study. In addition, the current study gives a general overview of the prevalence of chronic pain in patients with ARD, however, it does not address the nuances of the pain experience in individual patients, which are influenced by disease status, but also, contextual determinants. This study is foundational as an early exploration of the epidemiological use of EHR, with need for future refinement. In conclusion, this was the first empirical study to show that established common modes of phenotyping chronic pain can lead to substantially varying estimates of number of patients with chronic pain. As a consequence, we created a reference for biases in case definitions for chronic pain in EHR. These findings could be used to estimate the extent of possible misclassifications or corrections in using datasets that cannot include specific data elements.

Introduction

Up to 7% of adults in the United States live with autoimmune rheumatic diseases (ARDs), including systemic lupus erythematosus, Sjögren’s syndrome, systemic sclerosis, ankylosing spondylitis, and psoriatic arthritis. 1 Autoimmune rheumatic diseases are heterogeneous conditions disproportionately affecting women and are hypothesized to be caused by a dysregulated immune system triggering organ dysfunction and damage. 1 The American College of Rheumatology noted pain as the most salient patient-reported outcome in autoimmune rheumatic diseases because it significantly impacts health-related quality of life and disability. 2 Despite treatment advances in rheumatology, including the advent of highly specific biologics and disease-modifying antirheumatic drugs, pain remains an undertreated problem. 3 Depending on the ARD diagnosis, upwards of 65% of patients with ARDs are on long-term opioid therapies and have higher rates of opioid overdose hospitalizations than the general population. 4 We need to understand the burden of chronic pain in autoimmune rheumatic diseases. People with rheumatologic conditions experience moderate to severe pain associated with the condition itself, are likely to have medical comorbidities, and are likely to experience co-occurring chronic pain from non-rheumatologic conditions (e.g., migraine or irritable bowel disease). Specifically, reliable and detailed estimates of the full burden of chronic pain in ARDs are needed to establish a basis for population-wide interventions and estimate trends. Electronic health records (EHRs) offer great potential to reduce the need for data gathering and thereby promote efficient calculation of estimates of disease burden. While EHRs are a rich data resource, accurate case definitions and algorithms are needed to ensure precision in estimating the burden of chronic pain. Accordingly, computational phenotyping is a clinical data science method that leverages data-driven methods to subtype and characterize patient conditions from heterogeneous EHR data, 5 and yield phenotyping algorithms that could be applied across datasets. Indeed, the potential uses of the chronic pain phenotyping algorithm include: estimating incidence, prevalence, and trends in administrative databases/EHR; providing baseline measures to characterize pain trajectories in administrative databases; understanding the epidemiology of other chronic overlapping pain conditions; facilitating cost modeling and health utilization projections; and characterizing treatment response and effectiveness. Variability in case definitions can create error in estimations and inferences 6 by affecting the number and type of cases identified. As such, accuracy and consistency in case definitions are crucial as downstream applications of results potentially impact direct patient care, research, and health policy. Chronic pain is often considered a subjective symptom and can be clinically complex. The lack of objective tests, biomarkers and diagnostic criteria may complicate the use of EHR data for epidemiological inference. In 2012, Tian et al published a well-cited general chronic pain algorithm that applied three unimodal criteria from structured EHR data: pain scores, diagnosis codes, and prescription opioids. 7 , 8 However, several limitations call into question the utility of this algorithm in rheumatology, including that: (1) trends in the diagnosis and management of chronic pain have evolved since the 2011–2012 study period; (2) it was developed using data from primary care thereby limiting its generalizability in the rheumatic specialty care context; (3) it was not externally validated outside of the derivation dataset; and (4) it included opioid medications only and omitted other prescription analgesics. To create greater precision in case definitions for chronic pain in EHR, the aims of this study are to: (1) describe the extraction and prevalence of four unimodal pain criteria in EHR using pain scores, diagnostic codes associated with chronic pain, analgesic medications, and pain interventions; (2) detail the correlation between the unimodal pain criteria; (3) determine the prevalence of distinct combinations of the pain criteria into algorithmic phenotyping and compare them to the Tian et al algorithms; and (4) quantify the strength of association between several variables (as exposures) and chronic pain (as outcomes) using different case definitions based on multimodal phenotyping algorithms. We also examined sex differences for each of the four study aims. We expect our comparison of several different chronic pain phenotyping algorithms will yield useful information for future researchers depending on their data availability and research questions. The next step in this line of research will be to conduct validation of these algorithms to estimate their accuracy. The Institutional Review Board at the Stanford School of Medicine approved this study (IRB# 53750). Stanford Health Care has ~1 million outpatient visits per year and uses EPIC as its EHR vendor. EHR data fields available for analysis include demographics, diagnostic codes, problem list, medications, laboratory studies, procedures, and clinical encounter notes. EHR records are accessed via a backend relational database (STAnford Research Repository, or STARR) that can be queried. We retrospectively queried STARR to identify patients ages ≥18 years visiting any Stanford outpatient rheumatology clinic (including joint immunology-dermatology clinics) in 2019. Further, we identified patients with ≥2 visits with ARD diagnoses ≥3 months apart using the International Classification of Diseases, Clinical Modification (ICD-10-CM) codes (Supplementary Table 1) for the following ARDs: ankylosing spondylitis, psoriatic arthritis, Sjogren’s syndrome, systemic lupus erythematosus, and systemic sclerosis. We chose to include at least 2 ARD diagnoses of the same ARD to ensure that these patients were being actively managed for each condition. Current pain intensity scores are recorded at clinic visits using the 11-point numeric pain rating scale (i.e., 0 = no pain; 10 = worst pain imaginable). The pain score is recorded as structured data in the EHR. Patients who had two or more pain scores rated as ≥4 point at least three months apart ( Figure 1 ) were categorized as having met the unimodal chronic pain criterion. 7 We defined a list of ICD-10 codes for general and non-rheumatic chronic pain conditions through an extensive literature review and expert input (Supplementary Table 1, Figure 1 ). We extracted ICD-10 codes for the following painful conditions: abdominal pain, chest pain, pain in joints, pain in limbs, cervicalgia, fibromyalgia, irritable bowel syndrome, urologic chronic pelvic pain syndrome, vulvodynia, migraine, chronic tension-type headache, temporomandibular disorder, chronic low back pain, myalgic encephalomyelitis /chronic fatigue syndrome, and endometriosis. We labelled patients with chronic pain using the diagnosis codes phenotyping algorithm if they had ≥2 ICD codes in Supplementary Table 1 recorded at visits separated by ≥3 months. We extracted pain-related prescriptions: opioids, non-steroidal anti-inflammatory drugs (NSAIDs), acetaminophen, anticonvulsants, antidepressants, and muscle relaxants using previously described definitions. 9 We categorized patients as having chronic pain if they met the analgesic prescription criterion of 2+ prescriptions of any of the pain medications listed in this section recorded at visits separated by ≥3 months (Supplementary Table 2, Figure 1 ). We identified patients with the most common Current Procedural Terminology (CPT ® ) codes used in interventional pain management based on conversations with rheumatologists and pain clinicians ( Figure 1 , Supplementary Table 3). These codes include, but are not limited to physical therapy, acupuncture, and nerve blocks. We labelled patients as having chronic pain using the pain interventions phenotyping algorithm if they had any of the procedures in Supplementary Table 3. We examined the baseline distribution of patients with each of the five rheumatic conditions by sociodemographic (sex, race, ethnicity, age at baseline, marital status, and insurance status) and lifestyle (smoking status and body mass index) variables. We estimated the prevalence of each of the four unimodal phenotyping algorithms in patients with rheumatic diseases. We also determined the sociodemographic and lifestyle variables associated with the presence of each unimodal phenotyping algorithm. We reported the associations between the unimodal phenotyping algorithms using tetrachoric correlation coefficients, a measure of association between the binary variables. Tetrachoric correlation coefficients can range from −1 (indicating strong negative correlation between two unimodal phenotypes) to 1 (indicating a strong positive correlation between two unimodal phenotypes) 10 – 12 . A value of 0 indicates no correlation between two unimodal phenotypes. We also evaluated the prevalence of eleven combinations of the unimodal phenotyping algorithms and examined the extent of the over- or under-estimation of chronic pain prevalence compared to the Tian et al chronic pain algorithm 7 , 8 . Finally, we quantified the strength of association between sex, race and disease status with chronic pain using different case definitions based on the multimodal pain phenotyping algorithms by estimating odds ratios.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-23T09:30:01.253652+00:00
unpaywall
last seen: 2026-08-23T06:29:45.520198+00:00