Methods
As shown in Figure 1 , we used quantitative methods for psychometric testing. Both on-menses and off-menses versions were evaluated using anonymous surveys. Participants who were menstruating completed the on-menses version, while those not on menses completed the off-menses version. Those who were on days 1 to 3 of their cycle were invited to participate in a follow-up survey 24 hours after the initial survey. This follow-up survey allowed us to (1) test the DSI’s test-retest reliability, (2) evaluate the DSI’s responsiveness to detect change, and (3) estimate its minimally important difference (MID).
Similar to phase 1, eligibility criteria included being: (1) female, (2) aged 14–42, (3) able to read and converse in English, (4) currently living in the United States, and (5) menstrual pain or menstrual gastrointestinal symptoms in the last 6 months. In addition, on menses participants were defined as being on days 1 to 3 of their menstrual cycle, whereas off menses participants were defined as not menstruating at the time of the study. Participants who were on day 4+ of their menses were excluded from the Phase 2 study because for the test-retest reliability, they were less likely to still experience symptoms after 24 hours. Participants from Phase 2 did not overlap with those who participated in Phase 1. Participants for Phase 2 were recruited from January to March 2019.
Different guidelines are available for judging the adequacy of sample sizes for factor analysis. Some methodologists recommend having at least 10–20 cases for each item in the scale being used (e.g., 20 × 9 items = 180; Everitt, 1975 ; Hair, 1998 ), while others suggest obtaining a sample size of at least 500 for factor analysis ( Comrey & Lee, 2013 ). We opted for the larger and more conservative minimum sample size of at least 500, which our study exceeded.
The local institutional review board approved Phase 2 of the study. Participants were recruited from the opt-in survey panel registrants maintained by a web-based service (Qualtrics, Provo, UT). Eligibility criteria for Phase 2 were the same as Phase 1. The survey panel service provider sent email invitations to 65,625 women aged 14–42 years old. Potentially interested participants clicked a hyperlink to the survey that was embedded in the email invitation (n=3754). Potential participants were further screened for eligibility (n=1654). If eligible, a study information sheet explaining consent appeared. Survey completion implied informed consent. Of those eligible (n=1032), 836 responded to the survey, which after data cleaning, resulted in a final sample size of 686.
The initial anonymous online survey included questions on participants’ demographic and health information, the DSI scale, menstrual pain severity, perceived stress, and sleep disturbance.
The only difference between the on- and off-menses versions was the recall period. The on-menses version asked participants to recall their last 24-hour experiences (i.e., the instructions read, “Over the last 24 hours, how much did your menstrual pain and menstrual gastrointestinal (GI) symptoms interfere with…”). The off-menses version asked participants to recall their experiences from the last menstrual period (i.e., the instructions read, “During your last menstrual period, how much did your menstrual pain and menstrual gastrointestinal (GI) symptoms interfere with…”). The online survey allowed for branching logic. Participants responded to different versions based on whether they were on menses or off menses. They also rated each item as essential, useful but not essential, or not necessary ( Lawshe, 1975 ) based on their own, friends’, or family members’ experiences.
Both versions had the same nine items (see Table 3 ) with response options of 1 (not at all) to 5 (very much). Individual item scores were averaged to generate a total scale score (possible range 1–5) with higher scores indicating greater interference (more negative outcome).
We used the validated numerical rating scale to assess menstrual pain severity ( Chen et al., 2015 ). For participants who were on menses, we asked what number best described their worst, least, and average pain in the last 24 hours and what number best described their current pain. Response options were from 0 (“no pain”) to 10 (“extremely severe”). For participants who were not on menses, we asked them to rate their worst, least, and average menstrual pain in their last menstrual period from 0 (“no pain”) to 10 (“extremely severe”).
We measured perceived stress using the 10-item validated perceived stress scale (PSS; Cohen et al., 1983 ). Each question asks participants about their feelings and thoughts during the last month on a 0 “never” to 4 “very often” scale. Scale scores were calculated by reverse scoring four items and summing across all items.
We used the 8-item PROMIS ® sleep disturbance scale short form (8b), which has been validated with diverse samples ( Yu et al., 2011 ). Each of the eight questions has five response options ranging from 1 to 5. Following the scoring manual, item scores were totaled to generate the raw scores. Raw scores were further converted to a T-score using a conversion table ( Health Measures, 2019 ).
Only participants who were on days 1 to 3 of their menstrual cycle (i.e., on menses) were invited to complete a follow-up survey at 24 hours. The follow-up survey included the DSI on-menses version and an additional item asking participants to rate how their symptoms changed over the last 24 hours on a 7-point scale with the response options of much worse, moderately worse, a little worse, no change, a little better, moderately better, or much better. This global rating of change has been widely used to assess responsiveness to change and MID ( Revicki et al., 2008 ). We did not invite participants who were off-menses to complete the follow-up survey because our goal was to assess test-retest reliability and responsiveness to change using current ratings rather than to assess consistency in recalled ratings.
To safeguard data quality, attention filters (i.e., “trap questions”) were also used. Responses from those who failed the attention filters were removed as were responses from “speeders” who were defined as having survey completion times less than one-third of the median survey duration. In total, we removed 150 problematic responses (from those who failed a trap question or who “sped”) before analysis. This left a sample size of 686 for psychometric analysis.
Internal consistency was evaluated using Cronbach’s alpha for both versions of the DSI scale. A threshold of 0.90 was used for internal consistency ( Nunnally & Bernstein, 1994 ). For test-retest reliability, Pearson’s correlation coefficients were calculated for participants who were on menses and whose symptoms did not change over the 24 hours. As few standards exist for judging the minimum acceptable value for a test-retest estimate, we used 0.7 as the threshold ( Crocker & Algina, 2008 ).
Content validity was evaluated by the content validity ratio (CVR) and the percentage of participants who rated the item as essential, useful but not essential, or not necessary ( Lawshe, 1975 ). Item CVRs were calculated as CVR = (ne – n/2)/(n/2), where ne indicates the number of participants indicating “essential” and n indicates the total number of participants. For a sample size of N=686, a CVR higher than 0.064 indicates excellent content validity for a given item ( Lawshe, 1975 ; Wilson et al., 2012 ). This critical value of 0.064 was calculated based on Wilson et al.’s (2012) formula, with an alpha level of 0.05 and a sample size of 686. To complement CVR, we also calculated the percentage of participants who rated a given item as essential or useful. For a given item, if 90% of participants rated it as essential or useful (i.e., <10% rated as “not necessary”), we concluded that the item had reasonable content validity.
To assess construct validity, we performed confirmatory factor analysis as opposed to exploratory factor analysis, because items were expected to measure a unidimensional construct of dysmenorrhea interference. Acceptable model fit was noted as a Comparative Fit Index (CFI) > 0.90, Goodness of Fit Index (GFI) > 0.9, Bentler-Bonett Normed Fit Index > 0.9, Root Mean Square Error of Approximation (RMSEA) < 0.10, and Standardized Root Mean Square Error (SRMR) < 0.08 ( Bartholomew et al., 2008 ; Hooper et al., 2008 ).
Concurrent validity was assessed separately for on-menses and off-menses groups using Pearson’s correlations of symptom interference with (1) menstrual pain severity, (2) stress, and (3) sleep disturbance. We expected correlations to be weak to moderate (r= 0.1–0.6) and positive (i.e., greater interference with greater menstrual pain severity, perceived stress, and sleep disturbance) as the latter are conceptually different from dysmenorrhea symptom interference ( Iacovides et al., 2015 ; Ju et al., 2014 ). We expected stronger correlations with menstrual pain severity than with perceived stress and sleep disturbance as the construct of menstrual pain severity is conceptually closer to the construct of dysmenorrhea symptom interference than perceived stress and sleep disturbance.
A MID is “the smallest difference in score in the domain of interest that patients perceived as important, either beneficial or harmful, and that would lead the clinician to consider a change in the patient’s management” (p. 377, Guyatt et al., 2002 ). We used both distribution- and anchor-based approaches to estimate MID. Distribution-based approaches are based on the statistical distribution of the measure scores, while anchor-based approaches are based on external criteria (also referred to as anchors; Revicki et al., 2008 ). For distribution-based approaches, we analyzed only those on menses women who responded to the question about symptom change. We calculated the standard deviation (SD) and the standard error of measurement (SEM). Because 0.2 SD approximates a small effect size and 0.5 SD approximates a medium effect size, a score difference between those boundaries (e.g., 0.35 SD) was used as a reasonable MID estimate ( Chen, Kroenke, et al., 2018 ; Eton et al., 2004 ). The anchor-based analysis consisted of calculating the mean DSI change from baseline to 24-hour follow-up for one category shift in the symptom change score (e.g., between “no change” and “a little worse”; Amtmann et al., 2010 ).
Responsiveness to change was estimated by calculating the standardized response mean (SRM; Revicki et al., 2008 ), which is the DSI mean at baseline minus the DSI mean at follow up divided by the standard deviation of the DSI change score. An absolute SRM value of 0.2 to 0.5 is considered a small change, 0.5 to 0.8 is moderate, and ≥ 0.8 is large ( Revicki et al., 2008 ). Some researchers suggest that an absolute SRM value ≥ 0.3 indicates responsiveness ( Askew et al., 2016 ). In addition, we compared the amount of DSI change across menstrual pain improved, unchanged, and worsened groups based on the retrospective rating of change. Omnibus analysis of variance was used to compare mean change across improved, unchanged, and worsening groups. The Tukey-Kramer adjustment was used to control the type I error for pairwise comparisons of unchanged versus worse and unchanged versus better.
Results
Phase 2 participants were diverse in race/ethnicity, educational level, and employment status (See Table 1 ). Among 686 participants, 260 (37.9%) were on menses and responded to the on-menses version of the DSI scale.
For the on-menses subset, the mean age was 28.6 years (SD=6.9). The mean menstrual pain at its worst was 6.4 (SD=2.4), menstrual pain at its least was 3.0 (SD=2.3), menstrual pain on average was 5.0 (SD=2.3), and menstrual right now was 4.4 (SD=2.8). For participants who were on menses during the initial survey, 100 (38.5%) completed the second survey. Among participants who were invited to participate in the second survey, those who completed and those who did not complete the survey were not statistically different in demographic characteristics (age, race, ethnicity) and menstrual pain level.
The remaining 62.1% (n=426) were off menses and responded to the off-menses version of the DSI. For the off-menses subset, the mean age was 27.6 years (SD=8.1). The mean menstrual pain at its worst was 6.4 (SD=2.0), menstrual pain at its least was 2.6 (SD=2.2), and menstrual pain on average was 4.9 (SD=1.9).
For the on-menses version (i.e., 24-hour recall), Cronbach’s alpha was 0.93 at Time 1 and 0.95 at Time 2, respectively. For participants whose symptoms did not change in the last 24 hours (n=32), the test-retest reliability was satisfactory (r=0.79, p<.0001).
For the off-menses version (i.e., recalling the last menstrual period), the scale was internally consistent with a Cronbach’s alpha of 0.91.
The content validity ratios were satisfactory for most items. For two items, leisure activities and social activities, the CVR was slightly below the critical value of 0.06 (See Table 3 ). The negative CVR for these two items indicated less than 50% of participants indicated the items as essential. Because only 6% of participants rated the leisure activities and social activities items as “not necessary,” we retained these two items for comprehensiveness.
Confirmatory factor analysis supported unidimensionality. As shown in Table 4 , fit indices suggested a good fit of the one-factor model for both on-menses and off-menses versions. In addition, all items for both on-menses and off-menses versions had large factor loadings (See Table 3 ).
As shown in Table 2 , the DSI scale was significantly correlated in the expected directions with pre-specified measures of menstrual pain severity, perceived stress, and sleep disturbance.
As shown in Table 2 , the distribution-based MID estimates were between 0.27 to 0.36 for both on-menses and off-menses version. The anchor-based estimate was 0.28 for minimally important improvement and 0.18 for minimally important worsening. Taken together, on a 5-point scale, the MID estimate for DSI was in the vicinity of 0.3 points.
The DSI scale on-menses version was very responsive to detect menstrual pain improvement (SRM=0.72, large effect size) but was not as responsive to detect menstrual pain worsening (SRM =−0.06, small effect size). The DSI discriminated the pain improved, unchanged, and worsened groups (p< .01 for the omnibus test). Pairwise comparisons showed the DSI successfully detected differences among pain improved and unchanged groups (p=0.046).
Discussion
We developed the DSI scale from the perspectives of adolescent girls and women aged 14 to 42. When tested in a diverse large sample in the United States, the DSI was shown to be reliable, valid, and responsive to detect menstrual pain change. The rigorous scale development process and strong psychometric properties make the tool useful for research and clinical practice.
The DSI is advantageous over other existing pain interference measures because the DSI is specific to cyclic menstrual pain. In addition, it can be used for diverse age, race/ethnicity, education level, employment status, and menstrual pain severity.
The tool captures a concept not previously considered in the field (i.e., dysmenorrhea symptom interference). Other measures of dysmenorrhea pain have not fully captured symptom interference. For example, one measure assessed “working ability” as the only dysmenorrhea impact ( Teheran et al., 2018 ), while another only assessed impact on “things the person usually does” without asking what specific aspects of life are affected ( Wyrwich et al., 2018 ). Pain interference with daily activities has been acknowledged as a core outcome in pain research, especially in clinical trials ( Dworkin et al., 2005 ). As a valid outcome measure, the DSI can be used to further develop and test interventions for dysmenorrhea.
The DSI can be administered flexibly at different phases of the menstrual cycle, given that the two DSI versions (on-menses and off-menses) with different recall periods were shown to be both reliable and valid. Compared to the off-menses version, the on-menses version is likely less subject to recall bias due to the shorter recall period. However, when daily measurement during menstruation is not feasible, recalling dysmenorrhea symptom interference with the most recent menstrual period can be appropriate. Using longitudinal designs, future research can evaluate the concordance between the DSI on-menses version (with daily recall) and the off-menses version (recalling the most recent menstrual period).
The DSI can be used for women with and without other gynecological conditions (e.g., endometriosis, uterine fibroids). Conceptually, the measure was intended to address dysmenorrhea pain regardless of clinical diagnosis. Other research also supports women with and without comorbid gynecological conditions (e.g., endometriosis, uterine fibroids) had little differences in dysmenorrhea experiences ( Nguyen et al., 2015 ). Psychometrics from this mixed sample of women were strong, suggesting the measure is appropriate for broad clinical use.
The DSI was very responsive to detect menstrual pain improvement, as shown by a large effect size estimate. However, the scale was not responsive to pain worsening. This may be because participants in this study had restricted lower bounds for worsening. For the on-menses version, we collected baseline data when participants were on days 1 to 3 of their menstrual cycle and collected follow-up data 24 hours after. Most women experience their highest menstrual pain on days 1 to 2, which may have contributed to a small magnitude of change in the pain worsening groups.
We also estimated the MID for the scale to see how much of a change in the DSI score was clinically meaningful. A change of 0.3 points in this 5-point scale indicated a clinically meaningful change had occurred. This MID estimate can help clinicians interpret scores about whether a meaningful change in dysmenorrhea has occurred, which can then guide treatment decision-making. The MID estimate also can be used to inform power calculations for clinical studies.
There are several strengths to this study. First, we developed and tested the scale using rigorous methods. Second, we developed and tested the scale using diverse samples in terms of age, race/ethnicity, education level, lifestyle, and menstrual pain severity. Third, both on-menses and off-menses versions were developed and evaluated. The availability of two versions gives clinicians and researchers a choice in regards to the timing of DSI assessment.
We acknowledge several study limitations. First, two DSI items (i.e., the leisure activities and social activities items) had low CVRs. This may be because dysmenorrhea symptoms affected individual women differently. Some women might not have seen these two items as essential. However, as only a very small percentage of participants rated the two items as “not necessary,” we retained them for comprehensiveness. These items need to be further evaluated and possibly dropped in the future. Second, we did not assess test-retest reliability of the off-menses version, as we were less interested in assessing consistency in recall based on a longer recall period. Third, for the MID anchor-based analysis, the sample size was small in a few anchor categories. Estimating MID based on fewer observations may result in unstable estimates ( Yost et al., 2011 ). Given that the sample size of the “somewhat worse” and “much worse” categories were both below 10, we only performed MID anchor-based analysis on other categories. Fourth, the samples from both phases of the study were self-selected rather than randomly selected. We acknowledge coverage bias and self-section bias. Fifth, the clinical data were self-reported. Sixth, because this was a descriptive study, the DSI scale’s responsiveness to intervention effects needs to be further evaluated in clinical trials.
In conclusion, we developed and tested the DSI scale in this two-phase study. In phase 1, we developed this 9-item scale to measure how dysmenorrhea symptoms interfere with physical, mental, and social activities based on qualitative data from cognitive interviews. The DSI has two versions (i.e., on-menses and off-menses versions) with different recall periods. In phase 2, we evaluated the psychometric properties of both versions of DSI. The DSI was shown to be reliable, valid, and responsive to detect menstrual pain improvement. It can be adopted in research and clinical practice to facilitate the measurement and management of dysmenorrhea.