{"paper_id":"cac71fa2-8919-451a-9155-d5f9a3df55d0","body_text":"1 \nSelf-reported and tracker-estimated physical activity outcomes in women with chronic \npelvic pain disorders: A longitudinal evaluation of construct validity \n \nTanisee Nagaldinne,1 Samia Shahnawaz,1 Suzanne R Bakken,2 Noemie Elhadad,3 Emma N \nHoran, 3 Carol Ewing-Garber,4 Jovita Rodrigues,1 Matteo Danieletto,1 Kyle Landell,1 Ipek Ensari1 \n \n \n1Windreich Department of Artificial Intelligence and Human Health, Icahn School of Medicine at \nMount Sinai, New York, NY, 10027 \n2Columbia University School of Nursing, New York, NY, 10031, \n3Department of Biomedical Informatics, Columbia University Irving Medical Center, New York, \nNY, 10031 \n4Department of Biobehavioral Sciences, Teachers College Columbia University, New York, NY, \n10027 \n \n \nCorresponding Author: Ipek Ensari \nWindreich Department of Artificial Intelligence and Human Health \nIcahn School of Medicine at Mount Sinai \n3 East 101st Street, Room 1009, New York, 10029 \n \nWord count=4,380 \n \n \n  \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \nNOTE: This preprint reports new research that has not been certified by peer review and should not be used to guide clinical practice.\n\n 2 \nAbstract \nObjective: This study aims to evaluate the short form International Physical Activity \nQuestionnaire (IPAQ) for use in women with chronic pelvic pain disorders (CPPDs) by \ncomparing its scores against objectively-estimated physical activity (PA) outcomes. We \ninvestigated IPAQ components that are most consistently predictive of habitual PA behavior.  \nMethod: The study sample included 966 weeks of data from 112 women with CPPDs who \nenrolled in a 14-week mHealth-based self-tracking study. Participants wore Fitbit devices and \ncompleted the IPAQ every week. We compared the IPAQ-reported minutes of walking, total \nactivity, sitting, light-, moderate-, and vigorous intensity PA for concordance and divergence \nagainst their corresponding Fitbit estimates. We used linear mixed-effects regression models \n(MLMs) for all analyses and quantified the between-participant variance in the magnitude of \nagreement between the two methods via random slope terms. We further evaluated temporal \nconsistency in scores using intraclass correlation coefficients (ICCs). \nResults: IPAQ-reported walking minutes were strongly associated with Fitbit step counts (B = \n3952.36; p = 0.006), minutes of moderate PA (B = 15.498; p = 0.0113), and moderate-to-\nvigorous PA (MVPA; B = 28.973; p = 0.007). IPAQ total activity minutes were associated with \nFitbit minutes of vigorous PA (B = 15.183; p = 0.007) and MVPA (B = 25.658; p = 0.010). IPAQ \nmoderate activity minutes were predictive of Fitbit vigorous PA minutes (B = 9.060; SE = 3.719; \np = 0.0151). There was substantial between-individual variance in these point estimates based \non the significant random-effect terms, and average weekly PA level was a significant \nmoderator of the association between IPAQ-reported and Fitbit-estimated scores for these \nvariables. IPAQ-reported sitting minutes were inversely associated with Fitbit step counts (B = -\n3125.61; p = 0.004), and minutes of MVPA (B = -21.848; p = 0.007), vigorous AP (B = -10.854; \np = 0.042), and moderate PA (B = -10.985; p = 0.004).  \nConclusion: These findings provide support for using IPAQ-reported walking and total activity \nminutes to monitor several PA domains in women with CPPDs, given their concordance with \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 3 \nseveral tracker-estimated PA outcomes. However, the item on “sitting time” may not be a \nsuitable for assessing sedentary time.  \n \n  \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 4 \nIntroduction  \nChronic pelvic pain disorders (CPPDs) affect a significant proportion of women \nworldwide, with prevalence estimates ranging from 5.7-26.6%.1 This cluster of conditions, \nincluding endometriosis, uterine fibroids, interstitial cystitis/bladder pain syndrome, and other \nsimilar diseases characterized by pelvic pain, inflammation, and other related symptoms, \nsignificantly impair quality of life (QoL) and well-being.1–3 Though scarce, existing studies \ndemonstrate that women with CPPDs engage in less physical activity (PA) and spend more time \nin sedentary lifestyle.4,5 This could be due to numerous reasons, such as worsened physical \nfunctioning that hinders PA behavior, CPPD related fatigue, or worry that PA will exacerbate \nsymptoms. Nevertheless, this posits a significant health risk for this population as physical \ninactivity is considered a major public health risk that is associated with a number of adverse \nhealth outcomes including all-cause mortality, cardiovascular and metabolic diseases, obesity, \nand depression.3,6,7 Specifically for those with CPPDs, there is growing recognition of the \nimportance of PA in managing symptoms and improving health outcomes, such as pain \nreduction, improved mood, and enhanced physical function.6,8 Accurate and comprehensive \nmeasurement of PA in this population is therefore clinically relevant for monitoring adherence to \nPA recommendations, assessing intervention efficacy and the relationship of PA to symptom \nseverity, and planning personalized treatments that incorporate appropriate PA goals.  \nDespite its clinical importance, there is a lack of high-quality data on PA patterns and \nmeasurement in CPPDs. Most studies to date rely solely on cross-sectional self-report \nmeasures, which are convenient and cost-effective, but also subject to recall bias and potential \noverestimation of PA levels.9,10 Additionally, the accuracy and consistency of self-reported PA \nestimates may vary depending on the unique characteristics of different patient groups.11,12 For \nexample, various chronic disease related factors such as fatigue, sleep quality, or opioid use \ncould impact recall and self-report accuracy.13 Despite the increased use of wearables for PA \nmeasurement due to their advantages (e.g., low participant burden, accuracy, granularity of \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 5 \ndata), self-reported measures of PA are still widely used in research and national surveillance \nsystems due to their practicality, low cost, and ease of administration, along with providing \ncontextual information about PA that is not possible to infer from trackers.14,15 Systematic \nreviews conclude that both self-report surveys and consumer grade trackers can meet \nacceptable accuracy for measuring PA behavior, with variability depending the specific PA \ndomain and user population.16,17 This highlights the importance of context and population-\nspecific evaluation of the commonly used standardized PA measures.  \nThe present study aims to address this gap by evaluating the construct validity of the \nshort form International Physical Activity Questionnaire (IPAQ)18 among individuals with CPPDs \nand identify those IPAQ components that are most consistently predictive of habitual PA \nbehavior in this population. The questions on the IPAQ cover domains of walking, moderate- \nand vigorous-intensity PA, and sedentary time (i.e., weekday sitting minutes) in total weekly \nminutes. The responses to these questions are multiplied by their respective metabolic \nequivalent (MET) and a total PA activity score is computed. It is the most widely used self-report \nmeasure of PA and has been validated in various populations and settings.12 For example, \nprevious research assessing the IPAQ in patients with fibromyalgia and axial spondyloarthritis \n(axSpA) report varying degrees of weak to moderate agreement between IPAQ- versus \nSensewear armband or Actigraph-based estimates of PA (i.e., Pearson’s r = 0.04-0.19 for \nfibromyalgia; r = 0.032-0.367 for axSpA).19,20 Another study conducted with healthy adults21 \nreported small (sedentary, moderate, and walking minutes, r = 0.16, 0.09, and 0.17, \nrespectively, p < 0.001) to moderate (vigorous intensity and total PA; r = 0.29 and 0.40, p < \n0.001) correlations between synchronous paired measurements from Fitbits and the IPAQ. \nHowever, IPAQ has not been comprehensively evaluated for use among patients with CPPDs. \nThis is an important point of inquiry because the nature of the disease might differentially impact \nthe degree to which the IPAQ domains are predictive of their reference benchmark (e.g., tracker \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 6 \nestimates). Elucidating these potential differences is necessary for making informed decisions \non how to best measure PA behavior in this population.  \nAccordingly, the objectives of this study are threefold: 1) Evaluate construct validity \nthrough concordance and divergence between IPAQ-reported vs Fitbit-estimated scores on \ndifferent PA domains, 2) Examine between-individual variability in the agreement between the \ntwo measurement methods, and 3) Assess temporal consistency between the methods through \ncomparison of the magnitude of variability over time. From a measurement science perspective, \nthis evaluation requires consideration of the dynamic nature of PA behavior and any systematic \nindividual differences in measurement.22 We address these aspects through a longitudinal, \nrepeated-measures design and a flexible mixed-effects multilevel modeling (MLM) framework. \nWe hypothesized that: 1) IPAQ-reported minutes of walking and different intensities of PA would \npositively correlate with Fitbit-estimated step counts and moderate- and vigorous intensity PA \nscores, and negatively correlate with Fitbit-estimated sedentary minutes, 2) IPAQ-reported \nsitting minutes would positively correlate with Fitbit-estimated sedentary minutes, and negatively \ncorrelate with Fitbit-estimated minutes of all PA intensities, and 3) there would be substantial \nvariability in the between-individual variance in the point estimates and a moderating effect of \nhabitual PA levels on the association of IPAQ-reported to Fitbit-estimated outcomes.  \n \nMethods \nStudy Design and Data Collection \nThe study design and procedures were approved by the IRB of the Icahn School of \nMedicine at Mount Sinai (ISMMS; IRB# STUDY-22-01002). This is a secondary analysis of the \ndata from an ongoing larger study that aims to design, develop, and evaluate CPPD-specific \nmHealth measures from patient generated health data with high complexity and temporality \nusing non-linear distributed lag and functional data modeling (NIH/NICHD: R01HD108263). It \nuses an observational study design to collect 90 days of data on patient self-tracked symptoms \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 7 \nvia a research mHealth app ehive23 and passively collected activity data using activity trackers \nfrom participants. All participants used the ehive research study app for providing the baseline \nand weekly data on overall health, symptoms, well-being and health behaviors, as well as for \nreceiving prompts and reminders about the study.23 Participants were instructed to wear a Fitbit \nfor the duration of the study on their non-dominant wrist continuously for the 14-week study \nperiod, except during water-based activities and charging. They also completed the IPAQ short \nform at the end of each week, reporting their PA behavior for the preceding 7 days.  \n \nStudy Sample \nRecruitment and enrollment procedures are explained in detail elsewhere.24 Briefly, \nparticipants were recruited from all campuses of the Mount Sinai Health System (MSHS) and \nColumbia University Irving Medical Center (CUIMC) via email advertisements and on the \nmyChart by EPIC mobile app for MSHS patients. Inclusion criteria were: (1) individuals with a \nfemale reproductive anatomy and diagnosis of a CPPD (e.g., endometriosis, uterine fibroids, \ninterstitial cystitis, or chronic pelvic pain syndrome), (2) experiencing pelvic pain for at least 6 \nmonths, and (3) ability to read and understand English. Exclusion criteria included pregnancies \nand any major comorbidities. All participants provided written informed consent. \n \nData Processing \n Raw Fitbit data were downloaded and aggregated into weekly totals for steps, \nsedentary minutes, and minutes of light, moderate, and vigorous physical activity. IPAQ \nresponses were scored according to the IPAQ scoring protocol25 to derive weekly totals for \nminutes of walking, and moderate- and vigorous intensity PA. For data quality, weeks with fewer \nthan 4 valid days of Fitbit wear time (≥10 hours/day) were excluded from the analysis.26 IPAQ \nand Fitbit data were cleaned according to standard protocols, removing outliers and implausible \nvalues. Weekly MVPA for the Fitbit and IPAQ were calculated by adding up the weekly totals of \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 8 \nmoderate-intensity activity and vigorous-intensity activity. Total weekly activity for the IPAQ was \ncalculated by adding up minutes of walking, moderate-, and vigorous- intensity PA.27 \n \nData Analysis \nWe calculated descriptive statistics for all outcomes and demographics to characterize \nthe study sample. While bivariate correlational analyses and metrics are the traditionally used \nfor measure evaluation, they are not suitable for nested data where there is a distinct grouping \nstructure and when the outcome isn’t expected to be time-invariant.28 We instead used MLMs29 \nto compare the IPAQ-based and Fitbit-estimated scores on multiple PA domains for \nconcordance and divergence, and the between-individual variance in these associations. We \nused intraclass correlations (ICCs) within a MLM framework to evaluate variability in scores and \nconsistency in inter-method agreement, as this framework accommodates the nested structure \nof the data  \n \nTemporal Consistency and Intraclass Correlations. Given the non-time-invariant nature of PA \nbehavior, we assessed agreement across the 14 weeks through comparison of the amount of \nvariability in scores captured by the two methods. ICCs are well-suited for this purpose as they \ncan quantify the degree of absolute agreement among measurements and account for \nsystematic differences between measurement methods.30 First, we conducted MLMs with \nbootstrapping for significance testing31 to estimate within- and between group variance with 95% \nconfidence intervals (CIs). The resulting “repeatability” statistic is analogous to the ICC31,32 \nwhere higher R values indicate stronger group-level effects (i.e., between-group variance) and \ntherefore more consistent within-group measurements. Published recommendations categorize \nthe magnitude of the values as weak (0.0 - 0.3), moderate (0.3 - 0.7), and strong (0.7 - 1.0).29  \nNext, we calculated inter-rater reliability (ICC2r) and intra-rater reliability (ICC2a) via two-\nway random effects model (i.e., ICC(2,1))33 using the R irrICC library.34 Relying on the \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 9 \nmethodology by Gwet,34,35 they are computed using variance decomposition (i.e., σ²subject, \nσ²rater, σ²error) where; ICC2r = σ²s/(σ²s + σ²r + σ²e) and ICC2a = (σ²s + σ²r)/(σ²s + σ²r + σ²e). \nICC2r represents absolute agreement between methods (i.e., “raters”), while ICC2a represents \nthe pooled intra-method (i.e., “rater”) consistency over time. The key difference is that ICC2a \naccounts for systematic differences between the two methods by including the rater variance \n(σ²r) in the numerator. To assess how individual response patterns affect the agreement \nbetween the two methods (i.e., generalizability across the sample), we computed ICC2a in \nmodels with a subject-rater interaction (i.e., σ²sr).36 The no-interaction model assumes a \nconsistent relationship between IPAQ and Fitbit measurements across participants, while the \ninteraction model accounts for participant-specific patterns in how the two methods relate.34,36 \n \nConcordance and Divergence Analyses. To assess construct validity, we evaluated \nconcordance and divergence between IPAQ- vs Fitbit-based PA scores. We conducted \nseparate MLMs for each IPAQ domain (i.e., steps/walking, minutes of PA intensities, total PA \nminutes, sedentary/sitting) to evaluate their concordance and divergence with the Fitbit-\nestimated PA variables. For concordance (i.e., expected positive association) evaluation, we \nregressed the Fitbit-estimated outcome on its corresponding IPAQ-based outcome as the \nprimary predictor. For divergence evaluation (i.e., expected inverse association), we regressed \nFitbit-estimated PA scores on the IPAQ-based sitting minutes scores as the predictor. For all \nmodels, we scaled the IPAQ regressor (i.e., by subtracting their mean and dividing by their \nstandard deviations) in the models as this is standard and recommended for better \nconvergence. We further included the number of weeks in the study as a covariate to adjust for \nthe number of measurements per participant.  \nAll models included a random intercept term for participant to account for the hierarchical \ndata structure and correlated errors. The random intercept estimate allows comparison of the \nproportion of variance in the PA outcome attributable to between-participant differences. Based \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 10 \non our a priori hypothesis that the fixed effects would vary by participant, we also estimated a \nrandom slope term for the IPAQ-based predictor in each model based on recommendations for \nfitting MLMs.37 This allows the slope of the point estimate to vary across participants and \nevaluates the between-individual variance in the predictor point estimate. The model estimated \ncorrelation between the random intercept and slope indicates the relationship between the \nintercepts and coefficients. A statistically discernable random slope (p<0.05 based on the \nsingle-term deletion likelihood ratio test statistic; LRT) indicates significant between-individual \nvariance in the predictor point estimate. If the variance in the random slopes is estimated close \nto, or at, zero, the model containing random slopes will have a singular fit and therefore \nconvergence error. In these instances, the random slope component was removed from the final \nmodel, as per guidelines.38  \n \nModerator Analysis. We assessed habitual MVPA level as a potential moderator of the \nrelationship between self-reported IPAQ and Fitbit-estimated outcomes. In each MLM, we \nincluded an interaction term between each participant’s habitual weekly PA (i.e., weekly MVPA \nminutes averaged across their total number of weeks) and the IPAQ predictor variable, in \naddition to the random intercept and total number of weeks as co-variate as before. All analyses \nwere performed using R version 4.0.3, with the lme4, lmerTest, and rptR libraries in R/RStudio. \nStatistical significance was set at p < 0.05. \n  \nResults \nStudy Sample Characteristics \nSummary statistics for the study sample demographics and PA outcomes are provided \nin Table 1. The final analytic sample included 966 person-level weeks of data from 112 women \naged 18-54 years (mean age 35.59 ± 8.70 years). On average, participants provided 10.81 ± \n3.36 weeks of data each. The majority of participants (65.28%) had a diagnosis of \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 11 \nendometriosis, 16.67% had uterine fibroids, and 27.78% had other forms of chronic pelvic pain \ndisorder (Table 1). Body Mass Index (BMI) of the sample ranged from 15.96 to 48.14 kg/m² \n(mean 26.42 ± 5.91 kg/m²). The highest level of education was a college degree or higher \n(84.7%). \n \nTemporal Consistency and Intraclass Correlations  \nResults of the MLMs comparing temporal consistency for all PA domain scores are \nprovided in Table 2. For the IPAQ, the highest repeatability was observed for walking minutes \n(R = 0.766), followed by total PA Minutes (R = 0.642) and sitting minutes (R = 0.595). For Fitbit-\nestimated outcomes, step counts had the highest consistency (R = 0.601), followed closely by \nmoderate PA minutes (R = 0.596) and vigorous PA Minutes (R = 0.561). The repeatability \nvalues for the remaining PA domains were lower, though statistically significant. The results of \nthe inter- and intra-method reliability analyses are provided in Table 2 (Columns 5,6,7). Walking \nminutes/steps demonstrated the strongest agreement (i.e., 0.402 for both ICC2r and ICC2a), \nindicating more consistent absolute agreement between methods. Total PA was associated with \nlower inter-method agreement (i.e., ICC2r = 0.202) and higher overall intra-method consistency \n(ICC2a = 0.555). Sedentary behavior had the largest difference between the 2 agreement types \n(i.e., ICC2r = 0.101 vs ICC2a = 0.650). Models accounting for subject-method interaction \nyielded increased ICCa values for all PA domains, with the largest change observed for walking \nmins/step counts. \nConcordance and Divergence Analyses. \nThe results from the MLMs evaluating concordance of IPAQ scores to FitBit estimates \nare provided in Table 3. IPAQ-reported walking minutes showed significant positive associations \nwith several Fitbit-measured outcomes, including step counts (B = 10750.15, p = 0.0001), light \nPA minutes (B = 66.06, p = 0.0124), moderate PA minutes (B = 36.371, p = 0.0026), and MVPA \nminutes (B = 78.721, p = 0.0014). Reverse-scaling for easier interpretation, these point \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 12 \nestimates suggest that for every 1-minute increase in IPAQ-reported walking; Fitbit step counts \nincrease by approximately 8.93 steps, Fitbit light PA minutes increase by 0.055 minutes, Fitbit \nmoderate PA minutes increase by 0.030 minutes, and Fitbit MVPA minutes increase by 0.065 \nminutes. IPAQ-reported total activity minutes were significantly associated with Fitbit vigorous \nPA minutes (B = 29.532, p = 0.0106) and Fitbit MVPA minutes (B = 57.319, p = 0.0095). When \nreverse-scaled, these correspond to increases in Fitbit MVPA by 0.066 minutes and in Fitbit \nvigorous PA minutes by 0.034 minutes, for every 1-minute increase in IPAQ-reported total \nactivity. The models evaluating the association of IPAQ sitting minutes to Fitbit-estimated \nsedentary minutes were not statistically significant (p>0.05). Finally, all random slope terms \nwere positive and statistically significant (See Figures 1-5), indicating that individuals with higher \nbaseline levels of a given PA outcome (i.e., intercepts) also tend to have steeper increases in \nthe Fitbit-estimated scores as their IPAQ-reported PA increase (i.e., slopes).  \nThe results from the MLMs evaluating divergence (i.e., hypothesized inverse \nrelationship) between IPAQ-reported sitting minutes and the Fitbit-estimated PA scores are \nprovided in Table 4. We report those that indicate statistically significant associations. IPAQ-\nreported sitting minutes were negatively associated with Fitbit step counts (B = -5222.01, p = \n0.0061), MVPA minutes (B = -27.085, p = 0.0136), vigorous PA minutes (B = -13.717, p = \n0.0378), moderate PA minutes (B = -11.327, p = 0.0095), and light PA minutes (B = -82.55, p = \n0.0167). Reverse-scaling the point estimates, these correspond to decreases in Fitbit step \ncounts by ~1.36 steps, Fitbit MVPA minutes by 0.0070 minutes, Fitbit vigorous PA minutes by \n0.0036 minutes, and Fitbit light PA minutes by 0.021 minutes, for each 1-minute increase in \nIPAQ-reported sitting time. Finally, none of the random term correlations were significant (See \nTable 4), though they were all negative as expected.  \n \nModeration by habitual MVPA. Results of the mixed effects model estimating the moderator \neffect of average weekly (i.e., habitual) MVPA are provided in Table 5. The point estimate for \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 13 \nthe interaction term indicated that overall, habitual MVPA positively moderates the relationships \nbetween IPAQ total activity minutes and Fitbit vigorous PA minutes (B = 20.214, p < 0.0001), \nIPAQ walking minutes and Fitbit step counts (B = 5045.209, p < 0.0001), and IPAQ walking \nminutes and Fitbit moderate PA minutes (B = 34.080, p < 0.0001). Reverse-scaling, every 1-\nminute increase in IPAQ-reported total activity was associated with an increase in Fitbit vigorous \nPA minutes by 0.018 minutes, when controlling for habitual MVPA levels. Similarly, for every 1-\nminute increase in IPAQ-reported walking, Fitbit step counts increase by ~4.81 steps, Fitbit light \nPA minutes increase by 0.053 minutes, and Fitbit moderate PA minutes increase by 0.021 \nminutes, when controlling for habitual MVPA. The models controlled for habitual physical activity \nmeasuring the relationship of Fitbit vigorous, moderate, and sedentary minutes to their \ncorresponding IPAQ variable were not significant (p > 0.05).   \n \nDiscussion \n This study aimed to evaluate scores from the IPAQ against objectively estimated PA \nscores from Fitbit devices in women with CPPDs. Accurate, comprehensive PA measurement \nwhilst minimizing participant burden and resource requirements is important for managing \nsymptoms and improving quality of life, however, the utility and validity of self-report PA tools \nlike the IPAQ remains underexplored in CPPDs.1 The strong associations between several \nIPAQ- and Fitbit-based scores suggest that the IPAQ can provide reasonable estimates of some \nactivities (e.g., walking) in women with CPPDs. We further report systematic variations across \nindividuals in PA reporting, supported by variance over time and effect of habitual MVPA levels. \nTo our knowledge, this is the first study to evaluate the IPAQ in individuals with CPPDs and \nthrough longitudinal data. \nThe results from the ICC analyses provide additional insights into the temporal \nagreement patterns between the two methods. The walking minutes/steps domain \ndemonstrated the highest consistency for both the IPAQ and the Fitbit (i.e., R = 0.766 and \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 14 \nR=0.601, respectively), followed by total PA (i.e., R=0.646 and R=0.470, respectively). It is \npossible that walking behavior is easier to recall due to its routine and identifiable nature. \nPrevious studies report a wide range of ICCs for IPAQ-based walking minutes,19,39–43 though the \ndata collection period in those studies were 2 weeks or shorter. Next, we observed that IPAQ-\nreported MVPA, vigorous, and moderate PA minutes had lower consistency compared to their \nFitbit counterparts (Table 2). This may indicate difficulty in accurately categorizing higher \nintensity PA through self-report, a finding in line with other similar studies.44 The consistently \nhigher ICC2a values with the inclusion of the interaction term suggest systematic patterns in \nindividual PA reporting relative to Fitbit estimates, most evident in total PA and sedentary \nbehavior domains. Walking minutes/steps demonstrated the most consistent agreement (ICC2r \n= ICC2a = 0.402). The low ICC2r values for vigorous and moderate PA with improvements in \nICC2a when accounting for interaction indicate poor absolute agreement but consistent relative \npatterns within individuals. Finally, sitting minutes demonstrated moderate to strong temporal \nstability for both IPAQ (R = 0.595) and Fitbit (R = 0.810) measures, yet indicated poor inter-\nmethod agreement (ICC2r = 0.101). This demonstrates that while the measurements remain \nrelatively consistent over time, there are substantial differences in how this behavior is captured \nor defined between the two methods. This is expected given the IPAQ and the Fitbit measure \nsomewhat different constructs (self-reported sitting vs. any movement that falls below the \nthreshold for light intensity). As such, exact conceptual equivalence cannot be assumed.45 \nTaken together, these results suggest that IPAQ and Fitbit may be more suitable for tracking \nindividual-level changes over time, particularly based on walking behavior and total PA. They \nalso highlight the importance of considering person-specific measurement properties in PA \nassessment, a key principle in measurement science.46 \nThe results of the validity evaluation through MLMs indicated that IPAQ-reported walking \nminutes were positively associated with numerous Fitbit-measured outcomes including step \ncount, MVPA, moderate activity minutes, and light PA minutes. Similarly, IPAQ total activity \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 15 \nminutes indicated good concordance with vigorous PA and MVPA measured by Fitbit. These \nfindings suggest that the IPAQ item on walking minutes might be the most efficient for \nretrospective PA assessment in this population, whereas total PA might be useful in estimating \nhigher-intensity activities. As such, these two items may be possible replacements in instances \nwhere resource constraints limit the use of accelerometer devices,9 or when wearing continuous \nactivity monitors is not feasible. A study validating the IPAQ in visually impaired adults47 \nsimilarly reported that IPAQ walking minutes were predictive of steps, light PA and total PA, but \nnot MVPA. In contrast, another study42 in healthy college students (74% female) demonstrated \nthat whilst IPAQ-reported walking minutes did not correlate with any of the Fitbit-based PA \noutcome scores, moderate- and vigorous intensity PA minutes were associated with their Fitbit-\nbased counterparts as well as step counts. The discrepant findings could be due to the shorter \nstudy period (1 week), use of ActiGraph devices, or differences in sample demographics.48 \nNevertheless, these results underscore the importance of considering population-specific \nfactors when interpreting inferences made from self-reported PA data or choice of PA domain \nwhen designing interventions4,6 (e.g., MVPA for targeting pain management,49 step counts for \ntargeting depressive symptoms or physical function).50,51 \nWe further report individual variability in the relationship between IPAQ-reported and \nFitbit-measured PA outcomes, based on the random slope terms in the MLMs. These effects \nwere significant for step counts, MVPA, vigorous and moderate intensity PA (Table 3), and point \nto individual differences in PA reporting among women with CPPDs. The strong correlations \nbetween the random intercepts and slopes corroborate the non-significant point estimates \nobserved for the fixed-effect terms (i.e., average group effect) in these models. That is, the \nmagnitude of variance between groups (i.e., level - 2 random effects) will in principle be inverse \nto the magnitude of a fixed-effect term (i.e., level -1 predictor). Accordingly, the generalizability \nof these results appears strongest for walking-based assessments. In contrast, IPAQ-reported \nsitting minutes were not significantly associated with Fitbit-estimated sedentary minutes. They \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 16 \nwere, however; significantly and inversely associated with other domain scores, providing \nsupport for divergence between these constructs. Moreover, these associations were \nhomogenous across the sample based on the non-significant random slope term. Our findings \non the low concordance in sedentary time estimates align with those from several other studies \nusing a variety of devices (e.g.,GT1M ActiGraph accelerometer52,53, SenseWear Pro Armband; \nSWA19, Xiaomi Mi Band 214). Collectively, these results suggest that sedentary time is relatively \nmore challenging to accurately recall and report regardless of population demographics.  \nFinally, our analyses identified habitual MVPA levels as a moderator of the relationship \nbetween IPAQ walking minutes and Fitbit step counts and Fitbit moderate PA minutes, as well \nas between IPAQ total activity minutes and Fitbit vigorous PA minutes (See Table 5). It is \npossible that individuals who are more active have better awareness and recall of their activity \npatterns, based on the stronger associations between their IPAQ-reported and Fitbit-estimated \nPA scores. On the other hand, the interaction term was not significant for predicting Fitbit-based \nlight PA. This pattern of significance follows that of the random slope terms (Table 3). One \nexplanation is that the between-individual variance in those models is explained away once \nhabitual MVPA is included as an interaction term (Table 5). This would then suggest that the \nunobserved between-individual differences were largely due to habitual MVPA differences \nand/or closely linked factors. For example, habitual MVPA was demonstrated to predict weight \nchange, weight compensation, and changes in energy intake.54 Taken together, these results \ndemonstrate that an individual's typical activity level influences their reporting accuracy, and \ninterpretation of IPAQ-reported PA scores are more generalizable to individuals with similar \nhabitual activity levels within the CPPD population. \n \nLimitations \nWe acknowledge several limitations of this study. First, the study sample was relatively \nhomogenous in demographic factors, including education and employment status. These \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 17 \ndemographic characteristics may influence PA levels and self-reporting, potentially limiting the \ngeneralizability of the findings. We did not include these as additional moderators in our \nanalyses because of the relative homogeneity in the sample and they were beyond the scope of \nour a priori objectives. Future studies may benefit from additionally evaluating these factors \nbased on evidence for their influence on PA behavior and self-reporting.55,56 The use of Fitbit \ndevices also poses limitations, including potential under- or overestimation of various PA \nparameters.57–59 They may not accurately capture certain types of activities, such as cycling, \nswimming,60 weightlifting or stationary exercises.61 In contrast to prior studies using 1-2 weeks \nof data, the longer study period herein enables a more robust and comprehensive analyses. \nNevertheless, some participants in our sample had fewer weeks of data, prompting us to adjust \nfor this variability in our regression analyses. The effect of number of weeks were not significant \nfor most of the models, however; the magnitude of the point estimates could increase with more \ndata points, especially for those with higher variability (i.e., lower ICCs).  \n \nConclusion \nIn conclusion, this study provides insights into the strengths and limitations of using the \nIPAQ for assessing PA in women with CPPDs. The strong associations between IPAQ walking \nminutes and various Fitbit-measured PA scores suggest that self-reported walking might be a \nmore comprehensive indicator of overall PA than previously thought in this population. Our \nfindings further point to individual differences in PA reporting, which underscores the need for \ncaution in generalizing results across all individuals with CPPDs and considering systematic \ndifferences in reporting across individuals. Similarly, considering habitual activity levels when \ninterpreting self-reported PA data is warranted, as more active individuals demonstrated greater \nconcordance in their self-reports. The IPAQ can be a useful complementary tool for PA \nmeasurement in clinical and research settings, especially when resources limit the use of more \nexpensive accelerometer devices.  \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 18 \nFunding Statement: This study was supported by the Eunice Kennedy Shriver National \nInstitute Of Child Health & Human Development of the National Institutes of Health under Award \nNumber R01HD108263 (PI=Ensari). The content is solely the responsibility of the authors and \ndoes not necessarily represent the official views of the National Institutes of Health. \n \nData Availability Statement: The data used in this study can be made available after the active \ngrant period is over upon reasonable request to the corresponding author. \n \n \n \nReferences \n1.  Ahangari A. Prevalence of chronic pelvic pain among women: an updated review. Pain \nPhysician 2014; 17: E141-147. \n2.  Till SR, As-Sanie S, Schrepf A. Psychology of Chronic Pelvic Pain: Prevalence, \nNeurobiological Vulnerabilities, and Treatment. Clin Obstet Gynecol 2019; 62: 22–36. \n3.  Siqueira-Campos VM, de Deus MSC, Poli-Neto OB, et al. Current Challenges in the \nManagement of Chronic Pelvic Pain in Women: From Bench to Bedside. Int J Womens \nHealth 2022; 14: 225–244. \n4.  Sachs MK, Dedes I, El-Hadad S, et al. Physical Activity in Women with Endometriosis: \nLess or More Compared with a Healthy Control? Int J Environ Res Public Health 2023; \n20: 6659. \n5.  Tricoche B, Caceres B, Shaw L, et al. Physical Activity Phenotypes in Endometriosis \nUsing Unsupervised Learning via Functional Mixture Models. Rev. \n6.  Ambrose KR, Golightly YM. Physical exercise as non-pharmacological treatment of \nchronic pain: Why and when. Best Pract Res Clin Rheumatol 2015; 29: 120–130. \n7.  Poeta Do Couto C, Policiano C, Pinto FJ, et al. Endometriosis and cardiovascular \ndisease: A systematic review and meta-analysis. Maturitas 2023; 171: 45–52. \n8.  Awad E, Ahmed HAH, Yousef A, et al. Efficacy of exercise on pelvic pain and posture \nassociated with endometriosis: within subject design. J Phys Ther Sci 2017; 29: 2112–\n2115. \n9.  Prince SA, Adamo KB, Hamel ME, et al. A comparison of direct versus self-report \nmeasures for assessing physical activity in adults: a systematic review. Int J Behav \nNutr Phys Act 2008; 5: 56. \n10.  Adams SA, Matthews CE, Ebbeling CB, et al. The effect of social desirability and social \napproval on self-reports of physical activity. Am J Epidemiol 2005; 161: 389–398. \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 19 \n11.  Short ME, Goetzel RZ, Pei X, et al. How Accurate are Self-Reports? An Analysis of Self-\nReported Healthcare Utilization and Absence When Compared to Administrative \nData. J Occup Environ Med Am Coll Occup Environ Med 2009; 51: 786–796. \n12.  Lee PH, Macfarlane DJ, Lam TH, et al. Validity of the International Physical Activity \nQuestionnaire Short Form (IPAQ-SF): a systematic review. Int J Behav Nutr Phys Act \n2011; 8: 115. \n13.  Edgley K, Horne AW, Saunders PTK, et al. Symptom tracking in endometriosis using \ndigital technologies: Knowns, unknowns, and future prospects. Cell Rep Med 2023; 4: \n101192. \n14.  Domingos C, Correia Santos N, Pêgo JM. Association between Self-Reported and \nAccelerometer-Based Estimates of Physical Activity in Portuguese Older Adults. \nSensors 2021; 21: 2258. \n15.  Prince SA, Cardilli L, Reed JL, et al. A comparison of self-reported and device measured \nsedentary behaviour in adults: a systematic review and meta-analysis. Int J Behav \nNutr Phys Act 2020; 17: 31. \n16.  Feehan LM, Geldman J, Sayre EC, et al. Accuracy of Fitbit Devices: Systematic Review \nand Narrative Syntheses of Quantitative Data. JMIR MHealth UHealth 2018; 6: e10527. \n17.  Evenson KR, Goto MM, Furberg RD. Systematic review of the validity and reliability of \nconsumer-wearable activity trackers. Int J Behav Nutr Phys Act 2015; 12: 159. \n18.  Craig CL, Marshall AL, Sj??Str??M M, et al. International Physical Activity \nQuestionnaire: 12-Country Reliability and Validity: Med Sci Sports Exerc 2003; 35: \n1381–1395. \n19.  Segura-Jiménez V, Munguía-Izquierdo D, Camiletti-Moirón D, et al. Comparison of the \nInternational Physical Activity Questionnaire (IPAQ) with a multi-sensor armband \naccelerometer in women with fibromyalgia: the al-Ándalus project. Clin Exp \nRheumatol 2013; 31: S94-101. \n20.  Bayraktar D, Yuksel Karsli T, Ozer Kaya D, et al. Is the international physical activity \nquestionnaire (IPAQ) a valid assessment tool for measuring physical activity of \npatients with axial spondyloartritis? Musculoskelet Sci Pract 2021; 55: 102418. \n21.  Beagle AJ, Tison GH, Aschbacher K, et al. Comparison of the Physical Activity \nMeasured by a Consumer Wearable Activity Tracker and That Measured by Self-\nReport: Cross-Sectional Analysis of the Health eHeart Study. JMIR MHealth UHealth \n2020; 8: e22090. \n22.  Meredith W, Teresi JA. An Essay on Measurement and Factorial Invariance: Med Care \n2006; 44: S69–S77. \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 20 \n23.  Hirten RP, Danieletto M, Landell K, et al. Development of the ehive Digital Health App: \nProtocol for a Centralized Research Platform. JMIR Res Protoc 2023; 12: e49204. \n24.  Leventhal EL, Nukavarapu N, Elhadad N, et al. Trajectories of mHealth-tracked mental \nhealth symptoms and their predictors in chronic pelvic pain. medRxiv 2024; 2024–09. \n25.  Sjostrom M, Ainsworth B, Bauman A, et al. Guidelines for data processing analysis of \nthe International Physical Activity Questionnaire (IPAQ) - Short and long forms, \nhttps://www.semanticscholar.org/paper/Guidelines-for-data-processing-analysis-of-\nthe-and-Sjostrom-Ainsworth/efb9575f5c957b73c640f00950982e618e31a7be (2005, \naccessed 13 August 2024). \n26.  Chan A, Chan D, Lee H, et al. Reporting adherence, validity and physical activity \nmeasures of wearable activity trackers in medical research: A systematic review. Int J \nMed Inf 2022; 160: 104696. \n27.  Garashi NHJ, Kandari JRA, Ainsworth BE, et al. Weekly Physical Activity from IPAQ \n(Arabic) Recalls and from IDEEA Activity Meters. Health (N Y) 2020; 12: 598–611. \n28.  Rönkkö M, Cho E. An Updated Guideline for Assessing Discriminant Validity. Organ Res \nMethods 2022; 25: 6–14. \n29.  Bakdash JZ, Marusich LR. Repeated Measures Correlation. Front Psychol 2017; 8: 456. \n30.  Liu J, Tang W, Chen G, et al. Correlation and agreement: overview and clarification of \ncompeting concepts and measures. Shanghai Arch Psychiatry 2016; 28: 115–120. \n31.  Stoffel MA, Nakagawa S, Schielzeth H. rptR: repeatability estimation and variance \ndecomposition by generalized linear mixed-effects models. Methods Ecol Evol 2017; \n8: 1639–1644. \n32.  Nakagawa S, Johnson PCD, Schielzeth H. The coefficient of determination R2 and intra-\nclass correlation coefficient from generalized linear mixed-effects models revisited \nand expanded. J R Soc Interface 2017; 14: 20170213. \n33.  Shrout PE, Fleiss JL. Intraclass correlations: uses in assessing rater reliability. Psychol \nBull 1979; 86: 420–428. \n34.  Gwet KL. Package ‘irrICC’: Intraclass Correlations for Quantifying Inter-Rater \nReliability, https://CRAN.R-project.org/package=irrICC (2019). \n35.  Gwet KL. Handbook of Inter-Rater Reliability. 4th Edition. Advanced Analytics, LLC, \n2014. \n36.  McGraw KO, Wong SP. Forming inferences about some intraclass correlation \ncoefficients. Psychol Methods 1996; 1: 30–46. \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 21 \n37.  Bell A, Fairbrother M, Jones K. Fixed and random effects models: making an informed \nchoice. Qual Quant 2019; 53: 1051–1074. \n38.  Heisig JP, Schaeffer M. Why You Should Always Include a Random Slope for the Lower-\nLevel Variable Involved in a Cross-Level Interaction. Eur Sociol Rev 2019; 35: 258–279. \n39.  Tomioka K, Iwamoto J, Saeki K, et al. Reliability and Validity of the International Physical \nActivity Questionnaire (IPAQ) in Elderly Adults: The Fujiwara-kyo Study. J Epidemiol \n2011; 21: 459–465. \n40.  Blikman T, Stevens M, Bulstra SK, et al. Reliability and Validity of the Dutch Version of \nthe International Physical Activity Questionnaire in Patients After Total Hip \nArthroplasty or Total Knee Arthroplasty. J Orthop Sports Phys Ther 2013; 43: 650–659. \n41.  Carvalho FA, Morelhão PK, Franco MR, et al. Reliability and validity of two \nmultidimensional self-reported physical activity questionnaires in people with chronic \nlow back pain. Musculoskelet Sci Pract 2017; 27: 65–70. \n42.  Dinger MK, Behrens TK, Han JL. Validity and Reliability of the International Physical \nActivity Questionnaire in College Students. Am J Health Educ 2006; 37: 337–343. \n43.  Regaieg S, Charfi N, Yaich S, et al. The Reliability and Concurrent Validity of a Modified \nVersion of the International Physical Activity Questionnaire for Adolescents (IPAQ-A) \nin Tunisian Overweight and Obese Youths. Med Princ Pract 2015; 25: 227–232. \n44.  Dyrstad SM, Hansen BH, Holme IM, et al. Comparison of self-reported versus \naccelerometer-measured physical activity. Med Sci Sports Exerc 2014; 46: 99–106. \n45.  Herdman M, Fox-Rushby J, Badia X. A model of equivalence in the cultural adaptation \nof HRQoL instruments: the universalist approach. Qual Life Res Int J Qual Life Asp \nTreat Care Rehabil 1998; 7: 323–335. \n46.  Borsboom D. When does measurement invariance matter? Med Care 2006; 44: S176-\n181. \n47.  Marmeleira J, Laranjo L, Marques O, et al. Criterion-related Validity of the Short Form of \nthe International Physical Activity Questionnaire in Adults who are Blind. J Vis Impair \nBlind 2013; 107: 375–381. \n48.  Keating XD, Guan J, Piñero JC, et al. A Meta-Analysis of College Students’ Physical \nActivity Behaviors. J Am Coll Health 2005; 54: 116–126. \n49.  Fjeld MK, Årnes AP, Engdahl B, et al. Consistent pattern between physical activity \nmeasures and chronic pain levels: the Tromsø Study 2015 to 2016. Pain 2023; 164: \n838–847. \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 22 \n50.  Kaleth AS, Slaven JE, Ang DC. Increasing Steps/Day Predicts Improvement in Physical \nFunction and Pain Interference in Adults with Fibromyalgia. Arthritis Care Res 2014; \n66: 1887–1894. \n51.  Grunberg VA, Greenberg J, Mace RA, et al. Fitbit Activity, Quota-Based Pacing, and \nPhysical and Emotional Functioning Among Adults With Chronic Pain. J Pain 2022; 23: \n1933–1944. \n52.  Ekblom Ö, Ekblom-Bak E, Bolam KA, et al. Concurrent and predictive validity of \nphysical activity measurement items commonly used in clinical settings– data from \nSCAPIS pilot study. BMC Public Health 2015; 15: 978. \n53.  Kaleth AS, Ang DC, Chakr R, et al. Validity and reliability of community health activities \nmodel program for seniors and short-form international physical activity \nquestionnaire as physical activity assessment tools in patients with fibromyalgia. \nDisabil Rehabil 2010; 32: 353–359. \n54.  Höchsmann C, Dorling JL, Apolzan JW, et al. Baseline Habitual Physical Activity \nPredicts Weight Loss, Weight Compensation, and Energy Intake During Aerobic \nExercise. Obesity 2020; 28: 882–892. \n55.  Winckers ANE, Mackenbach JD, Compernolle S, et al. Educational differences in the \nvalidity of self-reported physical activity. BMC Public Health 2015; 15: 1299. \n56.  Quinlan C, Rattray B, Pryor D, et al. The accuracy of self-reported physical activity \nquestionnaires varies with sex and body mass index. PLoS ONE 2021; 16: e0256008. \n57.  O’Driscoll R, Turicchi J, Hopkins M, et al. The validity of two widely used commercial \nand research-grade activity monitors, during resting, household and activity \nbehaviours. Health Technol 2020; 10: 637–648. \n58.  Brewer W, Swanson BT, Ortiz A. Validity of Fitbit’s active minutes as compared with a \nresearch-grade accelerometer and self-reported measures. BMJ Open Sport Exerc \nMed 2017; 3: e000254. \n59.  Matlary RED, Holme PA, Glosli H, et al. Comparison of free-living physical activity \nmeasurements between ActiGraph GT3X-BT and Fitbit Charge 3 in young people with \nhaemophilia. Haemophilia; 28. Epub ahead of print November 2022. DOI: \n10.1111/hae.14624. \n60.  Dorn D, Gorzelitz J, Gangnon R, et al. Automatic Identification of Physical Activity Type \nand Duration by Wearable Activity Trackers: A Validation Study. JMIR MHealth \nUHealth 2019; 7: e13547. \n61.  Bassett DR. Validity and Reliability issues in Objective Monitoring of Physical Activity. \nRes Q Exerc Sport 2000; 71: 30–36. \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 23 \n  \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 24 \n \nTable 1. Study Sample Characteristics \nSample Demographics Mean ± SD Range \nAge (N=91) 36.11 (9.55)  18 - 60 \nBody Mass Index (kg/m^2) 26.47 ± 5.93  15.96 - 48.14 \nWeeks of Data 10.14 ± 3.46 3 - 15 \nDiagnosis: \nEndometriosis  \nUterine Fibroids \nOther CPPD \nFrequency (%)  \n74 (65) \n19 (17)  \n 31 (28) \n \nRace: \nWhite \nBlack \nUnknown \nAsian \nMixed Race \nFrequency (%)  \n54 (48) \n26 (23) \n14 (13) \n12 (11) \n4 (4) \n \nEthnicity: \n     Hispanic or Latino \n     Not Hispanic or Latino \n     Unknown \nFrequency (%)  \n29 (26) \n76 (68) \n7 (6) \n \nEducation: \n     College \n     Some College \n     GED \n     High School \nFrequency (%)  \n96 (86) \n9 (8) \n2 (2) \n5 (4) \n \nEmployment: \n    Employed \n    Student \n    Unemployed/Unable to work \n    Unknown \n    Caregiver \nFrequency (%)  \n89 (79) \n6 (5) \n14 (13) \n1 (0.8) \n2 (1) \n \nMarital Status \n    Married \n    Divorced \n    Single \nFrequency (%)  \n50 (44) \n9 (8) \n53 (47) \n \nAnnual household income: \n    $200,000 or more \n    $100,000-$199,999 \n    $50,000-$99,999 \n    Under $50,000 \n    Unknown \nFrequency (%)  \n16 (14) \n32 (28) \n28 (25) \n18 (16) \n18 (16) \n \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 25 \nFitbit PA Outcomes Mean ± SD  \nVigorous PA Minutes  125.68 ± 117.86 0 - 729  \nModerate PA Minutes  116.08 ± 126.55 0 – 1,384  \nMVPA Minutes 241.76 ± 217.53 0 – 2,042 \nLight PA Minutes 1368.65 ± 554.96 120 – 3,726 \nWeekday Sedentary Minutes* 1356.67 ± 687.83 0 – 4,734  \nSteps 51464.42 ± 27169.19 3833 – 224,661 \nTotal Activity Minutes 1538.35 ± 726.34 41 – 5,005  \nIPAQ Self-Reported Outcomes Mean ± SD  \nVigorous PA Minutes 99.51 ± 333.53   0 – 5,040 \nModerate PA Minutes 107.07 ± 286.20 0 – 3,360 \nMVPA Minutes  169.50 ± 517.31  0 – 8,400 \nWalking Minutes 564.41 ± 830.75  0 – 5,040  \nWeekday Sitting Minutes* 460.61 ± 259.90 60 – 1,440  \nTotal Activity Minutes 731.77 ± 1146.70 0 – 13,440  \n*The data for the amount of time spent sitting is based on daily minutes (not weekly, unlike the \nother PA data). It is derived from the IPAQ question, \"During the last 7 days, how much time did \nyou spend sitting on a week day?\". Accordingly, the provided corresponding Fitbit-based \nestimates of sedentary minutes include week days. \n \nTable 2. Results from repeatability (R) and Intraclass correlation (ICC) estimation. ICC2r and ICC2a values \ndenote inter- and intra-method reliability without subject-method interaction.   \nIPAQ Variable \n(N*, n**) \nR (SE) Fitbit Variable \n(N=966, n=112) \nR (SE) ICC2r ICC2a ICC2a w/ \ninteraction \nWalking Min \n(N=602, n=102) \n0.766 (0.031), \n95% CI = 0.697, 0.817 \nStep Counts \n \n0.601 (0.038) \n95% CI= 0.513, 0.666 \n0.402 0.402 0.656 \nTotal PA Min \n(N=423, n=88) \n0.642 (0.045), \n95% CI = 0.542, 0.721 \nTotal PA Min \n(N=926, n=109) \n0.470 (0.042) \n95% CI = 0.384, 0.547 \n0.202 0.555 0.769 \nMVPA Min \n(N=539, n=100) \n0.154 (0.043), \n95% CI= 0.070, \nMVPA Min 0.578 (0.038) \n95% CI = 0.497, 0.645 \n0.152 0.171 0.288 \nVigorous PA Min \n(N=803, n=112) \n0.250 (0.040), \n95% CI = 0.171, 0.326 \nVigorous PA \nMin \n0.561 (0.038) \n95%CI = 0.479, 0.630 \n0.147 0.151 0.309 \nModerate PA Min \n(N=748, n=109) \n0.224 (0.042) \n95% CI = 0.142, 0.308 \nModerate PA \nMin \n0.596 (0.037) \n95% CI = 0.518, 0.665 \n0.152 0.152 0.305 \nSitting Min \n(N=540, n=90) \n0.595 (0.050), \n95% CI = 0.486, 0.686 \nWeek day \nSedentary Min \n0.303 (0.040) \n95% CI= 0.225, 0.383 \n0.101 0.650 0.688 \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 26 \n  Light PA Min 0.547 (0.039) \n95% CI = 0.466, 0.621 \n   \n*Number of person-level weeks of data included in the model, **Number of participants \nMVPA=moderate-to-vigorous intensity physical activity. 95% CI= 95% confidence intervals \n \n \n \n \n \n \n \n \n \n  \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 27 \n \nTable 3. Concordance Analysis: Results from 6 separate linear mixed-effects regression \nmodels. Point estimates and associated statistics are reported for the fixed effects. \n \nOutcome \n(Random slope correlation) \nPredictors \n(Fixed-Effect Term) \nB Coefficient (SE) t-value p-value \nFitbit Step Counts \n(Cor = 0.50, p < 0.0001) \nIntercept \nIPAQ Walking Min \nnweeks \n43226.09 (5706.72) \n10750.15 (2477.63) \n925.79 (569.45) \n7.575 \n4.339 \n1.626 \n<0.0001 \n0.0001 \n0.106 \nFitbit Light PA Min \n(N/A) \nIntercept \nIPAQ Walking Min \nnweeks \n1050.54 (114.99) \n66.06 (26.33) \n30.00 (11.57) \n9.136 \n2.509 \n2.593 \n<0.0001 \n0.0124 \n0.0108 \nFitbit Moderate PA Min \n(Cor = 0.88, p < 0.001) \nIntercept \nIPAQ Walking Min \nnweeks \n87.785 (24.141) \n36.371 (9.146) \n3.005 (2.410) \n3.636 \n3.977 \n1.247 \n0.00004 \n0.0026 \n0.2150 \nFitbit Vigorous PA Min \n(Cor = 0.78, p < 0.001) \nIntercept \nIPAQ Total Activity Min \nnweeks \n77.749 (27.550) \n29.532 (9.964) \n4.174 (2.705) \n2.822 \n 2.964  \n1.543  \n0.0056  \n0.0106 \n0.1260 \nFitbit MVPA Min \n(Cor = 0.72, p < 0.001) \nIntercept \nIPAQ Walking Min \nnweeks \n186.269 (43.123) \n78.721 (22.258) \n6.872 (4.241) \n4.319 \n 3.537 \n1.620   \n<0.0001 \n0.0014 \n0.1079 \nFitbit MVPA Minutes \n(Cor = 0.97, p < 0.001) \nIntercept \nIPAQ Total Activity Min \nnweeks \n165.196 (50.017) \n57.319 (18.677) \n6.484 (4.896) \n3.303 \n3.069 \n1.324 \n0.0012 \n0.0095 \n0.1884 \n \n \n \n \n \n  \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 28 \nTable 4. Divergence analysis: Results from 5 separate linear mixed-effects regression \nmodels. Point estimates and associated statistics are reported for the fixed effects. \nModel Outcome \n(Random slope correlation) \nModel Predictor \n(Fixed-Effect Term) \nB Coefficient (SE) t-value p-value \nFitbit Step Counts \n(Cor = -0.05, p > 0.05) \nIntercept \nIPAQ Sitting Minutes \nnweeks \n39,536.11 (5187.82)  \n-5,222.01 (1535.28) \n1047.93 (507.98) \n7.621 \n-3.401 \n2.063   \n< 0.0001 \n0.0061 \n0.0416 \nFitbit Light PA Minutes \n(Cor = 0.09, p = 0.089) \n \nIntercept \nIPAQ Sitting Minutes \nnweeks \n1105.02 (110.05) \n-82.55 (31.33) \n20.74 (10.76) \n10.041 \n-2.635 \n1.927 \n<0.0001 \n0.0167 \n0.0568 \nFitbit Moderate PA Minutes \n(Cor = -0.97, p = 0.09) \nIntercept \nIPAQ Sitting Minutes \nnweeks \n71.128 (17.795) \n-11.327 (3.753) \n3.660 (1.730) \n3.997 \n-3.018 \n2.116 \n0.0001 \n0.0095 \n0.0369 \nFitbit Vigorous PA Minutes \n(Cor = -0.48, p = 0.235) \nIntercept \nIPAQ Sitting Minutes \nnweeks \n73.167 (24.677) \n-13.717 (6.026) \n5.692 (2.410) \n2.295 \n-2.276 \n2.362 \n0.0036 \n0.0378 \n0.0199 \nFitbit MVPA Minutes \n(Cor = -0.61, p = 0.091) \nIntercept \nIPAQ Sitting Minutes \nnweeks \n145.593 (38.493) \n-27.085 (9.250) \n9.178 (3.751) \n3.782 \n -2.928 \n2.447  \n0.0002 \n0.0136 \n0.0162 \n \n \n \n \n \n \n \n  \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 29 \nTable 5. Results from 4 separate linear mixed-effects regression models that are estimating IPAQ self-reported \nactivity on different types of Fitbit data (step count, lightly active minutes, and vigorous activity minutes \nminutes). \nModel \nOutcome \nModel Predictors \n(Fixed effect term) \nB Coefficient (SE) t-value p-value \nFitbit Step \nCounts \nIntercept \nIPAQ Walking Min \nHabitual MVPA \nIPAQ Walking Min * Habitual MVPA \nnweeks \n51317.551 (3401.414) \n5785.888 (977.868) \n20043.456 (1231.025) \n5045.209 (893.586) \n0.166 (330.233) \n15.087 \n5.917 \n16.282 \n5.646 \n0.001 \n< 0.0001 \n< 0.0001 \n< 0.0001 \n< 0.0001 \n1 \nFitbit Light PA \nMin \nIntercept \nIPAQ Walking Min \nHabitual MVPA \nIPAQ Walking Min * Habitual MVPA \nnweeks \n1174.25 (107.53)  \n63.56 (25.62) \n201.67 (42.13) \n19.59 (21.26) \n17.74 (10.75) \n10.921 \n2.481 \n4.787 \n0.921 \n1.650  \n < 0.0001 \n0.0134 \n< 0.0001 \n0.3573 \n0.1019   \nFitbit Moderate \nPA Min \nIntercept \nIPAQ Walking Min \nHabitual MVPA \nIPAQ Walking Min * Habitual MVPA \nnweeks \n123.322 (15.149) \n25.507 (4.231) \n92.991 (5.581) \n34.080 (3.785) \n-0.737 (1.481) \n8.140 \n6.029 \n16.661 \n9.003 \n-0.498 \n< 0.0001 \n< 0.0001 \n< 0.0001 \n< 0.0001 \n0.62 \nFitbit Vigorous \nPA Min \nIntercept \nIPAQ Walking Min \nHabitual MVPA \nIPAQ Walking Min * Habitual MVPA \nnweeks \n131.322 (14.873) \n9.167 (4.516) \n87.267 (5.131) \n18.133 (4.364) \n-0.408 (1.416) \n8.829 \n2.030 \n17.008 \n4.155 \n-0.288 \n< 0.0001 \n0.043 \n< 0.0001 \n< 0.0001 \n0.773 \nFitbit Vigorous \nPA Min \nIntercept \nIPAQ Total Activity Min \nHabitual MVPA \nIPAQ Total Activity Min * Habitual MVPA \nnweeks \n122.115 (17.539) \n15.446 (5.119) \n84.555 (5.933) \n20.214 (4.945) \n-0.062 (1.670) \n6.962 \n3.017 \n14.251 \n4.088 \n-0.038 \n< 0.0001 \n0.0027 \n< 0.0001 \n< 0.0001 \n0.9700 \n \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 30 \n \nFigure 1. Model-estimated random (i.e., person-level) slopes for the model predicting Fitbit  \nstep counts (y-axis) from IPAQ-based walking minutes. Each participant is represented by one  \ndotted grey line (n=102). Each colored circle represents one person-level week (N=602).   \n \n \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 31 \n \nFigure 2. Model-estimated random (i.e., person-level) slopes for the model predicting Fitbit  \nMVPA minutes (y-axis) from IPAQ-based walking minutes. Each participant is represented by  \none dotted grey line (n=102). Each colored circle represents one person-level week (N=602).   \n \n \n \nFigure 3. Model-estimated random (i.e., person-level) slopes for the model predicting Fitbit  \nVigorous PA minutes (y-axis) from IPAQ-based total activity minutes. Each participant is  \nrepresented by one dotted grey line (n=88). Each colored circle represents one person-level  \nweek (N=423).   \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint \n\n 32 \n \nFigure 4. Model-estimated random (i.e., person-level) slopes for the model predicting Fitbit  \nMVPA minutes (y-axis) from IPAQ-based total activity minutes. Each participant is  \nrepresented by one dotted grey line (n=88). Each colored circle represents one person-level  \nweek (N=423).   \n \n \n \nFigure 5. Model-estimated random (i.e., person-level) slopes for the model predicting Fitbit  \nMVPA minutes (y-axis) from IPAQ-based sitting minutes. Each participant is represented by one  \ndotted grey line (n=90). Each colored circle represents one person-level week (N=540). \n . CC-BY-NC-ND 4.0 International licenseIt is made available under a \n is the author/funder, who has granted medRxiv a license to display the preprint in perpetuity. (which was not certified by peer review)\nThe copyright holder for this preprint this version posted February 12, 2025. ; https://doi.org/10.1101/2025.02.10.25322006doi: medRxiv preprint","source_license":"CC0","license_restricted":false}