Robust forecasting of sedentary bouts in chronic pelvic pain disorders for on-device learning and real-time deployment

Other OA: gold
⚙ AI-generated summary by qwen3.7-flash, 2026-10-05 ⓘ

This study developed a real-time forecasting framework using Fitbit data from 134 women with chronic pelvic pain disorders to predict sedentary bouts, demonstrating that models leveraging recent activity patterns can generate timely alerts for just-in-time adaptive interventions.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

⚙ AI-generated deep summary by qwen3.7-flash, 2026-10-05 · read from full text ⓘ

This study developed a computationally lean, on-device learning framework to forecast physical activity scores and detect sedentary bouts in real-time using wearable data. The researchers compared recursive least squares (RLS), seasonal autoregressive integrated moving average (SARIMA), and long short-term memory (LSTM) models, finding that RLS and SARIMA achieved comparable accuracy while better supporting autonomous, privacy-preserving personalization than the more complex LSTM architecture. Although the primary analysis focused on general chronic pelvic pain disorders, the model's performance remained consistent when tested on a healthy control cohort, demonstrating robustness across different populations. Relevance to endometriosis: listed as one indication for GnRH antagonists, though the paper's main focus is uterine fibroids.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Reducing sedentary behavior through personalized digital interventions holds particular promise for individuals with chronic pelvic pain disorders (CPPDs), who face unique barriers to physical activity. We present a self-contained, missing data-resilient framework for real-time forecasting of a physical activity score (PAS) using wearable Fitbit data from 134 females with CPPDs. Comparing online and offline learning approaches, we demonstrate that models leveraging recent activity and daily recurring patterns perform the best. When applied to 15-min sedentary bouts (SBs), the PAS forecasts support timely alerts, yielding approximately one true alert per day vs 0.6 false alerts at a conservative operating point. By integrating real-time imputation, the system supports forecasting of SBs for potential use in just-in-time adaptive interventions. Our work demonstrates a privacy-preserving, scalable pathway for integrating precision forecasting into just-in-time adaptive interventions, laying the groundwork for more equitable and effective digital health solutions for women with CPPDs.
Full text 46,941 characters · extracted from pmc-nxml · 5 sections · click to expand

Methods

This study is a secondary analysis of data from an ongoing observational study approved by the Institutional Review Board (IRB) of the Icahn School of Medicine at Mount Sinai (IRB# STUDY-22-01002). The parent grant study aims to develop mHealth-based digital patient-reported outcome measures grounded in CPPDs (NIH/NICHD: R01HD108263). The study design involves daily and weekly data collection for 90 days and wearable activity trackers (Fitbit Inspire 2). Participants provided informed consent and received onboarding instructions, including guidance on device usage and data syncing, described in detail elsewhere 7 . Inclusion criteria for the parent study were: (1) assigned female at birth and menstruating, (2) age between 18 and 64 years, (3) surgical- or clinician diagnosis of CPPD based on self-report, and (4) chronic pelvic pain for ≥6 months. Exclusion criteria included current, recent or planned pregnancy within 6 months and active major co-morbidities (e.g., cancer, heart failure, acute coronary syndrome). Given our focus on wearable-estimated PA herein, we include participants from that parent project who had available Fitbit data in this analysis. The CPPD cohort was partitioned into tuning and evaluation sets based on data completeness. For each participant, we measured data completeness as the number of valid observations as a proportion of the total observation time (see Supplementary Fig. S1 ). Participants with at least 90% completeness were assigned to the tuning cohort ( N  = 19). The remaining participants ( N  = 135) were considered for inclusion in the evaluation cohort. To produce stable performance estimates following the initial 240 h training period, we required participants to have at least another 240 h of valid data, totaling 480 h. Applying this criterion yielded an evaluation cohort of N  = 115 participants. Out of N  = 71 healthy controls, the N  = 61 participants who had at least 20 valid days constituted the control cohort. We provide a table with demographic information about all cohorts in the Supplementary Materials (Supplementary Table 1 ). We worked with Fitbit estimated time-matched activity, heart-rate data and sleep data at minute-level resolution to enable short-term forecasting. To distinguish true sedentary behavior from device non-wear time, any minute lacking heart rate data was marked as missing. All other minutes were considered “valid.” Each valid minute was encoded using a numerical activity score based on Fitbit’s activity levels: sedentary (0), light (1/3), moderate (2/3), and vigorous (1). We aggregated minute-level scores into larger bins by averaging valid scores within each bin. If more than 70% of minutes within a bin were missing, the entire bin was labeled as missing to prevent bias from low-quality estimates. Otherwise, the bin was considered valid. This resulted in the 1-dimensional PAS, which forms the basis of the work. To obtain a binary sleep status per time bin, the binary, minute-level sleep data were processed analogously to Fitbit’s activity levels and subsequently rounded to either 0 or 1. For model development, we used 60-min bins, which balance temporal resolution with noise reduction while preserving clinically relevant patterns. For the analysis of prevented sedentary episodes, we additionally employed 15-min bins to capture shorter sedentary periods that are more amenable to intervention. After computing the PAS at a given temporal resolution, for each participant, we counted the number of valid ( v ) and missing ( m ) time points, excluding leading and trailing periods of complete missingness before the first wear and after the last wear. Using the total time ( t  =  v  +  m ), we then computed a missingness ratio as the proportion of missing time points relative to total time ( m / t ). We define data completeness analogously as the ratio of valid data points: v / t  = 1− m / t . Data quality is assessed by the total amount of valid data ( v ), measured in hours, and data completeness. While data completeness and missingness ratio are mathematically redundant, we use them to emphasize high data quality or missingness, respectively. The temporal pattern of missingness varies considerably across participants, as illustrated in Supplementary Fig. S2 and visualized for all participants (see Supplementary Section S1.3 and Data 1 ). We evaluated a range of forecasting approaches spanning simple baselines, offline-trained models, and fully adaptive online models. Both offline and online models have a limited number of hyperparameters, optimized on a tuning cohort (see “Model training and hyperparameter selection”), and maintain internal states such as lagged values. The distinguishing feature of online models is that their model parameters are updated continuously and treated analogously to internal states. In contrast, offline models require a dedicated training period to estimate parameters, which are subsequently held fixed and cannot adapt to non-stationary activity patterns over time. All models perform inference sequentially, issuing one-step-ahead predictions using only information available up to the current time point. Given our focus on privacy-preserving, per-participant modeling, we restricted attention to models of low to moderate complexity to match the limited data volume available at the individual level. We hypothesized that the PAS at any moment might be predicted from two potential sources of temporal structure: recent activity patterns and daily routines. Recent activity could provide predictive information through various mechanisms, including but not limited to simple autocorrelation of human activity 64 , 65 . For instance, periods of moderate-to-vigorous activity might be followed by compensatory rest, while prolonged sedentary periods might precede transitions to movement. Daily routines, such as regular commute times or scheduled breaks, could provide an independent source of predictability tied to time of day rather than recent history. If the time series is rearranged as a 24-h × day matrix, these hypotheses can be visualized as distinct temporal receptive fields anchored at the current time point, spanning recent history and the same time-of-day on previous days (see Supplementary Fig. S3 ). To test these hypotheses without imposing strong assumptions about functional form, we examined multiple model classes: linear models, polynomial feature maps for nonlinear interactions, and RNNs for complex temporal patterns. Empirical evaluation with systematic hyperparameter selection (Supplementary Material S5 ) revealed which temporal receptive fields and model architectures best captured the predictive structure in individual activity patterns. Two simple baselines were used for reference. The persistence model predicts the next value as equal to the most recent observation, representing a no-change forecast sensitive to short-term autocorrelation. The mean predictor forecasts all future values as the participant-specific mean estimated during the training period, serving as a stationary reference model. We evaluated RNN architectures of low to moderate complexity, including standard RNNs, gated recurrent units (GRUs), and LSTM networks 66 , 67 . Across all architectures, a key hyperparameter is the temporal context length. Short context lengths primarily capture recent activity dynamics, whereas longer context lengths (e.g., spanning 24 h or more) are required to encode daily routines. To keep the computational footprint small during learning, we restricted the RNNs to maximally 2 layers and 128 hidden units. We also evaluated SARIMA models 49 , 50 . SARIMA models assume that the PAS can be expressed as a linear combination of recent past values and recurring seasonal patterns, after optional differencing to remove slow-moving trends and inclusion of past prediction errors as predictors. Model hyperparameters define the shape of the temporal receptive field, allowing SARIMA to simultaneously capture short-term dependencies and regular daily structure, i.e., 24-h seasonality. For online learning, we evaluated RLS 30 , 31 . Like SARIMA, RLS models the PAS as an autoregressive process over a temporal receptive field. However, unlike SARIMA, both the prediction coefficients and their associated uncertainty are updated continuously as new data arrive. This formulation treats parameters as adaptive quantities, enabling the model to track gradual behavioral changes. Recent observations are weighted more strongly than distant history, facilitating responsiveness to non-stationarity without retraining. Parameters are updated based on closed-form solutions, keeping the computational demand of learning to a minimum. In addition to linear inputs, we evaluated second-order polynomial transformations of the PAS to test the hypothesis that nonlinear interactions among recent activity values improve predictive performance. Wearable data are frequently affected by intermittent missingness due to device removal, charging, or synchronization gaps. While post-hoc analysis often relies on excluding days that fail to meet wear-time thresholds or iterative statistical methods, like multiple imputation 22 , 23 , these approaches are incompatible with real-time forecasting on data streams. Consequently, we adopted pragmatic strategies tailored to each model’s learning paradigm. For baseline and offline models (SARIMA, LSTM), missing values were imputed using the participant-specific mean estimated from the training period. For the online RLS model, we employed forward imputation by substituting missing observations with the model’s most recent predictions. This strategy is consistent with the online learning paradigm, which precludes reliance on pre-specified training data, and is particularly effective when data exhibit strong autocorrelation 68 . Forward imputation only affects the model’s input history (the lagged values used as features); imputed values were never used as training targets, and evaluation labels remained unobserved until after predictions were made (see “Evaluation strategy and outcome definitions”). Further implementation details for all models are provided in the Supplementary Material Section S2 . All models were trained and evaluated at the participant level, consistent with our goal to preserve privacy. For each participant, the first valid, i.e., not missing, 240 h of data were used as the training period during which model parameters were estimated. For the RLS model, the training period was used to initialize its internal states. Forecasting performance was evaluated on the remaining data, with a minimum requirement of 240 h per participant to ensure stable performance estimates. Participants not meeting this criterion were excluded from model evaluation. This design reflects a realistic deployment setting in which models are personalized using an initial calibration period and then applied prospectively. Hyperparameters were selected separately for each of the five model architectures: RNN, LSTM, GRU, SARIMA and RLS. A dedicated tuning cohort was used. For each family, candidate configurations were evaluated based on median RMSE across participants, and the best-performing configuration was selected. This single configuration was then fixed and applied uniformly to all participants in the test cohort to avoid per-participant overfitting. The resulting models were used for both performance comparisons and downstream clinical endpoint analyses. Further details on the hyperparameter search spaces and selected values are provided in the Supplementary Material Section S5 . For time series forecasting, the ground truth is the subsequently observed value in the same time series used as input to the model. Models predict the PAS at time t using only information available up to time t −1, and these predictions are evaluated against the actual observed PAS at time t once it becomes available. This sequential evaluation paradigm ensures that models are tested on their ability to forecast future activity patterns, not to reconstruct past observations. Sequential evaluation was used to evaluate the PAS forecast and adapted for SB prediction. Forecast accuracy was evaluated using the RMSE between forecast and observed PAS. RMSE was computed at the participant level over the test period following the training phase. For each model class, the best-performing hyperparameter configuration was identified on the tuning cohort and subsequently evaluated on the evaluation and control cohorts. To mitigate the potential effect of outliers and, consistent with developing patient-specific models, we chose the median and interquartile range to summarize group-level performance. To examine whether performance differences were consistent across individuals, participant-level RMSE differences were computed relative to the highest-performing model. To test whether wearable data missingness could confound forecasting performance, we plotted the relationship between the missingness ratio and forecasting performance RMSE. After fitting a linear regression, we used the explained variance ( R ²) to indicate the extent to which missing data contributed to forecast error. To test whether high PA could explain forecasting performance, we computed the median and 95th percentile of the PAS for each participant. We then repeated the previous analysis using these summary statistics as an independent variable ( x -axis). To assess the clinical utility of the best model, we evaluated its forecast against SB endpoints. One minus the PAS forecast for time t was interpreted as a continuous risk estimate for sedentary behavior. Ground-truth SBs at time t were defined by thresholding the observed PAS using two clinically plausible definitions: complete inactivity (PAS = 0) and near-zero activity (PAS ≤ 0.03), with the latter accounting for micro-movements or minor postural adjustments. Consistent with the consensus requirement that SB must occur during waking hours 33 , evaluation was restricted to wake-time periods. We utilized the Fitbit-derived main-sleep, a method recently validated against polysomnography for high-fidelity sleep-wake detection 69 . Performance was summarized using the median to reflect the accuracy for a typical participant. Selecting an operating point to translate continuous risk estimates into discrete alerts involves an inherent trade-off between precision and recall. We therefore considered two complementary operating-point selection strategies that reflect distinct deployment priorities. First, a conservative operating point was defined by fixing recall at approximately 20%, prioritizing a low false-alert burden consistent with practical intervention constraints. Owing to the discrete nature of precision–recall curves in finite datasets, the exact target recall may not be attainable; accordingly, we reported the maximum precision among operating points achieving recall within ±5% of the target value. Second, a balanced operating point was defined by maximizing the F1-score on a per-participant basis, reflecting a regime in which missed events and false alerts are weighted symmetrically. For both strategies, we computed the mean number of true and false alerts per participant per day, along with precision and, for the balanced strategy, recall. While model selection was conducted at hourly resolution, SB evaluation was performed for a forecasting horizon of 1 h but at two levels of temporal granularity: 1-h and 15-min. This change in temporal granularity preserved the underlying forecasting model and learned temporal structure while reflecting different clinical decision intervals. The study was approved by the IRB of the Icahn School of Medicine at Mount Sinai (IRB# STUDY-22-01002). All participants provided written informed consent prior to participation. The study was conducted in accordance with the Declaration of Helsinki.

Results

We evaluated forecasting models across two categories: fully adaptive online learning models and static offline learning models. The online approach was represented by RLS, while offline models comprised seasonal autoregressive integrated moving average (SARIMA) and long short-term memory (LSTM). The models were selected based on having a modest computational footprint during training, ensuring that the entire learning lifecycle is suitable for autonomous, on-device execution. After hyperparameter optimization on the tuning cohort (see Supplementary Section S5 ), defined by its high level of data completeness, the best performing models were selected for further evaluation. Performance was assessed on the evaluation cohort, which has more typical patterns of missingness. We used the root mean squared error (RMSE) of the 1-h-ahead forecast, reflecting the accuracy of PAS, which supports subsequent evaluation of clinical decision thresholds for preventing SBs. As depicted in Fig. 2A , all three models substantially outperform two simple baselines—a mean predictor and a persistence model—confirming the presence of informative temporal structure in the PAS time series. SARIMA achieved the lowest median RMSE (0.090, IQR: 0.073–0.112), followed closely by RLS (0.091, IQR: 0.073–0.110) and LSTM (0.095, IQR: 0.076–0.117). In contrast, the persistence baseline had a median RMSE of 0.107 (IQR: 0.087–0.138), and the mean predictor had a median RMSE of 0.110 (IQR: 0.089–0.136). The observed interquartile ranges reflect substantial inter-participant variability in baseline predictability, motivating a paired comparison to assess relative model performance within individuals. Fig. 2 Forecasting performance across model classes on the evaluation cohort. A Absolute RMSE for 1-h-ahead PAS forecasts ( n  = 115). Each dot represents the RMSE for a participant. The RMSE values are summarized across participants with the median and IQR, shown as solid bars. B Participant-level RMSE differences relative to RLS; positive values indicate higher error than RLS. Aggregation as before. A Absolute RMSE for 1-h-ahead PAS forecasts ( n  = 115). Each dot represents the RMSE for a participant. The RMSE values are summarized across participants with the median and IQR, shown as solid bars. B Participant-level RMSE differences relative to RLS; positive values indicate higher error than RLS. Aggregation as before. Figure 2B presents participant-level performance relative to RLS. SARIMA demonstrated nearly identical performance to RLS (median ΔRMSE = 0.0004, mean = 0.001), with marginal differences distributed evenly across participants. LSTM demonstrated slightly higher error (median ΔRMSE = 0.004, mean = 0.007). Both RLS and SARIMA captured informative short-term dynamics and daily structure in PA, indicating that most predictive signals arise from recent behavior and daily patterns (Supplementary Fig. S11 for RLS Sweeps and Fig. S13 for SARIMA Sweeps). Given comparable accuracy between RLS and SARIMA, and RLS’s fully online formulation enabling continuous parameter adaptation without retraining, we selected RLS for subsequent analyses. All results were consistent when evaluated in a healthy control cohort (Supplementary Section S3 ). Substantial inter-participant variability in forecasting error was observed (ranging from 0.04 to 0.16) (see Fig. 2 ), motivating an investigation into its underlying drivers. We examined whether participant-level missingness or PA characteristics explain differences in forecasting accuracy. The RMSE of the RLS model did not vary systematically with the proportion of missing data across participants (see Fig. 3A ). A linear fit yielded an R ² of 0.002, indicating that missingness explained a negligible fraction of the variance in forecasting error. Similarly, median wake-time activity levels computed across time showed minimal association with RMSE ( R ² = 0.165, Fig. 3B ). This negative result suggests that across participants’ differences in typical PAS values do not impact forecasting. In contrast, the 95th percentile of wake-time activity, representing high-intensity activity episodes, was strongly predictive of forecasting error ( R ² = 0.835, Fig. 3C ). This pattern was consistent across all evaluated models (Supplementary Fig. S5 on controls and other models). Fig. 3 Drivers of inter-participant forecasting error variability. Each point represents the mean RMSE for an individual participant. A Missingness ratio shows negligible correlation with forecasting error ( R ² = 0.002). B Wake-time median activity shows modest correlation ( R ² = 0.165). C Wake-time 95th percentile activity strongly predicts error variability ( R ² = 0.835). High-intensity episodes, not missingness, drive error variability. Each point represents the mean RMSE for an individual participant. A Missingness ratio shows negligible correlation with forecasting error ( R ² = 0.002). B Wake-time median activity shows modest correlation ( R ² = 0.165). C Wake-time 95th percentile activity strongly predicts error variability ( R ² = 0.835). High-intensity episodes, not missingness, drive error variability. These findings indicate that inter-participant variability in RMSE is driven primarily by the presence and magnitude of high PAS rather than by typical activity levels or data missingness. This is consistent with the sensitivity of RMSE to outliers, i.e., high PAS values, while the median for most participants is in the range of 0.05 and 0.1 (see Fig. 3B ). Importantly, as long as the model predictions are directionally correct, errors in predicting the amplitude of high PAS values are unlikely to compromise the forecast of sedentary periods. Consistent with this interpretation, the RMSE only weakly predicts the precision of SB predictions (see Supplementary Section S 4.5 ). Forecasting the PAS becomes clinically relevant when it can support timely interventions. Here, we assessed whether PAS forecasts could be used to prevent SBs. We treated the RLS forecast of one minus the PAS as a continuous risk estimate for sedentary behavior and compared it against the observed presence or absence of SBs. We evaluated the model’s ability to predict two distinct clinical manifestations of sedentary behavior. The first, complete inactivity (PAS = 0), targets the “inactivity physiology” paradigm, representing periods of chronic muscular unloading linked to suppressed skeletal muscle lipoprotein lipase activity 32 . The second, near-zero activity (PAS ≤ 0.03), aligns with the broader metabolic definition of SB: any waking behavior under 1.5 METs in a seated or reclined posture 33 . By testing both thresholds, we demonstrate that a single personalized forecast can be effectively mapped to multiple clinical endpoints, ranging from absolute physiological stillness to general sedentary time. Direct prediction of SBs over a 1-h horizon resulted in an unfavorable trade-off between true and false alerts (Supplementary Section S4.1 ), reflecting the challenge of forecasting at more distant horizons and, in the case of the zero definition of SBs, the low prevalence of only 6% of 1-h SBs during wake time. We therefore evaluated the same forecasts at finer temporal resolutions, corresponding to four consecutive 15-min intervals within the next hour. This represents a more granular clinical decision process while relying on the same model class and learned temporal dependencies. The prevalence of SBs increases to 0.25 and 0.36 at the 15-min interval, yielding a more informative prediction task (Supplementary Fig. S7F ). At a conservative operating point of 20% recall, Fig. 4A reports the median number of true daily alerts per participant across forecast horizons, while Fig. 4B shows the corresponding false alerts. For the shortest horizon, i.e., 15 min, the model generated 0.67 and 0.99 true alerts per day under the respective SB definitions, with 0.49 and 0.40 false alerts. The precision was 60% and 72% for the two thresholds, respectively. At increasing horizons, the false alerts accumulated, and the precision decreased steadily, an operating regime likely to induce alarm fatigue. Fig. 4 Clinical utility of SB forecasting at 15-min resolution. Lines show median performance across participants; shaded regions indicate interquartile range. A Daily true alerts (SBs prevented) across prediction horizons at 20% recall. B Daily false alerts across prediction horizons. C Precision at 20% recall. Results shown for two SB definitions: zero activity (blue) and near-zero activity (red). Lines show median performance across participants; shaded regions indicate interquartile range. A Daily true alerts (SBs prevented) across prediction horizons at 20% recall. B Daily false alerts across prediction horizons. C Precision at 20% recall. Results shown for two SB definitions: zero activity (blue) and near-zero activity (red). Another strategy for selecting an operating point is to maximize the F1-score per participant, which balances precision and recall. At the 15-min horizon and for the near-zero definition of SB, the recall increased to 82% but reduced precision to 53%, resulting in a median of 3.89 true and 3.53 false alerts per day (Supplementary Section S4.2 ). In summary, these results indicate that activity forecasting can support clinically meaningful intervention strategies when applied at short temporal horizons and conservative operating points. The appropriate operating regime depends on the relative costs of missed SBs versus false alerts: in wellness or behavioral coaching applications, minimizing alarm fatigue may be paramount, whereas in clinically supervised PA prescriptions, higher recall may be acceptable despite increased false alerts.

Discussion

This study aimed to develop a self-contained, participant-specific framework for forecasting the PAS in women with CPPDs; to ensure robustness to missing data and enable privacy-preserving local training; and to assess whether such forecasts can reliably detect SBs suitable for JITAI deployment. Based on our findings, we identify three key novel contributions: (1) we introduce the PAS and demonstrate that simple linear models achieve robust per-participant forecasting performance using only 10 days (240 h) of training data; (2) we show that interpretable autoregressive models uncover the predictive signal and that performance remains stable despite missing data; and (3) we identify actionable operating points where approximately one SB can be prevented per day with fewer than 0.5 false alerts (conservative regime) versus 3.89 bouts prevented with 3.53 false alerts (balanced regime). To our knowledge, this is the first study to translate forecasting accuracy to intervention metrics under realistic constraints using primary data collected from the target clinical population. A central question of this study was whether the PAS contained statistical signals for forecasting future values. We found that RLS, an online method, and SARIMA, an offline method, converged to similar performance based on short-term dependencies over recent hours combined with stable daily recurrence. This convergence reveals three key characteristics of the forecasting task. First, the activity time series appears stationary. RLS did not benefit from its capacity to adapt to non-stationary processes, indicating that patterns present in the training period (the first 240 h) persisted throughout the test period. Corroborating this, RLS performance increased with longer adaptability times, approximating static behavior. Second, missingness had minimal impact on forecasting performance. RMSE remained independent of the missingness ratio across all models, and different imputation strategies (forward imputation in RLS versus mean imputation in SARIMA) yielded equivalent performance. This suggests that the missingness pattern did not substantially disrupt the temporal patterns exploited by the models, likely because they often occur in continuous episodes rather than scattered gaps. Third, the most successful models were autoregressive and linear, suggesting that the behavioral time series contained consistent, predictable patterns based on temporal correlations rather than complex nonlinear dynamics. Nevertheless, RLS’s online adaptability may become important when incorporating additional modalities alongside the PAS, such as location, social context 34 or survey data 35 . The predictive signal would likely become non-stationary following life events such as employment changes or joining structured exercise programs. In such scenarios, online learning could mitigate the need for explicit retraining while preserving privacy through per-participant modeling, instead of the overhead of federated learning approaches put forward recently 36 , 37 . Another advantage is that online methods provide immediate predictions and subsequently improve their performance over time. In contrast, offline methods require a period in which training data is collected and subsequent model training. This immediate feedback can enhance user engagement 38 and reduce early dropout, which is particularly important given that sustained engagement with wearable devices remains challenging 39 . The finding that RLS and SARIMA, both simple linear models with interpretable temporal receptive fields, outperformed the recurrent neural network (RNN) architectures challenges the trend towards highly parametrized models, such as deep learning-based approaches 40 , 41 . Simple models excel when data is limited (here constrained by the 240-h training period) and when the underlying statistical structure is relatively straightforward, e.g., the combination of recent activity and daily recurrence. Unlike natural language or images, which exhibit rich hierarchies and long-range dependencies, wearable sensor data captures a single behavioral dimension with inherent stochasticity. In such settings, the signal-to-noise ratio may be too low to justify architectures designed for highly structured data. Furthermore, these models offer a more efficient learning lifecycle for mobile health applications because they rely on direct mathematical solutions rather than the iterative, gradient-based optimization required by deep learning. This allows for a private-by-design architecture without the need for cross-participant data sharing or the logistical complexity of federated learning. Moreover, our participant-specific approach sidesteps the challenge of cross-participant generalization, as the models are required only to adapt to the individual’s future behavior. SB forecasting presents both an opportunity and a methodological challenge. SBs are meaningful from an intervention standpoint but are variably defined in detail 27 . We circumvented this issue by introducing a continuous activity score at a high temporal resolution, which enables flexible downstream evaluation as demonstrated by restricting the analyses to wake periods and applying two plausible sedentary thresholds. While the PAS is exact at 1-min resolution, it becomes increasingly lossy at coarser temporal scales, raising the question of the appropriate temporal granularity. At 1-h resolution, high-intensity activity bursts were the primary source of error, confirming that they were, despite their rarity, well represented with the PAS. However, favoring a more granular timescale, we found that the precision of SBs forecasts declined steadily across successive 15-min intervals within an hour. This decline confirms the presence of temporal structures at the 15-min resolution. We evaluated two operating points for converting continuous risk predictions into binary alerts, reflecting different trade-offs between intervention coverage and notification burden. The conservative operating point (20% recall) yields approximately one targeted SB per day with fewer than one false alert, consistent with prior mobile health studies reporting acceptable engagement at 1–3 daily messages and alert fatigue at higher frequencies 42 , 43 . This regime may be appropriate for low-intensity wellness or behavioral coaching applications. In contrast, the balanced operating point captures most SBs at the cost of additional false alerts, which may be acceptable in higher-intensity or clinically supervised interventions where sedentary reduction is a prescribed goal and false alerts are less consequential. Importantly, alert acceptability is not determined by frequency alone: JITAI design principles emphasize the joint role of user burden, contextual relevance, and goal framing in shaping intervention effectiveness 44 . Prior work suggests that gamified or goal-oriented mobile interventions can increase engagement 45 , 46 , potentially allowing higher alert rates when alerts are perceived as opportunities to progress toward explicit targets or open goals 47 . Finally, individual differences and contextual constraints suggest that fixed operating points may be suboptimal in practice (e.g., reducing temporal accessibility and consequently engagement over time) 48 , motivating adaptive thresholding strategies central to the adaptive trial (e.g., JITAI) frameworks. Beyond revealing tradeoffs in alert frequency, our results inform when interventions should be delivered. Our model achieved its highest predictive precision for imminent SBs occurring within the next 15 min, with performance degrading at longer horizons (15–30 and 30–45 min). This pattern reflects an empirical constraint rather than a design choice and is consistent with information decay in temporal prediction, whereby the predictive signal diminishes as the forecast horizon extends 49 , 50 . To our knowledge, this is the first study to characterize this decay at a sub-hour timescale for sedentary behavior. The existence of a short horizon with sufficient predictive precision creates a practical opportunity for intervention: accurate near-term predictions can be paired with brief, actionable activity prompts delivered close to the point of action, in line with core JITAI design principles emphasizing timely, contextually relevant support 44 . This timing also aligns with growing evidence that short activity breaks, often referred to as “exercise snacks,” can meaningfully interrupt sedentary time and improve cardiometabolic outcomes 51 , 52 , with recent work characterizing dose–response relationships for brief bouts and the effectiveness of frequent interruptions throughout the day 53 , 54 . Recommending short activity breaks may be more achievable than longer exercise prescriptions, particularly for populations with high baseline sedentary time or barriers to structured exercise, including underserved groups and individuals with chronic pain conditions 55 – 58 . Together, these findings suggest that a 15-min prediction horizon represents a feasible and action-relevant timescale for linking sedentary behavior prediction to low-burden, just-in-time micro-interventions. Future work is needed to compare intervention effectiveness across different forecast horizons, evaluating whether proximal predictions optimized for accuracy or advance predictions optimized for day-planning better support sustained behavior change. A noteworthy finding was that while overall forecasting error (RMSE) was strongly associated with the 95th percentile of wake activity ( R ² = 0.835), a marker of high-intensity bouts, this variability had minimal impact on the accuracy of SB prediction. This dissociation suggests that sedentary behavior follows more stable, routine-driven temporal dynamics, likely structured by daily obligations such as work schedules, mealtimes, and commuting patterns 59 , compared to moderate and vigorous activity, which may arise more spontaneously from contextual factors such as social events, weather, and motivation fluctuations. Population-specific factors include the current pain intensity and pain-activity regulation: recent work found that pain variability (i.e., unpredictability) drives activity avoidance (and disability) more than pain intensity itself 60 . Despite this source of forecasting error, our model’s performance for predicting SBs remained robust. Combined with robustness to missingness and privacy-preserving per-participant training, this makes our approach suitable for diverse real-world chronic pain populations where the daily burden of self-management is high but adherence to rigid activity protocols varies widely 57 . Our study used consumer-grade Fitbit devices to measure PA, which has been reported to provide estimates of sedentary behaviors comparable to research-grade activPAL devices 61 , though it might underestimate MVPA in free-living conditions, based on prior research comparing it to research-grade Actigraph devices 62 . However, a more recent comparison between Actigraph and a selection of commercial devices during an emulation of free-living behavior concluded that both types “provided similar strengths and weaknesses and similar PA estimates” 63 . Thus, caution is warranted when interpreting the PA data, as it most likely resembles a conservative approximation of true PA. We highlight that any systematic measurement bias has a limited impact on our forecasting framework because models are both trained and evaluated on the same instrument, our predictions and metrics (e.g., alerts per day) capture valid relative patterns even if absolute PA values are biased. External validation against intervention-relevant thresholds established with research-grade devices remains for future work. Despite measurement limitations, consumer-grade devices offer advantages in large-scale behavioral research, including ease of wear, charging and syncing, all of which improve participant adherence. Our preprocessing pipeline focused on selecting SBs with high specificity. Our priority was to avoid misclassification of non-wear time as sedentary behavior. Because the presence of the heart rate signal excludes non-wear, we excluded all time points where heart rate was missing. Thus, our approach may have led to the removal of valid activity data. Future work should evaluate alternative preprocessing strategies to assess their impact on SB robustness specifically and PA forecasting more generally. Finally, our study sample consisted mostly of urban-dwelling, full-time employed individuals with a college degree or higher educational status. Moreover, the participants did not have any major co-morbidities (e.g., active cancers, CVD, kidney disease, unmanaged cognitive or neurological disorders). As such, generalizability to other socio-demographic cohorts remains to be further explored, particularly as underserved populations face distinct structural barriers to PA 56 . Nonetheless, forecasting performance was consistent in both CPPD participants and healthy controls, suggesting that our modeling approach captures activity patterns with some degree of structural stability across health states.

Introduction

Higher levels of physical activity (PA) and reduced sedentary time are independently associated with decreased mortality risk and improved cardiometabolic health outcomes among adults 1 , as well as other favorable health outcomes (e.g., mental health, physical function) 2 . Despite awareness of these benefits, insufficient PA and sedentary behavior remain highly prevalent in Western populations due to numerous biological, psychological, sociocultural, and environmental factors 3 . Individuals with certain chronic pain disorders face a paradox: they are at increased risk for sedentary behavior due to pain-related interference, even though significant gains in symptom relief and health benefits from increasing their PA for symptom management have been demonstrated 4 . One such population that merits particular attention comprises women with chronic pelvic pain disorders (CPPDs; e.g., endometriosis, adenomyosis, fibroids, etc.). With an estimated prevalence of one in seven persons, people living with CPPDs comprise a clinically significant yet understudied group, with many unmet patient needs due to the currently inadequate treatment options 5 , 6 . Promising evidence suggests PA can improve pain and mental health symptoms in this population, without exacerbating pain 7 – 9 . Consequently, there is a pressing need for targeted strategies to reduce sedentary behavior, as well as for research to determine whether specific intervention time points can enhance effectiveness. A framework that is both scalable to large user populations and enables individualization remains unexplored to date, motivating the rationale for undertaking the analysis herein. The proliferation of consumer-grade wearable devices and mobile applications increasingly enables continuous monitoring of PA patterns at scale, which provides the potential for greater understanding of PA behaviors during free-living conditions. Wearables offer objective estimation capabilities that overcome limitations inherent in self-reported data (e.g., recall bias, lack of data granularity), minimize participant burden, and facilitate real-time intervention delivery. Recent systematic reviews indicate that wearable-based interventions can be effective for increasing daily step counts; however, their impact on reducing sedentary behavior remains limited 10 . Moreover, digital interventions have demonstrated variable efficacy across the socioeconomic spectrum 11 . This highlights the necessity for more personalized approaches tailored to individual behavioral patterns and contexts. Existing digital interventions include email-reminders 12 , video-coaching 13 , online support groups 14 and instruction-sessions with follow-up calls 15 , but they primarily rely on retrospective summaries of PA rather than real-time insights necessary for timely intervention delivery. Given the rapid advancement in mHealth technologies, wearables combined with mobile applications offer unprecedented opportunities for addressing some of these gaps. These technologies can typically be leveraged in just-in-time adaptive interventions (JITAIs) 16 – 18 and N-of-1 personalized trials 19 . They involve delivering timely support based on real-time data input from the participant’s behavior and environmental factors. However, in the context of targeting sedentary behavior, a critical prerequisite is the accurate and timely forecasting of the individual’s future behavior. Contemporary frameworks emphasize that movement behaviors should be conceptualized as an integrated 24-h continuum rather than isolated constructs 20 , necessitating forecasting approaches that jointly model the spectrum of activity intensities at high temporal resolution. Toward this endeavor, we identify three practical requirements for PA forecasting models: (1) robust handling of missing data within high-resolution, real-time streams; (2) protection of participant privacy; and (3) a self-contained, learning lifecycle to enable autonomous on-device personalization. Previous work has addressed these challenges through various lenses. In the context of missing data, strategies have evolved from rigid heuristics, such as the 10-h “valid day” threshold 21 , 22 , to complex statistical imputation for longitudinal phenotypes 23 and population-specific reliability rubrics 24 . Despite these advances, there remains no established consensus for handling missingness in 24-h data streams where post-hoc imputation is not feasible. Regarding privacy and complexity, researchers have investigated federated learning frameworks 25 and missing data using a transformer-based modeling strategy 26 , coinciding with broader interest in domain-agnostic time-series transformers 27 . Nevertheless, increasing complexity introduces new risks to data privacy 28 . Moreover, in time-series forecasting, a surprisingly simple one-layer model with normalization can outperform many of the more complex transformer architectures 29 . This is particularly relevant for mobile health applications, where the efficiency of the update step (learning) is more critical than the capacity for large-scale inference. While modern mobile hardware can support inference for larger models, the iterative gradient descent required by transformers and other deep neural networks makes the learning lifecycle prohibitively expensive computationally. In contrast, the entire lifecycle can be supported by more shallow networks and classical online learning approaches, such as classical recursive least squares (RLS) 30 , 31 , which can be trained without excessive data, time, or dependency on external infrastructure. Despite these practical advantages of simpler models, no prior work has investigated the combined challenges of missing data and privacy preservation in PA forecasting using a computationally lean approach suitable for on-device learning. To address this gap, the goal of this study is threefold: (1) to develop a computationally lean, self-contained and participant-specific forecasting framework for predicting hourly PA in individuals with CPPD; (2) to demonstrate that models can handle missing data robustly; and (3) to evaluate whether these forecasts enable accurate detection of clinically meaningful SBs. Here, we present a time-resolved modeling framework that enables continuous forecasting of the physical activity score (PAS) from wearable data and subsequent detection of sedentary behavior through threshold-based classification. We evaluate performance sequentially: first assessing continuous forecasting accuracy of the PAS, then evaluating binary classification performance when these predictions are used to detect SBs (see Fig. 1 ). Fig. 1 Forecasting framework for sedentary behavior prediction. A Minute-level wearable data streams. B Computing the PAS from Fitbit’s intensity levels. C Offline versus online learning paradigms, emphasizing how newly observed data (orange) is used for training data (light blue) in the online learning paradigm. D Cohort partitioning for hyperparameter tuning, model comparison and clinical evaluation. PAS forecasting is evaluated via the RMSE. The clinical utility of the PAS forecast for predicting SBs is evaluated with prediction-recall metrics. A Minute-level wearable data streams. B Computing the PAS from Fitbit’s intensity levels. C Offline versus online learning paradigms, emphasizing how newly observed data (orange) is used for training data (light blue) in the online learning paradigm. D Cohort partitioning for hyperparameter tuning, model comparison and clinical evaluation. PAS forecasting is evaluated via the RMSE. The clinical utility of the PAS forecast for predicting SBs is evaluated with prediction-recall metrics.

Supplementary Material

Supplementary Material Supplementary Material

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

SciLite annotations

organisms 3
human human human

Source provenance

europepmc
last seen: 2026-10-11T09:27:45.537177+00:00
pubmed
last seen: 2026-10-08T21:57:15.771029+00:00
scilite
last seen: 2026-10-04T09:59:34.739275+00:00