Explaining the variation in the attained power of a stepped-wedge trial with unequal cluster sizes | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research article Explaining the variation in the attained power of a stepped-wedge trial with unequal cluster sizes Yongdong Ouyang, Mohammad Ehsanul Karim, Paul Gustafson, Thalia S. Field, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-15483/v2 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 24 Jun, 2020 Read the published version in BMC Medical Research Methodology → Version 2 posted 8 You are reading this latest preprint version Show more versions Abstract Background In a cross-sectional stepped-wedge trial with unequal cluster sizes, attained power in the trial depends on the realized allocation of the clusters. This attained power may differ from the expected power calculated using standard formulae by averaging the attained powers over all allocations the randomization algorithm can generate. We investigated the effect of design factors and allocation characteristics on attained power and developed models to predict attained power based on allocation characteristics. Method Based on data simulated and analyzed using linear mixed-effects models, we evaluated the distribution of attained powers under different scenarios with varying intraclass correlation coefficient (ICC) of the responses, coefficient of variation (CV) of the cluster sizes, number of cluster-size groups, distributions of group sizes, and number of clusters. We explored the relationship between attained power and two allocation characteristics: the individual-level correlation between treatment status and time period, and the absolute treatment group imbalance. When computational time was excessive due to a scenario having a large number of possible allocations, we developed regression models to predict attained power using the treatment-vs-time period correlation and absolute treatment group imbalance as predictors. Results The risk of attained power falling more than 5% below the expected or nominal power decreased as the ICC or number of clusters increased and as the CV decreased. Attained power was strongly affected by the treatment-vs-time period correlation. The absolute treatment group imbalance had much less impact on attained power. The attained power for any allocation was predicted accurately using a logistic regression model with the treatment-vs-time period correlation and the absolute treatment group imbalance as predictors. Conclusion In a stepped-wedge trial with unequal cluster sizes, the risk that randomization yields an allocation with inadequate attained power depends on the ICC, the CV of the cluster sizes, and number of clusters. To reduce the computational burden of simulating attained power for allocations, the attained power can be predicted via regression modeling. Trial designers can reduce the risk of low attained power by restricting the randomization algorithm to avoid allocations with large treatment-vs-time period correlations. Health Economics & Outcomes Research Biostatistics cluster randomized trial power distribution treatment-time period correlation treatment group imbalance Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Background Use of the stepped-wedge cluster randomized controlled trial (SW-CRT) design has increased dramatically in recent years. 1–3 I In a standard stepped-wedge design (Figure 1), every cluster () begins (time-point ) with delivery of the existing standard-of-care (shown as white cells) to participants. Clusters then transition to delivery of a new intervention (shown as gray cells) at randomly determined time-points (), until all clusters are delivering the new intervention. The SW-CRT design often is used when there is a desire or need to implement and evaluate the intervention at the population level 4 , or when it would not be logistically feasible to implement the intervention in every cluster at the same time 3 , or when recruitment of clusters could be enhanced by ensuring all clusters eventually receive the new intervention. It has been implemented in trials exploring both single interventions as well as pathways of care in multiple settings. 5–8 Copas et al. 9 described three main SW-CRT designs: the “closed cohort”, the “open cohort”, and the “continuous recruitment short exposure” designs. In the first two designs, the treatment received by a participant during follow-up matches the treatment being delivered by his/her cluster at each time-point and each participant contributes multiple outcome measurements over time. In the third design, participants receive the treatment being delivered by his/her cluster at the time of study entry, remains on this treatment throughout follow-up, and contributes a single outcome measurement. For simplicity, we will refer to the continuous recruitment short exposure design as a “cross-sectional” design, a term which although less precise, is more commonly used. 3,10 The results presented throughout this paper were derived using a cross-sectional SW-CRT. Methods for calculating power and sample size in SW-CRTs with equal cluster sizes have been discussed widely in the literature. 2,11–14 Code for calculating the power of a SW-CRT with equal cluster sizes are also available for commonly used statistical software (R 15,16 , Stata 17 and SAS 18 ). Although it has been recognized that unequal cluster sizes in a parallel cluster randomized trial (P-CRT) leads to loss of power and efficiency 19,20 , limited work has been conducted on power and sample size calculations in unequal cluster size SW-CRTs. Hussey and Hughes 2 described a procedure for calculating approximate power based on a Wald test which allows for unequal cluster sizes in the SW-CRT, but they did not explore the impact of unequal cluster sizes. Kristunas et al. 21 conducted simulations that showed that the power of a SW-CRT with unequal cluster sizes was only slightly lower than if the cluster sizes had been equal. Girling 22 extended the method for evaluating the precision of regular single-period P-RCT with unequal cluster sizes to the stepped wedge design. Harrison et al. 23 derived first-order approximation formulae for the expected treatment effect variance given either known cluster sizes or the mean and coefficient of variation of the cluster sizes. They then used this treatment effect variance formula in a Wald test to conduct sample size calculations as well as to investigate the relative efficiency comparing equal and unequal cluster size SW-CRT designs and the design effect relative to individually randomized designs. However, it is important to recognize that these results focused on the expected power , that is, the power obtained through averaging over both 1) the randomness in the allocation of the clusters and 2) the random outcome variation across participants. Distinguishing between the two components is important. The latter source of variation is an intrinsic attribute of the population and outside the control of the investigator. However, the trial investigator has control over the former source through the choice of the randomization algorithm, and this choice impacts on trial power. Both Wong et al. 24 and Martin et al. 25 showed that the attained power (Wong et al.), also called the actual power (Martin et al.), defined as the power associated with a specific allocation, can vary substantially across realized allocations and can be much lower than the expected power. Through deriving an analytic approximation, Matthews 26 showed how the variance of the treatment effect estimate varies across allocations. Thus, a trial that was expected to achieve a specified power prior to allocation of the clusters might turn out to be underpowered due to an “unlucky” allocation. For example, we are involved in a SW-CRT 27 with 20 hospitals (clusters) in which the two largest hospitals were expected to enrol more than 200 participants each, while the four smallest hospitals were expected to enrol less than 10 participants each, with a cluster size coefficient of variation of 1.17. Power calculations obtained via simulation showed that although using an unrestricted randomization algorithm would yield an expected power of 80%, the attained power varied from a low of 68% to a high of 83% depending on the allocation. Martin et al. 25 observed larger variation in attained powers in designs with a smaller number of clusters, and a smaller intraclass correlation coefficient (ICC) . In addition, they observed that the absolute difference in the numbers of participants allocated to the two treatment groups, referred to hereafter as the “treatment group imbalance (TGI)”, explained only a small portion of the variability in attained power across allocations in the SW-CRT and that contrary to intuition, the attained power was relatively higher for allocations with a large TGI. In limited scenarios, Wong et al. 24 found that the (Pearson) correlation between the participant-level treatment assignment and the time-period with clustering ignored, referred to hereafter as the “treatment-vs-time period correlation (TTC)”, had a large impact on the attained power. They observed that the allocations with the greatest attained power were those in which the large clusters transitioned during the first or last steps (equally split) and that these allocations had the lowest TTC. In the SW-CRT, it is important to adjust for the time period to avoid bias in the treatment effect estimate because treatment group and time period inherently are highly correlated. They proposed that as TTC increases, this adjustment increases the standard error of the estimated treatment effect, which leads to a greater loss in attained power. However, we are not aware of a comprehensive study on these topics. In this study, we investigated how different allocation characteristics interacted with design factors to affect attained power and the risk of obtaining low attained power in cross-sectional SW-CRTs. This risk can be assessed by constructing the pre-randomization power distribution (PD) , 24 defined as the distribution of attained powers obtained from all possible allocations that a randomization algorithm can generate. A good randomization algorithm will ensure that the risk of obtaining a low attained-power allocation is acceptably small. Identifying such an algorithm requires an understanding of what factors cause low attained power. The first aim of this work was to gain an understanding of the factors that affect the risk of low attained power. We used simulation to evaluate the attained power across different allocations under a wide variety of scenarios. While it is possible to assess the attained power using approximate analytic formulae (e.g., Hussey & Hughes 2 ), the accuracy of those formulae has not been well investigated in the context of SW-CRTs, especially when the number of clusters or the cluster sizes are low. However, the computational time needed to simulate the attained powers for all possible allocations often is not feasible. Hence, our second aim was to develop regression models that could accurately predict the attained power for any allocation based on allocation characteristics (specifically TGI and TTC) and the attained powers from a sample of allocations. The results of this work provide guidance on how to assess attained power and avoid having an unacceptably low attained power when designing a SW-CRT with unequal cluster sizes. Methods Aim One: Evaluating attained power and the impact of different factors on the risk of low attained power In this study, we assumed that individual outcomes, , were generated from the following linear mixed-effects model: [Please see the supplementary files section to view the equations.] (1) where indexed the cluster, indexed the time period (denotes baseline), and indexed the individual within cluster and time period . The error terms were assumed to be independently sampled from a standard normal distribution (mean zero and variance ). treatment group indicator (0 = standard-of-care, 1 = intervention) was denoted as and the treatment effect, , was set to 0.26. As , the standardized effect size, , also equaled 0.26. Because the variance of the treatment effect estimate does not depend on the linear time effect coefficient, , its value was set to 1 without loss of generality. The cluster-specific effects, , were assumed to be independently sampled from a normal distribution with mean zero and variance . The value of was calculated from set values of the ICC, where ICC was defined as . The treatment-vs-time period correlation (TTC) was defined as the Pearson correlation coefficient between treatment and the time-period, , across all individual observations with clustering ignored. The treatment group imbalance (TGI) was defined as the difference in the numbers of participants who received the new intervention versus the standard-of-care. Computational challenge and a practical solution: Our aim was to obtain the set of attained powers by simulation and their corresponding power distribution for a specified randomization algorithm. First, we needed to obtain a list of all possible allocations that the randomization algorithm can generate, and then to evaluate the attained power associated with each possible allocation. However, the evaluation of attained power for all potential allocations can be computationally challenging, given the potentially huge number of possible allocations. For example, a twenty-cluster cross-sectional SW-CRT with unique cluster sizes and four clusters transitioning at each of five steps has more than 300 billion unique allocations. One strategy to reduce the number of unique allocations was by categorizing the clusters into size groups (e.g., (S)mall/(L)arge or (S)mall/(M)edium/(L)arge) and treating the clusters within each size group as identical, thereby reducing the number of unique allocations while increasing their multiplicities. Doing so will lead to only an approximate solution. However, this is often adequate in practice, given that the anticipated number of individuals that a cluster will enroll often can be only approximated at the start of a trial. In practice, we do not recommend having more than four size categories, as then the number of unique allocations likely will be too large to evaluate. Simulation specifications: As shown in Table 1, we investigated the impact of the total number of clusters (12, 24 or 48), the number of cluster size groups (2 – S/L or 3 – S/M/L), the distribution of clusters across cluster size groups (equal or unequal), the coefficient of variation (CV) of the cluster sizes (0.4, 0.7, 1.0, and 1.3), defined as the ratio of standard deviation of cluster sizes to the mean, and the ICC of the response variable (0.01, 0.05, and 0.10). A full factorial layout was investigated for all of these factors except the CV. An equal distribution of clusters across cluster size groups corresponded to an equal number of clusters in each cluster size group; an unequal distribution corresponded to distribution of clusters in a S:L ratio of 3:1 for scenarios with two cluster size groups or 3:2:1 to S:M:L for scenarios with three cluster size groups. For unequal distributions, CVs of 0.4. 0,7, 1.0, and 1.3 were investigated for all unequal distribution scenarios, but for equal distribution scenarios, CVs of only 0.4 and 0.7 were investigated as a CV larger than one cannot be obtained with equal distributions. (Thus, when the actual cluster sizes have a CV larger than one, trial designers need to use unequal distributions to apply the approach presented here.) Thus, there were a total of 108 scenarios. Table 1 Factors and their specific values explored in the simulation study. Factor Values Number of clusters 12, 24, 48 Number of cluster size groups 2, 3 Distribution of clusters to cluster size groups Equal (6S+6L, 12S+12L, 24S+24L, 4S+4M+4L, 8S+8M+8L, 16S+16M+16L), Unequal (9S+3L, 18S+6L, 36S+12L, 6S+4M+2L, 12S+8M+4L, 24S+16M+8L) ICC 0.01, 0.05, 0.1 CV 0.4, 0.7, 1.0*, 1.3* * For equal distribution of clusters to cluster size groups, only CVs of 0.4 and 0.7 were investigated as a CV of 1 or larger is not possible. Trial parameter values for all 108 scenarios are listed in the Additional file (Table A1). The number of transition steps was fixed at four. To determine the total number of individuals needed to obtain approximately 80% power in each scenario, we used Hussey and Hughes’s 2 sample size formula for a SW-CRT with equal cluster sizes as implemented by Baio et al. 11 in the R package “SWSamp 14 ” (implementation of the Hussey and Hughes’s formula). Then, we re-distributed the total number of individuals to create unequal sized clusters that matched the CVs we set under each scenario (see Additional file, A2 for details). 28 If a resulting cluster size was not an integer, we set it at random to one of the two bracketing integers such that the expected value matched the initially calculated cluster size. For example, if the calculated cluster size was 12.3, the cluster size was set at random to either 12 or 13, such that its expected value was 12.3. Within a cluster, individuals were allocated to the time periods using a multinomial distribution with equal probability for each period. The total number of unique allocations is presented in the third column from the right in Table A.1. Among the 108 scenarios, 72 scenarios (referred to as ‘completed scenarios’) had less than 2,000 unique allocations. For these scenarios, the attained power for every allocation was evaluated via simulation. For the other 36 scenarios (referred to as ‘sampled scenarios’), we evaluated the attained power for only a sample of 2,000 allocations. Subsequently, the attained power for two sampled scenarios, one with a relatively small number of allocations (scenario #64, with 8,623 allocations) and one with a relatively large number of allocations (scenario #103, with 113,949 allocations), had all of their allocations evaluated to enable validation of the prediction models that were constructed using only the original 2,000 allocations. Performance metrics: For each selected allocation, 10,000 datasets were simulated and analyzed using the model (1). Analysis was performed using the lme function in R. 29 This P-value is calculated through comparing the Wald statistic to quantiles from a t -distribution with degrees of freedom equivalent to what would be used in a balanced, multilevel ANOVA designs 30 . The attained power was estimated using the proportion of P-values less than 0.05. Assuming a power near 80%, the simulation standard error of each attained power was approximately 0.4%. The PD associated with the unrestricted randomization algorithm was then constructed from the estimated attained powers by weighting each attained power by the probability of the corresponding allocation. Two measures of the risk of low attained power were then obtained from the PD: (1) the probability that the attained power falls more than 5% below the expected power (obtained in the simulations), and (2) the probability that the attained power falls more than 5% below the nominal 80% power (i.e., less than 75%) that would be achieved in a trial with equal cluster sizes. The first measure provides an indication of whether potential low attained power is a concern when power is calculated using the approach presented here for handling unequal cluster sizes. The second measure would be of greater interest when one is assessing whether power calculations obtained under the assumption of equal cluster sizes are adequate. We set a 5% power loss as being meaningful, but trial designers may wish to choose a different value more relevant to their context. In this paper, the phrase “risk of low attained power” will be used as a short form to refer to either of these measures. All computations were conducted using R version 3.5.0 on Cedar, Compute Canada 31,32 . Aim Two: Explaining the variation in attained power across allocations and predicting attained power using allocation characteristics Using the simulation results from Aim One, we first constructed scatterplots to examine the bivariate relationship between the simulated attained power and each of TTC and TGI separately. Then, to predict the attained power for each scenario, we fitted logistic regression models of the proportion of simulations in each allocation that achieved statistical significance (at level 0.05) as a function of linear and quadratic terms for TTC and TGI. Based on the relationships observed in the bivariate scatterplots, we chose to investigate a sequence of four nested models. The terms included in the four models that were considered were: (Model 1: TTC), (Model 2: TTC, TGI), (Model 3: TTC, TGI, TTC 2 ), and (Model 4: TTC, TGI, TTC 2 , TGI 2 ). We measured the predictive accuracy of each model by using five-fold cross-validation, conducted using cv.glm 33 in R, to estimate the root mean squared prediction error (RMSPE), defined by , where represented the set of allocations in the -th partition, was the number of allocations in the -th partition, and were the simulated and predicted powers, respectively, for the -th allocation, and was the number of unique allocations in the scenario. From these models, we selected the one for which no meaningful improvement in RMSPE was achieved by adding another term to the model. For the selected model, we examined additional predictive performance measures for each scenario, including maximum and average absolute prediction errors (). For two sampled scenarios (#64, #103), we validated the selected model with respect to both prediction of attained power for individual allocations, and estimation of the risk of low attained power. For the former objective, we compared the predicted attained power to the simulated attained power for each allocation that was not in the set used to fit the predictive model. For the latter objective, we repeatedly sampled (10,000 times) 2000 allocations and estimated the risk of low attained power from the fitted model. We then compared these estimated risks to the “true” risk derived from the simulated attained powers. Results Aim One: Evaluating the attained powers and factors affecting the risk of low attained power In the body of this paper, we report on the scenarios with an unequal distribution of clusters to cluster size groups. The results for scenarios with an equal distribution of clusters to cluster size groups were similar when matched on the remaining factors (see Additional file, Figure A3.1 and A3.2). Substantial variation in attained powers was observed in all the scenarios. Across all simulations, attained power ranged from 0.62 to 0.86. As examples, the PDs for selected scenarios (#85 to #87) are displayed in Figure 2. The expected powers obtained from the simulations (Additional file, Table A1, the “Weighted expected power”) usually were quite close to the nominal 80% power associated with an equal cluster size design, but losses of up to roughly 4% were seen when the CV was large. ICC Figure 3 shows the risk of low attained power (for both measures) as a function of ICC for different combinations of the total number of clusters, the CV, and the distribution of clusters to the size groups. The risk of low attained power varied greatly with ICC. For both risk measures, with a total of 24 or 48 clusters, the risk of obtaining low power was smaller in the scenarios with larger ICCs. If nominal power is used as the reference, the risk was as high as 38% when the ICC was 0.01 but dropped to nearly zero even when the ICC was 0.1. When there were 12 clusters in total, the risk of low attained power often continued to be high even with larger ICCs, and the patterns were not as evident. However, it should be noted that in these scenarios, the number of unique allocations is quite low so the combined effect of simulation error and discretization effects may have disrupted the monotonic patterns that were expected. Coefficient of variation Figure 4 plots the risks of low attained power as a function of the CV. The risk of low attained power varied greatly across the CVs and depended on the ICC. Except for the scenarios with 12 clusters and two size groups, the risk was near zero when the CV was 0.4. However, as the CV increased above 0.7, the risk increased rapidly up to more than 30% when the CV was 1.3 and the ICC was 0.01. Aim Two: Explaining the variation in attained power across allocations and predicting attained power using allocation characteristics Relationship among transition time-points of large clusters, treatment-vs-time period correlation (TTC) and absolute treatment group imbalance (TGI) Understanding the relationship between the time-points when the large clusters transition, TTC, and TGI will be helpful for interpreting the predictive models presented later. Figure 5 summarizes the relationship typically seen between the transition time-points of the large clusters with TTC and TGI across simulation runs from one representative scenario (#37). Plotted symbols represent different allocations. The number used as the plotting symbol indicates the number of large clusters transitioning at the two ends (i.e., at either the first or last step). Setting aside the values of TGI, this plot shows that the number of large clusters transitioning at the ends strongly determines TTC, with TTC decreasing as this number increases. When the transition time-points of the large clusters are equally split between the two ends, both TTC and TGI are low (numbers plotted in red), but when these transitioning times are concentrated at one end, TTC remains low while TGI is high (numbers plotted in blue). When all of the large clusters transition at the middle steps (numbers plotted in green), TGI will be low and TTC is high. Impact of treatment-time period correlation on attained power FFigure 6 displays the relationship between the attained power for a given allocation and TTC for a selected set of 16 scenarios (two size groups , 12 or 48 clusters, ICC = 0.01 or 0.1). Within this grid of panels, columns correspond to different CV values (0.4, 0.7, 1.0, 1.3 from left to right) and rows correspond to the number of clusters (12 in the first row; 48 in the second row). Within each panel, points in blue correspond to an ICC of 0.1, while those in red correspond to an ICC of 0.01. A graph for all scenarios is included in the Additional file, Figure A4. Across all scenarios, TTC consistently exhibited a strong negative, and mildly non-linear, relationship with the attained power. Greater variation in attained power was observed in the scenarios with a larger CV, a lower ICC, and a smaller number of clusters. Impact of treatment group imbalance on attained power Figure 7 displays the relationship between attained power and TGI, using the same layout and for the same scenarios used in Figure 6. A graph for all scenarios is included in the Additional file, Figure A5. Across all scenarios, the association exhibited a “triangle” pattern with much weaker associations than were seen for TTC. In 90% (97/108) of the scenarios, univariate regression model fits showed that allocations with a large TGI were associated with a relatively higher attained power. This is counter-intuitive and opposite to the P-CRT setting, where attained power decreases with increasing TGI. When TGI was low, the attained power ranged over both low and high values across different allocations. This unexpected result is resolved in the following section. Essentially, the complicated relationship between TTC and TGI as described earlier leads to the true impact of TGI on attained power being confounded by its association with TTC, with the latter being a much stronger determinant of attained power. Predictive model Without any predictors (i.e., the null model with only an intercept term), the RMSPE of the predicted attained power ranged across the 108 scenarios from 0.0060 to 0.0454 with a median of 0.0157. When TTC alone (Model 1) was added to the model, the RMSPE decreased in every scenario with the median RMSPE shrinking down to 0.0037 (range: 0.0022, 0.0103). When TGI was then added (Model 2), the RMSPE decreased in 93 of the scenarios and the median RMSPE dropping down to 0.0035 (range 0.0020 to 0.0085). Note that in comparison to the crude effect of TGI, after adjustment for TTC, the percentage of scenarios with a positive TGI coefficient declined from 90% to 68% (74/108) and the distribution of TGI coefficients was compressed much closer to zero compared to the distribution of TGI coefficients from the unadjusted models (see Additional file, Figure A6). On adding the TTC 2 term (Model 3), the RMSPE decreased in 99 scenarios, with the median RMSPE shrinking to 0.0031 (range 0.0017, 0.0065). Further adding TGI 2 (Model 4), resulted in the RMSPE decreased in 71 scenarios but the median RMSPE changed negligibly (i.e., on the order of ). Hence, we selected the Model 3 as the final model. Figure 8 displays a contour plot of the predicted attained power derived from the regression model for one scenario (#79). The almost-vertical curves indicate that attained power decreased substantially as TTC increased. However, when TTC was fixed, the attained power decreased only slightly as TGI increased. RMSPE, maximum absolute prediction error, and average absolute prediction error for each scenario are shown in Table A7 in the Additional file. Validation of prediction model The medians of the differences between the predicted and the simulated attained powers for the allocations not used in model fitting for scenarios #64 and #103 were near zero and 95% of the predicted attained powers fell within 0.0069 (#64) and 0.0094 (#103) of the simulated attained powers (See Additional file, A8 for the distribution of differences). This magnitude of prediction error is comparable to the expected error in the simulated attained powers (i.e., 95% of simulated attained power within 2 x 0.004 = 0.008), suggesting the selected model performs nearly as well as simulation. The true risks of low attained power (> 5% below nominal 80%), calculated from the simulated attained powers, were 0.0201 (#64) and 0.0697 (#103). In repeated sampling and fitting of the selected model, the median absolute difference between the model-based estimate and the true risk of low attained power was 0.0010 (#64) and 0.0019 (#103), with 95% of the values falling within 0.0018 (#64) and 0.0034 (#103). Discussion Understanding the risk of obtaining low power Using simulation, we have examined how and why the attained power in an SW-CRT with unequal cluster sizes may be different from the expected power that typically is the focus of power calculations. We have argued that even if the expected power is adequate, the risk that the randomization leads to a trial with low attained power should be considered. The extension of the CONSORT 2010 statement 34 recommends that researchers report the unequal cluster sizes in trials. This study highlights the importance of knowing the cluster sizes to enable appropriate assessments of power. The power distribution constructed from the set of attained powers can be used to evaluate the risk that a given randomization algorithm will yield a trial with unacceptably low power. As shown in our study, the CV of the cluster sizes strongly impacted the risk of low attained power. For CV values below 0.7, the risk of the attained power falling more than 5% below the expected power across the investigated ranges of values for the other parameters was sufficiently low that this risk may not be a concern. As the CV increased above 0.7, the risk of low attained power increased substantially, especially when the ICC was low. As the ICC increased, the risk of low attained power decreased. This result is consistent with Matthews et al. 35 , who showed that in a “row-column” analysis of a SW-CRT, which corresponds to a linear mixed effects analysis, the variance of the treatment effect decreased as the ICC increased. Our heuristic explanation is that when the ICC was large, the effective sample sizes of large clusters were reduced greatly, which in turn reduced the effective CV of the cluster sizes. As expected, the risk of low attained power decreased as the number of clusters increased. However, even with 48 clusters, this risk was not ignorable when the CV was large, and the ICC was small. Explaining the variation in attained power A key contribution of this study is clarifying how attained power is impacted by the joint effects of treatment-vs-time period correlation and absolute treatment group imbalance. TTC was shown to be the dominating factor. Our results suggest that the observation that TGI on its own is only weakly related to attained power and in the direction opposite to what is expected, made by Martin et al 20 and confirmed here, is a consequence of confounding with TTC. After adjusting for TTC, the impact of TGI on the attained power was reduced or the direction reversed in nearly every scenario. That the direction was not reversed for every scenario may be due in part to simulation noise, but it could also reflect residual confounding due to the model mis-specifying the true dependence of attained power on TTC, or the omission of other important predictors. We have shown that a model containing TTC, TTC 2 , and TGI can accurately predicts attained power for any allocation, and hence to construct the power distribution for any randomization algorithm and to estimate the risk of low attained power, when the number of allocations is too large to evaluate their attained powers using simulation. The original rationale for creating size groups was to reduce the number of unique allocations that needed to be evaluated to obtain the power distribution. However, given the high accuracy seen in these predictive models, creating cluster size groups may not be necessary. Instead, a predictive model could be built based on a representative sample of allocations without grouping of clusters and this model could be used to predict the attained powers associated with each non-sampled allocation. The caveats here are that some effort may be needed to determine how large a sample is needed to achieve a target predictive accuracy, and when the number of unique allocations is large, simply listing out all possible unique allocations may require excessive (computational) effort in which case size groups will still be needed. Examination of attained power can assist in the identification of the allocation characteristics that lead to low attained power. For example, our study verified the results from previous work 24 showing that transitioning large clusters during the middle steps can increase the treatment-vs-time period correlation and lead to low attained power. This result is also consistent with Kasza and Forbes 36 , who showed that for an equal-cluster-size design, the sequences (i.e., allocations) that transition during the middle steps contribute less information than those that transition at the end steps. Therefore, allocating large clusters to transition during the middle steps should (as a heuristic) yield less information than if they were allocated to transition at the end steps. Identifying these allocations affords trialists the opportunity to take pre-emptive actions to mitigate this problem. The simplest option would be to restrict the randomization algorithm to exclude low power allocations. However, this approach risks violating the criteria needed for valid randomization inference, as it could result in some clusters having no chance of being randomized to transition at particular time-points (e.g., large clusters may not be allowed to transition during the middle step(s)). A more conservative option could be to stratify randomization based on cluster size. This would ensure that every large cluster has a chance to be allocated to any transition step, while ensuring even distribution of large clusters across all of the steps. This would tend to yield a low-variability power distribution, as the excluded allocations would tend to be both those with low and high power. Further exploration of potential restricted randomization algorithms and their validity, benefits and limitations is needed. Limitations The Wald test we used to evaluate statistical significance is known to yield higher-than-nominal Type I error rates 37 . Therefore, the powers that are reported here will tend to be higher than what would be obtained using a test that achieved the nominal size. We have no reason to expect that the patterns of findings that we have reported would be different if the test used achieved the nominal size. However, important future work would be to ascertain which test(s) provide the most accurate size and power values to ensure that real-life trial designs fulfill the desired statistical criteria. Some work towards this goal has been conducted by Tanner 38 , using various analytic models (GLMM, GEE, cluster-level analysis) for binary outcomes. Our study left several design factors as fixed values, including the number of transition steps at four, and a standardized effect size of 0.26. Others 25 have reported that the trial power is only weakly affected by the number of steps. Results from limited simulations we conducted with six steps were concordant with those reports (data not shown), so we did not pursue a comprehensive investigation in this regard. The standardized effect size of 0.26 arose from a calculation performed to achieve 80% power in one particular scenario. We adopted this value as being within the typical range seen in real trials 39 . Assuming a larger effect size would have led to smaller cluster sizes needed to achieve 80% power. Again, we have no reason to expect different patterns to our findings, but smaller sample sizes may lead to greater Type I error and inaccuracy in the estimated power. The numbers of clusters (12/24/48) were chosen primarily to reflect our belief that these are typical numbers seen in real SW-CRTs (but also for convenience as these choices allowed the numbers of clusters to be balanced across steps and cluster size groups). While we found that increasing the number of clusters reduces the risk of low attained power, we did not establish a bound above which this risk is ignorable. When the CV is large and the ICC is small, this risk will need to be evaluated even when a trial includes a larger number of clusters. Due to limits on available computational time, we validated the predictive model for only two scenarios. The selected predictive model was shown to predict the attained power with accuracy similar to that achieved using simulation with 10,000 replicates. However, because this model may be mis-specified, there is no guarantee that this model will yield sufficiently accurate predictions if greater precision is required. Alternative models may need to be developed, perhaps including other allocation characteristics. In addition, simulating attained power for 2000 allocations required substantial computational time. An attractive goal would be developing accurate analytic formulae, like the one derived by Matthews 26 for the treatment effect variance, to support faster calculation of attained power. This work considered only a continuous outcome within a cross-sectional SW-CRT. Further work is needed to extend the results to other types of outcomes (binary, count, survival, etc.) and to other SW-CRT designs. Conclusion In a stepped-wedge trial with unequal cluster sizes, the risk that randomization yields an allocation with inadequate attained power is a function of the ICC, the CV of the cluster sizes, and the total number of clusters. To reduce the computational burden of simulating attained power for every allocation, the attained power can be predicted from the treatment-vs-time period correlation and the treatment group imbalance via regression modeling. Trial designers can reduce the risk of low attained power by restricting the randomization algorithm to reduce the chance of obtaining an allocation which yields a large treatment-vs-time period correlation. List of Abbreviations CRT: cluster randomized trial CV: coefficient of variation ICC: intraclass correlation coefficient P-CRT: parallel-arm cluster randomized trial PD: (pre-randomization) power distribution RMSPE: root mean square prediction error SW-CRT: stepped-wedge cluster randomized trial TGI: treatment group imbalance TTC: treatment-vs-time period correlation Declarations Ethics approval and consent to participate Not applicable Consent for publication Not applicable Availability of data and materials Not applicable. Simulation codes are available from: https://github.com/douyangyd/swcrtpd Competing interests The authors declare that they have no conflict of interest with this manuscript. Funding This work was supported in part through funding from the BC SUPPORT Unit. Author Contributions HW conceived/developed the ideas of this paper and contributed to the critical revision of the manuscript. YO conducted all the computational work, drafted the manuscripts, developed the ideas and contributed to the critical revision of the manuscript. MEK and PG developed the ideas and contributed to the critical revision of the manuscript. TF contributed to the critical revision of the manuscript. Acknowledgement This research was enabled in part by support provided by WestGrid ( www.westgrid.ca ) and Compute Canada ( www.computecanada.ca ). MEK is supported in part by a Scholar Award from the Michael Smith Foundation for Health Research, partnered with Centre for Health Evaluation and Outcome Sciences (ID#: 17661). MEK has also received consulting fees from Biogen and participated in Advisory Boards and/or Satellite Symposia of Biogen Inc. TF is supported by the Michael Smith Foundation for Health Research, the Vancouver Coastal Health Research Institute, and the Heart and Stroke Foundation of Canada. TF also receives in-kind study medication from Bayer Canada and received an honorarium for speaking from Servier. Also, we want to thank Liang Xu for sharing her code for efficiently generating all unique allocations. Reference Hemming K, Haines TP, Chilton PJ, et al. The stepped wedge cluster randomised trial: rationale, design, analysis, and reporting. BMJ 2015; 350: h391. Hussey MA, Hughes JP. Design and analysis of stepped wedge cluster randomized trials. Contemp Clin Trials 2007; 28: 182–191. Brown CA, Lilford RJ. The stepped wedge trial design: a systematic review. BMC Medical Research Methodology 2006; 6: 54. Grayling MJ, Wason JMS, Mander AP. Stepped wedge cluster randomized controlled trial designs: a review of reporting quality and design features. Trials 2017; 18: 33. Durovni B, Saraceni V, Moulton LH, et al. Effect of improved tuberculosis screening and isoniazid preventive therapy on incidence of tuberculosis and death in patients with HIV in clinics in Rio de Janeiro, Brazil: a stepped wedge, cluster-randomised trial. Lancet Infect Dis 2013; 13: 852–858. Bacchieri G, Barros AJD, Santos JV dos, et al. A community intervention to prevent traffic accidents among bicycle commuters. Rev Saude Publica 2010; 44: 867–875. Tirlea L, Truby H, Haines TP. Investigation of the effectiveness of the “Girls on the Go!” program for building self-esteem in young women: trial protocol. Springerplus ; 2. Epub ahead of print 19 December 2013. DOI: 10.1186/2193-1801-2-683. Gruber JS, Reygadas F, Arnold BF, et al. A Stepped Wedge, Cluster-Randomized Trial of a Household UV-Disinfection and Safe Storage Drinking Water Intervention in Rural Baja California Sur, Mexico. Am J Trop Med Hyg 2013; 89: 238–245. Copas AJ, Lewis JJ, Thompson JA, et al. Designing a stepped wedge trial: three main designs, carry-over effects and randomisation approaches. Trials 2015; 16: 352. Barker D, McElduff P, D’Este C, et al. Stepped wedge cluster randomised trials: a review of the statistical methodology used and available. BMC Med Res Methodol 2016; 16: 69. Baio G, Copas A, Ambler G, et al. Sample size calculation for a stepped wedge trial. Trials 2015; 16: 354. Hemming K, Taljaard M. Sample size calculations for stepped wedge and cluster randomised trials: a unified approach. J Clin Epidemiol 2016; 69: 137–146. Woertman W, de Hoop E, Moerbeek M, et al. Stepped wedge designs could reduce the required sample size in cluster randomized trials. J Clin Epidemiol 2013; 66: 752–758. Zhou X, Liao X, Spiegelman D. “Cross-sectional” stepped wedge designs always reduce the required sample size when there is no time effect. J Clin Epidemiol 2017; 83: 108–109. Hughes J, Hakhu NR, Voldal E. swCRTdesign: stepped wedge cluster randomized trial (SW CRT) design, https://cran.r-project.org/web/packages/swCRTdesign/index.html (Accessed 20 Aug 2019). Baio G. SWSamp: Simulation-based sample size calculations for a Stepped Wedge Trial (and more). Gianluca Baio /gianluca/software/swsamp/ (2018, accessed 13 May 2019). Hemming K, Girling A. A Menu-Driven Facility for Power and Detectable-Difference Calculations in Stepped-Wedge Cluster-Randomized Trials. The Stata Journal 2014; 14: 363–380. Teerenstra S, Taljaard M, Haenen A, et al. Sample size calculation for stepped-wedge cluster-randomized trials with more than two levels of clustering. Clinical Trials 2019; 16: 225–236. Eldridge SM, Ashby D, Kerry S. Sample size for cluster randomized trials: effect of coefficient of variation of cluster size and analysis method. Int J Epidemiol 2006; 35: 1292–1300. Breukelen GJP van, Candel MJJM. Efficiency loss because of varying cluster size in cluster randomized trials is smaller than literature suggests. Statistics in Medicine 2012; 31: 397–400. Kristunas CA, Smith KL, Gray LJ. An imbalance in cluster sizes does not lead to notable loss of power in cross-sectional, stepped-wedge cluster randomised trials with a continuous outcome. Trials 2017; 18: 109. Girling AJ. Relative efficiency of unequal cluster sizes in stepped wedge and other trial designs under longitudinal or cross-sectional sampling. Stat Med 2018; 37: 4652–4664. Harrison LJ, Chen T, Wang R. Power calculation for cross-sectional stepped wedge cluster randomized trials with variable cluster sizes. Biometrics ; n/a. DOI: 10.1111/biom.13164. Wong H, Ouyang Y, Karim ME. The randomization-induced risk of a trial failing to attain its target power: assessment and mitigation. Trials 2019; 20: 360. Martin JT, Hemming K, Girling A. The impact of varying cluster size in cross-sectional stepped-wedge cluster randomised trials. BMC Med Res Methodol 2019; 19: 123. Matthews JNS. Highly efficient stepped wedge designs for clusters of unequal size. Biometrics ; 2020. DOI: 10.1111/biom.13218. ClinicalTrials.gov[Internet] Ho K, University of British Columbia, 2018 Feb 20. Identifier NCT03439384, TEC4Home Heart Failure: Using Home Health Monitoring to Support the Transition of Care; 2020 Mar 24 [cited 2020 Apr 13]; [about 6 screens]. Available from https://clinicaltrials.gov/ct2/show/NCT03439384. Hasselman B. nleqslv: Solve Systems of Nonlinear Equations, https://CRAN.R-project.org/package=nleqslv (2017, accessed 12 November 2019). Pinheiro J, Bates D, DebRoy S, Sarkar D, R Core Team (2019). nlme: Linear and Nonlinear Mixed Effects Models . R package version 3.1-142, https://CRAN.R-project.org/package=nlme Pinheiro JC, Bates DM (eds). Theory and Computational Methods for Linear Mixed-Effects Models. In: Mixed-Effects Models in S and S-PLUS . New York, NY: Springer, pp. 57–96. R Development Core Team, R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria 2008. ISBN 3-900051-07-0, URL http://www.R-project.org. Accessed 15 Mar 2019. Compute Canada Cedar - CC Doc, https://docs.computecanada.ca/wiki/Cedar (accessed 1 May 2019). Canty A, Ripley B. boot: Bootstrap Functions (Originally by Angelo Canty for S), https://CRAN.R-project.org/package=boot (2019, accessed 28 March 2020). Hemming K, Taljaard M, McKenzie JE, et al. Reporting of stepped wedge cluster randomised trials: extension of the CONSORT 2010 statement with explanation and elaboration. BMJ 2018; 363: k1614. Matthews JNS, Forbes AB. Stepped wedge designs: insights from a design of experiments perspective. Statistics in Medicine 2017; 36: 3772–3790. Kasza J, Forbes AB. Information content of cluster–period cells in stepped wedge trials. Biometrics 2019; 75: 144–152. Johnson JL, Kreidler SM, Catellier DJ, et al. Recommendations for choosing an analysis method that controls Type I error for unbalanced cluster sample designs with Gaussian outcomes. Statistics in Medicine 2015; 34: 3531–3545. Tanner W. Improved Standard Error Estimation for Maintaining the Validities of Inference in Small-Sample Cluster Randomized Trials and Longitudinal Studies. Theses and Dissertations--Epidemiology and Biostatistics . Epub ahead of print 1 January 2018. DOI: https://doi.org/10.13023/etd.2018.434. Rothwell JC, Julious SA, Cooper CL. A study of target effect sizes in randomised controlled trials published in the Health Technology Assessment journal. Trials 2018; 19: 544. Supplementary Files MethodswithEquations.docx Additionalfile.pdf SWCRTsupplementary.docx Cite Share Download PDF Status: Published Journal Publication published 24 Jun, 2020 Read the published version in BMC Medical Research Methodology → Version 2 posted Review # 2 received at journal 10 May, 2020 Review # 1 received at journal 04 May, 2020 Reviewers invited by journal 24 Apr, 2020 Reviewer # 1 agreed at journal 24 Apr, 2020 Reviewer # 2 agreed at journal 24 Apr, 2020 Editor assigned by journal 22 Apr, 2020 Submission checks completed at journal 21 Apr, 2020 Editor invited by journal 21 Apr, 2020 You are reading this latest preprint version Show more versions Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-15483","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research article","associatedPublications":[],"authors":[{"id":515518,"identity":"1e472e30-6bfb-4676-a208-6a34d6a9c48e","order_by":1,"name":"Yongdong Ouyang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5ElEQVRIiWNgGAWjYDCCAxBKjkECwkggWosx6VoSG4jWwne8x0zi547a9P7ZzQcf/mA4nMffwPzwAz4tkmfOmEn2njmeO+POsWRjHobDxRIH2Iwl8GkxuJFjJsHbdix3g0SOmTQDw+HEDQw8DAS1SP5tO5ZuIJH/TfIHRAvzD0JapHnbahIMJHLYJHggWtjw2iJ55lixtWzbAcMZN9KMjXkM0oslDrOZWeDTwne8eePNt2118vwzkh8+/FFhncff3vz4Bj4tQMACdMZhmDuBmJmAepASYCzUEVY2CkbBKBgFIxcAABd7SZ/oAPZtAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0002-8692-2991","institution":"University of British Columbia","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Yongdong","middleName":"","lastName":"Ouyang","suffix":""},{"id":515519,"identity":"2d18662b-a791-405b-af4a-4749e81fbd9c","order_by":2,"name":"Mohammad Ehsanul Karim","email":"","orcid":"","institution":"The University of British Columbia","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Mohammad","middleName":"Ehsanul","lastName":"Karim","suffix":""},{"id":515520,"identity":"3a88f455-29c1-49b6-970b-85a647631e46","order_by":3,"name":"Paul Gustafson","email":"","orcid":"","institution":"The University of British Columbia","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Paul","middleName":"","lastName":"Gustafson","suffix":""},{"id":515521,"identity":"a245e7f1-7427-4d24-9a72-22fc779ae97a","order_by":4,"name":"Thalia S. Field","email":"","orcid":"","institution":"The University of British Columbia","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Thalia","middleName":"S.","lastName":"Field","suffix":""},{"id":515522,"identity":"5f70e018-3c11-410c-b8a5-eb0278686d1a","order_by":5,"name":"Hubert Wong","email":"","orcid":"","institution":"The University of British Columbia","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hubert","middleName":"","lastName":"Wong","suffix":""}],"badges":[],"createdAt":"2020-02-27 12:31:05","currentVersionCode":2,"declarations":"","doi":"10.21203/rs.3.rs-15483/v2","doiUrl":"https://doi.org/10.21203/rs.3.rs-15483/v2","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12874-020-01036-5","type":"published","date":"2020-06-24T12:00:00+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":1043644,"identity":"39201485-f7c3-468b-9058-e79ba832028c","added_by":"auto","created_at":"2020-05-07 01:46:39","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":4879,"visible":true,"origin":"","legend":"Diagram for a standard stepped-wedge cluster randomized controlled trial design. Cells in white correspond to periods during which participants receive standard-of-care, cells in grey correspond to periods during which participants receive the new intervention.","description":"","filename":"fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-15483/v2/fig1.png"},{"id":1043646,"identity":"edf481dc-7812-4863-abd3-7d638c4348e4","added_by":"auto","created_at":"2020-05-07 01:46:40","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":18640,"visible":true,"origin":"","legend":"Examples of power distributions: scenario 85 (48 clusters: 36 L (size: 56.99), 12 S (size: 9); CV:1.0; ICC:0.01), scenario 86 (48 clusters: 36 L (size: 67.85), 12 S (size: 10.72); CV:1.0; ICC:0.05). scenario 87 (48 clusters: 36 L (size: 75.99), 12 S (size: 12); CV:1.0; ICC:0.1).","description":"","filename":"fig2.pdf.png","url":"https://assets-eu.researchsquare.com/files/rs-15483/v2/fig2.pdf.png"},{"id":1043648,"identity":"fad3557d-d57c-49ce-b04c-c8ea925d7e2c","added_by":"auto","created_at":"2020-05-07 01:46:40","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":78374,"visible":true,"origin":"","legend":"The probability that the attained power falls more than five percent below the nominal (left set of panels) or the expected (right set of panels) power by ICC. The risks decreased as ICC increased in all scenarios, except for the ones with 12 clusters. Panel labels identify the total number of clusters (first line) and the distribution of clusters to the cluster size groups S:L or S:M:L (second line).","description":"","filename":"fig3.pdf.png","url":"https://assets-eu.researchsquare.com/files/rs-15483/v2/fig3.pdf.png"},{"id":1043649,"identity":"b33ca629-eb0a-4fc6-b151-3998cf53cdd1","added_by":"auto","created_at":"2020-05-07 01:46:41","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":76985,"visible":true,"origin":"","legend":"The probability that the attained power falls more than five percent below the nominal (left set of panels) or the expected (right set of panels) power by CV. The risks increased as CV increased in all scenarios except the ones with 12 clusters. The risks were near zero when the CV was smaller than 0.7, except for the scenarios with 12 clusters and two size groups. Panel labels identify the total number of clusters (first line) and the distribution of clusters to the cluster size groups S:L or S:M:L (second line).","description":"","filename":"fig4.pdf.png","url":"https://assets-eu.researchsquare.com/files/rs-15483/v2/fig4.pdf.png"},{"id":1043650,"identity":"b59b1f1b-8376-43ea-b60e-de8a593c3b85","added_by":"auto","created_at":"2020-05-07 01:46:41","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":37236,"visible":true,"origin":"","legend":"The relationship between the transition time-points of the large clusters with TTC and TGI from one simulation run for a selected scenario (scenario 37: 24 clusters with 18S,6L). The trend is typical for other runs. The plotted number indicates the number of large clusters transitioning at end steps for that allocation. When all of the large clusters transition at the middle steps (number in green), TTC is high and TGI is low. When all of the large clusters transition at one end step (number in red), both TTC and TGI are low. When all of the large clusters transition equally split between the two ends (number in blue), TTC is low while TGI is high.","description":"","filename":"fig5.pdf.png","url":"https://assets-eu.researchsquare.com/files/rs-15483/v2/fig5.pdf.png"},{"id":1043651,"identity":"11da84f5-8ab7-4b89-8aee-9af76e6f5f09","added_by":"auto","created_at":"2020-05-07 01:46:41","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":141881,"visible":true,"origin":"","legend":"The relationship between treatment-vs-time period correlation and attained power among 16 scenarios. Columns represented difference CV values (from left to right: 0.4, 0.7, 1.0, 1.3). Rows correspond to the number of clusters (12 in the first row; 48 in second row). The colors of the plotted points correspond to ICC of 0.1 (blue) and ICC of 0.01 (red). The attained power decreases considerably as the treatment-vs-time period correlation increases.","description":"","filename":"fig6.png","url":"https://assets-eu.researchsquare.com/files/rs-15483/v2/fig6.png"},{"id":1043652,"identity":"20b9352a-82cc-44b3-b9e6-59cd959a0680","added_by":"auto","created_at":"2020-05-07 01:46:41","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":190976,"visible":true,"origin":"","legend":"The relationship between treatment group imbalance and attained powers among 16 scenarios. Columns represented difference CV values (from left to right: 0.4, 0.7, 1.0, 1.3). Rows correspond to the number of clusters (12 in the first row; 48 in second row). The colors of the plotted points correspond to ICC of 0.1 (blue) and ICC of 0.01 (red). A triangle pattern was observed in most scenarios. Large treatment group imbalances appeared to be associated with higher power.","description":"","filename":"fig7.png","url":"https://assets-eu.researchsquare.com/files/rs-15483/v2/fig7.png"},{"id":1043653,"identity":"d8eaf836-b10a-4666-980a-fde1a0a03c4f","added_by":"auto","created_at":"2020-05-07 01:46:41","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":63091,"visible":true,"origin":"","legend":"Contour plot of predicted attained power as a function of the treatment group imbalance and the treatment-time period correlation for scenario 79. The red dots correspond to the possible allocations in this scenario. The nearly vertical contour curves show that attained power is determined mainly by the treatment-time period correlation and the treatment group imbalance has only a small impact.","description":"","filename":"8.PNG","url":"https://assets-eu.researchsquare.com/files/rs-15483/v2/8.PNG"},{"id":13502092,"identity":"6c415833-27be-4030-8f3a-daa35b700041","added_by":"auto","created_at":"2021-09-16 23:13:18","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1655220,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-15483/v2/5f4dbd72-eaba-41df-b010-4c6ba6d20574.pdf"},{"id":1043645,"identity":"b12487bc-524d-4ac1-a47f-662b2ba33f6e","added_by":"auto","created_at":"2020-05-07 01:46:40","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":19116,"visible":true,"origin":"","legend":"","description":"","filename":"MethodswithEquations.docx","url":"https://assets-eu.researchsquare.com/files/rs-15483/v2/MethodswithEquations.docx"},{"id":1043647,"identity":"78b128c0-38be-4fd2-a8aa-b7989b0f62c7","added_by":"auto","created_at":"2020-05-07 01:46:40","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":12540121,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile.pdf","url":"https://assets-eu.researchsquare.com/files/rs-15483/v2/Additionalfile.pdf"},{"id":1043643,"identity":"8d8e9cbc-0dfd-4cfd-9ed5-cd66aca9dd8c","added_by":"auto","created_at":"2020-05-07 01:46:39","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":9853022,"visible":true,"origin":"","legend":"","description":"","filename":"SWCRTsupplementary.docx","url":"https://assets-eu.researchsquare.com/files/rs-15483/v2/SWCRTsupplementary.docx"}],"financialInterests":"","formattedTitle":"Explaining the variation in the attained power of a stepped-wedge trial with unequal cluster sizes","fulltext":[{"header":"Background","content":"\u003cp\u003eUse of the stepped-wedge cluster randomized controlled trial (SW-CRT) design has increased dramatically in recent years.\u003csup\u003e1\u0026ndash;3\u003c/sup\u003e I In a standard stepped-wedge design (Figure 1), every cluster () begins (time-point ) with delivery of the existing standard-of-care (shown as white cells) to participants. Clusters then transition to delivery of a new intervention (shown as gray cells) at randomly determined time-points (), until all clusters are delivering the new intervention. The SW-CRT design often is used when there is a desire or need to implement and evaluate the intervention at the population level\u003csup\u003e4\u003c/sup\u003e, or when it would not be logistically feasible to implement the intervention in every cluster at the same time\u003csup\u003e3\u003c/sup\u003e, or when recruitment of clusters could be enhanced by ensuring all clusters eventually receive the new intervention. It has been implemented in trials exploring both single interventions as well as pathways of care in multiple settings.\u003csup\u003e5\u0026ndash;8\u003c/sup\u003e Copas et al.\u003csup\u003e9\u003c/sup\u003e described three main SW-CRT designs: the \u0026ldquo;closed cohort\u0026rdquo;, the \u0026ldquo;open cohort\u0026rdquo;, and the \u0026ldquo;continuous recruitment short exposure\u0026rdquo; designs.\u0026nbsp; In the first two designs, the treatment received by a participant during follow-up matches the treatment being delivered by his/her cluster at each time-point and each participant contributes multiple outcome measurements over time. In the third design, participants receive the treatment being delivered by his/her cluster at the time of study entry, remains on this treatment throughout follow-up, and contributes a single outcome measurement. For simplicity, we will refer to the continuous recruitment short exposure design as a \u0026ldquo;cross-sectional\u0026rdquo; design, a term which although less precise, is more commonly used.\u003csup\u003e3,10\u003c/sup\u003e The results presented throughout this paper were derived using a cross-sectional SW-CRT.\u003c/p\u003e\n\u003cp\u003eMethods for calculating power and sample size in SW-CRTs with equal cluster sizes have been discussed widely in the literature.\u003csup\u003e2,11\u0026ndash;14\u003c/sup\u003e Code for calculating the power of a SW-CRT with equal cluster sizes are also available for commonly used statistical software (R\u003csup\u003e15,16\u003c/sup\u003e, Stata\u003csup\u003e17\u003c/sup\u003e and \u0026nbsp;SAS\u003csup\u003e18\u003c/sup\u003e). Although it has been recognized that unequal cluster sizes in a parallel cluster randomized trial (P-CRT) leads to loss of power and efficiency\u003csup\u003e19,20\u003c/sup\u003e, limited work has been conducted on power and sample size calculations in unequal cluster size SW-CRTs. Hussey and Hughes\u003csup\u003e2\u003c/sup\u003e described a procedure for calculating approximate power based on a Wald test which allows for unequal cluster sizes in the SW-CRT, but they did not explore the impact of unequal cluster sizes. Kristunas et al.\u003csup\u003e21\u003c/sup\u003e conducted simulations that showed that the power of a SW-CRT with unequal cluster sizes was only slightly lower than if the cluster sizes had been equal. Girling\u003csup\u003e22\u003c/sup\u003e extended the method for evaluating the precision of regular single-period P-RCT with unequal cluster sizes to the stepped wedge design. Harrison et al.\u003csup\u003e23\u003c/sup\u003e derived first-order approximation formulae for the expected treatment effect variance given either known cluster sizes or the mean and coefficient of variation of the cluster sizes. They then used this treatment effect variance formula in a Wald test to conduct sample size calculations as well as to investigate the relative efficiency comparing equal and unequal cluster size SW-CRT designs and the design effect relative to individually randomized designs. However, it is important to recognize that these results focused on the \u003cem\u003eexpected power\u003c/em\u003e, that is, the power obtained through averaging over both 1) the randomness in the allocation of the clusters and 2) the random outcome variation across participants. Distinguishing between the two components is important. The latter source of variation is an intrinsic attribute of the population and outside the control of the investigator. However, the trial investigator has control over the former source through the choice of the randomization algorithm, and this choice impacts on trial power.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBoth Wong et al.\u003csup\u003e24\u003c/sup\u003e and Martin et al.\u003csup\u003e25\u003c/sup\u003e showed that the \u003cem\u003eattained power\u003c/em\u003e (Wong et al.), also called the \u003cem\u003eactual power\u003c/em\u003e (Martin et al.), defined as the power associated with a specific allocation, can vary substantially across realized allocations and can be much lower than the expected power. Through deriving an analytic approximation, Matthews\u003csup\u003e26\u003c/sup\u003e showed how the variance of the treatment effect estimate varies across allocations. Thus, a trial that was expected to achieve a specified power prior to allocation of the clusters might turn out to be underpowered due to an \u0026ldquo;unlucky\u0026rdquo; allocation. For example, we are involved in a SW-CRT\u003csup\u003e27\u003c/sup\u003e with 20 hospitals (clusters) in which the two largest hospitals were expected to enrol more than 200 participants each, while the four smallest hospitals were expected to enrol less than 10 participants each, with a cluster size coefficient of variation of 1.17. Power calculations obtained via simulation showed that although using an unrestricted randomization algorithm would yield an expected power of 80%, the attained power varied from a low of 68% to a high of 83% depending on the allocation.\u003c/p\u003e\n\u003cp\u003eMartin et al.\u003csup\u003e25\u003c/sup\u003e observed larger variation in attained powers in designs with a smaller number of clusters, and a smaller intraclass correlation coefficient (ICC)\u003cem\u003e.\u003c/em\u003e In addition, they observed that the absolute difference in the numbers of participants allocated to the two treatment groups, referred to hereafter as the \u0026ldquo;treatment group imbalance (TGI)\u0026rdquo;, explained only a small portion of the variability in attained power across allocations in the SW-CRT and that contrary to intuition, the attained power was relatively higher for allocations with a large TGI. In limited scenarios, Wong et al.\u003csup\u003e24\u003c/sup\u003e found that the (Pearson) correlation between the participant-level treatment assignment and the time-period with clustering ignored, referred to hereafter as the \u0026ldquo;treatment-vs-time period correlation (TTC)\u0026rdquo;, had a large impact on the attained power. They observed that the allocations with the greatest attained power were those in which the large clusters transitioned during the first or last steps (equally split) and that these allocations had the lowest TTC. In the SW-CRT, it is important to adjust for the time period to avoid bias in the treatment effect estimate because treatment group and time period inherently are highly correlated. They proposed that as TTC increases, this adjustment increases the standard error of the estimated treatment effect, which leads to a greater loss in attained power. However, we are not aware of a comprehensive study on these topics.\u003c/p\u003e\n\u003cp\u003eIn this study, we investigated how different allocation characteristics interacted with design factors to affect attained power and the risk of obtaining low attained power in cross-sectional SW-CRTs. This risk can be assessed by constructing the \u003cem\u003epre-randomization power distribution (PD)\u003c/em\u003e,\u003csup\u003e24\u003c/sup\u003e defined as the distribution of attained powers obtained from all possible allocations that a randomization algorithm can generate. A good randomization algorithm will ensure that the risk of obtaining a low attained-power allocation is acceptably small. Identifying such an algorithm requires an understanding of what factors cause low attained power. The first aim of this work was to gain an understanding of the factors that affect the risk of low attained power. We used simulation to evaluate the attained power across different allocations under a wide variety of scenarios. While it is possible to assess the attained power using approximate analytic formulae (e.g., Hussey \u0026amp; Hughes\u003csup\u003e2\u003c/sup\u003e), the accuracy of those formulae has not been well investigated in the context of SW-CRTs, especially when the number of clusters or the cluster sizes are low.\u0026nbsp; However, the computational time needed to simulate the attained powers for all possible allocations often is not feasible. Hence, our second aim was to develop regression models that could accurately predict the attained power for any allocation based on allocation characteristics (specifically TGI and TTC) and the attained powers from a sample of allocations. The results of this work provide guidance on how to assess attained power and avoid having an unacceptably low attained power when designing a SW-CRT with unequal cluster sizes.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp style=\"line-height: 200%;\"\u003e\u003cstrong\u003eAim One: Evaluating attained power and the impact of different factors on the risk of low attained power\u003c/strong\u003e\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003eIn this study, we assumed that individual outcomes, , were generated from the following linear mixed-effects model:\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp style=\"text-align: center; line-height: 200%;\"\u003e[Please see the supplementary files section to view the equations.] \u0026nbsp;\u0026nbsp;\u0026nbsp; \u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; (1)\u0026nbsp;\u003c/p\u003e\n\u003cp style=\"text-align: center; line-height: 200%;\"\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003ewhere \u0026nbsp;indexed the cluster, \u0026nbsp;indexed the time period (denotes baseline), and \u0026nbsp;indexed the individual within cluster \u0026nbsp;and time period . The error terms \u0026nbsp;were assumed to be independently sampled from a standard normal distribution (mean zero and variance ). \u0026nbsp;treatment group indicator (0 = standard-of-care, 1 = intervention) was denoted as and the treatment effect, , was set to 0.26.\u0026nbsp; As , the standardized effect size, , also equaled 0.26. Because the variance of the treatment effect estimate does not depend on the linear time effect coefficient, , its value was set to 1 without loss of generality. The cluster-specific effects, , were assumed to be independently sampled from a normal distribution with mean zero and variance \u0026nbsp;. The value of \u0026nbsp;was calculated from set values of the ICC, where ICC was defined as . The treatment-vs-time period correlation (TTC) was defined as the Pearson correlation coefficient between treatment and the time-period, , across all individual observations with clustering ignored. The treatment group imbalance (TGI) was defined as the difference in the numbers of participants who received the new intervention versus the standard-of-care.\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003e\u003cstrong\u003e\u003cem\u003eComputational challenge and a practical solution:\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003eOur aim was to obtain the set of attained powers by simulation and their corresponding power distribution for a specified randomization algorithm. First, we needed to obtain a list of all possible allocations that the randomization algorithm can generate, and then to evaluate the attained power associated with each possible allocation.\u0026nbsp;\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003eHowever, the evaluation of attained power for all potential allocations can be computationally challenging, given the potentially huge number of possible allocations. For example, a twenty-cluster cross-sectional SW-CRT with unique cluster sizes and four clusters transitioning at each of five steps has more than 300 billion unique allocations.\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003eOne strategy to reduce the number of unique allocations was by categorizing the clusters into size groups (e.g., (S)mall/(L)arge or (S)mall/(M)edium/(L)arge) and treating the clusters within each size group as identical, thereby reducing the number of unique allocations while increasing their multiplicities. Doing so will lead to only an approximate solution. However, this is often adequate in practice, given that the anticipated number of individuals that a cluster will enroll often can be only approximated at the start of a trial. In practice, we do not recommend having more than four size categories, as then the number of unique allocations likely will be too large to evaluate.\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003e\u003cstrong\u003e\u003cem\u003eSimulation specifications:\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003eAs shown in Table 1, we investigated the impact of the total number of clusters (12, 24 or 48), the number of cluster size groups (2 \u0026ndash; S/L or 3 \u0026ndash; S/M/L), the distribution of clusters across cluster size groups (equal or unequal), the coefficient of variation (CV) of the cluster sizes (0.4, 0.7, 1.0, and 1.3), defined as the ratio of standard deviation of cluster sizes to the mean, and the ICC of the response variable (0.01, 0.05, and 0.10). A full factorial layout was investigated for all of these factors except the CV. An equal distribution of clusters across cluster size groups corresponded to an equal number of clusters in each cluster size group; an unequal distribution corresponded to distribution of clusters in a S:L ratio of 3:1 for scenarios with two cluster size groups or 3:2:1 to S:M:L for scenarios with three cluster size groups. For unequal distributions, CVs of 0.4. 0,7, 1.0, and 1.3 were investigated for all unequal distribution scenarios, but for equal distribution scenarios, CVs of only 0.4 and 0.7 were investigated as a CV larger than one cannot be obtained with equal distributions. (Thus, when the actual cluster sizes have a CV larger than one, trial designers need to use unequal distributions to apply the approach presented here.) Thus, there were a total of 108 scenarios.\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp style=\"margin-bottom: 10.0pt; page-break-after: avoid;\"\u003e\u003cstrong\u003e\u003cspan style=\"font-size: 10.0pt;\"\u003eTable 1 Factors and their specific values explored in the simulation study.\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003ctable style=\"width: 467.5pt; border-collapse: collapse; border: none;\" width=\"623\"\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd style=\"width: 148.6pt; border: none; border-bottom: solid black 1.0pt; padding: 0in 5.4pt 0in 5.4pt;\" width=\"198\"\u003e\n\u003cp style=\"line-height: 150%;\"\u003e\u003cstrong\u003eFactor\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 318.9pt; border: none; border-bottom: solid black 1.0pt; padding: 0in 5.4pt 0in 5.4pt;\" width=\"425\"\u003e\n\u003cp style=\"line-height: 150%;\"\u003e\u003cstrong\u003eValues\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd style=\"width: 148.6pt; border: none; border-bottom: solid black 1.0pt; padding: 0in 5.4pt 0in 5.4pt;\" width=\"198\"\u003e\n\u003cp style=\"line-height: 150%;\"\u003e\u003cstrong\u003eNumber of clusters\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 318.9pt; border: none; border-bottom: solid black 1.0pt; padding: 0in 5.4pt 0in 5.4pt;\" width=\"425\"\u003e\n\u003cp style=\"line-height: 150%;\"\u003e12, 24, 48\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd style=\"width: 148.6pt; border: none; border-bottom: solid black 1.0pt; padding: 0in 5.4pt 0in 5.4pt;\" width=\"198\"\u003e\n\u003cp style=\"line-height: 150%;\"\u003e\u003cstrong\u003eNumber of cluster size groups \u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 318.9pt; border: none; border-bottom: solid black 1.0pt; padding: 0in 5.4pt 0in 5.4pt;\" width=\"425\"\u003e\n\u003cp style=\"line-height: 150%;\"\u003e2, 3\u003c/p\u003e\n\u003cp style=\"line-height: 150%;\"\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd style=\"width: 148.6pt; border: none; border-bottom: solid black 1.0pt; padding: 0in 5.4pt 0in 5.4pt;\" width=\"198\"\u003e\n\u003cp style=\"line-height: 150%;\"\u003e\u003cstrong\u003eDistribution of clusters to cluster size groups\u003c/strong\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 318.9pt; border: none; border-bottom: solid black 1.0pt; padding: 0in 5.4pt 0in 5.4pt;\" width=\"425\"\u003e\n\u003cp style=\"line-height: 150%;\"\u003eEqual (6S+6L, 12S+12L, 24S+24L, 4S+4M+4L, 8S+8M+8L, 16S+16M+16L), Unequal (9S+3L, 18S+6L, 36S+12L, 6S+4M+2L, 12S+8M+4L, 24S+16M+8L)\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd style=\"width: 148.6pt; border: none; border-bottom: solid black 1.0pt; padding: 0in 5.4pt 0in 5.4pt;\" width=\"198\"\u003e\n\u003cp style=\"line-height: 150%;\"\u003e\u003cstrong\u003eICC\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 318.9pt; border: none; border-bottom: solid black 1.0pt; padding: 0in 5.4pt 0in 5.4pt;\" width=\"425\"\u003e\n\u003cp style=\"line-height: 150%;\"\u003e0.01, 0.05, 0.1\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd style=\"width: 148.6pt; border: none; border-bottom: solid black 1.0pt; padding: 0in 5.4pt 0in 5.4pt;\" width=\"198\"\u003e\n\u003cp style=\"line-height: 150%;\"\u003e\u003cstrong\u003eCV\u003c/strong\u003e\u003cstrong\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp; \u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"width: 318.9pt; border: none; border-bottom: solid black 1.0pt; padding: 0in 5.4pt 0in 5.4pt;\" width=\"425\"\u003e\n\u003cp style=\"line-height: 150%;\"\u003e0.4, 0.7, 1.0*, 1.3*\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp style=\"text-align: justify; text-justify: inter-ideograph;\"\u003e\u003cstrong\u003e\u003cspan style=\"font-size: 9pt; font-family: 'Times', serif;\"\u003e* \u003c/span\u003e\u003c/strong\u003e\u003cstrong\u003e\u003cspan style=\"font-size: 9.0pt; font-family: 'Times',serif;\"\u003eFor equal distribution of clusters to cluster size groups, only CVs of 0.4 and 0.7 were investigated as a CV of 1 or larger is not possible.\u003c/span\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp style=\"line-height: 150%;\"\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003eTrial parameter values for all 108 scenarios are listed in the Additional file (Table A1). The number of transition steps was fixed at four. To determine the total number of individuals needed to obtain approximately 80% power in each scenario, we used Hussey and Hughes\u0026rsquo;s\u003csup\u003e2\u003c/sup\u003e sample size formula for a SW-CRT with equal cluster sizes as implemented by Baio et al.\u003csup\u003e11\u003c/sup\u003e in the R package \u0026ldquo;SWSamp\u003csup\u003e14\u003c/sup\u003e\u0026rdquo;\u003csup\u003e \u0026nbsp;\u003c/sup\u003e(implementation of the Hussey and Hughes\u0026rsquo;s formula). Then, we re-distributed the total number of individuals to create unequal sized clusters that matched the CVs we set under each scenario (see Additional file, A2 for details).\u003csup\u003e28\u003c/sup\u003e If a resulting cluster size was not an integer, we set it at random to one of the two bracketing integers such that the expected value matched the initially calculated cluster size. For example, if the calculated cluster size was 12.3, the cluster size was set at random to either 12 or 13, such that its expected value was 12.3. Within a cluster, individuals were allocated to the time periods using a multinomial distribution with equal probability for each period. The total number of unique allocations is presented in the third column from the right in Table A.1. Among the 108 scenarios, 72 scenarios (referred to as \u0026lsquo;completed scenarios\u0026rsquo;) had less than 2,000 unique allocations. For these scenarios, the attained power for every allocation was evaluated via simulation. For the other 36 scenarios (referred to as \u0026lsquo;sampled scenarios\u0026rsquo;), we evaluated the attained power for only a sample of 2,000 allocations. Subsequently, the attained power for two sampled scenarios, one with a relatively small number of allocations (scenario #64, with 8,623 allocations) and one with a relatively large number of allocations (scenario #103, with 113,949 allocations), had all of their allocations evaluated to enable validation of the prediction models that were constructed using only the original 2,000 allocations.\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003e\u003cstrong\u003e\u003cem\u003ePerformance metrics:\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003eFor each selected allocation, 10,000 datasets were simulated and analyzed using the model (1). Analysis was performed using the \u003cem\u003elme\u003c/em\u003e function in R.\u003csup\u003e29\u003c/sup\u003e This P-value is calculated through comparing the Wald statistic to quantiles from a \u003cem\u003et\u003c/em\u003e-distribution with degrees of freedom equivalent to what would be used in a balanced, multilevel ANOVA designs\u003csup\u003e30\u003c/sup\u003e. The attained power was estimated using the proportion of P-values less than 0.05. Assuming a power near 80%, the simulation standard error of each attained power was approximately 0.4%. The PD associated with the unrestricted randomization algorithm was then constructed from the estimated attained powers by weighting each attained power by the probability of the corresponding allocation. Two measures of the risk of low attained power were then obtained from the PD: (1) the probability that the attained power falls more than 5% below the expected power (obtained in the simulations), and (2) the probability that the attained power falls more than 5% below the nominal 80% power (i.e., less than 75%) that would be achieved in a trial with equal cluster sizes. The first measure provides an indication of whether potential low attained power is a concern when power is calculated using the approach presented here for handling unequal cluster sizes. The second measure would be of greater interest when one is assessing whether power calculations obtained under the assumption of equal cluster sizes are adequate. We set a 5% power loss as being meaningful, but trial designers may wish to choose a different value more relevant to their context. In this paper, the phrase \u0026ldquo;risk of low attained power\u0026rdquo; will be used as a short form to refer to either of these measures. All computations were conducted using R version 3.5.0 on Cedar, Compute Canada\u003csup\u003e31,32\u003c/sup\u003e.\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003e\u003cstrong\u003eAim Two: Explaining the variation in attained power across allocations and predicting attained power using allocation characteristics\u003c/strong\u003e\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003eUsing the simulation results from Aim One, we first constructed scatterplots to examine the bivariate relationship between the simulated attained power and each of TTC and TGI separately. \u0026nbsp;Then, to predict the attained power for each scenario, we fitted logistic regression models of the proportion of simulations in each allocation that achieved statistical significance (at level 0.05) as a function of linear and quadratic terms for TTC and TGI.\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003eBased on the relationships observed in the bivariate scatterplots, we chose to investigate a sequence of four nested models. The terms included in the four models that were considered were: (Model 1: TTC), (Model 2: TTC, TGI), (Model 3: TTC, TGI, TTC\u003csup\u003e2\u003c/sup\u003e), and (Model 4: TTC, TGI, TTC\u003csup\u003e2\u003c/sup\u003e, TGI\u003csup\u003e2\u003c/sup\u003e).\u003c/p\u003e\n\u003cp style=\"line-height: 200%;\"\u003eWe measured the predictive accuracy of each model by using five-fold cross-validation, conducted using cv.glm\u003csup\u003e33\u003c/sup\u003e in R, to estimate the root mean squared prediction error (RMSPE), defined by , where \u0026nbsp;represented the set of allocations in the -th partition, \u0026nbsp;was the number of allocations in the -th partition, and \u0026nbsp;were the simulated and predicted powers, respectively, for the -th allocation, and \u0026nbsp;was the number of unique allocations in the scenario. From these models, we selected the one for which no meaningful improvement in RMSPE was achieved by adding another term to the model. For the selected model, we examined additional predictive performance measures for each scenario, including maximum and average absolute prediction errors (). For two sampled scenarios (#64, #103), we validated the selected model with respect to both prediction of attained power for individual allocations, and estimation of the risk of low attained power. For the former objective, we compared the predicted attained power to the simulated attained power for each allocation that was not in the set used to fit the predictive model. For the latter objective, we repeatedly sampled (10,000 times) 2000 allocations and estimated the risk of low attained power from the fitted model. We then compared these estimated risks to the \u0026ldquo;true\u0026rdquo; risk derived from the simulated attained powers.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eAim One: Evaluating the attained powers and factors affecting the risk of low attained power\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn the body of this paper, we report on the scenarios with an \u003cstrong\u003eunequal distribution\u003c/strong\u003e of clusters to cluster size groups. The results for scenarios with an equal distribution of clusters to cluster size groups were similar when matched on the remaining factors (see Additional file, Figure A3.1 and A3.2).\u003c/p\u003e\n\u003cp\u003eSubstantial variation in attained powers was observed in all the scenarios. Across all simulations, attained power ranged from 0.62 to 0.86. As examples, the PDs for selected scenarios (#85 to #87) are displayed in Figure 2. The expected powers obtained from the simulations (Additional file, Table A1, the \u0026ldquo;Weighted expected power\u0026rdquo;) usually were quite close to the nominal 80% power associated with an equal cluster size design, but losses of up to roughly 4% were seen when the CV was large.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eICC\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFigure 3 shows the risk of low attained power (for both measures) as a function of ICC for different combinations of the total number of clusters, the CV, and the distribution of clusters to the size groups. The risk of low attained power varied greatly with ICC. For both risk measures, with a total of 24 or 48 clusters, the risk of obtaining low power was smaller in the scenarios with larger ICCs. If nominal power is used as the reference, the risk was as high as 38% when the ICC was 0.01 but dropped to nearly zero even when the ICC was 0.1. When there were 12 clusters in total, the risk of low attained power often continued to be high even with larger ICCs, and the patterns were not as evident. However, it should be noted that in these scenarios, the number of unique allocations is quite low so the combined effect of simulation error and discretization effects may have disrupted the monotonic patterns that were expected.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eCoefficient of variation\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFigure 4 plots the risks of low attained power as a function of the CV. The risk of low attained power varied greatly across the CVs and depended on the ICC. Except for the scenarios with 12 clusters and two size groups, the risk was near zero when the CV was 0.4. However, as the CV increased above 0.7, the risk increased rapidly up to more than 30% when the CV was 1.3 and the ICC was 0.01.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAim Two: Explaining the variation in attained power across allocations and predicting attained power using allocation characteristics\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eRelationship among transition time-points of large clusters, treatment-vs-time period correlation (TTC) and absolute treatment group imbalance (TGI) \u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eUnderstanding the relationship between the time-points when the large clusters transition, TTC, and TGI will be helpful for interpreting the predictive models presented later. Figure 5 summarizes the relationship typically seen between the transition time-points of the large clusters with TTC and TGI across simulation runs from one representative scenario (#37). Plotted symbols represent different allocations. The number used as the plotting symbol indicates the number of large clusters transitioning at the two ends (i.e., at either the first or last step). Setting aside the values of TGI, this plot shows that the number of large clusters transitioning at the ends strongly determines TTC, with TTC decreasing as this number increases. When the transition time-points of the large clusters are equally split between the two ends, both TTC and TGI are low (numbers plotted in red), but when these transitioning times are concentrated at one end, TTC remains low while TGI is high (numbers plotted in blue). When all of the large clusters transition at the middle steps (numbers plotted in green), TGI will be low and TTC is high.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eImpact of treatment-time period correlation on attained power \u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFFigure 6 displays the relationship between the attained power for a given allocation and TTC for a selected set of 16 scenarios (two size groups , 12 or 48 clusters, ICC = 0.01 or 0.1). Within this grid of panels, columns correspond to different CV values (0.4, 0.7, 1.0, 1.3 from left to right) and rows correspond to the number of clusters (12 in the first row; 48 in the second row). Within each panel, points in blue correspond to an ICC of 0.1, while those in red correspond to an ICC of 0.01. A graph for all scenarios is included in the Additional file, Figure A4. Across all scenarios, TTC consistently exhibited a strong negative, and mildly non-linear, relationship with the attained power. Greater variation in attained power was observed in the scenarios with a larger CV, a lower ICC, and a smaller number of clusters.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eImpact of treatment group imbalance on attained power\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFigure 7 displays the relationship between attained power and TGI, using the same layout and for the same scenarios used in Figure 6.\u0026nbsp; A graph for all scenarios is included in the Additional file, Figure A5.\u0026nbsp; Across all scenarios, the association exhibited a \u0026ldquo;triangle\u0026rdquo; pattern with much weaker associations than were seen for TTC. In 90% (97/108) of the scenarios, univariate regression model fits showed that allocations with a large TGI were associated with a relatively higher attained power. This is counter-intuitive and opposite to the P-CRT setting, where attained power decreases with increasing TGI. When TGI was low, the attained power ranged over both low and high values across different allocations. This unexpected result is resolved in the following section. Essentially, the complicated relationship between TTC and TGI as described earlier leads to the true impact of TGI on attained power being confounded by its association with TTC, with the latter being a much stronger determinant of attained power.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003ePredictive model\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWithout any predictors (i.e., the null model with only an intercept term), the RMSPE of the predicted attained power ranged across the 108 scenarios from 0.0060 to 0.0454 with a median of 0.0157. When TTC alone (Model 1) was added to the model, the RMSPE decreased in every scenario with the median RMSPE shrinking down to 0.0037 (range: 0.0022, 0.0103). When TGI was then added (Model 2), the RMSPE decreased in 93 of the scenarios and the median RMSPE dropping down to 0.0035 (range 0.0020 to 0.0085). Note that in comparison to the crude effect of TGI, after adjustment for TTC, the percentage of scenarios with a positive TGI coefficient declined from 90% to 68% (74/108) and the distribution of TGI coefficients was compressed much closer to zero compared to the distribution of TGI coefficients from the unadjusted models (see Additional file, Figure A6).\u003c/p\u003e\n\u003cp\u003eOn adding the TTC\u003csup\u003e2\u003c/sup\u003e term (Model 3), the RMSPE decreased in 99 scenarios, with the median RMSPE shrinking to 0.0031 (range 0.0017, 0.0065). Further adding TGI\u003csup\u003e2\u003c/sup\u003e (Model 4), resulted in the RMSPE decreased in 71 scenarios but the median RMSPE changed negligibly (i.e., on the order of \u0026nbsp;). Hence, we selected the Model 3 as the final model.\u003c/p\u003e\n\u003cp\u003eFigure 8 displays a contour plot of the predicted attained power derived from the regression model for one scenario (#79). The almost-vertical curves indicate that attained power decreased substantially as TTC increased. However, when TTC was fixed, the attained power decreased only slightly as TGI increased.\u003c/p\u003e\n\u003cp\u003eRMSPE, maximum absolute prediction error, and average absolute prediction error for each scenario are shown in Table A7 in the Additional file.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003e\u0026nbsp;\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eValidation of prediction model\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe medians of the differences between the predicted and the simulated attained powers for the allocations not used in model fitting for scenarios #64 and #103 were near zero and 95% of the predicted attained powers fell within 0.0069 (#64) and 0.0094 (#103) of the simulated attained powers (See Additional file, A8 for the distribution of differences). This magnitude of prediction error is comparable to the expected error in the simulated attained powers (i.e., 95% of simulated attained power within 2 x 0.004 = 0.008), suggesting the selected model performs nearly as well as simulation.\u0026nbsp; The true risks of low attained power (\u0026gt; 5% below nominal 80%), calculated from the simulated attained powers, were 0.0201 (#64) and 0.0697 (#103). In repeated sampling and fitting of the selected model, the median absolute difference between the model-based estimate and the true risk of low attained power was 0.0010 (#64) and 0.0019 (#103), with 95% of the values falling within 0.0018 (#64) and 0.0034 (#103).\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003e\u003cstrong\u003eUnderstanding the risk of obtaining low power\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eUsing simulation, we have examined how and why the attained power in an SW-CRT with unequal cluster sizes may be different from the expected power that typically is the focus of power calculations. We have argued that even if the expected power is adequate, the risk that the randomization leads to a trial with low attained power should be considered. The extension of the CONSORT 2010 statement\u003csup\u003e34\u003c/sup\u003e recommends that researchers report the unequal cluster sizes in trials. This study highlights the importance of knowing the cluster sizes to enable appropriate assessments of power. The power distribution constructed from the set of attained powers can be used to evaluate the risk that a given randomization algorithm will yield a trial with unacceptably low power.\u003c/p\u003e\n\u003cp\u003eAs shown in our study, the CV of the cluster sizes strongly impacted the risk of low attained power. For CV values below 0.7, the risk of the attained power falling more than 5% below the expected power across the investigated ranges of values for the other parameters was sufficiently low that this risk may not be a concern. As the CV increased above 0.7, the risk of low attained power increased substantially, especially when the ICC was low. As the ICC increased, the risk of low attained power decreased. This result is consistent with Matthews et al.\u003csup\u003e35\u003c/sup\u003e, who showed that in a \u0026ldquo;row-column\u0026rdquo; analysis of a SW-CRT, which corresponds to a linear mixed effects analysis, the variance of the treatment effect decreased as the ICC increased. Our heuristic explanation is that when the ICC was large, the \u003cem\u003eeffective\u003c/em\u003e sample sizes of large clusters were reduced greatly, which in turn reduced the \u003cem\u003eeffective\u003c/em\u003e CV of the cluster sizes. As expected, the risk of low attained power decreased as the number of clusters increased. However, even with 48 clusters, this risk was not ignorable when the CV was large, and the ICC was small.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eExplaining the variation in attained power\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA key contribution of this study is clarifying how attained power is impacted by the joint effects of treatment-vs-time period correlation and absolute treatment group imbalance. TTC was shown to be the dominating factor. Our results suggest that the observation that TGI on its own is only weakly related to attained power and in the direction opposite to what is expected, made by Martin et al\u003csup\u003e20\u003c/sup\u003e and confirmed here, is a consequence of confounding with TTC. After adjusting for TTC, the impact of TGI on the attained power was reduced or the direction reversed in nearly every scenario. That the direction was not reversed for every scenario may be due in part to simulation noise, but it could also reflect residual confounding due to the model mis-specifying the true dependence of attained power on TTC, or the omission of other important predictors.\u003c/p\u003e\n\u003cp\u003eWe have shown that a model containing TTC, TTC\u003csup\u003e2\u003c/sup\u003e, and TGI can accurately predicts attained power for any allocation, and hence to construct the power distribution for any randomization algorithm and to estimate the risk of low attained power, when the number of allocations is too large to evaluate their attained powers using simulation.\u003c/p\u003e\n\u003cp\u003eThe original rationale for creating size groups was to reduce the number of unique allocations that needed to be evaluated to obtain the power distribution. However, given the high accuracy seen in these predictive models, creating cluster size groups may not be necessary. Instead, a predictive model could be built based on a representative sample of allocations without grouping of clusters and this model could be used to predict the attained powers associated with each non-sampled allocation. The caveats here are that some effort may be needed to determine how large a sample is needed to achieve a target predictive accuracy, and when the number of unique allocations is large, simply listing out all possible unique allocations may require excessive (computational) effort in which case size groups will still be needed.\u003c/p\u003e\n\u003cp\u003eExamination of attained power can assist in the identification of the allocation characteristics that lead to low attained power. For example, our study verified the results from previous work\u003csup\u003e24\u003c/sup\u003e showing that transitioning large clusters during the middle steps can increase the treatment-vs-time period correlation and lead to low attained power. This result is also consistent with Kasza and Forbes\u003csup\u003e36\u003c/sup\u003e, who showed that for an equal-cluster-size design, the sequences (i.e., allocations) that transition during the middle steps contribute less information than those that transition at the end steps. Therefore, allocating large clusters to transition during the middle steps should (as a heuristic) yield less information than if they were allocated to transition at the end steps. Identifying these allocations affords trialists the opportunity to take pre-emptive actions to mitigate this problem. The simplest option would be to restrict the randomization algorithm to exclude low power allocations. However, this approach risks violating the criteria needed for valid randomization inference, as it could result in some clusters having no chance of being randomized to transition at particular time-points (e.g., large clusters may not be allowed to transition during the middle step(s)). A more conservative option could be to stratify randomization based on cluster size. This would ensure that every large cluster has a chance to be allocated to any transition step, while ensuring even distribution of large clusters across all of the steps. This would tend to yield a low-variability power distribution, as the excluded allocations would tend to be both those with low and high power.\u0026nbsp; Further exploration of potential restricted randomization algorithms and their validity, benefits and limitations is needed.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003e\u0026nbsp;\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLimitations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Wald test we used to evaluate statistical significance is known to yield higher-than-nominal Type I error rates\u003csup\u003e37\u003c/sup\u003e. Therefore, the powers that are reported here will tend to be higher than what would be obtained using a test that achieved the nominal size. We have no reason to expect that the patterns of findings that we have reported would be different if the test used achieved the nominal size.\u0026nbsp; However, important future work would be to ascertain which test(s) provide the most accurate size and power values to ensure that real-life trial designs fulfill the desired statistical criteria.\u0026nbsp; Some work towards this goal has been conducted by Tanner\u003csup\u003e38\u003c/sup\u003e, using various analytic models (GLMM, GEE, cluster-level analysis) for binary outcomes.\u003c/p\u003e\n\u003cp\u003eOur study left several design factors as fixed values, including the number of transition steps at four, and a standardized effect size of 0.26. Others\u003csup\u003e25\u003c/sup\u003e have reported that the trial power is only weakly affected by the number of steps. Results from limited simulations we conducted with six steps were concordant with those reports (data not shown), so we did not pursue a comprehensive investigation in this regard. The standardized effect size of 0.26 arose from a calculation performed to achieve 80% power in one particular scenario.\u0026nbsp; We adopted this value as being within the typical range seen in real trials\u003csup\u003e39\u003c/sup\u003e. Assuming a larger effect size would have led to smaller cluster sizes needed to achieve 80% power. Again, we have no reason to expect different patterns to our findings, but smaller sample sizes may lead to greater Type I error and inaccuracy in the estimated power. The numbers of clusters (12/24/48) were chosen primarily to reflect our belief that these are typical numbers seen in real SW-CRTs (but also for convenience as these choices allowed the numbers of clusters to be balanced across steps and cluster size groups).\u0026nbsp; While we found that increasing the number of clusters reduces the risk of low attained power, we did not establish a bound above which this risk is ignorable. When the CV is large and the ICC is small, this risk will need to be evaluated even when a trial includes a larger number of clusters.\u003c/p\u003e\n\u003cp\u003eDue to limits on available computational time, we validated the predictive model for only two scenarios. The selected predictive model was shown to predict the attained power with accuracy similar to that achieved using simulation with 10,000 replicates. However, because this model may be mis-specified, there is no guarantee that this model will yield sufficiently accurate predictions if greater precision is required. Alternative models may need to be developed, perhaps including other allocation characteristics. In addition, simulating attained power for 2000 allocations required substantial computational time. An attractive goal would be developing accurate analytic formulae, like the one derived by Matthews\u003csup\u003e26\u003c/sup\u003e for the treatment effect variance, to support faster calculation of attained power.\u003c/p\u003e\n\u003cp\u003eThis work considered only a continuous outcome within a cross-sectional SW-CRT. Further work is needed to extend the results to other types of outcomes (binary, count, survival, etc.) and to other SW-CRT designs.\u0026nbsp;\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn a stepped-wedge trial with unequal cluster sizes, the risk that randomization yields an allocation with inadequate attained power is a function of the ICC, the CV of the cluster sizes, and the total number of clusters. To reduce the computational burden of simulating attained power for every allocation, the attained power can be predicted from the treatment-vs-time period correlation and the treatment group imbalance via regression modeling. Trial designers can reduce the risk of low attained power by restricting the randomization algorithm to reduce the chance of obtaining an allocation which yields a large treatment-vs-time period correlation.\u003c/p\u003e"},{"header":"List of Abbreviations","content":"\u003cp\u003e\u003cstrong\u003eCRT: \u003c/strong\u003ecluster randomized trial\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCV:\u003c/strong\u003e coefficient of variation\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eICC:\u003c/strong\u003e intraclass correlation coefficient\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eP-CRT:\u003c/strong\u003e parallel-arm cluster randomized trial\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePD: \u003c/strong\u003e(pre-randomization) power distribution\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRMSPE:\u003c/strong\u003e\u0026nbsp; root mean square prediction error\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSW-CRT: \u003c/strong\u003estepped-wedge cluster randomized trial\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTGI:\u003c/strong\u003e treatment group imbalance\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTTC:\u003c/strong\u003e treatment-vs-time period correlation\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003e\u0026nbsp;\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eConsent for publication\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003e\u0026nbsp;\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eAvailability of data and materials \u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u0026nbsp; Simulation codes are available from: \u003ca href=\"https://github.com/douyangyd/swcrtpd\"\u003ehttps://github.com/douyangyd/swcrtpd\u003c/a\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003e\u0026nbsp;\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eCompeting interests\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no conflict of interest with this manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003e\u0026nbsp;\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eFunding\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported in part through funding from the BC SUPPORT Unit.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003e\u0026nbsp;\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eAuthor Contributions\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eHW conceived/developed the ideas of this paper and contributed to the critical revision of the manuscript. YO conducted all the computational work, drafted the manuscripts, developed the ideas and contributed to the critical revision of the manuscript. MEK and PG developed the ideas and contributed to the critical revision of the manuscript. TF contributed to the critical revision of the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003e\u0026nbsp;\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eAcknowledgement\u003c/em\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eThis research was enabled in part by support provided by WestGrid (\u003c/em\u003ewww.westgrid.ca\u003cem\u003e) and Compute Canada (\u003c/em\u003e\u003ca href=\"http://www.computecanada.ca\"\u003ewww.computecanada.ca\u003c/a\u003e\u003cem\u003e). \u003c/em\u003eMEK is supported in part by a Scholar Award from the Michael Smith Foundation for Health Research, partnered with Centre for Health\u003cbr /\u003e Evaluation and Outcome Sciences (ID#: 17661). MEK has also received consulting fees from Biogen and participated in Advisory Boards and/or Satellite Symposia of Biogen Inc. TF is supported by the Michael Smith Foundation for Health Research, the Vancouver Coastal Health Research Institute, and the Heart and Stroke Foundation of Canada. TF also receives in-kind study medication from Bayer Canada and received an honorarium for speaking from Servier. Also, we want to thank Liang Xu for sharing her code for efficiently generating all unique allocations.\u003c/p\u003e"},{"header":"Reference","content":"\u003col\u003e\n\u003cli\u003eHemming K, Haines TP, Chilton PJ, et al. The stepped wedge cluster randomised trial: rationale, design, analysis, and reporting. \u003cem\u003eBMJ\u003c/em\u003e 2015; 350: h391.\u003c/li\u003e\n\u003cli\u003eHussey MA, Hughes JP. Design and analysis of stepped wedge cluster randomized trials. \u003cem\u003eContemp Clin Trials\u003c/em\u003e 2007; 28: 182\u0026ndash;191.\u003c/li\u003e\n\u003cli\u003eBrown CA, Lilford RJ. The stepped wedge trial design: a systematic review. \u003cem\u003eBMC Medical Research Methodology\u003c/em\u003e 2006; 6: 54.\u003c/li\u003e\n\u003cli\u003eGrayling MJ, Wason JMS, Mander AP. Stepped wedge cluster randomized controlled trial designs: a review of reporting quality and design features. \u003cem\u003eTrials\u003c/em\u003e 2017; 18: 33.\u003c/li\u003e\n\u003cli\u003eDurovni B, Saraceni V, Moulton LH, et al. Effect of improved tuberculosis screening and isoniazid preventive therapy on incidence of tuberculosis and death in patients with HIV in clinics in Rio de Janeiro, Brazil: a stepped wedge, cluster-randomised trial. \u003cem\u003eLancet Infect Dis\u003c/em\u003e 2013; 13: 852\u0026ndash;858.\u003c/li\u003e\n\u003cli\u003eBacchieri G, Barros AJD, Santos JV dos, et al. A community intervention to prevent traffic accidents among bicycle commuters. \u003cem\u003eRev Saude Publica\u003c/em\u003e 2010; 44: 867\u0026ndash;875.\u003c/li\u003e\n\u003cli\u003eTirlea L, Truby H, Haines TP. Investigation of the effectiveness of the \u0026ldquo;Girls on the Go!\u0026rdquo; program for building self-esteem in young women: trial protocol. \u003cem\u003eSpringerplus\u003c/em\u003e; 2. Epub ahead of print 19 December 2013. DOI: 10.1186/2193-1801-2-683.\u003c/li\u003e\n\u003cli\u003eGruber JS, Reygadas F, Arnold BF, et al. A Stepped Wedge, Cluster-Randomized Trial of a Household UV-Disinfection and Safe Storage Drinking Water Intervention in Rural Baja California Sur, Mexico. \u003cem\u003eAm J Trop Med Hyg\u003c/em\u003e 2013; 89: 238\u0026ndash;245.\u003c/li\u003e\n\u003cli\u003eCopas AJ, Lewis JJ, Thompson JA, et al. Designing a stepped wedge trial: three main designs, carry-over effects and randomisation approaches. \u003cem\u003eTrials\u003c/em\u003e 2015; 16: 352.\u003c/li\u003e\n\u003cli\u003eBarker D, McElduff P, D\u0026rsquo;Este C, et al. Stepped wedge cluster randomised trials: a review of the statistical methodology used and available. \u003cem\u003eBMC Med Res Methodol\u003c/em\u003e 2016; 16: 69.\u003c/li\u003e\n\u003cli\u003eBaio G, Copas A, Ambler G, et al. Sample size calculation for a stepped wedge trial. \u003cem\u003eTrials\u003c/em\u003e 2015; 16: 354.\u003c/li\u003e\n\u003cli\u003eHemming K, Taljaard M. Sample size calculations for stepped wedge and cluster randomised trials: a unified approach. \u003cem\u003eJ Clin Epidemiol\u003c/em\u003e 2016; 69: 137\u0026ndash;146.\u003c/li\u003e\n\u003cli\u003eWoertman W, de Hoop E, Moerbeek M, et al. Stepped wedge designs could reduce the required sample size in cluster randomized trials. \u003cem\u003eJ Clin Epidemiol\u003c/em\u003e 2013; 66: 752\u0026ndash;758.\u003c/li\u003e\n\u003cli\u003eZhou X, Liao X, Spiegelman D. \u0026ldquo;Cross-sectional\u0026rdquo; stepped wedge designs always reduce the required sample size when there is no time effect. \u003cem\u003eJ Clin Epidemiol\u003c/em\u003e 2017; 83: 108\u0026ndash;109.\u003c/li\u003e\n\u003cli\u003eHughes J, Hakhu NR, Voldal E. swCRTdesign: stepped wedge cluster randomized trial (SW CRT) design, https://cran.r-project.org/web/packages/swCRTdesign/index.html (Accessed 20 Aug 2019).\u003c/li\u003e\n\u003cli\u003eBaio G. SWSamp: Simulation-based sample size calculations for a Stepped Wedge Trial (and more). \u003cem\u003eGianluca Baio\u003c/em\u003e/gianluca/software/swsamp/ (2018, accessed 13 May 2019).\u003c/li\u003e\n\u003cli\u003eHemming K, Girling A. A Menu-Driven Facility for Power and Detectable-Difference Calculations in Stepped-Wedge Cluster-Randomized Trials. \u003cem\u003eThe Stata Journal\u003c/em\u003e 2014; 14: 363\u0026ndash;380.\u003c/li\u003e\n\u003cli\u003eTeerenstra S, Taljaard M, Haenen A, et al. Sample size calculation for stepped-wedge cluster-randomized trials with more than two levels of clustering. \u003cem\u003eClinical Trials\u003c/em\u003e 2019; 16: 225\u0026ndash;236.\u003c/li\u003e\n\u003cli\u003eEldridge SM, Ashby D, Kerry S. Sample size for cluster randomized trials: effect of coefficient of variation of cluster size and analysis method. \u003cem\u003eInt J Epidemiol\u003c/em\u003e 2006; 35: 1292\u0026ndash;1300.\u003c/li\u003e\n\u003cli\u003eBreukelen GJP van, Candel MJJM. Efficiency loss because of varying cluster size in cluster randomized trials is smaller than literature suggests. \u003cem\u003eStatistics in Medicine\u003c/em\u003e 2012; 31: 397\u0026ndash;400.\u003c/li\u003e\n\u003cli\u003eKristunas CA, Smith KL, Gray LJ. An imbalance in cluster sizes does not lead to notable loss of power in cross-sectional, stepped-wedge cluster randomised trials with a continuous outcome. \u003cem\u003eTrials\u003c/em\u003e 2017; 18: 109.\u003c/li\u003e\n\u003cli\u003eGirling AJ. Relative efficiency of unequal cluster sizes in stepped wedge and other trial designs under longitudinal or cross-sectional sampling. \u003cem\u003eStat Med\u003c/em\u003e 2018; 37: 4652\u0026ndash;4664.\u003c/li\u003e\n\u003cli\u003eHarrison LJ, Chen T, Wang R. Power calculation for cross-sectional stepped wedge cluster randomized trials with variable cluster sizes. \u003cem\u003eBiometrics\u003c/em\u003e; n/a. DOI: 10.1111/biom.13164.\u003c/li\u003e\n\u003cli\u003eWong H, Ouyang Y, Karim ME. The randomization-induced risk of a trial failing to attain its target power: assessment and mitigation. \u003cem\u003eTrials\u003c/em\u003e 2019; 20: 360.\u003c/li\u003e\n\u003cli\u003eMartin JT, Hemming K, Girling A. The impact of varying cluster size in cross-sectional stepped-wedge cluster randomised trials. \u003cem\u003eBMC Med Res Methodol\u003c/em\u003e 2019; 19: 123.\u003c/li\u003e\n\u003cli\u003eMatthews JNS. Highly efficient stepped wedge designs for clusters of unequal size. \u003cem\u003eBiometrics\u003c/em\u003e; 2020. DOI: 10.1111/biom.13218.\u003c/li\u003e\n\u003cli\u003eClinicalTrials.gov[Internet] Ho K, University of British Columbia, 2018 Feb 20. Identifier NCT03439384, TEC4Home Heart Failure: Using Home Health Monitoring to Support the Transition of Care; 2020 Mar 24 [cited 2020 Apr 13]; [about 6 screens]. Available from https://clinicaltrials.gov/ct2/show/NCT03439384.\u003c/li\u003e\n\u003cli\u003eHasselman B. nleqslv: Solve Systems of Nonlinear Equations, https://CRAN.R-project.org/package=nleqslv (2017, accessed 12 November 2019).\u003c/li\u003e\n\u003cli\u003ePinheiro J, Bates D, DebRoy S, Sarkar D, R Core Team (2019).\u0026nbsp;\u003cem\u003enlme: Linear and Nonlinear Mixed Effects Models\u003c/em\u003e. R package version 3.1-142,\u0026nbsp;https://CRAN.R-project.org/package=nlme\u003c/li\u003e\n\u003cli\u003ePinheiro JC, Bates DM (eds). Theory and Computational Methods for Linear Mixed-Effects Models. In: \u003cem\u003eMixed-Effects Models in S and S-PLUS\u003c/em\u003e. New York, NY: Springer, pp. 57\u0026ndash;96.\u003c/li\u003e\n\u003cli\u003eR Development Core Team, R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria 2008. ISBN 3-900051-07-0, URL http://www.R-project.org. Accessed 15 Mar 2019.\u003c/li\u003e\n\u003cli\u003eCompute Canada Cedar - CC Doc, https://docs.computecanada.ca/wiki/Cedar (accessed 1 May 2019).\u003c/li\u003e\n\u003cli\u003eCanty A, Ripley B. boot: Bootstrap Functions (Originally by Angelo Canty for S), https://CRAN.R-project.org/package=boot (2019, accessed 28 March 2020).\u003c/li\u003e\n\u003cli\u003eHemming K, Taljaard M, McKenzie JE, et al. Reporting of stepped wedge cluster randomised trials: extension of the CONSORT 2010 statement with explanation and elaboration. \u003cem\u003eBMJ\u003c/em\u003e 2018; 363: k1614.\u003c/li\u003e\n\u003cli\u003eMatthews JNS, Forbes AB. Stepped wedge designs: insights from a design of experiments perspective. \u003cem\u003eStatistics in Medicine\u003c/em\u003e 2017; 36: 3772\u0026ndash;3790.\u003c/li\u003e\n\u003cli\u003eKasza J, Forbes AB. Information content of cluster\u0026ndash;period cells in stepped wedge trials. \u003cem\u003eBiometrics\u003c/em\u003e 2019; 75: 144\u0026ndash;152.\u003c/li\u003e\n\u003cli\u003eJohnson JL, Kreidler SM, Catellier DJ, et al. Recommendations for choosing an analysis method that controls Type I error for unbalanced cluster sample designs with Gaussian outcomes. \u003cem\u003eStatistics in Medicine\u003c/em\u003e 2015; 34: 3531\u0026ndash;3545.\u003c/li\u003e\n\u003cli\u003eTanner W. Improved Standard Error Estimation for Maintaining the Validities of Inference in Small-Sample Cluster Randomized Trials and Longitudinal Studies. \u003cem\u003eTheses and Dissertations--Epidemiology and Biostatistics\u003c/em\u003e. Epub ahead of print 1 January 2018. DOI: https://doi.org/10.13023/etd.2018.434.\u003c/li\u003e\n\u003cli\u003eRothwell JC, Julious SA, Cooper CL. A study of target effect sizes in randomised controlled trials published in the Health Technology Assessment journal. \u003cem\u003eTrials\u003c/em\u003e 2018; 19: 544.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"cluster randomized trial, power distribution, treatment-time period correlation, treatment group imbalance","lastPublishedDoi":"10.21203/rs.3.rs-15483/v2","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-15483/v2","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eBackground In a cross-sectional stepped-wedge trial with unequal cluster sizes, attained power in the trial depends on the realized allocation of the clusters. This attained power may differ from the expected power calculated using standard formulae by averaging the attained powers over all allocations the randomization algorithm can generate. We investigated the effect of design factors and allocation characteristics on attained power and developed models to predict attained power based on allocation characteristics.\u0026nbsp;\u003c/p\u003e\u003cp\u003eMethod Based on data simulated and analyzed using linear mixed-effects models, we evaluated the distribution of attained powers under different scenarios with varying intraclass correlation coefficient (ICC) of the responses, coefficient of variation (CV) of the cluster sizes, number of cluster-size groups, distributions of group sizes, and number of clusters. We explored the relationship between attained power and two allocation characteristics: the individual-level correlation between treatment status and time period, and the absolute treatment group imbalance. When computational time was excessive due to a scenario having a large number of possible allocations, we developed regression models to predict attained power using the treatment-vs-time period correlation and absolute treatment group imbalance as predictors.\u0026nbsp;\u003c/p\u003e\u003cp\u003eResults The risk of attained power falling more than 5% below the expected or nominal power decreased as the ICC or number of clusters increased and as the CV decreased. Attained power was strongly affected by the treatment-vs-time period correlation. The absolute treatment group imbalance had much less impact on attained power. The attained power for any allocation was predicted accurately using a logistic regression model with the treatment-vs-time period correlation and the absolute treatment group imbalance as predictors.\u0026nbsp;\u003c/p\u003e\u003cp\u003eConclusion In a stepped-wedge trial with unequal cluster sizes, the risk that randomization yields an allocation with inadequate attained power depends on the ICC, the CV of the cluster sizes, and number of clusters. To reduce the computational burden of simulating attained power for allocations, the attained power can be predicted via regression modeling. Trial designers can reduce the risk of low attained power by restricting the randomization algorithm to avoid allocations with large treatment-vs-time period correlations.\u003c/p\u003e","manuscriptTitle":"Explaining the variation in the attained power of a stepped-wedge trial with unequal cluster sizes","msid":"","msnumber":"","nonDraftVersions":[{"code":2,"date":"2020-05-07 01:44:59","doi":"10.21203/rs.3.rs-15483/v2","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2020-05-10T12:00:00+00:00","index":2,"fulltext":"Recommendation: Reviewer's comments unavailable pending editorial decision\n"},{"type":"editorInvitedReview","content":"","date":"2020-05-04T12:00:00+00:00","index":1,"fulltext":"Recommendation: Reviewer's comments unavailable pending editorial decision\n"},{"type":"reviewersInvited","content":"","date":"2020-04-24T12:00:00+00:00","index":"","fulltext":""},{"type":"reviewerAgreed","content":"","date":"2020-04-24T12:00:00+00:00","index":1,"fulltext":""},{"type":"reviewerAgreed","content":"","date":"2020-04-24T12:00:00+00:00","index":2,"fulltext":""},{"type":"editorAssigned","content":"","date":"2020-04-22T12:00:00+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2020-04-21T12:00:00+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2020-04-21T12:00:00+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}},{"code":1,"date":"2020-02-28 21:47:25","doi":"10.21203/rs.3.rs-15483/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2020-03-26T12:00:00+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2020-03-25T12:00:00+00:00","index":2,"fulltext":"Recommendation: Major revisions required\nForm responses:\n---\n\nComments to Author:\n---\nPEER REVIEWER ASSESSMENTS:\n\nOBJECTIVE - Full research articles: is there a clear objective that addresses a testable research question(s) (brief or other article types: is there a clear objective)?\nYes - there is a clear objective\n\nDESIGN - Is the current approach (including controls and analysis protocols) appropriate for the objective?\nYes - the approach is appropriate\n\nEXECUTION - Are the experiments and analyses performed with technical rigor to allow confidence in the results?\nYes - experiments and analyses were performed appropriately\n\nSTATISTICS - Is the use of statistics in the manuscript appropriate?\nNo - there are issues with the statistics in the study\n\nINTERPRETATION - Is the current interpretation/discussion of the results reasonable and not overstated?\nNo - there are minor issues\n\nOVERALL MANUSCRIPT POTENTIAL - Is the current version of this work technically sound? If not, can revisions be made to make the work technically sound?\nMaybe - with major revisions\n\nPEER REVIEWER COMMENTS:\n\nGENERAL COMMENTS:\nThis is an exhaustive simulation of different designs of stepped wedge cross-sectional trials. It is an immense amount of work with 108 designs each simulated up to 10,000 times.\n1) The paper is very dense, with numerous charts to illustrate the findings. One question I have is: does the reader need all these charts? The main point of the charts appears to be to illustrate some general principles. Why not consign most of the charts to the supplementary material, and just describe the findings in words?\n\nREQUESTED REVISIONS:\n1) Some of the results are already known, e.g., Girling (2018) found that 'In practice, the loss of precision due to unequal cluster sizes is unlikely to exceed 12%'. Mathews (2020) found a useful simplification of the variance and gave advice on design for unequal cluster sizes.\n2) I am not sure how useful an applied researcher will find this paper. Often the cluster size is under the control of the trialist. I think the authors could give examples when clusters are, of necessity, variable (e.g., when organisations of different sizes are being randomised).\n3) The paper is very dry, and it would have helped to have a practical example.\n4) There are several other papers (Matthews 2017, Kaska et al, 2019) which could have been mentioned, as they cover similar areas.\n\n\n\nRefs\nGirling A .Relative efficiency of unequal cluster sizes in stepped wedge and other trial designs under longitudinal or cross‐sectional sampling 2018 Statistics in Medicine 37, 4652-4664\nKasza J, Hemming K, Hooper R, Matthews JNS, Forbes A. Impact of non-uniform correlation structure on sample size and power in multiple-period cluster randomised trials. Statistical Methods in Medical Research 2019, 28(3), 703-716.\nMatthews JNS. Highly Efficient Stepped Wedge Designs for Clusters of Unequal Size. Biometrics 2020, epub ahead of print.\nMatthews, J.N.S. and Forbes, A.B. (2017) Stepped wedge designs: insights from a design of experiments perspective. Statistics in Medicine, 36, 3772-3790\n\n\nThus I would expect a revision to have considered the above comments\nADDITIONAL REQUESTS/SUGGESTIONS:\n1) Line 206: It would be helpful to point out that the standard Normal distribution has σ2=1, so the relationship in line 207 holds true.\n2) Line 209: Why choose a treatment effect of 0.26?\n3) Line 333: The model is not one I would have chosen. Power is a proportion, so one would expect a logistic model to fit better than a linear one.* Are the methods appropriate and well described?: **Yes**\n* Does the work include the necessary controls?: **Yes**\n* Are the conclusions drawn adequately supported by the data shown?: **No**\n* Are you able to assess any statistics in the manuscript or would you recommend an additional statistical review?: **I am able to assess the statistics**\n* Quality of written English: **Acceptable**\n* Declaration of competing interests: **This reviewer has been recruited by a partner organization, Research Square. Reviewers with declared or apparent competing interests are not utilized for these reviews. This reviewer has agreed to publication of their comments online under a Creative Commons Attribution License attributed to Research Square and was paid a small honorarium for completing the review within a specified timeframe. Honoraria for reviews such as this are paid regardless of the reviewer recommendation.**\n* I agree to the open peer review policy of the journal. I understand that my name will be included on my report to the authors and, if the manuscript is accepted for publication, my named report including any attachments I upload will be posted on the website along with the authors' responses. I agree for my report to be made available under an Open Access Creative Commons CC-BY license (http://creativecommons.org/licenses/by/4.0/). I understand that any comments which I do not wish to be included in my named report can be included as confidential comments to the editors, which will not be published.: ** I agree to the open peer review policy of the journal**\n"},{"type":"reviewerAgreed","content":"","date":"2020-03-15T12:00:00+00:00","index":2,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2020-02-11T12:00:00+00:00","index":1,"fulltext":"Recommendation: Major revisions required\nForm responses:\n---\n\nComments to Author:\n---\nThank you for the opportunity to review this manuscript. It is unfortunate that {Martin et al (2019) doi: 10.1186/s12874-019-0760-6} was just published (which I presume was not the case when the authors began their work) otherwise this manuscript would have been far more substantial in its contributions to the literature on unequal cluster sizes in SW-CRTs. Nonetheless, there are additional insights here that may well justify publication (in particular relating to the treatment-vs-time period correlation's influence on attained power). However, I have a number of comments regarding the articles content that I believe would need to be addressed first.\n\n[1] This may not in fact be the case, because aspects of the given description are difficult to follow (specifically L344-345 on P17), but it appears the authors have used simulated power values to train a prediction model and then evaluated the performance of this prediction model by comparing the distribution of its predictions to those of the original simulated power values? If this is the case, even if the predictions are based on other possible allocations it is no surprise that the boxplots given in Figure 10 are similar for each scenario (given the model fits well). What needs to be done (if this isn't the case) is the model should be used to predict the power for other allocations and the simulated powers *of these other allocations* should then be compared to *their* predictions - if these are similar it suggests 2000 random allocations is enough to build a useful model. This issue needs to be clearly resolved, and also links with points [2] and [3] below.\n\n[2] Throughout the paper there are various instances in which the language used could be made more precise. E.g., there are many instances of the type \"A increased as B did\" and \"C decreased with D\". It would be helpful if the authors could revise their manuscript to use actual figures/summary statistics wherever possible.\n\n[3] Similarly, there are many cases where the mathematical details given should be spelled out using formulae and not words. E.g., the treatment-vs-time period correlation is never given a formal definition in terms of the parameters of the assumed model and the procedure for choosing the best prediction model is entirely described through prose. Similarly, no definition of the CV or the ICC is given. I suggest that the authors modify the article to include more equations.\n\n[4] In the Introduction, when the authors describe previous work on SW-CRTs with unequal cluster sizes, they miss two important papers that it would be useful to discuss {Girling (2018) doi: 10.1002/sim.7943; Matthews (2020) doi: 10.1111/biom.13218}.\n\n[5] In relation to the fact that \"the risk of obtaining low power decreases as the ICC increased\", I think this is implied by the results of {Matthews and Forbes (2017) doi: 10.1002/sim.7403}, who discuss the contributions of vertical and horizontal comparisons as the ICC varies. This should be noted. Similarly, I think the comments around when the large clusters switch needs to make reference to the results from {Kasza and Forbes (2019) doi: 10.1111/biom.12959}.\n\n[6] It would be more logical in my opinion to begin the Methods section with a description of the model and then describe the scenarios that are considered - it is a little confusing to have various model parameters discussed (e.g., ICC) before the model (c.f. [2] above).\n\n[7] More justification should be given to the design parameters. E.g., why consider 12/24/48 clusters? Why assume an effect size of 0.26? Are these in any way based on reality in SW-CRTs?\n\n[8] It wasn't entirely clear to me what variance is assumed for the error terms epsilon_ijk? A justification for why beta can be chosen \"without loss of generality\" as 1 also needs to be added.\n\n[9] Authors should make their code available, as is BMC Med Res Methodol practice [\"…should include a link to the most recent version of your software or code (e.g. GitHub or Sourceforge)…\"].\n\n[10] Keywords should be reserved for additional phrases that don't appear in the title, e.g. \"cluster randomised trial\", to assist future literature searches.\n\n[11] On L66 of P4, clusters aren't \"enrolling into\" the intervention group in the final period - they are in the intervention condition/receiving the intervention. Use of the word 'enrolling' suggests clusters are still being recruited at this point.\n\n[12] A better way to describe the types of SW-CRT would be to follow {Copas et al (2015) doi: 10.1186/s13063-015-0842-7}, who describe three categories, rather than listing just cohort and cross-sectional designs as is currently done in the introduction.\n\n[13] In the introduction, the authors list a few reasons for using a SW-CRT, but don't mention the fact that SW-CRTs are commonly used when there is a desire to provide the intervention to everyone. Given this has emerged as the most commonly cited reason for using a SW-CRT [as backed up by {Grayling et al (2017) doi: 10.1186/s13063-017-1783-0}], I would add it in.\n\n[14] The authors state there is code available for SAS, R, and Stata, but I am unaware of any code for SAS (none is listed here for example - https://clusterrandomisedtrials.qmul.ac.uk/) and they don't cite the relevant code for Stata {Hemming and Girling (2014) doi: 10.1177/1536867X1401400208}.\n\n[15] It would be helpful to more precisely describe the method used for assessing significance (c.f. L417) - in particular nlme::lme() assumes a specific form for the degrees of freedom that should be highlighted to the reader for non-R experts.\n\n[16] In relation to identifying tests that provide accurate error control (c.f. L422) some relevant uncited work is {Tanner (2018) Improved standard error estimation for maintaining the validities of inference in small-sample cluster randomized trials and longitudinal studies. PhD Thesis, University of Kentucky}.\n\n[17] Duplication of the word \"of\" on L400.\n\n[18] The list of abbreviations is missing several words e.g., TTC and TGI.* Are the methods appropriate and well described?: **Yes**\n* Does the work include the necessary controls?: **Yes**\n* Are the conclusions drawn adequately supported by the data shown?: **Yes**\n* Are you able to assess any statistics in the manuscript or would you recommend an additional statistical review?: **I am able to assess the statistics**\n* Quality of written English: **Acceptable**\n* Declaration of competing interests: **I declare that I have no competing interests**\n* I agree to the open peer review policy of the journal. I understand that my name will be included on my report to the authors and, if the manuscript is accepted for publication, my named report including any attachments I upload will be posted on the website along with the authors' responses. I agree for my report to be made available under an Open Access Creative Commons CC-BY license (http://creativecommons.org/licenses/by/4.0/). I understand that any comments which I do not wish to be included in my named report can be included as confidential comments to the editors, which will not be published.: ** I agree to the open peer review policy of the journal**\n"},{"type":"reviewerAgreed","content":"","date":"2020-01-29T12:00:00+00:00","index":1,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2020-01-16T12:00:00+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2019-12-12T12:00:00+00:00","index":"","fulltext":""},{"type":"submitted","content":"","date":"2019-12-11T12:00:00+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2019-12-11T12:00:00+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2019-12-11T12:00:00+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2106f21f-6def-48f3-8421-ea69781fbcfe","owner":[],"postedDate":"May 7th, 2020","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":95167,"name":"Health Economics \u0026 Outcomes Research"},{"id":95168,"name":"Biostatistics"}],"tags":[],"updatedAt":"2020-06-28T16:57:05+00:00","versionOfRecord":{"articleIdentity":"rs-15483","link":"https://doi.org/10.1186/s12874-020-01036-5","journal":{"identity":"bmc-medical-research-methodology","isVorOnly":false,"title":"BMC Medical Research Methodology"},"publishedOn":"2020-06-24 12:00:00","publishedOnDateReadable":"June 24th, 2020"},"versionCreatedAt":"2020-05-07 01:44:59","video":"","vorDoi":"10.1186/s12874-020-01036-5","vorDoiUrl":"https://doi.org/10.1186/s12874-020-01036-5","workflowStages":[]},"version":"v2","identity":"rs-15483","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-15483","identity":"rs-15483","version":["v2"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.