Employing Multiple Synchronous Outcome Samples Per Subject to Improve Study Efficiency. | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Software Employing Multiple Synchronous Outcome Samples Per Subject to Improve Study Efficiency. Roger A'Hern This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-358007/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 12 You are reading this latest preprint version Abstract Background: Accuracy can be improved by taking multiple synchronous samples from each subject in a study to estimate the endpoint of interest if sample values are not highly correlated. If feasible, it is useful to assess the value of this cluster approach when planning studies. Multiple assessments may be the only method to increase power to an acceptable level if the number of subjects is limited. Methods: The main aim is to estimate the difference in outcome between groups of subjects by taking one or more synchronous primary outcome samples or measurements. A summary statistic from multiple samples per subject will typically have a lower sampling error. The number of subjects can be balanced against the number of synchronous samples to minimize the sampling error, subject to design constraints. This approach can include estimating the optimum number of samples given the cost per subject and the cost per sample. Results: The accuracy improvement achieved by taking multiple samples depends on the intra-class correlation (ICC). The lower the ICC, the greater the benefit that can accrue. If the ICC is high, then a second sample will provide little additional information about the subject's true value. If the ICC is very low, adding a sample can be equivalent to adding an extra subject. Benefits of multiple samples include the ability to reduce the number of subjects in a study and increase both the power and the available alpha. If, for example, the ICC is 35%, adding a second measurement can be equivalent to adding 48% more subjects to a single measurement study. Conclusion: A study's design can sometimes be improved by taking multiple synchronous samples. It is useful to evaluate this strategy as an extension of a single sample design. An Excel workbook is provided to allow researchers to explore the most appropriate number of samples to take in a given setting. Health Economics & Outcomes Research Clinical Trials Sample Size Multiple Synchronous Samples Cluster Design Figures Figure 1 Figure 2 Figure 3 Background In some circumstances, it is possible to undertake more than one synchronous assessment of the same subject to estimate a measure of interest. The overall result might then be calculated as the average across the assessments or perhaps as the maximum or minimum value if they are more critical. It is natural to ask whether it is worth making multiple measurements to improve the assessment quality. This has to be weighed against any disadvantages - an assessment may be burdensome to the subject or clinician/scientist undertaking it or more costly. This paper offers a quantitative framework, including easy-to-use software, for making this decision, developed from first principles to outline its basis. The overall variance is the sum of the between-subject and within-subject variance, so reducing the effect of within-subject variability by performing repeat within-subject observations can be beneficial. As might be anticipated, this strategy is most useful when the within-subject variation is a large proportion of the overall variation, equivalent to a low to moderate intraclass correlation (ICC) between observations. Improving study design by taking repeated measurements may be particularly valuable if the number of subjects eligible for the study is limited, for example, in rare disease types or restricted subject groups. The degree of correlation between repeated measurements is important. For example, if the ICC were 85%, taking more than one sample would offer little additional information about the subject’s actual value. If the correlation is 10% or lower, multiple samples might be considered if there were reasonable grounds for this degree of independence. However, in other circumstances, it may be concluded the parameter being measured was of little value because of low repeatability. As a practical example, many human organs are bilateral, and if a sample were taken from each, these would show correlated results because of their shared genotypic and environmental background. Is it worth taking measurements from both organs to study organ function? The eye provides a useful illustration. Glynn and Rosner [ 1 ] presented alternative models for predicting the percent of normal visual field in 197 subjects being followed up for glaucoma. Inclusion of both eyes (N=394) in a mixed-effects analysis, which allows for correlation, was found to improve the accuracy of the estimates of potential predictive factors compared with the use of only one eye. The correlation between the eyes was 0.55 (55%). A median reduction in the standard error of the parameter estimates in the 394 ‘eye’ based analysis relative to the single eye analysis was 15% (for the factor hypertension) with a range of 12% (for gender) to 39% (for acuity loss). Reductions of this order are worthwhile, a 15% reduction in the standard error could, for example, increase the power to detect a real difference from 66% to 80%, and a 39% reduction could increase the power from 40% to 80%. From another viewpoint, employing a summary value from several measurements rather than the value of a single measure can be seen as improving the ICC by lowering the within-subject variation. For example, Lee et al. [ 2 ] studied the ICC's of 3 representative physical examinations - popliteal angle, Thomas test, and Staheli test, performed twice each by three orthopedic surgeons on 30 cerebral palsy subjects with a mean age of 12.5 years. The popliteal angle test is a measure of hamstring tightness; both the Thomas test and the Staheli test are methods of measuring hip flexion contracture. The single test ICC's were 0.71, 0.46, and 0.22. However, the ICC's of the averages of the three assessments were 0.88, 0.74, and 0.46, respectively, suggesting the popliteal test can be improved and a more satisfactory ICC for the Thomas test can be achieved if they are applied independently multiple times. Examples of specific contexts in which this technique might be used will now be considered. A sample size calculation based on the above Glynn and Rosner example suggests that a standardized difference of 0.4 in a binary predictive factor could be detected with 80% power using 197 single eyes; this decreases to a standardized difference of 0.25 if both eyes (N=394) are studied. Nicholson and Holmes [ 3 ] noted that a popular but improper method for assessing high-throughput assays' precision is by scatter-plotting data. This consists of equally dividing a sample and assaying the two halves separately, then plotting and correlating all analytes' results in the first half versus the second half. They concluded that precision should not be based on all analytes' plots. However, the repeatability of individual analytes and a variance inflation factor should be used to calculate appropriate sample sizes to detect changes in specific analyte levels that are the focus of the researcher's interest. The biased scatter-plotting method typically gives ‘excellent’ correlations of 0.95 or greater, but for four high throughput assays Nicholson and Holmes reported ICCs of 0.31 (0.10-0.53) [Median (IQ range)] for 1624 microRNA analytes, 0.59 (0.24-0.80) for 17,788 mRNA analytes, 0.31 (0.20-0.50) for 69 proteins and 0.94 (0.82-0.96) for 163 metabolites. This suggests ICCs for some analytes in high-throughput assays are at levels that may make repeat samples worthwhile. Ionan et al. [ 4 ] described the National Cancer Institute's Director's Challenge reproducibility study results. This examined the reproducibility of 22,283 features from the Affymetrix U133A Genechip across a collection of eleven frozen patient tissue samples. These were assayed at four different labs. Fifty-percent of the 22,283 ICC's were below 0.52, and 25% were below 0.23. The majority of ICC values derived from adjustment factors calculated in a study of UK Biobank data by Morgan et al. [ 5 ] were observed to be above 70%, but not all. The following are below 70% : Diastolic blood pressure (60% (95%CI:60-62)); Systolic blood pressure (65% (95%CI: 64-65)); Pulse rate (62% (95%CI:61-64)); Peak expiratory flow (60%(95%CI:59-61)) and Grip strength (65% (95%CI:63-67)). All hematological factors in this study and a meta-analysis by Coskuna et al. [ 6 ] had ICCs that were 70% or greater. More complex situations can arise if multiple measures with differing numbers of measures per subject are used to calculate an endpoint. In studies of advanced cancer, in which subjects may have cancer present at multiple sites, the within-participant sum of tumor lesion diameters is used for overall tumor response calculation in RECIST [ 7 ]. This sum will decrease if there is a positive response to therapy. Caution was initially exercised; before 2009, the recommendation was to measure up to 10 lesions per subject but subsequently [ 8 ], a maximum of 5 was found to be sufficient, supporting the correlation of the responses in different lesions from the same subject. The potential value of synchronous measurements is also supported by recognizing that slope (linear trend) estimation can be optimized by making as many observations as feasible at the extreme ends of the ranges of independent variables. For example, to estimate a linear relationship (decrease or increase) within a subject over two years by making six measurements, three measures at baseline and three at 24m would yield a more accurate estimate of change than spacing the measurements, such as one each at baseline, 4m, 8m, 12m, 18m and 24m. In this paper, subjects are considered the basic experimental unit of interest, and single or multiple assessments, measurements, or samples are taken from these subjects. The methodology is not new but is identical to that of cluster randomized trials. However, the focus is on samples within subjects rather than subjects within clusters. The software presented can also be used for cluster designs by considering the cluster units as subjects. For simplicity, the main focus is on situations where it is possible to average measures across repeated samples. Implementation The variance of the mean of several identically normally distributed random variables can be calculated by noting that for two variables, the variance Var(aX + bY) = a 2 Var(X) + b 2 Var(Y) + 2abCovar(X,Y). If X and Y have mean (X + Y)/2, ρ is their correlation and Var(X) = Var(Y) = s 2 , then Covar(X,Y) is ρs 2 and a = b = 1/2. Hence Var(Mean) = s 2 /4 + s2/4 + 2ρs 2 /4 = (s 2 /2)(1 + ρ). By extension, it can be shown that the variance of the mean of m samples is Var(Mean)=(s 2 /m)(1+(m-1)ρ). In the absence of correlation, the variance would be s 2 /m; hence the quantity (m-1)ρ is the variance increase due to the correlation, and (1+(m-1)ρ) is known as the Variance Inflation Factor (VIF). If m is not consistent across the subjects, it is typically replaced by the mean value across subjects; however, the coefficient of variation (CV) of m can also be incorporated to estimate the variance (see below). These formulae mirror those used in cluster randomized trials, programs performing sample size calculation for cluster randomized trials can therefore be used in the current context. A fundamental relationship, shown in Figs. 1 (A) and (B), is worth noting. If subjects have normally distributed average values with variance V subj , and samples have values normally distributed about these averages with variance V samp , then the intraclass correlation (ICC or ρ) is V subj /(V subj +V samp ). Figure 1 (A) shows this diagrammatically for a randomly simulated sample with V subj =1 and V samp =0.25, a correlation of 0.8. In this simulated data, the subject means have been defined to vary randomly about zero, and there are 20 subjects each with ten assessments. This correlation is also apparent when plotting within-subject observations, plotting pairs of values from the same subject from this simulated data (Fig. 1 (B)). The correlation in this context will typically be close to but not the same as the ICC. An Excel Workbook (STARS.xlsx) is provided, which includes worksheets to illustrate sample size calculation incorporating the number of measures, the ICC, and the CV of the number of measures, including details on the optimum ‘cost’ based choice of the number of samples. Sample size worksheets for estimating ICC are also included, as is are examples demonstrating techniques of analysis and associated commands in the programming language R. Results The relationship between sample size and ICC The measurement variance is composed of between-subject variability and within-subject variability (figure 1A). Hypothetically, if the overall variance is considered constant and the within-subject variance decreases, then the between-subject variance must increase proportionately (figure 2). The correlation (x-axis, figure 2) is where the specific between/within variance relationship occurs for the measure. The correlation (ICC) is the proportion of the overall variation attributable to between-subject differences, calculated as the between-subject variance divided by the overall variance (set to 1 here). The relationship between correlation and the % division of the two sources of variance is the centered ‘X’ in the plot. The sample size chosen for a clinical trial or other between-group study comparison is directly related to the variation of the outcome. For example, setting out to detect a 1 unit increase between two equal-sized groups and considering outcome variances of 1, 2, and 4, appropriate calculated sample sizes might be 146, 292, and 584 subjects, respectively; these also have a 1:2:4 ratio. In figure 2, the relative numbers of study subjects required for different correlations with m=2 and m=4 are shown as the upper diagonal line starting with 50% and 25% of subjects, respectively. The pattern shown in figure 2 is also shown in table 1, with alternative designs for specific correlations being shown in columns. %Measures/% Subjects Intra-Class Correlation Measures 0 0.15 0.35 0.5 0.65 0.85 0.9 1 1 100 / 100 100 / 100 100 / 100 100 / 100 100 / 100 100 / 100 100 / 100 100 / 100 2 100 / 50 116 / 58 136 / 68 150 / 75 166 / 83 186 / 93 190 / 95 200 / 100 3 99 / 33 129 / 43 171 / 57 201 / 67 231 / 77 270 / 90 279 / 93 300 / 100 4 100 / 25 144 / 36 204 / 51 252 / 63 296 / 74 356 / 89 372 / 93 400 / 100 5 100 / 20 160 / 32 240 / 48 300 / 60 360 / 72 440 / 88 460 / 92 500 / 100 Additional % Subjects required for same design characteristics with one measure Measures 1 0 0 0 0 0 0 0 0 2 100 74 48 33 21 8 5 0 3 200 131 76 50 30 11 7 0 4 300 176 95 60 36 13 8 0 5 400 213 108 67 39 14 9 0 Table 1. The relationship between the proportion of measures and subjects versus intra-class correlation relative to the proportion needed when only one measure is employed (top rows). The lower table shows the effective percentage increase in study efficiency (in terms of subjects) that can be achieved by taking multiple samples. For example, a two-sample study with a correlation of 0.5 is equivalent to a one-sample study with 33% more patients (as is also apparent from the first table in the 75 to 100 difference). For a correlation of 0.5, 100 samples could be taken from 100 patients (with m=1), or 201 samples could be taken from 67 subjects (with m=3). If the recommended study size with one sample were 160 patients, the corresponding figures for m=3 would be 1.6 times the values shown (324 samples in 108 patients). Some scenarios are unlikely to ever be of practical value, such as those with correlations of 0.85 or greater because of the small efficiency improvement. A similar table, with 5% correlation increments, is given in the Excel workbook (STARS.xlsx, ‘Introduction’ worksheet), together with an example of data analysis, with associated R code (STARS.xlsx, ‘Analysis Examples’ worksheet). The Excel workbook (STARS.xlsx), which accompanies this paper, contains worksheets that allow calculation of sample sizes for continuous normal and binomial outcomes, as well as for estimating intraclass correlation. These worksheets can be accessed via the initial ‘Contents’ worksheet. STARS stands for ‘Sample size calculations for Two-group comparisons with Repeated Synchronous sampling.’ All formulae employed can be found on associated ‘Calculations’ worksheets; this format has the advantage of transparency and ease of further development by interested researchers. The worksheets are protected to prevent inappropriate changes, but can be unprotected by using the supplied password. Four features apparent from the use of this program will now be discussed. Increasing power and the available alpha If the number of subjects available for a study is approximately known, then because the standard error of the endpoint is reduced by taking multiple samples, increasing the number of measures can be used to increase the power of the study or increase the amount of alpha available. The first four panels of figure 3 show power improvements possible for four different correlations. The final two panels illustrate changes in available alpha. Subjects (such as patients) might be willing to donate more samples or provide more assessments in a study if a benefit was that there were more interim futility or efficacy analyses, or greater power. As a simple example, if it is decided approximately 200 subjects could be entered into a trial with a single measurement that used an alpha error rate of P=5% for the primary comparison, taking two measurements with a correlation of ρ=0.5 would imply P=1.54% could be used for this comparison. The remaining 3.46% could be used for other purposes such as interim analyses, and the overall type I error rate of 5% would be maintained. Unequal number of samples per subject There may be circumstances in which there is an unequal number of samples per subject, for example, for logistic reasons or because of subject preference. If the variation in the number of samples is small (CV not greater than 0.23 [9]) the average number of samples per subject can be employed for sample size calculations; in other circumstances, an adjustment should be used. The variation can be summarised by its coefficient of variation (the CV is the SD of m divided by mean), and a correction based on this can be employed (Rutherford, Copas, and Eldridge [ 9 ]). If the CV of the number of samples is not known at the outset of a study, the study size could be adapted to allow for the observed variation, estimated from an early analysis. Summary Endpoints for Serial Measurements A linear trend across time is an example of a single measure that can be used to summarise serial within-subject measurements. As mentioned above, linear trends can be estimated more accurately by clustering measurements at the extreme range of independent factors. Matthews et al. [ 10 ] provides an informative introduction to summary endpoints. Simulation can be used to compare the accuracy of estimated effects using strategies to summarise multiple synchronous measurements. Simulation was used to mimic follow-up over two years to identify subjects with rapid visual field progression (-2 dB/year) by Crabb and Garway-Heath [ 11 ] ; these showed measurement either 2 or 3 times was superior to every six months or every four months. The ‘Latanoprost for open-angle glaucoma (UKGTS)’ trial [ 12 ] accordingly incorporated this approach. The methodology described in this paper can also be used to approximate the number of subjects required if a single overall outcome calculated from true serial measurements is being considered, in which each component measure is weighted equally. It should also be realistic to assume that the relationship between the measurements can be characterized by an overall representative single correlation. This approach could be employed to approximate sample sizes for studies based on more complex correlation matrices, these sample sizes can be refined using simulation employing programs such as Superpower [ 13 ]. Optimising Study Design on the basis of ‘Cost’ The number of subjects varies with the number of measures employed, so if a cost is assigned to each, then the overall cost can be compared. For example, if the cost per subject is 60 and the cost per sample is 10 (and using a specific design*), the costs for up to eight measures are: 1 - 12,180; 2 - 9,120; 3 – 8,640; 4 – 8,400; 5 – 8,580; 6 – 9,000; 7 – 9,360 and 8 – 9,660. This suggests the optimum number of measures is 4, but that for 3 is similar. Cost units could be arbitrary; for example, the assigned cost could represent a currency or an alternative such as a linear score combining cost to the subject in terms of inconvenience and risk, the cost to the staff undertaking the procedure, and financial cost. This result is similar to that obtained from the formula suggested for the optimum number of patients in the clusters of cluster randomized trials (m=√(c/s x (1- ρ)/ρ) [ 14 ], which yields m=3.74, where c is a cost per subject and s the cost per measure. This aspect of study design is also included in the Excel Workbook (STARS.xlsx, Sample Size calculation worksheets). *Difference to detect=0.5, SD=1, ρ=0.3, CV=0.5, Power=0.85, alpha=0.05, k=2. Discussion If there is an opportunity to repeat assessments in a study, it is useful to quantify the benefit of this strategy. There is a danger that the use of multiple assessments is dismissed too readily, perhaps merely on the basis that assessments will be correlated, without thoroughly evaluating the value of adding further samples or measures. The additional burden of extra samples needs to be considered. If a medical study is being prospectively defined, which requires assessments over and above those of standard care, public/participant involvement (PPI) could be employed to investigate participants' views on the provision of extra samples balanced against the benefits this could yield in design. These could include a smaller overall study size, shorter trial duration, increased power, or more interim analyses. In studies using laboratory animals, smaller experiments may be desirable to reduce the number of animals required, particularly if they have to be sacrificed. Note also that if an outcome is of interest, high within-subject variability does not necessarily preclude a study from being undertaken if it is possible to take multiple samples. A small, intensive study of the value of an outcome with high within-subject variability may sometimes be useful to evaluate whether it is worth refining measurement of the outcome to reduce within-subject variability. The cost of some samples or measures may be reduced with time as more efficient methods of obtaining and analyzing them are developed, making it more practical to obtain multiple samples. It may therefore be useful to review decisions when sample costs decrease. It is important to note that the scale of measurement can be critical when measuring ICCs. Repeatability is frequently assessed by plotting the values of two measurements on the same subject against each other as a scatterplot with a line of equality. However, the variability is more easily understood by plotting the difference in a subject’s measurements from the two methods against the mean of the measurements, known as a Bland–Altman plot [ 15 ]. These plots illustrate measurement error alongside the necessary ‘limits of agreement’, which give a range within which 95% of future differences in measurements would be expected to lie, the latter being calculated from the mean and SD of the paired differences [ 16 ]. However, this method assumes the SD is the same throughout the measurement range. It is common for the SD to increase with the mean, the coefficient of variation (CV) rather than the standard deviation often quoted to summarise variability for such measurements. This suggests the measurements have a lognormal rather than a normal distribution and a remedy that is frequently successful is to use the logarithm of the two measurements for analysis [16]. A further consideration is that studies that focus on subgroups for precision medicine may have a lower ICC than an unselected group of subjects, making repeated assessments more relevant. If a prognostic factor is used to select a subgroup with a more limited outcome range, then the between-subject variance of this range would be expected to be lower than that in all subjects. However, the within-subject variance might be expected to be similar, lowering the ICC. Focussing on subgroups may therefore change the relevance of multiple measurements per subject. Given the potentially high variability seen in high throughput assays referred to in the introduction, it is interesting to note that individual results of components of high throughput assays are sometimes aggregated to produce a single overall score to represent an underlying phenomenon of interest. This may offer a way to improve study efficiency without making more measurements because the averaging across several components could even out the effect of errors in the individual components. For example, a multigene assay score that is used to predict recurrence in Breast Cancer [ 17 ] has a proliferation (tumor growth) component that is composed of the expression of five genes (Ki67, STK15, Survivin, CCNB1, and MYBL2) combined by averaging the five gene scores. It would be expected that components will be correlated if such an approach is used. The need for a representative sample may override the desire to reduce the number of subjects by making multiple measurements. The majority of studies aim to obtain a typical sample of the population of subjects being examined to ensure the study results are generalizable. For example, a study of 40 subjects might not be considered large enough to represent the diversity seen in the population the study was chosen to represent. However, a study of 100 subjects may be considered more appropriate. Note that systematic differences between repeats do not necessarily invalidate the use of repeat samples. It may be possible to adjust the repeat measurements statistically to quantitatively remove such differences; this has the effect of increasing the ICC. Conclusion It may be beneficial to undertake multiple synchronous observations per subject in some circumstances. This option is part of the toolset available to researchers when planning effective studies. Both between-subject and within-subject variability are critical parameters for decision-making in this context. An Excel workbook is provided to aid exploration of the statistical background of this feature of study design. Abbreviations CV Coefficient of Variation ICC Intra-Class Correlation Coefficient STARS Sample size calculations for Two-group comparisons with Repeated Synchronous sampling Declarations Ethics approval and consent to participate Not Applicable. Consent for publication None required, apart from single author. Availability of data and materials One item, an Excel Workbook, submitted. It contains previously published anonymous data. Competing interests None. Funding None. Authors' contributions This is a single author submission. Acknowledgements None required. References Glynn RJ, Rosner B. Accounting for the correlation between fellow eyes in regression analysis. Arch Ophthalmol. 1992 Mar; 110(3):381-7. DOI: 10.1001/archopht.1992.01080150079033 Lee KM, Lee J, Chin Youb Chung CY et al. Pitfalls and Important Issues in Testing Reliability Using Intraclass Correlation Coefficients in Orthopaedic Research Clinics in Orthopedic Surgery 2012;4:149-155. http://dx.doi.org/10.4055/cios.2012.4.2.149 Nicholson G, Chris Holmes C. A note on statistical repeatability and study design for high‐throughput assays. Statistics in Medicine 2017;36 5:790-798. Ionan AC, Polley M-YC, McShane LM, Dobbin KK. Comparison of confidence interval methods for an intra-class correlation coefficient (ICC). BMC Medical Research Methodology 2014, 14:121 http://www.biomedcentral.com/1471-2288/14/121 Morgan KA, Cook S Leon DA, Frost C. Reflection on modern methods: calculating a sample size for a repeatability sub-study to correct for measurement error in a single continuous exposure. International Journal of Epidemiology 2019;48 5:1721-1726. Coskuna A, Bragab F, Carobenea A et al. Systematic review and meta-analysis of within-subject and between-subject biological variation estimates of 20 haematological parameters. Clin Chem Lab Med 2020; 58 1: 25–32. Therasse P, Arbuck SG, Eisenhauer EA, et al. New guidelines to evaluate the response to treatment in solid tumors. J Natl Cancer Inst 2000;92 3:205-16. doi: 10.1093/jnci/92.3.205. Bogaerts J, Ford R, Dan Sargent D et al. Individual patient data analysis to assess modifications to the RECIST criteria. Eur J Cancer. 2009 Jan;45 2:248-60. doi: 10.1016/j.ejca.2008.10.027. Rutterford C, Copas A, Eldridge S. Methods for sample size determination in cluster randomized trials. International Journal of Epidemiology 2015, 1051–1067. doi: 10.1093/ije/dyv113 Matthews JNS, Altman DG, Campbell MJ, Royston P. Analysis of serial measurements in medical research, British Medical Journal 1990; 300: 230-235. doi: 10.1136/bmj.300.6719.230. Crabb DP, Garway-Heath DF. Intervals between visual field tests when monitoring the glaucomatous patient: wait-and-see approach. Invest Ophthalmol Vis Sci 2012; 53 : 2770–76. doi: https://doi.org/10.1167/iovs.12-9476 Garway-Heath DF,Crabb DP,Bunce C et al. Latanoprost for open-angle glaucoma (UKGTS): a randomised, multicentre, placebo-controlled trial. Lancet 2015; 385: 1295-1304. Doi: https://doi,org/10.1016/S0140-6736(14)6211 Caldwell A, Laken D (2020). https://github.com/arcaldwell49/Superpower van Breukelen GJP, Candel MJJM. Calculating sample sizes for cluster randomized trials: we can keep it simple and efficient!J Clin Epidemiol. 2012; 65 11:1212-8. doi: 10.1016/j.jclinepi.2012.06.002 Bland JM, Altman DG. Statistical methods for assessing agreement between two methods of clinical measurement. Lancet 1986;327 8476: 307-310. https://doi.org/10.1016/S0140-6736(86)90837-8 Bartlett, JW, Frost C. Reliability, repeatability and reproducibility: analysis of measurement errors in continuous variables. Ultrasound Obstet Gynecol 2008; 31: 466-475 . Paik S, Shak S, Tang G et al. A multigene assay to predict recurrence of tamoxifen-treated, node-negative breast cancer. New England Journal of Medicine 2004; 351(27):2817-2826. DOI: 10.1056/nejmoa041588 Supplementary Files STARSV1.0.xlsx Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Major revision 15 Apr, 2021 Review # 1 received at journal 14 Apr, 2021 Review # 3 received at journal 30 Mar, 2021 Review # 2 received at journal 30 Mar, 2021 Reviewer # 3 agreed at journal 28 Mar, 2021 Reviewer # 2 agreed at journal 24 Mar, 2021 Editor assigned by journal 21 Mar, 2021 Reviewers invited by journal 21 Mar, 2021 Reviewer # 1 agreed at journal 21 Mar, 2021 Submission checks completed at journal 21 Mar, 2021 Editor invited by journal 21 Mar, 2021 First submitted to journal 14 Mar, 2021 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-358007","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Software","associatedPublications":[],"authors":[{"id":18234100,"identity":"a0126d40-1676-4bef-8cf5-58001e24cf35","order_by":0,"name":"Roger A'Hern","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA4UlEQVRIiWNgGAWjYBCDBCBmfJDAYAFlE6mF2SCBQYI0LWxA9URoMW9vv/i4gsEuj1/67LGKBzUScgbHExg/fMzBrUXmzJliwzMMycWSfXlpNxKOSRgbnHnALDlzG24tEhI5aZINDMyJG87wmN1IbJBI3HAjgY2ZF7+W9J8NDPWJ+4FaCojUkn6MsYHhcOIGHh4zBuK08JxhlmwwOJ444wyPsQTIL5JnHjbj9wt7+8OPDRXVif09PIYff9TYyPEdTz744SMeLQwMPAYMDAYoIkCH4gfsDwgoGAWjYBSMghEPAJYMTglelAYEAAAAAElFTkSuQmCC","orcid":"","institution":"NA","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Roger","middleName":"","lastName":"A'Hern","suffix":""}],"badges":[],"createdAt":"2021-03-24 11:36:23","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-358007/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-358007/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":7349608,"identity":"d5a477b6-ffaa-41f1-806c-298776bf9157","added_by":"auto","created_at":"2021-03-25 13:50:01","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":43654,"visible":true,"origin":"","legend":"(A) and (B) illustrate the joint effects of between and within subject variation and their contribution to correlation. Please see text for more detail.","description":"","filename":"Figure1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-358007/v1/23c3db9895284f05e15ea8e6.jpg"},{"id":7349886,"identity":"5ea9a056-e7a8-4acc-8554-1abb7087f5fd","added_by":"auto","created_at":"2021-03-25 13:53:01","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":70199,"visible":true,"origin":"","legend":"A schematic illustrating the relationship of overall variability, between-subject variability, within-subject variability, intra-class correlation (ICC), number of synchronous samples (m), and study size. For convenience, overall variability is set to 1, between-subject variability is the increasing diagonal (left to right), and within-subject variability the decreasing diagonal. The intra-class correlation (ICC) is on the x-axis and summarises the relationship between subject and within-subject variability. Study size is statistically linearly related to the two sources of variability and correlation and in addition to the number of synchronous samples (m).","description":"","filename":"Figure2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-358007/v1/7a7a039fda594ec77c8e59f5.jpg"},{"id":7349887,"identity":"839f9e25-91a0-42e7-9135-0701f05f8328","added_by":"auto","created_at":"2021-03-25 13:53:01","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":139194,"visible":true,"origin":"","legend":"Improvements in power and alpha can be achieved by increasing the number of measurements and their relationship to the correlation. For power (top four panels), a two-sided significance level of 5% has been used (and a CV of zero, see below). Critical alpha values (bottom two panels, power 85%) can be lowered by making multiple measurements (see text for explanation). The programs for generating these relationships are available in the Excel Web appendix (STARS.xls, under ‘Other’).","description":"","filename":"Figure3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-358007/v1/7fd8d48f7899e3174bbb483d.jpg"},{"id":13682405,"identity":"343cb0e2-d466-4cfe-b2e5-f5ec1904179a","added_by":"auto","created_at":"2021-09-17 11:57:08","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":482165,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-358007/v1/a1d3b9df-d152-4ae3-ba27-91ba31951182.pdf"},{"id":7349888,"identity":"9fe2e081-1bfb-4851-b2e0-e6ba02afcc4b","added_by":"auto","created_at":"2021-03-25 13:53:02","extension":"xlsx","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":5878622,"visible":true,"origin":"","legend":"","description":"","filename":"STARSV1.0.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-358007/v1/94978cc4b29a642597acc7da.xlsx"}],"financialInterests":"","formattedTitle":"\u003cp\u003eEmploying Multiple Synchronous Outcome Samples Per Subject to Improve Study Efficiency.\u003c/p\u003e","fulltext":[{"header":"Background","content":"\u003cp\u003eIn some circumstances, it is possible to undertake more than one synchronous assessment of the same subject to estimate a measure of interest. The overall result might then be calculated as the average across the assessments or perhaps as the maximum or minimum value if they are more critical. It is natural to ask whether it is worth making multiple measurements to improve the assessment quality. This has to be weighed against any disadvantages - an assessment may be burdensome to the subject or clinician/scientist undertaking it or more costly. This paper offers a quantitative framework, including easy-to-use software, for making this decision, developed from first principles to outline its basis. The overall variance is the sum of the between-subject and within-subject variance, so reducing the effect of within-subject variability by performing repeat within-subject observations can be beneficial. As might be anticipated, this strategy is most useful when the within-subject variation is a large proportion of the overall variation, equivalent to a low to moderate intraclass correlation (ICC) between observations.\u003c/p\u003e\n\u003cp\u003eImproving study design by taking repeated measurements may be particularly valuable if the number of subjects eligible for the study is limited, for example, in rare disease types or restricted subject groups. The degree of correlation between repeated measurements is important. For example, if the ICC were 85%, taking more than one sample would offer little additional information about the subject\u0026rsquo;s actual value.\u003c/p\u003e\n\u003cp\u003eIf the correlation is 10% or lower, multiple samples might be considered if there were reasonable grounds for this degree of independence. However, in other circumstances, it may be concluded the parameter being measured was of little value because of low repeatability.\u003c/p\u003e\n\u003cp\u003eAs a practical example, many human organs are bilateral, and if a sample were taken from each, these would show correlated results because of their shared genotypic and environmental background. Is it worth taking measurements from both organs to study organ function? The eye provides a useful illustration. Glynn and Rosner [\u003csup\u003e1\u003c/sup\u003e] presented alternative models for predicting the percent of normal visual field in 197 subjects being followed up for glaucoma. Inclusion of both eyes (N=394) in a mixed-effects analysis, which allows for correlation, was found to improve the accuracy of the estimates of potential predictive factors compared with the use of only one eye. The correlation between the eyes was 0.55 (55%). A median reduction in the standard error of the parameter estimates in the 394 \u0026lsquo;eye\u0026rsquo; based analysis relative to the single eye analysis was 15% (for the factor hypertension) with a range of 12% (for gender) to 39% (for acuity loss). Reductions of this order are worthwhile, a 15% reduction in the standard error could, for example, increase the power to detect a real difference from 66% to 80%, and a 39% reduction could increase the power from 40% to 80%.\u003c/p\u003e\n\u003cp\u003eFrom another viewpoint, employing a summary value from several measurements rather than the value of a single measure can be seen as improving the ICC by lowering the within-subject variation. For example, Lee et al. [\u003csup\u003e2\u003c/sup\u003e] studied the ICC's of 3 representative physical examinations - popliteal angle, Thomas test, and Staheli test, performed twice each by three orthopedic surgeons on 30 cerebral palsy subjects with a mean age of 12.5 years. The popliteal angle test is a measure of hamstring tightness; both the Thomas test and the Staheli test are methods of measuring hip flexion contracture. The single test ICC's were 0.71, 0.46, and 0.22. However, the ICC's of the averages of the three assessments were 0.88, 0.74, and 0.46, respectively, suggesting the popliteal test can be improved and a more satisfactory ICC for the Thomas test can be achieved if they are applied independently multiple times.\u003c/p\u003e\n\u003cp\u003eExamples of specific contexts in which this technique might be used will now be considered. A sample size calculation based on the above Glynn and Rosner example suggests that a standardized difference of 0.4 in a binary predictive factor could be detected with 80% power using 197 single eyes; this decreases to a standardized difference of 0.25 if both eyes (N=394) are studied. Nicholson and Holmes [\u003csup\u003e3\u003c/sup\u003e] noted that a popular but improper method for assessing high-throughput assays' precision is by scatter-plotting data. This consists of equally dividing a sample and assaying the two halves separately, then plotting and correlating all analytes' results in the first half versus the second half. They concluded that precision should not be based on all analytes' plots. However, the repeatability of individual analytes and a variance inflation factor should be used to calculate appropriate sample sizes to detect changes in specific analyte levels that are the focus of the researcher's interest. The biased scatter-plotting method typically gives \u0026lsquo;excellent\u0026rsquo; correlations of 0.95 or greater, but for four high throughput assays Nicholson and Holmes reported ICCs of 0.31 (0.10-0.53) [Median (IQ range)] for 1624 microRNA analytes, 0.59 (0.24-0.80) for 17,788 mRNA analytes, 0.31 (0.20-0.50) for 69 proteins and 0.94 (0.82-0.96) for 163 metabolites. This suggests ICCs for some analytes in high-throughput assays are at levels that may make repeat samples worthwhile. Ionan et al. [\u003csup\u003e4\u003c/sup\u003e] described the National Cancer Institute's Director's Challenge reproducibility study results. This examined the reproducibility of 22,283 features from the Affymetrix U133A Genechip across a collection of eleven frozen patient tissue samples. These were assayed at four different labs. Fifty-percent of the 22,283 ICC's were below 0.52, and 25% were below 0.23.\u003c/p\u003e\n\u003cp\u003eThe majority of ICC values derived from adjustment factors calculated in a study of UK Biobank data by Morgan et al. [\u003csup\u003e5\u003c/sup\u003e] were observed to be above 70%, but not all. The following are below 70% : Diastolic blood pressure (60% (95%CI:60-62)); Systolic blood pressure (65% (95%CI: 64-65)); Pulse rate (62% (95%CI:61-64)); Peak expiratory flow (60%(95%CI:59-61)) and Grip strength (65% (95%CI:63-67)). All hematological factors in this study and a meta-analysis by Coskuna et al. [\u003csup\u003e6\u003c/sup\u003e] had ICCs that were 70% or greater. More complex situations can arise if multiple measures with differing numbers of measures per subject are used to calculate an endpoint. In studies of advanced cancer, in which subjects may have cancer present at multiple sites, the within-participant sum of tumor lesion diameters is used for overall tumor response calculation in RECIST [\u003csup\u003e7\u003c/sup\u003e]. This sum will decrease if there is a positive response to therapy. Caution was initially exercised; before 2009, the recommendation was to measure up to 10 lesions per subject but subsequently [\u003csup\u003e8\u003c/sup\u003e], a maximum of 5 was found to be sufficient, supporting the correlation of the responses in different lesions from the same subject.\u003c/p\u003e\n\u003cp\u003eThe potential value of synchronous measurements is also supported by recognizing that slope (linear trend) estimation can be optimized by making as many observations as feasible at the extreme ends of the ranges of independent variables. For example, to estimate a linear relationship (decrease or increase) within a subject over two years by making six measurements, three measures at baseline and three at 24m would yield a more accurate estimate of change than spacing the measurements, such as one each at baseline, 4m, 8m, 12m, 18m and 24m.\u003c/p\u003e\n\u003cp\u003eIn this paper, subjects are considered the basic experimental unit of interest, and single or multiple assessments, measurements, or samples are taken from these subjects. The methodology is not new but is identical to that of cluster randomized trials. However, the focus is on samples within subjects rather than subjects within clusters. The software presented can also be used for cluster designs by considering the cluster units as subjects. For simplicity, the main focus is on situations where it is possible to average measures across repeated samples.\u003c/p\u003e"},{"header":"Implementation","content":"\u003cp\u003eThe variance of the mean of several identically normally distributed random variables can be calculated by noting that for two variables, the variance Var(aX\u0026thinsp;+\u0026thinsp;bY)\u0026thinsp;=\u0026thinsp;a\u003csup\u003e2\u003c/sup\u003eVar(X)\u0026thinsp;+\u0026thinsp;b\u003csup\u003e2\u003c/sup\u003eVar(Y)\u0026thinsp;+\u0026thinsp;2abCovar(X,Y). If X and Y have mean (X\u0026thinsp;+\u0026thinsp;Y)/2, \u0026rho; is their correlation and Var(X)\u0026thinsp;=\u0026thinsp;Var(Y)\u0026thinsp;=\u0026thinsp;s\u003csup\u003e2\u003c/sup\u003e, then Covar(X,Y) is \u0026rho;s\u003csup\u003e2\u003c/sup\u003e and a\u0026thinsp;=\u0026thinsp;b\u0026thinsp;=\u0026thinsp;1/2. Hence Var(Mean)\u0026thinsp;=\u0026thinsp;s\u003csup\u003e2\u003c/sup\u003e/4\u0026thinsp;+\u0026thinsp;s2/4\u0026thinsp;+\u0026thinsp;2\u0026rho;s\u003csup\u003e2\u003c/sup\u003e/4 = (s\u003csup\u003e2\u003c/sup\u003e/2)(1\u0026thinsp;+\u0026thinsp;\u0026rho;). By extension, it can be shown that the variance of the mean of m samples is Var(Mean)=(s\u003csup\u003e2\u003c/sup\u003e/m)(1+(m-1)\u0026rho;). In the absence of correlation, the variance would be s\u003csup\u003e2\u003c/sup\u003e/m; hence the quantity (m-1)\u0026rho; is the variance increase due to the correlation, and (1+(m-1)\u0026rho;) is known as the Variance Inflation Factor (VIF). If m is not consistent across the subjects, it is typically replaced by the mean value across subjects; however, the coefficient of variation (CV) of m can also be incorporated to estimate the variance (see below). These formulae mirror those used in cluster randomized trials, programs performing sample size calculation for cluster randomized trials can therefore be used in the current context.\u003c/p\u003e\n\u003cp\u003eA fundamental relationship, shown in Figs.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e (A) and (B), is worth noting. If subjects have normally distributed average values with variance V\u003csub\u003esubj\u003c/sub\u003e, and samples have values normally distributed about these averages with variance V\u003csub\u003esamp\u003c/sub\u003e, then the intraclass correlation (ICC or \u0026rho;) is V\u003csub\u003esubj\u003c/sub\u003e/(V\u003csub\u003esubj\u003c/sub\u003e+V\u003csub\u003esamp\u003c/sub\u003e). Figure\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e (A) shows this diagrammatically for a randomly simulated sample with V\u003csub\u003esubj\u003c/sub\u003e=1 and V\u003csub\u003esamp\u003c/sub\u003e=0.25, a correlation of 0.8. In this simulated data, the subject means have been defined to vary randomly about zero, and there are 20 subjects each with ten assessments. This correlation is also apparent when plotting within-subject observations, plotting pairs of values from the same subject from this simulated data (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e (B)). The correlation in this context will typically be close to but not the same as the ICC.\u003c/p\u003e\n\u003cp\u003eAn Excel Workbook (STARS.xlsx) is provided, which includes worksheets to illustrate sample size calculation incorporating the number of measures, the ICC, and the CV of the number of measures, including details on the optimum \u0026lsquo;cost\u0026rsquo; based choice of the number of samples. Sample size worksheets for estimating ICC are also included, as is are examples demonstrating techniques of analysis and associated commands in the programming language R.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eThe relationship between sample size and ICC\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe measurement variance is composed of between-subject variability and within-subject variability (figure 1A). Hypothetically, if the overall variance is considered constant and the within-subject variance decreases, then the between-subject variance must increase proportionately (figure 2). The correlation (x-axis, figure 2) is where the specific between/within variance relationship occurs for the measure. The correlation (ICC) is the proportion of the overall variation attributable to between-subject differences, calculated as the between-subject variance divided by the overall variance (set to 1 here). The relationship between correlation and the % division of the two sources of variance is the centered \u0026lsquo;X\u0026rsquo; in the plot.\u003c/p\u003e\n\u003cp\u003eThe sample size chosen for a clinical trial or other between-group study comparison is directly related to the variation of the outcome. For example, setting out to detect a 1 unit increase between two equal-sized groups and considering outcome variances of 1, 2, and 4, appropriate calculated sample sizes might be 146, 292, and 584 subjects, respectively; these also have a 1:2:4 ratio. In figure 2, the relative numbers of study subjects required for different correlations with m=2 and m=4 are shown as the upper diagonal line starting with 50% and 25% of subjects, respectively. The pattern shown in figure 2 is also shown in table 1, with alternative designs for specific correlations being shown in columns.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e%Measures/% Subjects\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eIntra-Class Correlation\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"1\" width=\"0\"\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd width=\"66\"\u003e\n\u003cp\u003e\u0026nbsp;\u003cstrong\u003e\u003cu\u003eMeasures\u003c/u\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0.15\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0.35\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0.5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0.65\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0.85\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0.9\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"66\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e100 / 100\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e100 / 100\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e100 / 100\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e100 / 100\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e100 / 100\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e100 / 100\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e100 / 100\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e100 / 100\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"66\"\u003e\n\u003cp\u003e2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e100 / 50\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e116 / 58\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e136 / 68\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e150 / 75\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e166 / 83\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e186 / 93\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e190 / 95\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e200 / 100\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"66\"\u003e\n\u003cp\u003e3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e99 / 33\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e129 / 43\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e171 / 57\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e201 / 67\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e231 / 77\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e270 / 90\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e279 / 93\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e300 / 100\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"66\"\u003e\n\u003cp\u003e4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e100 / 25\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e144 / 36\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e204 / 51\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e252 / 63\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e296 / 74\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e356 / 89\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e372 / 93\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e400 / 100\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"66\"\u003e\n\u003cp\u003e5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e100 / 20\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e160 / 32\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e240 / 48\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e300 / 60\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e360 / 72\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e440 / 88\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e460 / 92\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e500 / 100\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAdditional % Subjects required for same design characteristics with one measure\u003c/strong\u003e\u003c/p\u003e\n\u003ctable border=\"1\" width=\"0\"\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cu\u003eMeasures\u003c/u\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e100\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e74\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e48\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e33\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e21\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e200\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e131\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e76\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e50\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e30\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e11\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e7\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e300\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e176\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e95\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e60\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e36\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e13\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e400\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e213\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e108\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e67\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e39\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e14\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e9\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"65\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTable 1. The relationship between the proportion of measures and subjects versus intra-class correlation relative to the proportion needed when only one measure is employed (top rows). The lower table shows the effective percentage increase in study efficiency (in terms of subjects) that can be achieved by taking multiple samples. For example, a two-sample study with a correlation of 0.5 is equivalent to a one-sample study with 33% more patients (as is also apparent from the first table in the 75 to 100 difference).\u003c/p\u003e\n\u003cp\u003eFor a correlation of 0.5, 100 samples could be taken from 100 patients (with m=1), or 201 samples could be taken from 67 subjects (with m=3). If the recommended study size with one sample were 160 patients, the corresponding figures for m=3 would be 1.6 times the values shown (324 samples in 108 patients). Some scenarios are unlikely to ever be of practical value, such as those with correlations of 0.85 or greater because of the small efficiency improvement. A similar table, with 5% correlation increments, is given in the Excel workbook (STARS.xlsx, \u0026lsquo;Introduction\u0026rsquo; worksheet), together with an example of data analysis, with associated R code (STARS.xlsx, \u0026lsquo;Analysis Examples\u0026rsquo; worksheet).\u003c/p\u003e\n\u003cp\u003eThe Excel workbook (STARS.xlsx), which accompanies this paper, contains worksheets that allow calculation of sample sizes for continuous normal and binomial outcomes, as well as for estimating intraclass correlation. These worksheets can be accessed via the initial \u0026lsquo;Contents\u0026rsquo; worksheet. STARS stands for \u0026lsquo;Sample size calculations for Two-group comparisons with Repeated Synchronous sampling.\u0026rsquo; All formulae employed can be found on associated \u0026lsquo;Calculations\u0026rsquo; worksheets; this format has the advantage of transparency and ease of further development by interested researchers. The worksheets are protected to prevent inappropriate changes, but can be unprotected by using the supplied password. Four features apparent from the use of this program will now be discussed.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eIncreasing power and the available alpha\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIf the number of subjects available for a study is approximately known, then because the standard error of the endpoint is reduced by taking multiple samples, increasing the number of measures can be used to increase the power of the study or increase the amount of alpha available. The first four panels of figure 3 show power improvements possible for four different correlations. The final two panels illustrate changes in available alpha. Subjects (such as patients) might be willing to donate more samples or provide more assessments in a study if a benefit was that there were more interim futility or efficacy analyses, or greater power.\u003c/p\u003e\n\u003cp\u003eAs a simple example, if it is decided approximately 200 subjects could be entered into a trial with a single measurement that used an alpha error rate of P=5% for the primary comparison, taking two measurements with a correlation of \u0026rho;=0.5 would imply P=1.54% could be used for this comparison. The remaining 3.46% could be used for other purposes such as interim analyses, and the overall type I error rate of 5% would be maintained.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eUnequal number of samples per subject\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThere may be circumstances in which there is an unequal number of samples per subject, for example, for logistic reasons or because of subject preference. If the variation in the number of samples is small (CV not greater than 0.23 [9]) the average number of samples per subject can be employed for sample size calculations; in other circumstances, an adjustment should be used. The variation can be summarised by its coefficient of variation (the CV is the SD of m divided by mean), and a correction based on this can be employed (Rutherford, Copas, and Eldridge [\u003csup\u003e9\u003c/sup\u003e]). If the CV of the number of samples is not known at the outset of a study, the study size could be adapted to allow for the observed variation, estimated from an early analysis.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSummary Endpoints for Serial Measurements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA linear trend across time is an example of a single measure that can be used to summarise serial within-subject measurements. As mentioned above, linear trends can be estimated more accurately by clustering measurements at the extreme range of independent factors. Matthews et al. [\u003csup\u003e10\u003c/sup\u003e] provides an informative introduction to summary endpoints. Simulation can be used to compare the accuracy of estimated effects using strategies to summarise multiple synchronous measurements. Simulation was used to mimic follow-up over two years to identify subjects with rapid visual field progression (-2 dB/year) by Crabb and Garway-Heath [\u003csup\u003e11\u003c/sup\u003e] ; these showed measurement either 2 or 3 times was superior to every six months or every four months. The \u0026lsquo;Latanoprost for open-angle glaucoma (UKGTS)\u0026rsquo; trial [\u003csup\u003e12\u003c/sup\u003e] accordingly incorporated this approach.\u003c/p\u003e\n\u003cp\u003eThe methodology described in this paper can also be used to approximate the number of subjects required if a single overall outcome calculated from true serial measurements is being considered, in which each component measure is weighted equally. It should also be realistic to assume that the relationship between the measurements can be characterized by an overall representative single correlation. This approach could be employed to approximate sample sizes for studies based on more complex correlation matrices, these sample sizes can be refined using simulation employing programs such as Superpower [\u003csup\u003e13\u003c/sup\u003e].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eOptimising Study Design on the basis of \u0026lsquo;Cost\u0026rsquo; \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe number of subjects varies with the number of measures employed, so if a cost is assigned to each, then the overall cost can be compared. For example, if the cost per subject is 60 and the cost per sample is 10 (and using a specific design*), the costs for up to eight measures are:\u0026nbsp;\u003cstrong\u003e1\u003c/strong\u003e\u0026nbsp;- 12,180;\u0026nbsp;\u003cstrong\u003e2\u003c/strong\u003e\u0026nbsp;- 9,120;\u0026nbsp;\u003cstrong\u003e3\u003c/strong\u003e\u0026nbsp;\u0026ndash; 8,640;\u0026nbsp;\u003cstrong\u003e4\u003c/strong\u003e\u0026nbsp;\u0026ndash; 8,400;\u003cstrong\u003e\u0026nbsp;5\u003c/strong\u003e\u0026nbsp;\u0026ndash; 8,580;\u0026nbsp;\u003cstrong\u003e6\u003c/strong\u003e\u0026nbsp;\u0026ndash; 9,000;\u0026nbsp;\u003cstrong\u003e7\u003c/strong\u003e\u0026nbsp;\u0026ndash; 9,360 and\u0026nbsp;\u003cstrong\u003e8\u003c/strong\u003e\u0026nbsp;\u0026ndash; 9,660. This suggests the optimum number of measures is 4, but that for 3 is similar. Cost units could be arbitrary; for example, the assigned cost could represent a currency or an alternative such as a linear score combining cost to the subject in terms of inconvenience and risk, the cost to the staff undertaking the procedure, and financial cost. This result is similar to that obtained from the formula suggested for the optimum number of patients in the clusters of cluster randomized trials (m=\u0026radic;(c/s x (1- \u0026rho;)/\u0026rho;) [\u003csup\u003e14\u003c/sup\u003e], which yields m=3.74, where c is a cost per subject and s the cost per measure. This aspect of study design is also included in the Excel Workbook (STARS.xlsx, Sample Size calculation worksheets).\u003c/p\u003e\n\u003cp\u003e*Difference to detect=0.5, SD=1, \u0026rho;=0.3, CV=0.5, Power=0.85, alpha=0.05, k=2.\u0026nbsp;\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eIf there is an opportunity to repeat assessments in a study, it is useful to quantify the benefit of this strategy. There is a danger that the use of multiple assessments is dismissed too readily, perhaps merely on the basis that assessments will be correlated, without thoroughly evaluating the value of adding further samples or measures. The additional burden of extra samples needs to be considered. If a medical study is being prospectively defined, which requires assessments over and above those of standard care, public/participant involvement (PPI) could be employed to investigate participants' views on the provision of extra samples balanced against the benefits this could yield in design. These could include a smaller overall study size, shorter trial duration, increased power, or more interim analyses. In studies using laboratory animals, smaller experiments may be desirable to reduce the number of animals required, particularly if they have to be sacrificed. Note also that if an outcome is of interest, high within-subject variability does not necessarily preclude a study from being undertaken if it is possible to take multiple samples. A small, intensive study of the value of an outcome with high within-subject variability may sometimes be useful to evaluate whether it is worth refining measurement of the outcome to reduce within-subject variability. The cost of some samples or measures may be reduced with time as more efficient methods of obtaining and analyzing them are developed, making it more practical to obtain multiple samples. It may therefore be useful to review decisions when sample costs decrease.\u003c/p\u003e\n\u003cp\u003eIt is important to note that the scale of measurement can be critical when measuring ICCs. Repeatability is frequently assessed by plotting the values of two measurements on the same subject against each other as a scatterplot with a line of equality. However, the variability is more easily understood by plotting the difference in a subject\u0026rsquo;s measurements from the two methods against the mean of the measurements, known as a Bland\u0026ndash;Altman plot [\u003csup\u003e15\u003c/sup\u003e]. These plots illustrate measurement error alongside the necessary \u0026lsquo;limits of agreement\u0026rsquo;, which give a range within which 95% of future differences in measurements would be expected to lie, the latter being calculated from the mean and SD of the paired differences [\u003csup\u003e16\u003c/sup\u003e]. However, this method assumes the SD is the same throughout the measurement range. It is common for the SD to increase with the mean, the coefficient of variation (CV) rather than the standard deviation often quoted to summarise variability for such measurements. This suggests the measurements have a lognormal rather than a normal distribution and a remedy that is frequently successful is to use the logarithm of the two measurements for analysis [16].\u003c/p\u003e\n\u003cp\u003eA further consideration is that studies that focus on subgroups for precision medicine may have a lower ICC than an unselected group of subjects, making repeated assessments more relevant. If a prognostic factor is used to select a subgroup with a more limited outcome range, then the between-subject variance of this range would be expected to be lower than that in all subjects. However, the within-subject variance might be expected to be similar, lowering the ICC. Focussing on subgroups may therefore change the relevance of multiple measurements per subject.\u003c/p\u003e\n\u003cp\u003eGiven the potentially high variability seen in high throughput assays referred to in the introduction, it is interesting to note that individual results of components of high throughput assays are sometimes aggregated to produce a single overall score to represent an underlying phenomenon of interest. This may offer a way to improve study efficiency without making more measurements because the averaging across several components could even out the effect of errors in the individual components. For example, a multigene assay score that is used to predict recurrence in Breast Cancer [\u003csup\u003e17\u003c/sup\u003e] has a proliferation (tumor growth) component that is composed of the expression of five genes (Ki67, STK15, Survivin, CCNB1, and MYBL2) combined by averaging the five gene scores. It would be expected that components will be correlated if such an approach is used.\u003c/p\u003e\n\u003cp\u003eThe need for a representative sample may override the desire to reduce the number of subjects by making multiple measurements. The majority of studies aim to obtain a typical sample of the population of subjects being examined to ensure the study results are generalizable. For example, a study of 40 subjects might not be considered large enough to represent the diversity seen in the population the study was chosen to represent. However, a study of 100 subjects may be considered more appropriate.\u003c/p\u003e\n\u003cp\u003eNote that systematic differences between repeats do not necessarily invalidate the use of repeat samples. It may be possible to adjust the repeat measurements statistically to quantitatively remove such differences; this has the effect of increasing the ICC.\u003c/p\u003e"},{"header":"Conclusion","content":" \u003cp\u003eIt may be beneficial to undertake multiple synchronous observations per subject in some circumstances. This option is part of the toolset available to researchers when planning effective studies. Both between-subject and within-subject variability are critical parameters for decision-making in this context. An Excel workbook is provided to aid exploration of the statistical background of this feature of study design.\u003c/p\u003e "},{"header":"Abbreviations","content":"\u003cp\u003eCV Coefficient of Variation\u003c/p\u003e\n\u003cp\u003eICC Intra-Class Correlation Coefficient\u003c/p\u003e\n\u003cp\u003eSTARS Sample size calculations for Two-group comparisons with Repeated Synchronous sampling\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot Applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNone required, apart from single author.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eOne item, an Excel Workbook, submitted. It contains previously published anonymous data.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNone.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNone.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors' contributions \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis is a single author submission.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNone required.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eGlynn\u0026nbsp;RJ,\u0026nbsp;\u003cem\u003eRosner\u003c/em\u003e\u0026nbsp;B.\u0026nbsp;\u003cem\u003eAccounting\u003c/em\u003e\u0026nbsp;for the\u0026nbsp;\u003cstrong\u003ecorrelation\u003c/strong\u003e\u0026nbsp;between fellow eyes in regression analysis. Arch Ophthalmol. 1992 Mar; 110(3):381-7. DOI:\u0026nbsp;10.1001/archopht.1992.01080150079033\u003c/li\u003e\n\u003cli\u003eLee KM, Lee J, Chin Youb Chung CY et al. Pitfalls and Important Issues in Testing Reliability Using Intraclass Correlation Coefficients in Orthopaedic Research Clinics in Orthopedic Surgery 2012;4:149-155. \u003ca href=\"http://dx.doi.org/10.4055/cios.2012.4.2.149\"\u003ehttp://dx.doi.org/10.4055/cios.2012.4.2.149\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003eNicholson G, Chris Holmes C. A note on statistical repeatability and study design for high‐throughput assays. Statistics in Medicine 2017;36 5:790-798.\u003c/li\u003e\n\u003cli\u003eIonan AC, Polley M-YC, McShane LM, Dobbin KK. Comparison of confidence interval methods for an intra-class correlation coefficient (ICC). BMC Medical Research Methodology 2014, 14:121\u0026nbsp;\u003ca href=\"http://www.biomedcentral.com/1471-2288/14/121\"\u003ehttp://www.biomedcentral.com/1471-2288/14/121\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003eMorgan KA, Cook S Leon DA, Frost C. Reflection on modern methods: calculating a sample size for a repeatability sub-study to correct for measurement error in a single continuous exposure. International Journal of Epidemiology 2019;48 5:1721-1726.\u003c/li\u003e\n\u003cli\u003eCoskuna A, Bragab F, Carobenea A et al. Systematic review and meta-analysis of within-subject and between-subject biological variation estimates of 20 haematological parameters. Clin Chem Lab Med 2020; 58 1: 25\u0026ndash;32.\u003c/li\u003e\n\u003cli\u003eTherasse P, Arbuck SG, Eisenhauer EA, et al. New guidelines to evaluate the response to treatment in solid tumors. J Natl Cancer Inst 2000;92 3:205-16. doi: 10.1093/jnci/92.3.205.\u003c/li\u003e\n\u003cli\u003eBogaerts J, Ford R, Dan Sargent D et al. Individual patient data analysis to assess modifications to the RECIST criteria. Eur J Cancer. 2009 Jan;45 2:248-60. doi: 10.1016/j.ejca.2008.10.027.\u003c/li\u003e\n\u003cli\u003eRutterford C, Copas A, Eldridge S. Methods for sample size determination in cluster randomized trials. International Journal of Epidemiology 2015, 1051\u0026ndash;1067. doi: 10.1093/ije/dyv113\u003c/li\u003e\n\u003cli\u003eMatthews JNS, Altman DG, Campbell MJ, Royston P. Analysis of serial measurements in medical research, British Medical Journal 1990; 300: 230-235. doi: 10.1136/bmj.300.6719.230.\u003c/li\u003e\n\u003cli\u003eCrabb DP, Garway-Heath DF. Intervals between visual field tests when monitoring the glaucomatous patient: wait-and-see approach. Invest Ophthalmol Vis Sci 2012; 53\u003cstrong\u003e: \u003c/strong\u003e2770\u0026ndash;76. doi:\u003ca href=\"https://doi.org/10.1167/iovs.12-9476\"\u003ehttps://doi.org/10.1167/iovs.12-9476\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003eGarway-Heath DF,Crabb DP,Bunce C et al. Latanoprost for open-angle glaucoma (UKGTS): a randomised, multicentre, placebo-controlled trial. Lancet 2015; 385: 1295-1304. Doi: \u003ca href=\"https://doi,org/10.1016/S0140-6736(14)6211\"\u003ehttps://doi,org/10.1016/S0140-6736(14)6211\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003eCaldwell A, Laken D (2020). https://github.com/arcaldwell49/Superpower\u003c/li\u003e\n\u003cli\u003evan Breukelen GJP, Candel MJJM. Calculating sample sizes for cluster randomized trials: we can keep it simple and efficient!J Clin Epidemiol. 2012; 65 11:1212-8. doi: 10.1016/j.jclinepi.2012.06.002\u003c/li\u003e\n\u003cli\u003eBland JM, Altman DG. Statistical methods for assessing agreement between two methods of clinical measurement. Lancet 1986;327 8476: 307-310. \u003ca href=\"https://doi.org/10.1016/S0140-6736(86)90837-8\"\u003ehttps://doi.org/10.1016/S0140-6736(86)90837-8\u003c/a\u003e\u003c/li\u003e\n\u003cli\u003eBartlett, JW, Frost C. Reliability, repeatability and reproducibility: analysis of measurement errors in continuous variables. Ultrasound Obstet Gynecol 2008; 31: 466-475\u003cem\u003e.\u003c/em\u003e\u003c/li\u003e\n\u003cli\u003ePaik S, Shak S, Tang G et al. A multigene assay to predict recurrence of tamoxifen-treated, node-negative breast cancer. New England Journal of Medicine 2004; 351(27):2817-2826. DOI: 10.1056/nejmoa041588\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Clinical Trials, Sample Size, Multiple Synchronous Samples, Cluster Design","lastPublishedDoi":"10.21203/rs.3.rs-358007/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-358007/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground:\u003c/strong\u003e\u0026nbsp;Accuracy can be improved by taking multiple synchronous samples from each subject in a study to estimate the endpoint of interest if sample values are not highly correlated. If feasible, it is useful to assess the value of this cluster approach when planning studies. Multiple assessments may be the only method to increase power to an acceptable level if the number of subjects is limited.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eMethods:\u003c/strong\u003e\u0026nbsp;The main aim is to estimate the difference in outcome between groups of subjects by taking one or more synchronous primary outcome samples or measurements. A summary statistic from multiple samples per subject will typically have a lower sampling error. The number of subjects can be balanced against the number of synchronous samples to minimize the sampling error, subject to design constraints. This approach can include estimating the optimum number of samples given the cost per subject and the cost per sample.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eResults:\u0026nbsp;\u003c/strong\u003eThe accuracy improvement achieved by taking multiple samples depends on the intra-class correlation (ICC). The lower the ICC, the greater the benefit that can accrue. If the ICC is high, then a second sample will provide little additional information about the subject's true value. If the ICC is very low, adding a sample can be equivalent to adding an extra subject. Benefits of multiple samples include the ability to reduce the number of subjects in a study and increase both the power and the available alpha. If, for example, the ICC is 35%, adding a second measurement can be equivalent to adding 48% more subjects to a single measurement study.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConclusion:\u003c/strong\u003e\u0026nbsp;A study's design can sometimes be improved by taking multiple synchronous samples. It is useful to evaluate this strategy as an extension of a single sample design. An Excel workbook is provided to allow researchers to explore the most appropriate number of samples to take in a given setting.\u003c/p\u003e","manuscriptTitle":"Employing Multiple Synchronous Outcome Samples Per Subject to Improve Study Efficiency.","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2021-03-25 13:49:59","doi":"10.21203/rs.3.rs-358007/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2021-04-16T00:00:00+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2021-04-15T00:00:00+00:00","index":1,"fulltext":"Recommendation: Accept after discretionary revisions\nForm responses:\n---\n\nComments to Author:\n---\nI have been requested to review the manuscript entitled \"Employing multiple synchronous outcome samples per subject to improve study efficiency,\" and I accepted, intrigued by the title. However, I soon realized that the content of this manuscript exceeds my expertise. Nevertheless, once recruited, a soldier goes bravely to battle.\n\nThe manuscript appears as Article Type \"Software,\" and an Excel workbook is attached for use by trialists and statisticians. The manuscript itself is a methodological note that very elegantly describes a sampling strategy based on doing more samples in each subject to improve the accuracy of estimation of differences in outcomes between groups of subjects being compared—a sort of cluster approach to controlled trials. The author explains and promotes this strategy (so long as the intra-class correlation is low) to improve accuracy and maybe reduce the number of subjects a study must recruit to achieve power.\n\nThe manuscript flows easily, and the writing is impeccable. For BMC Medical Research Methodology, this article is very well-aligned with journal purpose. I recommend that peer statisticians be called in to assess the manuscript's technical merit, but it seems sound to the best of my knowledge.\n\nThe only suggestion I would make is for the author to include a box or panel with some real-world examples of applying this strategy in a randomized clinical trial and a prospective cohort study. These simple examples could compare traditional sample calculations with and without multiple synchronous samples. A narrative or info-graphic box display could catch a researcher's attention—those who might be interested in grasping the essence of the approach but who do not need to understand the nuts and bolts (we can leave that for the statisticians).\n\nAfter reading the manuscript, I could not find anything that should be corrected or improved. The author has done a great job reporting the technique and devising a very well-thought-out and comprehensive workbook.\n\nCongratulations. One can see the care and passion put into this manuscript.\n* Publons Reviewer Recognition. Springer Nature can send verification of this review directly to Publons (a subsidiary of Clarivate Analytics). If you would like to take advantage of this service, please click on the “Yes” option below. Your name, email address, title of the reviewed manuscript, name of the journal, and date of your review submission (the “Review Data”) will then be transmitted to Publons after the final decision on the manuscript has been made. If you have already registered at Publons, they will notify you of the receipt of this review and update your profile as per your settings and their policy. If you are not registered with Publons, you will receive an email from them asking you to register in order for them to be able to recognize your review on your new profile page. Publons may use the Review Data to generate derivative metadata for the benefit of Publons and you as a reviewer, carefully considering the sensitivity of such information. For example, Publons may verify your record as a reviewer by updating your profile published on its webservice if you have registered for such service or help editors to identify candidate reviewers. Please find the details of processing in Publons’ privacy policy https://publons.com/about/terms: **Yes**\n* Declaration of competing interests: **I declare that I have no competing interests**\n* Reviewer Publication Consent. I agree for my report to be made available under an Open Access Creative Commons CC-BY License (http://creativecommons.org/licenses/by/4.0) if this manuscript is accepted for publication. Any comments that I do not wish to be included in the published report have been included as confidential comments to the editor, which will not be published.: **I agree to the terms of the CC-BY 4.0 license; please publish my name with my report.**\n* Is the study design appropriate to answer the research question (including the use of appropriate controls), and are the conclusions supported by the evidence presented?: **Yes**\n* Are the methods sufficiently described to allow the study to be repeated?: **Yes**\n* Is the use of statistics and treatment of uncertainties appropriate?: **Yes**\n* Is the presentation of the work clear?: **Yes**\n* Are the images in this manuscript (including electrophoretic gels and blots) free from apparent manipulation?: **Yes**\n"},{"type":"editorInvitedReview","content":"","date":"2021-03-31T00:00:00+00:00","index":3,"fulltext":"Recommendation: Reviewer's comments unavailable pending editorial decision\n"},{"type":"editorInvitedReview","content":"","date":"2021-03-31T00:00:00+00:00","index":2,"fulltext":"Recommendation: Major revisions required\nForm responses:\n---\n\nComments to Author:\n---\nThe present study provides an overview and an excel tool for researchers of how multiple measurements from a subject can improve power or reduce sample size in studies. Although this approach is not new, I find the paper of relevance for researchers aiming to design their studies in the most efficient and acceptable way. Please find some specific comments below.\n\n1) Introduction. The use of multiple measurements obtained from each subject is a common strategy to reduce random measurement error, for example, as recommended by Hutcheon et al, 2010 (doi: 10.1136/bmj.c2289). Please consider citing this study in the introduction. Moreover, a range of other examples than those currently included may help to meet a broader audience. For example, it is a much used to strategy to obtain and use the mean of 3-4 measurements of blood pressure, typically 2 measurements of waist circumference, a minimum of 4 days of accelerometer monitoring for measuring physical activity, etc. Moreover, the same strategy is applied for questionnaires, where several items is often used to form a mean score. In these cases, the Cronbach's alpha is calculated with the aim to show how many items, days, measurements, etc are needed to arrive at a sufficient level of reliability. Actually, whether internal consistency, intra- or inter-observer reliability or test-retest reliability is the most appropriate term and context, the approach is in principle the same. Please consider broadening the context.\n\n2) I find the parallel to cluster randomized trials very useful.\n\n3) Implementation. No single reference is given. Please provide a few citations to support the concepts and formulas presented.\n\n4) Page 5 line 48-50. I agree with this hypothesis, however, it does not necessarily hold. Please see the study by Aadland et al, 2020 (doi: 10.1080/02640414.2020.1743054) Suppl Fig 1 regarding variance components and consider whether this finding inform the present work.\n\n5) Table 1. I think the second part of the figure is difficult to understand and recommend stating/using the headline, for example, \"Effective sample size\" and/or start out with 100 subjects in the first row and then show the \"effective\" number of subjects if increasing the number of measurements for consecutive rows.\n\n6) Page 6 line 46-48. I agree some scenarios are of little practical value and recommend adding/discussing which problems arise when the ICC is very low. Are you actually measuring the same phenomenon in such situations and would such scores be relevant to sum?\n\n7) Some places the term correlation is used. I expect ICC is meant. I recommend using ICC throughout to avoid confusion with the Pearson correlation, which readers will be very familiar with.\n\n8) The excel tool can be used for sample size calculation for between-group comparisons. However, in many cases studies have pre- and post-tests where the correlation between the measurements can be taken into account to reduce the number of subjects needed. How does this correction relate to this work? This point may also relate to comment 1, as this situation might exemplify what may be regarded different types of reliability (e.g., internal consistency versus test-retest)? My point is that the correlation between the pre-test and post-test summary measures is also needed to calculate the sample size in longitudinal designs.\n\n9) Page 9 line 12. I would not use the statement \"change the relevance\". I would say it is relevant, but careful consideration of the context is needed.\n* Publons Reviewer Recognition. Springer Nature can send verification of this review directly to Publons (a subsidiary of Clarivate Analytics). If you would like to take advantage of this service, please click on the “Yes” option below. Your name, email address, title of the reviewed manuscript, name of the journal, and date of your review submission (the “Review Data”) will then be transmitted to Publons after the final decision on the manuscript has been made. If you have already registered at Publons, they will notify you of the receipt of this review and update your profile as per your settings and their policy. If you are not registered with Publons, you will receive an email from them asking you to register in order for them to be able to recognize your review on your new profile page. Publons may use the Review Data to generate derivative metadata for the benefit of Publons and you as a reviewer, carefully considering the sensitivity of such information. For example, Publons may verify your record as a reviewer by updating your profile published on its webservice if you have registered for such service or help editors to identify candidate reviewers. Please find the details of processing in Publons’ privacy policy https://publons.com/about/terms: **No**\n* Declaration of competing interests: **I declare that I have no competing interests**\n* Reviewer Publication Consent. I agree for my report to be made available under an Open Access Creative Commons CC-BY License (http://creativecommons.org/licenses/by/4.0) if this manuscript is accepted for publication. Any comments that I do not wish to be included in the published report have been included as confidential comments to the editor, which will not be published.: **I agree to the terms of the CC-BY 4.0 license; please do not publish my name with my report. (default)**\n* Is the study design appropriate to answer the research question (including the use of appropriate controls), and are the conclusions supported by the evidence presented?: **Yes**\n* Are the methods sufficiently described to allow the study to be repeated?: **Yes**\n* Is the use of statistics and treatment of uncertainties appropriate?: **Yes**\n* Is the presentation of the work clear?: **Yes**\n* Are the images in this manuscript (including electrophoretic gels and blots) free from apparent manipulation?: **Yes**\n"},{"type":"reviewerAgreed","content":"","date":"2021-03-29T00:00:00+00:00","index":3,"fulltext":""},{"type":"reviewerAgreed","content":"","date":"2021-03-25T00:00:00+00:00","index":2,"fulltext":""},{"type":"editorAssigned","content":"","date":"2021-03-22T00:00:00+00:00","index":"","fulltext":""},{"type":"reviewersInvited","content":"","date":"2021-03-22T00:00:00+00:00","index":"","fulltext":""},{"type":"reviewerAgreed","content":"","date":"2021-03-22T00:00:00+00:00","index":1,"fulltext":""},{"type":"checksComplete","content":"","date":"2021-03-21T23:00:00+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2021-03-21T23:00:00+00:00","index":"","fulltext":""},{"type":"submitted","content":"","date":"2021-03-15T00:00:00+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"3577a70d-0f4e-4601-8cf6-c45fe0b6de99","owner":[],"postedDate":"March 25th, 2021","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":3215244,"name":"Health Economics \u0026 Outcomes Research"}],"tags":[],"updatedAt":"2021-03-25T13:49:59+00:00","versionOfRecord":[],"versionCreatedAt":"2021-03-25 13:49:59","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-358007","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-358007","identity":"rs-358007","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.