Confidence Intervals for the Difference Between Proportions after Multiple Imputation: A simulation study comparing different strategies

preprint OA: closed
Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-04 · read from full text

This paper studied how confidence interval methods for the absolute difference between two proportions behave after multiple imputation, using a simulation framework that generated two-group binomial outcomes with varying baseline proportions, effect sizes, and sample sizes. Across missing data mechanisms (missing completely at random and missing at random) with 10% and 30% missingness, it compared Newcombe-Wilson, Wald, and Agresti-Caffo confidence intervals after multiple imputation via MICE (m=10) with Rubin’s Rules for pooling. The main finding was that all three methods had empirical coverage probabilities close to the nominal 0.95 and produced similar interval lengths in all simulated scenarios, though the Wald intervals showed undercoverage and the authors highlighted Agresti-Caffo as a simpler alternative with strong performance. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Background The difference between proportions and their confidence intervals are important statistical methods to determine group differences in randomized trials and observational studies. In these studies missing data may severely bias results and should be handled carefully. The Newcombe-Wilson, Wald and Agresti-Caffo methods are commonly used to calculate confidence intervals between proportions. It is unclear how these methods behave in combination with multiple imputation to handle missing data. Methods A simulation study was conducted to compare the coverage probability and interval lengths of the Newcombe-Wilson, Wald and Agresti-Caffo methods to calculate the confidence intervals for the difference between proportions after multiple imputation. Evaluated was the performance in MCAR and MAR missing data of 10% and 30%, using large and small differences between proportions and small and large sample sizes. Results In all simulation scenarios the methods under study performed equally well and were close to the desired coverage probability of 0.95, with similar interval lengths. Conclusions It can be concluded that simple methods as the Wald and Agresti-Caffo confidence intervals perform similar as a more complex procedure as the Newcombe-Wilson intervals. The Wald intervals showed undercoverage which makes the Agresti-Caffo confidence intervals a strong alternative for the more complex Newcombe-Wilson intervals method.
Full text 79,595 characters · extracted from preprint-html · click to expand
Confidence Intervals for the Difference Between Proportions after Multiple Imputation: A simulation study comparing different strategies | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Confidence Intervals for the Difference Between Proportions after Multiple Imputation: A simulation study comparing different strategies MW Heymans, JL Brand This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8789678/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 17 You are reading this latest preprint version Abstract Background The difference between proportions and their confidence intervals are important statistical methods to determine group differences in randomized trials and observational studies. In these studies missing data may severely bias results and should be handled carefully. The Newcombe-Wilson, Wald and Agresti-Caffo methods are commonly used to calculate confidence intervals between proportions. It is unclear how these methods behave in combination with multiple imputation to handle missing data. Methods A simulation study was conducted to compare the coverage probability and interval lengths of the Newcombe-Wilson, Wald and Agresti-Caffo methods to calculate the confidence intervals for the difference between proportions after multiple imputation. Evaluated was the performance in MCAR and MAR missing data of 10% and 30%, using large and small differences between proportions and small and large sample sizes. Results In all simulation scenarios the methods under study performed equally well and were close to the desired coverage probability of 0.95, with similar interval lengths. Conclusions It can be concluded that simple methods as the Wald and Agresti-Caffo confidence intervals perform similar as a more complex procedure as the Newcombe-Wilson intervals. The Wald intervals showed undercoverage which makes the Agresti-Caffo confidence intervals a strong alternative for the more complex Newcombe-Wilson intervals method. Missing data Multiple Imputation Difference between Proportions Confidence Intervals Newcombe-Wilson method Agresti-Caffo method Wald method Figures Figure 1 Figure 2 Figure 3 Introduction In medical and epidemiological studies often a dichotomous outcome variable is used in relation to a dichotomous determinant or predictor variable. This outcome variable could be recovery following treatment. With this outcome the effect can be estimated between (treatment) groups by taking the difference between the proportions related to each (treatment) group. This measure is the absolute difference in proportions between groups and often interpreted as the absolute risk difference. For the confidence interval around this difference mostly the Wald interval is used (Wald 1943) because it is easily obtained as: \(\:{\widehat{p}}_{1}-{\widehat{p}}_{2}\:\pm\:z\:\sqrt{\frac{{\widehat{p}}_{1}\left(1-{\widehat{p}}_{1}\right)}{{n}_{1}}+\frac{{\widehat{p}}_{2}(1-{\widehat{p}}_{2})}{{n}_{2}}}\) , where \(\:{\widehat{p}}_{1}\) and \(\:{\widehat{p}}_{2}\) are the risks in each group, \(\:{n}_{1}\) and \(\:{n}_{2}\) the sample size in each group and \(\:z\) the 0.975 quantile of the standard normal distribution, which is 1.96. Although popular, the Wald intervals do not always lead to desirable confidence intervals coverage rates (Newcombe 1998). To generate more reliable results, other confidence interval procedures have been developed like the Newcombe-Wilson (Newcombe 1998) and Agresti-Caffo confidence intervals (Agresti and Caffo 2000), where the latter interval forms a simple adjustment to the Wald interval. Both intervals have shown to behave better in terms of coverage probability compared to the Wald interval (Newcombe 1998; Agresti and Caffo 2000; Brown 2005). Missing data plays a role in almost every study, although researchers do their best to avoid them. Complete case analysis may lead to underpowered results (Little 2019). A recommended approach to handle missing data is Multiple Imputation (MI). With MI several completed datasets are generated, pre-planned analyses conducted in the completed datasets and finally repeated analyses pooled into one summary estimate using Rubin’s Rules (RR) (Rubin 1987). MI can be used under the assumption that the data is Missing at random (MAR). MAR means that the missing data is related to observed variables, mostly assumed in practice and in combination with MI leads to correct parameter estimates and study outcomes (Rubin 1987). MCAR means that missing data is randomly distributed over study variables and is a rather strict assumption that you won’t see much in research practice. Recently, Lott and Reiter derived a Wilson confidence interval for single proportions that can be used in combination with MI and showed good performance in MCAR and MAR situations (Wilson 1927; Lott and Reiter 2020). This interval can be used to derive a Newcombe-Wilson confidence interval for the difference between two proportions in combination with MI as evaluated for its performance by Sidi and Harel (2022). They showed that this interval showed good coverage compared to the Wald confidence interval and recommended to use either the Wald or Newcombe-Wilson interval. A limitation of the Wilson interval may be that the procedure is complex and not easily implemented by applied researchers, especially not the method developed in combination with MI. A good alternative may be the Agresti-Caffo interval (2000). This interval is much easier obtained and implemented than the Wilson interval and showed improved coverage in complete data compared to the Wald interval (Brown and Li 2005). However, the performance of the Agresti-Caffo interval is never tested against the Wald or Wilson confidence intervals in combination with MI. The aim of this simulation study is to evaluate the properties of the Agresti-Caffo interval compared to the Wald and Newcombe-Wilson intervals to calculate the difference between two proportions after multiple imputation. Simulation setup Data generation For the simulation setup we followed the same steps as used by Lott and Reiter (2020) and Sidi and Harel (2022). Lott and Reiter generated one group of simulation data. We repeated their approach twice to generate data for two independent groups with binomial outcomes Y 1 and Y 2 with proportions p 1 and p 2 respectively. In the first group binomial data was generated using proportions p 1 = 0.3, 0.65, and 0.9 and in the second group using proportions p 2 = p1- \(\:\varDelta\:\) , with \(\:\varDelta\:\) = 0.1, 0.025. In both groups, two types of missing data assumptions were generated, one to mimic a missing completely at random (MCAR) and another to mimic a missing at random (MAR) situation. For both situations 10% and 30% of the data was made missing which we called the low and high missing data scenario’s respectively. To generate MCAR data, data was randomly set to missing in both simulated Y 1 and Y 2 outcomes. To simulate MAR data we applied the same procedure as described in Lott and Reiter within each group separately. We first generated a second Bernoulli variable x in each simulation sample to mimic a strong relation between the x and Y 1 and Y 2 variables, which we call strong relation. For this the following rule was used: \(\:\text{Pr}\left(x=1|{Y}_{\text{1,2}}=0\right)=0.6\:\) and \(\:\text{Pr}\left(x=1|{Y}_{\text{1,2}}=1\right)=0.2\) . We also generated data containing a weak correlation between x and and Y 1 and Y 2 according to: \(\:\text{Pr}\left(x=1|{Y}_{\text{1,2}}=0\right)=\text{Pr}\left(x=1|{Y}_{\text{1,2}}=1\right)=0.6\:\) . Subsequently, the data in Y 1 and Y 2 was made missing depending on x according to two missing data mechanisms with R = 1 if Y 1,2 is missing and R = 0 otherwise. The first missingness mechanism, which we call the low missingness mechanism, uses Pr (R = 1| x = 0 ) = 0.15 and Pr (R = 1| x = 1 ) = 0.07 and the second high missingness mechanism, uses Pr (R = 1| x = 0 ) = 0.48 and Pr (R = 1| x = 1 ) = 0.18. Sample sizes of 100 and 250 were used for each group. This resulted in 12 combined simulation scenario’s for the MCAR data (3 for each method evaluated, 3 for lower and higher proportions used in each group, 2 depending on the difference between proportions, 2 on the percentage of missing data and 2 on the sample size) and 14 for the MAR data (3 for each method evaluated, 3 for lower and higher proportions used in each group, 2 depending on the difference between proportions, 2 on the weak and strong correlation between variables, 2 depending on the percentage of missing data and 2 on the sample size). Simulated scenario’s were repeated 10,000 times. Imputation To multiply impute the missing data, the multivariate imputation by chained equations (MICE) algorithm was used to generate m = 10 imputed datasets using 10 iterations (Van Buuren 2018). For the MCAR missing data no additional variables were included into the imputation algorithm and for the MAR missing data only the x covariate was included as extra variable. Simulation measures The empirical nominal coverage rates are calculated for the complete and multiply imputed data respectively. The coverage rate is calculated as the number of times the confidence interval includes the value of \(\:{\widehat{p}}_{1}-{\widehat{p}}_{2}=0.025\) or \(\:{\widehat{p}}_{1}-{\widehat{p}}_{2}=0.1\) over al simulation samples. The length of the confidence intervals is calculated as the difference between the upper and lower limit of the confidence intervals. The confidence intervals are calculated as explained in the next paragraphs. Rubin’s Rules Rubin developed rules for inference from multiply imputed datasets. These rules work as follows (6). Assume we are interested in estimating a parameter \(\:Q\) . Than \(\:{\widehat{Q}}_{i}\:\) is the estimate of \(\:Q\) in each imputed dataset and we obtain \(\:\stackrel{-}{Q}=\frac{1}{m}\sum\:_{i}^{m}{\widehat{Q}}_{i}\) , where m is the number of imputed datasets and i refers to the estimate of \(\:\widehat{Q}\) in each imputed dataset. The variance is a combination of the within and between variance of \(\:{\widehat{Q}}_{i}\) . The within variance is the average of the complete-data variance \(\:{\widehat{U}}_{i}\) and is: \(\:\stackrel{-}{U}=\:\frac{1}{m}\sum\:_{i}^{m}{\widehat{U}}_{i}\) and the between variance is the variance between \(\:{\widehat{Q}}_{i}\) and is: \(\:B=\frac{1}{m-1}{\sum\:_{i}^{m}({\widehat{Q}}_{i}-\stackrel{-}{Q})}^{2}\) . Using these values, the total variance is calculated as: \(\:T=\left(1+\:\frac{1}{m}\right)\times\:B+\stackrel{-}{U}\) and the standard error is: \(\:SE=\:\sqrt{T}\) . The confidence interval is than estimated as \(\:\stackrel{-}{Q}\pm\:{t}_{v}\times\:SE\) . The value \(\:{t}_{v}\) is the quantile of the t-distribution with \(\:v\) degrees of freedom, calculated as \(\:v=(m-1){(1+1/r)}^{2}\) with \(\:{r}_{m}=\left(1+1/m\right)B/\stackrel{-}{U}\:\) . MI-Wald confidence intervals In the complete data situation the Wald interval is calculated as described in the introduction section. To derive the pooled Wald interval after MI, called MI-Wald, we apply Rubin’s Rules after the Wald interval is obtained repeatedly in the (completed) imputed data. First we take the average of the differences between proportions of \(\:\varDelta\:P=\:{\widehat{p}}_{1}-{\widehat{p}}_{2}\:\) in each imputed dataset as \(\:\stackrel{-}{Q}=\stackrel{-}{\varDelta\:P}=\:\frac{1}{m}\sum\:_{i}^{m}({{\widehat{p}}_{1}-{\widehat{p}}_{2})}_{i}\:\) . The within variance is the mean of the variance of the difference between proportions in each imputed datasset and obtained as: \(\:{W}_{var}=\:\frac{1}{m}\sum\:_{i}^{m}{{SE}^{2}}_{i}\) with \(\:{{SE}^{2}}_{i}=\frac{{\widehat{p}}_{1}\left(1-{\widehat{p}}_{1}\right)}{{n}_{1}}+\frac{{\widehat{p}}_{2}(1-{\widehat{p}}_{2})}{{n}_{2}}\) and the between variance is \(\:{B}_{var}=\sum\:_{i}^{m}\frac{(\varDelta\:{P}_{i}-\stackrel{-}{\varDelta\:P})}{m-1}\:.\:\) The total variance is than calculated as: \(\:{T}_{var}=\left(1+\frac{1}{m}\right)\times\:{B}_{var}+{W}_{var}\) and the standard error as \(\:SE=\:\sqrt{{T}_{var}}\) . The confidence intervals are than calculated as follows: \(\:\stackrel{-}{\varDelta\:P}\pm\:{t}_{v}\times\:SE\) as explained in the previous paragraph. MI-Agresti-Caffo (MI-AC) confidence intervals This interval is calculated in complete data as \(\:{\widehat{p}}_{1}-{\widehat{p}}_{2}\:\pm\:z\:\sqrt{\frac{{\widehat{p}}_{1}\left(1-{\widehat{p}}_{1}\right)}{{n}_{1}+2}+\frac{{\widehat{p}}_{2}(1-{\widehat{p}}_{2})}{{n}_{2}+2}}\) , $$\:\text{w}\text{h}\text{e}\text{r}\text{e}\:{\widehat{p}}_{1}=\:\frac{{x}_{1}+1}{{n}_{1}+2}\:\text{a}\text{n}\text{d}\:{\widehat{p}}_{2}=\:\frac{{x}_{2}+1}{{n}_{1}+2}$$ \(\:{\widehat{p}}_{1}\) \(\:\:\) and \(\:{\widehat{p}}_{2}\) are the risks in each group, \(\:{x}_{1}\) and \(\:{x}_{2}\:\) the events and \(\:{n}_{1}\) and \(\:{n}_{2}\) the sample size in each group and \(\:z\) the 0.975 quantile of the standard normal distribution, which is 1.96. To calculate the pooled Agresti-Caffo interval after MI, called the MI-AC interval, we use the same steps as for the MI-Wald but now use the adjusted version to calculate \(\:{\widehat{p}}_{1}\) and \(\:{\widehat{p}}_{2}\) . Note that these adjusted versions are used to calculate the confidence intervals but that the difference between proportions is calculated as in Wald, \(\:\varDelta\:P=\:{\widehat{p}}_{1}-{\widehat{p}}_{2}\) . MI Newcombe-Wilson (MI-NW) confidence interval For the Newcombe-Wilson interval in complete data we refer to (Lott and Reiter 2020). To derive the Newcombe-Wilson confidence intervals for the difference between two proportions across multiply imputed datasets we first applied the Wilson interval for the single proportion as described by Lott and Reiter (2020) to determine the upper and lower confidence limits for subgroup of \(\:{\widehat{p}}_{1}\) and \(\:{\widehat{p}}_{2}\) : $$\:{l}_{{\widehat{p}}_{1},\:{\widehat{p}}_{2}},\:{u}_{{\widehat{p}}_{1},\:{\widehat{p}}_{2}}=\frac{2{\stackrel{-}{Q}}_{m}+\frac{{{t}^{2}}_{v}}{n}+\frac{{{t}^{2}}_{v}{r}_{m}}{n}}{2(1+\frac{{{t}^{2}}_{v}}{n}+\frac{{{t}^{2}}_{v}{r}_{m}}{n})}\pm\:\sqrt{\frac{{(2{\stackrel{-}{Q}}_{m}+\frac{{{t}^{2}}_{v}}{n}+\frac{{{t}^{2}}_{v}{r}_{m}}{n})}^{2}}{{4(1+\frac{{{t}^{2}}_{v}}{n}+\frac{{{t}^{2}}_{v}{r}_{m}}{n})}^{2}}-\frac{{\stackrel{-}{Q}}_{m}^{2}}{1+\frac{{{t}^{2}}_{v}}{n}+\frac{{{t}^{2}}_{v}{r}_{m}}{n}}}$$ where \(\:\stackrel{-}{Q}\) is the mean proportion of \(\:{\widehat{p}}_{1}\) and \(\:{\widehat{p}}_{2}\) across imputed data sets, n the sample size, \(\:{t}^{2}\) is the quantile of the t distribution squared with v degrees of freedom as obtained by Rubin’s Rules, and \(\:{r}_{m}\) is defined as \(\:\left(1+1/m\right)B/\stackrel{-}{U}\:\) . Now we denote the multiply imputed confidence intervals for \(\:{\widehat{p}}_{1}\) as \(\:{(l}_{1},\:{u}_{1})\:\) and \(\:{\widehat{p}}_{2}\) as \(\:{(l}_{2},\:{u}_{2})\) . The Wilson interval can than be calculated as: $$\:\stackrel{-}{\varDelta\:P}-\sqrt{{({\stackrel{-}{p}}_{1}-{l}_{1})}^{2}+{({u}_{2}-{\stackrel{-}{p}}_{2})}^{2}}\:\text{a}\text{n}\text{d}\:\stackrel{-}{\varDelta\:P}+\sqrt{{({\stackrel{-}{p}}_{2}-{l}_{2})}^{2}+{({u}_{1}-{\stackrel{-}{p}}_{1})}^{2}}$$ Where \(\:\varDelta\:P={\widehat{p}}_{1}-{\widehat{p}}_{2}\) and \(\:\stackrel{-}{\varDelta\:P}\) is the mean \(\:\varDelta\:P\) across the multiply imputed datasets and \(\:{\stackrel{-}{p}}_{1}\) and \(\:{\stackrel{-}{p}}_{2}\) are the mean proportions across multiply imputed datasets for respectively group 1 and 2. Results Complete data The results of the complete data are presented in Figure S1 in the supplementary section. It can be concluded that for all the studied methods NW, AC and Wald and for lower and higher proportions and smaller and larger differences between proportions the confidence intervals are close to the nominal coverage rates of 0.95. This is independent of using a lower or higher sample size. MCAR missing data Figure 1 shows the coverage probabilities for the MCAR missing data situations when 10% and 30% of the data was made missing in the y variable in groups of sample sizes of n = 100 and n = 250. The differences between the coverage rates of the MI-NW, MI-AC and MI-Wald intervals are small. They are all close to the nominal coverage rates of 0.95. This is the case for lower and higher missing data rates, for lower and higher proportion values and also for smaller and larger differences in proportions between the groups. The average length of the confidence intervals (Supplement Table 1) is slightly higher after multiple imputation compared to the complete data for all methods, but are similar between the three methods. MAR weak correlation In Fig. 2 the coverage probabilities are presented of the MI-Wald, MI-NW and MI-AC confidence intervals with MAR missing data for the scenario that the x and y variables are weakly correlated. It can be shown that the coverage rates for all methods are close to the nominal coverage probability of 0.95, for a lower and higher missing data rate of 10% and 30% respectively and a lower and higher sample size of 100 and 250 respectively. This is also the case for lower and higher proportions and smaller and larger differences in proportions between groups. The average confidence interval lengths follow the same pattern as in the MCAR situation, they are a little bit larger after multiple imputation but are well in agreement with each other (Supplement Table 2). MAR strong correlation Figure 3 . Coverage probabilities for MAR missing data with a strong correlation between the Y and x variables. (the dashed line is the desired coverage probability of 0.95). For both missing data rates of 10% and 30% and sample sizes of 100 and 250, the MI-NW, MI-AC and MI-Wald methods generate confidence intervals that are close to the nominal level of 95%. For proportions around 0.9 the coverage rates are sometimes a little bit higher for the MI-NW method. For other proportion values and smaller and larger differences between proportions the coverage rates are close tot he nominal level. The average confidence lengths are presented in Table 3 and are well in agreement between all methods (Supplement Table 3). Discussion We have shown in this study that the MI-AC and MI-Wald confidence intervals coverage rates and length of confidence intervals for the difference between proportions in combination with MI perform equally well as the more complicated MI-NW method. A new finding is that the MI-AC method performs equally well as the MI-NW method. This is an advantage for applied researchers because the MI-AC method is not yet available in standard software packages, but is much easier to use and implement for the own study situation. The results in the current study confirm the results in the Study of Sidi and Harel (2022). They also compared the MI-Wald and MI-NW intervals and found a similar performance in terms of coverage probability and confidence interval length of both methods. They also used two other procedures, that are called MI-plug and MI-Li (Harel and Zhou 2006; Li et al. 2006). However, they did not recommend these methods for further use and therefore we did not evaluate them in the current study. An interesting difference between the studies of Lott and Reiter (2020) and Sidi and Harel (2022) and the current study is that they applied a multiple imputation procedure for binary data that is not implemented as a common multiple imputation procedure in software packages. To impute the missing binary data they used a beta binomial posterior predictive distribution. In our study, missing binary values were multiply imputed using the Multivariate Imputation by Chained Equations (MICE) method. This methods is frequently used and implemented in several software packages. MICE uses logistic regression models to model the outcome probability and to generate imputation on the 0, 1 scale (Van Buuren 2018). Using this imputation algorithm we found similar results in terms of performance between the Newcombe-Wilson and Wald method as in the study of Sidi and Harel (2022). We can therefore conclude that MICE is a valid multiple imputation procedure to handle missing data in studies that include binary variables and use the difference between proportions as main effect size. There are various methods developed to calculate confidence intervals for the difference between proportions. The method according to Newcombe-Wilson was developed as a reaction to the Wald method, because the latter procedure showed worse coverage in simulation studies of complete data (Newcombe 1998; Brown and Li 2005). Although the Newcombe-Wilson method shows better results, the method is also more complex and harder to implement for own use by applied researchers. That is especially the case for the Newcombe-Wilson procedure in combination with MI. An alternative procedure that is much easier to use and showed good performance in complete data is the Agresti-Caffo method (2000), which needs a slight adjustment to the Wald’s method. We showed that the confidence intervals according to Agresti-Caffo performed equally well as the Newcombe-Wilson method. We therefore recommend to use the Agresti-Caffo method to calculate confidence intervals between proportions after multiple imputation. This method as well as the Newcombe-Wilson and Wald method are implemented in the R package miceafter (Heymans 2022). Multiple imputation is a method that is applied more and more often by applied researchers. One of the reasons is that the application of MI is now available in various standard software programs as SPSS, SAS and Stata and therefore within reach of many applied researchers. Also the literature on multiple imputation applications is growing as well as guidelines on how to apply it (Lee et al. 2021; Little et al. 2012; Sterne et al. 2009, Li et al 2015). What remains behind is research on the application of statistical methods after MI. This is especially useful for applied researchers that use MI more often but are looking for specific and easy to use pooling methods that are not available within standard software packages yet, as calculating confidence intervals between proportions. This paper will add to that missing link between applying MI and pooling statistical methods in a valid way that is can easily be implemented in software packages. Some strengths of this study are that we used a lot of different realistic data scenario’s using lower and higher proportions and larger and smaller differences between proportions that can also be found in the literature. Another strong point is that we used a high number of simulation runs to test the performance of the methods. A limitation may be that we did not study the performance of the methods under Missing Not At Random (MNAR) data situations as Sidi and Harel (2022) did in their simulation study. However, by evaluating the performance in MCAR and MAR missing data situations we cover the most frequently accepted missing data types that are seen in the literature and accepted in practice and therefore think that our results are informative for many applied researchers. Conclusion We have shown that an easy to apply method as the Agresti-Caffo method to calculate confidence intervals between proportions can be validly applied after Multiple Imputation. This method is much easier to implement and performs equally well as the more complex Newcombe-Wilson method in multiply imputed datasets. Declarations Ethics approval and consent to participate Ethics, Consent to Participate, and Consent to Publish declarations: not applicable. Consent for publication Consent to for publication: not applicable Competing interests the authors have no competing interests that might be perceived to influence the results and/or discussion reported in this paper. Funding This research was not supported by any funding sources Author Contribution M.W.H. and J.L.B. conducted the analysis and wrote the main manuscript text . M.W.H. and J.L.B. reviewed and approved the manuscript. Data Availability The code to generate the simulated data of this study is available from the corresponding author upon request. The code to obtain the confidence intervals are available as separate functions in the R package miceafter and can be found on https://github.com/mwheymans/miceafter/ References Agresti, Alan, and Brian Caffo. Simple and Effective Confidence Intervals for Proportions and Differences of Proportions Result from Adding Two Successes and Two Failures. The American Statistician. 2000;54(4):280–88. Brown, L., and Li, X. Confidence Intervals for Two Sample Binomial Distribution. Journal of Statistical Planning and Inference. 2005;130:359–375. Buuren, S.V. (2018). Flexible Imputation of Missing Data (2nd ed.). Chapman and Hall/CRC.Stat. Harel,O. and Zhou, X. H. Multiple Imputation for CorrectingVerification Bias. Statistics in Medicine. 2006;25, 3769–3786. Heymans M.W. (2022). miceafter: Data Analysis and Pooling after Multiple Imputation. R package version 0.5.0. https://cran.r-project.org/web/packages/miceafter/index.html and https://mwheymans.github.io/miceafter/ Lee KJ, Tilling KM, Cornish RP, Little RJA, Bell ML, Goetghebeur E, Hogan JW, Carpenter JR; STRATOS initiative. Framework for the treatment and reporting of missing data in observational studies: The Treatment And Reporting of Missing data in Observational Studies framework. J Clin Epidemiol. 2021;134:79–88. Li, X., Mehrotra, D. V., and Barnard, J. Analysis of Incomplete Longitudinal Binary Data Using Multiple Imputation. Statistics in Medicine. 2006;25:2107–2124. Li P, Stuart EA, Allison DB. Multiple Imputation: A Flexible Tool for Handling Missing Data. JAMA. 2015;314(18):1966–1967. Little RJ, D'Agostino R, Cohen ML, Dickersin K, Emerson SS, Farrar JT, Frangakis C, Hogan JW, Molenberghs G, Murphy SA, Neaton JD, Rotnitzky A, Scharfstein D, Shih WJ, Siegel JP, Stern H. The prevention and treatment of missing data in clinical trials. N Engl J Med. 2012;367(14):1355-60. Little, Roderick, and Donald Rubin. (2019). Statistical Analysis with Missing Data, Third Edition. Hoboken, NJ, USA: Wiley. Lott, A., and Reiter, J. P. Wilson Confidence Intervals for Binomial Proportions With Multiple Imputation for Missing Data. The American Statistician. 2020;74:109–115. Newcombe, R. G. Interval Estimation for the Difference Between Independent Proportions: Comparison of Eleven Methods. Statistics in Medicine. 1998;17:873–890. Rubin, D. B. (1987). Multiple Imputation for Nonresponse in Surveys, New York: Wiley. Sidi, Y. & Ofer Harel. Difference Between Binomial Proportions Using Newcombe’s Method With Multiple Imputation for Incomplete Data. The American Statistician. 2022;76(1):29–36. Sterne JA, White IR, Carlin JB, Spratt M, Royston P, Kenward MG, Wood AM, Carpenter JR. Multiple imputation for missing data in epidemiological and clinical research: potential and pitfalls. BMJ. 2009;338:b2393. Wald, A. Tests of Statistical Hypotheses Concerning Several Parameters When the Number of Observations is Large. Transactions of the American Mathematical Society. 1943;54:426–482. Wilson, E. B. Probable Inference, the Law of Succession, and Statistical Inference. Journal of the American Statistical Association. 1927;22: 209–212. Additional Declarations No competing interests reported. Supplementary Files Supplementarymaterial.docx Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 19 Feb, 2026 Reviews received at journal 19 Feb, 2026 Reviews received at journal 15 Feb, 2026 Reviewers agreed at journal 14 Feb, 2026 Reviews received at journal 14 Feb, 2026 Reviews received at journal 13 Feb, 2026 Reviewers agreed at journal 11 Feb, 2026 Reviewers agreed at journal 11 Feb, 2026 Reviewers agreed at journal 10 Feb, 2026 Reviewers agreed at journal 10 Feb, 2026 Reviewers agreed at journal 09 Feb, 2026 Reviewers agreed at journal 09 Feb, 2026 Reviewers invited by journal 09 Feb, 2026 Editor invited by journal 09 Feb, 2026 Editor assigned by journal 05 Feb, 2026 Submission checks completed at journal 05 Feb, 2026 First submitted to journal 04 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8789678","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":589847923,"identity":"b1b170a7-7ba9-4e2f-a1c7-5c85a952f69f","order_by":0,"name":"MW Heymans","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABBklEQVRIiWNgGAWjYHACNijNfADKSIDS7AS1sIGUGiBpYSaohceAOC26MxLYHnxs2yZvzn/m46cbf/7ImbcnP93Mm2PHYI5Di9mNBHbDmW23DXfOyN0sndtmYCxz5pnZbd5tyQyWzTi03E5gk+Ztu8244QbvBuncBoPEGRIJIC0HGAwO49Hyt+22/YbzZx7/zvljUD9DIv0bYS2MbbcTNxzIYZPOYTNIkJDIIWDL/Ydtkj3nbidvuJFmZp3bZmw4g+dN2c2525J5cGo5c/iYxI+y27Ybzh9+fDvnj5y8BHv6thtvt9nJGRxvwK6HgbGBgZENizgPDvVQ8Ae/9CgYBaNgFIxwAADpeGN20zOVawAAAABJRU5ErkJggg==","orcid":"","institution":"Amsterdam UMC Location VUmc","correspondingAuthor":true,"prefix":"","firstName":"MW","middleName":"","lastName":"Heymans","suffix":""},{"id":589847926,"identity":"93f5029e-3070-4a40-9e69-2091f93032df","order_by":1,"name":"JL Brand","email":"","orcid":"","institution":"UWV (Employee Insurance Agency)","correspondingAuthor":false,"prefix":"","firstName":"JL","middleName":"","lastName":"Brand","suffix":""}],"badges":[],"createdAt":"2026-02-04 18:53:31","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8789678/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8789678/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":102541470,"identity":"bca59bd3-a880-4799-a898-76f48c568771","added_by":"auto","created_at":"2026-02-12 19:12:08","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":11407,"visible":true,"origin":"","legend":"\u003cp\u003eCoverage probabilities for lower and higher rates of MCAR data (the dashed line is the desired coverage probability of 0.95).\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8789678/v1/6bce53c876cedb54953cb007.png"},{"id":102541466,"identity":"03428656-a387-42bb-b010-ec54ee1e04e5","added_by":"auto","created_at":"2026-02-12 19:12:05","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":328372,"visible":true,"origin":"","legend":"\u003cp\u003eCoverage probabilities for MAR missing data with a weak correlation between the Y and x variables. (the dashed line is the desired coverage probability of 0.95).\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8789678/v1/bc57598057e608d8a5733c3e.jpeg"},{"id":102541446,"identity":"8eb6a9ec-584d-4677-bb01-7bcbaf73cbbe","added_by":"auto","created_at":"2026-02-12 19:12:00","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":11494,"visible":true,"origin":"","legend":"\u003cp\u003eCoverage probabilities for MAR missing data with a strong correlation between the Y and x variables. (the dashed line is the desired coverage probability of 0.95).\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8789678/v1/ac24bff2c2acb8771e95d396.png"},{"id":103049264,"identity":"c3d8b258-f588-451f-8ace-fb1b97b9ecbe","added_by":"auto","created_at":"2026-02-20 07:39:18","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1261128,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8789678/v1/1820b917-5b5b-4d49-80f6-801706ff5c4f.pdf"},{"id":102541449,"identity":"0aa4d31e-9652-4658-93a9-9accca07c7f8","added_by":"auto","created_at":"2026-02-12 19:12:01","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":140826,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementarymaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-8789678/v1/dcad4383ef50539b5a6d22de.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Confidence Intervals for the Difference Between Proportions after Multiple Imputation: A simulation study comparing different strategies","fulltext":[{"header":"Introduction","content":"\u003cp\u003eIn medical and epidemiological studies often a dichotomous outcome variable is used in relation to a dichotomous determinant or predictor variable. This outcome variable could be recovery following treatment. With this outcome the effect can be estimated between (treatment) groups by taking the difference between the proportions related to each (treatment) group. This measure is the absolute difference in proportions between groups and often interpreted as the absolute risk difference. For the confidence interval around this difference mostly the Wald interval is used (Wald 1943) because it is easily obtained as: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{1}-{\\widehat{p}}_{2}\\:\\pm\\:z\\:\\sqrt{\\frac{{\\widehat{p}}_{1}\\left(1-{\\widehat{p}}_{1}\\right)}{{n}_{1}}+\\frac{{\\widehat{p}}_{2}(1-{\\widehat{p}}_{2})}{{n}_{2}}}\\)\u003c/span\u003e\u003c/span\u003e, where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{1}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{2}\\)\u003c/span\u003e\u003c/span\u003e are the risks in each group, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{n}_{1}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{n}_{2}\\)\u003c/span\u003e\u003c/span\u003e the sample size in each group and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:z\\)\u003c/span\u003e\u003c/span\u003e the 0.975 quantile of the standard normal distribution, which is 1.96.\u003c/p\u003e \u003cp\u003eAlthough popular, the Wald intervals do not always lead to desirable confidence intervals coverage rates (Newcombe 1998). To generate more reliable results, other confidence interval procedures have been developed like the Newcombe-Wilson (Newcombe 1998) and Agresti-Caffo confidence intervals (Agresti and Caffo 2000), where the latter interval forms a simple adjustment to the Wald interval. Both intervals have shown to behave better in terms of coverage probability compared to the Wald interval (Newcombe 1998; Agresti and Caffo 2000; Brown 2005).\u003c/p\u003e \u003cp\u003eMissing data plays a role in almost every study, although researchers do their best to avoid them. Complete case analysis may lead to underpowered results (Little 2019). A recommended approach to handle missing data is Multiple Imputation (MI). With MI several completed datasets are generated, pre-planned analyses conducted in the completed datasets and finally repeated analyses pooled into one summary estimate using Rubin\u0026rsquo;s Rules (RR) (Rubin 1987). MI can be used under the assumption that the data is Missing at random (MAR). MAR means that the missing data is related to observed variables, mostly assumed in practice and in combination with MI leads to correct parameter estimates and study outcomes (Rubin 1987). MCAR means that missing data is randomly distributed over study variables and is a rather strict assumption that you won\u0026rsquo;t see much in research practice.\u003c/p\u003e \u003cp\u003eRecently, Lott and Reiter derived a Wilson confidence interval for single proportions that can be used in combination with MI and showed good performance in MCAR and MAR situations (Wilson 1927; Lott and Reiter 2020). This interval can be used to derive a Newcombe-Wilson confidence interval for the difference between two proportions in combination with MI as evaluated for its performance by Sidi and Harel (2022). They showed that this interval showed good coverage compared to the Wald confidence interval and recommended to use either the Wald or Newcombe-Wilson interval. A limitation of the Wilson interval may be that the procedure is complex and not easily implemented by applied researchers, especially not the method developed in combination with MI.\u003c/p\u003e \u003cp\u003eA good alternative may be the Agresti-Caffo interval (2000). This interval is much easier obtained and implemented than the Wilson interval and showed improved coverage in complete data compared to the Wald interval (Brown and Li 2005). However, the performance of the Agresti-Caffo interval is never tested against the Wald or Wilson confidence intervals in combination with MI. The aim of this simulation study is to evaluate the properties of the Agresti-Caffo interval compared to the Wald and Newcombe-Wilson intervals to calculate the difference between two proportions after multiple imputation.\u003c/p\u003e"},{"header":"Simulation setup","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eData generation\u003c/h2\u003e \u003cp\u003eFor the simulation setup we followed the same steps as used by Lott and Reiter (2020) and Sidi and Harel (2022). Lott and Reiter generated one group of simulation data. We repeated their approach twice to generate data for two independent groups with binomial outcomes Y\u003csub\u003e1\u003c/sub\u003e and Y\u003csub\u003e2\u003c/sub\u003e with proportions p\u003csub\u003e1\u003c/sub\u003e and p\u003csub\u003e2\u003c/sub\u003e respectively. In the first group binomial data was generated using proportions p\u003csub\u003e1\u003c/sub\u003e\u0026thinsp;=\u0026thinsp;0.3, 0.65, and 0.9 and in the second group using proportions p\u003csub\u003e2\u003c/sub\u003e\u0026thinsp;=\u0026thinsp;p1- \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varDelta\\:\\)\u003c/span\u003e\u003c/span\u003e, with \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varDelta\\:\\)\u003c/span\u003e\u003c/span\u003e = 0.1, 0.025. In both groups, two types of missing data assumptions were generated, one to mimic a missing completely at random (MCAR) and another to mimic a missing at random (MAR) situation. For both situations 10% and 30% of the data was made missing which we called the low and high missing data scenario\u0026rsquo;s respectively.\u003c/p\u003e \u003cp\u003eTo generate MCAR data, data was randomly set to missing in both simulated Y\u003csub\u003e1\u003c/sub\u003e and Y\u003csub\u003e2\u003c/sub\u003e outcomes. To simulate MAR data we applied the same procedure as described in Lott and Reiter within each group separately. We first generated a second Bernoulli variable x in each simulation sample to mimic a strong relation between the x and Y\u003csub\u003e1\u003c/sub\u003e and Y\u003csub\u003e2\u003c/sub\u003e variables, which we call strong relation. For this the following rule was used: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\text{Pr}\\left(x=1|{Y}_{\\text{1,2}}=0\\right)=0.6\\:\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\text{Pr}\\left(x=1|{Y}_{\\text{1,2}}=1\\right)=0.2\\)\u003c/span\u003e\u003c/span\u003e. We also generated data containing a weak correlation between x and and Y\u003csub\u003e1\u003c/sub\u003e and Y\u003csub\u003e2\u003c/sub\u003e according to: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\text{Pr}\\left(x=1|{Y}_{\\text{1,2}}=0\\right)=\\text{Pr}\\left(x=1|{Y}_{\\text{1,2}}=1\\right)=0.6\\:\\)\u003c/span\u003e\u003c/span\u003e. Subsequently, the data in Y\u003csub\u003e1\u003c/sub\u003e and Y\u003csub\u003e2\u003c/sub\u003e was made missing depending on \u003cem\u003ex\u003c/em\u003e according to two missing data mechanisms with R\u0026thinsp;=\u0026thinsp;1 if Y\u003csub\u003e1,2\u003c/sub\u003e is missing and R\u0026thinsp;=\u0026thinsp;0 otherwise. The first missingness mechanism, which we call the low missingness mechanism, uses Pr\u003cem\u003e(R\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1|\u003cem\u003ex\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0\u003cem\u003e)\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.15 and Pr\u003cem\u003e(R\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1|\u003cem\u003ex\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1\u003cem\u003e)\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.07 and the second high missingness mechanism, uses Pr\u003cem\u003e(R\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1|\u003cem\u003ex\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0\u003cem\u003e)\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.48 and Pr\u003cem\u003e(R\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1|\u003cem\u003ex\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1\u003cem\u003e)\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.18. Sample sizes of 100 and 250 were used for each group. This resulted in 12 combined simulation scenario\u0026rsquo;s for the MCAR data (3 for each method evaluated, 3 for lower and higher proportions used in each group, 2 depending on the difference between proportions, 2 on the percentage of missing data and 2 on the sample size) and 14 for the MAR data (3 for each method evaluated, 3 for lower and higher proportions used in each group, 2 depending on the difference between proportions, 2 on the weak and strong correlation between variables, 2 depending on the percentage of missing data and 2 on the sample size). Simulated scenario\u0026rsquo;s were repeated 10,000 times.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eImputation\u003c/h3\u003e\n\u003cp\u003eTo multiply impute the missing data, the multivariate imputation by chained equations (MICE) algorithm was used to generate m\u0026thinsp;=\u0026thinsp;10 imputed datasets using 10 iterations (Van Buuren 2018). For the MCAR missing data no additional variables were included into the imputation algorithm and for the MAR missing data only the x covariate was included as extra variable.\u003c/p\u003e\n\u003ch3\u003eSimulation measures\u003c/h3\u003e\n\u003cp\u003eThe empirical nominal coverage rates are calculated for the complete and multiply imputed data respectively. The coverage rate is calculated as the number of times the confidence interval includes the value of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{1}-{\\widehat{p}}_{2}=0.025\\)\u003c/span\u003e\u003c/span\u003e or \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{1}-{\\widehat{p}}_{2}=0.1\\)\u003c/span\u003e\u003c/span\u003e over al simulation samples. The length of the confidence intervals is calculated as the difference between the upper and lower limit of the confidence intervals. The confidence intervals are calculated as explained in the next paragraphs.\u003c/p\u003e\n\u003ch3\u003eRubin’s Rules\u003c/h3\u003e\n\u003cp\u003eRubin developed rules for inference from multiply imputed datasets. These rules work as follows (6). Assume we are interested in estimating a parameter \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:Q\\)\u003c/span\u003e\u003c/span\u003e. Than \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{Q}}_{i}\\:\\)\u003c/span\u003e\u003c/span\u003eis the estimate of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:Q\\)\u003c/span\u003e\u003c/span\u003e in each imputed dataset and we obtain \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{-}{Q}=\\frac{1}{m}\\sum\\:_{i}^{m}{\\widehat{Q}}_{i}\\)\u003c/span\u003e\u003c/span\u003e, where m is the number of imputed datasets and i refers to the estimate of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\widehat{Q}\\)\u003c/span\u003e\u003c/span\u003e in each imputed dataset. The variance is a combination of the within and between variance of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{Q}}_{i}\\)\u003c/span\u003e\u003c/span\u003e. The within variance is the average of the complete-data variance \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{U}}_{i}\\)\u003c/span\u003e\u003c/span\u003e and is: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{-}{U}=\\:\\frac{1}{m}\\sum\\:_{i}^{m}{\\widehat{U}}_{i}\\)\u003c/span\u003e\u003c/span\u003e and the between variance is the variance between \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{Q}}_{i}\\)\u003c/span\u003e\u003c/span\u003e and is: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:B=\\frac{1}{m-1}{\\sum\\:_{i}^{m}({\\widehat{Q}}_{i}-\\stackrel{-}{Q})}^{2}\\)\u003c/span\u003e\u003c/span\u003e. Using these values, the total variance is calculated as: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:T=\\left(1+\\:\\frac{1}{m}\\right)\\times\\:B+\\stackrel{-}{U}\\)\u003c/span\u003e\u003c/span\u003e and the standard error is: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:SE=\\:\\sqrt{T}\\)\u003c/span\u003e\u003c/span\u003e. The confidence interval is than estimated as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{-}{Q}\\pm\\:{t}_{v}\\times\\:SE\\)\u003c/span\u003e\u003c/span\u003e. The value \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{t}_{v}\\)\u003c/span\u003e\u003c/span\u003e is the quantile of the t-distribution with \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:v\\)\u003c/span\u003e\u003c/span\u003e degrees of freedom, calculated as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:v=(m-1){(1+1/r)}^{2}\\)\u003c/span\u003e\u003c/span\u003e with \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{r}_{m}=\\left(1+1/m\\right)B/\\stackrel{-}{U}\\:\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e\n\u003ch3\u003eMI-Wald confidence intervals\u003c/h3\u003e\n\u003cp\u003eIn the complete data situation the Wald interval is calculated as described in the introduction section. To derive the pooled Wald interval after MI, called MI-Wald, we apply Rubin\u0026rsquo;s Rules after the Wald interval is obtained repeatedly in the (completed) imputed data. First we take the average of the differences between proportions of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varDelta\\:P=\\:{\\widehat{p}}_{1}-{\\widehat{p}}_{2}\\:\\)\u003c/span\u003e\u003c/span\u003ein each imputed dataset as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{-}{Q}=\\stackrel{-}{\\varDelta\\:P}=\\:\\frac{1}{m}\\sum\\:_{i}^{m}({{\\widehat{p}}_{1}-{\\widehat{p}}_{2})}_{i}\\:\\)\u003c/span\u003e\u003c/span\u003e. The within variance is the mean of the variance of the difference between proportions in each imputed datasset and obtained as: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{W}_{var}=\\:\\frac{1}{m}\\sum\\:_{i}^{m}{{SE}^{2}}_{i}\\)\u003c/span\u003e\u003c/span\u003e with \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{{SE}^{2}}_{i}=\\frac{{\\widehat{p}}_{1}\\left(1-{\\widehat{p}}_{1}\\right)}{{n}_{1}}+\\frac{{\\widehat{p}}_{2}(1-{\\widehat{p}}_{2})}{{n}_{2}}\\)\u003c/span\u003e\u003c/span\u003e and the between variance is \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{B}_{var}=\\sum\\:_{i}^{m}\\frac{(\\varDelta\\:{P}_{i}-\\stackrel{-}{\\varDelta\\:P})}{m-1}\\:.\\:\\)\u003c/span\u003e\u003c/span\u003eThe total variance is than calculated as: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{T}_{var}=\\left(1+\\frac{1}{m}\\right)\\times\\:{B}_{var}+{W}_{var}\\)\u003c/span\u003e\u003c/span\u003e and the standard error as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:SE=\\:\\sqrt{{T}_{var}}\\)\u003c/span\u003e\u003c/span\u003e. The confidence intervals are than calculated as follows: \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{-}{\\varDelta\\:P}\\pm\\:{t}_{v}\\times\\:SE\\)\u003c/span\u003e\u003c/span\u003e as explained in the previous paragraph.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eMI-Agresti-Caffo (MI-AC) confidence intervals\u003c/h2\u003e \u003cp\u003eThis interval is calculated in complete data as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{1}-{\\widehat{p}}_{2}\\:\\pm\\:z\\:\\sqrt{\\frac{{\\widehat{p}}_{1}\\left(1-{\\widehat{p}}_{1}\\right)}{{n}_{1}+2}+\\frac{{\\widehat{p}}_{2}(1-{\\widehat{p}}_{2})}{{n}_{2}+2}}\\)\u003c/span\u003e\u003c/span\u003e,\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:\\text{w}\\text{h}\\text{e}\\text{r}\\text{e}\\:{\\widehat{p}}_{1}=\\:\\frac{{x}_{1}+1}{{n}_{1}+2}\\:\\text{a}\\text{n}\\text{d}\\:{\\widehat{p}}_{2}=\\:\\frac{{x}_{2}+1}{{n}_{1}+2}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{1}\\)\u003c/span\u003e \u003c/span\u003e \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\:\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{2}\\)\u003c/span\u003e\u003c/span\u003e are the risks in each group, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{x}_{1}\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{x}_{2}\\:\\)\u003c/span\u003e\u003c/span\u003ethe events and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{n}_{1}\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{n}_{2}\\)\u003c/span\u003e\u003c/span\u003e the sample size in each group and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:z\\)\u003c/span\u003e\u003c/span\u003e the 0.975 quantile of the standard normal distribution, which is 1.96.\u003c/p\u003e \u003cp\u003eTo calculate the pooled Agresti-Caffo interval after MI, called the MI-AC interval, we use the same steps as for the MI-Wald but now use the adjusted version to calculate \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{1}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{2}\\)\u003c/span\u003e\u003c/span\u003e. Note that these adjusted versions are used to calculate the confidence intervals but that the difference between proportions is calculated as in Wald, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varDelta\\:P=\\:{\\widehat{p}}_{1}-{\\widehat{p}}_{2}\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eMI Newcombe-Wilson (MI-NW) confidence interval\u003c/p\u003e \u003cp\u003eFor the Newcombe-Wilson interval in complete data we refer to (Lott and Reiter 2020). To derive the Newcombe-Wilson confidence intervals for the difference between two proportions across multiply imputed datasets we first applied the Wilson interval for the single proportion as described by Lott and Reiter (2020) to determine the upper and lower confidence limits for subgroup of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{1}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{2}\\)\u003c/span\u003e\u003c/span\u003e:\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:{l}_{{\\widehat{p}}_{1},\\:{\\widehat{p}}_{2}},\\:{u}_{{\\widehat{p}}_{1},\\:{\\widehat{p}}_{2}}=\\frac{2{\\stackrel{-}{Q}}_{m}+\\frac{{{t}^{2}}_{v}}{n}+\\frac{{{t}^{2}}_{v}{r}_{m}}{n}}{2(1+\\frac{{{t}^{2}}_{v}}{n}+\\frac{{{t}^{2}}_{v}{r}_{m}}{n})}\\pm\\:\\sqrt{\\frac{{(2{\\stackrel{-}{Q}}_{m}+\\frac{{{t}^{2}}_{v}}{n}+\\frac{{{t}^{2}}_{v}{r}_{m}}{n})}^{2}}{{4(1+\\frac{{{t}^{2}}_{v}}{n}+\\frac{{{t}^{2}}_{v}{r}_{m}}{n})}^{2}}-\\frac{{\\stackrel{-}{Q}}_{m}^{2}}{1+\\frac{{{t}^{2}}_{v}}{n}+\\frac{{{t}^{2}}_{v}{r}_{m}}{n}}}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{-}{Q}\\)\u003c/span\u003e\u003c/span\u003e is the mean proportion of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{1}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{2}\\)\u003c/span\u003e\u003c/span\u003e across imputed data sets, n the sample size, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{t}^{2}\\)\u003c/span\u003e\u003c/span\u003e is the quantile of the t distribution squared with v degrees of freedom as obtained by Rubin\u0026rsquo;s Rules, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{r}_{m}\\)\u003c/span\u003e\u003c/span\u003e is defined as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\left(1+1/m\\right)B/\\stackrel{-}{U}\\:\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eNow we denote the multiply imputed confidence intervals for \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{1}\\)\u003c/span\u003e\u003c/span\u003e as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{(l}_{1},\\:{u}_{1})\\:\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\widehat{p}}_{2}\\)\u003c/span\u003e\u003c/span\u003e as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{(l}_{2},\\:{u}_{2})\\)\u003c/span\u003e\u003c/span\u003e. The Wilson interval can than be calculated as:\u003cdiv id=\"Equc\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$$\\:\\stackrel{-}{\\varDelta\\:P}-\\sqrt{{({\\stackrel{-}{p}}_{1}-{l}_{1})}^{2}+{({u}_{2}-{\\stackrel{-}{p}}_{2})}^{2}}\\:\\text{a}\\text{n}\\text{d}\\:\\stackrel{-}{\\varDelta\\:P}+\\sqrt{{({\\stackrel{-}{p}}_{2}-{l}_{2})}^{2}+{({u}_{1}-{\\stackrel{-}{p}}_{1})}^{2}}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eWhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varDelta\\:P={\\widehat{p}}_{1}-{\\widehat{p}}_{2}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{-}{\\varDelta\\:P}\\)\u003c/span\u003e\u003c/span\u003e is the mean \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\varDelta\\:P\\)\u003c/span\u003e\u003c/span\u003e across the multiply imputed datasets and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\stackrel{-}{p}}_{1}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\stackrel{-}{p}}_{2}\\)\u003c/span\u003e\u003c/span\u003e are the mean proportions across multiply imputed datasets for respectively group 1 and 2.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003eComplete data\u003c/h2\u003e \u003cp\u003eThe results of the complete data are presented in Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e in the supplementary section. It can be concluded that for all the studied methods NW, AC and Wald and for lower and higher proportions and smaller and larger differences between proportions the confidence intervals are close to the nominal coverage rates of 0.95. This is independent of using a lower or higher sample size.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eMCAR missing data\u003c/h2\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the coverage probabilities for the MCAR missing data situations when 10% and 30% of the data was made missing in the y variable in groups of sample sizes of n\u0026thinsp;=\u0026thinsp;100 and n\u0026thinsp;=\u0026thinsp;250.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe differences between the coverage rates of the MI-NW, MI-AC and MI-Wald intervals are small. They are all close to the nominal coverage rates of 0.95. This is the case for lower and higher missing data rates, for lower and higher proportion values and also for smaller and larger differences in proportions between the groups. The average length of the confidence intervals (Supplement Table\u0026nbsp;1) is slightly higher after multiple imputation compared to the complete data for all methods, but are similar between the three methods.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eMAR weak correlation\u003c/h2\u003e \u003cp\u003eIn Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e the coverage probabilities are presented of the MI-Wald, MI-NW and MI-AC confidence intervals with MAR missing data for the scenario that the x and y variables are weakly correlated.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIt can be shown that the coverage rates for all methods are close to the nominal coverage probability of 0.95, for a lower and higher missing data rate of 10% and 30% respectively and a lower and higher sample size of 100 and 250 respectively. This is also the case for lower and higher proportions and smaller and larger differences in proportions between groups. The average confidence interval lengths follow the same pattern as in the MCAR situation, they are a little bit larger after multiple imputation but are well in agreement with each other (Supplement Table\u0026nbsp;2).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eMAR strong correlation\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. Coverage probabilities for MAR missing data with a strong correlation between the Y and x variables. (the dashed line is the desired coverage probability of 0.95).\u003c/p\u003e \u003cp\u003eFor both missing data rates of 10% and 30% and sample sizes of 100 and 250, the MI-NW, MI-AC and MI-Wald methods generate confidence intervals that are close to the nominal level of 95%. For proportions around 0.9 the coverage rates are sometimes a little bit higher for the MI-NW method. For other proportion values and smaller and larger differences between proportions the coverage rates are close tot he nominal level. The average confidence lengths are presented in Table\u0026nbsp;3 and are well in agreement between all methods (Supplement Table\u0026nbsp;3).\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eWe have shown in this study that the MI-AC and MI-Wald confidence intervals coverage rates and length of confidence intervals for the difference between proportions in combination with MI perform equally well as the more complicated MI-NW method. A new finding is that the MI-AC method performs equally well as the MI-NW method. This is an advantage for applied researchers because the MI-AC method is not yet available in standard software packages, but is much easier to use and implement for the own study situation.\u003c/p\u003e \u003cp\u003eThe results in the current study confirm the results in the Study of Sidi and Harel (2022). They also compared the MI-Wald and MI-NW intervals and found a similar performance in terms of coverage probability and confidence interval length of both methods. They also used two other procedures, that are called MI-plug and MI-Li (Harel and Zhou 2006; Li et al. 2006). However, they did not recommend these methods for further use and therefore we did not evaluate them in the current study.\u003c/p\u003e \u003cp\u003eAn interesting difference between the studies of Lott and Reiter (2020) and Sidi and Harel (2022) and the current study is that they applied a multiple imputation procedure for binary data that is not implemented as a common multiple imputation procedure in software packages. To impute the missing binary data they used a beta binomial posterior predictive distribution. In our study, missing binary values were multiply imputed using the Multivariate Imputation by Chained Equations (MICE) method. This methods is frequently used and implemented in several software packages. MICE uses logistic regression models to model the outcome probability and to generate imputation on the 0, 1 scale (Van Buuren 2018). Using this imputation algorithm we found similar results in terms of performance between the Newcombe-Wilson and Wald method as in the study of Sidi and Harel (2022). We can therefore conclude that MICE is a valid multiple imputation procedure to handle missing data in studies that include binary variables and use the difference between proportions as main effect size.\u003c/p\u003e \u003cp\u003eThere are various methods developed to calculate confidence intervals for the difference between proportions. The method according to Newcombe-Wilson was developed as a reaction to the Wald method, because the latter procedure showed worse coverage in simulation studies of complete data (Newcombe 1998; Brown and Li 2005). Although the Newcombe-Wilson method shows better results, the method is also more complex and harder to implement for own use by applied researchers. That is especially the case for the Newcombe-Wilson procedure in combination with MI. An alternative procedure that is much easier to use and showed good performance in complete data is the Agresti-Caffo method (2000), which needs a slight adjustment to the Wald\u0026rsquo;s method. We showed that the confidence intervals according to Agresti-Caffo performed equally well as the Newcombe-Wilson method. We therefore recommend to use the Agresti-Caffo method to calculate confidence intervals between proportions after multiple imputation. This method as well as the Newcombe-Wilson and Wald method are implemented in the R package miceafter (Heymans 2022).\u003c/p\u003e \u003cp\u003eMultiple imputation is a method that is applied more and more often by applied researchers. One of the reasons is that the application of MI is now available in various standard software programs as SPSS, SAS and Stata and therefore within reach of many applied researchers. Also the literature on multiple imputation applications is growing as well as guidelines on how to apply it (Lee et al. 2021; Little et al. 2012; Sterne et al. 2009, Li et al 2015). What remains behind is research on the application of statistical methods after MI. This is especially useful for applied researchers that use MI more often but are looking for specific and easy to use pooling methods that are not available within standard software packages yet, as calculating confidence intervals between proportions. This paper will add to that missing link between applying MI and pooling statistical methods in a valid way that is can easily be implemented in software packages.\u003c/p\u003e \u003cp\u003eSome strengths of this study are that we used a lot of different realistic data scenario\u0026rsquo;s using lower and higher proportions and larger and smaller differences between proportions that can also be found in the literature. Another strong point is that we used a high number of simulation runs to test the performance of the methods. A limitation may be that we did not study the performance of the methods under Missing Not At Random (MNAR) data situations as Sidi and Harel (2022) did in their simulation study. However, by evaluating the performance in MCAR and MAR missing data situations we cover the most frequently accepted missing data types that are seen in the literature and accepted in practice and therefore think that our results are informative for many applied researchers.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eWe have shown that an easy to apply method as the Agresti-Caffo method to calculate confidence intervals between proportions can be validly applied after Multiple Imputation. This method is much easier to implement and performs equally well as the more complex Newcombe-Wilson method in multiply imputed datasets.\u003c/p\u003e"},{"header":"Declarations","content":" \u003cp\u003e \u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e \u003cp\u003eEthics, Consent to Participate, and Consent to Publish declarations: not applicable.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eConsent for publication\u003c/strong\u003e \u003cp\u003eConsent to for publication: not applicable\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eCompeting interests\u003c/h2\u003e \u003cp\u003ethe authors have no competing interests that might be perceived to influence the results and/or discussion reported in this paper.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eThis research was not supported by any funding sources\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eM.W.H. and J.L.B. conducted the analysis and wrote the main manuscript text . M.W.H. and J.L.B. reviewed and approved the manuscript.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe code to generate the simulated data of this study is available from the corresponding author upon request. The code to obtain the confidence intervals are available as separate functions in the R package miceafter and can be found on https://github.com/mwheymans/miceafter/\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAgresti, Alan, and Brian Caffo. Simple and Effective Confidence Intervals for Proportions and Differences of Proportions Result from Adding Two Successes and Two Failures. The American Statistician. 2000;54(4):280–88.\u003c/li\u003e\n\u003cli\u003eBrown, L., and Li, X. Confidence Intervals for Two Sample Binomial Distribution. Journal of Statistical Planning and Inference. 2005;130:359–375.\u003c/li\u003e\n\u003cli\u003eBuuren, S.V. (2018). Flexible Imputation of Missing Data (2nd ed.). Chapman and Hall/CRC.Stat.\u003c/li\u003e\n\u003cli\u003eHarel,O. and Zhou, X. H. Multiple Imputation for CorrectingVerification Bias. Statistics in Medicine. 2006;25, 3769–3786.\u003c/li\u003e\n\u003cli\u003eHeymans M.W. (2022). miceafter: Data Analysis and Pooling after Multiple Imputation. R package version 0.5.0. https://cran.r-project.org/web/packages/miceafter/index.html and https://mwheymans.github.io/miceafter/\u003c/li\u003e\n\u003cli\u003eLee KJ, Tilling KM, Cornish RP, Little RJA, Bell ML, Goetghebeur E, Hogan JW, Carpenter JR; STRATOS initiative. Framework for the treatment and reporting of missing data in observational studies: The Treatment And Reporting of Missing data in Observational Studies framework. J Clin Epidemiol. 2021;134:79–88.\u003c/li\u003e\n\u003cli\u003eLi, X., Mehrotra, D. V., and Barnard, J. Analysis of Incomplete Longitudinal Binary Data Using Multiple Imputation. Statistics in Medicine. 2006;25:2107–2124.\u003c/li\u003e\n\u003cli\u003eLi P, Stuart EA, Allison DB. Multiple Imputation: A Flexible Tool for Handling Missing Data. JAMA. 2015;314(18):1966–1967.\u003c/li\u003e\n\u003cli\u003eLittle RJ, D'Agostino R, Cohen ML, Dickersin K, Emerson SS, Farrar JT, Frangakis C, Hogan JW, Molenberghs G, Murphy SA, Neaton JD, Rotnitzky A, Scharfstein D, Shih WJ, Siegel JP, Stern H. The prevention and treatment of missing data in clinical trials. N Engl J Med. 2012;367(14):1355-60.\u003c/li\u003e\n\u003cli\u003eLittle, Roderick, and Donald Rubin. (2019). Statistical Analysis with Missing Data, Third Edition. Hoboken, NJ, USA: Wiley.\u003c/li\u003e\n\u003cli\u003eLott, A., and Reiter, J. P. Wilson Confidence Intervals for Binomial Proportions With Multiple Imputation for Missing Data. The American Statistician. 2020;74:109–115.\u003c/li\u003e\n\u003cli\u003eNewcombe, R. G. Interval Estimation for the Difference Between Independent Proportions: Comparison of Eleven Methods. Statistics in Medicine. 1998;17:873–890.\u003c/li\u003e\n\u003cli\u003eRubin, D. B. (1987). Multiple Imputation for Nonresponse in Surveys, New York: Wiley.\u003c/li\u003e\n\u003cli\u003eSidi, Y. \u0026amp; Ofer Harel. Difference Between Binomial Proportions Using Newcombe’s Method With Multiple Imputation for Incomplete Data. The American Statistician. 2022;76(1):29–36.\u003c/li\u003e\n\u003cli\u003eSterne JA, White IR, Carlin JB, Spratt M, Royston P, Kenward MG, Wood AM, Carpenter JR. Multiple imputation for missing data in epidemiological and clinical research: potential and pitfalls. BMJ. 2009;338:b2393.\u003c/li\u003e\n\u003cli\u003eWald, A. Tests of Statistical Hypotheses Concerning Several Parameters When the Number of Observations is Large. Transactions of the American Mathematical Society. 1943;54:426–482.\u003c/li\u003e\n\u003cli\u003eWilson, E. B. Probable Inference, the Law of Succession, and Statistical Inference. Journal of the American Statistical Association. 1927;22: 209–212.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Missing data, Multiple Imputation, Difference between Proportions, Confidence Intervals, Newcombe-Wilson method, Agresti-Caffo method, Wald method","lastPublishedDoi":"10.21203/rs.3.rs-8789678/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8789678/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eThe difference between proportions and their confidence intervals are important statistical methods to determine group differences in randomized trials and observational studies. In these studies missing data may severely bias results and should be handled carefully. The Newcombe-Wilson, Wald and Agresti-Caffo methods are commonly used to calculate confidence intervals between proportions. It is unclear how these methods behave in combination with multiple imputation to handle missing data.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eA simulation study was conducted to compare the coverage probability and interval lengths of the Newcombe-Wilson, Wald and Agresti-Caffo methods to calculate the confidence intervals for the difference between proportions after multiple imputation. Evaluated was the performance in MCAR and MAR missing data of 10% and 30%, using large and small differences between proportions and small and large sample sizes.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eIn all simulation scenarios the methods under study performed equally well and were close to the desired coverage probability of 0.95, with similar interval lengths.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eIt can be concluded that simple methods as the Wald and Agresti-Caffo confidence intervals perform similar as a more complex procedure as the Newcombe-Wilson intervals. The Wald intervals showed undercoverage which makes the Agresti-Caffo confidence intervals a strong alternative for the more complex Newcombe-Wilson intervals method.\u003c/p\u003e","manuscriptTitle":"Confidence Intervals for the Difference Between Proportions after Multiple Imputation: A simulation study comparing different strategies","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-02-12 19:11:17","doi":"10.21203/rs.3.rs-8789678/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-02-19T15:14:00+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-19T12:53:15+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-16T01:16:45+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"11253807529632054824042739172838868915","date":"2026-02-14T22:40:59+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-14T16:38:23+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-13T16:13:46+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"263969901783624490392489725287288679075","date":"2026-02-11T20:21:02+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"104805844598014746978612482222957257871","date":"2026-02-11T11:03:02+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"103421024659007203426169494759001803521","date":"2026-02-10T18:58:13+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"133243822661651797466620505309182266380","date":"2026-02-10T09:45:46+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"140824345241938297780252664989950194584","date":"2026-02-09T17:44:22+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"117487980406203607140156610890192343417","date":"2026-02-09T17:39:15+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-02-09T17:12:41+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-02-09T13:57:47+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-02-05T11:36:59+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-02-05T11:34:06+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Research Methodology","date":"2026-02-04T18:37:15+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"f72e371c-6e6a-4167-b642-8765337aed9e","owner":[],"postedDate":"February 12th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[],"tags":[],"updatedAt":"2026-02-19T15:32:05+00:00","versionOfRecord":[],"versionCreatedAt":"2026-02-12 19:11:17","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8789678","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8789678","identity":"rs-8789678","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00