Evaluation of Multiple Imputation Approaches for Handling Missing Covariate Information in a Case-Cohort Study. | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Evaluation of Multiple Imputation Approaches for Handling Missing Covariate Information in a Case-Cohort Study. Melissa Middleton, Cattram Nguyen, Margarita Moreno-Betancur, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-922973/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 9 You are reading this latest preprint version Abstract Background In case-cohort studies a random subcohort is selected from the inception cohort and acts as the sample of controls for several outcome investigations. Analysis is conducted using only the cases and the subcohort, with inverse probability weighting (IPW) used to account for the unequal sampling probabilities resulting from the study design. Like all epidemiological studies, case-cohort studies are susceptible to missing data. Multiple imputation (MI) has become increasingly popular for addressing missing data in epidemiological studies. It is currently unclear how best to incorporate the weights from a case-cohort analysis in MI procedures used to address missing covariate data. Method A simulation study was conducted with missingness in two covariates, motivated by a case study within the Barwon Infant Study. MI methods considered were: using the outcome, a proxy for weights in the simple case-cohort design considered, as a predictor in the imputation model, with and without exposure and covariate interactions; imputing separately within each weight category; and using a weighted imputation model. These methods were compared to a complete case analysis (CCA) within the context of a standard IPW analysis model estimating either the risk or odds ratio. The strength of associations, missing data mechanism, proportion of observations with incomplete covariate data, and subcohort selection probability varied across the simulation scenarios. Methods were also applied to the case study. Results There was similar performance in terms of relative bias and precision with all MI methods across the scenarios considered, with expected improvements compared with the CCA. Slight underestimation of the standard error was seen throughout but the nominal level of coverage (95%) was generally achieved. All MI methods showed a similar increase in precision as the subcohort selection probability increased, irrespective of the scenario. A similar pattern of results was seen in the case study. Conclusions How weights were incorporated into the imputation model had minimal effect on the performance of MI; this may be due to case-cohort studies only having two weight categories. In this context, inclusion of the outcome in the imputation model was sufficient to account for the unequal sampling probabilities in the analysis model. Biochemical Research Methods Multiple Imputation Case-Cohort Study Simulation Study Missing Data Unequal Sampling Probability Inverse Probability Weighting Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Full Text Additional Declarations No competing interests reported. Supplementary Files AdditionalFile1p1.docx AdditionalFile2p1.docx AdditionalFile3p1.docx Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Major revision 29 Oct, 2021 Reviews received at journal 19 Oct, 2021 Reviews received at journal 18 Oct, 2021 Reviewers agreed at journal 07 Oct, 2021 Reviewers invited by journal 07 Oct, 2021 Editor assigned by journal 06 Oct, 2021 Editor invited by journal 05 Oct, 2021 Submission checks completed at journal 05 Oct, 2021 First submitted to journal 20 Sep, 2021 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-922973","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":55600775,"identity":"f21216e6-dfed-4550-91d5-c01292776f04","order_by":0,"name":"Melissa Middleton","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABGElEQVRIiWNgGAWjYFAC5oYDEEaCIZjBD2YzMDA24NTCCNdicACkVLKBCC0MKFoMDqCJowPd9sbGgz/+2AEZyRsOfPxhk2d87fCzDw8YbGQ3HGB+JoFFi9mZgw2HeduSgYxnBQdnJKQVm91OM56RwJBmvOEAmxlWLTcSGw4zNjADGTkGh3kSDiduu53DDHTg4cQNBxiwa7n/sAHosHqIlj9ALZtng7X8B2ph/4bdFmCI8bAdhmhhAGrZIA3WcgCohQe7LWcSQX45zgP2S09aWuIMoF+AoZdsPPMwT7EFNi3HDx/++ONPtZzZ8eSND37Y2CT2z05+zPijwk6273j7xhs4AhoEeND4BkDMjEf9KBgFo2AUjAK8AABd9HP0DXVONQAAAABJRU5ErkJggg==","orcid":"","institution":"University of Melbourne","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Melissa","middleName":"","lastName":"Middleton","suffix":""},{"id":55600776,"identity":"2eb184a7-8c41-492b-9a91-0da363b7473a","order_by":1,"name":"Cattram Nguyen","email":"","orcid":"","institution":"Murdoch Children's Research Institute","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Cattram","middleName":"","lastName":"Nguyen","suffix":""},{"id":55600777,"identity":"d2871ebd-2cf6-4c23-9d0e-ce363daf5440","order_by":2,"name":"Margarita Moreno-Betancur","email":"","orcid":"","institution":"University of Melbourne","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Margarita","middleName":"","lastName":"Moreno-Betancur","suffix":""},{"id":55600778,"identity":"9652bdbc-62ca-4d09-909b-d26792ae1aa3","order_by":3,"name":"John B Carlin","email":"","orcid":"","institution":"Murdoch Children's Research Institute","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"John","middleName":"B","lastName":"Carlin","suffix":""},{"id":55600779,"identity":"f61f0cbb-d753-4dc2-a66a-21c2978959f1","order_by":4,"name":"Katherine J Lee","email":"","orcid":"","institution":"Murdoch Children's Research Institute","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Katherine","middleName":"J","lastName":"Lee","suffix":""}],"badges":[],"createdAt":"2021-09-20 05:44:09","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-922973/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-922973/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":14299919,"identity":"3dbdaa92-ca46-4007-898f-61b80f7087cc","added_by":"auto","created_at":"2021-10-06 20:36:41","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":62641,"visible":true,"origin":"","legend":"Missing data directed acyclic graph (m-DAG) depicting the assumed causal structure between simulated variables and missingness indicators under the dependent missing mechanisms. For the independent missing mechanism, the dashed lines are absent. For simplicity, associations between baseline covariates have not been shown.","description":"","filename":"Fig01.png","url":"https://assets-eu.researchsquare.com/files/rs-922973/v1/bfdc041aa13ffe673d26561c.png"},{"id":14300021,"identity":"e2d699de-7046-4744-b6b0-937c8884887a","added_by":"auto","created_at":"2021-10-06 20:39:41","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":42825,"visible":true,"origin":"","legend":"Relative bias in the coefficient under the extreme scenarios with 30% missing covariate information. Error bars represent 1.96xMonte Carlo standard errors.","description":"","filename":"Fig02.png","url":"https://assets-eu.researchsquare.com/files/rs-922973/v1/0e983a7bcc0d26b9074a8590.png"},{"id":14299918,"identity":"585c5cfe-9e1c-4838-8863-bfb53368a528","added_by":"auto","created_at":"2021-10-06 20:36:41","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":45069,"visible":true,"origin":"","legend":"Empirical standard error and model based standard error under the extreme scenarios with 30% missing covariate information. Error bars represent 1.96xMonte Carlo standard errors.","description":"","filename":"Fig03.png","url":"https://assets-eu.researchsquare.com/files/rs-922973/v1/06919856bf5ad8ba5152c050.png"},{"id":14299922,"identity":"43fce188-5238-4e16-9baa-b466fd5b7140","added_by":"auto","created_at":"2021-10-06 20:36:41","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":40153,"visible":true,"origin":"","legend":"Coverage probability across 2000 simulations under the extreme scenarios with 30% missing covariate information. Error bars represent 1.96xMonte Carlo standard errors.","description":"","filename":"Fig04.png","url":"https://assets-eu.researchsquare.com/files/rs-922973/v1/f418c420a51892941bd8e3c9.png"},{"id":14299924,"identity":"75cb242e-f275-443b-aadd-26d672f83edb","added_by":"auto","created_at":"2021-10-06 20:36:42","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":31303,"visible":true,"origin":"","legend":"Estimated parameter value with 95% confidence interval in case study dataset","description":"","filename":"Fig05.png","url":"https://assets-eu.researchsquare.com/files/rs-922973/v1/afa60b9104a0d5b8c18f83d5.png"},{"id":14300025,"identity":"e5a93927-b1c6-43fd-b70f-af7674b3ec1e","added_by":"auto","created_at":"2021-10-06 20:39:49","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":553918,"visible":true,"origin":"","legend":"","description":"","filename":"00Project1FinalPaper130921.pdf","url":"https://assets-eu.researchsquare.com/files/rs-922973/v1_covered.pdf"},{"id":14299925,"identity":"de1ad18d-20e4-46d5-bcb7-5f4739f795ea","added_by":"auto","created_at":"2021-10-06 20:36:42","extension":"docx","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":24769,"visible":true,"origin":"","legend":"","description":"","filename":"AdditionalFile1p1.docx","url":"https://assets-eu.researchsquare.com/files/rs-922973/v1/bf4accaa342294abdd9617db.docx"},{"id":14299921,"identity":"85aadd96-2617-482f-ba25-22ec9562204d","added_by":"auto","created_at":"2021-10-06 20:36:41","extension":"docx","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":26744,"visible":true,"origin":"","legend":"","description":"","filename":"AdditionalFile2p1.docx","url":"https://assets-eu.researchsquare.com/files/rs-922973/v1/60d0c4e9f7fd513065848d3a.docx"},{"id":14299923,"identity":"4845f7b2-b891-45a7-abe9-d8f0f7c787f4","added_by":"auto","created_at":"2021-10-06 20:36:42","extension":"docx","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":174218,"visible":true,"origin":"","legend":"","description":"","filename":"AdditionalFile3p1.docx","url":"https://assets-eu.researchsquare.com/files/rs-922973/v1/e5cb74fbf6b23192fede485a.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"\u003cp\u003eEvaluation of Multiple Imputation Approaches for Handling Missing Covariate Information in a Case-Cohort Study.\u003c/p\u003e","fulltext":[{"header":"Full Text","content":"This preprint is available for \u003ca href='/article/rs-922973/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e."}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Multiple Imputation, Case-Cohort Study, Simulation Study, Missing Data, Unequal Sampling Probability, Inverse Probability Weighting","lastPublishedDoi":"10.21203/rs.3.rs-922973/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-922973/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e In case-cohort studies a random subcohort is selected from the inception cohort and acts as the sample of controls for several outcome investigations. Analysis is conducted using only the cases and the subcohort, with inverse probability weighting (IPW) used to account for the unequal sampling probabilities resulting from the study design. Like all epidemiological studies, case-cohort studies are susceptible to missing data. Multiple imputation (MI) has become increasingly popular for addressing missing data in epidemiological studies. It is currently unclear how best to incorporate the weights from a case-cohort analysis in MI procedures used to address missing covariate data.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eMethod\u003c/strong\u003e A simulation study was conducted with missingness in two covariates, motivated by a case study within the Barwon Infant Study. MI methods considered were: using the outcome, a proxy for weights in the simple case-cohort design considered, as a predictor in the imputation model, with and without exposure and covariate interactions; imputing separately within each weight category; and using a weighted imputation model. These methods were compared to a complete case analysis (CCA) within the context of a standard IPW analysis model estimating either the risk or odds ratio. The strength of associations, missing data mechanism, proportion of observations with incomplete covariate data, and subcohort selection probability varied across the simulation scenarios. Methods were also applied to the case study.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eResults \u003c/strong\u003eThere was similar performance in terms of relative bias and precision with all MI methods across the scenarios considered, with expected improvements compared with the CCA. Slight underestimation of the standard error was seen throughout but the nominal level of coverage (95%) was generally achieved. All MI methods showed a similar increase in precision as the subcohort selection probability increased, irrespective of the scenario. A similar pattern of results was seen in the case study.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConclusions \u003c/strong\u003eHow weights were incorporated into the imputation model had minimal effect on the performance of MI; this may be due to case-cohort studies only having two weight categories. In this context, inclusion of the outcome in the imputation model was sufficient to account for the unequal sampling probabilities in the analysis model.\u003c/p\u003e","manuscriptTitle":"Evaluation of Multiple Imputation Approaches for Handling Missing Covariate Information in a Case-Cohort Study.","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2021-10-06 20:36:39","doi":"10.21203/rs.3.rs-922973/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2021-10-29T12:27:50+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2021-10-19T09:01:33+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2021-10-18T20:39:45+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"b64ac384-af4e-4f7f-8bcd-ce9f2b9d635e","date":"2021-10-08T02:51:26+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2021-10-07T09:05:05+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2021-10-06T05:46:11+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2021-10-05T20:56:28+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2021-10-05T14:36:50+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Research Methodology","date":"2021-09-20T05:42:14+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-medical-research-methodology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bmrm","sideBox":"Learn more about [BMC Medical Research Methodology](http://bmcmedresmethodol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bmrm/default.aspx","title":"BMC Medical Research Methodology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ee3f2538-6aaf-46c6-a16f-f6132178b3f7","owner":[],"postedDate":"October 6th, 2021","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":7687866,"name":"Biochemical Research Methods"}],"tags":[],"updatedAt":"2021-12-15T02:14:03+00:00","versionOfRecord":[],"versionCreatedAt":"2021-10-06 20:36:39","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-922973","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-922973","identity":"rs-922973","version":["v1"]},"buildId":"ApUGefWb6u5IBVtyqm6d5","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.