Robust Multipe Imputation with GAM | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Robust Multipe Imputation with GAM Matthias Templ This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3576079/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 22 May, 2024 Read the published version in Statistics and Computing → Version 1 posted You are reading this latest preprint version Abstract Multiple imputation of missing values is a key step in data analytics and a standard process in data mining. Non-linear imputation methods ones comes into play whenever the linear relationship between a response and predictors cannot be linearized. One kind of popular non-linear methods are Generalized Additive Models (GAM) and an extension of GAM, namely GAMLSS, where each parameter of the distribution (e.g., mean, variance, skewness, kurtosis) can be modeled as a function of predictors. However, non-robust methods such as standard GAM's and GAMLSS's can be swayed by outliers, leading to outlier-driven imputations. This can apply concerning both representative outliers - those true yet unusual values of your population - and non-representative outliers, which are mere measurement errors. Robust (imputation) methods effectively manage outliers and exhibit remarkable resistance to their influence, providing a more reliable approach to dealing with missing data. A new robust imputation algorithm is introduced. This innovative solution addresses three significant challenges with robustness. (1) It uses a robust bootstrap to manage model uncertainty when imputing a random sample, (2) it incorporates robust fitting to reinforce accuracy, and (3) it takes into account imputation uncertainty in a resilient manner. Furthermore, any complex model for any variable with missingness can be considered and run through the algorithm. For the employed real-world datasets and the conducted simulation study, the novel algorithm imputeRobust demonstrates superior performance in comparison to other prevalent methods. Missing values Multiple imputation Generalized additive models Robust estimation Robust bootstrap Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 22 May, 2024 Read the published version in Statistics and Computing → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3576079","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":247246918,"identity":"cc4aa681-3dc5-4257-81a5-4f1656919329","order_by":0,"name":"Matthias Templ","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA6klEQVRIiWNgGAWjYJACZgYDGLPAhh/MZmzAq4GxGayFDcQxSJPcQJwWBriWw4S16LafP/64oMDOnkG++ejmAoPzEubSPQYMP3fg1mJ2JpmxeYZBcmIDG1va7RkGtyUs55wxYOw9g0fLAaAWHgPmBAY2HrPbPAa36wxu5G5gZmzDo+X8Y5CWensGNv5vQC3nJAhruQG25TBjAxsPG1DLAWK0PDaczWNwPLGNLc0M6JdkoJb8Dwd78Tos8cFnnj/V9vzMh5/dLqiwA2pJS3zwE48WOABFCzOMc4AIDRDATFjJKBgFo2AUjEQAAMqRTq/pzE+YAAAAAElFTkSuQmCC","orcid":"","institution":"University of Applied Sciences and Arts Northwestern Switzerland","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Matthias","middleName":"","lastName":"Templ","suffix":""}],"badges":[],"createdAt":"2023-11-07 19:14:10","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3576079/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3576079/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s11222-024-10429-1","type":"published","date":"2024-05-22T07:12:13+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":60209849,"identity":"dd481a11-5dd2-4a55-8525-ce4a5d877284","added_by":"auto","created_at":"2024-07-13 07:12:18","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":643422,"visible":true,"origin":"","legend":"","description":"","filename":"paperijdsasubmissionfiles.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3576079/v1_covered_cc604652-08cd-43a8-b639-59869730cd4d.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Robust Multipe Imputation with GAM","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Missing values, Multiple imputation, Generalized additive models, Robust estimation, Robust bootstrap","lastPublishedDoi":"10.21203/rs.3.rs-3576079/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3576079/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eMultiple imputation of missing values is a key step in data analytics and a standard process in data mining. Non-linear imputation methods ones comes into play whenever the linear relationship between a response and predictors cannot be linearized. One kind of popular non-linear methods are Generalized Additive Models (GAM) and an extension of GAM, namely GAMLSS, where each parameter of the distribution (e.g., mean, variance, skewness, kurtosis) can be modeled as a function of predictors.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eHowever, non-robust methods such as standard GAM's and GAMLSS's can be swayed by outliers, leading to outlier-driven imputations. This can apply concerning both representative outliers - those true yet unusual values of your population - and non-representative outliers, which are mere measurement errors.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eRobust (imputation) methods effectively manage outliers and exhibit remarkable resistance to their influence, providing a more reliable approach to dealing with missing data.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eA new robust imputation algorithm is introduced. This innovative solution addresses three significant challenges with robustness. (1) It uses a robust bootstrap to manage model uncertainty when imputing a random sample, (2) it incorporates robust fitting to reinforce accuracy, and (3) it takes into account imputation uncertainty in a resilient manner. Furthermore, any complex model for any variable with missingness can be considered and run through the algorithm.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFor the employed real-world datasets and the conducted simulation study, the novel algorithm imputeRobust demonstrates superior performance in comparison to other prevalent methods.\u003c/p\u003e","manuscriptTitle":"Robust Multipe Imputation with GAM","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-11-10 19:46:34","doi":"10.21203/rs.3.rs-3576079/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"a1baffd4-62ec-464f-907b-5dcd2d56548d","owner":[],"postedDate":"November 10th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-07-13T07:12:13+00:00","versionOfRecord":{"articleIdentity":"rs-3576079","link":"https://doi.org/10.1007/s11222-024-10429-1","journal":{"identity":"statistics-and-computing","isVorOnly":false,"title":"Statistics and Computing"},"publishedOn":"2024-05-22 07:12:13","publishedOnDateReadable":"May 22nd, 2024"},"versionCreatedAt":"2023-11-10 19:46:34","video":"","vorDoi":"10.1007/s11222-024-10429-1","vorDoiUrl":"https://doi.org/10.1007/s11222-024-10429-1","workflowStages":[]},"version":"v1","identity":"rs-3576079","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3576079","identity":"rs-3576079","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.