Imputation Validation Mean Squared Error: a novel measurement to evaluate the performance of missing value imputation
preprint
OA: closed
CC-BY-4.0
Abstract
Missing value imputation is essential in data pre-processing for statistical model building and data mining. The problem with using a machine learning model to impute missing values is that it is hard to evaluate the performance of imputation and choose the optimal model and tuning parameters. Although many missing values imputation models have been proposed, the precise understanding of the imputation performance evaluation is missing in the literature. Therefore, this study aims to develop a new method named ‘imputation validation mean square error’ which would potentially improve the knowledge of the missing values imputation. The imputation validation dataset is obtained by randomly assigning some attributes as NA in the complete dataset according to the information on the percentage of missing values for each attribute and how often each combination of missing values occurs. This literature compares four imputation models (Clustering, KNN, PCA, and MCMC) for multivariate missing values imputation on the Pima Indian Diabetes dataset. The imputation validation MSE range is from 0 to 1. These four methods with imputation validation MSE equal to 0.1447, 0.1360, 0.1129, and 0.1312 are better than the mean (0.2270) and median (0.2306) imputation methods. The performance of different imputation models can be compared clearly based on this method.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00
- unpaywall
- last seen: 2026-05-29T02:00:03.542394+00:00
License: CC-BY-4.0