A Novel Algorithm for Imputing the Missing Values in Incomplete Datasets

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-17

This paper introduces the novel IMV-RE algorithm, which sets an upper limit for classes with missing values, outperforming existing methods in accuracy, RMSE, and R² on benchmark datasets.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

In today’s world, we completely rely on digital devices to collect data; a failure in such digital devices may result in huge information loss thereby making data mining a more tedious job for a Data Analyst. Missingness to a greater extent in a dataset subsequently comes out with inappropriate results and incomplete data analysis. Therefore, a need to develop an algorithm that can predict the missing values efficiently and accurately. This research paper proposes a novel splitting-based IMV-RE ( I mputing the M issing V alues in R eal-Time E nvironment) algorithm to impute different missing values within a dataset. In the proposed IMV-RE algorithm, an upper limit is set for every class containing missing values that assist the algorithm to predict the missing values more accurately. The experimentation is performed on ten benchmark datasets that include completely numerical values as well as mixed data. Comparative experimental analysis indicates that the proposed IMV-RE algorithm outperforms the existing techniques in sensitivity to Accuracy, Root Mean Square Error (RMSE) and Coefficient of Determination (R 2 ).
Full text 137,695 characters · extracted from preprint-html · click to expand
A Novel Algorithm for Imputing the Missing Values in Incomplete Datasets | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Novel Algorithm for Imputing the Missing Values in Incomplete Datasets Hutashan Vishal Bhagat, Manminder Singh This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1729251/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract In today’s world, we completely rely on digital devices to collect data; a failure in such digital devices may result in huge information loss thereby making data mining a more tedious job for a Data Analyst. Missingness to a greater extent in a dataset subsequently comes out with inappropriate results and incomplete data analysis. Therefore, a need to develop an algorithm that can predict the missing values efficiently and accurately. This research paper proposes a novel splitting-based IMV-RE ( I mputing the M issing V alues in R eal-Time E nvironment) algorithm to impute different missing values within a dataset. In the proposed IMV-RE algorithm, an upper limit is set for every class containing missing values that assist the algorithm to predict the missing values more accurately. The experimentation is performed on ten benchmark datasets that include completely numerical values as well as mixed data. Comparative experimental analysis indicates that the proposed IMV-RE algorithm outperforms the existing techniques in sensitivity to Accuracy, Root Mean Square Error (RMSE) and Coefficient of Determination (R 2 ). Imputation Imputing values Missingness Mechanisms Missing Values Data Missingness Data Imputation Model Incomplete datasets Root Mean Square Error Figures Figure 1 Figure 2 1. Introduction We are living in a digital world where information can easily be acquired with the help of smart devices, sensors etc. The advent of IoT makes it possible to collect data without the physical intervention of humans. But sometimes, a failure in such devices may result in data loss and hence, affects the subsequent in-depth analysis and data interpretation that provides erroneous results. The rapid increase in the size of datasets has led to emerging of various data mining techniques. To ensure data mining results to be effective and valuable, data scientists must ensure the quality of the collected data. In real-time scenarios, it is usually the case that collected datasets for data analysis may contain some missing values [ 1 ]. Therefore, it is not possible for most of the data mining algorithms to directly handle these incomplete datasets. The simplest solution to such problems is case deletion which means removing data having missing values. However, case deletion can be appropriate if the missing rate is small, e.g., 5% and if the missing rate is somehow larger, say 25%, then using 75% of the original dataset might be insufficient to completely reflect the real-world problem, which could also affect the mining results [ 2 ]. To ensure the quality of collected datasets, data scientists first pre-process the datasets [ 3 ]. In concern to the relationship between the missing data and other data values of the variables in a dataset, missing data mechanism can be assorted as Missing at Random (MAR), Missing Completely at Random (MCAR) and Missing Not at Random (MNAR) [ 4 , 5 ]. If the probability of missing data depends upon the observed responses and there exist no relation among the missing values itself, the missing data mechanism is referred to as MAR. Hence, in MAR the reason for missingness in a feature Q of a dataset mainly depends upon the rest of the features within the dataset rather than the Q itself. If there is no relation between the missing values and the set of observed responses, the missingness mechanism is said to be MCAR. In MCAR, the reason for missing in a feature Q depends neither on the other features within the dataset nor on Q itself. If the probability of the missingness in a feature Q depends either on Q itself or on the other features that also contains missing values, the missing mechanism is referred to as MNAR. Missing value imputation (MVI) techniques provide the best solutions to impute the missing values within the datasets. MVI techniques can be categorized into two categories- MVI techniques based on Statistical Methods (Mean, Mode, Median and Regression) and MVI techniques based on Machine Learning (Neural Networks, Support Vector Machines and Deep Learning Models) [ 6 ]. Although, machine learning techniques being complex in nature produce better imputation results, yet they are computationally expensive. For high dimensional datasets, such techniques show high computation time as compared to statistical methods because of which they are not deployed to mission-critical systems [ 7 ]. Being inspired from the lower complexity and computation time of statistical MVI techniques, this paper proposed a novel splitting-based imputation approach, namely, IMV-RE. To justify the superiority of the proposed IMV-RE approach, ten benchmark datasets containing both numerical as well as mixed values with varying dimensionality take into consideration and is compared with five existing imputation techniques. Classification accuracy, RMSE and Coefficient of determination (R 2 ) are considered as the performance metrics with which the proposed IMV-RE approach is evaluated. 2. Literature Review Missingness within a dataset is a very common problem in statistical analysis. Researchers have come out with numerous MVI techniques depending upon the missing rate, data characteristics and pattern and data correlation. This section gives a brief description of the recent MVI techniques. Regularized Expectation Maximization (EM) algorithm for imputing missing value which is based on the iterated analysis of the linear regressions among the missing and non-missing variables is proposed in [ 8 ]. The experimental test is carried out using regularized EM technique over the climate data. The regularized EM technique is further enhanced in [ 9 ] and two novel methods kEMI and kEMI + are proposed by the authors. The proposed techniques are based on the information fusion mechanism that uses Dempster-Shafer fusion to fuse the most appropriate estimates. Due to the high computation time of the missForest [ 10 ], the authors in [ 11 ] proposed a new technique known as mForest that can achieve ten times less computation time than missForest. A Column-wise Guided Data Imputation (cGDI) technique that divides the complete samples from the incomplete samples and selects the most suitable imputation method separately for each feature is proposed in [ 12 ]. In [ 13 ], the authors proposed a CBC-IM (Class Based Clustering approach for Imputation) technique for missing values imputation especially for medical datasets. CBC-IM technique first partitioned the dataset into complete and incomplete variables and then uses the Euclidean distance and Fuzzy measure to find out the similarity between the two records. Another similar approach known as CCMVI (Class Center based Missing Value Imputation) is proposed in [ 7 ]. CCMVI is a two-step process that first defines the threshold value for every class containing missing values and then finds the appropriate estimate from the complete samples considering the threshold values. An approach based on the linear regression technique known as CLR (Cumulative Linear Regression) is proposed in [ 14 ]. The incomplete variables are first cumulated and incorporated in the linear regression equation to replace the missing values in the next incomplete variable. The KNNI (K Nearest Neighbor Imputation) is the most extensively used MVI technique. In [ 15 ], the authors proposed a feature weighted grey KNNI technique that uses the combination of relevant feature information and grey rational based k nearest neighbors to impute the missing values. Another approach that makes combine use of Multi-Layer Perceptron and KNN to impute the multiple missing values simultaneously is proposed in [ 16 ]. This combination results in an increase in the performance with an increase in the time complexity of the algorithm. In [ 17 ], the authors proposed two techniques CBRL (Cumulative Bayesian Ridge with Less NaN) and CBRC (Cumulative Bayesian Ridge with high Correlation). CBRL has used the most appropriate features within the dataset that contain a lesser number of missing values whereas CBRC is used to select the best feature that gives a high correlation with the target feature. Correlation Maximization-based Imputation Methods (CMIM) are proposed in [ 18 ] that first find the highly correlated segments of the data and use linear regression estimator to impute the missing values. The authors use a two-step method to handle large missing gaps in [ 19 ] known as Ratio-Based Imputation (RBI). In RBI the MVI is done by using machine learning models whereas the analysis is done by data fusion technique in CPS (Cyber-Physical Systems) datasets. Considering the missing values in a dataset an optimization problem, the authors in [ 20 ] proposed an optimal method known as BNII. The BNII technique is a two-stage approach: firstly, using the Bayesian Network relationship among different attributes is calculated and secondly, in an iterative manner imputation is done till local maximum posterior probability is reached. A novel architecture based on Particle Swarm Optimization (PSO) for cleaning the data and then utilizing the K-means to calculate the fitness value as well as to narrow down the search space is proposed in [ 21 ]. The ontology makes PSO replace missing values more accurately but in worst scenarios, this approach has a time complexity of Θ(n 3 ). A novel method known as Modulo 9 proposed in [ 22 ] impute the missing values within the interval of [0–9] and then use congruency with addition and multiplication to make an appropriate estimation. The authors compared the Modulo 9 approach with the eleven robust MVI techniques and outperform all of them. Considering generalization for datasets that contain mixed-type data, the authors in [ 23 ] proposed a tuple-oriented region splitting imputation technique known as RESI (Region-Splitting Imputation). RESI technique first uses the entropy weight method to assign weights to the attributes and split the data into complete and incomplete subsets based on their integrity rate. The model is trained over a complete subset that iteratively imputes the next incomplete subset. From the last few years, researchers have come out with numerous imputation techniques based on meta-heuristic techniques ([ 24 – 27 ],association mining rules[ 28 , 29 ], dynamic programming techniques[ 30 ] and various such hybrid techniques [ 31 – 34 ] to efficiently impute the missing values. The authors in [ 35 ] analyzed that there is no such MVI technique that could be considered as a master technique for distinct problems. The key factors that can influence the performance of MVI techniques are the data distribution within the dataset, the missingness rate and characteristics of a dataset [ 36 ]. The main contribution of this paper is to overcome the limitations of state-of-the-art imputation techniques based on the statistical methods. The authors proposed a CBC-IM [ 13 ] (Class Based Clustering approach for Imputation) technique for missing values imputation especially for medical datasets. CBC-IM technique first partitioned the dataset into complete and incomplete variables and then uses the Euclidean distance and Fuzzy measure to find out the similarity between the two records. Another similar approach known as CCMVI is proposed [ 7 ]. CCMVI is a two-step process that first defines the threshold value for every class containing missing values and then finds the appropriate estimate from the complete samples considering the threshold values. Although, the CCMVI and IM-CBC techniques have less time complexity, yet their dependency on complete samples make them unsuitable for the imputation especially when a dataset has high missingness rate or when there is an unavailability of at least one data sample without missing values within a class. The major goal of any MVI technique is to replace the missing values in such a manner that the overall data integrity, structure, and trends of the data are maintained. The proposed IMV-RE MVI technique shows better performance than the commonly used existing single as well as multiple imputation techniques. Unlike the CCMVI technique, the IMV-RE technique successfully imputes the missing values even the missing rate is high. Unlike the missForest [ 10 ], MICE (Multiple Imputation using Chained Equations) and other meta-heuristic techniques [ 24 – 27 ], the proposed IMV-RE has low time complexity. Because of the simplicity of the proposed IMV-RE algorithm, it is significant for high dimensional datasets also. 3. Imputing The Missing Values In Real-time Environment The proposed IMV-RE is a splitting-based data imputation algorithm that is the robust statistical MVI technique. The proposed algorithm first decides the upper limit for each feature within a cluster which is further used for imputing the missing values so that the imputed values lie much nearer to the exact values. The following sub-section discussed the step-by-step process of the proposed imputation algorithm. 3.1 Imputation Algorithm (IMV-RE) Figure 1 shows the whole process from start to the end of IMV-RE algorithm and each step is described as below: Step 1 Initially a dataset X D is given that has D dimensions and C N number of classes. Step 2 Cluster the data samples into their respective classes along with the missing data samples using the labels given in case of labeled datasets and for unlabeled datasets, K-Means clustering is used to cluster the similar data samples such that C 1 , C 2 , C 3 , …… C N numbers of classes are there. Step 3 For each class C i (where i = 1 to N ), replace all the missing values with zero. Calculate the centered value ( Cent i ) and standard deviation ( SDev i ) for class C i . Step 4 Calculate the L 2 norm between the Cent i and each data sample in class C i . Select the mid-point value as an upper limit ( UpLimC i ) for class C i . Step 5 For each class C i obtained from the Step 2 , separate all the data samples containing missing values from the complete data samples into a cluster such that cluster C i contains L miss_sam numbers of missing samples. Step 6 For each missing sample L miss_sam containing N_miss number of missing values, impute each missing value with Cent i calculated for the class C i in Step 3 . Now, calculate the L 2 norm between the Cent i and the imputed L miss_sam sample. If the value is smaller than the UpLimC i (calculated in Step 4 ), then imputed value is finalized else +/- standard deviation ( SDev i ) is applied to new imputed value in a sequential order and corresponding the L 2 norm between the Cent i and the imputed value is calculated. Step 7 The final value to be imputed is the value in correspond to smaller value calculated for the L 2 norm between the Cent i and the new imputed L miss_sam sample. 4. Experimental Implementation This section is divided into three subsections. The first subsection represents the benchmark datasets, the second subsection represents the experimental setup and the third subsection represents the metrics used to evaluate the proposed IMV-RE algorithm. 4.1 Datasets Table 1 depicts the brief description of the datasets used for the experimentation. A total of ten datasets, that includes 8 numerical and 2 of mixed data type, from the UCI machine learning repository are collected with the number of samples ranging from 70 to 5000, the number of features ranging from 4 to 206 and the number of classes ranging from 2 to 10. 4.2 Experimental Setup The MCAR (Missing Completely at Random) missing value mechanism is used to intentionally add missing values in the above datasets with missingness rates of 10%, 20%, 30%, 40% and 50% of the total data within datasets. A fixed random seed is used for generating the missing values in every dataset to avoid biased results. The proposed IMV-RE approach is used to impute the missing values and is compared with five baseline approaches: Class Center based Missing Value Imputation (CCMVI), K-Nearest Neighbor Imputation (KNNI), MiceForest [ 37 ], IterativeImputer [ 38 ]and SimpleImputer [ 39 ]. The experimental work is performed on python IDE Spyder version 5.2.1 in Anaconda Navigator usinglaptop PC DELL G5, Intel Core i7 processor, RAM 12GB, 512GB SSD. Table 1 Datasets Description Datasets Datasets type No. of instances No. of features No. of classes Waveform Numerical 5000 21 3 Glass Numerical 214 9 7 Wheat-Seed Numerical 210 7 3 Digits Numerical 1797 64 10 Wine Numerical 178 13 3 Iris Numerical 150 4 3 Seeds Numerical 210 7 3 Ionosphere Numerical 351 34 2 SCADI Mixed 70 206 7 Ecoli Mixed 336 7 8 4.3 Evaluation Metrics A stratified 5-fold cross-validation using RandomForest Classifier is performed to evaluate the classification accuracy on the imputed datasets. For every dataset, there are five different missing rates that correspond to five different classification accuracies and average of five accuracies are taken into consideration for comparison among different imputation techniques. In addition to the classification accuracy, two more performance metrics are used to evaluate the performance of the proposed IMV-RE approach. Root Mean Square Error (RMSE): If x i and \(\widehat{x}\) are the original value and the imputed value of the i th observation respectively, n is total number of samples, then, RMSE is given by the equation: $$RMSE=\sqrt{\frac{{\sum }_{i=1}^{n}{\left({x}_{i}-\widehat{{x}_{i}}\right)}^{2}}{n}}$$ 1 Coefficient of Determination (R 2 ): Calculated using the following equation: $${ R}^{2}\left(x, \widehat{x}\right)=1-\frac{{\sum }_{i=1}^{n}{\left({x}_{i}-\widehat{{x}_{i}}\right)}^{2}}{{\sum }_{i=1}^{n}{\left({x}_{i}-\stackrel{-}{x}\right)}^{2}}$$ 2 where, $$\stackrel{-}{x}= \frac{1}{n}\sum _{i=1}^{n}{x}_{i}$$ Moreover, the mean and standard deviation of the imputed datasets are compared with the mean and standard deviation of the actual datasets and percentage error is calculated for all datasets. 5. Results And Discussions Three metrics are used to evaluate the performance of the proposed IMV-RE algorithm. The IMV-RE algorithm is applied to ten benchmark datasets and the classification results obtained using RandomForest Classifier, RMSE values and Coefficient of determination (R 2 ) values are compared to CCMVI, KNNI, MiceForest, IterativeImputer, SimpleImputer MVI techniques. Table 2 depicts the average classification accuracies obtained on the ten benchmark datasets Waveform, Glass, Wheat-Seed, Digits, Wine, Iris, Seeds, Ionosphere, SCADI and Ecoli having 10%, 20%, 30%, 40% and 50% missing rates. The average accuracy of the proposed IMV-RE algorithm achieved for all the ten datasets is 91.58% which is better than the other five MVI approaches MiceForest (81.91%), KNNI (80.95%), SimpleImputer (81.7%) and IterativeImputer (82.81%) For Waveform, Glass, Wheat-Seed, Digits and Wine datasets the CCMVI technique is not applicable beyond 10% missing rate because of its dependency on the samples without missing values to calculate the threshold value that is utilized to estimate the missing values in a particular class. Similarly, Table 3 compares the average RMSE values obtained for ten datasets. The results show that the proposed IMV-RE algorithm achieves the lowest RMSE value 0.73 in comparison to the other five MVI techniques. RMSE for CCMVI technique cannot be calculated because of the same above-mentioned reason. Table 4 compares the average value of coefficient of determination obtained for all the ten datasets. The proposed IMV-RE algorithm achieves average Coefficient of Determination value 0.886 which is higher than all other MVI techniques. Table 2 Average classification accuracies obtained from distinct MVI techniques Datasets IMV-RE MiceForest KNNI CCMVI SimpleImputer IterativeImputer Waveform 0.96443 0.78015 0.76447 - 0.77695 0.79092 Glass 0.80498 0.54197 0.52140 - 0.55801 0.57763 Wheat-Seed 0.95944 0.78583 0.78957 - 0.81168 0.82681 Digits 0.98664 0.77682 0.73919 - 0.65356 0.71446 Wine 0.95883 0.95197 0.93429 0.96990 0.96860 0.96537 Iris 0.97333 0.92267 0.90667 0.97267 0.92933 0.94800 Seeds 0.89143 0.87143 0.88095 0.90286 0.87714 0.87810 Ionosphere 0.91851 0.91059 0.90365 0.92029 0.91285 0.90766 SCADI 0.85275 0.84440 0.83846 0.84725 0.85626 0.84703 Ecoli 0.84753 0.80530 0.81605 0.85595 0.82558 0.82551 Average 0.91579 0.81911 0.80947 - 0.81700 0.82815 Table 3 Average RMSE obtained from distinct MVI techniques Datasets IMV-RE MiceForest KNNI CCMVI SimpleImputer IterativeImputer Waveform 0.68948 0.70871 0.73801 - 0.80543 0.71955 Glass 0.32751 0.37197 0.38146 - 0.38244 0.39661 Wheat-Seed 0.32499 0.46663 0.48078 - 0.63877 0.42437 Digits 2.02541 2.07970 2.25400 - 2.60753 2.53113 Wine 3.54842 4.65323 4.98252 3.79756 7.21529 6.15099 Iris 0.15431 0.21552 0.21250 0.17177 0.40314 0.21158 Seeds 0.14982 0.11402 0.15104 0.15284 0.30043 0.13956 Ionosphere 0.07780 0.05712 0.04592 0.07728 0.06871 0.06354 SCADI 0.00102 0.00082 0.00041 0.00100 0.00182 0.00159 Ecoli 0.02252 0.03085 0.02623 0.02361 0.02984 0.02654 Average 0.73213 0.86986 0.92729 - 1.24534 1.06655 Table 4 Average Coefficient of Determination (R 2 ) obtained from distinct MVI techniques Datasets IMV-RE MiceForest KNNI CCMVI SimpleImputer IterativeImputer Waveform 0.78181 0.75131 0.73399 - 0.71399 0.76373 Glass 0.71782 0.63663 0.60446 - 0.65406 0.65720 Wheat-Seed 0.85907 0.68851 0.71047 - 0.57340 0.77365 Digits 0.67990 0.67332 0.56739 - 0.53716 0.50253 Wine 0.96225 0.96254 0.93562 0.95724 0.93715 0.95333 Iris 0.95632 0.93200 0.92954 0.94407 0.82545 0.93378 Seeds 0.97092 0.97381 0.96951 0.96438 0.90623 0.95981 Ionosphere 0.97309 0.98352 0.98874 0.97279 0.97986 0.98238 SCADI 0.99958 0.99940 0.99981 0.99929 0.99934 0.99948 Ecoli 0.96658 0.94247 0.94557 0.96319 0.95191 0.95639 Average 0.88673 0.85435 0.83851 - 0.80785 0.84823 Figure 2 graphically represents the performance comparison of proposed IMV-RE algorithm over different evaluation metrics. Table 5 Percentage Error between mean values of actual datasets and imputed datasets Datasets Ecoli Glass Wheat-Seed Digit Wine Iris Seed Ionosphere SCADI Waveform Actual Mean 0.4996 0.0968 6.8967 4.8843 0.0437 3.4645 0.0820 0.2477 0.2033 1.7123 Imputed Mean 0.4988 0.0790 6.9166 4.8887 0.0371 3.4657 0.0749 0.2438 0.2035 1.7153 Percentage Error 0.1520 18.3642 0.2890 0.0890 15.0280 0.0332 8.5944 1.5801 0.0914 0.1745 Table 6 Percentage Error between calculated standard deviation of actual datasets and imputed datasets Datasets Ecoli Glass Wheat-Seed Digit Wine Iris Seed Ionosphere SCADI Waveform Actual Stdev 0.144 1.048 1.010 3.684 0.702 0.948 0.632 0.510 0.263 1.520 Imputed Stdev 0.156 0.961 5.311 5.549 0.704 1.968 0.634 0.573 0.951 1.758 Percentage Error 8.470 8.343 425.986 50.616 0.366 107.660 0.375 12.176 261.444 15.672 Table 5 represents the percentage error calculated between the mean values of actual datasets and imputed datasets by considering the average of total mean value calculated at different missing rates 10%, 20%, 30%, 40% and 50% using proposed IMV-RE algorithm. Similarly, the Table 6 represents the percentage error calculated between the standard deviation of actual datasets and imputed datasets by considering the average of total standard deviation value calculated at different missing rates 10%, 20%, 30%, 40% and 50% using proposed IMV-RE algorithm. 6. Conclusion In this paper, a novel splitting-based IMV-RE algorithm is proposed. The proposed IMV-RE algorithm is a two-step process. In the first step, an upper limit is calculated for every class containing missing values and the second step utilizes the upper limit to impute missing values efficiently. The IMV-RE algorithm only searches within the class to calculate the centered value for the final imputation, unlike the other techniques that need to go through the whole data within the dataset. Hence, the proposed algorithm can be utilized for real-time problems. The proposed IMV-RE algorithm prosperously abstracts the dependency of the CCMVI technique on the complete samples within a class to estimate the missing values. As a result, the proposed IMV-RE algorithm successfully imputes the missing values for higher data missing. Classification accuracy, RMSE and coefficient of determination (R 2 ) are the evaluation metrics used to evaluate the performance of the proposed IMV-RE algorithm. The experimental results depict better performance of the proposed IMV-RE algorithm for imputing the missing values under distinct missing rates in comparison to the other state-of-the-art techniques. Declarations Conflict of Interest The authors of this publication declare there is no conflict of interest. Funding This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Code/Data Availability N/A. Authors' Contributions Hutashan Vishal Bhagat: Conceptualization, Methodology, Software, Formal analysis, Writing - Original Draft, Validation, Data Curation and Visualization. Manminder Singh: Writing Review, Editing and Supervision References Kang, H. (2013). The prevention and handling of the missing data.Korean journal of anesthesiology, 64(5),402. https://doi.org/10.4097/kjae.2013.64.5.402 Kalkan, Ö. K., Yusuf, K. A. R. A., & Kelecioğlu, H. (2018). Evaluating performance of missing data imputation methods in IRT analyses.International Journal of Assessment Tools in Education, 5(3),403–416. https://doi.org/10.21449/ijate.430720 García, S., Luengo, J., & Herrera, F. (2015). Data preprocessing in data mining (Vol. 72, pp.59–139). Cham, Switzerland:Springer International Publishing Kelkar, B. A. (2022). Missing Data Imputation: A Survey. International Journal of Decision Support System Technology (IJDSST) , 14 (1), 1–20. DOI: 10.4018/IJDSST.292446 Little, R. J., & Rubin, D. B. (2019). Statistical analysis with missing data (Vol. 793).John Wiley & Sons Baraldi, A. N., & Enders, C. K. (2010). An introduction to modern missing data analyses. Journal of school psychology , 48 (1), 5–37 Tsai, C. F., Li, M. L., & Lin, W. C. (2018). A class center based approach for missing value imputation. Knowledge-Based Systems , 151 , 124–135 Schneider, T. (2001). Analysis of incomplete climate data: Estimation of mean values and covariance matrices and imputation of missing values. Journal of climate , 14 (5), 853–871 Razavi-Far, R., Cheng, B., Saif, M., & Ahmadi, M. (2020). Similarity-learning information-fusion schemes for missing data imputation. Knowledge-Based Systems , 187 , 104805 Probst, P., Wright, M. N., & Boulesteix, A. L. (2019). Hyperparameters and tuning strategies for random forest. Wiley Interdisciplinary Reviews: data mining and knowledge discovery , 9 (3), e1301 Tang, F., & Ishwaran, H. (2017). Random forest missing data algorithms. Statistical Analysis and Data Mining: The ASA Data Science Journal , 10 (6), 363–377 Petrozziello, A., & Jordanov, I. (2017). Column-wise guided data imputation. Procedia Computer Science , 108 , 2282–2286 Sammulal, P., Usha Rani, Y., & Yepuri, A. (2017). A class based clustering approach for imputation and mining of medical records (CBC-IM). IADIS International Journal on Computer Science & Information Systems , 12 (1), 61–74 Mostafa, S. M. (2019). Imputing missing values using cumulative linear regression.CAAI Transactions on Intelligence Technology, 4(3),182–200. https://doi.org/10.1049/trit.2019.0032 Pan, R., Yang, T., Cao, J., Lu, K., & Zhang, Z. (2015). Missing data imputation by K nearest neighbours based on grey relational structure and mutual information.Applied Intelligence, 43(3),614–632. https://doi.org/10.1007/s10489-015-0666-x Silva-Ramírez, E. L., Pino-Mejías, R., & López-Coello, M. (2015). Single imputation with multilayer perceptron and multiple imputation combining multilayer perceptron and k-nearest neighbours for monotone patterns. Applied Soft Computing , 29 , 65–74 Mostafa, M., Eladimy, S. S., Hamad, A., S., & Amano, H. (2020). CBRL and CBRC: Novel algorithms for improving missing value imputation accuracy based on Bayesian ridge regression.Symmetry, 12(10),1594. https://doi.org/10.3390/sym12101594 Sefidian, A. M., & Daneshpour, N. (2020). Estimating missing data using novel correlation maximization based methods. Applied Soft Computing , 91 , 106249 Adhikari, D., Jiang, W., & Zhan, J. (2021). Imputation using information fusion technique for sensor generated incomplete data with high missing gap. Microprocessors and Microsystems . https://doi.org/10.1016/j.micpro.2020.103636103636 Lan, Q., Xu, X., Ma, H., & Li, G. (2020). Multivariable data imputation for the analysis of incomplete credit data. Expert Systems with Applications , 141 , 112926 Kamkhad, N., Jampachaisri, K., Siriyasatien, P., & Kesorn, K. (2020). Toward semantic data imputation for a dengue dataset.Knowledge-Based Systems, 196,105803. https://doi.org/10.1016/j.knosys.2020.105803 Ngueilbaye, A., Wang, H., Mahamat, D. A., & Junaidu, S. B. (2021). Modulo 9 model-based learning for missing data imputation.Applied Soft Computing, 103,107167. https://doi.org/10.1016/j.asoc.2021.107167 Peng, D., Zou, M., Liu, C., & Lu, J. (2021). RESI: a region-splitting imputation method for different types of missing data.Expert Systems with Applications, 168,114425. https://doi.org/10.1016/j.eswa.2020.114425 Austin, P. C., White, I. R., Lee, D. S., & van Buuren, S. (2021). Missing data in clinical research: a tutorial on multiple imputation. Canadian Journal of Cardiology , 37 (9), 1322–1331 Gautam, C., & Ravi, V. (2015). Data imputation via evolutionary computation, clustering and a neural network. Neurocomputing, 156, 134–142. https://doi.org/10.1016/j.neucom.2014.12.073 Priya, R. D., Sivaraj, R., & &Priyaa, N. S. (2017). Heuristically repopulated Bayesian ant colony optimization for treating missing values in large databases. Knowledge-Based Systems , 133 , 107–121 Lobato, F., Sales, C., Araujo, I., Tadaiesky, V., Dias, L., Ramos, L., & Santana, A. (2015). Multi-objective genetic algorithm for missing data imputation.Pattern Recognition Letters, 68,126–131. https://doi.org/10.1016/j.patrec.2015.08.023 Wu, C. H., Wun, C. H., & Chou, H. J. (2004, December). Using association rules for completing missing data. In Fourth International Conference on Hybrid Intelligent Systems (HIS'04) (pp. 236–241). IEEE. https://doi.org/10.1109/ICHIS.2004.91 Wu, J., Song, Q., & Shen, J. (2007, July). An novel association rule mining based missing nominal data imputation method. In Eighth ACIS International Conference on Software Engineering, Artificial Intelligence, Networking, and Parallel/Distributed Computing (SNPD 2007) (Vol. 3, pp. 244–249). IEEE. https://doi.org/10.1109/SNPD.2007.93 Nelwamondo, F. V., Golding, D., & Marwala, T. (2013). A dynamic programming approach to missing data estimation using neural networks.Information Sciences, 237,49–58. https://doi.org/10.1016/j.ins.2009.10.008 Tang, J., Zhang, G., Wang, Y., Wang, H., & Liu, F. (2015). A hybrid approach to integrate fuzzy C-means based imputation method with genetic algorithm for missing traffic volume data estimation. Transportation Research Part C: Emerging Technologies , 51 , 29–40 Aydilek, I. B., & Arslan, A. (2013). A hybrid method for imputation of missing values using optimized fuzzy c-means with support vector regression and a genetic algorithm. Information Sciences , 233 , 25–35 Vazifehdan, M., Moattar, M. H., & Jalali, M. (2019). A hybrid Bayesian network and tensor factorization approach for missing value imputation to improve breast cancer recurrence prediction. Journal of King Saud University-Computer and Information Sciences , 31 (2), 175–184 Choudhary, A., Kumar, S., Sharma, M., & Sharma, K. P. (2022). A Framework for Data Prediction and Forecasting in WSN with Auto ARIMA. Wireless Personal Communications , 123 (3), 2245–2259. https://doi.org/10.1007/s11277-021-09237-x Kwon, O., & Sim, J. M. (2013). Effects of data set features on the performances of classification algorithms. Expert Systems with Applications , 40 (5), 1847–1857 Sim, J., Kwon, O., & Lee, K. C. (2016). Adaptive pairing of classifier and imputation methods based on the characteristics of missing values in data sets. Expert Systems with Applications , 46 , 485–493 Shah, A. D., Bartlett, J. W., Carpenter, J., Nicholas, O., & Hemingway, H. (2014). Comparison of random forest and parametric imputation models for imputing missing data using MICE: a CALIBER study. American journal of epidemiology , 179 (6), 764–774 Van Buuren, S., & Groothuis-Oudshoorn, K. (2011). mice: Multivariate imputation by chained equations in R. Journal of statistical software , 45 , 1–67 Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., & Duchesnay, E. (2011). Scikit-learn: Machine learning in Python.The Journal of machine Learning research, 12,2825–2830 Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1729251","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":113580224,"identity":"f56196bf-70ce-4ffa-a5c2-4b26b0b96a28","order_by":0,"name":"Hutashan Vishal Bhagat","email":"","orcid":"","institution":"Sant Longowal Institute of Engineering and Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hutashan","middleName":"Vishal","lastName":"Bhagat","suffix":""},{"id":113580225,"identity":"04860a7c-1440-4ba2-b672-734c3f005a67","order_by":1,"name":"Manminder Singh","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA40lEQVRIiWNgGAWjYLCCBAZmBgbmAwwMH0Bs4rWwJTAwziBaCwNUCzMPMVp0288e3fBwh7WceRv7s8+2bXZ5/OwNjB8+5uDWYnYmL+1G4pl0Y5ljDMmzc9uSiyV7DjBLztyGR8uBHLMbiW2HE2fINxxmzm1jTtxwI4GNmReflvNvwFrqZ7AxNjNbttUToeUGxJYECTZmZmZGoHVEaAHbkm44g42NmbHn3PHEmT0Hm/H75XyO2c2fbdbyEmzsjxl+lFUn9rM3H/zwEY8WVMDIBiYbiFUPAn9IUTwKRsEoGAUjBQAAWcJTNH7AUA4AAAAASUVORK5CYII=","orcid":"","institution":"Sant Longowal Institute of Engineering and Technology","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Manminder","middleName":"","lastName":"Singh","suffix":""}],"badges":[],"createdAt":"2022-06-06 07:43:24","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1729251/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1729251/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":22738268,"identity":"0cedb84f-e0fe-440f-8113-710a4b87dbb8","added_by":"auto","created_at":"2022-06-16 16:17:33","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":104148,"visible":true,"origin":"","legend":"\u003cp\u003eThe two step IMV-RE process for Imputation\u003c/p\u003e","description":"","filename":"Fig1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-1729251/v1/652cbd25c019e9c9c1b99fe1.jpeg"},{"id":22738267,"identity":"5437134a-a8e7-4140-bc21-1258859fb56d","added_by":"auto","created_at":"2022-06-16 16:17:33","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":33007,"visible":true,"origin":"","legend":"\u003cp\u003eThe performance comparison of the proposed IMV-RE algorithm\u003c/p\u003e","description":"","filename":"Fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-1729251/v1/b9a2361fa10adb37bb363460.png"},{"id":26439163,"identity":"ac5757c0-fd9a-4753-b801-92ede9422345","added_by":"auto","created_at":"2022-09-14 07:20:07","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":522476,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1729251/v1/c0916dae-5765-4433-acab-bff8ae5146ad.pdf"}],"financialInterests":"","formattedTitle":"A Novel Algorithm for Imputing the Missing Values in Incomplete Datasets","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eWe are living in a digital world where information can easily be acquired with the help of smart devices, sensors etc. The advent of IoT makes it possible to collect data without the physical intervention of humans. But sometimes, a failure in such devices may result in data loss and hence, affects the subsequent in-depth analysis and data interpretation that provides erroneous results. The rapid increase in the size of datasets has led to emerging of various data mining techniques. To ensure data mining results to be effective and valuable, data scientists must ensure the quality of the collected data. In real-time scenarios, it is usually the case that collected datasets for data analysis may contain some missing values [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Therefore, it is not possible for most of the data mining algorithms to directly handle these incomplete datasets. The simplest solution to such problems is case deletion which means removing data having missing values. However, case deletion can be appropriate if the missing rate is small, e.g., 5% and if the missing rate is somehow larger, say 25%, then using 75% of the original dataset might be insufficient to completely reflect the real-world problem, which could also affect the mining results [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. To ensure the quality of collected datasets, data scientists first pre-process the datasets [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn concern to the relationship between the missing data and other data values of the variables in a dataset, missing data mechanism can be assorted as Missing at Random (MAR), Missing Completely at Random (MCAR) and Missing Not at Random (MNAR) [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. If the probability of missing data depends upon the observed responses and there exist no relation among the missing values itself, the missing data mechanism is referred to as MAR. Hence, in MAR the reason for missingness in a feature \u003cem\u003eQ\u003c/em\u003e of a dataset mainly depends upon the rest of the features within the dataset rather than the \u003cem\u003eQ\u003c/em\u003e itself. If there is no relation between the missing values and the set of observed responses, the missingness mechanism is said to be MCAR. In MCAR, the reason for missing in a feature \u003cem\u003eQ\u003c/em\u003e depends neither on the other features within the dataset nor on \u003cem\u003eQ\u003c/em\u003e itself. If the probability of the missingness in a feature \u003cem\u003eQ\u003c/em\u003e depends either on \u003cem\u003eQ\u003c/em\u003e itself or on the other features that also contains missing values, the missing mechanism is referred to as MNAR.\u003c/p\u003e \u003cp\u003eMissing value imputation (MVI) techniques provide the best solutions to impute the missing values within the datasets. MVI techniques can be categorized into two categories- MVI techniques based on Statistical Methods (Mean, Mode, Median and Regression) and MVI techniques based on Machine Learning (Neural Networks, Support Vector Machines and Deep Learning Models) [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Although, machine learning techniques being complex in nature produce better imputation results, yet they are computationally expensive. For high dimensional datasets, such techniques show high computation time as compared to statistical methods because of which they are not deployed to mission-critical systems [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eBeing inspired from the lower complexity and computation time of statistical MVI techniques, this paper proposed a novel splitting-based imputation approach, namely, IMV-RE. To justify the superiority of the proposed IMV-RE approach, ten benchmark datasets containing both numerical as well as mixed values with varying dimensionality take into consideration and is compared with five existing imputation techniques. Classification accuracy, RMSE and Coefficient of determination (R\u003csup\u003e2\u003c/sup\u003e) are considered as the performance metrics with which the proposed IMV-RE approach is evaluated.\u003c/p\u003e "},{"header":"2. Literature Review","content":"\u003cp\u003eMissingness within a dataset is a very common problem in statistical analysis. Researchers have come out with numerous MVI techniques depending upon the missing rate, data characteristics and pattern and data correlation. This section gives a brief description of the recent MVI techniques.\u003c/p\u003e \u003cp\u003eRegularized Expectation Maximization (EM) algorithm for imputing missing value which is based on the iterated analysis of the linear regressions among the missing and non-missing variables is proposed in [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. The experimental test is carried out using regularized EM technique over the climate data. The regularized EM technique is further enhanced in [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] and two novel methods kEMI and kEMI\u003csup\u003e+\u003c/sup\u003e are proposed by the authors. The proposed techniques are based on the information fusion mechanism that uses Dempster-Shafer fusion to fuse the most appropriate estimates. Due to the high computation time of the missForest [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], the authors in [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] proposed a new technique known as mForest that can achieve ten times less computation time than missForest. A Column-wise Guided Data Imputation (cGDI) technique that divides the complete samples from the incomplete samples and selects the most suitable imputation method separately for each feature is proposed in [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e], the authors proposed a CBC-IM (Class Based Clustering approach for Imputation) technique for missing values imputation especially for medical datasets. CBC-IM technique first partitioned the dataset into complete and incomplete variables and then uses the Euclidean distance and Fuzzy measure to find out the similarity between the two records. Another similar approach known as CCMVI (Class Center based Missing Value Imputation) is proposed in [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. CCMVI is a two-step process that first defines the threshold value for every class containing missing values and then finds the appropriate estimate from the complete samples considering the threshold values. An approach based on the linear regression technique known as CLR (Cumulative Linear Regression) is proposed in [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. The incomplete variables are first cumulated and incorporated in the linear regression equation to replace the missing values in the next incomplete variable.\u003c/p\u003e \u003cp\u003eThe KNNI (K Nearest Neighbor Imputation) is the most extensively used MVI technique. In [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e], the authors proposed a feature weighted grey KNNI technique that uses the combination of relevant feature information and grey rational based k nearest neighbors to impute the missing values. Another approach that makes combine use of Multi-Layer Perceptron and KNN to impute the multiple missing values simultaneously is proposed in [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. This combination results in an increase in the performance with an increase in the time complexity of the algorithm. In [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], the authors proposed two techniques CBRL (Cumulative Bayesian Ridge with Less NaN) and CBRC (Cumulative Bayesian Ridge with high Correlation). CBRL has used the most appropriate features within the dataset that contain a lesser number of missing values whereas CBRC is used to select the best feature that gives a high correlation with the target feature.\u003c/p\u003e \u003cp\u003eCorrelation Maximization-based Imputation Methods (CMIM) are proposed in [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] that first find the highly correlated segments of the data and use linear regression estimator to impute the missing values. The authors use a two-step method to handle large missing gaps in [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e] known as Ratio-Based Imputation (RBI). In RBI the MVI is done by using machine learning models whereas the analysis is done by data fusion technique in CPS (Cyber-Physical Systems) datasets. Considering the missing values in a dataset an optimization problem, the authors in [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] proposed an optimal method known as BNII. The BNII technique is a two-stage approach: firstly, using the Bayesian Network relationship among different attributes is calculated and secondly, in an iterative manner imputation is done till local maximum posterior probability is reached. A novel architecture based on Particle Swarm Optimization (PSO) for cleaning the data and then utilizing the K-means to calculate the fitness value as well as to narrow down the search space is proposed in [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. The ontology makes PSO replace missing values more accurately but in worst scenarios, this approach has a time complexity of \u003cem\u003eΘ(n\u003c/em\u003e\u003csup\u003e\u003cem\u003e3\u003c/em\u003e\u003c/sup\u003e\u003cem\u003e).\u003c/em\u003e\u003c/p\u003e \u003cp\u003eA novel method known as Modulo 9 proposed in [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e] impute the missing values within the interval of [0\u0026ndash;9] and then use congruency with addition and multiplication to make an appropriate estimation. The authors compared the Modulo 9 approach with the eleven robust MVI techniques and outperform all of them. Considering generalization for datasets that contain mixed-type data, the authors in [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] proposed a tuple-oriented region splitting imputation technique known as RESI (Region-Splitting Imputation). RESI technique first uses the entropy weight method to assign weights to the attributes and split the data into complete and incomplete subsets based on their integrity rate. The model is trained over a complete subset that iteratively imputes the next incomplete subset.\u003c/p\u003e \u003cp\u003eFrom the last few years, researchers have come out with numerous imputation techniques based on meta-heuristic techniques ([\u003cspan additionalcitationids=\"CR25 CR26\" citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e],association mining rules[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e], dynamic programming techniques[\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e] and various such hybrid techniques [\u003cspan additionalcitationids=\"CR32 CR33\" citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e] to efficiently impute the missing values. The authors in [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e] analyzed that there is no such MVI technique that could be considered as a master technique for distinct problems. The key factors that can influence the performance of MVI techniques are the data distribution within the dataset, the missingness rate and characteristics of a dataset [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe main contribution of this paper is to overcome the limitations of state-of-the-art imputation techniques based on the statistical methods. The authors proposed a CBC-IM [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] (Class Based Clustering approach for Imputation) technique for missing values imputation especially for medical datasets. CBC-IM technique first partitioned the dataset into complete and incomplete variables and then uses the Euclidean distance and Fuzzy measure to find out the similarity between the two records. Another similar approach known as CCMVI is proposed [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. CCMVI is a two-step process that first defines the threshold value for every class containing missing values and then finds the appropriate estimate from the complete samples considering the threshold values. Although, the CCMVI and IM-CBC techniques have less time complexity, yet their dependency on complete samples make them unsuitable for the imputation especially when a dataset has high missingness rate or when there is an unavailability of at least one data sample without missing values within a class. The major goal of any MVI technique is to replace the missing values in such a manner that the overall data integrity, structure, and trends of the data are maintained. The proposed IMV-RE MVI technique shows better performance than the commonly used existing single as well as multiple imputation techniques. Unlike the CCMVI technique, the IMV-RE technique successfully imputes the missing values even the missing rate is high. Unlike the missForest [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], MICE (Multiple Imputation using Chained Equations) and other meta-heuristic techniques [\u003cspan additionalcitationids=\"CR25 CR26\" citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e], the proposed IMV-RE has low time complexity. Because of the simplicity of the proposed IMV-RE algorithm, it is significant for high dimensional datasets also.\u003c/p\u003e"},{"header":"3. Imputing The Missing Values In Real-time Environment","content":"\u003cp\u003eThe proposed IMV-RE is a splitting-based data imputation algorithm that is the robust statistical MVI technique. The proposed algorithm first decides the upper limit for each feature within a cluster which is further used for imputing the missing values so that the imputed values lie much nearer to the exact values. The following sub-section discussed the step-by-step process of the proposed imputation algorithm.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Imputation Algorithm (IMV-RE)\u003c/h2\u003e \u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the whole process from start to the end of IMV-RE algorithm and each step is described as below:\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStep 1\u003c/strong\u003e \u003cp\u003eInitially a dataset \u003cem\u003eX\u003c/em\u003e\u003csub\u003e\u003cem\u003eD\u003c/em\u003e\u003c/sub\u003e is given that has \u003cem\u003eD\u003c/em\u003e dimensions and \u003cem\u003eC\u003c/em\u003e\u003csub\u003e\u003cem\u003eN\u003c/em\u003e\u003c/sub\u003e number of classes.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStep 2\u003c/strong\u003e \u003cp\u003eCluster the data samples into their respective classes along with the missing data samples using the labels given in case of labeled datasets and for unlabeled datasets, K-Means clustering is used to cluster the similar data samples such that \u003cem\u003eC\u003c/em\u003e\u003csub\u003e\u003cem\u003e1\u003c/em\u003e\u003c/sub\u003e, \u003cem\u003eC\u003c/em\u003e\u003csub\u003e\u003cem\u003e2\u003c/em\u003e\u003c/sub\u003e, \u003cem\u003eC\u003c/em\u003e\u003csub\u003e\u003cem\u003e3\u003c/em\u003e\u003c/sub\u003e,\u003cem\u003e\u0026hellip;\u0026hellip; C\u003c/em\u003e\u003csub\u003e\u003cem\u003eN\u003c/em\u003e\u003c/sub\u003e numbers of classes are there.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStep 3\u003c/strong\u003e \u003cp\u003eFor each class \u003cem\u003eC\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e (where \u003cem\u003ei\u0026thinsp;=\u0026thinsp;1 to N\u003c/em\u003e), replace all the missing values with zero. Calculate the centered value (\u003cem\u003eCent\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e) and standard deviation (\u003cem\u003eSDev\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e) for class \u003cem\u003eC\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStep 4\u003c/strong\u003e \u003cp\u003eCalculate the L\u003csub\u003e2\u003c/sub\u003e norm between the \u003cem\u003eCent\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e and each data sample in class \u003cem\u003eC\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e. Select the mid-point value as an upper limit (\u003cem\u003eUpLimC\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e) for class \u003cem\u003eC\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStep 5\u003c/strong\u003e \u003cp\u003eFor each class \u003cem\u003eC\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e obtained from the \u003cb\u003eStep 2\u003c/b\u003e, separate all the data samples containing missing values from the complete data samples into a cluster such that cluster \u003cem\u003eC\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e contains \u003cem\u003eL\u003c/em\u003e\u003csub\u003e\u003cem\u003emiss_sam\u003c/em\u003e\u003c/sub\u003e numbers of missing samples.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStep 6\u003c/strong\u003e \u003cp\u003eFor each missing sample \u003cem\u003eL\u003c/em\u003e\u003csub\u003e\u003cem\u003emiss_sam\u003c/em\u003e\u003c/sub\u003e containing \u003cem\u003eN_miss\u003c/em\u003e number of missing values, impute each missing value with \u003cem\u003eCent\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e calculated for the class \u003cem\u003eC\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003ein \u003cb\u003eStep 3\u003c/b\u003e. Now, calculate the L\u003csub\u003e2\u003c/sub\u003e norm between the \u003cem\u003eCent\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e and the imputed \u003cem\u003eL\u003c/em\u003e\u003csub\u003e\u003cem\u003emiss_sam\u003c/em\u003e\u003c/sub\u003e sample. If the value is smaller than the \u003cem\u003eUpLimC\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e (calculated in \u003cb\u003eStep 4\u003c/b\u003e), then imputed value is finalized else +/- standard deviation (\u003cem\u003eSDev\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e) is applied to new imputed value in a sequential order and corresponding the L\u003csub\u003e2\u003c/sub\u003e norm between the \u003cem\u003eCent\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e and the imputed value is calculated.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eStep 7\u003c/strong\u003e \u003cp\u003eThe final value to be imputed is the value in correspond to smaller value calculated for the L\u003csub\u003e2\u003c/sub\u003e norm between the \u003cem\u003eCent\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e and the new imputed \u003cem\u003eL\u003c/em\u003e\u003csub\u003e\u003cem\u003emiss_sam\u003c/em\u003e\u003c/sub\u003e sample.\u003c/p\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"4. Experimental Implementation","content":"\u003cp\u003eThis section is divided into three subsections. The first subsection represents the benchmark datasets, the second subsection represents the experimental setup and the third subsection represents the metrics used to evaluate the proposed IMV-RE algorithm.\u003c/p\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Datasets\u003c/h2\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e depicts the brief description of the datasets used for the experimentation. A total of ten datasets, that includes 8 numerical and 2 of mixed data type, from the UCI machine learning repository are collected with the number of samples ranging from 70 to 5000, the number of features ranging from 4 to 206 and the number of classes ranging from 2 to 10.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Experimental Setup\u003c/h2\u003e \u003cp\u003eThe MCAR (Missing Completely at Random) missing value mechanism is used to intentionally add missing values in the above datasets with missingness rates of 10%, 20%, 30%, 40% and 50% of the total data within datasets. A fixed random seed is used for generating the missing values in every dataset to avoid biased results.\u003c/p\u003e \u003cp\u003eThe proposed IMV-RE approach is used to impute the missing values and is compared with five baseline approaches: Class Center based Missing Value Imputation (CCMVI), K-Nearest Neighbor Imputation (KNNI), MiceForest [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e], IterativeImputer [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]and SimpleImputer [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. The experimental work is performed on python IDE Spyder version 5.2.1 in Anaconda Navigator usinglaptop PC DELL G5, Intel Core i7 processor, RAM 12GB, 512GB SSD.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDatasets Description\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDatasets\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDatasets type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo. of instances\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNo. of features\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNo. of classes\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWaveform\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumerical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGlass\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumerical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e214\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWheat-Seed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumerical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e210\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDigits\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumerical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1797\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWine\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumerical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e178\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIris\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumerical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e150\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSeeds\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumerical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e210\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIonosphere\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumerical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e351\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e34\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSCADI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMixed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e206\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEcoli\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMixed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e336\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e4.3 Evaluation Metrics\u003c/h2\u003e \u003cp\u003eA stratified 5-fold cross-validation using RandomForest Classifier is performed to evaluate the classification accuracy on the imputed datasets. For every dataset, there are five different missing rates that correspond to five different classification accuracies and average of five accuracies are taken into consideration for comparison among different imputation techniques.\u003c/p\u003e \u003cp\u003eIn addition to the classification accuracy, two more performance metrics are used to evaluate the performance of the proposed IMV-RE approach.\u003c/p\u003e \u003cp\u003eRoot Mean Square Error (RMSE): If \u003cem\u003ex\u003c/em\u003e\u003csub\u003e\u003cem\u003ei\u003c/em\u003e\u003c/sub\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\widehat{x}\\)\u003c/span\u003e\u003c/span\u003e are the original value and the imputed value of the \u003cem\u003ei\u003c/em\u003e\u003csup\u003e\u003cem\u003eth\u003c/em\u003e\u003c/sup\u003e observation respectively, \u003cem\u003en\u003c/em\u003e is total number of samples, then, RMSE is given by the equation:\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$RMSE=\\sqrt{\\frac{{\\sum }_{i=1}^{n}{\\left({x}_{i}-\\widehat{{x}_{i}}\\right)}^{2}}{n}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eCoefficient of Determination (R\u003csup\u003e2\u003c/sup\u003e): Calculated using the following equation:\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$${ R}^{2}\\left(x, \\widehat{x}\\right)=1-\\frac{{\\sum }_{i=1}^{n}{\\left({x}_{i}-\\widehat{{x}_{i}}\\right)}^{2}}{{\\sum }_{i=1}^{n}{\\left({x}_{i}-\\stackrel{-}{x}\\right)}^{2}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere,\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\stackrel{-}{x}= \\frac{1}{n}\\sum _{i=1}^{n}{x}_{i}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eMoreover, the mean and standard deviation of the imputed datasets are compared with the mean and standard deviation of the actual datasets and percentage error is calculated for all datasets.\u003c/p\u003e \u003c/div\u003e"},{"header":"5. Results And Discussions","content":"\u003cp\u003eThree metrics are used to evaluate the performance of the proposed IMV-RE algorithm. The IMV-RE algorithm is applied to ten benchmark datasets and the classification results obtained using RandomForest Classifier, RMSE values and Coefficient of determination (R\u003csup\u003e2\u003c/sup\u003e) values are compared to CCMVI, KNNI, MiceForest, IterativeImputer, SimpleImputer MVI techniques.\u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e depicts the average classification accuracies obtained on the ten benchmark datasets Waveform, Glass, Wheat-Seed, Digits, Wine, Iris, Seeds, Ionosphere, SCADI and Ecoli having 10%, 20%, 30%, 40% and 50% missing rates. The average accuracy of the proposed IMV-RE algorithm achieved for all the ten datasets is 91.58% which is better than the other five MVI approaches MiceForest (81.91%), KNNI (80.95%), SimpleImputer (81.7%) and IterativeImputer (82.81%)\u003c/p\u003e \u003cp\u003eFor Waveform, Glass, Wheat-Seed, Digits and Wine datasets the CCMVI technique is not applicable beyond 10% missing rate because of its dependency on the samples without missing values to calculate the threshold value that is utilized to estimate the missing values in a particular class.\u003c/p\u003e \u003cp\u003eSimilarly, Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e compares the average RMSE values obtained for ten datasets. The results show that the proposed IMV-RE algorithm achieves the lowest RMSE value 0.73 in comparison to the other five MVI techniques. RMSE for CCMVI technique cannot be calculated because of the same above-mentioned reason.\u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e compares the average value of coefficient of determination obtained for all the ten datasets. The proposed IMV-RE algorithm achieves average Coefficient of Determination value 0.886 which is higher than all other MVI techniques.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAverage classification accuracies obtained from distinct MVI techniques\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDatasets\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eIMV-RE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMiceForest\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eKNNI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCCMVI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSimpleImputer\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eIterativeImputer\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWaveform\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.96443\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.78015\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.76447\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.77695\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.79092\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGlass\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.80498\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.54197\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.52140\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.55801\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.57763\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWheat-Seed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.95944\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.78583\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.78957\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.81168\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.82681\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDigits\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.98664\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.77682\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.73919\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.65356\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.71446\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWine\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.95883\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.95197\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.93429\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.96990\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.96860\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.96537\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIris\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.97333\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.92267\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.90667\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.97267\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.92933\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.94800\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSeeds\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.89143\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.87143\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.88095\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.90286\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.87714\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.87810\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIonosphere\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.91851\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.91059\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.90365\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.92029\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.91285\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.90766\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSCADI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.85275\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.84440\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.83846\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.84725\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.85626\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.84703\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEcoli\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.84753\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.80530\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.81605\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.85595\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.82558\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.82551\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAverage\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.91579\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.81911\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.80947\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.81700\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e0.82815\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAverage RMSE obtained from distinct MVI techniques\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDatasets\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eIMV-RE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMiceForest\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eKNNI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCCMVI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSimpleImputer\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eIterativeImputer\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWaveform\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.68948\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.70871\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.73801\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.80543\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.71955\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGlass\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.32751\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.37197\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.38146\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.38244\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.39661\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWheat-Seed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.32499\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.46663\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.48078\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.63877\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.42437\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDigits\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2.02541\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.07970\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2.25400\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.60753\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e2.53113\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWine\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e3.54842\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.65323\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e4.98252\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3.79756\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e7.21529\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e6.15099\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIris\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.15431\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.21552\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.21250\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.17177\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.40314\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.21158\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSeeds\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.14982\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.11402\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.15104\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.15284\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.30043\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.13956\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIonosphere\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.07780\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.05712\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.04592\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.07728\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.06871\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.06354\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSCADI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.00102\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.00082\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.00041\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.00100\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.00182\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.00159\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEcoli\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.02252\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.03085\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.02623\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.02361\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.02984\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.02654\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAverage\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.73213\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.86986\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.92729\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e1.24534\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e1.06655\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAverage Coefficient of Determination (R\u003csup\u003e2\u003c/sup\u003e) obtained from distinct MVI techniques\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDatasets\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eIMV-RE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMiceForest\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eKNNI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCCMVI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSimpleImputer\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eIterativeImputer\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWaveform\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.78181\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.75131\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.73399\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.71399\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.76373\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGlass\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.71782\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.63663\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.60446\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.65406\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.65720\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWheat-Seed\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.85907\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.68851\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.71047\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.57340\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.77365\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDigits\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.67990\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.67332\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.56739\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.53716\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.50253\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWine\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.96225\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.96254\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.93562\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.95724\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.93715\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.95333\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIris\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.95632\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.93200\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.92954\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.94407\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.82545\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.93378\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSeeds\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.97092\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.97381\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.96951\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.96438\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.90623\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.95981\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIonosphere\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.97309\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.98352\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.98874\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.97279\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.97986\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.98238\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSCADI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.99958\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.99940\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.99981\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.99929\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.99934\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.99948\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEcoli\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.96658\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.94247\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.94557\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.96319\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.95191\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.95639\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAverage\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.88673\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.85435\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.83851\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.80785\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e0.84823\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e graphically represents the performance comparison of proposed IMV-RE algorithm over different evaluation metrics.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePercentage Error between mean values of actual datasets and imputed datasets\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"11\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c11\" colnum=\"11\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDatasets\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eEcoli\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGlass\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eWheat-Seed\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDigit\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eWine\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eIris\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eSeed\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eIonosphere\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c10\"\u003e \u003cp\u003eSCADI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c11\"\u003e \u003cp\u003eWaveform\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eActual Mean\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.4996\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.0968\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e6.8967\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e4.8843\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.0437\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e3.4645\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.0820\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.2477\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.2033\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e1.7123\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eImputed Mean\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.4988\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.0790\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e6.9166\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e4.8887\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.0371\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e3.4657\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.0749\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.2438\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.2035\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e1.7153\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ePercentage Error\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e0.1520\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e18.3642\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.2890\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e0.0890\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e15.0280\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e0.0332\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e8.5944\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e\u003cb\u003e1.5801\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e\u003cb\u003e0.0914\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e\u003cb\u003e0.1745\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePercentage Error between calculated standard deviation of actual datasets and imputed datasets\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"11\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c11\" colnum=\"11\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDatasets\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eEcoli\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGlass\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eWheat-Seed\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDigit\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eWine\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eIris\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eSeed\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eIonosphere\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c10\"\u003e \u003cp\u003eSCADI\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c11\"\u003e \u003cp\u003eWaveform\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eActual Stdev\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.144\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.048\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1.010\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e3.684\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.702\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.948\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.632\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.510\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.263\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e1.520\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eImputed Stdev\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.156\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.961\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e5.311\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5.549\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.704\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.968\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.634\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e0.573\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e0.951\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e1.758\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ePercentage Error\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e8.470\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e8.343\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e425.986\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e50.616\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e0.366\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e107.660\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e0.375\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e\u003cb\u003e12.176\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c10\"\u003e \u003cp\u003e\u003cb\u003e261.444\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c11\"\u003e \u003cp\u003e\u003cb\u003e15.672\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e represents the percentage error calculated between the mean values of actual datasets and imputed datasets by considering the average of total mean value calculated at different missing rates 10%, 20%, 30%, 40% and 50% using proposed IMV-RE algorithm.\u003c/p\u003e \u003cp\u003eSimilarly, the Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e represents the percentage error calculated between the standard deviation of actual datasets and imputed datasets by considering the average of total standard deviation value calculated at different missing rates 10%, 20%, 30%, 40% and 50% using proposed IMV-RE algorithm.\u003c/p\u003e"},{"header":"6. Conclusion","content":"\u003cp\u003eIn this paper, a novel splitting-based IMV-RE algorithm is proposed. The proposed IMV-RE algorithm is a two-step process. In the first step, an upper limit is calculated for every class containing missing values and the second step utilizes the upper limit to impute missing values efficiently. The IMV-RE algorithm only searches within the class to calculate the centered value for the final imputation, unlike the other techniques that need to go through the whole data within the dataset. Hence, the proposed algorithm can be utilized for real-time problems. The proposed IMV-RE algorithm prosperously abstracts the dependency of the CCMVI technique on the complete samples within a class to estimate the missing values. As a result, the proposed IMV-RE algorithm successfully imputes the missing values for higher data missing.\u003c/p\u003e \u003cp\u003eClassification accuracy, RMSE and coefficient of determination (R\u003csup\u003e2\u003c/sup\u003e) are the evaluation metrics used to evaluate the performance of the proposed IMV-RE algorithm. The experimental results depict better performance of the proposed IMV-RE algorithm for imputing the missing values under distinct missing rates in comparison to the other state-of-the-art techniques.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eConflict of Interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors of this publication declare there is no conflict of interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode/Data Availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eN/A.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors' Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eHutashan Vishal Bhagat: Conceptualization, Methodology, Software, Formal analysis, Writing - Original Draft, Validation, Data Curation and Visualization.\u003c/p\u003e\n\u003cp\u003eManminder Singh: Writing Review, Editing and Supervision\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eKang, H. (2013). The prevention and handling of the missing data.Korean journal of anesthesiology, 64(5),402. https://doi.org/10.4097/kjae.2013.64.5.402\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKalkan, \u0026Ouml;. K., Yusuf, K. A. R. A., \u0026amp; Kelecioğlu, H. (2018). Evaluating performance of missing data imputation methods in IRT analyses.International Journal of Assessment Tools in Education, 5(3),403\u0026ndash;416. https://doi.org/10.21449/ijate.430720\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGarc\u0026iacute;a, S., Luengo, J., \u0026amp; Herrera, F. (2015). Data preprocessing in data mining (Vol. 72, pp.59\u0026ndash;139). Cham, Switzerland:Springer International Publishing\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKelkar, B. A. (2022). Missing Data Imputation: A Survey. \u003cem\u003eInternational Journal of Decision Support System Technology (IJDSST)\u003c/em\u003e, \u003cem\u003e14\u003c/em\u003e(1), 1\u0026ndash;20. DOI: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.4018/IJDSST.292446\u003c/span\u003e\u003cspan address=\"10.4018/IJDSST.292446\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLittle, R. J., \u0026amp; Rubin, D. B. (2019). Statistical analysis with missing data (Vol. 793).John Wiley \u0026amp; Sons\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBaraldi, A. N., \u0026amp; Enders, C. K. (2010). An introduction to modern missing data analyses. \u003cem\u003eJournal of school psychology\u003c/em\u003e, \u003cem\u003e48\u003c/em\u003e(1), 5\u0026ndash;37\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTsai, C. F., Li, M. L., \u0026amp; Lin, W. C. (2018). A class center based approach for missing value imputation. \u003cem\u003eKnowledge-Based Systems\u003c/em\u003e, \u003cem\u003e151\u003c/em\u003e, 124\u0026ndash;135\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchneider, T. (2001). Analysis of incomplete climate data: Estimation of mean values and covariance matrices and imputation of missing values. \u003cem\u003eJournal of climate\u003c/em\u003e, \u003cem\u003e14\u003c/em\u003e(5), 853\u0026ndash;871\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRazavi-Far, R., Cheng, B., Saif, M., \u0026amp; Ahmadi, M. (2020). Similarity-learning information-fusion schemes for missing data imputation. \u003cem\u003eKnowledge-Based Systems\u003c/em\u003e, \u003cem\u003e187\u003c/em\u003e, 104805\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eProbst, P., Wright, M. N., \u0026amp; Boulesteix, A. L. (2019). Hyperparameters and tuning strategies for random forest. \u003cem\u003eWiley Interdisciplinary Reviews: data mining and knowledge discovery\u003c/em\u003e, \u003cem\u003e9\u003c/em\u003e(3), e1301\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTang, F., \u0026amp; Ishwaran, H. (2017). Random forest missing data algorithms. \u003cem\u003eStatistical Analysis and Data Mining: The ASA Data Science Journal\u003c/em\u003e, \u003cem\u003e10\u003c/em\u003e(6), 363\u0026ndash;377\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePetrozziello, A., \u0026amp; Jordanov, I. (2017). Column-wise guided data imputation. \u003cem\u003eProcedia Computer Science\u003c/em\u003e, \u003cem\u003e108\u003c/em\u003e, 2282\u0026ndash;2286\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSammulal, P., Usha Rani, Y., \u0026amp; Yepuri, A. (2017). A class based clustering approach for imputation and mining of medical records (CBC-IM). \u003cem\u003eIADIS International Journal on Computer Science \u0026amp; Information Systems\u003c/em\u003e, \u003cem\u003e12\u003c/em\u003e(1), 61\u0026ndash;74\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMostafa, S. M. (2019). Imputing missing values using cumulative linear regression.CAAI Transactions on Intelligence Technology, 4(3),182\u0026ndash;200. https://doi.org/10.1049/trit.2019.0032\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePan, R., Yang, T., Cao, J., Lu, K., \u0026amp; Zhang, Z. (2015). Missing data imputation by K nearest neighbours based on grey relational structure and mutual information.Applied Intelligence, 43(3),614\u0026ndash;632. https://doi.org/10.1007/s10489-015-0666-x\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSilva-Ram\u0026iacute;rez, E. L., Pino-Mej\u0026iacute;as, R., \u0026amp; L\u0026oacute;pez-Coello, M. (2015). Single imputation with multilayer perceptron and multiple imputation combining multilayer perceptron and k-nearest neighbours for monotone patterns. \u003cem\u003eApplied Soft Computing\u003c/em\u003e, \u003cem\u003e29\u003c/em\u003e, 65\u0026ndash;74\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMostafa, M., Eladimy, S. S., Hamad, A., S., \u0026amp; Amano, H. (2020). CBRL and CBRC: Novel algorithms for improving missing value imputation accuracy based on Bayesian ridge regression.Symmetry, 12(10),1594. https://doi.org/10.3390/sym12101594\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSefidian, A. M., \u0026amp; Daneshpour, N. (2020). Estimating missing data using novel correlation maximization based methods. \u003cem\u003eApplied Soft Computing\u003c/em\u003e, \u003cem\u003e91\u003c/em\u003e, 106249\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAdhikari, D., Jiang, W., \u0026amp; Zhan, J. (2021). Imputation using information fusion technique for sensor generated incomplete data with high missing gap. \u003cem\u003eMicroprocessors and Microsystems\u003c/em\u003e. \u0026lt;background-color:#cfbfb1;uvertical-align:super;\u0026gt;https://doi.org/10.1016/j.micpro.2020.103636\u0026lt;/background-color:#cfbfb1;uvertical-align:super;\u0026gt;103636\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLan, Q., Xu, X., Ma, H., \u0026amp; Li, G. (2020). Multivariable data imputation for the analysis of incomplete credit data. \u003cem\u003eExpert Systems with Applications\u003c/em\u003e, \u003cem\u003e141\u003c/em\u003e, 112926\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKamkhad, N., Jampachaisri, K., Siriyasatien, P., \u0026amp; Kesorn, K. (2020). Toward semantic data imputation for a dengue dataset.Knowledge-Based Systems, 196,105803. https://doi.org/10.1016/j.knosys.2020.105803\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNgueilbaye, A., Wang, H., Mahamat, D. A., \u0026amp; Junaidu, S. B. (2021). Modulo 9 model-based learning for missing data imputation.Applied Soft Computing, 103,107167. https://doi.org/10.1016/j.asoc.2021.107167\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeng, D., Zou, M., Liu, C., \u0026amp; Lu, J. (2021). RESI: a region-splitting imputation method for different types of missing data.Expert Systems with Applications, 168,114425. https://doi.org/10.1016/j.eswa.2020.114425\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAustin, P. C., White, I. R., Lee, D. S., \u0026amp; van Buuren, S. (2021). Missing data in clinical research: a tutorial on multiple imputation. \u003cem\u003eCanadian Journal of Cardiology\u003c/em\u003e, \u003cem\u003e37\u003c/em\u003e(9), 1322\u0026ndash;1331\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGautam, C., \u0026amp; Ravi, V. (2015). Data imputation via evolutionary computation, clustering and a neural network. Neurocomputing, 156, 134\u0026ndash;142. https://doi.org/10.1016/j.neucom.2014.12.073\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePriya, R. D., Sivaraj, R., \u0026amp; \u0026amp;Priyaa, N. S. (2017). Heuristically repopulated Bayesian ant colony optimization for treating missing values in large databases. \u003cem\u003eKnowledge-Based Systems\u003c/em\u003e, \u003cem\u003e133\u003c/em\u003e, 107\u0026ndash;121\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLobato, F., Sales, C., Araujo, I., Tadaiesky, V., Dias, L., Ramos, L., \u0026amp; Santana, A. (2015). Multi-objective genetic algorithm for missing data imputation.Pattern Recognition Letters, 68,126\u0026ndash;131. https://doi.org/10.1016/j.patrec.2015.08.023\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu, C. H., Wun, C. H., \u0026amp; Chou, H. J. (2004, December). Using association rules for completing missing data. In \u003cem\u003eFourth International Conference on Hybrid Intelligent Systems (HIS'04)\u003c/em\u003e (pp.\u0026nbsp;236\u0026ndash;241). IEEE. https://doi.org/10.1109/ICHIS.2004.91\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu, J., Song, Q., \u0026amp; Shen, J. (2007, July). An novel association rule mining based missing nominal data imputation method. In \u003cem\u003eEighth ACIS International Conference on Software Engineering, Artificial Intelligence, Networking, and Parallel/Distributed Computing (SNPD 2007)\u003c/em\u003e (Vol.\u0026nbsp;3, pp.\u0026nbsp;244\u0026ndash;249). IEEE. https://doi.org/10.1109/SNPD.2007.93\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNelwamondo, F. V., Golding, D., \u0026amp; Marwala, T. (2013). A dynamic programming approach to missing data estimation using neural networks.Information Sciences, 237,49\u0026ndash;58. https://doi.org/10.1016/j.ins.2009.10.008\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTang, J., Zhang, G., Wang, Y., Wang, H., \u0026amp; Liu, F. (2015). A hybrid approach to integrate fuzzy C-means based imputation method with genetic algorithm for missing traffic volume data estimation. \u003cem\u003eTransportation Research Part C: Emerging Technologies\u003c/em\u003e, \u003cem\u003e51\u003c/em\u003e, 29\u0026ndash;40\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAydilek, I. B., \u0026amp; Arslan, A. (2013). A hybrid method for imputation of missing values using optimized fuzzy c-means with support vector regression and a genetic algorithm. \u003cem\u003eInformation Sciences\u003c/em\u003e, \u003cem\u003e233\u003c/em\u003e, 25\u0026ndash;35\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVazifehdan, M., Moattar, M. H., \u0026amp; Jalali, M. (2019). A hybrid Bayesian network and tensor factorization approach for missing value imputation to improve breast cancer recurrence prediction. \u003cem\u003eJournal of King Saud University-Computer and Information Sciences\u003c/em\u003e, \u003cem\u003e31\u003c/em\u003e(2), 175\u0026ndash;184\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChoudhary, A., Kumar, S., Sharma, M., \u0026amp; Sharma, K. P. (2022). A Framework for Data Prediction and Forecasting in WSN with Auto ARIMA. \u003cem\u003eWireless Personal Communications\u003c/em\u003e, \u003cem\u003e123\u003c/em\u003e(3), 2245\u0026ndash;2259. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s11277-021-09237-x\u003c/span\u003e\u003cspan address=\"10.1007/s11277-021-09237-x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKwon, O., \u0026amp; Sim, J. M. (2013). Effects of data set features on the performances of classification algorithms. \u003cem\u003eExpert Systems with Applications\u003c/em\u003e, \u003cem\u003e40\u003c/em\u003e(5), 1847\u0026ndash;1857\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSim, J., Kwon, O., \u0026amp; Lee, K. C. (2016). Adaptive pairing of classifier and imputation methods based on the characteristics of missing values in data sets. \u003cem\u003eExpert Systems with Applications\u003c/em\u003e, \u003cem\u003e46\u003c/em\u003e, 485\u0026ndash;493\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShah, A. D., Bartlett, J. W., Carpenter, J., Nicholas, O., \u0026amp; Hemingway, H. (2014). Comparison of random forest and parametric imputation models for imputing missing data using MICE: a CALIBER study. \u003cem\u003eAmerican journal of epidemiology\u003c/em\u003e, \u003cem\u003e179\u003c/em\u003e(6), 764\u0026ndash;774\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVan Buuren, S., \u0026amp; Groothuis-Oudshoorn, K. (2011). mice: Multivariate imputation by chained equations in R. \u003cem\u003eJournal of statistical software\u003c/em\u003e, \u003cem\u003e45\u003c/em\u003e, 1\u0026ndash;67\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., \u0026amp; Duchesnay, E. (2011). Scikit-learn: Machine learning in Python.The Journal of machine Learning research, 12,2825\u0026ndash;2830\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Imputation, Imputing values, Missingness Mechanisms, Missing Values, Data Missingness, Data Imputation Model, Incomplete datasets, Root Mean Square Error","lastPublishedDoi":"10.21203/rs.3.rs-1729251/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1729251/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eIn today\u0026rsquo;s world, we completely rely on digital devices to collect data; a failure in such digital devices may result in huge information loss thereby making data mining a more tedious job for a Data Analyst. Missingness to a greater extent in a dataset subsequently comes out with inappropriate results and incomplete data analysis. Therefore, a need to develop an algorithm that can predict the missing values efficiently and accurately. This research paper proposes a novel splitting-based IMV-RE (\u003cb\u003eI\u003c/b\u003emputing the \u003cb\u003eM\u003c/b\u003eissing \u003cb\u003eV\u003c/b\u003ealues in \u003cb\u003eR\u003c/b\u003eeal-Time \u003cb\u003eE\u003c/b\u003environment) algorithm to impute different missing values within a dataset. In the proposed IMV-RE algorithm, an upper limit is set for every class containing missing values that assist the algorithm to predict the missing values more accurately. The experimentation is performed on ten benchmark datasets that include completely numerical values as well as mixed data. Comparative experimental analysis indicates that the proposed IMV-RE algorithm outperforms the existing techniques in sensitivity to Accuracy, Root Mean Square Error (RMSE) and Coefficient of Determination (R\u003csup\u003e2\u003c/sup\u003e).\u003c/p\u003e","manuscriptTitle":"A Novel Algorithm for Imputing the Missing Values in Incomplete Datasets","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-06-16 16:17:31","doi":"10.21203/rs.3.rs-1729251/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"eb7c2c3c-2c89-4f4e-b99f-b7059d4613bb","owner":[],"postedDate":"June 16th, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2022-09-14T07:20:01+00:00","versionOfRecord":[],"versionCreatedAt":"2022-06-16 16:17:31","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-1729251","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1729251","identity":"rs-1729251","version":["v1"]},"buildId":"FbvkV6FR0MCFSLy54lSbu","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00