Application of machine learning technics to similar case imputation for missing values in Demographic and Health Survey (DHS) data | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Systematic Review Application of machine learning technics to similar case imputation for missing values in Demographic and Health Survey (DHS) data Mr. Cyprien HABINSHUTI, Prof. François NIRAGIRE This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9380937/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: The study sought to assess the performance of machine learning algorithms for handling missing data in demographic and Health Survey (DHS) dataset using similar case imputation method. In quantitative data analysis, imputation has the potential to greatly improve the knowledge available for mining high-quality compounds by supplying accurate predictions to fill in the gaps. Methods: This study used data from three rounds of the Rwanda Demographic and Health Survey (RDHS) conducted between 2015 and 2020. The study was conducted using three datasets with a total of 400,656 to 459,102 observations, each including 9,002, 7,856, and 8,092 rows with 51 columns. Twenty numerical variables and thirty-one categorical variables made up the merged dataset after concatenation. The Multiple Imputation Regression for ‘m’ iteration, Decision Tree based, Support Vector Machine (SVM) as classification algorithm, K-nearest neighbor (KNN) as Clustering algorithms, and Random Forest (RF) machine learning algorithms were applied and compared using performance metrics to identify the best algorithm to impute missing values. Results: It was found that Support Vector Machine (SVM) ranked first for imputing categorical variables, with an accuracy of 100% and precision of 100%. It was followed by Decision Tree, which achieved an accuracy of 79.9% and precision of 100%, Random Forest came in third with an accuracy of 78.0% and precision of 99.2%, while KNN ranked last with an accuracy of 67.1% and precision of 69.4%. For imputing numerical data, Random Forest performed the best, with both MSE and MAE values of 0. It was followed by Multiple Imputation by Chained Equations (MICE), which had an MSE of 1.53e^09 and an R² of 0.81. Finally, Support Vector Machine (SVM) regression ranked last among the machine learning models used for numerical data imputation. Conclusions: Performance measurements revealed that Random Forest is the optimal algorithm model for numerical variables, while Support Vector Machine classification, and Decision Tree, Random Forest are the best for categorical variables. It was concluded that the best machine learning algorithm for managing missing values, both categorically and numerically, is Random Forest. Information Retrieval and Management Machine learning Imputation SVM KNN MICE Random Forest Decision tree 1. Introduction Without data, nothing is possible in the contemporary digital era, and enormous volumes are used every second. When data is complete, it can be processed to yield the expected results. However, incomplete data makes it challenging to achieve accurate outcomes. This incomplete data is often referred to as missing data or unknown values. Consequently, researchers, academicians, industry professionals, and those in the medical field strive to avoid incomplete data. To address this, it must be converted into complete data before processing. Several scholars have investigated diverse techniques for managing absent data in datasets ( 1 ). In research, controlling missing values is crucial since their existence might introduce bias, diminish the reliability of results, and call into question the accuracy of conclusions. Missing data can cause statistical studies to be distorted, concealing the underlying relationships between variables and leading to inaccurate findings if not handled correctly ( 2 ). The research sample loses important information due to missing values. The efficiency and clarity of study findings are reduced the more missing data there are since more information is lost. Missing data can impair the validity and dependability of findings and, in certain situations, cause disruptions to the entire data structure, making analysis more difficult ( 2 ). Missing data is a significant challenge in both research and practical applications involving real-world datasets. For example, in the fields of educational and psychological research, Peugh and Enders discovered that, on average, 9.7% of the data was missing, with some datasets having up to 67% missing values. Similarly, Rombach et al. reported that missing data percentages varied from 1% to over 70%, with a median of 25%. Additionally, missing values are prevalent in over 45% of datasets, highlighting the widespread nature of this issue even in the University of California Irvine machine learning repository (UCI repository), one of the most widely used sources of datasets for machine learning and data science researchers ( 3 ). Addressing missing values in a dataset has been a persistent challenge across various disciplines. A key issue is that many algorithms cannot work effectively with incomplete datasets ( 4 ). The issue of missing values negatively impacts analysis results, particularly by causing biased parameter estimates ( 5 ). Because of things like human error, malfunctioning instruments, and other problems, real-world datasets are frequently noisy, redundant, incomplete, and inconsistent. In order to manage incomplete data efficiently, sophisticated algorithms are suggested to fill in the missing information. Numerous machine learning methods have been utilized to approximate and supplement these absent variables with logical substitutes ( 6 ). Ensuring data quality has become a significant difficulty in modern software programs as data pipelines become more sophisticated and important. Maintaining stringent standards for data quality is essential for machine learning (ML) applications to provide accurate prediction outcomes and thoughtful automated judgment. A prevalent problem in data quality is the absence of values. If incomplete datasets are not adequately treated, they can cause havoc with data pipelines and negatively affect downstream machine learning applications ( 7 ). Missing values represent one of the numerous difficulties using data from the real world. Since the quality of the data has a major impact on the accuracy and efficacy of machine learning models, researchers and data analysts need to use pertinent methods to deal with these inescapably missing values ( 8 , 9 ). This study is impactfully to the different researchers in the domain of health. It plays an important role in improving the research findings for students, scholars, and scientific researchers. 2. Related works and Materials 2.1. Related work The results from the first paper on a Classifier Ensemble Machine Learning Approach to Improve Efficiency for Missing Value Imputation settled the most accurate model was Support Vector Machine (SVM), and the stacking strategy utilizing GLM and to combine predictions from many models, such as Random Forest with svmRadial, rpart, glm, lda, and kNN was applied. The correlation between the two sub-models with the highest correlation, kNN and Logistic Regression (GLM), was just 0.517, which is still considered low because it falls significantly below the 0.75 criteria, even though it was the strongest ( 10 ). The second papers on a survey on missing data in machine learning, on the Iris dataset, it was found that KNN imputation performed better than RF imputation when RMSE was used as the assessment metric at two different missingness ratios. On the ID fan dataset, RF imputation, however, performed better across all missingness ratios ( 11 ). And the last paper, on Proper Imputation Techniques for Missing Values concluded that the HD imputation greatly improves prediction accuracy, especially for large datasets; C5.0 achieves strong classification accuracy across different types of data, even though it does not use all of the dataset's attributes; Expectation Maximization (EM) takes less time than KNN, especially when working with huge datasets. Mean imputation also violates assumptions about normality and weakens connections with other variables, but it might be appropriate in situations when fewer than 5% of the data are missing ( 12 ). Refereeing to the results from the above results researcher came to decide on the use of some of the algorithms in this study in order to apply them for handling missing data. More specifically, the researcher used KNN, SVM, Random Forest (RF), decision Tree and Multiple imputation algorithms. 2.2. Methods The study used the longitudinal study design. It used 3 cross-section data from RDHS that were combined to make one dataset (DHS 2015–2020). The specific dataset used was children or kids record (KR) as the one with almost all variables of interest to respond to child mortality. The combined dataset was used as standardized and allow comparability according to the variables of interest. The datasets were accessed from the official website for Demographic and Health Survey (DHS) for different countries 1 . The population of the study was all children whose 5 years at the time of interview as indicated in children or kids record (KR). The study variables are identified in the dataset of children/kids that are related to child mortality. The Table 1 gives variable name and its code in demographic health Survey kids dataset. Table 1 List of study variables S/N Children or kids record (KR) Variable code 1 Smokes/uses other V463X 2 Age at death B6 3 Age at death (months, imputed) B7 4 Age at first birth V212 5 Age at first cohabitation V511 6 Age of household head V152 7 Age of respondent at 1st birth V212 8 Age of respondent in 5-years group V013 9 Antenatal care visits M14 10 Birth order BORD 11 Birth weight in kilograms M19 12 Births in past year V209 13 Breastfeeding status (Whether the child currently breastfeed at the time of survey) M4 14 Child is twin B0 15 Cluster number V001 16 Current breastfeeding V404 17 Current contraceptive method V312 18 Current marital status V501 19 Currently residing with husband/partner V504 20 Date of birth B3 21 De jure region of residence V139 22 De jure type of place of residence V140 23 Delivery by caesarean section M17 24 Does not use tobacco V463Z 25 Highest education level S105 26 Highest educational level V106 27 Household has: electricity V119 28 Household head V151 29 Household number V002 30 Husband/partner’s education V701 31 Husband's desire for children V621 32 live with father (Living with one of the lists) B9 33 Number of household members V136 34 Number of living children V218 35 Place of delivery (health facility, resp home, elsewhere) M15 36 Preceding birth interval B11 37 Region V101, V024 38 Religion V130 39 Respondent currently working (employment) V714 40 Respondent slept under mosquito bed net V461 41 Respondent's year of birth V010 42 Selected for Domestic Violence module V044 43 Sex of child B4 44 Sex of household head: V151 45 Size of child at birth M18 46 Survival status (Child is alive, 1 = Yes, 0 = No) B5 47 Total children ever born V201 48 Type of place of residence V025 49 Type of place of residence V102 50 Wealth index combined: Wealth index combined V190 51 Wealth index for urban/rural V190A 52 When child put to breast V426 53 Women's individual sample weight (6 decimals) V005 54 Year of first cohabitation V508 55 Year of interview V007 [ Table 1 : List of study variables] The methods of dealing with missing data is the matching pattern and imputation method. The data is separated into groups based on similarities, and the missing value is imputed by choosing an algorithm-based set of observed values from each category. A number of procedures were used during the data analysis process, including data splitting, which divides the data into two halves. The training and test part where they are used to train machine learning model whereas the test part is used to evaluate the performance of machine learning model. The model was evaluated based on performance metrics. The ones for classification models are accuracy, precision, recall or sensitivity, and F-score while the performance metrics for numerical variables were mean absolute error (MAE), mean square error (MSE), R-squared value, and adjusted R 2 . The performance metrics for the model are compared and decide the best model accordingly. The algorithms used in the study are Multiple Imputation regression, Decision Tree based), classification (SVM), Clustering (KNN), and Random Forest (RF) and their performance metrics were compared to rank the top performing algorithm to impute missing values. 3. Results 3.1. Missingness in the dataset The Table 2 below gives the percentage of missing per the variable in the dataset. Table 2 Percentage of missing values per variables SN Variables Missing % 1. Women's individual sample weight (6 decimals) 0.000000 2. Date of interview (CMC) 0.000000 3. Age 0.000000 4. Age group 0.000000 5. Region 0.000000 6. Type of place of residence 0.000000 7. Highest education level 0.000000 8. Source of drinking water 0.020295 9. Type of toilet facility 0.069002 10. Household has: electricity 0.028413 11. Religion 0.064943 12. Number of household members (listed) 0.000000 13. Number of living children 0.000000 14. Sex of household head: 0.000000 15. Type of cooking fuel 0.000000 16. Wealth index combined: Wealth index combined 0.000000 17. Total children ever born 0.000000 18. Births in last five years 0.000000 19. Births in past year 0.000000 20. Age at first birth 0.000000 21. Number of living children 0.000000 22. Current contraceptive method 0.000000 23. Contraceptive use and intention 0.000000 24. Wanted last child 0.020295 25. Current breastfeeding 0.000000 26. When child put to breast 1.286683 27. Respondent slept under mosquito bed net 0.000000 28. Does not use cigarettes and tobacco 0.004059 29. Current marital status 0.000000 30. Currently residing with husband/partner 16.118034 31. Age at first cohabitation 7.768803 32. Husband's desire for children 24.909689 33. Husband/partner's educational attainment 10.687178 34. Respondent currently working 0.056825 35. Birth order 0.000000 36. Child is twin 0.000000 37. Date of birth (CMC) 0.000000 38. Sex of child 0.000000 39. Survival status (Child is alive, 1 = Yes, 0 = No) 0.000000 40. Age at death 95.575760 41. Age at death (months, imputed) 95.575760 42. Current age of child 4.424240 43. Currently residing with husband/partner 4.424240 44. Preceding birth interval (months) 26.748387 45. Live birth between births 26.427731 46. Duration of breastfeeding 0.227300 47. Number of antenatal visits during pregnancy 26.147664 48. Place of delivery (health facility, resp home, elsewhere) 0.069002 49. Delivery by caesarean section 0.069002 50. Size of child at birth 0.560133 51. Baby postnatal check within 2 months 52.628161 52. Birth to current child age 0.000000 [ Table 2 : Percentage of missing values per variables] 3.2. Machine learning algorithms for categorical variables The 4 types of machine learning algorithms were used in this study to treat missingness from categorical variables. These are algorithms for classification like K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Decision Tree, and Random Forest. The Table 3 gives the summary of performance metrics for each machine learning algorithm to treat missing values from the categorical variables. [ Table 3 : Machine learning algorithms for categorical data] Table 3 Performance of machine learning algorithms for categorical data Performance metrics Algorithms Accuracy Precision Recall F1 Score SVM 1.000 1.000 1.000 1.000 KNN 0.671 0.694 0.671 0.668 Random forest 0.780 0.992 0.780 0.992 Decision tree 0.800 1.000 0.800 0.889 3.3. Machine learning algorithms for numerical variables The 4 machine learning algorithms used to treat missingness from numerical variables are Multivariate Imputation by Chained Equations (MICE), Random Forest, and Support Vector Machines (SVM) for regression. The Table 4 gives the performance for each machine learning model. [ Table 4 : Machine learning algorithms for numerical data] Table 4 Performance of machine learning algorithms for numerical data Performance metrics Algorithm Mean Squared Error (MSE) Mean Absolute Error (MAE) R2 Score best value = 0; worst value = +∞ best value = 0; worst value = +∞ [0–1] Multivariate Imputation 1.53E + 09 3248.788 0.810034 Random Forest 0 0 - Support Vector Machines 8.586746 677.0923 0.186535 3.4. The evaluation of machine learning algorithms Based on performance metrics, the machine learning algorithm to treat missingness from categorical variables is the SVM classification with the accuracy of 100% and precision of 100%, and the Decision Tree with the accuracy of 79.9% and precision of 100%, and Random forest with 0.78 of accuracy and 0.992 precision while the one to treat missing values from numerical variables is Random Forest with MSE and MAE value of 0 and the Multivariate Imputation by Chained Equations (MICE) with MSE of 1.53. e^09 and R 2 of 0.81. In addition, the machine learning algorithm to treat missingness in both numerical and categorical variables is Random Forest. 4. Discussion In this study, it was revealed that the best algorithm model for imputing categorical variables is Support Vector Machine classification, Decision Tree and Random Forest while the one for numerical variables is Random Forest. It was also revealed that Random Forest is the most effective machine learning algorithm for handling both categorical and numerical missing values. According to the research titled “Machine Learning Based Missing Values Imputation in Categorical Datasets” by Muhammad Ishaq and collaborators, machine learning algorithms demonstrated strong performance in predicting and imputing missing values. The effectiveness of these algorithms varied depending on the dataset and the pattern of missing values ( 13 ). The study evaluated a range of machine learning models, including ensemble models that employ the Error Correcting Output Codes (ECOC) architecture. These models included ensembles based on SVM and KNN, as well as an ensemble classifier that integrated both SVM and KNN. The results validate the efficaciousness of various machine learning techniques in handling incomplete data. Another study by Dimitris Bertsimas titled "From Predictive Methods to Missing Data Imputation: An Optimization Approach" lends support to this, framing the traditional missing data problem as a non-convex optimization problem and utilizing a variety of prediction models. Using effective first-order optimization approaches, the research presents a new family of imputation algorithms dubbed Opt. Impute, which offers high-quality solutions to this challenge. To validate these techniques, numerous computational experiments were carried out ( 14 ). The findings in the study are supported by others research in the same field. One of the documented on categorical variables revealed that the random forest was the effective machine learning algorithm to impute the missing values ( 13 ). The researches indicated that the performance metrics indicated that random forest is the best methods for predicting comprehensive strength with the numerical dataset. it is evident that machine learning can help to impute missing values. The researcher suggests to use machine learning algorithms to impute missing values before proceeding with other data analysis steps for further future research to improve the prediction accuracy. The study was limited to a slight number of machine learning algorithms and specific dataset related to data generated for responding to health conditions. The researcher would suggest to other researchers to explore other machine learning algorithms using the dataset from the domain of their interest. 5. Conclusions The research concludes that the best algorithm to treat missing for categorical values is Support Vector Machine for classification, Decision Tree and Random Forest while the best to treat numerical values is Random Forest and the Multivariate Imputation by Chained Equations (MICE) algorithms. It is evidence that the Random Forest is the most effective machine learning algorithm for handling both categorical and numerical missing values. Abbreviations RDHS Rwanda Demographic and Health Survey SVM Support Vector Machine KNN K-nearest neighbor RF Random Forest MICE Multiple Imputation by Chained Equations MSE Mean Absolute Error MAE Mean Square Error UCI University of California Irvine GLM Logistic Regression DBMSs Database Management systems ML Maximum Learning ECOC Correcting Output Codes Declarations 6.2. Ethical and approval and consent to participate I hereby declare that this research was conducted in full accordance with the University’s established academic, ethical, and professional standards required for the completion of a Master’s degree in Data Science. All methods, procedures, and analytical techniques employed in this study adhere to the principles of integrity, transparency, and responsible research conduct. The datasets used in this study were ethically cleared and accessed from the DHS Program (dhsprogram.com) under authorized use. No attempt was made to identify individual respondents, and all data handling complied with applicable data protection regulations and best practices for confidentiality, privacy, and responsible data stewardship. This manuscript represents my original work and has not been submitted, in whole or in part, toward any other journal. All sources of information, data, and prior research have been properly acknowledged through accurate citation and referencing. Any form of plagiarism, data falsification, data manipulation, or unauthorized reuse of materials has been strictly avoided. No additional institutional ethical approval was required for this study. All analyses were conducted under academic supervision and in alignment with recognized standards of fairness, accountability, and transparency in Data Science research. I affirm that this study upholds the principles of academic integrity, respect for data subjects, and responsible knowledge production, as expected within a Master’s degree program in Data Science. 6.3. Consent for publication Not applicable 6.4. Availability of data and materials The data used in this study is publicly available at The DHS Program. The specific dataset analyzed in this study can be accessed directly through the DHS dataset portal. The code for data analysis are available on: google colab. 6.5. Competing interest The researcher declares that there are no competing interests associated with this study. 6.6. Funding Declaration The researcher declares that this study was conducted without any financial support. 6.7. Authors' contributions Mr. HABINSHUTI Cyprien accessed the study data and led the proposal development. He contributed to the data collection strategy, methodological design, data cleaning, data analysis, and interpretation of the results, and drafted the manuscript. Email address: [email protected] Prof. François NIRAGIRE contributed to the methodological development and conceptualization of the study and reviewed and edited the manuscript. Email address: [email protected] 6.8. Acknowledgements The author sincerely acknowledges the administrative support of the University of Rwanda and expresses appreciation to all staff members of the African Center of Excellence in Data Science (AC-DS) for their continuous guidance and support during this study and the Master’s program in Data Science. References Sivakani R, Ansari GA (2020) Imputation Using Machine Learning Techniques. 4th International Conference on Computer, Communication and Signal Processing, ICCCSP 2020. ;0–5 Setiawan I, Gernowo R, Warsito B (2023) A Systematic Literature Review on Missing Values: Research Trends, Datasets, Methods and Frameworks. E3S Web of Conferences. ;448 Dong Y, Peng CYJ (2013) Principled missing data methods for researchers (Expectation Maximaization explained). Springerplus 2(1):1–17 Platias C, Petasis G (2020) A comparison of machine learning methods for data imputation. ACM International Conference Proceeding Series. ;150–9 Nadzurah ZA, Amelia Ritahani I, Nurul A (2018) Performance Analysis of Machine Learning Algorithms for Missing Value Imputation. Int J Adv Comput Sci Appl. ;9(6) Ismail AR, Abidin NZ, Maen MK (2022) Systematic Review on Missing Data Imputation Techniques with Machine Learning Algorithms for Healthcare. J Rob Control (JRC) 3(2):143–152 Jäger S, Allhorn A, Bießmann F (2021) A Benchmark for Data Imputation Methods. Front Big Data 4(July):1–16 Cappelletti L, Fontana T, Di Donato GW, Di Tucci L, Casiraghi E, Valentini G (2020) Complex data imputation by auto-encoders and convolutional neural networks—A case study on genome gap-filling. Computers. ;9(2) Oluwaseye L, J* A, Doorsamy W, Paul S (2022) A Review of Missing Data Handling Techniques for Machine Learning. International Journal of Innovative Technology and Interdisciplinary Sciences wwwIJITIS.org [Internet]. ;5(3):971–1005. Available from: https://doi.org/10.15157/IJITIS.2022.5.3.971-1005 Chhabra G, Vashisht V, Ranjan J (2018) A classifier ensemble machine learning approach to improve efficiency for missing value imputation. 2018 International Conference on Computing, Power and Communication Technologies, GUCON. 2019;23–7 Emmanuel T, Maupong T, Mpoeleng D, Semong T, Mphago B (2021) A survey on missing data in machine learning [Internet]. Journal of Big Data. Springer International Publishing; Available from: https://doi.org/10.1186/s40537-021-00516-9 Aljuaid T, Sasi S (2016) Proper imputation techniques for missing values in data sets. Proceedings of the 2016 International Conference on Data Science and Engineering, ICDSE. 2017 Ishaq M, Zahir S, Iftikhar L, Bulbul MF, Rho S, Lee MY (2024) Machine Learning Based missing data Imputation in Categorical Datasets. IEEE Access. ;1–29 Bertsimas D, Pawlowski C, Zhuo YD (2018) From predictive methods to missing data imputation: An optimization approach. J Mach Learn Res 18:1–39 Footnotes https://dhsprogram.com/data/dataset_admin/login_main.cfm;jsessionid=A8E4992909AED729EABAB7C9D5579E85.cfusion?CFID=306774449&CFTOKEN=c153a0a8d77f8b19-F7DAA648-D7AD-AB05-AAB740663C775C2B Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9380937","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Systematic Review","associatedPublications":[],"authors":[{"id":621080905,"identity":"97a20f68-82f5-46e8-a87c-8fbd0e6f9f5e","order_by":0,"name":"Mr. Cyprien HABINSHUTI","email":"","orcid":"","institution":"University of Rwanda","correspondingAuthor":false,"prefix":"Mr.","firstName":"Cyprien","middleName":"","lastName":"HABINSHUTI","suffix":""},{"id":621080906,"identity":"8ac0a706-0f5e-44f5-b50f-37a862268edf","order_by":1,"name":"Prof. François NIRAGIRE","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA2UlEQVRIiWNgGAWjYBAC9gbGBiAlAWIBaQMLwlp4DsC08BwAaZEgRguMJZHAANFLUIt0c9uHH78s8vkln1/d8KNAgoG/vTsBvxaZg80ze/skLGfOzim72QN0mMSZsxvwarGXSGxm4O2RMDC4nZN2gweoxUAiF78WHqAWxr9ALfY3z6Td/EOsFmaeH0BbJNiP3SbaFmbZBgkDiTM5bLdlDCR4CPqFRyL9MeObP3UG/O3Hn91888dGjr+9F78WMGBsA+s2AJOElYPBHxDB/oBI1aNgFIyCUTDSAAC5p0PSr7yhJAAAAABJRU5ErkJggg==","orcid":"","institution":"University of Rwanda","correspondingAuthor":true,"prefix":"","firstName":"Prof.","middleName":"François","lastName":"NIRAGIRE","suffix":""}],"badges":[],"createdAt":"2026-04-10 15:04:17","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-9380937/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9380937/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":106961070,"identity":"fa5a76db-89fb-4ebf-a4ef-e1ea161a8153","added_by":"auto","created_at":"2026-04-15 09:24:10","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":902608,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9380937/v1/733b1135-97ec-4be9-aab7-c4ad95a06e9e.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003eApplication of machine learning technics to similar case imputation for missing values in Demographic and Health Survey (DHS) data\u003c/p\u003e","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eWithout data, nothing is possible in the contemporary digital era, and enormous volumes are used every second. When data is complete, it can be processed to yield the expected results. However, incomplete data makes it challenging to achieve accurate outcomes. This incomplete data is often referred to as missing data or unknown values. Consequently, researchers, academicians, industry professionals, and those in the medical field strive to avoid incomplete data. To address this, it must be converted into complete data before processing. Several scholars have investigated diverse techniques for managing absent data in datasets (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eIn research, controlling missing values is crucial since their existence might introduce bias, diminish the reliability of results, and call into question the accuracy of conclusions. Missing data can cause statistical studies to be distorted, concealing the underlying relationships between variables and leading to inaccurate findings if not handled correctly (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). The research sample loses important information due to missing values. The efficiency and clarity of study findings are reduced the more missing data there are since more information is lost. Missing data can impair the validity and dependability of findings and, in certain situations, cause disruptions to the entire data structure, making analysis more difficult (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eMissing data is a significant challenge in both research and practical applications involving real-world datasets. For example, in the fields of educational and psychological research, Peugh and Enders discovered that, on average, 9.7% of the data was missing, with some datasets having up to 67% missing values. Similarly, Rombach et al. reported that missing data percentages varied from 1% to over 70%, with a median of 25%. Additionally, missing values are prevalent in over 45% of datasets, highlighting the widespread nature of this issue even in the University of California Irvine machine learning repository (UCI repository), one of the most widely used sources of datasets for machine learning and data science researchers (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eAddressing missing values in a dataset has been a persistent challenge across various disciplines. A key issue is that many algorithms cannot work effectively with incomplete datasets (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e). The issue of missing values negatively impacts analysis results, particularly by causing biased parameter estimates (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e). Because of things like human error, malfunctioning instruments, and other problems, real-world datasets are frequently noisy, redundant, incomplete, and inconsistent. In order to manage incomplete data efficiently, sophisticated algorithms are suggested to fill in the missing information. Numerous machine learning methods have been utilized to approximate and supplement these absent variables with logical substitutes (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eEnsuring data quality has become a significant difficulty in modern software programs as data pipelines become more sophisticated and important. Maintaining stringent standards for data quality is essential for machine learning (ML) applications to provide accurate prediction outcomes and thoughtful automated judgment. A prevalent problem in data quality is the absence of values. If incomplete datasets are not adequately treated, they can cause havoc with data pipelines and negatively affect downstream machine learning applications (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eMissing values represent one of the numerous difficulties using data from the real world. Since the quality of the data has a major impact on the accuracy and efficacy of machine learning models, researchers and data analysts need to use pertinent methods to deal with these inescapably missing values (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThis study is impactfully to the different researchers in the domain of health. It plays an important role in improving the research findings for students, scholars, and scientific researchers.\u003c/p\u003e"},{"header":"2. Related works and Materials","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. Related work\u003c/h2\u003e \u003cp\u003eThe results from the first paper on a Classifier Ensemble Machine Learning Approach to Improve Efficiency for Missing Value Imputation settled the most accurate model was Support Vector Machine (SVM), and the stacking strategy utilizing GLM and to combine predictions from many models, such as Random Forest with svmRadial, rpart, glm, lda, and kNN was applied. The correlation between the two sub-models with the highest correlation, kNN and Logistic Regression (GLM), was just 0.517, which is still considered low because it falls significantly below the 0.75 criteria, even though it was the strongest (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe second papers on a survey on missing data in machine learning, on the Iris dataset, it was found that KNN imputation performed better than RF imputation when RMSE was used as the assessment metric at two different missingness ratios. On the ID fan dataset, RF imputation, however, performed better across all missingness ratios (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eAnd the last paper, on Proper Imputation Techniques for Missing Values concluded that the HD imputation greatly improves prediction accuracy, especially for large datasets; C5.0 achieves strong classification accuracy across different types of data, even though it does not use all of the dataset's attributes; Expectation Maximization (EM) takes less time than KNN, especially when working with huge datasets. Mean imputation also violates assumptions about normality and weakens connections with other variables, but it might be appropriate in situations when fewer than 5% of the data are missing (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e). Refereeing to the results from the above results researcher came to decide on the use of some of the algorithms in this study in order to apply them for handling missing data. More specifically, the researcher used KNN, SVM, Random Forest (RF), decision Tree and Multiple imputation algorithms.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2. Methods\u003c/h2\u003e \u003cp\u003eThe study used the longitudinal study design. It used 3 cross-section data from RDHS that were combined to make one dataset (DHS 2015\u0026ndash;2020). The specific dataset used was children or kids record (KR) as the one with almost all variables of interest to respond to child mortality. The combined dataset was used as standardized and allow comparability according to the variables of interest. The datasets were accessed from the official website for Demographic and Health Survey (DHS) for different countries\u003csup\u003e1\u003c/sup\u003e. The population of the study was all children whose 5 years at the time of interview as indicated in children or kids record (KR).\u003c/p\u003e \u003cp\u003eThe study variables are identified in the dataset of children/kids that are related to child mortality. The Table \u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e gives variable name and its code in demographic health Survey kids dataset.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eList of study variables\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eS/N\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eChildren or kids record (KR)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eVariable code\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSmokes/uses other\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV463X\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge at death\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eB6\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge at death (months, imputed)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eB7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge at first birth\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV212\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge at first cohabitation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV511\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge of household head\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV152\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge of respondent at 1st birth\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV212\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge of respondent in 5-years group\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV013\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAntenatal care visits\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eM14\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBirth order\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBORD\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBirth weight in kilograms\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eM19\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBirths in past year\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV209\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBreastfeeding status (Whether the child currently breastfeed at the time of survey)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eM4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eChild is twin\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eB0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCluster number\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCurrent breastfeeding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV404\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCurrent contraceptive method\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV312\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCurrent marital status\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV501\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCurrently residing with husband/partner\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV504\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDate of birth\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eB3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDe jure region of residence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV139\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDe jure type of place of residence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV140\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e23\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDelivery by caesarean section\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eM17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDoes not use tobacco\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV463Z\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHighest education level\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eS105\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHighest educational level\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV106\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHousehold has: electricity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV119\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHousehold head\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV151\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e29\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHousehold number\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV002\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHusband/partner\u0026rsquo;s education\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV701\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHusband's desire for children\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV621\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003elive with father (Living with one of the lists)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eB9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber of household members\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV136\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e34\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber of living children\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV218\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e35\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePlace of delivery (health facility, resp home, elsewhere)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eM15\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePreceding birth interval\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eB11\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e37\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRegion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV101, V024\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e38\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eReligion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV130\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRespondent currently working (employment)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV714\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRespondent slept under mosquito bed net\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV461\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e41\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRespondent's year of birth\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV010\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e42\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSelected for Domestic Violence module\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV044\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e43\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSex of child\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eB4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e44\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSex of household head:\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV151\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSize of child at birth\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eM18\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e46\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSurvival status (Child is alive, 1\u0026thinsp;=\u0026thinsp;Yes, 0\u0026thinsp;=\u0026thinsp;No)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eB5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTotal children ever born\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV201\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eType of place of residence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV025\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eType of place of residence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV102\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWealth index combined: Wealth index combined\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV190\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e51\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWealth index for urban/rural\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV190A\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWhen child put to breast\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV426\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWomen's individual sample weight (6 decimals)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV005\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eYear of first cohabitation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV508\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eYear of interview\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eV007\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003e[\u003c/em\u003eTable\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e: \u003cem\u003eList of study variables]\u003c/em\u003e\u003c/p\u003e \u003cp\u003eThe methods of dealing with missing data is the matching pattern and imputation method. The data is separated into groups based on similarities, and the missing value is imputed by choosing an algorithm-based set of observed values from each category. A number of procedures were used during the data analysis process, including data splitting, which divides the data into two halves. The training and test part where they are used to train machine learning model whereas the test part is used to evaluate the performance of machine learning model.\u003c/p\u003e \u003cp\u003eThe model was evaluated based on performance metrics. The ones for classification models are accuracy, precision, recall or sensitivity, and F-score while the performance metrics for numerical variables were mean absolute error (MAE), mean square error (MSE), R-squared value, and adjusted R\u003csup\u003e2\u003c/sup\u003e. The performance metrics for the model are compared and decide the best model accordingly. The algorithms used in the study are Multiple Imputation regression, Decision Tree based), classification (SVM), Clustering (KNN), and Random Forest (RF) and their performance metrics were compared to rank the top performing algorithm to impute missing values.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e3.1. Missingness in the dataset\u003c/h2\u003e \u003cp\u003eThe Table \u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e below gives the percentage of missing per the variable in the dataset.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePercentage of missing values per variables\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSN\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eVariables\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMissing %\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1.\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWomen's individual sample weight (6 decimals)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDate of interview (CMC)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge group\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRegion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eType of place of residence\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e7.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHighest education level\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e8.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSource of drinking water\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.020295\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e9.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eType of toilet facility\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.069002\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e10.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHousehold has: electricity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.028413\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e11.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eReligion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.064943\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber of household members (listed)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e13.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber of living children\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e14.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSex of household head:\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e15.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eType of cooking fuel\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e16.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWealth index combined: Wealth index combined\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e17.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTotal children ever born\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e18.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBirths in last five years\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e19.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBirths in past year\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e20.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge at first birth\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e21.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber of living children\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e22.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCurrent contraceptive method\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e23.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eContraceptive use and intention\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e24.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWanted last child\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.020295\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e25.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCurrent breastfeeding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e26.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eWhen child put to breast\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.286683\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e27.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRespondent slept under mosquito bed net\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e28.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDoes not use cigarettes and tobacco\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.004059\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e29.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCurrent marital status\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e30.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCurrently residing with husband/partner\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e16.118034\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e31.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge at first cohabitation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e7.768803\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e32.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHusband's desire for children\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e24.909689\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e33.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHusband/partner's educational attainment\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e10.687178\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e34.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRespondent currently working\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.056825\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e35.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBirth order\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e36.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eChild is twin\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e37.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDate of birth (CMC)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e38.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSex of child\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e39.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSurvival status (Child is alive, 1\u0026thinsp;=\u0026thinsp;Yes, 0\u0026thinsp;=\u0026thinsp;No)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e40.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge at death\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e95.575760\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e41.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAge at death (months, imputed)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e95.575760\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e42.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCurrent age of child\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.424240\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e43.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCurrently residing with husband/partner\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.424240\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e44.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePreceding birth interval (months)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e26.748387\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e45.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLive birth between births\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e26.427731\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e46.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDuration of breastfeeding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.227300\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e47.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber of antenatal visits during pregnancy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e26.147664\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e48.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePlace of delivery (health facility, resp home, elsewhere)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.069002\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e49.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDelivery by caesarean section\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.069002\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e50.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSize of child at birth\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.560133\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e51.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBaby postnatal check within 2 months\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e52.628161\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e52.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eBirth to current child age\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.000000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003e[\u003c/em\u003eTable\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e: \u003cem\u003ePercentage of missing values per variables]\u003c/em\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e3.2. Machine learning algorithms for categorical variables\u003c/h2\u003e \u003cp\u003eThe 4 types of machine learning algorithms were used in this study to treat missingness from categorical variables. These are algorithms for classification like K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Decision Tree, and Random Forest. The Table \u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e gives the summary of performance metrics for each machine learning algorithm to treat missing values from the categorical variables.\u003c/p\u003e \u003cp\u003e \u003cem\u003e[\u003c/em\u003eTable\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e: \u003cem\u003eMachine learning algorithms for categorical data]\u003c/em\u003e\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance of machine learning algorithms for categorical data\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"5\" nameend=\"c5\" namest=\"c1\"\u003e \u003cp\u003ePerformance metrics\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlgorithms\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eF1 Score\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSVM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.671\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.694\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.671\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.668\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.780\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.992\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.780\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.992\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDecision tree\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.800\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.800\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.889\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e3.3. Machine learning algorithms for numerical variables\u003c/h2\u003e \u003cp\u003eThe 4 machine learning algorithms used to treat missingness from numerical variables are Multivariate Imputation by Chained Equations (MICE), Random Forest, and Support Vector Machines (SVM) for regression. The Table \u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e gives the performance for each machine learning model.\u003c/p\u003e \u003cp\u003e \u003cem\u003e[\u003c/em\u003eTable\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e: \u003cem\u003eMachine learning algorithms for numerical data]\u003c/em\u003e\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance of machine learning algorithms for numerical data\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"4\" nameend=\"c4\" namest=\"c1\"\u003e \u003cp\u003ePerformance metrics\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eAlgorithm\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMean Squared Error (MSE)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMean Absolute Error (MAE)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eR2 Score\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ebest value\u0026thinsp;=\u0026thinsp;0; worst value = +\u0026infin;\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ebest value\u0026thinsp;=\u0026thinsp;0; worst value = +\u0026infin;\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e[0\u0026ndash;1]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMultivariate Imputation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.53E\u0026thinsp;+\u0026thinsp;09\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3248.788\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.810034\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSupport Vector Machines\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e8.586746\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e677.0923\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.186535\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e3.4. The evaluation of machine learning algorithms\u003c/h2\u003e \u003cp\u003eBased on performance metrics, the machine learning algorithm to treat missingness from categorical variables is the SVM classification with the accuracy of 100% and precision of 100%, and the Decision Tree with the accuracy of 79.9% and precision of 100%, and Random forest with 0.78 of accuracy and 0.992 precision while the one to treat missing values from numerical variables is Random Forest with MSE and MAE value of 0 and the Multivariate Imputation by Chained Equations (MICE) with MSE of 1.53. e^09 and R\u003csup\u003e2\u003c/sup\u003e of 0.81. In addition, the machine learning algorithm to treat missingness in both numerical and categorical variables is Random Forest.\u003c/p\u003e \u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eIn this study, it was revealed that the best algorithm model for imputing categorical variables is Support Vector Machine classification, Decision Tree and Random Forest while the one for numerical variables is Random Forest. It was also revealed that Random Forest is the most effective machine learning algorithm for handling both categorical and numerical missing values.\u003c/p\u003e \u003cp\u003eAccording to the research titled \u0026ldquo;Machine Learning Based Missing Values Imputation in Categorical Datasets\u0026rdquo; by Muhammad Ishaq and collaborators, machine learning algorithms demonstrated strong performance in predicting and imputing missing values. The effectiveness of these algorithms varied depending on the dataset and the pattern of missing values (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). The study evaluated a range of machine learning models, including ensemble models that employ the Error Correcting Output Codes (ECOC) architecture. These models included ensembles based on SVM and KNN, as well as an ensemble classifier that integrated both SVM and KNN. The results validate the efficaciousness of various machine learning techniques in handling incomplete data.\u003c/p\u003e \u003cp\u003eAnother study by Dimitris Bertsimas titled \"From Predictive Methods to Missing Data Imputation: An Optimization Approach\" lends support to this, framing the traditional missing data problem as a non-convex optimization problem and utilizing a variety of prediction models. Using effective first-order optimization approaches, the research presents a new family of imputation algorithms dubbed Opt. Impute, which offers high-quality solutions to this challenge. To validate these techniques, numerous computational experiments were carried out (\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe findings in the study are supported by others research in the same field. One of the documented on categorical variables revealed that the random forest was the effective machine learning algorithm to impute the missing values (\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). The researches indicated that the performance metrics indicated that random forest is the best methods for predicting comprehensive strength with the numerical dataset. it is evident that machine learning can help to impute missing values.\u003c/p\u003e \u003cp\u003eThe researcher suggests to use machine learning algorithms to impute missing values before proceeding with other data analysis steps for further future research to improve the prediction accuracy.\u003c/p\u003e \u003cp\u003eThe study was limited to a slight number of machine learning algorithms and specific dataset related to data generated for responding to health conditions. The researcher would suggest to other researchers to explore other machine learning algorithms using the dataset from the domain of their interest.\u003c/p\u003e"},{"header":"5. Conclusions","content":"\u003cp\u003eThe research concludes that the best algorithm to treat missing for categorical values is Support Vector Machine for classification, Decision Tree and Random Forest while the best to treat numerical values is Random Forest and the Multivariate Imputation by Chained Equations (MICE) algorithms. It is evidence that the Random Forest is the most effective machine learning algorithm for handling both categorical and numerical missing values.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eRDHS\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eRwanda Demographic and Health Survey\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eSVM\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eSupport Vector Machine\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eKNN\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eK-nearest neighbor\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eRF\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eMICE\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMultiple Imputation by Chained Equations\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eMSE\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMean Absolute Error\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eMAE\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMean Square Error\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eUCI\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eUniversity of California Irvine\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eGLM\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eLogistic Regression\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eDBMSs\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eDatabase Management systems\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eML\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMaximum Learning\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eECOC\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eCorrecting Output Codes\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003e6.2. Ethical and approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eI hereby declare that this research was conducted in full accordance with the University\u0026rsquo;s established academic, ethical, and professional standards required for the completion of a Master\u0026rsquo;s degree in Data Science. All methods, procedures, and analytical techniques employed in this study adhere to the principles of integrity, transparency, and responsible research conduct.\u003c/p\u003e\n\u003cp\u003eThe datasets used in this study were ethically cleared and accessed from the DHS Program (dhsprogram.com) under authorized use. No attempt was made to identify individual respondents, and all data handling complied with applicable data protection regulations and best practices for confidentiality, privacy, and responsible data stewardship.\u003c/p\u003e\n\u003cp\u003eThis manuscript represents my original work and has not been submitted, in whole or in part, toward any other journal. All sources of information, data, and prior research have been properly acknowledged through accurate citation and referencing. Any form of plagiarism, data falsification, data manipulation, or unauthorized reuse of materials has been strictly avoided.\u003c/p\u003e\n\u003cp\u003eNo additional institutional ethical approval was required for this study. All analyses were conducted under academic supervision and in alignment with recognized standards of fairness, accountability, and transparency in Data Science research.\u003c/p\u003e\n\u003cp\u003eI affirm that this study upholds the principles of academic integrity, respect for data subjects, and responsible knowledge production, as expected within a Master\u0026rsquo;s degree program in Data Science.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e6.3. Consent for publication\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e6.4. Availability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data used in this study is publicly available at The DHS Program. The specific dataset analyzed in this study can be accessed directly through the DHS dataset portal. The code for data analysis are available on: google colab.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e6.5. Competing interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe researcher declares that there are no competing interests associated with this study.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e6.6. Funding Declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe researcher declares that this study was conducted without any financial support.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e6.7. Authors\u0026apos; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMr. HABINSHUTI Cyprien\u003c/strong\u003e accessed the study data and led the proposal development. He contributed to the data collection strategy, methodological design, data cleaning, data analysis, and interpretation of the results, and drafted the manuscript.\u003c/p\u003e\n\u003cp\u003eEmail address:
[email protected]\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eProf. Fran\u0026ccedil;ois NIRAGIRE\u003c/strong\u003e contributed to the methodological development and conceptualization of the study and reviewed and edited the manuscript.\u003c/p\u003e\n\u003cp\u003eEmail address:
[email protected]\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e6.8. Acknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author sincerely acknowledges the administrative support of the University of Rwanda and expresses appreciation to all staff members of the African Center of Excellence in Data Science (AC-DS) for their continuous guidance and support during this study and the Master\u0026rsquo;s program in Data Science.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eSivakani R, Ansari GA (2020) Imputation Using Machine Learning Techniques. 4th International Conference on Computer, Communication and Signal Processing, ICCCSP 2020. ;0\u0026ndash;5\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSetiawan I, Gernowo R, Warsito B (2023) A Systematic Literature Review on Missing Values: Research Trends, Datasets, Methods and Frameworks. E3S Web of Conferences. ;448\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDong Y, Peng CYJ (2013) Principled missing data methods for researchers (Expectation Maximaization explained). Springerplus 2(1):1\u0026ndash;17\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePlatias C, Petasis G (2020) A comparison of machine learning methods for data imputation. ACM International Conference Proceeding Series. ;150\u0026ndash;9\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNadzurah ZA, Amelia Ritahani I, Nurul A (2018) Performance Analysis of Machine Learning Algorithms for Missing Value Imputation. Int J Adv Comput Sci Appl. ;9(6)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIsmail AR, Abidin NZ, Maen MK (2022) Systematic Review on Missing Data Imputation Techniques with Machine Learning Algorithms for Healthcare. J Rob Control (JRC) 3(2):143\u0026ndash;152\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJ\u0026auml;ger S, Allhorn A, Bie\u0026szlig;mann F (2021) A Benchmark for Data Imputation Methods. Front Big Data 4(July):1\u0026ndash;16\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCappelletti L, Fontana T, Di Donato GW, Di Tucci L, Casiraghi E, Valentini G (2020) Complex data imputation by auto-encoders and convolutional neural networks\u0026mdash;A case study on genome gap-filling. Computers. ;9(2)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOluwaseye L, J* A, Doorsamy W, Paul S (2022) A Review of Missing Data Handling Techniques for Machine Learning. International Journal of Innovative Technology and Interdisciplinary Sciences wwwIJITIS.org [Internet]. ;5(3):971\u0026ndash;1005. Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.15157/IJITIS.2022.5.3.971-1005\u003c/span\u003e\u003cspan address=\"10.15157/IJITIS.2022.5.3.971-1005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChhabra G, Vashisht V, Ranjan J (2018) A classifier ensemble machine learning approach to improve efficiency for missing value imputation. 2018 International Conference on Computing, Power and Communication Technologies, GUCON. 2019;23\u0026ndash;7\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEmmanuel T, Maupong T, Mpoeleng D, Semong T, Mphago B (2021) A survey on missing data in machine learning [Internet]. Journal of Big Data. Springer International Publishing; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s40537-021-00516-9\u003c/span\u003e\u003cspan address=\"10.1186/s40537-021-00516-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAljuaid T, Sasi S (2016) Proper imputation techniques for missing values in data sets. Proceedings of the 2016 International Conference on Data Science and Engineering, ICDSE. 2017\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIshaq M, Zahir S, Iftikhar L, Bulbul MF, Rho S, Lee MY (2024) Machine Learning Based missing data Imputation in Categorical Datasets. IEEE Access. ;1\u0026ndash;29\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBertsimas D, Pawlowski C, Zhuo YD (2018) From predictive methods to missing data imputation: An optimization approach. J Mach Learn Res 18:1\u0026ndash;39\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Footnotes","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://dhsprogram.com/data/dataset_admin/login_main.cfm;jsessionid=A8E4992909AED729EABAB7C9D5579E85.cfusion?CFID=306774449\u0026amp;CFTOKEN=c153a0a8d77f8b19-F7DAA648-D7AD-AB05-AAB740663C775C2B\u003c/span\u003e\u003cspan address=\"https://dhsprogram.com/data/dataset_admin/login_main.cfm;jsessionid=A8E4992909AED729EABAB7C9D5579E85.cfusion?CFID=306774449\u0026amp;CFTOKEN=c153a0a8d77f8b19-F7DAA648-D7AD-AB05-AAB740663C775C2B\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Machine learning, Imputation, SVM, KNN, MICE, Random Forest, Decision tree","lastPublishedDoi":"10.21203/rs.3.rs-9380937/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9380937/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground: \u003c/strong\u003eThe study sought to assess the performance of machine learning algorithms for handling missing data in demographic and Health Survey (DHS) dataset using similar case imputation method. In quantitative data analysis,\u003cstrong\u003e \u003c/strong\u003eimputation has the potential to greatly improve the knowledge available for mining high-quality compounds by supplying accurate predictions to fill in the gaps.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods: \u003c/strong\u003eThis study used data from three rounds of the Rwanda Demographic and Health Survey (RDHS) conducted between 2015 and 2020. The study was conducted using three datasets with a total of 400,656 to 459,102 observations, each including 9,002, 7,856, and 8,092 rows with 51 columns. Twenty numerical variables and thirty-one categorical variables made up the merged dataset after concatenation.\u003c/p\u003e\n\u003cp\u003eThe Multiple Imputation Regression for ‘m’ iteration, Decision Tree based, Support Vector Machine (SVM) as classification algorithm, K-nearest neighbor (KNN) as Clustering algorithms, and Random Forest (RF) machine learning algorithms were applied and compared using performance metrics to identify the best algorithm to impute missing values.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults: \u003c/strong\u003eIt was found that Support Vector Machine (SVM) ranked first for imputing categorical variables, with an accuracy of 100% and precision of 100%. It was followed by Decision Tree, which achieved an accuracy of 79.9% and precision of 100%, Random Forest came in third with an accuracy of 78.0% and precision of 99.2%, while KNN ranked last with an accuracy of 67.1% and precision of 69.4%.\u003c/p\u003e\n\u003cp\u003eFor imputing numerical data, Random Forest performed the best, with both MSE and MAE values of 0. It was followed by Multiple Imputation by Chained Equations (MICE), which had an MSE of 1.53e^09 and an R² of 0.81. Finally, Support Vector Machine (SVM) regression ranked last among the machine learning models used for numerical data imputation.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions: \u003c/strong\u003ePerformance measurements revealed that Random Forest is the optimal algorithm model for numerical variables, while Support Vector Machine classification, and Decision Tree, Random Forest are the best for categorical variables. It was concluded that the best machine learning algorithm for managing missing values, both categorically and numerically, is Random Forest.\u003c/p\u003e","manuscriptTitle":"Application of machine learning technics to similar case imputation for missing values in Demographic and Health Survey (DHS) data","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-14 18:10:47","doi":"10.21203/rs.3.rs-9380937/v1","editorialEvents":[{"type":"communityComments","content":1}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"274b1503-6b11-43e0-9ed9-bd0e7dce7f3d","owner":[],"postedDate":"April 14th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":66101548,"name":"Information Retrieval and Management"}],"tags":[],"updatedAt":"2026-04-14T18:10:47+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-14 18:10:47","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9380937","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9380937","identity":"rs-9380937","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.