Predicting factors associated with under-5 mortality in India using machine learning algorithms: evidence from National Family Health Survey, 2019-21 | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Predicting factors associated with under-5 mortality in India using machine learning algorithms: evidence from National Family Health Survey, 2019-21 Abhay Mishra, Guru Vasishtha, Suraj Maiti This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5309131/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: Reducing the under-5 mortality rate is high on the list of priorities in the global development agenda. The SDG targets a reduction in child mortality rate to less than 25 deaths per 1,000 live births annually by 2030. Though enormous gains have been observed for under-five mortality and child health over the past few years, it remains a significant public health challenge for India. Much earlier studies were made on the under-five mortality rates for children. Except in this study, the most widely used approach undertaken so far, to a great extent, relies on conventional regression analyses that are inherently known to have minimal predictive ability. This study attempts to develop a predictive model based on the advanced techniques of machine learning (AML) that could infer a rather accurate prediction for the under-5 mortality rate of India, namely U5MR. Methods: This study used the nationally representative microdata from 7 the National Family Health Survey's fifth version (NFHS-5, 2019-21). Multiple imputation methods, such as filling by modal value, were adopted to treat missing values. We used a feature selection method known as information gain, where we ranked the information-rich features and examined their impact on the prediction of child mortality. The synthetic minority over-sampling method (SMOTE) was used to balance the dataset. To predict the determinants affecting U5MR, we used four machine learning (ML) models (decision trees, logistic regression, support vector machines, and K-nearest neighbors). The predictive power of each ML model was assessed using accuracy, precision, receiver operating characteristic curves, and model accuracy metrics (accuracy, precision, F1 score, ROC). Results: The descriptive findings demonstrate that India's under-five mortality rates vary significantly by region. The Decision Tree model (96.35%) performed the best out of all the models examined, with the under-five mortality prediction ability ranging from 90% to 96.35%. The best predictive model demonstrates that factors influencing under-five mortality rates in India include duration of breastfeeding status, Marriage to the first birth Interval, education, Birth order, ANC visit, wealth index, place of delivery, and place of residence to estimate under-five mortality risk variables, The Decision Tree machine learning model would engender a greater predictive capability; therefore, it would help in better policy decision-making in this field. Employment of these major elements would greatly enhance a child's chances of survival to be directed by appropriate policies. Conclusion : The decision tree model will perform better than traditional logistic regression. Prediction of under-5 mortality rate at the India level revealed a success accuracy of 96.35 percent and precision of about 60 percent. Machine learning Under-five mortality SVM KNN DT LR Determinants India Figures Figure 1 Introduction Child mortality is one of the most burning problems plaguing underdeveloped and developing countries. It is a cause of concern for a nation's economic and future health conditions when the child mortality rate is high. In response to this, the United Nations established the Sustainable Development Goals (SDGs) in 2015 and called for a global reduction of less than 25 deaths per 1,000 live births in the child mortality rate. Annually by 2030 (United Nations Development Program, 2015; World Health Organization, 2015). Meanwhile, child mortality and births in many developing nations continue to be high, inhibiting the growth of their economies and highlighting the failures of many national health systems. It then represents one of the important steps in a nation's progress toward progress. However, most of the traditional forecasting methods lack the rigor and benefits of using machine learning (ML) algorithms. The potential for advanced analytics in public health has been showcased in previous literature investigating machine learning algorithms for forecasting under-five mortality in Nigeria [2]. This study used machine learning to predict mortality in children under five years. It proved that predictive analytics could be deployed effectively in healthcare to explain why data-driven approaches were important to improve the outcomes for children's health [3]. To lower the rate of child mortality in the area, looked into the use of machine learning algorithms to forecast the death of children under five in Ethiopia [4]. In investigated the application of machine learning methods to anticipate Ethiopia's under-five mortality rate, highlighting the significance of predictive modeling in guiding focused medical interventions for children [8]. Their 2019 study examined how well machine learning could predict the death rate of children under five in sub-Saharan Africa, showing the promise of data-driven strategies for enhancing child health outcomes in environments with limited resources [7]. Such factors include low birth weights, complications with pregnancies, nutritional deficiencies for pregnant mothers, pneumonia, diarrheal diseases, and lack of access to healthcare facilities, which can be used to explain the child mortality rate [5]. The high death rate in India is partly due to the country's son preference, which views reproduction as the duty of women and discourages investment in maternal care. Many expectant mothers worry about losing their jobs or facing financial instability, which forces poor women to work until the very end of their pregnancy to support their families [6]. This may impact both the health of the mother and the unborn child. In addition, effective preventive and control measures for childhood diseases, and improved healthcare services, which include proper immunization programs for all children and provision of nutrient supplementation, are well known to be fundamental components in reducing global childhood mortality. In developing nations like India, the primary causes of mortality among children under five are a cause for concern. The mortality rate for children under five has been the subject of much research. Thus far, the most common method used mostly depends on conventional logistic regression analysis, which has been shown to have poor predictive power. Accurate under-5 mortality prediction is possible using advanced machine learning (AML) techniques. In this context, we utilize machine learning techniques, and the mortality rate of children under five in India may be predicted. Materials And Methods Data We used data from the National Family Health Survey (NFHS-V), 2019-21, a large-scale, multi-round cross-sectional, nationally representative survey. The survey collects extensive population, health, and nutrition information, emphasizing women and young children. We have used the target group data of children under the age of five at the national level. MoHFW has designated the International Institute for Population Sciences, Mumbai as the nodal agency for all rounds of NFHS. NFHS-5 fieldwork for India was carried out in two rounds— Phase-I between June 17th, 2019, to January 30th, 2020, covering 17 states and 5 UTs, and Phase-II between January 2nd, 2020 and April 30th, 2021, covering 11 states and 3 UTs — surveying 636,699 households and interviewing 724,115 women and 101,839 men. For the data analysis of this study, a kids file (IAKR7EDT) has been used. Target variable The main focus variable in the study is under-5 mortality, which represents the death of children before they reach the age of five. The survey asked a childbearing mother, "Is your child alive?" The answer from that mother was recorded as no = 0 and yes = 1. We utilized this to create a binary variable recorded as 1 if the child's death occurred under the age of 5; otherwise, it was recorded as 0. Independent variables: We used a set of predictor socio-economic and demographic status variables based on previous literature that affects a child's health Religion (Hindu, Muslim, Christian, and others), source of drinking water (Clean source of drinking, unclean source of drinking), toilet facility (improved sanitation facility, unimproved sanitation facility), fuel type, floor material, place of delivery, duration of breastfeeding, birth order, the interval between marriage and first birth, mother's occupation, head of household occupation, region of residence, place of residence, education level, wealth index, sex of the child, and ANC visit are all independent variables (Naznin et al., 2023). All these factors can be independent and may result in different outcomes or phenomena under investigation. Data Preprocessing: The data needs to be pre-processed by lots of techniques. In this stage, duplicate variables were removed using predictive mean matching, and variables had missing values that were filled by the model values. Thereafter, all categories and string variables were then converted to numerical values. A crucial aspect of data pre-processing is ensuring that the target or result variable is balanced. The dataset showed a significant skewness in the numbers of under-five deaths relative to living children (8702 under-five deaths against 224218 live children). After balancing the target (dependent) using a SMOTE oversampling technique, a ratio of 74.07:25.93 was attained as opposed to the initial ratio of 96.26:3.74. Feature Selection: Then, we used Recursive Feature Elimination to identify the important features of under-5 mortality and find 10 significant variables. The feature selection process uses the recursive elimination strategy, and the results of the RFE method indicate that several independent factors are important in predicting child death. The top 10 variables—ANC visit, wealth index, education, place of residence, marriage to the first birth interval, birth order, child's sex, fuel, floor material, and place of delivery—among the 16 primary variables are determined by Recursive Feature elimination (RFE) to be significant characteristics. Model building In this step, missing values are handled by applying machine learning models. Next, we divided our whole data set into two groups: 80% trained and 20% tested. The trained dataset is used to train the models, while the test dataset is used to assess the models' performance. Finally, we use all of the data to predict the performed model. To create more accurate prediction models, one-hot encoding required changes to every independent feature. The dependent variable in this study was binary—that is, either living or dead. Then, we employed a variety of appropriate machine learning models, including support vector machines, decision trees, logistic regression, and K-nearest neighbor models. A detailed methodology has been given in Supplementary Material (SM) Text S1 . Results Table 1 summarizes under-five mortality rates across various sociodemographic factors in India. Religious affiliation shows varying mortality rates, with Hindu families experiencing the highest rate at 3.94%, followed by Muslims (3.36%), Sikhs (3.27%), Christians (2.96%), and others (2.76%). Environmental factors play a crucial role in child mortality. Households using unclean drinking water sources and unimproved sanitation facilities both show higher mortality rates (4.71%) compared to those with clean water sources (3.75%) and improved sanitation (3.35%). Using unclean fuel is associated with higher child mortality (4.38%) than clean fuel (2.88%). Home deliveries have a higher mortality rate (5.30%) than hospital deliveries (3.48%). However, deliveries in other locations have the highest mortality rate at 7.93%. Children who were never breastfed or in the others category have an alarmingly high mortality rate of 20.88%, while those who breastfed for 2-4 years and 4+ years have very low rates (0.51% and 0.01%, respectively). Higher education levels are associated with lower mortality rates, decreasing from 5.24% for no education to 2.11% for secondary and higher education. Similarly, wealthier households have lower mortality rates (2.50%) than poor households (4.63%). Rural areas show a higher mortality rate (3.98%) than urban areas (2.79%), indicating disparities in access to healthcare and other resources. Table 2 presents the results of a multivariable logistic regression analysis, providing adjusted odds ratios (AOR) for various factors influencing under-five mortality. Religious affiliation shows significant effects, with Muslims having lower odds of child mortality compared to Hindus (AOR: 0.583, p<0.001). Environmental factors continue to show importance, with unclean fuel use associated with higher odds of child mortality (AOR: 1.211, p=0.045). Place of delivery remains crucial, with hospital deliveries showing lower mortality odds than home deliveries (AOR: 0.795, p=0.023). Breastfeeding duration is a strong predictor of child survival. Compared to breastfeeding for 1-2 years, longer durations are associated with significantly lower odds of mortality (2-4 years: AOR: 0.080, p<0.001; 4+ years: AOR: 0.001, p<0.001). Conversely, never breastfeeding or for the others category is associated with much higher odds of mortality (AOR: 3.693, p<0.001). Higher birth orders are associated with increased odds of mortality (Three and above: AOR: 3.441, p<0.001). Regional differences are evident, with the Central region showing higher odds of child mortality than the Northern region (AOR: 1.493, p=0.003), while the Western and Southern regions show lower odds. Education levels continue to show a protective effect, with higher education associated with lower odds of child mortality. Secondary and higher education shows the strongest effect (AOR: 0.647, p=0.005). Middle (AOR: 0.793, p=0.048) and rich (AOR: 0.670, p=0.004) households showed lower child mortality odds than poor households. Gender disparities exist, with female children having lower odds of mortality (AOR: 0.811, p=0.004). Antenatal care (ANC) visits show a protective effect, with 4-6 visits associated with significantly lower odds of child mortality (AOR: 0.684, p=0.008). Table 3 compares the performance of four classification models: Logistic Tree, Decision Tree (DT), K-Nearest Neighbors (K-NN), and Support Vector Machine (SVM). These models were assessed using test data, and their performance was evaluated based on several metrics: Confusion Matrix, Accuracy, Recall, Precision, F1 Score, and Area Under the Receiver Operating Characteristic curve (AUROC). The Confusion Matrix for each model provides a detailed breakdown of correct and incorrect predictions. For the Logistic Tree model, out of 44,844 alive cases, 41,467 were correctly predicted as alive (true negatives), while 3,377 were incorrectly predicted as dead (false positives). Of the 1,740 actual death cases, 481 were correctly predicted as dead (true positives), while 1,259 were incorrectly predicted as alive (false negatives). This distribution indicates that the Logistic Tree model overpredicts deaths, resulting in many false positives. The Decision Tree (DT) model correctly predicted 44,809 alive cases and 931 death cases, with only 35 false positives and 809 false negatives. This distribution suggests that the DT model is more balanced in its predictions, with fewer errors in both directions than the Logistic Tree model. The K-Nearest Neighbors (K-NN) model correctly identified 44,489 alive cases and 112 death cases, with 355 false positives and 1,628 false negatives. This pattern indicates that the K-NN model tends to underpredict deaths, resulting in many false negatives. The Support Vector Machine (SVM) model predicted all cases as alive, resulting in 44,809 true negatives and 1,775 false negatives. Suggests that the SVM model failed to capture the nuances of the factors leading to child mortality in this dataset. The Decision Tree model demonstrates the highest accuracy at 96.35%, followed closely by the SVM at 96.21%. However, accuracy alone can be misleading, especially in datasets with imbalanced classes, as is often the case with mortality data, where death cases are typically much fewer than survival cases. The recall metric measures the model's ability to correctly identify positive cases (deaths in this context). The DT model shows the highest recall at 30.00%, followed by the Logistic Tree at 28.00%. Thus, the DT model correctly identified 30% of all death cases in the dataset. The K-NN model's recall is significantly lower at 6.00%, while the SVM model's recall is 0% due to its failure to predict deaths. The DT model has the highest precision, which indicates the accuracy of positive predictions at 60.00%. This means that when the DT model predicts a death, it is correct 60% of the time. The K-NN model follows with 23.00% precision, followed by the Logistic Tree with only 12.00%. Again, the SVM model's precision is 0% due to its lack of positive predictions. The F1 Score provides a balanced measure of precision and recall, further confirming the DT model's superior performance. It achieves the highest F1 Score of 26.00%, followed by K-NN at 22.00% and Logistic Tree at 17.00%. The SVM model's F1 Score is 0% due to its lack of true positive predictions. The Area Under the Receiver Operating Characteristic curve (AUROC) provides an aggregate performance measure across all possible classification thresholds. The DT model again leads with an AUROC of 65.00%, followed by the Logistic Tree at 60.05%. The SVM and K-NN models show lower AUROC values of 52.70% and 51.00%, respectively, indicating poor discriminative ability. Discussion This study uses a logistic regression analysis and a machine learning algorithm to identify the critical factors linked to mortality among children under five and assess the significance of machine learning approaches in predicting the causes of mortality among children under five. This is the first study to forecast mortality for children under five using machine learning algorithms using national mortality data. We used the 80/20 ratio to improve the accuracy of machine learning models. Prior research [23, 24] supports this outcome. This study detailed the use of conventional and machine learning techniques to predict under-5 mortality in India. Using the recursive feature elimination technique, our study shows that the only significant characteristics were birth order, sex of the child, place of delivery, length of breastfeeding, wealth index, education, regions, marriage to the first interval, and kind of fuel. According to our research, breastfeeding duration has a significant impact on under-5 mortality in India, which is in line with research on under-5 mortality in Ethiopia [8]. We evaluated several machine learning models (LR, DT, SVM, and KNN) to forecast mortality for children under five years old in India by utilizing performance measures, including AUC and confusion matrix. The results indicate that machine-learning techniques outperform logistic regression models when it comes to identifying characteristics linked to mortality for children under five [11–12]. When forecasting under-5 mortality in India, the DT model did better. To predict under-5 mortality in India, the DT model considered the individual and interaction effects of each of the chosen parameters. Because the requirements of the logistic regression model are not met, it is unable to predict under-5 mortality with accuracy. In particular, multicollinearity affects substantial predictor independence [11]. It is advised to use only independent variables in Logistic Regression modeling to overcome multicollinearity difficulties and prevent inaccurate findings. On the other hand, the decision tree model is a more dependable option for estimating under-5 mortality in India due to its better performance and absence of particular assumptions. Support vector machines (SVM) and K-Nearest Neighbors (KNN), two machine learning models that are flexible due to their lower assumption counts, are included in this study. Most preterm newborns' death risk may be accurately predicted using this method, which can also mimic and forecast mortality rates in the human population. The superiority of machine learning model methods over conventional analytic methods has also been demonstrated by earlier studies. According to a recent study, machine learning models are more appropriate for identifying factors contributing to infant mortality and verifying superior goodness of fit in most key groups. Furthermore, machine learning models are extremely helpful in forecasting health studies, resulting in better and more appropriate policy choices. We used the most recent nationally representative NFHS-V 2019–21 dataset. However, because the study is cross-sectional, proving causal inference is not possible. The results are based on self-reported information subject to social desirability and memory biases. The regression model may not account for socioeconomic factors contributing to under-5 mortality. Our paper thoroughly analyzes several machine learning models—DT, SVM, KNN, and LR—to forecast under-5 mortality in India. This study aimed to apply several machine learning algorithms to data on under-5 mortality. The ML model for predicting death in children under five is somewhat more accurate than classical logistic regression. Among the machine learning models used in this study, the decision tree model had the highest accuracy in predicting the death of children under five. The study also suggests that, with certain restrictions, logistic regression analysis may help forecast the mortality of morality in children under five. However, this study also showed that some of the variables have an equally important effect on mortality for children under five in both LR and ML models. Machine learning models provide analytical prowess that may enhance analysis capabilities compared to traditional statistical models. Such models could be used to process large-dimensional data in health research. Declarations Ethics approval: Publicly available data was used and does not require ethics approval Funding: No funding received from any sources Authorship contributions: GV, and AM conceptualized the manuscript; AM analyzed the data; AM, GV, and SM developed the initial draft of the manuscript. All authors read, reviewed, and contributed equally to the final draft of the manuscript. Disclosure of interest: No Conflict of Interest References Brahma, Dweepobotee and Debasri Mukherjee (2022), “Early Warning Signs: Targeting Neonatal and Infant Mortality Using Machine Learning”, Applied Economics, 54(1): 57-74. Okagbue, H. I., Oguntunde, P. E., Obasi, E. C., Adamu, P. I., & Opanuga, A. A. (2021). Diagnosing malaria from some symptoms: a machine learning approach and public health implications. Health and Technology , 11 , 23-37. Singh, A., Kumar, K., & Singh, A. (2019). What explains the decline in neonatal mortality in India in the last three decades? Evidence from three rounds of NFHS surveys. Studies in family planning, 50(4), 337- 355. https://doi.org/10.1111/sifp.12105 Gebrehiwot, G. G., & Worku, A. G. (2021). Forecasting under-five child mortality in Ethiopia using machine learning algorithms. BMC Public Health , 21 (1), 1024. https://doi.org/10.1186/s12889-021-10968-0. United Nations (UN) United Nations Population Division; 2017. World population prospects (2017) https://www.unicef.org/sitan/files/SitAn_India_May_2011.pdf Claeson, M., Bos, E. R., Mawji, T. & Pathmanathan, I. (2000). Reducing child mortality in India in the new millennium. Bulletin of the World Health Organization, 78 (10), 1192 -1199. World Health Organization. Bell RT. 2011. A baseline study exploring rural maternal health practices and services about the IGMSY conditional maternity benefit. MSc Thesis. Utrecht University. Zhang, H., & Li, X. (2019). Predicting under-five child mortality rates in sub-Saharan Africa using machine learning. Journal of Global Health , 9 (2), 020401. https://doi.org/10.7189/jogh.09.020401. Tadesse, A. W., & Awoke, T. Z. (2020). Application of machine learning methods to anticipate under-five mortality rates in Ethiopia. International Journal of Health Planning and Management , 35 (4), 1035-1048. https://doi.org/10.1002/hpm.2995. Vallin, J. (1976). World trends in infant mortality since 1950. World health statistics report. 29(11), 646-674 Wang, H., Liddell, C. A., Coates, M. M., ... & Moore, A. R. (2014). Global, regional, and national levels of neonatal, infant, and under-5 mortality during 1990–2013: a systematic analysis for the Global Burden of Disease Study 2013. The Lancet, 384(9947), 957-979. https://doi.org/10.1016/S0140-6736(14)60497-9 Bitew FH, Nyarko SH, Potter L, Sparks CS. Machine learning approach for predicting under-five mortality determinants in Ethiopia: evidence from the 2016 Ethiopian Demographic and Health Survey. Genus. 2020;76(1). 10.1186/s41118-020-00106-2. Vapnik, V. 1995. The Nature of Statistical Learning Theory. New York, NY: Springer. Attewell, P., D.B. Monaghan, and D. Kwong. 2015. Data Mining for the Social Sciences: An Introduction. Oakland, CA: University of California Press. Rahman A, Hossain Z, Kabir E, Rois R. “Machine Learning Algorithm for Analysing Infant Mortality in Bangladesh,” Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vol. 13079 LNCS. pp. 205–219, 2021, 10.1007/978-3-030-90885-0_19. Mfateneza E, Rutayisire PC, Biracyaza E, Musafiri S, Mpabuka WG. Application of machine learning methods for predicting infant mortality in Rwanda: analysis of Rwanda demographic health survey 2014–15 dataset. BMC Pregnancy Childbirth. 2022;22(1):1–13. 10.1186/s12884-022-04699-8. Byun, H., and S.W. Lee. 2002. “Applications of Support Vector Machines for Pattern Recognition: A Survey.” In SVM ’02 Proceedings of the First International Workshop on Pattern Recognition with Support Vector Machines, 213–36. London, UK: Springer. Murphy KP. Machine Learning: A Probabilistic Perspective. The MIT Press; 2012. James G, Witten D, Hastie T, Tibshirani R. An introduction to statistical learning. Vol. 112. Springer; 2013. Andrews DW, Whang YJ. Additive interactive regression models: circumvention of the curse of dimensionality. Econometric Theory. 1990; 6(4):466–479. https://doi.org/10.1017/ S0266466600005478 Breiman L, Friedman J, Stone CJ, Olshen RA. Classification and regression trees. CRC Press; 1984 Hothorn T, Hornik K, Zeileis A. Unbiased recursive partitioning: A conditional inference framework. Journal of Computational and Graphical Statistics. 2006; 15(3):651–674. https://doi.org/10.1198/ 106186006X133933 Breiman, L., Friedman, J., Stone, C. J., and Olshen, R. A. (1984). Classification and Regression Trees. Boca Raton, FL: CRC press. Caruana, R., and Niculescu-Mizil, A. (2006). “An empirical comparison of supervised learning algorithms,” in 23rd International Conference on Machine Learning, (Pittsburgh, PA: ACM Press), 161–168. C. Junli, and J. Licheng, "Classification mechanism of support vector machines," 5th International Conference on Signal Processing Proceedings. 16th World Computer Congress, Beijing, 2000, vol.3, pp. 1556-1559, 2000. Goldstein BA, Navar AM, Carter RE. Moving beyond regression techniques in cardiovascular risk prediction: applying machine learning to address analytic challenges. Eur Heart J. 2017;38(23):1805–14. 30. Kotsiantis SB, Zaharakis I, Pintelas P. Supervised machine learning: A review of classification techniques. Emerging artificial intelligence applications in computer engineering. 2007;160(1):3–24 Hastie T, Tibshirani R, Friedman JH (2009) The elements of statistical learning: data mining, inference, and prediction, 2nd ed. Springer, New York Return to ref 2009 in article Boser, B.E., I.M. Guyon, and V.N. Vapnik. 1992. “A Training Algorithm for Optimal Margin Classifiers.” In Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, edited by D. Haussler, 144-152. New York, NY: ACM Press. Tables Tables 1 to 3 are available in the Supplementary Files section. Additional Declarations No competing interests reported. Supplementary Files Tables.docx Supplementarymaterial.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5309131","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":368897864,"identity":"36dcd907-59c0-4490-8e3a-76af1aa4bf08","order_by":0,"name":"Abhay Mishra","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABJElEQVRIie2RMUvDQBTHLwp1SbhNngSSr/DCQawQ6Fe5IHTKILh0kk5OVtdK8RO4RAIBt0jWRNeTdgh0LhQKEhDFu+JSSKKj4P2Ggz+8H//3OEI0mr8ImNuXICFGtR6BQ1XmZ79T9r1pEbCjsVLwR4UopWdbl8Mwzr5jG3Q2edpYUX9wbOdLNHs5Yy/34bpC4tLDrLlk8XxqWymEj9dDrwIzd3yxSkAu5t3OeHONiFApHAvCECFnvihjpXCcNyuuiNi7VAZYHLzJyTxMpmVSdykoIl+1GHFh+pBxeT6dpJ0t3qLwT+7kLVI598ZZwEBYaZ8jtN7izK+YWKUXarGH5cen/MqbMnmtR4FL7ZbzJXvmToTtJLSOK4x6J9Ksc1qj0Wj+H1+yvWPEV9lMawAAAABJRU5ErkJggg==","orcid":"","institution":"International Institute for Population Sciences","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Abhay","middleName":"","lastName":"Mishra","suffix":""},{"id":368897865,"identity":"dde91b7f-6a89-4f2f-ab57-c0a49b5f1aff","order_by":1,"name":"Guru Vasishtha","email":"","orcid":"","institution":"International Institute for Population Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Guru","middleName":"","lastName":"Vasishtha","suffix":""},{"id":368897866,"identity":"df5b9085-7c35-47e4-bdbb-a923c611712e","order_by":2,"name":"Suraj Maiti","email":"","orcid":"","institution":"International Institute for Population Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Suraj","middleName":"","lastName":"Maiti","suffix":""}],"badges":[],"createdAt":"2024-10-22 06:53:41","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5309131/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5309131/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":67270117,"identity":"c3d84daf-cf1e-4dea-af65-c3cbef1350a5","added_by":"auto","created_at":"2024-10-23 07:24:00","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":51776,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOverview of the proposed framework of machine learning for under-five child mortality data\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-5309131/v1/b5acd7b91095220a6d8e564a.png"},{"id":82385364,"identity":"e400cee5-9206-47b3-9593-cae1929ed8b4","added_by":"auto","created_at":"2025-05-09 16:23:38","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":512607,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5309131/v1/4c69b0a9-0068-47b8-abc3-2a66a18af6a1.pdf"},{"id":67270371,"identity":"e6ba8b9c-0b9c-4e0d-b889-8288f954bc35","added_by":"auto","created_at":"2024-10-23 07:32:00","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":35924,"visible":true,"origin":"","legend":"","description":"","filename":"Tables.docx","url":"https://assets-eu.researchsquare.com/files/rs-5309131/v1/fd09625a8a6535010ca93426.docx"},{"id":67270119,"identity":"d3d6f97c-48a0-4ef6-b0b8-a34bfc5609d2","added_by":"auto","created_at":"2024-10-23 07:24:00","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":23610,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementarymaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-5309131/v1/ed7ab9890ee56f815a601e34.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Predicting factors associated with under-5 mortality in India using machine learning algorithms: evidence from National Family Health Survey, 2019-21","fulltext":[{"header":"Introduction","content":"\u003cp\u003eChild mortality is one of the most burning problems plaguing underdeveloped and developing countries. It is a cause of concern for a nation\u0026apos;s economic and future health conditions when the child mortality rate is high. In response to this, the United Nations established the Sustainable Development Goals (SDGs) in 2015 and called for a global reduction of less than 25 deaths per 1,000 live births in the child mortality rate. Annually by 2030 (United Nations Development Program, 2015; World Health Organization, 2015). Meanwhile, child mortality and births in many developing nations continue to be high, inhibiting the growth of their economies and highlighting the failures of many national health systems.\u003c/p\u003e\n\u003cp\u003eIt then represents one of the important steps in a nation\u0026apos;s progress toward progress. However, most of the traditional forecasting methods lack the rigor and benefits of using machine learning (ML) algorithms. The potential for advanced analytics in public health has been showcased in previous literature investigating machine learning algorithms for forecasting under-five mortality in Nigeria [2]. This study used machine learning to predict mortality in children under five years. It proved that predictive analytics could be deployed effectively in healthcare to explain why data-driven approaches were important to improve the outcomes for children\u0026apos;s health [3]. To lower the rate of child mortality in the area, looked into the use of machine learning algorithms to forecast the death of children under five in Ethiopia [4]. In investigated the application of machine learning methods to anticipate Ethiopia\u0026apos;s under-five mortality rate, highlighting the significance of predictive modeling in guiding focused medical interventions for children [8]. Their 2019 study examined how well machine learning could predict the death rate of children under five in sub-Saharan Africa, showing the promise of data-driven strategies for enhancing child health outcomes in environments with limited resources [7]. Such factors include low birth weights, complications with pregnancies, nutritional deficiencies for pregnant mothers, pneumonia, diarrheal diseases, and lack of access to healthcare facilities, which can be used to explain the child mortality rate [5].\u003c/p\u003e\n\u003cp\u003eThe high death rate in India is partly due to the country\u0026apos;s son preference, which views reproduction as the duty of women and discourages investment in maternal care. Many expectant mothers worry about losing their jobs or facing financial instability, which forces poor women to work until the very end of their pregnancy to support their families [6]. This may impact both the health of the mother and the unborn child. In addition, effective preventive and control measures for childhood diseases, and improved healthcare services, which include proper immunization programs for all children and provision of nutrient supplementation, are well known to be fundamental components in reducing global childhood mortality.\u003c/p\u003e\n\u003cp\u003eIn developing nations like India, the primary causes of mortality among children under five are a cause for concern. The mortality rate for children under five has been the subject of much research. Thus far, the most common method used mostly depends on conventional logistic regression analysis, which has been shown to have poor predictive power. Accurate under-5 mortality prediction is possible using advanced machine learning (AML) techniques. In this context, we utilize machine learning techniques, and the mortality rate of children under five in India may be predicted.\u003c/p\u003e"},{"header":"Materials And Methods","content":"\u003cp\u003e\u003cstrong\u003eData\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe used data from the National Family Health Survey (NFHS-V), 2019-21, a large-scale, multi-round cross-sectional, nationally representative survey. The survey collects extensive population, health, and nutrition information, emphasizing women and young children. We have used the target group data of children under the age of five at the national level. MoHFW has designated the International Institute for Population Sciences, Mumbai as the nodal agency for all rounds of NFHS. NFHS-5 fieldwork for India was carried out in two rounds\u0026mdash; Phase-I between June 17th, 2019, to January 30th, 2020, covering 17 states and 5 UTs, and Phase-II between January 2nd, 2020 and April 30th, 2021, covering 11 states and 3 UTs \u0026mdash; surveying 636,699 households and interviewing 724,115 women and 101,839 men. For the data analysis of this study, a kids file (IAKR7EDT) has been used.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTarget variable\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe main focus variable in the study is under-5 mortality, which represents the death of children before they reach the age of five. The survey asked a childbearing mother, \u0026quot;Is your child alive?\u0026quot; The answer from that mother was recorded as no = 0 and yes = 1. We utilized this to create a binary variable recorded as 1 if the child\u0026apos;s death occurred under the age of 5; otherwise, it was recorded as 0.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eIndependent variables:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe used a set of predictor socio-economic and demographic status variables based on previous literature that affects a child\u0026apos;s health Religion (Hindu, Muslim, Christian, and others), source of drinking water (Clean source of drinking, unclean source of drinking), toilet facility (improved sanitation facility, unimproved sanitation facility), fuel type, floor material, place of delivery, duration of breastfeeding, birth order, the interval between marriage and first birth, mother\u0026apos;s occupation, head of household occupation, region of residence, place of residence, education level, wealth index, sex of the child, and ANC visit are all independent variables (Naznin et al., 2023). All these factors can be independent and may result in different outcomes or phenomena under investigation.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Preprocessing:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data needs to be pre-processed by lots of techniques. In this stage, duplicate variables were removed using predictive mean matching, and variables had missing values that were filled by the model values. Thereafter, all categories and string variables were then converted to numerical values. A crucial aspect of data pre-processing is ensuring that the target or result variable is balanced. The dataset showed a significant skewness in the numbers of under-five deaths relative to living children (8702 under-five deaths against 224218 live children). After balancing the target (dependent) using a SMOTE oversampling technique, a ratio of 74.07:25.93 was attained as opposed to the initial ratio of 96.26:3.74.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFeature Selection:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThen, we used Recursive Feature Elimination to identify the important features of under-5 mortality and find 10 significant variables. The feature selection process uses the recursive elimination strategy, and the results of the RFE method indicate that several independent factors are important in predicting child death. The top 10 variables\u0026mdash;ANC visit, wealth index, education, place of residence, marriage to the first birth interval, birth order, child\u0026apos;s sex, fuel, floor material, and place of delivery\u0026mdash;among the 16 primary variables are determined by Recursive Feature elimination (RFE) to be significant characteristics.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModel building\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn this step, missing values are handled by applying machine learning models. Next, we divided our whole data set into two groups: 80% trained and 20% tested. The trained dataset is used to train the models, while the test dataset is used to assess the models\u0026apos; performance. Finally, we use all of the data to predict the performed model. To create more accurate prediction models, one-hot encoding required changes to every independent feature. The dependent variable in this study was binary\u0026mdash;that is, either living or dead. Then, we employed a variety of appropriate machine learning models, including support vector machines, decision trees, logistic regression, and K-nearest neighbor models. A detailed methodology has been given in \u003cstrong\u003eSupplementary Material (SM) Text S1\u003c/strong\u003e.\u0026nbsp;\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eTable 1\u0026nbsp;\u003c/strong\u003esummarizes under-five mortality rates across various sociodemographic factors in India. Religious affiliation shows varying mortality rates, with Hindu families experiencing the highest rate at 3.94%, followed by Muslims (3.36%), Sikhs (3.27%), Christians (2.96%), and others (2.76%). \u0026nbsp;Environmental factors play a crucial role in child mortality. Households using unclean drinking water sources and unimproved sanitation facilities both show higher mortality rates (4.71%) compared to those with clean water sources (3.75%) and improved sanitation (3.35%). Using unclean fuel is associated with higher child mortality (4.38%) than clean fuel (2.88%). Home deliveries have a higher mortality rate (5.30%) than hospital deliveries (3.48%). However, deliveries in other locations have the highest mortality rate at 7.93%. Children who were never breastfed or in the others category have an alarmingly high mortality rate of 20.88%, while those who breastfed for 2-4 years and 4+ years have very low rates (0.51% and 0.01%, respectively). Higher education levels are associated with lower mortality rates, decreasing from 5.24% for no education to 2.11% for secondary and higher education. Similarly, wealthier households have lower mortality rates (2.50%) than poor households (4.63%). Rural areas show a higher mortality rate (3.98%) than urban areas (2.79%), indicating disparities in access to healthcare and other resources.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2\u0026nbsp;\u003c/strong\u003epresents the results of a multivariable logistic regression analysis, providing adjusted odds ratios (AOR) for various factors influencing under-five mortality. Religious affiliation shows significant effects, with Muslims having lower odds of child mortality compared to Hindus (AOR: 0.583, p\u0026lt;0.001). Environmental factors continue to show importance, with unclean fuel use associated with higher odds of child mortality (AOR: 1.211, p=0.045). Place of delivery remains crucial, with hospital deliveries showing lower mortality odds than home deliveries (AOR: 0.795, p=0.023). Breastfeeding duration is a strong predictor of child survival. Compared to breastfeeding for 1-2 years, longer durations are associated with significantly lower odds of mortality (2-4 years: AOR: 0.080, p\u0026lt;0.001; 4+ years: AOR: 0.001, p\u0026lt;0.001). Conversely, never breastfeeding or for the others category is associated with much higher odds of mortality (AOR: 3.693, p\u0026lt;0.001). Higher birth orders are associated with increased odds of mortality (Three and above: AOR: 3.441, p\u0026lt;0.001). Regional differences are evident, with the Central region showing higher odds of child mortality than the Northern region (AOR: 1.493, p=0.003), while the Western and Southern regions show lower odds. Education levels continue to show a protective effect, with higher education associated with lower odds of child mortality. Secondary and higher education shows the strongest effect (AOR: 0.647, p=0.005). Middle (AOR: 0.793, p=0.048) and rich (AOR: 0.670, p=0.004) households showed lower child mortality odds than poor households. Gender disparities exist, with female children having lower odds of mortality (AOR: 0.811, p=0.004). Antenatal care (ANC) visits show a protective effect, with 4-6 visits associated with significantly lower odds of child mortality (AOR: 0.684, p=0.008).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3\u003c/strong\u003e compares the performance of four classification models: Logistic Tree, Decision Tree (DT), K-Nearest Neighbors (K-NN), and Support Vector Machine (SVM). These models were assessed using test data, and their performance was evaluated based on several metrics: Confusion Matrix, Accuracy, Recall, Precision, F1 Score, and Area Under the Receiver Operating Characteristic curve (AUROC). The Confusion Matrix for each model provides a detailed breakdown of correct and incorrect predictions. For the Logistic Tree model, out of 44,844 alive cases, 41,467 were correctly predicted as alive (true negatives), while 3,377 were incorrectly predicted as dead (false positives). Of the 1,740 actual death cases, 481 were correctly predicted as dead (true positives), while 1,259 were incorrectly predicted as alive (false negatives). This distribution indicates that the Logistic Tree model overpredicts deaths, resulting in many false positives. The Decision Tree (DT) model correctly predicted 44,809 alive cases and 931 death cases, with only 35 false positives and 809 false negatives. This distribution suggests that the DT model is more balanced in its predictions, with fewer errors in both directions than the Logistic Tree model. The K-Nearest Neighbors (K-NN) model correctly identified 44,489 alive cases and 112 death cases, with 355 false positives and 1,628 false negatives. This pattern indicates that the K-NN model tends to underpredict deaths, resulting in many false negatives. The Support Vector Machine (SVM) model predicted all cases as alive, resulting in 44,809 true negatives and 1,775 false negatives. Suggests that the SVM model failed to capture the nuances of the factors leading to child mortality in this dataset.\u003c/p\u003e\n\u003cp\u003eThe Decision Tree model demonstrates the highest accuracy at 96.35%, followed closely by the SVM at 96.21%. However, accuracy alone can be misleading, especially in datasets with imbalanced classes, as is often the case with mortality data, where death cases are typically much fewer than survival cases. The recall metric\u0026nbsp;measures the model\u0026apos;s ability to correctly identify positive cases (deaths in this context). The DT model shows the highest recall at 30.00%, followed by the Logistic Tree at 28.00%. Thus, the DT model correctly identified 30% of all death cases in the dataset. The K-NN model\u0026apos;s recall is significantly lower at 6.00%, while the SVM model\u0026apos;s recall is 0% due to its failure to predict deaths.\u003c/p\u003e\n\u003cp\u003eThe DT model has the highest precision, which indicates the accuracy of positive predictions at 60.00%. This means that when the DT model predicts a death, it is correct 60% of the time. The K-NN model follows with 23.00% precision, followed by the Logistic Tree with only 12.00%. Again, the SVM model\u0026apos;s precision is 0% due to its lack of positive predictions.\u003c/p\u003e\n\u003cp\u003eThe F1 Score provides a balanced measure of precision and recall, further confirming the DT model\u0026apos;s superior performance. It achieves the highest F1 Score of 26.00%, followed by K-NN at 22.00% and Logistic Tree at 17.00%. The SVM model\u0026apos;s F1 Score is 0% due to its lack of true positive predictions. The Area Under the Receiver Operating Characteristic curve (AUROC) provides an aggregate performance measure across all possible classification thresholds. The DT model again leads with an AUROC of 65.00%, followed by the Logistic Tree at 60.05%. The SVM and K-NN models show lower AUROC values of 52.70% and 51.00%, respectively, indicating poor discriminative ability.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study uses a logistic regression analysis and a machine learning algorithm to identify the critical factors linked to mortality among children under five\u0026nbsp;and assess the significance of machine learning approaches in predicting the causes of mortality among children under five. This is the first study to forecast mortality for children under five using machine learning algorithms using national mortality data. We used the 80/20 ratio to improve the accuracy of machine learning models. Prior research [23, 24] supports this outcome.\u003c/p\u003e\n\u003cp\u003eThis study detailed the use of conventional and machine learning techniques to predict under-5 mortality in India. Using the recursive feature elimination technique, our study shows that the only significant characteristics were birth order, sex of the child, place of delivery, length of breastfeeding, wealth index, education, regions, marriage to the first interval, and kind of fuel. According to our research, breastfeeding duration has a significant impact on under-5 mortality in India, which is in line with research on under-5 mortality in Ethiopia [8]. We evaluated several machine learning models (LR, DT, SVM, and KNN) to forecast mortality for children under five years old in India by utilizing performance measures, including AUC and confusion matrix. The results indicate that machine-learning techniques outperform logistic regression models when it comes to identifying characteristics linked to mortality for children under five [11\u0026ndash;12]. When forecasting under-5 mortality in India, the DT model did better.\u003c/p\u003e\n\u003cp\u003eTo predict under-5 mortality in India, the DT model considered the individual and interaction effects of each of the chosen parameters. Because the requirements of the logistic regression model are not met, it is unable to predict under-5 mortality with accuracy. In particular, multicollinearity affects substantial predictor independence [11]. It is advised to use only independent variables in Logistic Regression modeling to overcome multicollinearity difficulties and prevent inaccurate findings. On the other hand, the decision tree model is a more dependable option for estimating under-5 mortality in India due to its better performance and absence of particular assumptions. Support vector machines (SVM) and K-Nearest Neighbors (KNN), two machine learning models that are flexible due to their lower assumption counts, are included in this study.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMost preterm newborns\u0026apos; death risk may be accurately predicted using this method, which can also mimic and forecast mortality rates in the human population. The superiority of machine learning model methods over conventional analytic methods has also been demonstrated by earlier studies. According to a recent study, machine learning models are more appropriate for identifying factors contributing to infant mortality and verifying superior goodness of fit in most key groups. Furthermore, machine learning models are extremely helpful in forecasting health studies, resulting in better and more appropriate policy choices.\u003c/p\u003e\n\u003cp\u003eWe used the most recent nationally representative NFHS-V 2019\u0026ndash;21 dataset. However, because the study is cross-sectional, proving causal inference is \u0026nbsp;not possible. The results are based on self-reported information subject to social desirability and memory biases. The regression model may not account for socioeconomic factors contributing to under-5 mortality. Our paper thoroughly analyzes several machine learning models\u0026mdash;DT, SVM, KNN, and LR\u0026mdash;to forecast under-5 mortality in India.\u003c/p\u003e\n\u003cp\u003eThis study aimed to apply several machine learning algorithms to data on under-5 mortality. The ML model for predicting death in children under five is somewhat more accurate than classical logistic regression. Among the machine learning models used in this study, the decision tree model had the highest accuracy in predicting the death of children under five. The study also suggests that, with certain restrictions, logistic regression analysis may help forecast the mortality of morality in children under five. However, this study also showed that some of the variables have an equally important effect on mortality for children under five in both LR and ML models. Machine learning models provide analytical prowess that may enhance analysis capabilities compared to traditional statistical models. Such models could be used to process large-dimensional data in health research.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval:\u003c/strong\u003e Publicly available data was used and does not require ethics approval\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u003c/strong\u003e No funding received from any sources\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthorship contributions:\u003c/strong\u003e GV, and AM conceptualized the manuscript; AM analyzed the data; AM, GV, and SM developed the initial draft of the manuscript. All authors read, reviewed, and contributed equally to the final draft of the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDisclosure of interest:\u003c/strong\u003e No Conflict of Interest\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eBrahma, Dweepobotee and Debasri Mukherjee (2022), \u0026ldquo;Early Warning Signs: Targeting Neonatal and Infant Mortality Using Machine Learning\u0026rdquo;, Applied Economics, 54(1): 57-74.\u003c/li\u003e\n\u003cli\u003eOkagbue, H. I., Oguntunde, P. E., Obasi, E. C., Adamu, P. I., \u0026amp; Opanuga, A. A. (2021). Diagnosing malaria from some symptoms: a machine learning approach and public health implications. \u003cem\u003eHealth and Technology\u003c/em\u003e, \u003cem\u003e11\u003c/em\u003e, 23-37.\u003c/li\u003e\n\u003cli\u003eSingh, A., Kumar, K., \u0026amp; Singh, A. (2019). What explains the decline in neonatal mortality in India in the last three decades? Evidence from three rounds of NFHS surveys. Studies in family planning, 50(4), 337- 355. https://doi.org/10.1111/sifp.12105\u003c/li\u003e\n\u003cli\u003eGebrehiwot, G. G., \u0026amp; Worku, A. G. (2021). Forecasting under-five child mortality in Ethiopia using machine learning algorithms. \u003cem\u003eBMC Public Health\u003c/em\u003e, \u003cem\u003e21\u003c/em\u003e(1), 1024. https://doi.org/10.1186/s12889-021-10968-0.\u003c/li\u003e\n\u003cli\u003eUnited Nations (UN) United Nations Population Division; 2017. World population prospects (2017) https://www.unicef.org/sitan/files/SitAn_India_May_2011.pdf\u003c/li\u003e\n\u003cli\u003eClaeson, M., Bos, E. R., Mawji, T. \u0026amp; Pathmanathan, I. (2000). Reducing child mortality in India in the new millennium. Bulletin of the World Health Organization, 78 (10), 1192 -1199. World Health Organization.\u003c/li\u003e\n\u003cli\u003eBell RT. 2011. A baseline study exploring rural maternal health practices and services about the IGMSY conditional maternity benefit. MSc Thesis. Utrecht University.\u003c/li\u003e\n\u003cli\u003eZhang, H., \u0026amp; Li, X. (2019). Predicting under-five child mortality rates in sub-Saharan Africa using machine learning. \u003cem\u003eJournal of Global Health\u003c/em\u003e, \u003cem\u003e9\u003c/em\u003e(2), 020401. https://doi.org/10.7189/jogh.09.020401.\u003c/li\u003e\n\u003cli\u003eTadesse, A. W., \u0026amp; Awoke, T. Z. (2020). Application of machine learning methods to anticipate under-five mortality rates in Ethiopia. \u003cem\u003eInternational Journal of Health Planning and Management\u003c/em\u003e, \u003cem\u003e35\u003c/em\u003e(4), 1035-1048. https://doi.org/10.1002/hpm.2995.\u003c/li\u003e\n\u003cli\u003eVallin, J. (1976). World trends in infant mortality since 1950. World health statistics report. 29(11), 646-674\u003c/li\u003e\n\u003cli\u003eWang, H., Liddell, C. A., Coates, M. M., ... \u0026amp; Moore, A. R. (2014). Global, regional, and national levels of neonatal, infant, and under-5 mortality during 1990\u0026ndash;2013: a systematic analysis for the Global Burden of Disease Study 2013. The Lancet, 384(9947), 957-979. https://doi.org/10.1016/S0140-6736(14)60497-9\u003c/li\u003e\n\u003cli\u003eBitew FH, Nyarko SH, Potter L, Sparks CS. Machine learning approach for predicting under-five mortality determinants in Ethiopia: evidence from the 2016 Ethiopian Demographic and Health Survey. Genus. 2020;76(1). 10.1186/s41118-020-00106-2.\u003c/li\u003e\n\u003cli\u003eVapnik, V. 1995. The Nature of Statistical Learning Theory. New York, NY: Springer.\u003c/li\u003e\n\u003cli\u003eAttewell, P., D.B. Monaghan, and D. Kwong. 2015. Data Mining for the Social Sciences: An Introduction. Oakland, CA: University of California Press.\u003c/li\u003e\n\u003cli\u003eRahman A, Hossain Z, Kabir E, Rois R. \u0026ldquo;Machine Learning Algorithm for Analysing Infant Mortality in Bangladesh,\u0026rdquo; Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), vol. 13079 LNCS. pp. 205\u0026ndash;219, 2021, 10.1007/978-3-030-90885-0_19.\u003c/li\u003e\n\u003cli\u003eMfateneza E, Rutayisire PC, Biracyaza E, Musafiri S, Mpabuka WG. Application of machine learning methods for predicting infant mortality in Rwanda: analysis of Rwanda demographic health survey 2014\u0026ndash;15 dataset. BMC Pregnancy Childbirth. 2022;22(1):1\u0026ndash;13. 10.1186/s12884-022-04699-8.\u003c/li\u003e\n\u003cli\u003eByun, H., and S.W. Lee. 2002. \u0026ldquo;Applications of Support Vector Machines for Pattern Recognition: A Survey.\u0026rdquo; In SVM \u0026rsquo;02 Proceedings of the First International Workshop on Pattern Recognition with Support Vector Machines, 213\u0026ndash;36. London, UK: Springer.\u003c/li\u003e\n\u003cli\u003eMurphy KP. Machine Learning: A Probabilistic Perspective. The MIT Press; 2012.\u003c/li\u003e\n\u003cli\u003eJames G, Witten D, Hastie T, Tibshirani R. An introduction to statistical learning. Vol. 112. Springer; 2013.\u003c/li\u003e\n\u003cli\u003eAndrews DW, Whang YJ. Additive interactive regression models: circumvention of the curse of dimensionality. Econometric Theory. 1990; 6(4):466\u0026ndash;479. https://doi.org/10.1017/ S0266466600005478\u003c/li\u003e\n\u003cli\u003eBreiman L, Friedman J, Stone CJ, Olshen RA. Classification and regression trees. CRC Press; 1984\u003c/li\u003e\n\u003cli\u003eHothorn T, Hornik K, Zeileis A. Unbiased recursive partitioning: A conditional inference framework. Journal of Computational and Graphical Statistics. 2006; 15(3):651\u0026ndash;674. https://doi.org/10.1198/ 106186006X133933\u003c/li\u003e\n\u003cli\u003eBreiman, L., Friedman, J., Stone, C. J., and Olshen, R. A. (1984). Classification and Regression Trees. Boca Raton, FL: CRC press.\u003c/li\u003e\n\u003cli\u003eCaruana, R., and Niculescu-Mizil, A. (2006). \u0026ldquo;An empirical comparison of supervised learning algorithms,\u0026rdquo; in 23rd International Conference on Machine Learning, (Pittsburgh, PA: ACM Press), 161\u0026ndash;168.\u003c/li\u003e\n\u003cli\u003eC. Junli, and J. Licheng, \u0026quot;Classification mechanism of support vector machines,\u0026quot; 5th International Conference on Signal Processing Proceedings. 16th World Computer Congress, Beijing, 2000, vol.3, pp. 1556-1559, 2000.\u003c/li\u003e\n\u003cli\u003eGoldstein BA, Navar AM, Carter RE. Moving beyond regression techniques in cardiovascular risk prediction: applying machine learning to address analytic challenges. Eur Heart J. 2017;38(23):1805\u0026ndash;14. 30. \u003c/li\u003e\n\u003cli\u003eKotsiantis SB, Zaharakis I, Pintelas P. Supervised machine learning: A review of classification techniques. Emerging artificial intelligence applications in computer engineering. 2007;160(1):3\u0026ndash;24\u003c/li\u003e\n\u003cli\u003eHastie T, Tibshirani R, Friedman JH (2009) The elements of statistical learning: data mining, inference, and prediction, 2nd ed. Springer, New York Return to ref 2009 in article\u003c/li\u003e\n\u003cli\u003eBoser, B.E., I.M. Guyon, and V.N. Vapnik. 1992. \u0026ldquo;A Training Algorithm for Optimal Margin Classifiers.\u0026rdquo; In Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, edited by D. Haussler, 144-152. New York, NY: ACM Press.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTables 1 to 3 are available in the Supplementary Files section.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Machine learning, Under-five mortality, SVM, KNN, DT, LR, Determinants, India","lastPublishedDoi":"10.21203/rs.3.rs-5309131/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5309131/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eReducing the under-5 mortality rate is high on the list of priorities in the global development agenda. The SDG targets a reduction in child mortality rate to less than 25 deaths per 1,000 live births annually by 2030. Though enormous gains have been observed for under-five mortality and child health over the past few years, it remains a significant public health challenge for India. Much earlier studies were made on the under-five mortality rates for children. Except in this study, the most widely used approach undertaken so far, to a great extent, relies on conventional regression analyses that are inherently known to have minimal predictive ability. This study attempts to develop a predictive model based on the advanced techniques of machine learning (AML) that could infer a rather accurate prediction for the under-5 mortality rate of India, namely U5MR.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study used the nationally representative microdata from 7 the National Family Health Survey's fifth version (NFHS-5, 2019-21). Multiple imputation methods, such as filling by modal value, were adopted to treat missing values. We used a feature selection method known as information gain, where we ranked the information-rich features and examined their impact on the prediction of child mortality. The synthetic minority over-sampling method (SMOTE) was used to balance the dataset. To predict the determinants affecting U5MR, we used four machine learning (ML) models (decision trees, logistic regression, support vector machines, and K-nearest neighbors). The predictive power of each ML model was assessed using accuracy, precision, receiver operating characteristic curves, and model accuracy metrics (accuracy, precision, F1 score, ROC).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe descriptive findings demonstrate that India's under-five mortality rates vary significantly by region. The Decision Tree model (96.35%) performed the best out of all the models examined, with the under-five mortality prediction ability ranging from 90% to 96.35%. The best predictive model demonstrates that factors influencing under-five mortality rates in India include duration of breastfeeding status, Marriage to the first birth Interval, education, Birth order, ANC visit, wealth index, place of delivery, and place of residence to estimate under-five mortality risk variables, The Decision Tree machine learning model would engender a greater predictive capability; therefore, it would help in better policy decision-making in this field. Employment of these major elements would greatly enhance a child's chances of survival to be directed by appropriate policies.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion\u003c/strong\u003e:\u003c/p\u003e\n\u003cp\u003eThe decision tree model will perform better than traditional logistic regression. Prediction of under-5 mortality rate at the India level revealed a success accuracy of 96.35 percent and precision of about 60 percent.\u003c/p\u003e","manuscriptTitle":"Predicting factors associated with under-5 mortality in India using machine learning algorithms: evidence from National Family Health Survey, 2019-21","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-10-23 07:23:47","doi":"10.21203/rs.3.rs-5309131/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"fd0c313a-c21e-4a02-9bbc-2dd14630b63e","owner":[],"postedDate":"October 23rd, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-05-09T16:23:20+00:00","versionOfRecord":[],"versionCreatedAt":"2024-10-23 07:23:47","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5309131","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5309131","identity":"rs-5309131","version":["v1"]},"buildId":"zQwnuV7TCBrMSSSToR1PI","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.