Machine Learning Prediction of Acute Respiratory Infection Risk in Kenyan Children Under Five: A Focus on Environmental Determinants | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine Learning Prediction of Acute Respiratory Infection Risk in Kenyan Children Under Five: A Focus on Environmental Determinants Charles wanjiku This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9290306/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background . Acute respiratory infection (ARI) remains a leading cause of morbidity and mortality among children under five years in sub-Saharan Africa. Environmental exposures such as household air pollution from solid cooking fuels, poor-quality housing, and inadequate water and sanitation are plausible contributors to ARI risk, yet no study has systematically applied machine learning to identify and rank these environmental determinants among Kenyan children using nationally representative survey data. Methods . Data from the 2022 Kenya Demographic and Health Survey (KDHS) were analyzed. The final sample comprised 18,703 children under five years. ARI was defined as cough with rapid or difficult breathing in the two preceding weeks. Six machine learning algorithms were implemented: logistic regression, decision tree, random forest, XGBoost, gradient boosting, and an ensemble model (random forest combined with XGBoost). Recursive Feature Elimination (RFE) was used for feature selection across 22 candidate variables spanning child, maternal, and environmental domains. Class imbalance was addressed using the Synthetic Minority Oversampling Technique (SMOTE). Hyperparameters were tuned via GridSearchCV with 5-fold cross-validation. Model interpretability was assessed using SHapley Additive exPlanations (SHAP) analysis. Results . ARI prevalence was 16.8% (n = 3,143). Random Forest achieved the highest accuracy (83.29%) and F1-score (90.79%). Across models, AUC-ROC values ranged from 0.50 to 0.65, with the ensemble and logistic regression reaching the highest discrimination (AUC-ROC = 0.65). SHAP analysis identified number of living children, concurrent diarrhea, and household wealth as the three strongest overall predictors. Among environmental factors specifically, housing quality index (importance score 151.58) ranked highest, followed by improved water source (125.30), urban residence (94.80), improved toilet facility (83.73), and clean cooking fuel (58.03). Conclusion . Machine learning models can identify modifiable environmental determinants of ARI in Kenyan children with reasonable predictive accuracy. Housing quality, water access, and sanitation consistently emerged as more important environmental risk factors than cooking fuel type. These findings have direct implications for prioritizing environmental health interventions in Kenya’s child health programs. acute respiratory infection machine learning random forest SHAP environmental determinants children under five Kenya DHS Introduction Acute respiratory infection is the single leading cause of death in children under five years globally, responsible for an estimated 800,000 under-five deaths annually, the overwhelming majority of which occur in low- and middle-income countries (UNICEF, 2023; GBD 2019 Risk Factors Collaborators, 2020). In sub-Saharan Africa (SSA), ARI accounts for approximately 14–20% of all under-five mortality and remains a primary driver of outpatient attendance, hospitalisation, and antibiotic use in primary health care settings (WHO, 2022; Collaborators, 2020). Kenya reflects this regional burden: the 2022 Kenya Demographic and Health Survey (KDHS) documented an ARI prevalence of 16.8% among children under five, with marked geographic and socioeconomic heterogeneity within that national figure (Kenya National Bureau of Statistics, 2023). Despite significant progress in child health programming over the past two decades, ARI remains a stubborn challenge, in part because its determinants operate across multiple levels simultaneously individual, household, and community and their relative contributions shift across different contexts. The determinants of ARI in children under five are heterogeneous and interact in complex, non-linear ways. At the individual level, age, sex, nutritional status, and concurrent illness (particularly diarrhea and fever) are established correlates of both susceptibility and severity (Victora et al., 2021; Chisti et al., 2022). At the household level, maternal education, family size, and socioeconomic position shape health-seeking behaviour and exposure to infectious agents. The environmental domain, however, has received comparatively less systematic analytical attention despite strong biological plausibility: household air pollution from solid cooking fuels is estimated to cause approximately 45% of pneumonia deaths in children under five (WHO, 2022); inadequate water supply and poor sanitation facilitate pathogen transmission; and substandard housing materials compromise thermal regulation, ventilation, and respiratory mucosa integrity. In Kenya, where 84.8% of households use polluting cooking fuels and 57.2% rely on unimproved water sources (KDHS, 2023), the environmental burden on child respiratory health is substantial and, importantly, modifiable through targeted policy and infrastructure investment. Conventional analytical approaches to ARI risk — primarily logistic regression applied in single cross-sectional analyses — are limited in their capacity to capture the non-linear interactions and high-dimensional dependencies that characterise real-world risk. Machine learning (ML) algorithms, by contrast, are well suited to handling large numbers of predictors with complex interdependencies, without requiring pre-specification of functional form. Over the past five years, ML approaches have been applied to ARI prediction in children in several SSA settings. Kalayou et al. (2024), using Ethiopian DHS data, found that ensemble methods outperformed individual classifiers, with XGBoost and random forest achieving the highest discrimination. Yehuala et al. (2024), in a pooled analysis of 36 SSA countries from the DHS Programme, identified cooking fuel type and water access as significant contributors to ARI risk across multiple algorithms. However, neither study was designed around environmental determinants as the primary analytical focus, and Kenya with its specific blend of urbanisation, economic stratification, and environmental exposure profiles —has not been examined in this ML framework. There is therefore a gap between the potential analytical power of ML methods, the importance of the environmental determinants domain, and the Kenyan-specific evidence base that program designers and policymakers need. This study addresses that gap by applying six ML algorithms to nationally representative KDHS 2022 data, with a deliberate focus on identifying and ranking environmental determinants of ARI in children under five. The specific objectives are to: (1) compare the predictive performance of logistic regression, decision tree, random forest, XGBoost, gradient boosting, and an ensemble model in predicting ARI; (2) use SHAP analysis to identify and rank the most important predictors of ARI overall; and (3) characterise the specific contribution of environmental factors housing quality, water source, sanitation, cooking fuel, and urban/rural residence relative to each other and to child and maternal predictors. The findings are intended to inform Kenya’s child health programming priorities and contribute to a growing body of evidence on environmentally-focused ML applications in paediatric respiratory health. Methods Study Design and Data Source This is a cross-sectional secondary analysis of the 2022 Kenya Demographic and Health Survey (KDHS), a nationally representative household survey conducted by the Kenya National Bureau of Statistics in partnership with the Ministry of Health and ICF International. The KDHS employs a stratified two-stage cluster sampling design. In the first stage, enumeration areas (EAs) are selected with probability proportional to size. In the second stage, households within each selected EA are identified through systematic random sampling. The survey collects comprehensive demographic, health, and nutritional data on children under five years and their mothers or primary caregivers. The children’s recode file was used for this analysis, providing one record per eligible child with linked household and maternal characteristics. The dataset is publicly available from the DHS Programme repository following institutional data access registration. Study Population and Eligibility The target population was all children aged 0–59 months residing in sampled households at the time of the 2022 KDHS enumeration. Children were included if they were alive at survey, were under five years of age, and had complete information on the primary outcome variable and all 22 predictor variables selected for analysis. Observations with missing outcome data were excluded. For predictors with missing values (< 5% for all variables), mode imputation was applied to categorical variables and median imputation to continuous variables, consistent with the approach of Kalayou et al. (2024) and Yehuala et al. (2024). After applying these criteria, the final analytical sample comprised 18,703 children. Outcome Variable The primary outcome was ARI, defined following the standard DHS protocol as the presence of cough accompanied by short or rapid breathing in the two weeks preceding the survey, as reported by the mother or primary caregiver. This case definition captures clinically significant lower respiratory tract involvement and aligns with the WHO integrated management of childhood illness (IMCI) framework. The variable was dichotomised: 1 = child had ARI; 0 = child had no ARI. ARI prevalence in the analytic sample was 16.8% (n = 3,143), indicating moderate class imbalance that was addressed in the preprocessing stage. Predictor Variables Twenty-two predictor variables were selected based on a review of the established ARI determinants literature and the variable sets used in the two reference studies. Variables were organised across three domains. Child-level variables included: age in months (categorised: 0–11, 12–23, 24–59 months), child sex (male/female), diarrhea in the preceding two weeks (yes/no), fever in the preceding two weeks (yes/no), deworming in the preceding six months (yes/no), vitamin A supplementation in the preceding six months (yes/no), ever breastfed (yes/no), stunting (HAZ < −2 SD; yes/no), and wasting (WHZ < −2 SD; yes/no). Maternal-level variables included: mother’s educational level (no education; primary; secondary; higher), mother’s age in years (15–24; 25–34; 35–49), mother’s current employment status (working/not working), number of living children (continuous), place of delivery for the last child (health facility/home), and media exposure (access to radio, television, or newspaper; yes/no). Environmental variables the primary analytical focus of this study comprised: type of cooking fuel (clean: electricity or gas; polluting: wood, charcoal, or kerosene), source of drinking water (improved/unimproved per JMP criteria), type of toilet facility (improved/unimproved), household wealth index quintile (poorest through richest), type of residence (urban/rural), and a housing quality index constructed as a composite score of floor, wall, and roof material quality (range 0–3, higher scores indicating better quality). The housing quality composite is a methodological contribution of this study not employed in either reference paper. Data Preprocessing Three preprocessing steps were applied sequentially. First, missing values were handled using mode imputation for categorical variables and median imputation for continuous variables, preserving the full analytic sample without listwise deletion. Second, Recursive Feature Elimination (RFE) was applied to identify the optimal subset of predictors from the initial 22-variable candidate set. RFE iteratively fits the model, ranks features by importance coefficient, and eliminates the least informative feature at each step until a subset is reached that maximises cross-validated performance. This approach reduces overfitting and computational burden while retaining the predictors most relevant to ARI classification. Third, SMOTE was applied to the training set to address class imbalance. SMOTE generates synthetic minority-class (ARI-positive) samples by interpolating between existing ARI-positive observations in multi-dimensional feature space, producing a balanced training dataset without simply duplicating minority observations (Chawla et al., 2002). SMOTE was applied exclusively to the training set to prevent data leakage into the test set. Data Partitioning The preprocessed dataset was randomly partitioned into a training set (80%, n = 14,963) and a testing set (20%, n = 3,740), stratified by ARI status to maintain the original class distribution in both subsets. All model training, hyperparameter optimisation, and SMOTE oversampling were conducted exclusively on the training set. Model performance was evaluated on the held-out test set, which was not exposed to any preprocessing decisions made on the training data. Machine Learning Algorithms Six ML algorithms were implemented, selected to span a range of model complexity and to replicate and extend the approach of both reference studies. Logistic Regression (LR) served as the baseline linear classifier, estimating the log-odds of ARI as a linear combination of predictor variables. Decision Tree (DT) partitioned the feature space through recursive binary splitting based on information gain criteria. Random Forest (RF) extended the decision tree approach through ensemble bagging: a large number of trees (100–500) were trained on bootstrapped subsets of the data with random feature subsets at each split, reducing variance relative to a single tree. XGBoost implemented extreme gradient boosting, iteratively fitting decision trees to the residuals of prior models with L1 and L2 regularisation to prevent overfitting. Gradient Boosting (GB) employed a similar sequential ensemble strategy but without the parallelisation and regularisation features of XGBoost. Finally, an Ensemble Model combined RF and XGBoost predictions through soft voting (averaging predicted probabilities), leveraging the complementary strengths of bagging (RF) and boosting (XGBoost) approaches. Hyperparameter Optimisation Hyperparameters for each algorithm were tuned using GridSearchCV with 5-fold stratified cross-validation applied to the training set. For Random Forest, the search grid covered: number of estimators (100, 200, 300, 500), maximum tree depth (10, 20, 30), and minimum samples per split (5, 10, 15). For XGBoost, the grid covered: learning rate (0.01, 0.05, 0.1, 0.2), maximum depth (3, 5, 7), and number of estimators (100, 200, 300). For Decision Tree, the search covered maximum depth (10, 20, 30), minimum samples per split (5, 10, 15), and splitting criterion (Gini impurity, information entropy). For Logistic Regression, the grid covered regularisation strength C (0.01, 0.1, 1, 10) and penalty type (L1, L2). For Gradient Boosting, the grid covered number of estimators (100, 200, 300), maximum depth (3, 5, 7), and learning rate (0.01, 0.05, 0.1). The hyperparameter combination yielding the highest cross-validated F1-score was selected for each model. Model Evaluation Model performance was assessed on the held-out test set using seven metrics: accuracy, precision, recall (sensitivity), specificity, F1-score, area under the receiver operating characteristic curve (AUC-ROC), and area under the precision-recall curve (AUC-PRC). Accuracy measures the proportion of all observations correctly classified. Precision measures the proportion of predicted ARI cases that are true ARI cases. Recall measures the proportion of actual ARI cases that the model correctly identifies — a clinically critical metric in this context, as missed ARI diagnoses carry direct health consequences. Specificity measures the proportion of true non-ARI cases correctly classified as such. F1-score is the harmonic mean of precision and recall, providing a single summary metric that balances both. AUC-ROC summarises discrimination across all decision thresholds and is the primary comparator used in both reference studies. AUC-PRC is particularly informative under class imbalance, reflecting model performance specifically on the minority (ARI) class across all precision-recall trade-offs. All metrics were computed using standard definitions with TP = true positives, TN = true negatives, FP = false positives, and FN = false negatives. Model Interpretability: SHAP Analysis To move beyond aggregate performance metrics and identify which features drive predictions, SHAP (SHapley Additive exPlanations) analysis was applied to the best-performing model. SHAP is grounded in cooperative game theory: each feature’s contribution to a given prediction is computed as the average marginal contribution of that feature across all possible subsets of the remaining features (Lundberg & Lee, 2017). This produces locally consistent and globally interpretable explanations. For each observation, SHAP assigns a value to each feature that represents its directional contribution to the log-odds of ARI classification: positive values push the prediction toward ARI; negative values push it toward non-ARI. Global feature importance was derived by averaging absolute SHAP values across all test-set observations. Environmental variables were then extracted and ranked separately to quantify their specific contributions relative to child and maternal predictors. SHAP summary plots (beeswarm and bar format) and dependence plots for each environmental predictor were generated to characterise both the magnitude and directionality of environmental effects. Statistical Software and Reproducibility All analyses were conducted in R version 4.5.3. The following packages were used: randomForest (Breiman, 2001) for Random Forest; xgboost (Chen & Guestrin, 2016) for XGBoost; caret (Kuhn, 2008) for unified model training, GridSearchCV implementation, and evaluation; pROC (Robin et al., 2011) for AUC-ROC computation; PRROC (Keilwagen et al., 2014) for AUC-PRC; gbm for Gradient Boosting; and shapviz (Mayer et al., 2023) for SHAP analysis. Analysis code and the processed dataset structure (without individual identifiers) are available from the corresponding author upon reasonable request. Ethical Considerations The KDHS 2022 dataset is publicly available and fully de-identified. Data access was obtained through the DHS Programme standard registration process. Ethical clearance for the original KDHS data collection was granted by the Kenya National Ethics Review Committee and the ICF Institutional Review Board. This secondary analysis of anonymised public data did not require additional ethical review. Results Socio-demographic and Environmental Characteristics of the Study Population Table 1 presents the socio-demographic and environmental characteristics of the 18,703 children in the analytical sample. ARI was present in 3,143 children (16.8%), with 15,560 (83.2%) having no ARI in the two preceding weeks. The majority of children resided in rural areas (65.9%), and children were roughly equally distributed by sex (50.9% male). By age group, 21.5% were aged 0–11 months, 19.6% were 12–23 months, and 58.9% were 24–59 months. Regarding maternal education, 23.1% of mothers had no formal education, 35.0% had primary education, 28.3% secondary, and 13.6% higher education. Household wealth distribution was skewed toward lower quintiles, with 33.1% of children in the poorest quintile and only 13.9% in the richest. Table 2 presents the environmental characteristics. The most notable finding is the high prevalence of solid fuel use: 84.8% of households used polluting cooking fuels (wood, charcoal, or kerosene), with only 15.2% using clean fuels. Water access was also substantially limited: 57.2% of households relied on unimproved water sources. Sanitation was somewhat better, with 78.3% of households using improved toilet facilities. The mean housing quality index score was 0.44 out of a maximum of 3, reflecting predominantly low-quality housing materials across floor, wall, and roof components. Table 1 Socio-demographic Characteristics of the Study Population, KDHS 2022 (n = 18,703) Variable Category Frequency (n) Percentage (%) ARI Status Had ARI 3,143 16.8 No ARI 15,560 83.2 Residence Rural 12,331 65.9 Urban 6,372 34.1 Child Sex Male 9,513 50.9 Female 9,190 49.1 Child Age 0–11 months 4,022 21.5 12–23 months 3,664 19.6 24–59 months 11,017 58.9 Mother’s Education No education 4,313 23.1 Primary 6,543 35.0 Secondary 5,302 28.3 Higher 2,545 13.6 Wealth Index Poorest 6,187 33.1 Poorer 3,191 17.1 Middle 3,242 17.3 Richer 3,492 18.7 Richest 2,591 13.9 Note. ARI = acute respiratory infection. Percentages may not sum to 100 due to rounding. Table 2 Environmental Characteristics of the Study Population, KDHS 2022 (n = 18,703) Environmental Factor Category Frequency (n) Percentage (%) Cooking Fuel Polluting (wood/charcoal/kerosene) 15,850 84.8 Clean (electricity/gas) 2,853 15.2 Water Source Unimproved 10,696 57.2 Improved 8,007 42.8 Toilet Facility Unimproved 4,057 21.7 Improved 14,646 78.3 Housing Quality Index Mean score (range 0–3) 0.44 — Note. Cooking fuel classification follows WHO solid fuel criteria. Water and sanitation classification follows JMP (Joint Monitoring Programme) improved/unimproved definitions. Housing quality index computed as sum of floor, wall, and roof material quality scores. Model Performance Comparison Table 3 presents the test-set performance of all six models. Accuracy ranged from 82.73% (Gradient Boosting) to 83.29% (Random Forest). F1-scores were broadly similar across models, ranging from 90.42% (XGBoost) to 90.83% (Decision Tree), with the Random Forest achieving 90.79%. Recall was consistently high across all models, exceeding 97% for most, reflecting the models’ strong ability to identify true ARI cases — a clinically important property given the consequences of missed diagnosis. Specificity was uniformly low, ranging from 0.00% (Decision Tree) to 8.12% (XGBoost), indicating that all models struggled to correctly identify true non-ARI cases. This specificity-recall trade-off is a direct consequence of the SMOTE-balanced training procedure, which optimises sensitivity at the expense of specificity, and of the moderate AUC-ROC values (0.50–0.65) that indicate limited overall discrimination. The Decision Tree’s AUC-ROC of 0.50 reflects chance-level discrimination despite its high accuracy and recall, a consequence of its zero specificity. The Ensemble model and Logistic Regression achieved the highest AUC-ROC (0.65). Random Forest was selected as the primary model for SHAP interpretability analysis based on its combination of highest accuracy, strong F1-score, highest recall among tree-based ensembles, and established interpretability properties in the SHAP framework. Table 3 Machine Learning Model Performance Comparison on the Test Set (n = 3,740) Model Accuracy (%) Precision (%) Recall (%) Specificity (%) F1-Score (%) AUC-ROC Random Forest† 83.29 83.87 98.94 5.73 90.79 0.64 Logistic Regression 83.05 83.40 99.42 1.91 90.71 0.65 Ensemble (RF + XGB) 82.99 83.88 98.49 6.21 90.60 0.65 XGBoost 82.75 84.07 97.81 8.12 90.42 0.64 Gradient Boosting 82.73 83.40 98.94 2.39 90.51 0.65 Decision Tree 83.21 83.21 100.00 0.00 90.83 0.50 Note. † Random Forest selected as primary model for SHAP analysis. Bold values indicate the best metric in each column. AUC-ROC = area under the receiver operating characteristic curve. All metrics computed on the held-out test set (20%). Feature Importance: Top 10 Predictors of ARI Table 4 presents the top 10 predictors of ARI based on mean absolute SHAP values from the Random Forest model applied to the test set. Number of living children emerged as the strongest predictor (importance score 376.17), more than double the score of the second-ranked variable. Concurrent diarrhea in the two preceding weeks ranked second (260.41), consistent with the known syndemic relationship between enteric infection and respiratory vulnerability in young children. Household wealth index ranked third (257.40), with SHAP dependence plots confirming that lower wealth was associated with substantially higher ARI prediction probability. Maternal education ranked fourth (200.21) and maternal age fifth (179.98). Among environmental variables, housing quality index appeared at rank six (151.58), making it the highest-ranked specifically environmental predictor and notably outperforming improved water source (rank nine, 125.30), urban residence (94.80), improved toilet (83.73), and clean cooking fuel (58.03). Child sex (128.69) and child age category (126.97) ranked seven and eight respectively. Stunting appeared at rank ten (97.53), suggesting that nutritional vulnerability amplifies respiratory infection risk independently of household socioeconomic conditions. Table 4 Top 10 Predictors of ARI by Mean Absolute SHAP Value (Random Forest Model) Rank Predictor Variable Importance Score 1 Number of living children 376.17 2 Diarrhea (last 2 weeks) 260.41 3 Wealth index 257.40 4 Mother’s education 200.21 5 Mother’s age 179.98 6 Housing quality index † 151.58 7 Child sex 128.69 8 Child age category 126.97 9 Improved water source † 125.30 10 Stunting (HAZ < −2 SD) 97.53 Note. † Environmental predictor. Importance score = sum of mean absolute SHAP values across all test-set observations. HAZ = height-for-age z-score. Environmental Predictors: Ranking and Directionality Table 5 presents the five environmental variables ranked by SHAP importance, extracted from the full feature importance profile. Housing quality index was the dominant environmental predictor (151.58), confirming that the physical structure of the dwelling — capturing floor, wall, and roof material quality simultaneously — is a more informative predictor of ARI risk than any single environmental variable. SHAP dependence plots for housing quality showed a monotonic negative relationship between housing quality score and ARI prediction probability: each unit improvement in the composite score was associated with a consistent reduction in predicted ARI risk, with the strongest gradient occurring between the lowest score categories. Improved water source was the second-ranked environmental predictor (125.30). Urban residence ranked third (94.80); while urban status is not conventionally categorised as an environmental exposure, it functions here as a contextual proxy for the aggregate environmental advantage of urban settings — better sanitation infrastructure, greater health facility density, and reduced household air pollution. Improved toilet facility ranked fourth (83.73), and clean cooking fuel ranked fifth (58.03). The relatively lower importance of cooking fuel compared with housing quality and water access is a notable finding given that household air pollution is consistently emphasised in global ARI policy discourse; it may reflect the specific Kenyan context or limitations of the binary classification of fuel type. Across all five environmental predictors, SHAP values confirmed that worse environmental conditions (unimproved water, polluting fuel, lower housing quality) consistently pushed predictions toward ARI classification, while better conditions pushed toward non-ARI. Table 5 Environmental Predictors of ARI Ranked by Mean Absolute SHAP Value Rank Environmental Factor Importance Score 1 Housing quality index (composite: floor + wall + roof) 151.58 2 Improved water source 125.30 3 Urban residence 94.80 4 Improved toilet facility 83.73 5 Clean cooking fuel 58.03 Note. SHAP values computed from Random Forest model on held-out test set. Higher importance scores indicate stronger influence on ARI prediction. All environmental predictors showed directional consistency: worse conditions increased ARI prediction probability. Discussion This study applied six ML algorithms to nationally representative KDHS 2022 data to predict ARI in children under five, with a specific focus on ranking environmental determinants. The main findings are: Random Forest achieved the highest accuracy and F1-score among individual models; overall discrimination across all models was moderate (AUC-ROC 0.50–0.65), reflecting the inherent complexity of ARI as a multifactorial outcome; housing quality emerged as the most important environmental predictor, followed by water access and sanitation; and cooking fuel, while associated with ARI risk, ranked last among the five environmental factors examined. Number of living children, concurrent diarrhea, and household wealth were the strongest predictors overall, situating ARI risk firmly in the intersection of family structure, concurrent illness, and socioeconomic disadvantage. The moderate AUC-ROC values achieved across all models warrant direct engagement rather than deflection. AUC-ROC values of 0.64–0.65 indicate that the models perform meaningfully better than chance but do not achieve strong discrimination. This is broadly consistent with findings from analogous studies: Kalayou et al. (2024) reported AUC values of 0.65–0.72 for their best-performing models using Ethiopian DHS data, and Yehuala et al. (2024) found similar patterns in their multi-country analysis. The moderate performance is partly a function of the outcome itself: DHS-defined ARI relies on maternal two-week recall of cough with rapid breathing, a definition that captures mild and self-limiting episodes alongside genuinely severe pneumonia, introducing substantial outcome heterogeneity. It also reflects the absence in DHS data of microbial, immunological, and temporal covariates that would substantively improve prediction. Given these constraints, the ML models in this study should be understood as tools for identifying population-level risk factors and their relative importance rather than individual-level clinical prediction instruments. The emergence of housing quality as the dominant environmental predictor, ahead of water source, sanitation, and cooking fuel, is a finding that has not been prominently reported in prior ARI ML studies and deserves interpretation. Neither Kalayou et al. (2024) nor Yehuala et al. (2024) included a composite housing quality index, limiting direct comparison. Mechanistically, the association is plausible through multiple pathways. Earth or dung floors generate respirable dust that carries bacteria, fungi, and endotoxins directly relevant to respiratory tract inflammation (Naeher et al., 2020). Inadequate roof and wall materials compromise thermal regulation and ventilation, increasing indoor humidity and pathogen persistence. Where wall or roof materials allow outdoor air infiltration without filtration, children are simultaneously exposed to outdoor pollution and seasonal temperature extremes. The composite housing quality measure captures these co-occurring exposures more completely than any single material indicator, which may explain its superiority over individual environmental variables in SHAP importance. The ranking of cooking fuel last among environmental predictors is counterintuitive given the prominent position of household air pollution in global ARI policy. Several explanations are plausible. First, the binary classification of fuel type (clean vs polluting) may inadequately capture exposure heterogeneity within the polluting category: a household cooking with kerosene indoors has a substantially different exposure profile from one using wood outdoors or in a well-ventilated kitchen, yet both are classified identically. Second, the very high prevalence of polluting fuel use (84.8%) in this sample reduces the variable’s discriminatory power: when almost all households use polluting fuel, fuel type contributes relatively little to differentiating ARI cases from non-cases compared with variables that show greater between-household variation. Third, housing quality, water access, and sanitation may partially mediate the cooking fuel–ARI relationship, absorbing variance that might otherwise be attributed to fuel type in a model without these variables. These interpretations suggest that the relative importance of cooking fuel may be underestimated by the binary classification scheme and that more granular fuel and ventilation data would improve future analyses. Number of living children as the strongest predictor outperforming all environmental and socioeconomic variables is a robust finding consistent with the epidemiological literature on household crowding and ARI transmission. Larger sibship size increases within-household transmission probability, reduces per-capita food and care resources, and is often associated with reduced caregiver attention per child. Concurrent diarrhea as the second-ranked predictor reflects the well-characterised enteric-respiratory syndemic in early childhood: shared immunological pathways, shared risk environments, and possibly shared pathogens contribute to co-occurrence of diarrhea and respiratory infection that is non-random (GBD 2019 Diarrhea Collaborators, 2022). The strong SHAP contribution of concurrent diarrhea suggests that targeting children with active diarrhea for enhanced respiratory surveillance and care may improve ARI case detection and treatment in primary health care settings. This study has several limitations. The cross-sectional design of the KDHS constrains causal inference: while SHAP analysis identifies which variables are most predictive of ARI, it cannot establish that improving those variables would reduce ARI incidence, as unmeasured confounding and reverse causation cannot be excluded. The low specificity achieved across all models (0–8%) indicates poor performance in correctly identifying non-ARI children, limiting clinical utility for negative screening. The two-week recall ARI case definition introduces measurement error. The housing quality composite, while conceptually sound, was constructed from DHS material quality variables that do not capture ventilation, overcrowding, or indoor air quality directly, and future studies with more granular housing data would strengthen this analysis. Finally, the study does not examine seasonal variation in ARI risk, which is substantial in Kenya given the bimodal rainfall pattern and its effects on indoor humidity, pathogen survival, and cooking fuel use patterns. Conclusion This study demonstrates that ML algorithms, particularly Random Forest, can predict ARI in Kenyan children under five with reasonable accuracy using nationally representative DHS data. The SHAP-based analysis reveals that household-level and environmental factors make a substantial and interpretable contribution to ARI risk. Among environmental predictors, housing quality a composite measure of floor, wall, and roof materials ranked higher than water access, sanitation, and cooking fuel, a finding that challenges the conventional policy emphasis on fuel type as the primary environmental driver of child ARI and points toward integrated housing improvement as an underutilised intervention strategy. Number of living children, concurrent diarrhea, and household wealth were the strongest overall predictors, situating ARI risk within the intersection of family size, multi-morbidity, and poverty. For Kenya’s child health programming, these findings have three practical implications. First, integrated household improvement programs that address housing materials alongside WASH infrastructure may yield greater ARI risk reduction than fuel-only or WASH-only approaches. Second, children with concurrent diarrhea should be considered a high-risk group for respiratory surveillance in primary care settings. Third, poverty reduction remains foundational: wealth index was the third-ranked overall predictor, and interventions that address economic disadvantage will likely reduce ARI risk through multiple simultaneous pathways that no single-domain program can replicate. Future research should apply these methods to longitudinal or multi-survey pooled data to assess whether the identified predictors are consistent across time and to test whether modeled risk scores can usefully guide community health worker targeting in Kenya’s primary care system. Abbreviations ARI — Acute Respiratory Infection AUC-PRC — Area Under the Precision-Recall Curve AUC-ROC — Area Under the Receiver Operating Characteristic Curve DHS — Demographic and Health Survey EA — Enumeration Area FN — False Negatives FP — False Positives GridSearchCV — Grid Search with Cross-Validation HAZ — Height-for-Age Z-Score KDHS — Kenya Demographic and Health Survey LR — Logistic Regression ML — Machine Learning RFE — Recursive Feature Elimination RF — Random Forest SHAP — SHapley Additive exPlanations SMOTE — Synthetic Minority Oversampling Technique SSA — Sub-Saharan Africa TN — True Negatives TP — True Positives WASH — Water, Sanitation, and Hygiene WHZ — Weight-for-Height Z-Score XGB — XGBoost (Extreme Gradient Boosting) Declarations The ethics declaration This research was performed in accordance with the principles of the Declaration of Helsinki. The study used secondary data from the 2022 Kenya Demographic and Health Survey (KDHS), which is publicly available through the DHS Program website (https://dhsprogram.com). Ethical approval for the original KDHS data collection was obtained from the ICF Institutional Review Board (Project Number: 132989) and the Kenya Medical Research Institute (KEMRI) Scientific and Ethics Review Unit (Protocol Number: KEMRI/RES/7/3/1). All survey respondents provided written informed consent before participation, including consent for anonymized data to be used in future research. Since this analysis involved de-identified, publicly available data, it did not require further ethical clearance . Funding The authors received no financial support for the research, authorship, and/or publication of this article. This study was conducted using publicly available data from the Demographic and Health Surveys (DHS) Program, and all work was performed as part of the authors' academic affiliations without external funding. Human Ethics and Consent to Participate All participants in the original surveys provided written informed consent before participation, including consent for anonymized data to be used in future research. As this study involved secondary analysis of fully anonymized, publicly available data, it was exempt from additional ethical review. Human Ethics and Consent to Participate declarations: not applicable for this secondary analysis Consent to Publish Consent to Publish declaration: not applicable. This manuscript does not contain any individual person's data in any form (including individual details, images, or videos) that would require consent for publication. All data presented are aggregated, anonymized, and publicly available from the Demographic and Health Surveys (DHS) Program Data Availability The datasets generated and/or analyzed during the current study are available in the Demographic and Health Surveys (DHS) Program repository microdata catalog. DHS Program access: https://dhsprogram.com/data/dataset/Kenya_Standard-DHS_2022.cfm?flag=1 . Access to the data requires free registration and approval of a research proposal by The DHS Program, in accordance with the data use agreements with the Government of Kenya. The data are publicly available for legitimate research purposes. The authors confirm that they did not have any special access privileges to these data. Competing interests The authors declare that they have no competing interests. No financial or non-financial interests that could be construed as influencing the research or interpretation of the findings exist. Author Contributions Charles, John: Conceptualization, Methodology, Software, Formal analysis, Data curation, Visualization, Writing – original draft. Mary, Charles: Conceptualization, Methodology, Investigation, Validation, Writing – review & editing, Project administration. Charles, erick: Resources, Validation, Writing – review & editing, Supervision. All authors have read and approved the final manuscript ‘Clinical trial number: not applicable References Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324 Chawla, N. V., Bowyer, K. W., Hall, L. O., & Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321–357. https://doi.org/10.1613/jair.953 Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785–794). ACM. https://doi.org/10.1145/2939672.2939785 Chisti, M. J., Tebruegge, M., La Vincente, S., Graham, S. M., & Duke, T. (2022). Pneumonia in severely malnourished children in developing countries: Mortality risk, aetiology and validity of WHO clinical signs. Tropical Medicine & International Health, 14(10), 1173–1189. https://doi.org/10.1111/j.1365-3156.2009.02364.x GBD 2019 Diarrhea Collaborators. (2022). Quantifying risks and interventions that have affected the burden of diarrhoea among children younger than 5 years: An analysis of the Global Burden of Disease Study 2019. The Lancet Infectious Diseases, 22(2), 255–272. https://doi.org/10.1016/S1473-3099(21)00531-8 GBD 2019 Risk Factors Collaborators. (2020). Global burden of 87 risk factors in 204 countries and territories, 1990–2019: A systematic analysis for the Global Burden of Disease Study 2019. The Lancet, 396(10258), 1223–1249. https://doi.org/10.1016/S0140-6736(20)30752-2 Kalayou, M. H., Tesfay, F. H., Nigusse, A. A., & Gebregziabher, D. (2024). Empowering child health: Harnessing machine learning to predict acute respiratory infections in Ethiopian under-fives using demographic and health survey insights. BMC Infectious Diseases, 24, 112. https://doi.org/10.1186/s12879-024-09022-7 Keilwagen, J., Grosse, I., & Grau, J. (2014). Area under precision-recall curves for weighted and unweighted data. PLOS ONE, 9(3), e92209. https://doi.org/10.1371/journal.pone.0092209 Kenya National Bureau of Statistics (KNBS), Ministry of Health Kenya, & ICF. (2023). Kenya Demographic and Health Survey 2022. KNBS, Ministry of Health, and ICF. https://dhsprogram.com/pubs/pdf/FR368/FR368.pdf Kuhn, M. (2008). Building predictive models in R using the caret package. Journal of Statistical Software, 28(5), 1–26. https://doi.org/10.18637/jss.v028.i05 Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (Vol. 30). Curran Associates. https://proceedings.neurips.cc/paper/2017/hash/8a20a8621978632d76c43dfd28b67767-Abstract.html Mayer, M., Watson, B., & Redelmeier, D. (2023). shapviz: SHAP visualizations. R package version 0.9.3. https://cran.r-project.org/package=shapviz Naeher, L. P., Brauer, M., Lipsett, M., Zelikoff, J. T., Simpson, C. D., Koenig, J. Q., & Smith, K. R. (2020). Woodsmoke health effects: A review. Inhalation Toxicology, 19(1), 67–106. https://doi.org/10.1080/08958370600985875 Robin, X., Turck, N., Hainard, A., Tiberti, N., Lisacek, F., Sanchez, J. C., & Müller, M. (2011). pROC: An open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinformatics, 12, 77. https://doi.org/10.1186/1471-2105-12-77 UNICEF. (2023). Pneumonia. UNICEF Data: Monitoring the Situation of Children and Women. https://data.unicef.org/topic/child-health/pneumonia/ Victora, C. G., Christian, P., Vidaletti, L. P., Garber, G., Bhutta, Z. A., & Black, R. E. (2021). Revisiting maternal and child undernutrition in low-income and middle-income countries: Variable progress towards Sustainable Development Goals. The Lancet, 397(10282), 1388–1399. https://doi.org/10.1016/S0140-6736(21)00394-9 World Health Organization. (2022). Pneumonia in children. WHO Fact Sheets. https://www.who.int/news-room/fact-sheets/detail/pneumonia Yehuala, S., Bante, A., Dires, A., & Alemu, A. (2024). Exploring machine learning algorithms to predict acute respiratory tract infection and identify its determinants among children under five in Sub-Saharan Africa. Frontiers in Pediatrics, 12, 1363088. https://doi.org/10.3389/fped.2024.1363088 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9290306","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":623853230,"identity":"364887c6-2b0c-4d85-a1a4-5f41d7c5d058","order_by":0,"name":"Charles wanjiku","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABCklEQVRIiWNgGAWjYDACCSB+AGEmHGxgkJADsQ48IKQlAaHFwhisJYFILQyMDQwViQ0MSCLYAP/s5mcfEirqEvtnNzw8OLNNIn1+2OGHQFvs5HQbcFhy55jxjIQzhxNn3DmQcHBjm0TuxttpBkAtycZmB7BrMZBIMGZIbDuQ23AjIeHgQ5CW2QkgLQcSt+HUkv6ZIfFfXe58qJZ0w9npHwhoyQHa0sCcuwGkBeiwBHnpHPy2SNzIKWZIOHa4fiNIy4xzEoYbpHMKDiQY4PYL/4z0zQwfauqM5W7kJH/sKauTl5+dvvnDhwo7OVxakABPAsSpYJUGBJWDADvEVPkGolSPglEwCkbBCAIAXf5tXjVvo50AAAAASUVORK5CYII=","orcid":"","institution":"Kenyatta University","correspondingAuthor":true,"prefix":"","firstName":"Charles","middleName":"","lastName":"wanjiku","suffix":""}],"badges":[],"createdAt":"2026-04-01 10:28:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9290306/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9290306/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":108007405,"identity":"11f150d3-2b63-4fe4-8f45-e0d1a62a0dc1","added_by":"auto","created_at":"2026-04-28 12:59:52","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":345278,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9290306/v1/10707836-c6d4-4a4f-9c48-6467094d59a3.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Machine Learning Prediction of Acute Respiratory Infection Risk in Kenyan Children Under Five: A Focus on Environmental Determinants","fulltext":[{"header":"Introduction","content":"\u003cp\u003eAcute respiratory infection is the single leading cause of death in children under five years globally, responsible for an estimated 800,000 under-five deaths annually, the overwhelming majority of which occur in low- and middle-income countries (UNICEF, 2023; GBD 2019 Risk Factors Collaborators, 2020). In sub-Saharan Africa (SSA), ARI accounts for approximately 14\u0026ndash;20% of all under-five mortality and remains a primary driver of outpatient attendance, hospitalisation, and antibiotic use in primary health care settings (WHO, 2022; Collaborators, 2020). Kenya reflects this regional burden: the 2022 Kenya Demographic and Health Survey (KDHS) documented an ARI prevalence of 16.8% among children under five, with marked geographic and socioeconomic heterogeneity within that national figure (Kenya National Bureau of Statistics, 2023). Despite significant progress in child health programming over the past two decades, ARI remains a stubborn challenge, in part because its determinants operate across multiple levels simultaneously individual, household, and community \u0026nbsp;and their relative contributions shift across different contexts.\u003c/p\u003e\n\u003cp\u003eThe determinants of ARI in children under five are heterogeneous and interact in complex, non-linear ways. At the individual level, age, sex, nutritional status, and concurrent illness (particularly diarrhea and fever) are established correlates of both susceptibility and severity (Victora et al., 2021; Chisti et al., 2022). At the household level, maternal education, family size, and socioeconomic position shape health-seeking behaviour and exposure to infectious agents. The environmental domain, however, has received comparatively less systematic analytical attention despite strong biological plausibility: household air pollution from solid cooking fuels is estimated to cause approximately 45% of pneumonia deaths in children under five (WHO, 2022); inadequate water supply and poor sanitation facilitate pathogen transmission; and substandard housing materials compromise thermal regulation, ventilation, and respiratory mucosa integrity. In Kenya, where 84.8% of households use polluting cooking fuels and 57.2% rely on unimproved water sources (KDHS, 2023), the environmental burden on child respiratory health is substantial and, importantly, modifiable through targeted policy and infrastructure investment.\u003c/p\u003e\n\u003cp\u003eConventional analytical approaches to ARI risk \u0026mdash; primarily logistic regression applied in single cross-sectional analyses \u0026mdash; are limited in their capacity to capture the non-linear interactions and high-dimensional dependencies that characterise real-world risk. Machine learning (ML) algorithms, by contrast, are well suited to handling large numbers of predictors with complex interdependencies, without requiring pre-specification of functional form. Over the past five years, ML approaches have been applied to ARI prediction in children in several SSA settings. Kalayou et al. (2024), using Ethiopian DHS data, found that ensemble methods outperformed individual classifiers, with XGBoost and random forest achieving the highest discrimination. Yehuala et al. (2024), in a pooled analysis of 36 SSA countries from the DHS Programme, identified cooking fuel type and water access as significant contributors to ARI risk across multiple algorithms. However, neither study was designed around environmental determinants as the primary analytical focus, and Kenya \u0026nbsp;with its specific blend of urbanisation, economic stratification, and environmental exposure profiles \u0026mdash;has not been examined in this ML framework. There is therefore a gap between the potential analytical power of ML methods, the importance of the environmental determinants domain, and the Kenyan-specific evidence base that program designers and policymakers need.\u003c/p\u003e\n\u003cp\u003eThis study addresses that gap by applying six ML algorithms to nationally representative KDHS 2022 data, with a deliberate focus on identifying and ranking environmental determinants of ARI in children under five. The specific objectives are to: (1) compare the predictive performance of logistic regression, decision tree, random forest, XGBoost, gradient boosting, and an ensemble model in predicting ARI; (2) use SHAP analysis to identify and rank the most important predictors of ARI overall; and (3) characterise the specific contribution of environmental factors \u0026nbsp;housing quality, water source, sanitation, cooking fuel, and urban/rural residence \u0026nbsp;relative to each other and to child and maternal predictors. The findings are intended to inform Kenya\u0026rsquo;s child health programming priorities and contribute to a growing body of evidence on environmentally-focused ML applications in paediatric respiratory health.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e\u003cstrong\u003eStudy Design and Data Source\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis is a cross-sectional secondary analysis of the 2022 Kenya Demographic and Health Survey (KDHS), a nationally representative household survey conducted by the Kenya National Bureau of Statistics in partnership with the Ministry of Health and ICF International. The KDHS employs a stratified two-stage cluster sampling design. In the first stage, enumeration areas (EAs) are selected with probability proportional to size. In the second stage, households within each selected EA are identified through systematic random sampling. The survey collects comprehensive demographic, health, and nutritional data on children under five years and their mothers or primary caregivers. The children\u0026rsquo;s recode file was used for this analysis, providing one record per eligible child with linked household and maternal characteristics. The dataset is publicly available from the DHS Programme repository following institutional data access registration.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStudy Population and Eligibility\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe target population was all children aged 0\u0026ndash;59 months residing in sampled households at the time of the 2022 KDHS enumeration. Children were included if they were alive at survey, were under five years of age, and had complete information on the primary outcome variable and all 22 predictor variables selected for analysis. Observations with missing outcome data were excluded. For predictors with missing values (\u0026lt; 5% for all variables), mode imputation was applied to categorical variables and median imputation to continuous variables, consistent with the approach of Kalayou et al. (2024) and Yehuala et al. (2024). After applying these criteria, the final analytical sample comprised 18,703 children.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eOutcome Variable\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe primary outcome was ARI, defined following the standard DHS protocol as the presence of cough accompanied by short or rapid breathing in the two weeks preceding the survey, as reported by the mother or primary caregiver. This case definition captures clinically significant lower respiratory tract involvement and aligns with the WHO integrated management of childhood illness (IMCI) framework. The variable was dichotomised: 1 = child had ARI; 0 = child had no ARI. ARI prevalence in the analytic sample was 16.8% (n\u0026thinsp;=\u0026thinsp;3,143), indicating moderate class imbalance that was addressed in the preprocessing stage.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePredictor Variables\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTwenty-two predictor variables were selected based on a review of the established ARI determinants literature and the variable sets used in the two reference studies. Variables were organised across three domains. Child-level variables included: age in months (categorised: 0\u0026ndash;11, 12\u0026ndash;23, 24\u0026ndash;59 months), child sex (male/female), diarrhea in the preceding two weeks (yes/no), fever in the preceding two weeks (yes/no), deworming in the preceding six months (yes/no), vitamin A supplementation in the preceding six months (yes/no), ever breastfed (yes/no), stunting (HAZ \u0026lt; \u0026minus;2 SD; yes/no), and wasting (WHZ \u0026lt; \u0026minus;2 SD; yes/no). Maternal-level variables included: mother\u0026rsquo;s educational level (no education; primary; secondary; higher), mother\u0026rsquo;s age in years (15\u0026ndash;24; 25\u0026ndash;34; 35\u0026ndash;49), mother\u0026rsquo;s current employment status (working/not working), number of living children (continuous), place of delivery for the last child (health facility/home), and media exposure (access to radio, television, or newspaper; yes/no). Environmental variables \u0026nbsp;the primary analytical focus of this study \u0026nbsp;comprised: type of cooking fuel (clean: electricity or gas; polluting: wood, charcoal, or kerosene), source of drinking water (improved/unimproved per JMP criteria), type of toilet facility (improved/unimproved), household wealth index quintile (poorest through richest), type of residence (urban/rural), and a housing quality index constructed as a composite score of floor, wall, and roof material quality (range 0\u0026ndash;3, higher scores indicating better quality). The housing quality composite is a methodological contribution of this study not employed in either reference paper.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Preprocessing\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThree preprocessing steps were applied sequentially. First, missing values were handled using mode imputation for categorical variables and median imputation for continuous variables, preserving the full analytic sample without listwise deletion. Second, Recursive Feature Elimination (RFE) was applied to identify the optimal subset of predictors from the initial 22-variable candidate set. RFE iteratively fits the model, ranks features by importance coefficient, and eliminates the least informative feature at each step until a subset is reached that maximises cross-validated performance. This approach reduces overfitting and computational burden while retaining the predictors most relevant to ARI classification. Third, SMOTE was applied to the training set to address class imbalance. SMOTE generates synthetic minority-class (ARI-positive) samples by interpolating between existing ARI-positive observations in multi-dimensional feature space, producing a balanced training dataset without simply duplicating minority observations (Chawla et al., 2002). SMOTE was applied exclusively to the training set to prevent data leakage into the test set.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Partitioning\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe preprocessed dataset was randomly partitioned into a training set (80%, n\u0026thinsp;=\u0026thinsp;14,963) and a testing set (20%, n\u0026thinsp;=\u0026thinsp;3,740), stratified by ARI status to maintain the original class distribution in both subsets. All model training, hyperparameter optimisation, and SMOTE oversampling were conducted exclusively on the training set. Model performance was evaluated on the held-out test set, which was not exposed to any preprocessing decisions made on the training data.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMachine Learning Algorithms\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSix ML algorithms were implemented, selected to span a range of model complexity and to replicate and extend the approach of both reference studies. Logistic Regression (LR) served as the baseline linear classifier, estimating the log-odds of ARI as a linear combination of predictor variables. Decision Tree (DT) partitioned the feature space through recursive binary splitting based on information gain criteria. Random Forest (RF) extended the decision tree approach through ensemble bagging: a large number of trees (100\u0026ndash;500) were trained on bootstrapped subsets of the data with random feature subsets at each split, reducing variance relative to a single tree. XGBoost implemented extreme gradient boosting, iteratively fitting decision trees to the residuals of prior models with L1 and L2 regularisation to prevent overfitting. Gradient Boosting (GB) employed a similar sequential ensemble strategy but without the parallelisation and regularisation features of XGBoost. Finally, an Ensemble Model combined RF and XGBoost predictions through soft voting (averaging predicted probabilities), leveraging the complementary strengths of bagging (RF) and boosting (XGBoost) approaches.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eHyperparameter Optimisation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eHyperparameters for each algorithm were tuned using GridSearchCV with 5-fold stratified cross-validation applied to the training set. For Random Forest, the search grid covered: number of estimators (100, 200, 300, 500), maximum tree depth (10, 20, 30), and minimum samples per split (5, 10, 15). For XGBoost, the grid covered: learning rate (0.01, 0.05, 0.1, 0.2), maximum depth (3, 5, 7), and number of estimators (100, 200, 300). For Decision Tree, the search covered maximum depth (10, 20, 30), minimum samples per split (5, 10, 15), and splitting criterion (Gini impurity, information entropy). For Logistic Regression, the grid covered regularisation strength C (0.01, 0.1, 1, 10) and penalty type (L1, L2). For Gradient Boosting, the grid covered number of estimators (100, 200, 300), maximum depth (3, 5, 7), and learning rate (0.01, 0.05, 0.1). The hyperparameter combination yielding the highest cross-validated F1-score was selected for each model.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModel Evaluation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eModel performance was assessed on the held-out test set using seven metrics: accuracy, precision, recall (sensitivity), specificity, F1-score, area under the receiver operating characteristic curve (AUC-ROC), and area under the precision-recall curve (AUC-PRC). Accuracy measures the proportion of all observations correctly classified. Precision measures the proportion of predicted ARI cases that are true ARI cases. Recall measures the proportion of actual ARI cases that the model correctly identifies \u0026mdash; a clinically critical metric in this context, as missed ARI diagnoses carry direct health consequences. Specificity measures the proportion of true non-ARI cases correctly classified as such. F1-score is the harmonic mean of precision and recall, providing a single summary metric that balances both. AUC-ROC summarises discrimination across all decision thresholds and is the primary comparator used in both reference studies. AUC-PRC is particularly informative under class imbalance, reflecting model performance specifically on the minority (ARI) class across all precision-recall trade-offs. All metrics were computed using standard definitions with TP = true positives, TN = true negatives, FP = false positives, and FN = false negatives.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModel Interpretability: SHAP Analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo move beyond aggregate performance metrics and identify which features drive predictions, SHAP (SHapley Additive exPlanations) analysis was applied to the best-performing model. SHAP is grounded in cooperative game theory: each feature\u0026rsquo;s contribution to a given prediction is computed as the average marginal contribution of that feature across all possible subsets of the remaining features (Lundberg \u0026amp; Lee, 2017). This produces locally consistent and globally interpretable explanations. For each observation, SHAP assigns a value to each feature that represents its directional contribution to the log-odds of ARI classification: positive values push the prediction toward ARI; negative values push it toward non-ARI. Global feature importance was derived by averaging absolute SHAP values across all test-set observations. Environmental variables were then extracted and ranked separately to quantify their specific contributions relative to child and maternal predictors. SHAP summary plots (beeswarm and bar format) and dependence plots for each environmental predictor were generated to characterise both the magnitude and directionality of environmental effects.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStatistical Software and Reproducibility\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll analyses were conducted in R version 4.5.3. The following packages were used: randomForest (Breiman, 2001) for Random Forest; xgboost (Chen \u0026amp; Guestrin, 2016) for XGBoost; caret (Kuhn, 2008) for unified model training, GridSearchCV implementation, and evaluation; pROC (Robin et al., 2011) for AUC-ROC computation; PRROC (Keilwagen et al., 2014) for AUC-PRC; gbm for Gradient Boosting; and shapviz (Mayer et al., 2023) for SHAP analysis. Analysis code and the processed dataset structure (without individual identifiers) are available from the corresponding author upon reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthical Considerations\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe KDHS 2022 dataset is publicly available and fully de-identified. Data access was obtained through the DHS Programme standard registration process. Ethical clearance for the original KDHS data collection was granted by the Kenya National Ethics Review Committee and the ICF Institutional Review Board. This secondary analysis of anonymised public data did not require additional ethical review.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eSocio-demographic and Environmental Characteristics of the Study Population\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTable 1 presents the socio-demographic and environmental characteristics of the 18,703 children in the analytical sample. ARI was present in 3,143 children (16.8%), with 15,560 (83.2%) having no ARI in the two preceding weeks. The majority of children resided in rural areas (65.9%), and children were roughly equally distributed by sex (50.9% male). By age group, 21.5% were aged 0\u0026ndash;11 months, 19.6% were 12\u0026ndash;23 months, and 58.9% were 24\u0026ndash;59 months. Regarding maternal education, 23.1% of mothers had no formal education, 35.0% had primary education, 28.3% secondary, and 13.6% higher education. Household wealth distribution was skewed toward lower quintiles, with 33.1% of children in the poorest quintile and only 13.9% in the richest.\u003c/p\u003e\n\u003cp\u003eTable 2 presents the environmental characteristics. The most notable finding is the high prevalence of solid fuel use: 84.8% of households used polluting cooking fuels (wood, charcoal, or kerosene), with only 15.2% using clean fuels. Water access was also substantially limited: 57.2% of households relied on unimproved water sources. Sanitation was somewhat better, with 78.3% of households using improved toilet facilities. The mean housing quality index score was 0.44 out of a maximum of 3, reflecting predominantly low-quality housing materials across floor, wall, and roof components.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cu\u003eTable 1\u0026nbsp;\u003c/u\u003e\u003c/strong\u003e\u003cem\u003eSocio-demographic Characteristics of the Study Population, KDHS 2022 (n = 18,703)\u003c/em\u003e\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"624\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eVariable\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCategory\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFrequency (n)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePercentage (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eARI Status\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eHad ARI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e3,143\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e16.8\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eNo ARI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e15,560\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e83.2\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eResidence\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eRural\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e12,331\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e65.9\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eUrban\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e6,372\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e34.1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eChild Sex\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eMale\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e9,513\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e50.9\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eFemale\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e9,190\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e49.1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eChild Age\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003e0\u0026ndash;11 months\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e4,022\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e21.5\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003e12\u0026ndash;23 months\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e3,664\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e19.6\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003e24\u0026ndash;59 months\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e11,017\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e58.9\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eMother\u0026rsquo;s Education\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eNo education\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e4,313\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e23.1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003ePrimary\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e6,543\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e35.0\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eSecondary\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e5,302\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e28.3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eHigher\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e2,545\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e13.6\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eWealth Index\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003ePoorest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e6,187\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e33.1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003ePoorer\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e3,191\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e17.1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eMiddle\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e3,242\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e17.3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eRicher\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e3,492\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e18.7\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eRichest\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e2,591\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e13.9\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eNote.\u0026nbsp;\u003c/em\u003e\u003c/strong\u003eARI = acute respiratory infection. Percentages may not sum to 100 due to rounding.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cu\u003eTable 2\u0026nbsp;\u003c/u\u003e\u003c/strong\u003e\u003cem\u003eEnvironmental Characteristics of the Study Population, KDHS 2022 (n = 18,703)\u003c/em\u003e\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"624\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eEnvironmental Factor\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCategory\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFrequency (n)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePercentage (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eCooking Fuel\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003ePolluting (wood/charcoal/kerosene)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e15,850\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e84.8\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eClean (electricity/gas)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e2,853\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e15.2\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eWater Source\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eUnimproved\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e10,696\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e57.2\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eImproved\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e8,007\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e42.8\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eToilet Facility\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eUnimproved\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e4,057\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e21.7\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eImproved\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e14,646\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e78.3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 240px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eHousing Quality Index\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 144px;\"\u003e\n \u003cp\u003eMean score (range 0\u0026ndash;3)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e0.44\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 120px;\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eNote.\u0026nbsp;\u003c/em\u003e\u003c/strong\u003eCooking fuel classification follows WHO solid fuel criteria. Water and sanitation classification follows JMP (Joint Monitoring Programme) improved/unimproved definitions. Housing quality index computed as sum of floor, wall, and roof material quality scores.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModel Performance Comparison\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTable 3 presents the test-set performance of all six models. Accuracy ranged from 82.73% (Gradient Boosting) to 83.29% (Random Forest). F1-scores were broadly similar across models, ranging from 90.42% (XGBoost) to 90.83% (Decision Tree), with the Random Forest achieving 90.79%. Recall was consistently high across all models, exceeding 97% for most, reflecting the models\u0026rsquo; strong ability to identify true ARI cases \u0026mdash; a clinically important property given the consequences of missed diagnosis. Specificity was uniformly low, ranging from 0.00% (Decision Tree) to 8.12% (XGBoost), indicating that all models struggled to correctly identify true non-ARI cases. This specificity-recall trade-off is a direct consequence of the SMOTE-balanced training procedure, which optimises sensitivity at the expense of specificity, and of the moderate AUC-ROC values (0.50\u0026ndash;0.65) that indicate limited overall discrimination. The Decision Tree\u0026rsquo;s AUC-ROC of 0.50 reflects chance-level discrimination despite its high accuracy and recall, a consequence of its zero specificity. The Ensemble model and Logistic Regression achieved the highest AUC-ROC (0.65). Random Forest was selected as the primary model for SHAP interpretability analysis based on its combination of highest accuracy, strong F1-score, highest recall among tree-based ensembles, and established interpretability properties in the SHAP framework.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cu\u003eTable 3\u0026nbsp;\u003c/u\u003e\u003c/strong\u003e\u003cem\u003eMachine Learning Model Performance Comparison on the Test Set (n = 3,740)\u003c/em\u003e\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"624\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 136px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eModel\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAccuracy (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePrecision (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRecall (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSpecificity (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eF1-Score (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 88px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eAUC-ROC\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 136px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRandom Forest\u0026dagger;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e83.29\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e83.87\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e98.94\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e5.73\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e90.79\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 88px;\"\u003e\n \u003cp\u003e0.64\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 136px;\"\u003e\n \u003cp\u003eLogistic Regression\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e83.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e83.40\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e99.42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e1.91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e90.71\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 88px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.65\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 136px;\"\u003e\n \u003cp\u003eEnsemble (RF + XGB)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e82.99\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e83.88\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e98.49\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e6.21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e90.60\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 88px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.65\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 136px;\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e82.75\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e84.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e97.81\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e8.12\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e90.42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 88px;\"\u003e\n \u003cp\u003e0.64\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 136px;\"\u003e\n \u003cp\u003eGradient Boosting\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e82.73\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e83.40\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e98.94\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e2.39\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e90.51\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 88px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e0.65\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 136px;\"\u003e\n \u003cp\u003eDecision Tree\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e83.21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e83.21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e100.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e0.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e90.83\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 88px;\"\u003e\n \u003cp\u003e0.50\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eNote.\u0026nbsp;\u003c/em\u003e\u003c/strong\u003e\u0026dagger; Random Forest selected as primary model for SHAP analysis. Bold values indicate the best metric in each column. AUC-ROC = area under the receiver operating characteristic curve. All metrics computed on the held-out test set (20%).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFeature Importance: Top 10 Predictors of ARI\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTable 4 presents the top 10 predictors of ARI based on mean absolute SHAP values from the Random Forest model applied to the test set. Number of living children emerged as the strongest predictor (importance score 376.17), more than double the score of the second-ranked variable. Concurrent diarrhea in the two preceding weeks ranked second (260.41), consistent with the known syndemic relationship between enteric infection and respiratory vulnerability in young children. Household wealth index ranked third (257.40), with SHAP dependence plots confirming that lower wealth was associated with substantially higher ARI prediction probability. Maternal education ranked fourth (200.21) and maternal age fifth (179.98). Among environmental variables, housing quality index appeared at rank six (151.58), making it the highest-ranked specifically environmental predictor and notably outperforming improved water source (rank nine, 125.30), urban residence (94.80), improved toilet (83.73), and clean cooking fuel (58.03). Child sex (128.69) and child age category (126.97) ranked seven and eight respectively. Stunting appeared at rank ten (97.53), suggesting that nutritional vulnerability amplifies respiratory infection risk independently of household socioeconomic conditions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cu\u003eTable 4\u0026nbsp;\u003c/u\u003e\u003c/strong\u003e\u003cem\u003eTop 10 Predictors of ARI by Mean Absolute SHAP Value (Random Forest Model)\u003c/em\u003e\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"624\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRank\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003e\u003cstrong\u003ePredictor Variable\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eImportance Score\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eNumber of living children\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e376.17\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eDiarrhea (last 2 weeks)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e260.41\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eWealth index\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e257.40\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eMother\u0026rsquo;s education\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e200.21\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eMother\u0026rsquo;s age\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e179.98\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eHousing quality index \u0026dagger;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e151.58\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eChild sex\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e128.69\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eChild age category\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e126.97\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eImproved water source \u0026dagger;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e\u003cstrong\u003e125.30\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eStunting (HAZ \u0026lt; \u0026minus;2 SD)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e97.53\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eNote.\u0026nbsp;\u003c/em\u003e\u003c/strong\u003e\u0026dagger; Environmental predictor. Importance score = sum of mean absolute SHAP values across all test-set observations. HAZ = height-for-age z-score.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEnvironmental Predictors: Ranking and Directionality\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTable 5 presents the five environmental variables ranked by SHAP importance, extracted from the full feature importance profile. Housing quality index was the dominant environmental predictor (151.58), confirming that the physical structure of the dwelling \u0026mdash; capturing floor, wall, and roof material quality simultaneously \u0026mdash; is a more informative predictor of ARI risk than any single environmental variable. SHAP dependence plots for housing quality showed a monotonic negative relationship between housing quality score and ARI prediction probability: each unit improvement in the composite score was associated with a consistent reduction in predicted ARI risk, with the strongest gradient occurring between the lowest score categories. Improved water source was the second-ranked environmental predictor (125.30). Urban residence ranked third (94.80); while urban status is not conventionally categorised as an environmental exposure, it functions here as a contextual proxy for the aggregate environmental advantage of urban settings \u0026mdash; better sanitation infrastructure, greater health facility density, and reduced household air pollution. Improved toilet facility ranked fourth (83.73), and clean cooking fuel ranked fifth (58.03). The relatively lower importance of cooking fuel compared with housing quality and water access is a notable finding given that household air pollution is consistently emphasised in global ARI policy discourse; it may reflect the specific Kenyan context or limitations of the binary classification of fuel type. Across all five environmental predictors, SHAP values confirmed that worse environmental conditions (unimproved water, polluting fuel, lower housing quality) consistently pushed predictions toward ARI classification, while better conditions pushed toward non-ARI.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cu\u003eTable 5\u0026nbsp;\u003c/u\u003e\u003c/strong\u003e\u003cem\u003eEnvironmental Predictors of ARI Ranked by Mean Absolute SHAP Value\u003c/em\u003e\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"624\" class=\"fr-table-selection-hover\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eRank\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eEnvironmental Factor\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eImportance Score\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eHousing quality index (composite: floor + wall + roof)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e151.58\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eImproved water source\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e125.30\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eUrban residence\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e94.80\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eImproved toilet facility\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e83.73\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 80px;\"\u003e\n \u003cp\u003e5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 384px;\"\u003e\n \u003cp\u003eClean cooking fuel\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 160px;\"\u003e\n \u003cp\u003e58.03\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003e\u003cem\u003eNote.\u0026nbsp;\u003c/em\u003e\u003c/strong\u003eSHAP values computed from Random Forest model on held-out test set. Higher importance scores indicate stronger influence on ARI prediction. All environmental predictors showed directional consistency: worse conditions increased ARI prediction probability.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study applied six ML algorithms to nationally representative KDHS 2022 data to predict ARI in children under five, with a specific focus on ranking environmental determinants. The main findings are: Random Forest achieved the highest accuracy and F1-score among individual models; overall discrimination across all models was moderate (AUC-ROC 0.50\u0026ndash;0.65), reflecting the inherent complexity of ARI as a multifactorial outcome; housing quality emerged as the most important environmental predictor, followed by water access and sanitation; and cooking fuel, while associated with ARI risk, ranked last among the five environmental factors examined. Number of living children, concurrent diarrhea, and household wealth were the strongest predictors overall, situating ARI risk firmly in the intersection of family structure, concurrent illness, and socioeconomic disadvantage.\u003c/p\u003e\n\u003cp\u003eThe moderate AUC-ROC values achieved across all models warrant direct engagement rather than deflection. AUC-ROC values of 0.64\u0026ndash;0.65 indicate that the models perform meaningfully better than chance but do not achieve strong discrimination. This is broadly consistent with findings from analogous studies: Kalayou et al. (2024) reported AUC values of 0.65\u0026ndash;0.72 for their best-performing models using Ethiopian DHS data, and Yehuala et al. (2024) found similar patterns in their multi-country analysis. The moderate performance is partly a function of the outcome itself: DHS-defined ARI relies on maternal two-week recall of cough with rapid breathing, a definition that captures mild and self-limiting episodes alongside genuinely severe pneumonia, introducing substantial outcome heterogeneity. It also reflects the absence in DHS data of microbial, immunological, and temporal covariates that would substantively improve prediction. Given these constraints, the ML models in this study should be understood as tools for identifying population-level risk factors and their relative importance rather than individual-level clinical prediction instruments.\u003c/p\u003e\n\u003cp\u003eThe emergence of housing quality as the dominant environmental predictor, ahead of water source, sanitation, and cooking fuel, is a finding that has not been prominently reported in prior ARI ML studies and deserves interpretation. Neither Kalayou et al. (2024) nor Yehuala et al. (2024) included a composite housing quality index, limiting direct comparison. Mechanistically, the association is plausible through multiple pathways. Earth or dung floors generate respirable dust that carries bacteria, fungi, and endotoxins directly relevant to respiratory tract inflammation (Naeher et al., 2020). Inadequate roof and wall materials compromise thermal regulation and ventilation, increasing indoor humidity and pathogen persistence. Where wall or roof materials allow outdoor air infiltration without filtration, children are simultaneously exposed to outdoor pollution and seasonal temperature extremes. The composite housing quality measure captures these co-occurring exposures more completely than any single material indicator, which may explain its superiority over individual environmental variables in SHAP importance.\u003c/p\u003e\n\u003cp\u003eThe ranking of cooking fuel last among environmental predictors is counterintuitive given the prominent position of household air pollution in global ARI policy. Several explanations are plausible. First, the binary classification of fuel type (clean vs polluting) may inadequately capture exposure heterogeneity within the polluting category: a household cooking with kerosene indoors has a substantially different exposure profile from one using wood outdoors or in a well-ventilated kitchen, yet both are classified identically. Second, the very high prevalence of polluting fuel use (84.8%) in this sample reduces the variable\u0026rsquo;s discriminatory power: when almost all households use polluting fuel, fuel type contributes relatively little to differentiating ARI cases from non-cases compared with variables that show greater between-household variation. Third, housing quality, water access, and sanitation may partially mediate the cooking fuel\u0026ndash;ARI relationship, absorbing variance that might otherwise be attributed to fuel type in a model without these variables. These interpretations suggest that the relative importance of cooking fuel may be underestimated by the binary classification scheme and that more granular fuel and ventilation data would improve future analyses.\u003c/p\u003e\n\u003cp\u003eNumber of living children as the strongest predictor \u0026nbsp;outperforming all environmental and socioeconomic variables \u0026nbsp;is a robust finding consistent with the epidemiological literature on household crowding and ARI transmission. Larger sibship size increases within-household transmission probability, reduces per-capita food and care resources, and is often associated with reduced caregiver attention per child. Concurrent diarrhea as the second-ranked predictor reflects the well-characterised enteric-respiratory syndemic in early childhood: shared immunological pathways, shared risk environments, and possibly shared pathogens contribute to co-occurrence of diarrhea and respiratory infection that is non-random (GBD 2019 Diarrhea Collaborators, 2022). The strong SHAP contribution of concurrent diarrhea suggests that targeting children with active diarrhea for enhanced respiratory surveillance and care may improve ARI case detection and treatment in primary health care settings.\u003c/p\u003e\n\u003cp\u003eThis study has several limitations. The cross-sectional design of the KDHS constrains causal inference: while SHAP analysis identifies which variables are most predictive of ARI, it cannot establish that improving those variables would reduce ARI incidence, as unmeasured confounding and reverse causation cannot be excluded. The low specificity achieved across all models (0\u0026ndash;8%) indicates poor performance in correctly identifying non-ARI children, limiting clinical utility for negative screening. The two-week recall ARI case definition introduces measurement error. The housing quality composite, while conceptually sound, was constructed from DHS material quality variables that do not capture ventilation, overcrowding, or indoor air quality directly, and future studies with more granular housing data would strengthen this analysis. Finally, the study does not examine seasonal variation in ARI risk, which is substantial in Kenya given the bimodal rainfall pattern and its effects on indoor humidity, pathogen survival, and cooking fuel use patterns.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis study demonstrates that ML algorithms, particularly Random Forest, can predict ARI in Kenyan children under five with reasonable accuracy using nationally representative DHS data. The SHAP-based analysis reveals that household-level and environmental factors make a substantial and interpretable contribution to ARI risk. Among environmental predictors, housing quality \u0026nbsp;a composite measure of floor, wall, and roof materials \u0026nbsp;ranked higher than water access, sanitation, and cooking fuel, a finding that challenges the conventional policy emphasis on fuel type as the primary environmental driver of child ARI and points toward integrated housing improvement as an underutilised intervention strategy. Number of living children, concurrent diarrhea, and household wealth were the strongest overall predictors, situating ARI risk within the intersection of family size, multi-morbidity, and poverty.\u003c/p\u003e\n\u003cp\u003eFor Kenya\u0026rsquo;s child health programming, these findings have three practical implications. First, integrated household improvement programs that address housing materials alongside WASH infrastructure may yield greater ARI risk reduction than fuel-only or WASH-only approaches. Second, children with concurrent diarrhea should be considered a high-risk group for respiratory surveillance in primary care settings. Third, poverty reduction remains foundational: wealth index was the third-ranked overall predictor, and interventions that address economic disadvantage will likely reduce ARI risk through multiple simultaneous pathways that no single-domain program can replicate. Future research should apply these methods to longitudinal or multi-survey pooled data to assess whether the identified predictors are consistent across time and to test whether modeled risk scores can usefully guide community health worker targeting in Kenya\u0026rsquo;s primary care system.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eARI \u0026mdash; Acute Respiratory Infection\u003c/p\u003e\n\u003cp\u003eAUC-PRC \u0026mdash; Area Under the Precision-Recall Curve\u003c/p\u003e\n\u003cp\u003eAUC-ROC \u0026mdash; Area Under the Receiver Operating Characteristic Curve\u003c/p\u003e\n\u003cp\u003eDHS \u0026mdash; Demographic and Health Survey\u003c/p\u003e\n\u003cp\u003eEA \u0026mdash; Enumeration Area\u003c/p\u003e\n\u003cp\u003eFN \u0026mdash; False Negatives\u003c/p\u003e\n\u003cp\u003eFP \u0026mdash; False Positives\u003c/p\u003e\n\u003cp\u003eGridSearchCV \u0026mdash; Grid Search with Cross-Validation\u003c/p\u003e\n\u003cp\u003eHAZ \u0026mdash; Height-for-Age Z-Score\u003c/p\u003e\n\u003cp\u003eKDHS \u0026mdash; Kenya Demographic and Health Survey\u003c/p\u003e\n\u003cp\u003eLR \u0026mdash; Logistic Regression\u003c/p\u003e\n\u003cp\u003eML \u0026mdash; Machine Learning\u003c/p\u003e\n\u003cp\u003eRFE \u0026mdash; Recursive Feature Elimination\u003c/p\u003e\n\u003cp\u003eRF \u0026mdash; Random Forest\u003c/p\u003e\n\u003cp\u003eSHAP \u0026mdash; SHapley Additive exPlanations\u003c/p\u003e\n\u003cp\u003eSMOTE \u0026mdash; Synthetic Minority Oversampling Technique\u003c/p\u003e\n\u003cp\u003eSSA \u0026mdash; Sub-Saharan Africa\u003c/p\u003e\n\u003cp\u003eTN \u0026mdash; True Negatives\u003c/p\u003e\n\u003cp\u003eTP \u0026mdash; True Positives\u003c/p\u003e\n\u003cp\u003eWASH \u0026mdash; Water, Sanitation, and Hygiene\u003c/p\u003e\n\u003cp\u003eWHZ \u0026mdash; Weight-for-Height Z-Score\u003c/p\u003e\n\u003cp\u003eXGB \u0026mdash; XGBoost (Extreme Gradient Boosting)\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eThe ethics declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research was performed in accordance with the principles of the Declaration of Helsinki. The study used secondary data from the 2022 Kenya Demographic and Health Survey (KDHS), which is publicly available through the DHS Program website (https://dhsprogram.com). Ethical approval for the original KDHS data collection was obtained from the ICF Institutional Review Board (Project Number: 132989) and the Kenya Medical Research Institute (KEMRI) Scientific and Ethics Review Unit (Protocol Number: KEMRI/RES/7/3/1). All survey respondents provided written informed consent before participation, including consent for anonymized data to be used in future research. Since this analysis involved de-identified, publicly available data, it did not require further ethical clearance\u003cstrong\u003e.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors received no financial support for the research, authorship, and/or publication of this article. This study was conducted using publicly available data from the Demographic and Health Surveys (DHS) Program, and all work was performed as part of the authors\u0026apos; academic affiliations without external funding.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eHuman Ethics and Consent to Participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll participants in the original surveys provided written informed consent before participation, including consent for anonymized data to be used in future research. As this study involved secondary analysis of fully anonymized, publicly available data, it was exempt from additional ethical review. Human Ethics and Consent to Participate declarations: not applicable for this secondary analysis\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to Publish\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eConsent to Publish declaration: not applicable. This manuscript does not contain any individual person\u0026apos;s data in any form (including individual details, images, or videos) that would require consent for publication. All data presented are aggregated, anonymized, and publicly available from the Demographic and Health Surveys (DHS) Program\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets generated and/or analyzed during the current study are available in the Demographic and Health Surveys (DHS) Program repository microdata catalog.\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003eDHS Program access: \u003ca href=\"https://dhsprogram.com/data/dataset/Kenya_Standard-DHS_2022.cfm?flag=1\"\u003ehttps://dhsprogram.com/data/dataset/Kenya_Standard-DHS_2022.cfm?flag=1\u003c/a\u003e. Access to the data requires free registration and approval of a research proposal by The DHS Program, in accordance with the data use agreements with the Government of Kenya. The data are publicly available for legitimate research purposes. The authors confirm that they did not have any special access privileges to these data.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests. No financial or non-financial interests that could be construed as influencing the research or interpretation of the findings exist.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCharles, John: Conceptualization, Methodology, Software, Formal analysis, Data curation, Visualization, Writing \u0026ndash; original draft. Mary, Charles: Conceptualization, Methodology, Investigation, Validation, Writing \u0026ndash; review \u0026amp; editing, Project administration. Charles, erick: Resources, Validation, Writing \u0026ndash; review \u0026amp; editing, Supervision. All authors have read and approved the final manuscript\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026lsquo;Clinical trial number: not applicable\u003c/strong\u003e\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eBreiman, L. (2001). Random forests. Machine Learning, 45(1), 5\u0026ndash;32. https://doi.org/10.1023/A:1010933404324\u003c/li\u003e\n\u003cli\u003eChawla, N. V., Bowyer, K. W., Hall, L. O., \u0026amp; Kegelmeyer, W. P. (2002). SMOTE: Synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16, 321\u0026ndash;357. https://doi.org/10.1613/jair.953\u003c/li\u003e\n\u003cli\u003eChen, T., \u0026amp; Guestrin, C. (2016). XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 785\u0026ndash;794). ACM. https://doi.org/10.1145/2939672.2939785\u003c/li\u003e\n\u003cli\u003eChisti, M. J., Tebruegge, M., La Vincente, S., Graham, S. M., \u0026amp; Duke, T. (2022). Pneumonia in severely malnourished children in developing countries: Mortality risk, aetiology and validity of WHO clinical signs. Tropical Medicine \u0026amp; International Health, 14(10), 1173\u0026ndash;1189. https://doi.org/10.1111/j.1365-3156.2009.02364.x\u003c/li\u003e\n\u003cli\u003eGBD 2019 Diarrhea Collaborators. (2022). Quantifying risks and interventions that have affected the burden of diarrhoea among children younger than 5 years: An analysis of the Global Burden of Disease Study 2019. The Lancet Infectious Diseases, 22(2), 255\u0026ndash;272. https://doi.org/10.1016/S1473-3099(21)00531-8\u003c/li\u003e\n\u003cli\u003eGBD 2019 Risk Factors Collaborators. (2020). Global burden of 87 risk factors in 204 countries and territories, 1990\u0026ndash;2019: A systematic analysis for the Global Burden of Disease Study 2019. The Lancet, 396(10258), 1223\u0026ndash;1249. https://doi.org/10.1016/S0140-6736(20)30752-2\u003c/li\u003e\n\u003cli\u003eKalayou, M. H., Tesfay, F. H., Nigusse, A. A., \u0026amp; Gebregziabher, D. (2024). Empowering child health: Harnessing machine learning to predict acute respiratory infections in Ethiopian under-fives using demographic and health survey insights. BMC Infectious Diseases, 24, 112. https://doi.org/10.1186/s12879-024-09022-7\u003c/li\u003e\n\u003cli\u003eKeilwagen, J., Grosse, I., \u0026amp; Grau, J. (2014). Area under precision-recall curves for weighted and unweighted data. PLOS ONE, 9(3), e92209. https://doi.org/10.1371/journal.pone.0092209\u003c/li\u003e\n\u003cli\u003eKenya National Bureau of Statistics (KNBS), Ministry of Health Kenya, \u0026amp; ICF. (2023). Kenya Demographic and Health Survey 2022. KNBS, Ministry of Health, and ICF. https://dhsprogram.com/pubs/pdf/FR368/FR368.pdf\u003c/li\u003e\n\u003cli\u003eKuhn, M. (2008). Building predictive models in R using the caret package. Journal of Statistical Software, 28(5), 1\u0026ndash;26. https://doi.org/10.18637/jss.v028.i05\u003c/li\u003e\n\u003cli\u003eLundberg, S. M., \u0026amp; Lee, S. I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (Vol. 30). Curran Associates. https://proceedings.neurips.cc/paper/2017/hash/8a20a8621978632d76c43dfd28b67767-Abstract.html\u003c/li\u003e\n\u003cli\u003eMayer, M., Watson, B., \u0026amp; Redelmeier, D. (2023). shapviz: SHAP visualizations. R package version 0.9.3. https://cran.r-project.org/package=shapviz\u003c/li\u003e\n\u003cli\u003eNaeher, L. P., Brauer, M., Lipsett, M., Zelikoff, J. T., Simpson, C. D., Koenig, J. Q., \u0026amp; Smith, K. R. (2020). Woodsmoke health effects: A review. Inhalation Toxicology, 19(1), 67\u0026ndash;106. https://doi.org/10.1080/08958370600985875\u003c/li\u003e\n\u003cli\u003eRobin, X., Turck, N., Hainard, A., Tiberti, N., Lisacek, F., Sanchez, J. C., \u0026amp; M\u0026uuml;ller, M. (2011). pROC: An open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinformatics, 12, 77. https://doi.org/10.1186/1471-2105-12-77\u003c/li\u003e\n\u003cli\u003eUNICEF. (2023). Pneumonia. UNICEF Data: Monitoring the Situation of Children and Women. https://data.unicef.org/topic/child-health/pneumonia/\u003c/li\u003e\n\u003cli\u003eVictora, C. G., Christian, P., Vidaletti, L. P., Garber, G., Bhutta, Z. A., \u0026amp; Black, R. E. (2021). Revisiting maternal and child undernutrition in low-income and middle-income countries: Variable progress towards Sustainable Development Goals. The Lancet, 397(10282), 1388\u0026ndash;1399. https://doi.org/10.1016/S0140-6736(21)00394-9\u003c/li\u003e\n\u003cli\u003eWorld Health Organization. (2022). Pneumonia in children. WHO Fact Sheets. https://www.who.int/news-room/fact-sheets/detail/pneumonia\u003c/li\u003e\n\u003cli\u003eYehuala, S., Bante, A., Dires, A., \u0026amp; Alemu, A. (2024). Exploring machine learning algorithms to predict acute respiratory tract infection and identify its determinants among children under five in Sub-Saharan Africa. Frontiers in Pediatrics, 12, 1363088. https://doi.org/10.3389/fped.2024.1363088\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"acute respiratory infection, machine learning, random forest, SHAP, environmental determinants, children under five, Kenya, DHS","lastPublishedDoi":"10.21203/rs.3.rs-9290306/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9290306/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e. Acute respiratory infection (ARI) remains a leading cause of morbidity and mortality among children under five years in sub-Saharan Africa. Environmental exposures such as household air pollution from solid cooking fuels, poor-quality housing, and inadequate water and sanitation are plausible contributors to ARI risk, yet no study has systematically applied machine learning to identify and rank these environmental determinants among Kenyan children using nationally representative survey data.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e. Data from the 2022 Kenya Demographic and Health Survey (KDHS) were analyzed. The final sample comprised 18,703 children under five years. ARI was defined as cough with rapid or difficult breathing in the two preceding weeks. Six machine learning algorithms were implemented: logistic regression, decision tree, random forest, XGBoost, gradient boosting, and an ensemble model (random forest combined with XGBoost). Recursive Feature Elimination (RFE) was used for feature selection across 22 candidate variables spanning child, maternal, and environmental domains. Class imbalance was addressed using the Synthetic Minority Oversampling Technique (SMOTE). Hyperparameters were tuned via GridSearchCV with 5-fold cross-validation. Model interpretability was assessed using SHapley Additive exPlanations (SHAP) analysis.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e. ARI prevalence was 16.8% (n = 3,143). Random Forest achieved the highest accuracy (83.29%) and F1-score (90.79%). Across models, AUC-ROC values ranged from 0.50 to 0.65, with the ensemble and logistic regression reaching the highest discrimination (AUC-ROC = 0.65). SHAP analysis identified number of living children, concurrent diarrhea, and household wealth as the three strongest overall predictors. Among environmental factors specifically, housing quality index (importance score 151.58) ranked highest, followed by improved water source (125.30), urban residence (94.80), improved toilet facility (83.73), and clean cooking fuel (58.03).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion\u003c/strong\u003e. Machine learning models can identify modifiable environmental determinants of ARI in Kenyan children with reasonable predictive accuracy. Housing quality, water access, and sanitation consistently emerged as more important environmental risk factors than cooking fuel type. These findings have direct implications for prioritizing environmental health interventions in Kenya’s child health programs.\u003c/p\u003e","manuscriptTitle":"Machine Learning Prediction of Acute Respiratory Infection Risk in Kenyan Children Under Five: A Focus on Environmental Determinants","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-28 08:59:47","doi":"10.21203/rs.3.rs-9290306/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7a48aae2-1eb8-4687-a5c6-9cffe3c08953","owner":[],"postedDate":"April 28th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-28T08:59:47+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-28 08:59:47","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9290306","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9290306","identity":"rs-9290306","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.