Forecasting HIV/AIDS Incidence in Ghana: A Retrospective Observational Study Using Ensemble Machine Learning Models 

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Background: In Ghana, the precise forecasting of Human Immunodeficiency Virus (HIV) infection and Acquired Immunodeficiency Syndrome (AIDS) is essential for public health strategies due to the intricate socio-structural factors that affect the transmission patterns of the virus. Public health planning becomes challenging because conventional linear statistical models do not take into account disjointed or multifactorial data. Methods: This used retrospective observational data covering the ten administrative regions of Ghana from 2000 to 2022. Four machine learning algorithms (Random Forest, XGBoost, Ridge Regression, and Support Vector Regression) were applied to forecast HIV/AIDS incidence. The dataset incorporated epidemiological trends, demographic profiles, and healthcare infrastructure indicators. Preprocessing steps included KNN imputation for missing healthcare infrastructure values and winsorization of disease incidence variables to reduce outlier bias. Results: Ridge Regression and Support Vector Regression performed poorly compared to XGBoost and Random Forest, with correlation coefficients above 0.98, indicating high predictive accuracy. HIV incidence in Ghana was foretasted to stabilize at 231 cases per 100,000 individuals by 2030. The SHAP (SHapley Additive exPlanations) analysis revealed that HIV awareness, access to antiretroviral therapy, poverty rates, and access to education were significant factors influencing incidence trends. Discussion: Ensemble machine learning models yielded more reliable predictions than conventional linear models. The predicted incidence plateau indicates that current intervention strategies may not reach the national and global reduction targets, thereby underscoring the need for more vigorous public health initiatives. Data limitations restrict real-time predictions and necessitate ongoing enhancements to the data infrastructure. Conclusion: This study demonstrates that ensemble machine learning is a viable and valuable tool for predicting HIV incidence rates in Ghana, providing reliable results to guide public health decision making and resource management. However, its effectiveness may be influenced by data quality limitations and contextual complexities, underscoring the need for expanded data-sharing infrastructure, cautious interpretation and continuous model validation in low-resource settings, as emphasized by ongoing calls for transparency in ML epidemiology.
Full text 192,150 characters · extracted from preprint-html · click to expand
Forecasting HIV/AIDS Incidence in Ghana: A Retrospective Observational Study Using Ensemble Machine Learning Models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Forecasting HIV/AIDS Incidence in Ghana: A Retrospective Observational Study Using Ensemble Machine Learning Models Valentine Golden Ghanem This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6639193/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 13 You are reading this latest preprint version Abstract Background: In Ghana, the precise forecasting of Human Immunodeficiency Virus (HIV) infection and Acquired Immunodeficiency Syndrome (AIDS) is essential for public health strategies due to the intricate socio-structural factors that affect the transmission patterns of the virus. Public health planning becomes challenging because conventional linear statistical models do not take into account disjointed or multifactorial data. Methods: This used retrospective observational data covering the ten administrative regions of Ghana from 2000 to 2022. Four machine learning algorithms (Random Forest, XGBoost, Ridge Regression, and Support Vector Regression) were applied to forecast HIV/AIDS incidence. The dataset incorporated epidemiological trends, demographic profiles, and healthcare infrastructure indicators. Preprocessing steps included KNN imputation for missing healthcare infrastructure values and winsorization of disease incidence variables to reduce outlier bias. Results: Ridge Regression and Support Vector Regression performed poorly compared to XGBoost and Random Forest, with correlation coefficients above 0.98, indicating high predictive accuracy. HIV incidence in Ghana was foretasted to stabilize at 231 cases per 100,000 individuals by 2030. The SHAP (SHapley Additive exPlanations) analysis revealed that HIV awareness, access to antiretroviral therapy, poverty rates, and access to education were significant factors influencing incidence trends. Discussion: Ensemble machine learning models yielded more reliable predictions than conventional linear models. The predicted incidence plateau indicates that current intervention strategies may not reach the national and global reduction targets, thereby underscoring the need for more vigorous public health initiatives. Data limitations restrict real-time predictions and necessitate ongoing enhancements to the data infrastructure. Conclusion: This study demonstrates that ensemble machine learning is a viable and valuable tool for predicting HIV incidence rates in Ghana, providing reliable results to guide public health decision making and resource management. However, its effectiveness may be influenced by data quality limitations and contextual complexities, underscoring the need for expanded data-sharing infrastructure, cautious interpretation and continuous model validation in low-resource settings, as emphasized by ongoing calls for transparency in ML epidemiology. HIV/AIDS Ghana Machine Learning Forecasting Public Health Epidemiology Ensemble Models XGBoost Random Forest SHAP Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 1. INTRODUCTION The Human Immunodeficiency Virus (HIV) and Acquired Immunodeficiency Syndrome (AIDS) pose worldwide health challenges which need multifaceted approaches to combat. These strategies focus on precisely forecasting public health HIV/AIDS interventions, particularly in sub-Saharan Africa where healthcare resources are limited and the rate of transmission is high. Like most of its regional counterparts, Ghana is fighting the HIV/AIDS pandemic, which has a national prevalence of about 2%. Women have a higher prevalence than men (2.5% and 1.1% respectively) [ 1 ]. The incidence of Ghana’s HIV/AIDS epidemic is greatly affected by spatial and socio-behavioral inequities. Greater HIV awareness is linked to higher rates of HIV testing, reduced stigma, and higher levels of social discrimination. In contrast, illiteracy and poverty is associated with greater transmission risks and higher social discrimination [ 2 , 3 ]. Moreover, rural areas have greater incidence because of migration patterns and limited access to health services. Although urban areas are often seen as transmission hotspots, suburban areas tend to have lower incidence rates [ 4 , 5 ]. Gender and marital status have been shown to be important factors that influence the dynamics of HIV transmission [ 6 , 7 ]. Younger women’s (15–39 years) testing rates and prevalence were both higher, indicating the impact of gender disparity and sociocultural dynamics on the epidemic [ 6 , 7 ]. Including such socio-spatial differences into disease modeling, adds relevance to the context of the predictive insights regarding targeted intervention strategies. With these considerations in mind, public health policies, as well as the sociological frameworks relied on for designing interventions, must more accurately calibrate regionally and demographically sensitive policies tailored to the cross-cutting socio-spatial heterogeneity, as examined through advanced statistical models that integrate multidimensional analytic frameworks to enhance predictive relevance in interventions. Ensemble machine learning models are utilized in public health, especially in the epidemiological surveillance of infectious diseases, to enhance forecasting accuracy and trends. These models can adapt to dynamic patterns in data from health surveillance systems. Random Forest and XGboost outperform traditional statistical models, such as Autoregressive Integrated Moving Average (ARIMA). Ensemble methods exceeded traditional ones for predicting sepsis-related mortality and the severity of Corona Virus Disease 2019 (COVID-19) pandemic [ 8 , 9 ]. Explainable Artificial Intelligence (xAI) frameworks, such as SHapley Additive exPlanations (SHAP), enhance the interpretability and transparency of ensemble models by identifying critical features and explaining their importance, thus promoting trust in the outputs by ML models [ 10 ]. The multifactorial transmission dynamics of infectious diseases like HIV, which is inherently complex, requires the development and integration of sophisticated explainable models to enhance public health forecasting. Despite the growing curiosity about applying Artificial Intelligence (AI) and Machine Learning (ML) to public health interventions, the persistent challenges of poor data quality continue to impede model effectiveness. These issues comprise of inadequate migration and behavioral data, under-reporting incidence, and sparse surveillance coverage. Moreover, the absence of explainability in ML models generates distrust among stakeholders which discourages their adoption in public health.When health professionals are trained in the use of ensemble ML models, their confidence improves, which drives the wide acceptance and integration of this technology into healthcare. Through the analysis of demographic and complex behavioral data, models like the Random Forest and XGboost excel in HIV risk prediction. This is supported by a recent study that reports the high accuracy of ensemble ML models in identifying high-risk individuals, notably among target groups such as men who have sex with men (MSM) [ 11 ]. Nevertheless, by forecasting clinic attendance and testing behavior after reminder messages, ML improves service delivery during HIV epidemic by optimizing patent care and resource allocation [ 12 ]. Model optimization enhances reliability, generalizability, and cost-effectiveness, which are crucial in resource-limited settings [ 13 ]. Forecasting HIV incidence remains an underexplored area in current research. Most machine learning applications priotized classifying individuals by risk level or predicting HIV serostatus, with limited focus on long-term incidence forecasting—a critical component for effective public health planning and resource allocation. Recent advances suggest that hybrid models, particularly those combining Extreme Gradient Boosting (XGBoost) with Long Short-Term Memory (LSTM) networks, demonstrate superior performance in disease forecasting tasks. These models have outperformed traditional methods in predicting infectious disease trends, including COVID-19 outbreaks, especially in settings with sparse historical data [ 14 , 15 ]. The performance of this model depends on the complete datasets, which are often scarce in African settings. Traditional time-series models like ARIMA and exponential smoothing are commonly used in lower- and middle-income countries (LIMCs) , such as Ghana. These statistical models often do not fully capture complex nonlinear trends in HIV transmission [ 16 , 17 ]. This leads to poor predictability, particularly in settings that involve multiple disease incidence drivers. Ensemble models achieve high efficiency in disease forecasting and surveillance by integrating several data sources to achieve a better performance. LMIC’s with infrastructural and data challenges such as Ghana can benefit immensely from using an ensemble model when SHAP and counterfactual simulations are used to enhance transparency and interpretability. This illustrates the role of variables such as ART (Antiretroviral Therapy) coverage and awareness in shaping HIV incidence and transmission, and supports evidence-based conclusions in public health, thereby fostering trust among stakeholders. The global health objective, as well as Ghana's broader future HIV response , will depend heavily on ML-driven forecasting, which will guide strategic planning at both the national and international levels. To ensure maximum impact, it is imperative to improve data quality and model interpretability, particularly when designing and using ML models in resource-constrained environments. Constant monitoring and hyperparameter adjustments of these models are needed in real-world applications given the evolving nature of variables in the HIV epidemic. Using publicly available data from multiple domains, this study seeks to assess the predictive accuracy of four supervised machine-learning algorithms (Random Forest, XGBoost, Ridge Regression, and SVR) for HIV/AIDS incidence trends in Ghana from 2000 to 2022. 2. METHODS 2.1 Study Design and Data Sources This observational study retrospectively examined HIV/AIDS incidence trends using a unified dataset consisting of annual HIV/AIDS case data from Ghana's ten administrative regions from 2000 to 2022. Information was compiled from several trustworthy public sources, including the Ghana Health Service, Ghana AIDS Commission, Ministry of Health, Ghana Statistical Service, UNAIDS, World Bank, and Humanitarian Data Exchange. The dataset includes various variables, such as demographic, epidemiologic al, health system, socio-behavioral, and geospatial variables, together with detailed information on data domains and sources (Table 1 ). To preserve temporal consistency with historical records, a ten-region format was selected, while post-2018 administrative boundary changes were omitted. Furthermore, the dataset incorporated socio-behavioral and spatial metrics, such as HIV awareness indices, gender-based vulnerability markers, and regional health service accessibility, to preserve the predictive value and contextual relevance of ML models. These features reflect the structural and behavioral determinants of HIV transmission in Ghana, enabling models to account for the complex, multilayered nature of the epidemic [ 2 , 3 ]. No participants were enrolled in this study and no interviews were conducted. The analysis exclusively draws on aggregate regional data from Ghana. Therefore, standard items regarding participant eligibility, matching, or calculation of sample size were not applied. 2.2 Data Preparation and Feature Engineering A comprehensive data quality assessment pipeline encompassing metadata screening, indicator validation, and outlier identification via Forest Isolation and Winsorization at the 1st and 99th percentiles was used to minimize noise and outliers. Missing data were estimated using the KNN algorithm. Temporal dependencies and trend smoothing were achieved using feature-engineering techniques, including the creation of lag variables and moving averages. Indices combining HIV awareness and regional socioeconomic factors have been developed to encapsulate the various factors that affect HIV incidence rates. All features were normalized using the MinMax method to allow comparable scales (Tables 2 – 4 ). To reduce measurement and selection bias, the data were validated using external metadata, including outlier detection and imputation, which are highly effective on their own. Enhanced inclusivity strengthened the sources used . 2.3 Model Selection Theory and Development The choice of model was based on the theoretical and empirical advantages of various supervised machine-learning techniques to identify intricate, possibly nonlinear patterns within the data. A comparative evaluation of the four ML models (Table 5 ). Random Forest (RF) and XGBoost are ensemble tree-based methods that are renowned for their robust capacity to approximate nonlinear functions and capture variable interactions. Ridge Regression provides a form of regularized linear modeling for evaluating baseline linear relationships. SVR ( Support Vector Regression) functions as a kernel-based approach that enables flexible nonlinear regression by using kernel transformations. The selection rationale strikes a strategic balance between model complexity, interpretability, predictive capability, and computational efficiency, as dictated by the established best practices in model selection theory, specifically regarding model traits and tuning parameters (Table 6 ). 2.4 Model Training, Hyperparameter Tuning, and Evaluation The dataset was divided into two subsets (80/20 split): a training set that spanned the years 2000 to 2017 and a testing set that covered 2018 to 2022, with the chronological order maintained to avoid data contamination and to mimic real-world forecasting situations. Grid and Random Search methods were employed within the cross-validation folds of the training data to determine the optimal model settings through hyperparameter optimization. The performance of the model was assessed on the test dataset using various evaluation criteria including the coefficient of determination (R²), root mean squared error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE). Ensemble models, notably XGBoost and Random Forest, showed higher effectiveness with R² values surpassing 0.98 and low prediction errors (Tables 6 – 8 ; Figs. 1 and 2 ). 2.4.1 Sensitivity Analysis Sensitivity analyses were not formally conducted . Nonetheless, model robustness was evaluated across ensemble, linear, and kernel-based methods, as well as through interpretability tools such as SHAP and counterfactuals, to confirm the consistency of recorded values. 2.5 Model Interpretability: Justification for SHAP and Counterfactuals Interpreting complex ensemble machine-learning models is essential for translating predictions into practical insights that can be used in public health initiatives. Feature interactions, including subgroup behaviors by region, were analyzed using th e SHAP dependence plots and counterfactual simulations.SHAP values offer a theoretically sound game theory approach for both local and global interpretations, allowing the model output to be broken down according to feature influence while accounting for feature interactions and correlations. In the context of Ghana, this interpretability is vital for transforming model outputs into actionable public health strategies—by revealing, for instance, how marginal increases in ART coverage or HIV awareness disproportionately affect incidence in specific regions. This enables policymakers to prioritize modifiable drivers in high-burden districts, thereby improving the precision and impact of HIV interventions [ 18 ]. Counterfactual simulation methods were used to investigate hypothetical situations, examining how changes in key modifiable factors (e.g., HIV awareness and condom use) might impact projected HIV incidence patterns. This approach enables an actionable inference by pinpointing the intervention points of influence. Consequently, the combination of SHAP and counterfactual analysis surpasses the traditional feature importance, bringing model transparency in line with the current interpretability benchmarks in health analytics ( Table 8 ; Figs. 3 – 6 ). Although machine learning models do not adjust for confounding in a traditional statistical sense, SHAP values illuminate the influence of key predictors and highlight features that buffer key confounding relationships. 2.6 Statistical and Software Tools All analyses were performed in Python 3.10 using libraries such as pandas, numpy, and scikit-learn for data handling and model development, and SHAP for interpretability analysis. Visualization libraries such as matplotlib and seaborn enabled the development of interpretive plots to support model evaluation and insight generation. 3. RESULTS 3.1 Predictive Performance and Model Comparison The predictive abilities of four supervised machine learning algorithms, Random Forest, XGBoost, Ridge Regression, and SVR, were thoroughly examined on a test dataset reserved prior to model training, covering the period from 2018 to 2022. The dataset comprised of a harmonized panel of observations spanning 2000 to 2022, thereby maintaining temporal consistency and illustrating significant epidemiological trends in the incidence of HIV/AIDS in Ghana. Given that this study was conducted using aggregated data, aspects concerning individual participants’ identification, eligibility, or attrition did not apply. Table 2 shows missing value percentages across all variables before KNN imputation was performed. By the last step of the preprocessing pipeline, all variables had 0% missing data, which means that the provided steps for data cleaning were sufficient to resolve issues of missing data prior to any form of advanced imputation techniques. The performance assessment employed multiple standard metrics, like the coefficient of determination (R ²), Root Mean Square error (RMSE), Mean Absolute Error (MAE), and Mean Absolute Percentage Error (MAPE), thereby facilitating a comprehensive evaluation of the error magnitudes and relative fit quality. Ensemble models, including Random Forest and XGBoost, showed enhanced predictive abilities, as outlined in Table 9. The Random Forest model produced the best results, with an R² of 0.9821, RMSE of 6.39, MAE of 4.64, and MAPE of 2.56%, and eclipsing XGBoost with an R² of 0.9813, RMSE of 6.54, MAE of 4.77, and MAPE of 2.64%. Unlike the baseline models, Ridge Regression and SVR produced significantly lower performance outcomes, with R² values of 0.9640 and 0.9686, respectively, indicating their inability to effectively model the complex nonlinear relationships observed in HIV incidence patterns (Table 7). Cross-validation and hyperparameter tuning led to XGBoost's slight technical edge in long-term forecasting from 2023 to 2030, with an RMSE of 0.47, outperforming Random Forest's RMSE of 0.55 (Table 10). According to the data, the ensemble forecasts predicted stabilization of HIV incidence at a rate of approximately 231 cases per 100,000 population by 2030, as illustrated in Fig. 7. Linear models exhibited contradictory trends, with Ridge Regression showing an upward trend in incidence and SVR forecast showing a slight decline, although with a reduced accuracy. The model-averaged forecast indicates a gradual increase, highlighting the need for more aggressive corrective measures (Table 10; Figs. 1, 2, and 7). The models used were not regression-based; therefore, the confidence intervals and p-values for the effect estimation were not reported. The performance metrics reflected the predictive accuracy through the R², RMSE, MAE, and MAPE. 3.2 Model Interpretability Through SHAP and Counterfactual Analysis Some techniques have been applied to increase the accuracy of models and to derive meaningful public health insights. The SHapley Additive exPlanations (SHAP) method was used to calculate the specific impact of each input variable on the model predictions at both overall and individual levels. This game-theoretic method enables a thorough breakdown of the prediction effects, while considering the relationships between different features. HIV awareness, ART coverage, and education access were among the most influential drivers in forecasting HIV incidence, as shown in the SHAP summary plot (Fig. 3). This finding is consistent with regional HIV prediction efforts in sub-Saharan Africa, where ensemble models identified these factors as critical for targeting screening and retention strategies [18, 19].Furthermore, the SHAP analysis was influenced by spatial variables such as access to health services across rural-urban contexts, as demonstrated in previous studies [4, 5]. This indicates that due to population and migration dynamics, urban cities have higher incidence rates, while rural resource-poor communities suffer from hidden transmission. Additional analysis in Fig. 8 using Partial Dependence Plots revealed complex interplay of socioeconomic factors alongside HIV incidence, which in addition to ART coverage and HIV awareness involved nonlinear inverse relationships. Counterfactual simulation methods were applied to devise hypothetical situations where certain key variables were altered to explore their effect on HIV trends and outcomes. This possibility provides a space for causal reasoning by demonstrating what would happen with the use of specific interventions, thus linking predictive modeling with policy-driven decision-making processes. Mulit-dimensional SHAP and counterfactual analyses directly enhance the clarity and rigor of trust placed in complex ensemble models, while upholding the principles of designed-for-purpose interpretability within the epidemiological machine-learning framework. 3.3 Conclusion The ensemble machine learning models, especially XGBoost, achieved optimal performance in the predictive modeling of HIV/AIDS cases within the Ghanaian epidemiological context due to their parameter tuning capabilities. The performance of the Random Forest algorithm was also fairly accurate. The study further illustrated the ensemble tree methods’ superiority over the older, linear and kernel-based models which are inadequate in dealing with intricate data dependencies. SHAP and counterfactual interpretation techniques are critical for model explanation to enhance evidence-based public health responses as well as proactive planning and targeted interventions. These conclusions endorse the application of complex interpretive frameworks of ensemble machine learning techniques to sharpen epidemiological surveillance, optimize the allocation of resources, and enhance planning and responsiveness to the HIV epidemic in Ghana. 4. DISCUSSION The study showed that ensemble machine learning (ML) models, particularly Random Forest and XGBoost, achieved the highest predictive accuracy (R² > 0.98) in forecasting HIV/AIDS incidence in Ghana over a period of two decades. Ensemble ML models outperform Support Vector Regression (SVM) and Ridge Regression due to their inherent capabilities to capture nonlinear and multivariate time-dependent relationships. Such characteristics make them effective for forecasting infectious diseases, especially when socio-behavioral, health, and demographic factors overlap. These findings align with observations that documented similar results for other infectious diseases [ 8 , 9 , 10 ]. This study stands out because it attempts to explain the effects of major predictors like ART coverage and HIV awareness using SHAP and counterfactual simulations, thus bridging the gap between predictive analytics and actionable public health policy. These methods dispel any uncertainty relating to the “black box” nature of ML models which, do not integrate counterfactual simulations, nor SHAP analysis, when forecasting trends. This research conclusion supports epidemiological evidence indicating that increasing ART and treatment coverage rates inadvertently leads to viral load suppression, thereby reducing HIV transmission and incidence rates [ 9 , 18 , 20 , 21 ] The identification of alterable parameters, like awareness of HIV and ART coverage, as important determinants aligns with the epidemiological understanding that HIV transmission is influenced by the behavior, structure, and healthcare delivery systems of a region. The importance of socioeconomic factors and health system infrastructure is connect ed to broader evidence that call s for holistic HIV prevention. Strategies that integrate biomedical, behavioral, and social actions [ 1 , 18 , 21 ]. Such gaps in targeted public health interventions strengthen the existing theories to enhance awareness initiatives and improve the range and accessibility of ART services. The study results reveal an alarming trend signaling a realistic forecast: stabilization of HIV incidence is projected to reach 231 cases per 100,000 population by 2030, which is drastically lower than the national and global targets set for meaningful reduction in new infections. This aligns with the broader downward trends forecasted in the Global Burden of Disease Study 2021, which provides country-level HIV incidence projections to 2050 [ 22 ]. It exemplifies the persistent challenge of HIV epidemic control in Ghana and emphasizes the crucial need to shift approaches to prevention intensification beyond the current trajectories. These models provide tools designed for public health planning and fundamentally quantify the impact of scaling up targeted resources and implementing interventions. This finding is consistent with the broader trends observed in sub-Saharan Africa, where renewed attempts to reduce HIV incidence have led to stagnation in new cases, most likely as a result of socio-behavioral determinants, systematic health resource bottlenecks and unequal resource distribution s [ 23 , 24 ]. These results have several direct implications for public health policy in Ghana. First, the high predictive value of ART coverage and HIV awareness suggests that intervention efforts should prioritize scaling up ART accessibility and education programs, particularly in rural and underserved regions. This aligns with findings in Mozambique and Nigeria where ML was used to identify ART clients at high risk of treatment interruption, effectively guiding outreach efforts [ 19 ]. Second, targeted outreach campaigns focused on condom use and stigma reduction in high-burden districts, especially in Greater Accra and Ashanti regions, are warranted. SHAP analysis revealed non-linear relationships suggesting diminishing returns in some interventions—echoing the importance of precision public health over blanket strategies [ 18 ]. Regional health plans integrating ML-based risk stratification may optimize resource allocation and increase intervention efficiency. Another strength of this study is the use of multi-source datasets that integrate epidemiological, demographical, behavioral, and infrastructural indicators from 2000 to 2022, providing deeper insights via enhanced data-driven analysis than a single-domain dataset. The random forest and XGboost models consistently outperformed the other models in terms of accuracy across all regions, with ART coverage and HIV awareness being the most influential predictors. Persisting regional disparities and structural factors affect incidence trends, which are further influenced by socioeconomic factors, especially in the Greater Accra and Ashanti regions. The ecological design of this study and lack of res ou rces in LMIC’s including Ghana,imposes constraints on data quality, which includes incomplete reports, reporting delays, and residual bias. Further studies support this claim [ 23 , 25 ]. In China , traditional models, such as ARIMA and exponential smoothing, have been successfully applied because of the availability of high-frequency monthly incidence data [ 16 , 17 ]. These models are plagued by linearity and stationarity, which affect their ability to capture complex predictors of HIV transmission in LMICs. In contrast, ensemble models , such as XGBoost, are more adept at handling such complexities. Despite the strengths of this study, several limitations could impact the generalizability and robustness of its findings. First, national HIV datasets are subject to under-reporting, especially in rural districts, due to incomplete surveillance coverage, delayed reporting cycles, and limited behavioral indicators [ 12 , 25 ]. Second, the reliance on aggregate region-level data introduces risks of ecological fallacy, where individual-level causal relationships cannot be inferred [ 18 ]. Third, although imputation techniques were used, missing data may bias model outputs if the missingness is not completely at random,a well-recognized challenge in machine learning applications within public health settings [ 26 ]. Finally, data sparsity in sub-populations—such as key communities or mobile groups—limits the precision of risk predictions and impedes targeted interventions [ 27 ]. These limitations suggest a cautious interpretation of the results and emphasize the need for further model validation in prospective settings, ideally through formal sensitivity analyses and robustness testing methods such as perturbation, bootstrapping, or Monte Carlo simulations, in line with best practices for predictive modeling in public health research [ 28 ]. Moreover, while current models demonstrate performance over a retrospective period, their ability to suddenly predict shifts like the rapid scaling up of new prevention technologies during social behavior changes, or System Implosions and Pandemics raises concern. This is where real-time monitoring and continual adjustment of models becomes essential [ 13 , 14 ]. In an attempt to fight these issues, hybrid methods like XGBoost integrated with LSTM are effective because they include non-linear feature interaction as well as temporal features like movement and behavioral change [ 14 , 29 ]. In LMIC’s, however, these advanced models cannot be applied owing to resource limitations. Results from this research corroborate with several investigations validating the accuracy and reliability of ensemble machine-learning techniques which outperform individual models or single-method approaches, by aggregating forecasts from several base learners.Such uniformity has been observed across infectious diseases, including dengue and COVID-19 [ 30 , 31 ]. The gradient boosting models, particularly XGBoost, are highly accurate due to their adjustable hyperparameters and proven ability to model sophisticated interactions among features on a myriad of factors in theoretical and applied epidemiology [ 30 , 32 ]. In summary, this study illustrates the invaluable predictive power of the ensemble ML models in HIV incidence forecasting, while pinpointing relevant considerations for policymakers. The Random Forest and XGBoost models showed high accuracy and provided meaningful insights that are advantageous for public health planning in Ghana and other similar contexts. Nonetheless, these advantages depend on the existence of detailed real-time data, which remain a major problem in most sub-Saharan African health systems. Continuous funding and policy support are essential to enhance the efficacy of these ensemble models. Additionally, the following policies and measures can be implemented: Enhancing health data infrastructure : The availability of detailed, complete, and timely data is critical for developing and testing useful models. Keeping up with changing circumstances : Integrating emerging data on mobility and Pre-Exposure Prophylaxis(PrEP) uptake is vital to prevent unpredictable epidemiological shifts. Structural strategies : Supplementing biomedical approaches with strategies focused on education, economic support, equality, and inclusion is crucial [ 33 ]. Equity-focused resource allocation : Addressing specific areas and populations with the highest spatial and demographic risk levels as defined by spatial and demographic analytics is vital[ 34 ]. CONCLUSION This study demonstrates that ensemble machine learning models, particularly XGBoost and Random Forest, can accurately and precisely forecast HIV incidence trends in Ghana. Leveraging advanced interpretability techniques such as SHAP and counterfactual analysis, these models translate complex data patterns into actionable insights, effectively connecting technical findings with public health policy needs. Key predictors identified, notably ART coverage and HIV awareness, represent critical targets for intervention efforts. Integrating these findings into Ghana’s public health strategies can enhance disease surveillance and enable more efficient allocation of resources, supporting national and global objectives for HIV reduction even within the constraints of limited resources. Despite their high predictive performance, the utility of these ensemble models is inherently limited by data quality, ecological variability, and changing epidemiological dynamics. Predictive models carry unavoidable uncertainty and require ongoing validation, incorporation of real-time data, and adaptive retraining to maintain their relevance and accuracy in evolving contexts. Future research should explore hybrid modeling approaches that combine machine learning with mechanistic models to jointly capture statistical trends and underlying transmission dynamics. Real-time adaptive modeling frameworks would also be valuable for responding dynamically to shifts in HIV epidemiology. By embedding such integrated modeling tools within broader public health frameworks, Ghana and similar low-resource settings can better harness predictive analytics to inform targeted interventions and accelerate progress toward epidemic control. While these findings are most directly applicable to Ghana and comparable contexts with similar socio-epidemiological profiles, caution should be exercised when generalizing to settings with substantially different health systems or demographics. Abbreviations AI: Artificial Intelligence AIDS: Acquired Immunodeficiency Syndrome ARIMA: AutoRegressive Integrated Moving Average ART: Antiretroviral Therapy CC-BY: Creative Commons Attribution (license type) COVID-19: Corona Virus Disease 2019 GAC: Ghana AIDS Commission GHS: Ghana Health Service GSS: Ghana Statistical Service HDX: Humanitarian Data Exchange HIV: Human Immunodeficiency Virus KNN: K-Nearest Neighbors LIMCs: Lower and Middle Income Countries LSTM: Long Short-Term Memory MAE: Mean Absolute Error MAPE: Mean Absolute Percentage Error ML: Machine Learning MoH: Ministry of Health MSM: Men who have Sex with Men ORCID: Open Researcher and Contributor ID PrEp: Pre - Exposure Prophylaxis R²: Coefficient of Determination RF: Random Forest RMSE: Root Mean Squared Error SHAP: SHapley Additive exPlanations SVR: Support Vector Regression UNAIDS: Joint United Nations Programme on HIV/AIDS WHO: World Health Organization xAI: Explainable Artificial Intelligence XGBoost: Extreme Gradient Boosting Declarations Ethics approval and consent to participate: Not applicable. This study used only publicly available aggregated data and did not involve human participants, human data, or directly collected human tissue. Consent for publication: Not applicable. This manuscript did not contain any individual data. Competing interests: The author have no competing interests to declare. Funding: The author declares no funding for this study. Author Contribution VG: Conceptualization, Methodology, Data curation, formal analysis, software development, Validation, Visualization, Writing – original draft, and writing – review and editing. Acknowledgements: The author s wish to thank the Ghana Health Service, Ghana AIDS Commission, Ghana Statistical Service, and the Humanitarian Data Exchange (HDX) platform for providing access to critical epidemiological and demographic datasets that supported the development of this research. Appreciation has also been extended to contributors to open geospatial data via GeoBoundaries. This work was dedicated to the memory of my beloved sister, Imelda Farr, whose encouragement inspired my pursuit of this academic endeavor. Authors' information : VG is a Ghana-based biomedical scientist at the Cocoa Clinic (Ghana Cocoa Board’s Medical Department) in Accra. He holds an MSc in Data Science from the University of East London and is currently completing an MSc in Public Health at the University of Suffolk through distance learning. His work focuses on infectious disease modeling, epidemiologic al forecasting, and the application of machine learning to public health policies. Data Availability The full dataset, cleaned indicators, spatial files, and model code supporting this study are publicly available at the Zenodo repository under license CC-BY 4.0. Available at: https://doi.org/10.5281/zenodo.15292209. References Ba DM, Ssentongo P, Sznajder KK. Prevalence, behavioral and socioeconomic factors associated with human immunodeficiency virus in Ghana: a population-based cross-sectional study. J Glob Health Rep. 2019;3:e2019092. https://doi.org/10.29392/joghr.3.e2019092 . Nutor J, Duah H, Duodu P, et al. Geographical variations and factors associated with recent HIV testing prevalence in Ghana. BMJ Open. 2021;11:e045458. https://doi.org/10.1136/bmjopen-2020-045458 . Melkam M, Fente BM. Multilevel analysis of discrimination of people living with HIV/AIDS and associated factors in Ghana. Front Public Health. 2024;12:1379487. https://doi.org/10.3389/fpubh.2024.1379487 . Wand H, Morris N, Reddy T. Temporal and spatial monitoring of HIV prevalence and incidence rates using geospatial models. Spat Spatiotemporal Epidemiol. 2021;37:100413. https://doi.org/10.1016/j.sste.2021.100413 . Dias B, Rodrigues T, Botelho E, Oliveira M, Feijão A, Polaro S. Integrative review on the incidence of HIV infection and its socio-spatial determinants. Rev Bras Enferm. 2021;74(2):e20200905. https://doi.org/10.1590/0034-7167-2020-0905 . Sia D, Onadja Y, Hajizadeh M, et al. What explains gender inequalities in HIV/AIDS prevalence in sub-Saharan Africa? BMC Public Health. 2016;16:1136. https://doi.org/10.1186/s12889-016-3783-5 . Adetokunboh O, Are E. Spatial distribution and determinants of HIV high burden in the Southern African sub-region. PLoS ONE. 2024;19:e0301850. https://doi.org/10.1371/journal.pone.0301850 . Hou N, Li M, He L, Xie B, Wang L, Zhang R, et al. Predicting 30-days mortality for MIMIC-III patients with sepsis-3: a machine learning approach using XGBoost. J Transl Med. 2020;18(1):NA. https://doi.org/10.1186/s12967-020-02620-5 . Hu CA, Chen CM, Fang YC, Liang SJ, Wang HC, Fang WF, et al. Using a machine learning approach to predict mortality in critically ill influenza patients: a cross-sectional retrospective multicentre study in Taiwan. BMJ Open. 2020;10(2):e033898. https://doi.org/10.1136/bmjopen-2019-033898 . Hong W, Zhou X, Jin S, Lu Y, Pan J, Lin Q, et al. A comparison of XGBoost, Random Forest, and nomograph for the prediction of disease severity in patients with COVID-19 pneumonia: implications of cytokine and immune cell profile. Front Cell Infect Microbiol. 2022;12:819267. https://doi.org/10.3389/fcimb.2022.819267 . Ji X, Tang Z, Osborne SR, Van Nguyen TP, Mullens AB, Dean JA, et al. STI/HIV risk prediction model development—A novel use of public data to forecast STIs/HIV risk for men who have sex with men. Front Public Health. 2025;12:1511689. https://doi.org/10.3389/fpubh.2024.1511689 . Xu X, Fairley CK, Chow EPF, Lee D, Aung ET, Zhang L, et al. Using machine learning approaches to predict timely clinic attendance and the uptake of HIV/STI testing post clinic reminder messages. Sci Rep. 2022;12(1):12033. https://doi.org/10.1038/s41598-022-12033-7 . Xu D, Chan WH, Haron H. Enhancing infectious disease prediction model selection with multi-objective optimization: an empirical study. PeerJ Comput Sci. 2024;10:e2217. https://doi.org/10.7717/peerj-cs.2217 . Guo K, Shen C, Zhou X, et al. Traffic data-empowered XGBoost-LSTM framework for infectious disease prediction. IEEE Trans Intell Transp Syst. 2022;PP(99):1–12. https://doi.org/10.1109/TITS.2022.3172206 . Lucas B, Vahedi B. A spatiotemporal machine learning approach to forecasting COVID-19 incidence at the county level in the United States. arXiv . Preprint. 2021. https://doi.org/10.48550/arXiv.2109.12094 . Tang D, Jin Y, Hu X, Lin D, Kapar A, Wang Y, et al. Study on the prediction performance of AIDS monthly incidence in Xinjiang based on time series and deep learning models. BMC Public Health. 2025;25(1):21982. https://doi.org/10.1186/s12889-025-21982-3 . Xu B, Li J, Wang M. Epidemiological and time series analysis on the incidence and death of AIDS and HIV in China. BMC Public Health. 2020;20(1):9977. https://doi.org/10.1186/s12889-020-09977-8 . Mutai CK, McSharry PE, Ngaruye I, Musabanganji E. Use of machine learning techniques to identify HIV predictors for screening in sub-Saharan Africa. BMC Med Res Methodol. 2021;21(1):268. https://doi.org/10.1186/s12874-021-01346-2 . Stockman J, Friedman J, Sundberg J, Harris E, Bailey L. Predictive analytics using machine learning to identify ART clients at health system level at greatest risk of treatment interruption in Mozambique and Nigeria. J Acquir Immune Defic Syndr. 2022;90(2):154–60. https://pubmed.ncbi.nlm.nih.gov/35262514/ . Ayamah R, Awuitor GK, Puotier Z. Mathematically modeling the spread of HIV/AIDS infection after the introduction of antiretroviral therapy in Ghana. Am J Comput Eng. 2020;3(1):24–39. https://doi.org/10.47672/ajce.602 . Fieggen J, Smith E, Arora L, Segal B. The role of machine learning in HIV risk prediction. Front Reprod Health. 2022;4:1062387. https://doi.org/10.3389/frph.2022.1062387 . Carter A, Zhang M, Tram KH, Walters MK, Jahagirdar D, Brewer ED et al. Global, regional, and national burden of HIV/AIDS, 1990–2021, and forecasts to 2050, for 204 countries and territories: the Global Burden of Disease Study 2021. The Lancet HIV [Internet]. 2024; Available from: https://www.thelancet.com/journals/lanhiv/article/PIIS2352-3018(24)00212-1/fulltext Ebulue NCC, Ekkeh NOV, Ebulue NOR, Ekesiobi NCS. Machine learning insights into HIV outbreak predictions in Sub-Saharan Africa. Int Med Sci Res J. 2024;4(5):558–78. https://fepbl.com/index.php/imsrj/article/view/1121 . Dzinamarira T, Mbunge E, Chingombe I, Cuadros DF, Moyo E, Chitungo I, et al. Using machine learning models to plan HIV services: Emerging opportunities in design, implementation and evaluation. S Afr Med J. 2024;114(6b):e1439. https://pubmed.ncbi.nlm.nih.gov/39041524 . Ahlström MG, Ronit A, Omland LH, Vedel S, Obel N. Algorithmic prediction of HIV status using nationwide electronic registry data. EClinicalMedicine. 2019;17:100203. https://doi.org/10.1016/j.eclinm.2019.10.016 . Peiffer-Smadja N, Maatoug R, Lescure FX, D’Ortenzio E, Pineau J, King JR. Machine learning for COVID-19 needs global collaboration and data-sharing. Nat Mach Intell. 2020;2(6):293–4. https://www.nature.com/articles/s42256-020-0181-6 . Khalid S, Yang C, Blacketer C, Duarte-Salles T, Fernández-Bertolín S, Kim C, et al. A standardized analytics pipeline for reliable and rapid development and validation of prediction models using observational health data. Comput Methods Programs Biomed. 2021;211:106394. https://doi.org/10.1016/j.cmpb.2021.106394 . Mooney SJ, Keil AP, Westreich DJ. Thirteen questions about using machine learning in causal research (You won't believe the answer to number 10!). Am J Epidemiol. 2021;190(8):1476–82. https://doi.org/10.1093/aje/kwab047 . Wang G, Wei W, Jiang J, Ning C, Chen H, Huang J, et al. Application of a long short-term memory neural network: a burgeoning method of deep learning in forecasting HIV incidence in Guangxi, China. Epidemiol Infect. 2019;147:e194. https://doi.org/10.1017/S095026881900075X . Sebastianelli A, Spiller D, Carmo R, Wheeler J, Nowakowski A, Jacobson LV, et al. A reproducible ensemble machine learning approach to forecast dengue outbreaks. Sci Rep. 2024;14(1):52796. https://doi.org/10.1038/s41598-024-52796-9 . Sherratt K, Gruson H, Grah R, Johnson H, Niehus R, Prasse B, et al. Predictive performance of multi-model ensemble forecasts of COVID-19 across European nations. eLife. 2023;12:e81916. https://doi.org/10.7554/eLife.81916 . Wang H, Kwok KO, Riley S. Forecasting influenza incidence as an ordinal variable using machine learning. medRxiv Preprint. 2023 Feb;10. https://doi.org/10.1101/2023.02.09.23285705 . Jahagirdar D, Walters MK, Novotney A, Brewer ED, Frank TD, Carter A, et al. Global, regional, and national sex-specific burden and control of the HIV epidemic, 1990–2019, for 204 countries and territories: the Global Burden of Diseases Study 2019. Lancet HIV. 2021;8(10):e633–51. https://doi.org/10.1016/S2352-3018(21)00152-1 . Maranhão TA, Alencar CH, De Avelar Figueiredo Mafra Magalhães M, Sousa GJB, Ribeiro LM, De Abreu WC, et al. Mortality due to acquired immunodeficiency syndrome and associated social factors: a spatial analysis. Rev Bras Enferm. 2020;73(Suppl 5):e20200002. https://www.scielo.br/j/reben/a/KtDh5ZsfwRjmNgg3tpq54rB/?lang=en . Tables Table 1 Dataset Ocerview File Name Description health_gha.csv Raw dataset used for data quality checks and feature engineering ghana_infectious_disease_model_dataset_cleaned.csv Fully cleaned dataset used for EDA and predictive modeling model_performance_metrics.csv Table storing model accuracy results across algorithms for HIV/AIDS Table 2 Detailed Data Quality Assessment Summary Fetures Data Type Missing Values Missing % Unique Values Outliers Detcted Post Winsorisation region object 0 0 10 64 0 date object 0 0 612 9 0 tb_outlier bool 0 0 1 0 0 malaria_outlier bool 0 0 2 379 0 urban_rural_sum_pct float64 0 0 1 0 0 age_sum_pct float64 0 0 3 57 0 migration_rate float64 0 0 9261 N/A 0 regional_stigma_index float64 0 0 6579 0 0 urbanization_level float64 0 0 9792 112 0 health_facility_density float64 0 0 9761 N/A 0 testing_coverage_pct float64 0 0 9792 N/A 0 access_to_art_pct float64 0 0 9792 0 0 hiv_awareness_index float64 0 0 9792 0 0 youth_unemployment_rate float64 0 0 9792 0 0 female_literacy_rate float64 0 0 9792 52 0 condom_use_rate float64 0 0 9792 69 0 education_access_index float64 0 0 9792 45 0 region_fixed object 0 0 10 0 0 population_rural_pct float64 0 0 9531 52 0 population_urban_pct float64 0 0 9531 N/A 0 population_15_64_pct float64 0 0 9792 334 0 population_65_plus_pct float64 0 0 9781 N/A 0 population_0_14_pct float64 0 0 9751 0 0 population_total float64 0 0 9631 44 0 hiv_incidence float64 0 0 9598 53 0 tb_incidence float64 0 0 9598 0 0 malaria_incidence float64 0 0 9598 59 0 month int64 0 0 12 N/A 0 year int64 0 0 51 92 0 hiv_outlier bool 0 0 1 2 0 Table 3 Additional Features integrated to enrich health_gha.csv Feature Name Description education_access_index % of population with secondary education or higher condom_use_rate Percentage consistently using condoms female_literacy_rate Literacy among women aged 15+ youth_unemployment_rate Youth (15–24) unemployment rate hiv_awareness_index Composite score measuring HIV knowledge and awareness access_to_art_pct % of HIV-positive individuals receiving ART testing_coverage_pct % of population tested for HIV health_facility_density Number of health facilities per 10,000 people regional_stigma_index 0–1 index quantifying HIV-related stigma in regions urbanization_level % of population living in urban areas migration_rate Net migration rate per 1,000 people Table 4 Feature Engineering and Pre-processing Techniques Technique Description Category Lagged Variables Created 1–12 month lags for hiv_incidence to capture autocorrelation and temporal dependencies. Useful for time series models (ARIMA, LSTM). Feature Engineering Composite Indices Normalized and combined related variables (e.g., ART access, stigma index, education) into single indices scaled from 0 to 1 for comparability. Feature Engineering Binned Features Transformed continuous variables like urbanization_leve l and access_to_art_pct into categories (low, medium, high) for stratified analysis. Feature Engineering Rolling Averages Applied 6-month moving averages to smooth out noise in volatile behavioral variables (e.g., condom_use_rate , hiv_awareness_index ). Preprocessing Table 5 Model Selection Model Rationale Random Forest Regressor Selected as the baseline model due to its robustness with non-linear patterns, interpretability via feature importance, and ability to handle missing values and noisy inputs. XGBoost Regressor Chosen for its superior accuracy and performance with structured datasets. It includes regularization (L1 and L2) to prevent overfitting and handles missing data natively. Ridge Regressor Used as a linear benchmark with L2 regularization to manage multicollinearity and improve generalization in high-dimensional feature spaces. Support Vector Regressor (SVR) Implemented for its robustness with small-to-medium sized datasets and capability to model non-linear relationships via kernel functions. Particularly useful for exploring residual variance and generalization. LSTM (Optional – Future Work) Although not implemented in this final version, LSTM models are ideal for modeling long-term dependencies in time series data and are recommended for future iterations with univariate or multivariate temporal windows. Table 6 Hyperparameter Tuning and Optimal CV Strategy Model Tuned Hyperparameters Values Tried Optimal CV Strategy Random Forest n_estimators, max_depth, min_samples_split [100, 200], [10, 20, None] , [ 2 , 5 ] 5-Fold Stratified Cross-Validation XGBoost n_estimators, learning_rate, max_depth [100, 200], [0.01, 0.1] , [ 3 , 6 , 10 ] 5-Fold Stratified Cross-Validation Ridge Regression alpha [0.01, 0.1, 1.0, 10] 10-Fold Cross-Validation Support Vector Regression (SVR) C, epsilon, kernel [ 1 , 10 ], [0.1, 0.2], ['rbf'] 5-Fold Cross-Validation Table 7 Model Performance Results Rank Model RMSE R² MAE MAPE Best Hyper parameter 1st Random Forest 6.39 0.9821 4.64 2.56% {'max_depth': 20, 'min_samples_split': 2, 'n_estimators': 200} 2nd XGBoost 6.54 0.9813 4.77 2.64% {'learning_rate': 0.1, 'max_depth': 10, 'n_estimators': 100} 3rd Ridge Regression 9.08 0.9640 6.69 3.69% {'alpha': 1} 4th Support Vector Regressor (SVR) 8.47 0.9686 5.90 3.17% {'C': 10, 'epsilon': 0.2, 'kernel': 'rbf'} Table 8 Key Feature Importance and SHAP Summary for hiv_incidence Feature SHAP Score Key Interaction Highlighted SHAP for condom_use_rate 0.127 Positive effect increases with literacy SHAP for education_access_index 0.241 Nonlinear jump effect around score ~55 SHAP for hiv_awareness_index 0.198 Steep increase above awareness ~65 SHAP for tb_incidence 0.113 Moderate rise, esp. with higher urbanization Table 9 2030 Forecast and Result Summary Model RMSE 2030 Forecast Value Trend Direction (→ ↑ ↓) Notes Random Forest 0.55 231.64 → Stable forecast; moderate accuracy XGBoost 0.47 231.32 → Most accurate; closely aligns w/ SDG SVR 1.30 230.34 ↓ Slight downward trend; less precise Ridge Regression 1.24 247.62 ↑ Overestimates 2030 value Ensemble — 235.23 ↑ Balanced output across models Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviewers agreed at journal 10 May, 2026 Reviews received at journal 09 May, 2026 Reviewers agreed at journal 07 May, 2026 Reviewers agreed at journal 07 May, 2026 Reviewers agreed at journal 24 Jun, 2025 Reviews received at journal 23 Jun, 2025 Reviewers agreed at journal 20 Jun, 2025 Reviewers agreed at journal 19 Jun, 2025 Reviewers invited by journal 12 Jun, 2025 Editor invited by journal 16 May, 2025 Editor assigned by journal 15 May, 2025 Submission checks completed at journal 15 May, 2025 First submitted to journal 11 May, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6639193","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":458728766,"identity":"0b5b4205-66ae-4f6c-ac0f-4f521b502580","order_by":0,"name":"Valentine Golden Ghanem","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABEklEQVRIiWNgGAWjYFAC5gYgIQHEPGCuHIg48ACvFkaglgSEFmOwlgTCWhjgWhJBljLg02Jw/GDj48IfFnn8/GsPfi6o2JY+P+zwQ6AtdnK6DTi0nElsNp6RIFEsOeNdsvSMM7dzN95OMwBqSTY2O4Bdi2RDYps0T4JE4oYbZwykeduAWmYngLQcSNyGS0v/w/bfIC37b5wx/s3773a64ez0D3i18EsktjGDbeHvMZPmbbidIC+dg98WfomHzdI8aRKJM27wmFnzHLttuEE6p+BAggFuv7DxJx/8zGNTl9jff8b4Nk/NbXn52embP3yosJPDpQUBJBIgtAFYpQEh5WAnQg2VbyBG9SgYBaNgFIwkAACV4WPWctVeaAAAAABJRU5ErkJggg==","orcid":"","institution":"Ghana Health Service","correspondingAuthor":true,"prefix":"","firstName":"Valentine","middleName":"Golden","lastName":"Ghanem","suffix":""}],"badges":[],"createdAt":"2025-05-11 11:08:12","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6639193/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6639193/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":83609058,"identity":"32bfc228-4f2b-4fbb-a51f-bb5ae348b6ad","added_by":"auto","created_at":"2025-05-29 11:52:00","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1200835,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eModel Performance Radar Chart (Scaled)\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-6639193/v1/a4eb93cef6a8b6a4c63936a0.png"},{"id":83609766,"identity":"9c2a9df4-eb35-4774-908f-650473fa73c6","added_by":"auto","created_at":"2025-05-29 12:00:00","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":625183,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eModel Performance Heatmap\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-6639193/v1/225640070a42a22b28cdff69.png"},{"id":83609063,"identity":"e63bff75-c7ef-4718-9452-7164d9a9ec0e","added_by":"auto","created_at":"2025-05-29 11:52:00","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":745221,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eFeature Importance (Random Forest)\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image3.png","url":"https://assets-eu.researchsquare.com/files/rs-6639193/v1/70623303761db937c27cea07.png"},{"id":83609767,"identity":"fa4a3610-32a8-4ed1-8a7c-ec2d6399be85","added_by":"auto","created_at":"2025-05-29 12:00:00","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":308878,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eSHAP Dependence For Education Access Index\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-6639193/v1/7cd06a520e80d29b261c9137.png"},{"id":83609062,"identity":"df4000cc-44c0-4063-8014-c60240f2cefd","added_by":"auto","created_at":"2025-05-29 11:52:00","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":436121,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eSHAP Dependence For Condom Rate\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image5.png","url":"https://assets-eu.researchsquare.com/files/rs-6639193/v1/bd553e4876084f71d1867fb2.png"},{"id":83610135,"identity":"4b71d757-5e91-45fe-bf9e-79b311e3e335","added_by":"auto","created_at":"2025-05-29 12:08:00","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":3243956,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eCounterfactual simulation - Education and HIV Awareness\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image6.png","url":"https://assets-eu.researchsquare.com/files/rs-6639193/v1/1e6008b3aa9041cae2da2424.png"},{"id":83609067,"identity":"1b2c0a63-3bbc-4cab-9277-71e1f24fbedf","added_by":"auto","created_at":"2025-05-29 11:52:00","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":443296,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eForecasting HIV Incidence to 2030\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image7.png","url":"https://assets-eu.researchsquare.com/files/rs-6639193/v1/be73d4d96137907c519b0b51.png"},{"id":83609071,"identity":"8111091d-99d9-4990-aab2-8946e2831f8c","added_by":"auto","created_at":"2025-05-29 11:52:00","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":1548798,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003ePartial Dependence plots (Top 4 Features)\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image8.png","url":"https://assets-eu.researchsquare.com/files/rs-6639193/v1/038d20f72f13d13eafccdf00.png"},{"id":83610913,"identity":"2a1a75b9-a459-496f-90e6-6b591c701cd1","added_by":"auto","created_at":"2025-05-29 12:16:07","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":9569116,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6639193/v1/0de08c89-05c6-416a-bea5-e92b8b6472e8.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Forecasting HIV/AIDS Incidence in Ghana: A Retrospective Observational Study Using Ensemble Machine Learning Models ","fulltext":[{"header":"1. INTRODUCTION","content":"\u003cp\u003eThe Human Immunodeficiency Virus (HIV) and Acquired Immunodeficiency Syndrome (AIDS) pose worldwide health challenges which need multifaceted approaches to combat. These strategies focus on precisely forecasting public health HIV/AIDS interventions, particularly in sub-Saharan Africa where healthcare resources are limited and the rate of transmission is high. Like most of its regional counterparts, Ghana is fighting the HIV/AIDS pandemic, which has a national prevalence of about 2%. Women have a higher prevalence than men (2.5% and 1.1% respectively) [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The incidence of Ghana\u0026rsquo;s HIV/AIDS epidemic is greatly affected by spatial and socio-behavioral inequities. Greater HIV awareness is linked to higher rates of HIV testing, reduced stigma, and higher levels of social discrimination. In contrast, illiteracy and poverty is associated with greater transmission risks and higher social discrimination [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Moreover, rural areas have greater incidence because of migration patterns and limited access to health services. Although urban areas are often seen as transmission hotspots, suburban areas tend to have lower incidence rates [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eGender and marital status have been shown to be important factors that influence the dynamics of HIV transmission [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Younger women\u0026rsquo;s (15\u0026ndash;39 years) testing rates and prevalence were both higher, indicating the impact of gender disparity and sociocultural dynamics on the epidemic [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Including such socio-spatial differences into disease modeling, adds relevance to the context of the predictive insights regarding targeted intervention strategies. With these considerations in mind, public health policies, as well as the sociological frameworks relied on for designing interventions, must more accurately calibrate regionally and demographically sensitive policies tailored to the cross-cutting socio-spatial heterogeneity, as examined through advanced statistical models that integrate multidimensional analytic frameworks to enhance predictive relevance in interventions.\u003c/p\u003e \u003cp\u003eEnsemble machine learning models are utilized in public health, especially in the epidemiological surveillance of infectious diseases, to enhance forecasting accuracy and trends. These models can adapt to dynamic patterns in data from health surveillance systems. Random Forest and XGboost outperform traditional statistical models, such as Autoregressive Integrated Moving Average (ARIMA). Ensemble methods exceeded traditional ones for predicting sepsis-related mortality and the severity of Corona Virus Disease 2019 (COVID-19) pandemic [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Explainable Artificial Intelligence (xAI) frameworks, such as SHapley Additive exPlanations (SHAP), enhance the interpretability and transparency of ensemble models by identifying critical features and explaining their importance, thus promoting trust in the outputs by ML models [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. The multifactorial transmission dynamics of infectious diseases like HIV, which is inherently complex, requires the development and integration of sophisticated explainable models to enhance public health forecasting.\u003c/p\u003e \u003cp\u003eDespite the growing curiosity about applying Artificial Intelligence (AI) and Machine Learning (ML) to public health interventions, the persistent challenges of poor data quality continue to impede model effectiveness. These issues comprise of inadequate migration and behavioral data, under-reporting incidence, and sparse surveillance coverage. Moreover, the absence of explainability in ML models generates distrust among stakeholders which discourages their adoption in public health.When\u003c/p\u003e \u003cp\u003ehealth professionals are trained in the use of ensemble ML models, their confidence\u003c/p\u003e \u003cp\u003eimproves, which drives the wide acceptance and integration of this technology into\u003c/p\u003e \u003cp\u003ehealthcare.\u003c/p\u003e \u003cp\u003eThrough the analysis of demographic and complex behavioral data, models like \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003ethe Random Forest and XGboost excel in HIV risk prediction. This is supported by a recent study that reports the high accuracy of ensemble ML models in identifying high-risk individuals, notably among target groups such as men who have sex with men (MSM)\u003c/span\u003e[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eNevertheless, by forecasting clinic attendance and testing behavior after reminder messages, ML improves service delivery during HIV epidemic by optimizing patent care and resource allocation\u003c/span\u003e [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eModel optimization enhances reliability, generalizability, and cost-effectiveness, which are crucial in resource-limited settings\u003c/span\u003e [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eForecasting HIV incidence remains an underexplored area in current research. Most machine learning applications priotized classifying individuals by risk level or predicting HIV serostatus, with limited focus on long-term incidence forecasting\u0026mdash;a critical component for effective public health planning and resource allocation. Recent advances suggest that hybrid models, particularly those combining Extreme Gradient Boosting (XGBoost) with Long Short-Term Memory (LSTM) networks, demonstrate superior performance in disease forecasting tasks. These models have outperformed traditional methods in predicting infectious disease trends, including COVID-19 outbreaks, especially in settings with sparse historical data [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eThe performance of this model depends on the complete datasets, which are often scarce in African settings. Traditional time-series models\u003c/span\u003e like ARIMA and exponential smoothing \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eare commonly used in lower- and middle-income countries (LIMCs)\u003c/span\u003e, such as Ghana. \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eThese statistical models often do not fully capture complex nonlinear trends in HIV transmission\u003c/span\u003e [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eThis leads to poor predictability, particularly in settings that\u003c/span\u003e involve multiple disease incidence drivers.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eEnsemble models achieve high efficiency in disease forecasting and surveillance by integrating several data sources to achieve \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003ea better performance. LMIC\u0026rsquo;s with infrastructural and data challenges such as Ghana can benefit immensely from using an ensemble model when SHAP and counterfactual simulations are used to enhance transparency and interpretability. This illustrates the role of variables such as ART (Antiretroviral Therapy) coverage and awareness in shaping HIV incidence and transmission, and supports evidence-based conclusions in public health, thereby fostering trust among stakeholders.\u003c/span\u003eThe global health objective, as well as Ghana's broader \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003efuture HIV response\u003c/span\u003e, will depend heavily on ML-driven forecasting, which will guide strategic planning at both \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003ethe national and international levels. To ensure maximum impact, it is imperative to improve data quality and model interpretability, particularly when designing and using ML models in resource-constrained environments. Constant monitoring and hyperparameter adjustments of these models are needed in real-world applications given the evolving nature of variables in the HIV epidemic.\u003c/span\u003e\u003c/p\u003e \u003cp\u003eUsing publicly available data from multiple domains, this study seeks to assess the predictive accuracy of four supervised machine-learning algorithms (Random Forest, XGBoost, Ridge Regression, and SVR) for HIV/AIDS incidence trends in Ghana from 2000 to 2022.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"2. METHODS","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Study Design and Data Sources\u003c/h2\u003e \u003cp\u003eThis observational study retrospectively examined HIV/AIDS incidence trends using a unified dataset consisting of annual HIV/AIDS case data from Ghana's ten administrative regions from 2000 to 2022. Information was compiled from several trustworthy public sources, including the Ghana Health Service, Ghana AIDS Commission, Ministry of Health, Ghana Statistical Service, UNAIDS, World Bank, and \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eHumanitarian Data Exchange.\u003c/span\u003e The dataset includes various variables, such as demographic, epidemiologic\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eal, health system, socio-behavioral, and geospatial variables, together with detailed\u003c/span\u003e information on data domains and sources (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e1\u003c/span\u003e). To preserve temporal consistency with historical records, a ten-region format was selected, while post-2018 administrative boundary changes were omitted.\u003c/p\u003e \u003cp\u003eFurthermore, the dataset incorporated socio-behavioral and spatial metrics, such as HIV awareness indices, gender-based vulnerability markers, and regional health service accessibility, to preserve the predictive value and contextual relevance of ML models. These features reflect the structural and behavioral determinants of HIV transmission in Ghana, enabling models to account for the complex, multilayered nature of the epidemic [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. No participants were enrolled in this study and no interviews were conducted. The analysis exclusively \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003edraws on aggregate\u003c/span\u003e regional data from Ghana. Therefore, standard items regarding participant eligibility, matching, or calculation of \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003esample size\u003c/span\u003e were not applied.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Data Preparation and Feature Engineering\u003c/h2\u003e \u003cp\u003eA comprehensive data quality assessment pipeline encompassing metadata screening, indicator validation, and outlier identification via Forest Isolation and Winsorization at the 1st and 99th percentiles was used to minimize noise and outliers. Missing data were estimated using the KNN algorithm. Temporal dependencies and trend smoothing were achieved using feature-engineering techniques, including the creation of lag variables and moving averages. Indices combining HIV awareness and regional socioeconomic factors have been developed to encapsulate \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003ethe various factors that affect\u003c/span\u003e HIV incidence rates. All features were normalized using the MinMax method to allow comparable scales (Tables\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTo reduce measurement and selection bias, the data were validated using external metadata, including outlier detection and imputation, which are highly effective on their own. Enhanced inclusivity strengthened the sources \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eused\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Model Selection Theory and Development\u003c/h2\u003e \u003cp\u003eThe choice of model was based on the theoretical and empirical advantages of various supervised machine-learning techniques to identify intricate, possibly \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003enonlinear patterns within the data. A comparative evaluation of the four\u003c/span\u003e ML models (Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eRandom Forest (RF) and XGBoost are ensemble tree-based methods \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003ethat are renowned for their robust capacity to approximate nonlinear functions and capture variable interactions.\u003c/span\u003e\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRidge Regression provides a form of regularized linear modeling for evaluating baseline linear relationships.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eSVR (\u003c/span\u003eSupport Vector Regression) \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003efunctions as a kernel-based approach that enables flexible nonlinear regression by using kernel transformations.\u003c/span\u003e\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThe selection rationale strikes a strategic balance between model complexity, interpretability, predictive capability, and computational efficiency, as dictated by the established best practices in \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003emodel selection theory, specifically\u003c/span\u003e regarding model traits and tuning parameters (Table\u0026nbsp;\u003cspan refid=\"Tab7\" class=\"InternalRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Model Training, Hyperparameter Tuning, and Evaluation\u003c/h2\u003e \u003cp\u003eThe dataset was divided into two subsets (80/20 split): a training set that spanned the years 2000 to 2017 and a testing set that covered 2018 to 2022, with the chronological order maintained to avoid data contamination and \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eto mimic real-world forecasting situations. Grid and Random Search methods were employed within the cross-validation folds of the training data to\u003c/span\u003e determine the optimal model settings through hyperparameter optimization. The performance \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eof the model was assessed on the test dataset\u003c/span\u003e using various evaluation criteria including the coefficient of determination (R\u0026sup2;), root mean squared error (RMSE), mean absolute error (MAE), and mean absolute percentage error (MAPE). Ensemble models, notably XGBoost and Random Forest, showed higher effectiveness with R\u0026sup2; values surpassing 0.98 and low prediction errors (Tables\u0026nbsp;\u003cspan refid=\"Tab7\" class=\"InternalRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Tab9\" class=\"InternalRef\"\u003e8\u003c/span\u003e; Figs.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e1\u003c/span\u003e and \u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e \u003ch2\u003e2.4.1 Sensitivity Analysis\u003c/h2\u003e \u003cp\u003eSensitivity analyses were not \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eformally conducted\u003c/span\u003e. Nonetheless, model robustness was evaluated across ensemble, linear, and kernel-based methods, as well as through interpretability tools \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003esuch as SHAP and counterfactuals, to confirm the consistency of\u003c/span\u003e recorded values.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.5 Model Interpretability: Justification for SHAP and Counterfactuals\u003c/h2\u003e \u003cp\u003eInterpreting complex ensemble machine-learning models is essential for translating predictions into practical insights that can be used in public health initiatives. Feature interactions, including subgroup behaviors by region, were analyzed \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eusing th\u003c/span\u003ee SHAP \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003edependence plots and counterfactual simulations.SHAP values offer a theoretically sound game theory approach for both local and global interpretations, allowing the model output to be broken down according to feature influence while accounting for feature interactions and correlations. In the context of Ghana, this interpretability is vital for transforming model outputs into actionable public health strategies\u0026mdash;by revealing, for instance, how marginal increases in ART coverage or HIV awareness disproportionately affect incidence in specific regions. This enables policymakers to prioritize modifiable drivers in high-burden districts, thereby improving the precision and impact of HIV interventions\u003c/span\u003e [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eCounterfactual simulation methods were used to investigate hypothetical situations, examining how changes in key modifiable factors (e.g., HIV awareness and condom use) might impact \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eprojected HIV incidence patterns. This approach enables an actionable inference by pinpointing the intervention points of influence. Consequently, the combination of SHAP and counterfactual analysis surpasses the traditional feature importance, bringing model transparency in line with the current interpretability benchmarks in health analytics (\u003c/span\u003eTable\u0026nbsp;\u003cspan refid=\"Tab9\" class=\"InternalRef\"\u003e8\u003c/span\u003e; Figs.\u0026nbsp;\u003cspan refid=\"Fig11\" class=\"InternalRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Fig14\" class=\"InternalRef\"\u003e6\u003c/span\u003e). Although machine learning models do not adjust for confounding in a traditional statistical sense, SHAP values illuminate the influence of key predictors and highlight features that buffer key confounding relationships.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.6 Statistical and Software Tools\u003c/h2\u003e \u003cp\u003e \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eAll analyses were performed in Python 3.10 using libraries such as pandas, numpy, and scikit-learn for data handling and model development, and SHAP for interpretability analysis. Visualization libraries such as matplotlib and seaborn enabled the development of interpretive plots to support model evaluation and insight generation.\u003c/span\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"3. RESULTS","content":"\u003cdiv id=\"Sec11\"\u003e\n \u003ch2\u003e3.1 Predictive Performance and Model Comparison\u003c/h2\u003e\n \u003cp\u003eThe predictive abilities of four supervised machine learning algorithms, Random Forest, XGBoost, Ridge Regression, and SVR, were thoroughly examined on a test dataset reserved prior to model training, covering the period from 2018 to 2022. The dataset comprised of a harmonized panel of observations spanning 2000 to 2022, thereby maintaining temporal consistency and illustrating significant epidemiological trends in the incidence of HIV/AIDS in Ghana. Given that this study was conducted using aggregated data, aspects concerning individual participants\u0026rsquo; identification, eligibility, or attrition did not apply.\u003c/p\u003e\n \u003cp\u003eTable\u0026nbsp;2 shows missing value percentages across all variables before KNN imputation was performed. By the last step of the preprocessing pipeline, all variables had 0% missing data, which means that the provided steps for data cleaning were sufficient to resolve issues of missing data prior to any form of advanced imputation techniques. The performance assessment employed multiple standard metrics, like the coefficient of determination (R \u0026sup2;), Root Mean Square error (RMSE), Mean Absolute Error (MAE), and Mean Absolute Percentage Error (MAPE), thereby facilitating a comprehensive evaluation of the error magnitudes and relative fit quality.\u003c/p\u003e\n \u003cp\u003eEnsemble models, including Random Forest and XGBoost, showed enhanced predictive abilities, as outlined in Table\u0026nbsp;9. The Random Forest model produced the best results, with an R\u0026sup2; of 0.9821, RMSE of 6.39, MAE of 4.64, and MAPE of 2.56%, and eclipsing XGBoost with an R\u0026sup2; of 0.9813, RMSE of 6.54, MAE of 4.77, and MAPE of 2.64%. Unlike the baseline models, Ridge Regression and SVR produced significantly lower performance outcomes, with R\u0026sup2; values of 0.9640 and 0.9686, respectively, indicating their inability to effectively model the complex nonlinear relationships observed in HIV incidence patterns (Table\u0026nbsp;7).\u003c/p\u003e\n \u003cdiv\u003eCross-validation and hyperparameter tuning led to XGBoost\u0026apos;s slight technical edge in long-term forecasting from 2023 to 2030, with an RMSE of 0.47, outperforming Random Forest\u0026apos;s RMSE of 0.55 (Table 10). According to the data, the ensemble forecasts predicted stabilization of HIV incidence at a rate of approximately 231 cases per 100,000 population by 2030, as illustrated in Fig. 7. Linear models exhibited contradictory trends, with Ridge Regression showing an upward trend in incidence and SVR forecast showing a slight decline, although with a reduced accuracy. The model-averaged forecast indicates a gradual increase, highlighting the need for more aggressive corrective measures (Table 10; Figs. 1, 2, and 7). The models used were not regression-based; therefore, the confidence intervals and p-values for the effect estimation were not reported. The performance metrics reflected the predictive accuracy through the R\u0026sup2;, RMSE, MAE, and MAPE.\u003ctable id=\"Tab1\" border=\"1\"\u003e\u003c/table\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec12\"\u003e\n \u003ch2\u003e3.2 Model Interpretability Through SHAP and Counterfactual Analysis\u003c/h2\u003e\n \u003cp\u003eSome techniques have been applied to increase the accuracy of models and to derive meaningful public health insights. The SHapley Additive exPlanations (SHAP) method was used to calculate the specific impact of each input variable on the model predictions at both overall and individual levels. This game-theoretic method enables a thorough breakdown of the prediction effects, while considering the relationships between different features.\u003c/p\u003e\n \u003cp\u003eHIV awareness, ART coverage, and education access were among the most influential drivers in forecasting HIV incidence, as shown in the SHAP summary plot (Fig.\u0026nbsp;3). This finding is consistent with regional HIV prediction efforts in sub-Saharan Africa, where ensemble models identified these factors as critical for targeting screening and retention strategies [18, 19].Furthermore, the SHAP analysis was influenced by spatial variables such as access to health services across rural-urban contexts, as demonstrated in previous studies [4, 5]. This indicates that due to population and migration dynamics, urban cities have higher incidence rates, while rural resource-poor communities suffer from hidden transmission. Additional analysis in Fig.\u0026nbsp;8 using Partial Dependence Plots revealed complex interplay of socioeconomic factors alongside HIV incidence, which in addition to ART coverage and HIV awareness involved nonlinear inverse relationships.\u003c/p\u003e\n \u003cp\u003eCounterfactual simulation methods were applied to devise hypothetical situations where certain key variables were altered to explore their effect on HIV trends and outcomes. This possibility provides a space for causal reasoning by demonstrating what would happen with the use of specific interventions, thus linking predictive modeling with policy-driven decision-making processes. Mulit-dimensional SHAP and counterfactual analyses directly enhance the clarity and rigor of trust placed in complex ensemble models, while upholding the principles of designed-for-purpose interpretability within the epidemiological machine-learning framework.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\"\u003e\n \u003ch2\u003e3.3 Conclusion\u003c/h2\u003e\n \u003cp\u003eThe ensemble machine learning models, especially XGBoost, achieved optimal performance in the predictive modeling of HIV/AIDS cases within the Ghanaian epidemiological context due to their parameter tuning capabilities. The performance of the Random Forest algorithm was also fairly accurate. The study further illustrated the ensemble tree methods\u0026rsquo; superiority over the older, linear and kernel-based models which are inadequate in dealing with intricate data dependencies.\u003c/p\u003e\n \u003cp\u003eSHAP and counterfactual interpretation techniques are critical for model explanation to enhance evidence-based public health responses as well as proactive planning and targeted interventions. These conclusions endorse the application of complex interpretive frameworks of ensemble machine learning techniques to sharpen epidemiological surveillance, optimize the allocation of resources, and enhance planning and responsiveness to the HIV epidemic in Ghana.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"4. DISCUSSION","content":"\u003cp\u003eThe study showed that ensemble machine learning (ML) models, particularly Random Forest and XGBoost, achieved the highest predictive accuracy (R² \u0026gt; 0.98) in forecasting HIV/AIDS incidence in Ghana over a period of two decades. Ensemble ML models outperform Support Vector Regression (SVM) and Ridge Regression due to their inherent capabilities to capture nonlinear and multivariate time-dependent relationships. Such characteristics make them effective for forecasting infectious diseases, especially when socio-behavioral, health, and demographic factors overlap. These findings align with observations that documented similar results for other infectious diseases [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThis study stands out because it attempts to explain the effects of major predictors like ART coverage and HIV awareness using SHAP and counterfactual simulations, thus bridging the gap between predictive analytics and actionable public health policy. These methods dispel any uncertainty relating to the “black box” nature of ML models which, do not integrate counterfactual simulations, nor SHAP analysis, when forecasting trends. This research conclusion supports epidemiological evidence indicating that increasing ART and treatment coverage rates inadvertently leads to viral load suppression, thereby reducing HIV transmission and incidence rates [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]\u003c/p\u003e \u003cp\u003eThe identification of alterable parameters, like awareness of HIV and ART\u003c/p\u003e \u003cp\u003ecoverage, as important determinants aligns with the epidemiological understanding\u003c/p\u003e \u003cp\u003ethat HIV transmission is influenced by the behavior, structure, and healthcare delivery\u003c/p\u003e \u003cp\u003esystems of a region. The importance of socioeconomic factors and health system\u003c/p\u003e \u003cp\u003einfrastructure \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eis connect\u003c/span\u003eed to broader evidence \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003ethat call\u003c/span\u003es for holistic HIV prevention. Strategies that integrate biomedical, behavioral, and social actions [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Such gaps in targeted public health interventions strengthen the existing theories to enhance awareness initiatives and improve the range and accessibility of ART services.\u003c/p\u003e \u003cp\u003eThe study results reveal an alarming trend signaling a realistic forecast:\u003c/p\u003e \u003cp\u003estabilization of HIV incidence is projected to reach 231 cases per 100,000 population\u003c/p\u003e \u003cp\u003eby 2030, which is drastically lower than the national and global targets set for\u003c/p\u003e \u003cp\u003emeaningful reduction in new infections. This aligns with the broader downward trends forecasted in the Global Burden of Disease Study 2021, which provides country-level HIV incidence projections to 2050 [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. It exemplifies the persistent challenge of HIV epidemic control in Ghana and emphasizes the crucial need to shift approaches to prevention intensification beyond the current trajectories. These models provide tools designed for public health planning and fundamentally\u003c/p\u003e \u003cp\u003equantify the impact of scaling up targeted resources and implementing interventions.\u003c/p\u003e \u003cp\u003eThis finding is consistent with the broader trends observed in sub-Saharan Africa,\u003c/p\u003e \u003cp\u003ewhere renewed attempts to reduce HIV incidence have led to stagnation in new cases,\u003c/p\u003e \u003cp\u003emost likely as a result of socio-behavioral determinants, systematic health resource\u003c/p\u003e \u003cp\u003ebottlenecks and unequal resource distribution\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003es\u003c/span\u003e [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThese results have several direct implications for public health policy in Ghana. First, the high predictive value of ART coverage and HIV awareness suggests that intervention efforts should prioritize scaling up ART accessibility and education programs, particularly in rural and underserved regions. This aligns with findings in Mozambique and Nigeria where ML was used to identify ART clients at high risk of treatment interruption, effectively guiding outreach efforts [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Second, targeted outreach campaigns focused on condom use and stigma reduction in high-burden districts, especially in Greater Accra and Ashanti regions, are warranted. SHAP analysis revealed non-linear relationships suggesting diminishing returns in some interventions—echoing the importance of precision public health over blanket strategies [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Regional health plans integrating ML-based risk stratification may optimize resource allocation and increase intervention efficiency.\u003c/p\u003e \u003cp\u003eAnother strength of this study is the use of multi-source datasets that integrate epidemiological, demographical, behavioral, and infrastructural indicators from 2000 to 2022, providing deeper insights via enhanced data-driven analysis than a single-domain dataset. The random forest and XGboost models consistently outperformed \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003ethe other models in terms of\u003c/span\u003e accuracy across all regions, with ART coverage and HIV awareness being the most influential predictors. Persisting regional disparities and structural factors affect \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eincidence trends, which are further influenced by socioeconomic factors, especially in the Greater Accra and Ashanti regions. The ecological design of this study and lack of res\u003c/span\u003eou\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003erces in LMIC’s including Ghana,imposes constraints on data quality, which includes incomplete reports, reporting delays, and residual bias. Further studies support this claim\u003c/span\u003e [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. In \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eChina\u003c/span\u003e, traditional models, \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003esuch as ARIMA and exponential smoothing, have been successfully applied because of the availability of high-frequency monthly incidence data\u003c/span\u003e [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eThese models are plagued by linearity and stationarity, which affect their ability to capture complex predictors of HIV transmission in LMICs. In contrast, ensemble models\u003c/span\u003e, such as XGBoost, \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eare more adept at handling such complexities.\u003c/span\u003e\u003c/p\u003e \u003cp\u003e \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eDespite the strengths of this study, several limitations could impact the generalizability and robustness of its findings. First, national HIV datasets are subject to under-reporting, especially in rural districts, due to incomplete surveillance coverage, delayed reporting cycles, and limited behavioral indicators\u003c/span\u003e [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eSecond, the reliance on aggregate region-level data introduces risks of ecological fallacy, where individual-level causal relationships cannot be inferred\u003c/span\u003e [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eThird, although imputation techniques were used, missing data may bias model outputs if the missingness is not completely at random,a well-recognized challenge in machine learning applications within public health settings\u003c/span\u003e [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eFinally, data sparsity in sub-populations—such as key communities or mobile groups—limits the precision of risk predictions and impedes targeted interventions\u003c/span\u003e [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eThese limitations suggest a cautious interpretation of the results and emphasize the need for further model validation in prospective settings, ideally through formal sensitivity analyses and robustness testing methods such as perturbation, bootstrapping, or Monte Carlo simulations, in line with best practices for predictive modeling in public health research\u003c/span\u003e [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eMoreover, while current models demonstrate performance over a retrospective period, their ability to suddenly predict shifts like the rapid scaling up of new prevention technologies during social behavior changes, or System Implosions and Pandemics raises concern. This is where real-time monitoring and continual adjustment of models becomes essential [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. In an attempt to fight these issues, hybrid methods like XGBoost integrated with LSTM are effective because they include non-linear feature interaction as well as temporal features like movement and behavioral change [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. In LMIC’s, however, these advanced models cannot be applied owing to resource limitations.\u003c/p\u003e \u003cp\u003eResults from this research corroborate with several investigations validating the accuracy and reliability of ensemble machine-learning techniques which outperform individual models or single-method approaches, by aggregating forecasts from several base learners.Such uniformity has been observed across infectious diseases, including dengue and COVID-19 [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. The gradient boosting models, particularly XGBoost, are highly accurate due to their adjustable hyperparameters and proven ability to model sophisticated interactions among features on a myriad of factors in theoretical and applied epidemiology [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn summary, this study illustrates the invaluable predictive power of \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003ethe ensemble ML models in HIV incidence forecasting, while pinpointing relevant considerations for policymakers. The Random Forest and XGBoost models showed high accuracy and provided meaningful insights that are advantageous for public health planning in Ghana and other similar contexts. Nonetheless, these advantages depend on the existence of detailed\u003c/span\u003e real-time data, which remain a major problem in most sub-Saharan African health systems. Continuous funding and policy support are essential to enhance the efficacy of these ensemble models. Additionally, the following policies and measures can be implemented:\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\u003cul\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eEnhancing health data infrastructure\u003c/b\u003e: The availability of detailed, complete, and timely data is critical for developing and testing useful models.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eKeeping up with changing circumstances\u003c/b\u003e: Integrating emerging data on mobility and Pre-Exposure Prophylaxis(PrEP) uptake is vital to prevent unpredictable epidemiological shifts.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eStructural strategies\u003c/b\u003e: Supplementing biomedical approaches with strategies focused on education, economic support, equality, \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eand inclusion is crucial\u003c/span\u003e [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e].\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eEquity-focused resource allocation\u003c/b\u003e: Addressing specific areas and populations with the highest spatial and demographic risk levels as defined by spatial and demographic analytics is vital[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e].\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003cp\u003e\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"CONCLUSION","content":"\u003cp\u003eThis study demonstrates that ensemble machine learning models, particularly XGBoost and Random Forest, can accurately and precisely forecast HIV incidence trends in Ghana. Leveraging advanced interpretability techniques such as SHAP and counterfactual analysis, these models translate complex data patterns into actionable insights, effectively connecting technical findings with public health policy needs. Key predictors identified, notably ART coverage and HIV awareness, represent critical targets for intervention efforts. Integrating these findings into Ghana’s public health strategies can enhance disease surveillance and enable more efficient allocation of resources, supporting national and global objectives for HIV reduction even within the constraints of limited resources.\u003c/p\u003e\u003cp\u003eDespite their high predictive performance, the utility of these ensemble models is inherently limited by data quality, ecological variability, and changing epidemiological dynamics. Predictive models carry unavoidable uncertainty and require ongoing validation, incorporation of real-time data, and adaptive retraining to maintain their relevance and accuracy in evolving contexts.\u003c/p\u003e\u003cp\u003eFuture research should explore hybrid modeling approaches that combine machine learning with mechanistic models to jointly capture statistical trends and underlying transmission dynamics. Real-time adaptive modeling frameworks would also be valuable for responding dynamically to shifts in HIV epidemiology. By embedding such integrated modeling tools within broader public health frameworks, Ghana and similar low-resource settings can better harness predictive analytics to inform targeted interventions and accelerate progress toward epidemic control. While these findings are most directly applicable to Ghana and comparable contexts with similar socio-epidemiological profiles, caution should be exercised when generalizing to settings with substantially different health systems or demographics.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003col\u003e\n \u003cli\u003eAI: Artificial Intelligence\u003c/li\u003e\n \u003cli\u003eAIDS: Acquired Immunodeficiency Syndrome\u003c/li\u003e\n \u003cli\u003eARIMA: AutoRegressive Integrated Moving Average\u003c/li\u003e\n \u003cli\u003eART: Antiretroviral Therapy\u003c/li\u003e\n \u003cli\u003eCC-BY: Creative Commons Attribution (license type)\u003c/li\u003e\n \u003cli\u003eCOVID-19: Corona Virus Disease 2019\u003c/li\u003e\n \u003cli\u003eGAC: Ghana AIDS Commission\u003c/li\u003e\n \u003cli\u003eGHS: Ghana Health Service\u003c/li\u003e\n \u003cli\u003eGSS: Ghana Statistical Service\u003c/li\u003e\n \u003cli\u003eHDX: Humanitarian Data Exchange\u003c/li\u003e\n \u003cli\u003eHIV: Human Immunodeficiency Virus\u003c/li\u003e\n \u003cli\u003eKNN: K-Nearest Neighbors\u003c/li\u003e\n \u003cli\u003eLIMCs: Lower and Middle Income Countries\u003c/li\u003e\n \u003cli\u003eLSTM: Long Short-Term Memory\u003c/li\u003e\n \u003cli\u003eMAE: Mean Absolute Error\u003c/li\u003e\n \u003cli\u003eMAPE: Mean Absolute Percentage Error\u003c/li\u003e\n \u003cli\u003eML: Machine Learning\u003c/li\u003e\n \u003cli\u003eMoH: Ministry of Health\u003c/li\u003e\n \u003cli\u003eMSM: Men who have Sex with Men\u003c/li\u003e\n \u003cli\u003eORCID: Open Researcher and Contributor ID\u003c/li\u003e\n \u003cli\u003ePrEp: Pre - Exposure Prophylaxis\u003c/li\u003e\n \u003cli\u003eR²: Coefficient of Determination\u003c/li\u003e\n \u003cli\u003eRF: Random Forest\u003c/li\u003e\n \u003cli\u003eRMSE: Root Mean Squared Error\u003c/li\u003e\n \u003cli\u003eSHAP: SHapley Additive exPlanations\u003c/li\u003e\n \u003cli\u003eSVR: Support Vector Regression\u003c/li\u003e\n \u003cli\u003eUNAIDS: Joint United Nations Programme on HIV/AIDS\u003c/li\u003e\n \u003cli\u003eWHO: World Health Organization\u003c/li\u003e\n \u003cli\u003exAI: Explainable Artificial Intelligence\u003c/li\u003e\n \u003cli\u003eXGBoost: Extreme Gradient Boosting\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003cstrong\u003eEthics approval and consent to participate:\u003c/strong\u003e \u003cp\u003eNot applicable. This study used only publicly available aggregated data and did not involve human participants, human data, or \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003edirectly collected human tissue.\u003c/span\u003e\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eConsent for publication:\u003c/strong\u003e \u003cp\u003eNot applicable. This manuscript did not contain any individual data.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eCompeting interests:\u003c/h2\u003e \u003cp\u003eThe author have no competing interests to declare.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding:\u003c/h2\u003e \u003cp\u003eThe author \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003edeclares\u003c/span\u003e no funding for this study.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eVG: Conceptualization, Methodology, Data curation, formal analysis, software development, Validation, Visualization, Writing \u0026ndash; original draft, and writing \u0026ndash; review and editing.\u003c/p\u003e\u003ch2\u003eAcknowledgements:\u003c/h2\u003e \u003cp\u003eThe author\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003es wish\u003c/span\u003e to thank the Ghana Health Service, Ghana AIDS Commission, Ghana Statistical Service, and the Humanitarian Data Exchange (HDX) platform for providing access to critical epidemiological and demographic datasets that supported the development of this research. Appreciation has also \u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003ebeen extended to contributors\u003c/span\u003e to open geospatial data via GeoBoundaries. This work was dedicated to the memory of my beloved sister, Imelda Farr, whose encouragement inspired my pursuit of this academic endeavor.\u003c/p\u003e \u003cp\u003e \u003cb\u003eAuthors' information\u003c/b\u003e: VG is a Ghana-based biomedical scientist at the Cocoa Clinic (Ghana Cocoa Board\u0026rsquo;s Medical Department) in Accra. He holds an MSc in Data Science from the University of East London and is currently completing an MSc in Public Health at the University of Suffolk through distance learning. His work focuses on infectious disease modeling, epidemiologic\u003cspan type=\"SmallCaps\" class=\"SmallCaps\" name=\"Emphasis\"\u003eal forecasting, and the application of machine learning to public health policies.\u003c/span\u003e\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe full dataset, cleaned indicators, spatial files, and model code supporting this study are publicly available at the Zenodo repository under license CC-BY 4.0. Available at: https://doi.org/10.5281/zenodo.15292209.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eBa DM, Ssentongo P, Sznajder KK. Prevalence, behavioral and socioeconomic factors associated with human immunodeficiency virus in Ghana: a population-based cross-sectional study. J Glob Health Rep. 2019;3:e2019092. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.29392/joghr.3.e2019092\u003c/span\u003e\u003cspan address=\"10.29392/joghr.3.e2019092\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNutor J, Duah H, Duodu P, et al. Geographical variations and factors associated with recent HIV testing prevalence in Ghana. BMJ Open. 2021;11:e045458. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1136/bmjopen-2020-045458\u003c/span\u003e\u003cspan address=\"10.1136/bmjopen-2020-045458\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMelkam M, Fente BM. Multilevel analysis of discrimination of people living with HIV/AIDS and associated factors in Ghana. Front Public Health. 2024;12:1379487. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/fpubh.2024.1379487\u003c/span\u003e\u003cspan address=\"10.3389/fpubh.2024.1379487\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWand H, Morris N, Reddy T. Temporal and spatial monitoring of HIV prevalence and incidence rates using geospatial models. Spat Spatiotemporal Epidemiol. 2021;37:100413. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.sste.2021.100413\u003c/span\u003e\u003cspan address=\"10.1016/j.sste.2021.100413\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDias B, Rodrigues T, Botelho E, Oliveira M, Feij\u0026atilde;o A, Polaro S. Integrative review on the incidence of HIV infection and its socio-spatial determinants. Rev Bras Enferm. 2021;74(2):e20200905. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1590/0034-7167-2020-0905\u003c/span\u003e\u003cspan address=\"10.1590/0034-7167-2020-0905\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSia D, Onadja Y, Hajizadeh M, et al. What explains gender inequalities in HIV/AIDS prevalence in sub-Saharan Africa? BMC Public Health. 2016;16:1136. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12889-016-3783-5\u003c/span\u003e\u003cspan address=\"10.1186/s12889-016-3783-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAdetokunboh O, Are E. Spatial distribution and determinants of HIV high burden in the Southern African sub-region. PLoS ONE. 2024;19:e0301850. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1371/journal.pone.0301850\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0301850\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHou N, Li M, He L, Xie B, Wang L, Zhang R, et al. Predicting 30-days mortality for MIMIC-III patients with sepsis-3: a machine learning approach using XGBoost. J Transl Med. 2020;18(1):NA. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12967-020-02620-5\u003c/span\u003e\u003cspan address=\"10.1186/s12967-020-02620-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu CA, Chen CM, Fang YC, Liang SJ, Wang HC, Fang WF, et al. Using a machine learning approach to predict mortality in critically ill influenza patients: a cross-sectional retrospective multicentre study in Taiwan. BMJ Open. 2020;10(2):e033898. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1136/bmjopen-2019-033898\u003c/span\u003e\u003cspan address=\"10.1136/bmjopen-2019-033898\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHong W, Zhou X, Jin S, Lu Y, Pan J, Lin Q, et al. A comparison of XGBoost, Random Forest, and nomograph for the prediction of disease severity in patients with COVID-19 pneumonia: implications of cytokine and immune cell profile. Front Cell Infect Microbiol. 2022;12:819267. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/fcimb.2022.819267\u003c/span\u003e\u003cspan address=\"10.3389/fcimb.2022.819267\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJi X, Tang Z, Osborne SR, Van Nguyen TP, Mullens AB, Dean JA, et al. STI/HIV risk prediction model development\u0026mdash;A novel use of public data to forecast STIs/HIV risk for men who have sex with men. Front Public Health. 2025;12:1511689. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/fpubh.2024.1511689\u003c/span\u003e\u003cspan address=\"10.3389/fpubh.2024.1511689\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu X, Fairley CK, Chow EPF, Lee D, Aung ET, Zhang L, et al. Using machine learning approaches to predict timely clinic attendance and the uptake of HIV/STI testing post clinic reminder messages. Sci Rep. 2022;12(1):12033. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41598-022-12033-7\u003c/span\u003e\u003cspan address=\"10.1038/s41598-022-12033-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu D, Chan WH, Haron H. Enhancing infectious disease prediction model selection with multi-objective optimization: an empirical study. PeerJ Comput Sci. 2024;10:e2217. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.7717/peerj-cs.2217\u003c/span\u003e\u003cspan address=\"10.7717/peerj-cs.2217\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo K, Shen C, Zhou X, et al. Traffic data-empowered XGBoost-LSTM framework for infectious disease prediction. IEEE Trans Intell Transp Syst. 2022;PP(99):1\u0026ndash;12. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/TITS.2022.3172206\u003c/span\u003e\u003cspan address=\"10.1109/TITS.2022.3172206\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLucas B, Vahedi B. A spatiotemporal machine learning approach to forecasting COVID-19 incidence at the county level in the United States. \u003cem\u003earXiv\u003c/em\u003e. Preprint. 2021. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arXiv.2109.12094\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.2109.12094\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTang D, Jin Y, Hu X, Lin D, Kapar A, Wang Y, et al. Study on the prediction performance of AIDS monthly incidence in Xinjiang based on time series and deep learning models. BMC Public Health. 2025;25(1):21982. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12889-025-21982-3\u003c/span\u003e\u003cspan address=\"10.1186/s12889-025-21982-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu B, Li J, Wang M. Epidemiological and time series analysis on the incidence and death of AIDS and HIV in China. BMC Public Health. 2020;20(1):9977. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12889-020-09977-8\u003c/span\u003e\u003cspan address=\"10.1186/s12889-020-09977-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMutai CK, McSharry PE, Ngaruye I, Musabanganji E. Use of machine learning techniques to identify HIV predictors for screening in sub-Saharan Africa. BMC Med Res Methodol. 2021;21(1):268. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s12874-021-01346-2\u003c/span\u003e\u003cspan address=\"10.1186/s12874-021-01346-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStockman J, Friedman J, Sundberg J, Harris E, Bailey L. Predictive analytics using machine learning to identify ART clients at health system level at greatest risk of treatment interruption in Mozambique and Nigeria. J Acquir Immune Defic Syndr. 2022;90(2):154\u0026ndash;60. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pubmed.ncbi.nlm.nih.gov/35262514/\u003c/span\u003e\u003cspan address=\"https://pubmed.ncbi.nlm.nih.gov/35262514/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAyamah R, Awuitor GK, Puotier Z. Mathematically modeling the spread of HIV/AIDS infection after the introduction of antiretroviral therapy in Ghana. Am J Comput Eng. 2020;3(1):24\u0026ndash;39. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.47672/ajce.602\u003c/span\u003e\u003cspan address=\"10.47672/ajce.602\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFieggen J, Smith E, Arora L, Segal B. The role of machine learning in HIV risk prediction. Front Reprod Health. 2022;4:1062387. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3389/frph.2022.1062387\u003c/span\u003e\u003cspan address=\"10.3389/frph.2022.1062387\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCarter A, Zhang M, Tram KH, Walters MK, Jahagirdar D, Brewer ED et al. Global, regional, and national burden of HIV/AIDS, 1990\u0026ndash;2021, and forecasts to 2050, for 204 countries and territories: the Global Burden of Disease Study 2021. \u003cem\u003eThe Lancet HIV\u003c/em\u003e [Internet]. 2024; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.thelancet.com/journals/lanhiv/article/PIIS2352-3018(24)00212-1/fulltext\u003c/span\u003e\u003cspan address=\"https://www.thelancet.com/journals/lanhiv/article/PIIS2352-3018(24)00212-1/fulltext\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEbulue NCC, Ekkeh NOV, Ebulue NOR, Ekesiobi NCS. Machine learning insights into HIV outbreak predictions in Sub-Saharan Africa. Int Med Sci Res J. 2024;4(5):558\u0026ndash;78. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://fepbl.com/index.php/imsrj/article/view/1121\u003c/span\u003e\u003cspan address=\"https://fepbl.com/index.php/imsrj/article/view/1121\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDzinamarira T, Mbunge E, Chingombe I, Cuadros DF, Moyo E, Chitungo I, et al. Using machine learning models to plan HIV services: Emerging opportunities in design, implementation and evaluation. S Afr Med J. 2024;114(6b):e1439. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pubmed.ncbi.nlm.nih.gov/39041524\u003c/span\u003e\u003cspan address=\"https://pubmed.ncbi.nlm.nih.gov/39041524\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAhlstr\u0026ouml;m MG, Ronit A, Omland LH, Vedel S, Obel N. Algorithmic prediction of HIV status using nationwide electronic registry data. EClinicalMedicine. 2019;17:100203. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.eclinm.2019.10.016\u003c/span\u003e\u003cspan address=\"10.1016/j.eclinm.2019.10.016\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeiffer-Smadja N, Maatoug R, Lescure FX, D\u0026rsquo;Ortenzio E, Pineau J, King JR. Machine learning for COVID-19 needs global collaboration and data-sharing. Nat Mach Intell. 2020;2(6):293\u0026ndash;4. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.nature.com/articles/s42256-020-0181-6\u003c/span\u003e\u003cspan address=\"https://www.nature.com/articles/s42256-020-0181-6\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhalid S, Yang C, Blacketer C, Duarte-Salles T, Fern\u0026aacute;ndez-Bertol\u0026iacute;n S, Kim C, et al. A standardized analytics pipeline for reliable and rapid development and validation of prediction models using observational health data. Comput Methods Programs Biomed. 2021;211:106394. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.cmpb.2021.106394\u003c/span\u003e\u003cspan address=\"10.1016/j.cmpb.2021.106394\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMooney SJ, Keil AP, Westreich DJ. Thirteen questions about using machine learning in causal research (You won't believe the answer to number 10!). Am J Epidemiol. 2021;190(8):1476\u0026ndash;82. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/aje/kwab047\u003c/span\u003e\u003cspan address=\"10.1093/aje/kwab047\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang G, Wei W, Jiang J, Ning C, Chen H, Huang J, et al. Application of a long short-term memory neural network: a burgeoning method of deep learning in forecasting HIV incidence in Guangxi, China. Epidemiol Infect. 2019;147:e194. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1017/S095026881900075X\u003c/span\u003e\u003cspan address=\"10.1017/S095026881900075X\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSebastianelli A, Spiller D, Carmo R, Wheeler J, Nowakowski A, Jacobson LV, et al. A reproducible ensemble machine learning approach to forecast dengue outbreaks. Sci Rep. 2024;14(1):52796. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41598-024-52796-9\u003c/span\u003e\u003cspan address=\"10.1038/s41598-024-52796-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSherratt K, Gruson H, Grah R, Johnson H, Niehus R, Prasse B, et al. Predictive performance of multi-model ensemble forecasts of COVID-19 across European nations. eLife. 2023;12:e81916. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.7554/eLife.81916\u003c/span\u003e\u003cspan address=\"10.7554/eLife.81916\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang H, Kwok KO, Riley S. Forecasting influenza incidence as an ordinal variable using machine learning. medRxiv Preprint. 2023 Feb;10. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1101/2023.02.09.23285705\u003c/span\u003e\u003cspan address=\"10.1101/2023.02.09.23285705\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJahagirdar D, Walters MK, Novotney A, Brewer ED, Frank TD, Carter A, et al. Global, regional, and national sex-specific burden and control of the HIV epidemic, 1990\u0026ndash;2019, for 204 countries and territories: the Global Burden of Diseases Study 2019. Lancet HIV. 2021;8(10):e633\u0026ndash;51. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/S2352-3018(21)00152-1\u003c/span\u003e\u003cspan address=\"10.1016/S2352-3018(21)00152-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMaranh\u0026atilde;o TA, Alencar CH, De Avelar Figueiredo Mafra Magalh\u0026atilde;es M, Sousa GJB, Ribeiro LM, De Abreu WC, et al. Mortality due to acquired immunodeficiency syndrome and associated social factors: a spatial analysis. Rev Bras Enferm. 2020;73(Suppl 5):e20200002. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.scielo.br/j/reben/a/KtDh5ZsfwRjmNgg3tpq54rB/?lang=en\u003c/span\u003e\u003cspan address=\"https://www.scielo.br/j/reben/a/KtDh5ZsfwRjmNgg3tpq54rB/?lang=en\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cdiv class=\"gridtable\"\u003e\n \u003cdiv class=\"colspec\" align=\"left\"\u003e\u003cbr\u003e\u003c/div\u003e\n \u003ctable id=\"Tab2\" border=\"1\"\u003e\n \u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Italic\"\u003eDataset Ocerview\u003c/span\u003e\u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eFile Name\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eDescription\u003c/div\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003ehealth_gha.csv\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRaw dataset used for data quality checks and feature engineering\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eghana_infectious_disease_model_dataset_cleaned.csv\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eFully cleaned dataset used for EDA and predictive modeling\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003emodel_performance_metrics.csv\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eTable storing model accuracy results across algorithms for HIV/AIDS\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003ctable id=\"Tab3\" border=\"1\"\u003e\n \u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Italic\"\u003eDetailed Data Quality Assessment Summary\u003c/span\u003e\u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eFetures\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eData Type\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMissing Values\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMissing %\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eUnique Values\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOutliers Detcted\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003ePost Winsorisation\u003c/div\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eregion\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eobject\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e10\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003edate\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eobject\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e612\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003etb_outlier\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003ebool\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003emalaria_outlier\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003ebool\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e2\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e379\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eurban_rural_sum_pct\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eage_sum_pct\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e3\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e57\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003emigration_rate\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9261\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eN/A\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eregional_stigma_index\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6579\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eurbanization_level\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9792\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e112\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003ehealth_facility_density\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9761\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eN/A\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003etesting_coverage_pct\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9792\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eN/A\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eaccess_to_art_pct\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9792\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003ehiv_awareness_index\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9792\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eyouth_unemployment_rate\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9792\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003efemale_literacy_rate\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9792\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e52\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003econdom_use_rate\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9792\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e69\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eeducation_access_index\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9792\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e45\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eregion_fixed\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eobject\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e10\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003epopulation_rural_pct\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9531\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e52\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003epopulation_urban_pct\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9531\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eN/A\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003epopulation_15_64_pct\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9792\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e334\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003epopulation_65_plus_pct\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9781\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eN/A\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003epopulation_0_14_pct\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9751\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003epopulation_total\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9631\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e44\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003ehiv_incidence\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9598\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e53\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003etb_incidence\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9598\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003emalaria_incidence\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003efloat64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9598\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e59\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003emonth\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eint64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e12\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eN/A\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eyear\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eint64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e51\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e92\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003ehiv_outlier\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003ebool\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e2\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable id=\"Tab4\" border=\"1\"\u003e\n \u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Italic\"\u003eAdditional Features integrated to enrich\u003c/span\u003e \u003cspan class=\"BoldItalic\"\u003ehealth_gha.csv\u003c/span\u003e\u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eFeature Name\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eDescription\u003c/div\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eeducation_access_index\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e% of population with secondary education or higher\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003econdom_use_rate\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003ePercentage consistently using condoms\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003efemale_literacy_rate\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eLiteracy among women aged 15+\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eyouth_unemployment_rate\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eYouth (15\u0026ndash;24) unemployment rate\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003ehiv_awareness_index\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eComposite score measuring HIV knowledge and awareness\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eaccess_to_art_pct\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e% of HIV-positive individuals receiving ART\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003etesting_coverage_pct\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e% of population tested for HIV\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003ehealth_facility_density\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eNumber of health facilities per 10,000 people\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eregional_stigma_index\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0\u0026ndash;1 index quantifying HIV-related stigma in regions\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eurbanization_level\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e% of population living in urban areas\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003emigration_rate\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eNet migration rate per 1,000 people\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u0026nbsp;\u003ctable id=\"Tab5\" border=\"1\"\u003e\n \u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Italic\"\u003eFeature Engineering and Pre-processing Techniques\u003c/span\u003e\u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eTechnique\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eDescription\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eCategory\u003c/div\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eLagged Variables\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eCreated 1\u0026ndash;12 month lags for hiv_incidence to capture autocorrelation and temporal dependencies. Useful for time series models (ARIMA, LSTM).\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eFeature Engineering\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eComposite Indices\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eNormalized and combined related variables (e.g., ART access, stigma index, education) into single indices scaled from 0 to 1 for comparability.\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eFeature Engineering\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eBinned Features\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eTransformed continuous variables like \u003cspan class=\"Bold\"\u003eurbanization_leve\u003c/span\u003el and \u003cspan class=\"Bold\"\u003eaccess_to_art_pct\u003c/span\u003e into categories (low, medium, high) for stratified analysis.\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eFeature Engineering\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRolling Averages\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eApplied 6-month moving averages to smooth out noise in volatile behavioral variables (e.g., \u003cspan class=\"Bold\"\u003econdom_use_rate\u003c/span\u003e, \u003cspan class=\"Bold\"\u003ehiv_awareness_index\u003c/span\u003e).\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003ePreprocessing\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u0026nbsp;\u003ctable id=\"Tab6\" border=\"1\"\u003e\n \u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Italic\"\u003eModel Selection\u003c/span\u003e\u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eModel\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRationale\u003c/div\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRandom Forest Regressor\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eSelected as the baseline model due to its robustness with non-linear patterns, interpretability via feature importance, and ability to handle missing values and noisy inputs.\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eXGBoost Regressor\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eChosen for its superior accuracy and performance with structured datasets. It includes regularization (L1 and L2) to prevent overfitting and handles missing data natively.\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRidge Regressor\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eUsed as a linear benchmark with L2 regularization to manage multicollinearity and improve generalization in high-dimensional feature spaces.\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eSupport Vector Regressor (SVR)\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eImplemented for its robustness with small-to-medium sized datasets and capability to model non-linear relationships via kernel functions. Particularly useful for exploring residual variance and generalization.\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eLSTM (Optional \u0026ndash; Future Work)\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eAlthough not implemented in this final version, LSTM models are ideal for modeling long-term dependencies in time series data and are recommended for future iterations with univariate or multivariate temporal windows.\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cdiv class=\"gridtable\"\u003e\n \u003cdiv class=\"colspec\" align=\"left\"\u003e\u0026nbsp;\u003c/div\u003e\u0026nbsp;\u003ctable id=\"Tab7\" border=\"1\"\u003e\n \u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Italic\"\u003eHyperparameter Tuning and Optimal CV Strategy\u003c/span\u003e\u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eModel\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eTuned Hyperparameters\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eValues Tried\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOptimal CV Strategy\u003c/div\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRandom Forest\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003en_estimators, max_depth, min_samples_split\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003e[100, 200], [10, 20, None]\u003c/span\u003e, [\u003cspan class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e5\u003c/span\u003e]\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e5-Fold Stratified Cross-Validation\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eXGBoost\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003en_estimators, learning_rate, max_depth\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003e[100, 200], [0.01, 0.1]\u003c/span\u003e, [\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e10\u003c/span\u003e]\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e5-Fold Stratified Cross-Validation\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRidge Regression\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003ealpha\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003e[0.01, 0.1, 1.0, 10]\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e10-Fold Cross-Validation\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eSupport Vector Regression (SVR)\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003eC, epsilon, kernel\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e[\u003cspan class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e10\u003c/span\u003e], \u003cspan class=\"Bold\"\u003e[0.1, 0.2], [\u0026apos;rbf\u0026apos;]\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e5-Fold Cross-Validation\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cdiv class=\"gridtable\"\u003e\n \u003cdiv class=\"colspec\" align=\"char\"\u003e\u0026nbsp;\u003c/div\u003e\u003cbr\u003e\u0026nbsp;\u003ctable id=\"Tab8\" border=\"1\"\u003e\n \u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Italic\"\u003eModel Performance Results\u003c/span\u003e\u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRank\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eModel\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRMSE\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eR\u0026sup2;\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMAE\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMAPE\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eBest Hyper parameter\u003c/div\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1st\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRandom Forest\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.39\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0.9821\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e4.64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e2.56%\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003e{\u0026apos;max_depth\u0026apos;: 20, \u0026apos;min_samples_split\u0026apos;: 2, \u0026apos;n_estimators\u0026apos;: 200}\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e2nd\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eXGBoost\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.54\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0.9813\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e4.77\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e2.64%\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003e{\u0026apos;learning_rate\u0026apos;: 0.1, \u0026apos;max_depth\u0026apos;: 10, \u0026apos;n_estimators\u0026apos;: 100}\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e3rd\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRidge Regression\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e9.08\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0.9640\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e6.69\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e3.69%\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003e{\u0026apos;alpha\u0026apos;: 1}\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e4th\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eSupport Vector Regressor (SVR)\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e8.47\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0.9686\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e5.90\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e3.17%\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003e{\u0026apos;C\u0026apos;: 10, \u0026apos;epsilon\u0026apos;: 0.2, \u0026apos;kernel\u0026apos;: \u0026apos;rbf\u0026apos;}\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003c/caption\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003e8\u003c/strong\u003e\u003cem\u003e\u0026nbsp;Key Feature Importance and SHAP Summary for hiv_incidence\u003c/em\u003e\u003c/p\u003e\n\u003ctable border=\"1\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 208px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eFeature\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 100px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eSHAP Score\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 248px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eKey Interaction Highlighted\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 208px;\"\u003e\n \u003cp\u003eSHAP for \u003cstrong\u003econdom_use_rate\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 100px;\"\u003e\n \u003cp\u003e0.127\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 248px;\"\u003e\n \u003cp\u003ePositive effect increases with literacy\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 208px;\"\u003e\n \u003cp\u003eSHAP for \u003cstrong\u003eeducation_access_index\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 100px;\"\u003e\n \u003cp\u003e0.241\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 248px;\"\u003e\n \u003cp\u003eNonlinear jump effect around score ~55\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 208px;\"\u003e\n \u003cp\u003eSHAP for \u003cstrong\u003ehiv_awareness_index\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 100px;\"\u003e\n \u003cp\u003e0.198\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 248px;\"\u003e\n \u003cp\u003eSteep increase above awareness ~65\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 208px;\"\u003e\n \u003cp\u003eSHAP for \u003cstrong\u003etb_incidence\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 100px;\"\u003e\n \u003cp\u003e0.113\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 248px;\"\u003e\n \u003cp\u003eModerate rise, esp. with higher urbanization\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003ctable id=\"Tab10\" border=\"1\"\u003e\n \u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 9\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Italic\"\u003e2030 Forecast and Result Summary\u003c/span\u003e\u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eModel\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRMSE\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e2030 Forecast Value\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eTrend Direction (\u0026rarr; \u0026uarr; \u0026darr;)\u003c/div\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eNotes\u003c/div\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRandom Forest\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0.55\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e231.64\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003e\u0026rarr;\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eStable forecast; moderate accuracy\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eXGBoost\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e0.47\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e231.32\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003e\u0026rarr;\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eMost accurate; closely aligns w/ SDG\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eSVR\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1.30\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e230.34\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003e\u0026darr;\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eSlight downward trend; less precise\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eRidge Regression\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e1.24\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e247.62\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003e\u0026uarr;\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eOverestimates 2030 value\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eEnsemble\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u0026mdash;\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e235.23\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003e\u003cspan class=\"Bold\"\u003e\u0026uarr;\u003c/span\u003e\u003c/div\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cdiv class=\"SimplePara\"\u003eBalanced output across models\u003c/div\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-public-health","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"pubh","sideBox":"Learn more about [BMC Public Health](http://bmcpublichealth.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/pubh/default.aspx","title":"BMC Public Health","twitterHandle":"@BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"HIV/AIDS, Ghana, Machine Learning, Forecasting, Public Health, Epidemiology, Ensemble Models, XGBoost, Random Forest, SHAP","lastPublishedDoi":"10.21203/rs.3.rs-6639193/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6639193/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground:\u003c/strong\u003e In Ghana, the precise forecasting of Human Immunodeficiency Virus (HIV) infection and Acquired Immunodeficiency Syndrome (AIDS) is essential for public health strategies due to the intricate socio-structural factors that affect the transmission patterns of the virus. Public health planning becomes challenging because conventional linear statistical models do not take into account disjointed or multifactorial data.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods:\u003c/strong\u003e This used retrospective observational data covering the ten administrative regions of Ghana from 2000 to 2022. Four machine learning algorithms (Random Forest, XGBoost, Ridge Regression, and Support Vector Regression) were applied to forecast HIV/AIDS incidence. The dataset incorporated epidemiological trends, demographic profiles, and healthcare infrastructure indicators. Preprocessing steps included KNN imputation for missing healthcare infrastructure values and winsorization of disease incidence variables to reduce outlier bias.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults:\u003c/strong\u003e Ridge Regression and Support Vector Regression performed poorly compared to XGBoost and Random Forest, with correlation coefficients above 0.98, indicating high predictive accuracy. HIV incidence in Ghana was foretasted to stabilize at 231 cases per 100,000 individuals by 2030. The SHAP (SHapley Additive exPlanations) analysis revealed that HIV awareness, access to antiretroviral therapy, poverty rates, and access to education were significant factors influencing incidence trends.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDiscussion: \u003c/strong\u003eEnsemble machine learning models yielded more reliable predictions than conventional linear models. The predicted incidence plateau indicates that current intervention strategies may not reach the national and global reduction targets, thereby underscoring the need for more vigorous public health initiatives. Data limitations restrict real-time predictions and necessitate ongoing enhancements to the data infrastructure.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion: \u003c/strong\u003eThis study demonstrates that ensemble machine learning is a viable and valuable tool for predicting HIV incidence rates in Ghana, providing reliable results to guide public health decision making and resource management. However, its effectiveness may be influenced by data quality limitations and contextual complexities, underscoring the need for expanded data-sharing infrastructure, cautious interpretation and continuous model validation in low-resource settings, \u0026nbsp;as emphasized by ongoing calls for transparency in ML epidemiology.\u003c/p\u003e","manuscriptTitle":"Forecasting HIV/AIDS Incidence in Ghana: A Retrospective Observational Study Using Ensemble Machine Learning Models ","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-05-29 11:51:55","doi":"10.21203/rs.3.rs-6639193/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"176103242372054291839202105591725956659","date":"2026-05-10T17:35:29+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-09T18:04:41+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"102298232569817261210416626891577721043","date":"2026-05-07T12:29:22+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"324893108947270495221268278015170832381","date":"2026-05-07T10:34:32+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"62184644583345523779182873680436870422","date":"2025-06-24T09:50:53+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-06-23T09:17:30+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"242385206044691805619371567448099828432","date":"2025-06-20T09:20:18+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"137103891425301296955338406600604637727","date":"2025-06-19T06:34:19+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-06-12T04:57:18+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-05-16T08:22:31+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-05-15T06:59:01+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-05-15T06:54:40+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Public Health","date":"2025-05-11T10:57:23+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-public-health","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"pubh","sideBox":"Learn more about [BMC Public Health](http://bmcpublichealth.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/pubh/default.aspx","title":"BMC Public Health","twitterHandle":"@BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"082e6fba-de8e-45cd-86e0-8c724e97f61e","owner":[],"postedDate":"May 29th, 2025","published":true,"recentEditorialEvents":[{"type":"reviewerAgreed","content":"176103242372054291839202105591725956659","date":"2026-05-10T17:35:29+00:00","index":94,"fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-09T18:04:41+00:00","index":92,"fulltext":""},{"type":"reviewerAgreed","content":"102298232569817261210416626891577721043","date":"2026-05-07T12:29:22+00:00","index":91,"fulltext":""},{"type":"reviewerAgreed","content":"324893108947270495221268278015170832381","date":"2026-05-07T10:34:32+00:00","index":90,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2025-06-12T05:08:18+00:00","versionOfRecord":[],"versionCreatedAt":"2025-05-29 11:51:55","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6639193","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6639193","identity":"rs-6639193","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0