Predicting Tuberculosis Incidence in Adult HIV Patients on ART in Debre Markos, Ethiopia: A Machine Learning Approach

preprint OA: closed
📄 Open PDF Full text JSON View at publisher

Abstract

ABSTRACT Tuberculosis (TB) is the commonest comorbidity among individuals with HIV/AIDS, especially in low- and middle-income nations such as Ethiopia. Early diagnosis of TB infection in HIV-infected patients is crucial for effective management of opportunistic infections that can result in mortality. Early identification of TB in HIV/AIDS patients plays a significant role in reducing morbidity and mortality. Machine learning algorithms have a significant role in detecting TB occurrences among HIV/AIDS patients. In this study, we used 5,392 HIV-infected individuals’ medical records. Techniques such as SMOTE and ADASYN were employed to adjust data imbalance between positive and negative TB status. Random forest, decision tree, logistic regression, gradient boosting, K-nearest neighbors, and XGBoost were evaluated to predict TB incidence. Of the total records, 3,440 (63.8%) were female patients, and the remaining 1,952 (36.2%) were male patients. 3,715 (68.9%) were labeled as green records addresses, while 1,677 (31.1%) had yellow records. The XGBoost algorithm is the best-performing model to predict TB incidence. Among the features included in this study is the most important classifier for feature selection. Among all the features, CD4 count and patient age were found to be the most important predictors of TB incidence among adult HIV patients. This study demonstrates that the XGBoost model was the most effective model for predicting tuberculosis incidence among HIV patients, utilizing features such as low CD4 counts, age, duration on ART, weight, sex, and WHO clinical stage. Author Summary Tuberculosis is the main opportunistic infection among HIV-infected individuals. The comorbidities of HIV and TB increase the risk of mortality among HIV patients. Early detection of TB infection from people living with HIV is a crucial step for increasing HIV patients’ life expectancy. A machine learning approach plays a vital role in detecting TB infection among HIV patients. This study aims to predict tuberculosis (TB) occurrence in adult HIV patients on antiretroviral medication (ART) in Debre Markos City, Ethiopia. Using a retrospective dataset of 5,392 patients, researchers compared seven methods. The XGBoost model fared best, earning 82% accuracy and a 90% AUC after resolving class imbalance using MOTE+ENN. Key predictors revealed were low CD4 count, age, time on ART, sex, WHO clinical stage, address status, DSD category, and TB preventative treatment (TPT status). Machine learning models will help health providers to predict the risk of TB infection and make early intervention in high TB-HIV co-infection loads.
Full text 42,176 characters · extracted from preprint-html · click to expand
Predicting Tuberculosis Incidence in Adult HIV Patients on ART in Debre Markos, Ethiopia: A Machine Learning Approach | medRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-P4HH5NV'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search Predicting Tuberculosis Incidence in Adult HIV Patients on ART in Debre Markos, Ethiopia: A Machine Learning Approach Desalegn Meseret Tadele , View ORCID Profile Getaye Tizazu Biwota , Lijalem Megibaru Enyew , Maru Meseret Tadele , Gizaw Hailiye Teferi doi: https://doi.org/10.1101/2025.08.24.25334330 Desalegn Meseret Tadele 1 Health informatics department, college of health science, Debre Markos University , Ethiopia Find this author on Google Scholar Find this author on PubMed Search for this author on this site Getaye Tizazu Biwota 1 Health informatics department, college of health science, Debre Markos University , Ethiopia Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Getaye Tizazu Biwota For correspondence: getaye_tizazu{at}dmu.edu.et Lijalem Megibaru Enyew 2 Institute of technology, Debre Markos University , Ethiopia Find this author on Google Scholar Find this author on PubMed Search for this author on this site Maru Meseret Tadele 1 Health informatics department, college of health science, Debre Markos University , Ethiopia Find this author on Google Scholar Find this author on PubMed Search for this author on this site Gizaw Hailiye Teferi 1 Health informatics department, college of health science, Debre Markos University , Ethiopia Find this author on Google Scholar Find this author on PubMed Search for this author on this site Abstract Full Text Info/History Metrics Data/Code Preview PDF ABSTRACT Tuberculosis (TB) is the commonest comorbidity among individuals with HIV/AIDS, especially in low- and middle-income nations such as Ethiopia. Early diagnosis of TB infection in HIV-infected patients is crucial for effective management of opportunistic infections that can result in mortality. Early identification of TB in HIV/AIDS patients plays a significant role in reducing morbidity and mortality. Machine learning algorithms have a significant role in detecting TB occurrences among HIV/AIDS patients. In this study, we used 5,392 HIV-infected individuals’ medical records. Techniques such as SMOTE and ADASYN were employed to adjust data imbalance between positive and negative TB status. Random forest, decision tree, logistic regression, gradient boosting, K-nearest neighbors, and XGBoost were evaluated to predict TB incidence. Of the total records, 3,440 (63.8%) were female patients, and the remaining 1,952 (36.2%) were male patients. 3,715 (68.9%) were labeled as green records addresses, while 1,677 (31.1%) had yellow records. The XGBoost algorithm is the best-performing model to predict TB incidence. Among the features included in this study is the most important classifier for feature selection. Among all the features, CD4 count and patient age were found to be the most important predictors of TB incidence among adult HIV patients. This study demonstrates that the XGBoost model was the most effective model for predicting tuberculosis incidence among HIV patients, utilizing features such as low CD4 counts, age, duration on ART, weight, sex, and WHO clinical stage. Author Summary Tuberculosis is the main opportunistic infection among HIV-infected individuals. The comorbidities of HIV and TB increase the risk of mortality among HIV patients. Early detection of TB infection from people living with HIV is a crucial step for increasing HIV patients’ life expectancy. A machine learning approach plays a vital role in detecting TB infection among HIV patients. This study aims to predict tuberculosis (TB) occurrence in adult HIV patients on antiretroviral medication (ART) in Debre Markos City, Ethiopia. Using a retrospective dataset of 5,392 patients, researchers compared seven methods. The XGBoost model fared best, earning 82% accuracy and a 90% AUC after resolving class imbalance using MOTE+ENN. Key predictors revealed were low CD4 count, age, time on ART, sex, WHO clinical stage, address status, DSD category, and TB preventative treatment (TPT status). Machine learning models will help health providers to predict the risk of TB infection and make early intervention in high TB-HIV co-infection loads. Introduction HIV is an infectious virus that primarily undermines the immune system by infecting specifically the CD4 cells, or T cells, which are responsible for combating various infections. The weakening of the body’s defense system through this process predisposes one to opportunistic diseases (OIs), such as tuberculosis [ 1 ]. Acquired Immune Deficiency Syndrome (AIDS) is a condition that occurs secondary to HIV infection. Tuberculosis (TB) is one of the common opportunistic infections that affect people with HIV/AIDS when compared with non-HIV-infected individuals [ 2 , 3 ]. Tuberculosis is the primary opportunistic illness in individuals infected with HIV, and the likelihood of developing active tuberculosis markedly increases owing to the HIV virus. HIV compromises the immune system, increasing the likelihood of tuberculosis infection advancing from latent to active phases, hence rendering co-infected patients more susceptible to mortality [ 4 ]. Globally, HIV-positive adults are highly susceptible to TB, and an estimated 208,000 ALWHIV died from TB in 2020. Sub-Saharan Africa, including Ethiopia, experiences high co-infection rates due to the synergy between HIV and TB, where declining immunity from HIV increases the risk of developing active TB [ 5 ]. Several studies locally have been conducted to determine the incidence and predictors of TB among adult HIV patients using classical statistical methods. Which primarily focus on learning factors behind TB cases that already exist. While they identify general risk factors, they do not offer personalized risk assessments. Additionally, they are static in nature because any patient data or new information does not update the analysis automatically. Moreover, the results are not as easily translated into real-time clinical applications. To overcome these challenges, this study used machine learning (ML) models to predict the incidence and identify the predictors of TB among adult HIV patients. ML emerged as a promising approach for analyzing complex health data and forecasting disease incidence [ 6 ]. Debre Markos City hosts numerous public health facilities providing ART services to adult HIV patients. Even though the facilities in the city have a strong electronic health record, none of them used the historical data to develop a model for predicting the risk of incidence of TB prior to the occurrence of the event. Therefore, the aim of this research is to develop a model for predicting the risk of TB in adult HIV patients on ART in public health facilities in Debre Markos City using past data. Therefore, the aim of this research is to develop a model for predicting the risk of TB in adult HIV patients on ART in public health facilities in Debre Markos city using past data. Adults on ART who develop TB may also have more advanced disease and complications, including respiratory distress and systemic illness, due to the dual burden of infection. Mortality risk was considerably high in TB patients who were left untreated compared to those who did not have TB [ 7 ]. HIV-TB co-infection considerably diminishes the health-related quality of life in adult HIV patients in the areas of physical health, mental health, and social functioning. Both diseases can cause stigma, which may result in social isolation and mental illness. The onset of TB may disrupt adherence to antiretroviral treatment (ART) because of the complexities of managing many drugs and the side effects of TB treatment [ 8 ]. Active TB patients are still infectious while on ART, which makes it even more community transmitted. In high HIV prevalence settings, the co-occurrence of TB can cause outbreaks, thus making it a public health issue. It requires a lot of healthcare services to treat dual-diagnosed patients, including longer hospital stays, additional diagnostics, and specialized care teams. TB development on ART patients presents significant challenges, both to individual health outcomes and to public health efforts at large. Addressing these issues requires coordinated healthcare strategies, ongoing research, and comprehensive support systems for affected individuals. This study aims to predict the incidence of TB among adult HIV patients and evaluate which model predicts accurately. Machine learning has emerged as a powerful force in healthcare as a predictor of disease outcomes. The innovation in machine learning and data analytics within the healthcare sector creates new possibilities in the optimization of patient care. Machine learning algorithms can process intricate data sets, thereby detecting patterns and connections that standard statistical techniques may fail to capture. Through the integration of various data sources and the detection of complicated patterns, the algorithms provide a robust alternative that results in higher predictive power. Hence, these innovative methods can facilitate a more precise targeting of individuals in adult HIV cohorts who are at risk of developing tuberculosis [ 9 ]. Capacity to process big and intricate datasets enables detection of non-linear relationships and interdependencies between variables, which might not be caught using conventional statistical techniques. Machine learning in medicine has gained speed; most recently, for disease prediction, it demonstrated the potential of machine learning models for predicting TB in adult HIV patients, indicating that these technologies can enhance clinical decision-making [ 10 ]. Machine learning models can process big datasets to identify patterns that are not visible to traditional approaches, representing a promising avenue for patient outcome enhancement. Ethiopia is one of the low-income countries that is heavily affected by both tuberculosis (TB) and HIV, with the country featuring in the top 30 of the world’s TB incidence rate, in addition to a high coincidence of TB and HIV cases [ 11 ]. Within this framework, the Debre Markos health facility functions as an essential health provider, addressing the needs of at-risk populations for these conditions. Researchers have revealed that there are significant challenges in the management of patients with co-infections, ranging from restrictions on resources, inadequate screening protocols, and limitations on data collection. The literature demands that there is a pressing need to tackle the TB and HIV co-infection with special reference to resource-poor settings such as Ethiopia [ 12 ]. The application of machine learning techniques is a promising way forward to add to predictive effectiveness and optimize patient care [ 13 ]. Methods Study design, setting, and study population An institution-based retrospective cross-sectional study design was employed at Debre Markos Comprehensive Specialized Hospital and Debre Markos Health Center in Debre Markos City, Ethiopia. Records of all adult HIV patients who ever started ART in public health facilities of Debre Markos city were the source population for this study, while complete records of all adult HIV patients in selected public health facilities at Debre Markos city were the study population for this study. Records of all adult HIV patients who ever started ART in selected public health facilities and patients who received/are receiving ART for at least 6 months were included in the study. Meanwhile, records of all adult HIV patients who already developed TB at entry to ART and had an incomplete record were excluded from the study. Data source The data source for this study was an institutional database usually called Smart Care. Healthcare providers manually record patient information on intake and follow-up forms at the time of ART enrollment and during follow-up, respectively. Data clerks also entered patient information from both the intake and follow-up forms into Smart Care. Description of the dataset The dataset included records of all HIV-positive individuals who have ever initiated ART at the respective public health facilities. It contained socio-demographic, clinical, and medication-related information for all HIV patients who had ever started ART in these facilities. In total, 3500 adults have ever started HIV care at DMCSH, and 1892 at Debre Markos Health Center. Therefore, the study focuses on adult HIV patients who have ever initiated ART at these two centers, with a combined total of 5392 individuals who have started ART. Study variables The outcome of interest is the TB status of the patient, which is described as positive (1) or negative (2). Variables that affect the incidence of TB among HIV patients are grouped as sociodemographic, baseline clinical and laboratory characteristics such as nutritional status, and medication-related characteristics, which include type of ART regimen, TPT status, duration on ART, and adherence to ART. Operational definitions Viral load: Suppressed if the viral load is 1000 copies per ml for clinical intervention [ 14 ]. Level of adherence to ART drugs in this study is classified as “good” (≥95% adherence or missing 1 out of 30 doses or missing 2 out of 60 doses). Fair (85–94% adherence or missing 2–4 out of 30 doses or missing 4–9 out of 60 doses), Poor (less than 85% or missing ≥5 doses out of 30 doses or ≥10 doses out of 60 doses) [ 14 ]. CD4: severe immune suppression, 200–349 cells/mm³ moderate immune suppression, 350–499 cells/mm³ mild immune suppression, and CD4 counts of 500 cells/mm³ or higher as normal immune function [ 14 ]. Modeling framework Data collection Data collection began immediately after receiving ethical approval from the Debre Markos University ethical review board. Information was gathered from Debre Markos Health Center and Debre Markos Comprehensive Specialized Hospital. Two data collectors were assigned to extract data from the HIV database at these facilities. The data was initially extracted in Excel format and then converted to CSV format for compatibility with the Python environment. Data Cleaning and Preprocessing Handling Missing values: Columns such as fluconazole start date, fluconazole end date, and specimen types contained no records and were consequently removed. Meanwhile, columns for address status, WHO clinical stage, months, and ICT status had a combined 1.8% missing values, which were filled using mean and mode imputation. Data encoding: To enhance data readability, one-hot encoding was applied to categorical variables, including sex, address status, TPT status, viral load status, DSD category, ICT status, follow-up, WHO stage, adherence, nutritional status, and TB incidence. Handling class imbalance Class imbalance in machine learning refers to a situation where the distribution of instances across different classes is not uniform, leading to a disproportionate representation of one class over another. This imbalance can significantly affect the performance of machine learning models, particularly in classification tasks, as models may become biased towards the majority class and fail to accurately predict the minority class [ 15 ]. The dataset contained 4,178 instances (77.5%) in the negative class, while the remaining 1,214 instances (22.5%) were in the positive class. According to the study conducted by Mulugeta et al., if the percentage of data in the minority class is between 20-40%, 1-20%, or less than 1% of the dataset, the class imbalance is considered mild, moderate, or extreme, respectively [ 16 ]. Based on this evidence, the dataset exhibited a mild class imbalance, which required the application of techniques to address this issue. Under-sampling, SMOTE, SMOTE+ENN, SMOTE+Tomek, Tomek Links, and ADASYN are all techniques available to handle class imbalance in datasets, particularly in machine learning [ 17 , 18 ]. As a result, this study compared each technique with the random forest algorithm using accuracy and AUC as evaluation metrics. Feature Selection Feature selection refers to the process of selecting a subset of the most relevant features (variables) from the original set of features to use in model training. The goal is to reduce the dimensionality of the data, improve model efficiency, and avoid overfitting [ 19 ]. Filter methods, including the use of a correlation matrix, are commonly employed in machine learning to select relevant features. In this research, the correlation matrix was utilized as part of the filter method to identify and eliminate highly correlated features, ensuring that only the most informative and non-redundant features were retained for model training. This approach helps to improve model performance and reduce overfitting by focusing on the most relevant variables. Data Split The technique of splitting a dataset into distinct subsets for testing and training is known as “data splitting” in machine learning. Data is usually divided into two sets: a training set for training the model and a test set for evaluating the model’s performance and ability to generalize to new data. Depending on the size of the dataset and the issue at hand, split ratios can vary, but typically they are 80% for training and 20% for testing [ 20 ]. Consequently, this research divided the dataset into 80% (4314) for training and 20% (1078) for testing. Model selection Model selection in machine learning is the process of choosing the most appropriate algorithm or model for a given problem based on factors like the nature of the data, the task at hand, and performance metrics [ 17 ]. In this study, seven supervised classification algorithms, including support vector machine, random forest, decision tree, logistic regression, gradient boosting, K-nearest neighbors, and XGBoost, were selected. After the selection process, each model was trained, and the one with the highest performance metrics was chosen ( Figure 1 ). Download figure Open in new tab Figure 1: ML modeling framework for predicting incidence of TB among adult HIV patients on ART Download figure Open in new tab Figure 2: Heat map displaying feature correlation in a study conducted in public health facilities in Debremarkos city, Ethiopia Download figure Open in new tab Figure 3: Description of classes before and after applying SMOTE+ENN technique in the adult HIV patients’ dataset from public health facilities of Debre Markos city, 2024 Evaluation measures After model training, the model’s performances were evaluated and compared to each other. The performance of the prediction models was evaluated based on the confusion matrix. This study used precision, sensitivity, specificity, the F1-score, and the area under the receiver operating characteristic (AUC-ROC) to assess the model performance. AUC-ROC is a popular and powerful performance metric for assessing the performance of binary classifiers and is used to evaluate the model’s predictive classification capacity [ 17 ]. The matrix is made up of predictions that have been summarized into a total number of correct and incorrect predictions [ 6 ]. Results Description of socio-demographic predictors The study included 5,392 records of adult HIV patients who had ever initiated ART. The average age of the participants was 42 years, with a standard deviation of ±11 years. Among the records, 4,359 (80.8%) were of patients aged 44 years or older, while 1,033 (19.2%) were between the ages of 18 and 43. Of the total records, 3,440 (63.8%) were from female patients, and 1,952 (36.2%) were from male patients. In terms of the patients’ addresses, 3,715 (68.9%) were associated with green records, and 1,677 (31.1%) had yellow records. Description of clinical predictors Out of the total records of adult HIV patients, 4,402 (81.6%) had a normal nutritional status, while 630 (11.7%) were classified as overweight and 360 (6.7%) as undernourished. In terms of WHO clinical stages, 4085 (75.8%) were at stage 1, 549 (10.2%) at stage 2, 538 (10%) at stage 3, and 220 (4%) at stage 4. Regarding CD4 count, 2478 (46%) had counts of ≥500 cells/mm³, 2374 (44%) had counts between 350 and 499 cells/mm³, and 540 (10%) had counts between 200 and 349 cells/mm³. Of the total, 5148 (95.5%) had a suppressed viral load, while 244 (4.5%) had an unsuppressed viral load. Description of ART and medication-related predictors In terms of the regimen type, 98% (5278) of the participants were on first-line drugs, while 2% (114) were on second-line drugs. Regarding the TPT (Tuberculosis Preventive Treatment) status, the records for 51.7% (2781) of participants were categorized as gold, 47.8% (2579) as bronze, and 0.6% (32) as silver. Exploratory data analysis For feature selection, a correlation matrix was used to identify and visualize features that were highly correlated ( table 2 ). The findings clearly show a class imbalance between the two instances (positive and negative). As a result, this study evaluated several techniques—under-sampling, SMOTE, SMOTE+ENN, SMOTE+Tomek, Tomek Links, and ADASYN—based on accuracy and AUC. Ultimately, the SMOTE+ENN technique was chosen to balance the classes ( Table 1 ). View this table: View inline View popup Download powerpoint Table 1: Comparison of class imbalance handling techniques for predicting TB incidence among adult HIV patients Model training and evaluation This study employed 7 machine learning algorithms, using 80% of the dataset for training. Among them, XGBoost performed better than the other algorithms ( Table 2 ). View this table: View inline View popup Download powerpoint Table 2: Comparison of algorithms based on accuracy, precision, recall, and F1 score before and after applying SMOTE +ENN using the adult HIV patients’ dataset from public health facilities of Debre Markos City, 2024 The relationship between false positive and true positive rates is shown in the ROC curve, which states that a predictive model was more accurate for the prediction of TB incidence ( Figure 4 ). Download figure Open in new tab Figure 4: ROC Curve comparing the performance of algorithms using the adult HIV patients’ dataset from public health facilities in Debre Markos City, 2024 Model testing To assess the performance of the selected model, it was tested on unseen data, which comprised 1078 instances (20% of the dataset). The model generated a confusion matrix, which included predicted positive vs. actual positive, predicted negative vs. actual negative, and predicted negative vs. actual positive instances ( Figure 5 ). Download figure Open in new tab Figure 5: Confusion matrix for XGBoost to predict incidence of TB among adult HIV patient on ART in public health facilities of Debre Markos city, 2024 Download figure Open in new tab Figure 6: Important features selected by XGBoost feature importance to predict TB incidence among adult HIV patients on ART in public health facilities of Debre Markos, 2024 Predictors of incidence of TB This study employed the XGBoost classifier for feature importance selection. Among all the features, CD4 count, age, months on ART, sex, WHO clinical stage, address status, DSD category, and TPT were found to be predictors of incidence TB Discussion Predicting the occurrence of TB and determining its predictors among adult HIV patients at public health institutions in Debre Markos, Ethiopia, was the goal of this study. This was accomplished by evaluating seven machine learning methods for TB prediction among HIV-infected people; XGBoost was found to outperform over other techniques. These results were consistent with research conducted in Taiwan and China, where the XGBoost algorithm was found to be the most successful classifier [ 13 , 21 ]. The model that was selected for this study, XGBoost, correctly identified 82% of all instances in the training dataset—both positive and negative—as actual positive or actual negative. This outcome nearly aligns with a Chinese study in which 84% of the training dataset’s occurrences were correctly classified by a similar model [ 21 ]. Nonetheless, it surpasses a Taiwanese study in which 70.5% of the training dataset was properly classified by the system [ 13 ]. Differences in the source population, algorithms included in the training, data splitting techniques, included characteristics, and dataset size may all be responsible for the discrepancies in classification performance. The XGBoost model in this research correctly classified 84.2% of predicted positives out of all actual positives, achieving a recall (sensitivity) of 84.2%. This performance surpasses the outcome of a Chinese study in which 71% of the true positives were correctly classified by the algorithm [ 21 ]. However, the outcome is comparable to a study conducted in Taiwan, where 81.2% of the training dataset’s actual positives were correctly classified as predicted positives [ 13 ]. Given that both this study and the Taiwanese study used training datasets that were comparatively bigger than the Chinese study, the discrepancies in the results between these three studies could be the result of disparities in dataset sizes. The AUC in this study is 90%, indicating a 90% probability that the model will correctly rank a randomly selected positive instance higher than a randomly selected negative one. This result is slightly higher than a study conducted in Taiwan, where the model demonstrated an 86.2% probability of correctly distinguishing a randomly chosen positive instance from a randomly chosen negative one [ 13 ]. This difference could be explained by variations in the dataset size. The dataset used in this study was approximately twice as large as the one used in the Taiwanese research. CD4 count, age, duration on ART, sex, WHO clinical stage, address status, DSD category, and TPT were identified as predictors of tuberculosis incidence among adult HIV patients in public health facilities in Debre Markos city. Machine learning methods used for predicting tuberculosis risk factors highlighted age and length of ART as significant risk factors for TB [ 22 ]. Additional research conducted in Ethiopia using statistical modeling found that a low CD4 count, not taking IPT prophylaxis, and being in WHO stage III or IV at baseline were associated with an increased risk of tuberculosis [ 23 ]. Similarly, a study in Tanzania identified male sex, lower CD4 counts, and advanced WHO disease stages as predictors of TB incidence [ 24 ]. Conclusion This study developed and validated machine learning models to predict TB incidence in adult HIV/AIDS patients who are at ART centers. In our study, support vector machine, random forest, decision tree, logistic regression, gradient boosting, K-nearest neighbors, and XGBoost models have been tested. The XGBoost models have provided strong predictive performance. The current study demonstrated that machine learning-based prediction models are promising tools in accurately predicting tuberculosis incidence in adult HIV/AIDS patients. Further external validation and the inclusion of additional factors are needed to enhance the model’s robustness and generalizability. Limitations of the research A limitation of this research is that the data was sourced from secondary datasets originally collected for different purposes, which resulted in the absence of certain features, such as BMI. Additionally, the available records contained 1.8% missing values, which were imputed, and this could also have influenced the study’s outcomes. Data Availability Data is available from PI and will be avail based on request Ethical Approval Ethical approval was granted by the Institutional Research Ethics Review Committee of Debre Markos University. A letter of cooperation was received from the Amhara Public Health Institute, Debre Markos Branch. Additionally, permission was obtained from both Debre Markos Comprehensive Specialized Hospital and Debre Markos Health Center. CRediT authorship contribution description Desalegn Meseret Tadele: Writing original draft, Methodology, Investigation, Formal analysis, Conceptualization. Getaye Tizazu Biwota: Supervision, Methodology, review and editing, and manuscript preparation. Lijalem Megbaru Enyew : Supervision, Methodology, review and editing. Maru Meseret Tadele: Resources, Methodology and Formal analysis. Gizaw Hailiye Teferi: Supervision and formal analysis and validation . Funding This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors. Conflicts of interest No conflict interest claim among authors Acknowledgments None Abbreviations and acronyms AIDS Acquired Immune Deficiency syndrome ALWHIV Adults Living with HIV ART Antiretroviral Therapy AUC Area Under the Curve BMI Body Mass Index CD4 Cluster of differentiation 4 CDC Center for Disease Control CPT Cotrimoxazole Prophylactic Therapy DMCSH Debre Markos Comprehensive Specialized Hospital DSD Differentiated Service Delivery ENN Edited Nearest Neighbor HER Electronic Health Record HIV Human Immune Virus ICT Index Case Test IPT Isoniazid Preventive Therapy ML Machine Learning MoH Ministry of Health OIs Opportunistic Infections ROC Receiver Operating Characteristic SMOTE Synthetic Minority Over Sampling Technique TB Tuberculosis TPT Tuberculosis Preventive Therapy WHO World Health Organization References 1. ↵ Olivier , C. and L. Luies , WHO Goals and Beyond: Managing HIV/TB Co-infection in South Africa . SN Comprehensive Clinical Medicine , 2023 . 5 ( 1 ): p. 251 . OpenUrl 2. ↵ Azevedo-Pereira , J.M. , et al. HIV/Mtb Co-Infection: From the Amplification of Disease Pathogenesis to an “Emerging Syndemic” . Microorganisms , 2023 . 11 , DOI: 10.3390/microorganisms11040853 . OpenUrl CrossRef PubMed 3. ↵ Shen , Y . Mycobacterium tuberculosis and HIV Co-Infection: A Public Health Problem That Requires Ongoing Attention . Viruses , 2024 . 16 , DOI: 10.3390/v16091375 . OpenUrl CrossRef 4. ↵ Tornheim Jeffrey , A. and E. Dooley Kelly , Tuberculosis Associated with HIV Infection . Microbiology Spectrum , 2017 . 5 ( 1 ): p. doi: 10.1128/microbiolspec.tnmi7-0028-2016 . OpenUrl CrossRef 5. ↵ Bizuneh , F.K. , et al. , Tuberculosis-associated mortality and risk factors for HIV-infected population in Ethiopia: a systematic review and meta-analysis . Front Public Health , 2024 . 12 : p. 1386113 . OpenUrl PubMed 6. ↵ Rodrigues , M.M.S. , et al. , Machine learning algorithms using national registry data to predict loss to follow-up during tuberculosis treatment . BMC Public Health , 2024 . 24 ( 1 ): p. 1385 . OpenUrl PubMed 7. ↵ Chu , R. , et al. , Impact of tuberculosis on mortality among HIV-infected patients receiving antiretroviral therapy in Uganda: a prospective cohort analysis . AIDS Research and Therapy , 2013 . 10 ( 1 ): p. 19 . OpenUrl 8. ↵ Pintassilgo , I. , et al. , The Lisbon patient: exceptional longevity with HIV suggests healthy aging as an ultimate goal for HIV care . BMC Infectious Diseases , 2020 . 20 ( 1 ): p. 290 . OpenUrl PubMed 9. ↵ Siqueira Santos , L.F. , et al. , Tuberculosis/HIV co-infection in Northeastern Brazil: Prevalence trends, spatial distribution, and associated factors . J Infect Dev Ctries , 2022 . 16 ( 9 ): p. 1490 – 1499 . OpenUrl PubMed 10. ↵ Chen , J. , et al. , LSTM-Based Prediction Model for Tuberculosis Among HIV-Infected Patients Using Structured Electronic Medical Records: A Retrospective Machine Learning Study . J Multidiscip Healthc , 2024 . 17 : p. 3557 – 3573 . OpenUrl PubMed 11. ↵ Gisso , B.T. , M.W. Hordofa , and M.D. Ormago , Prevalence of pulmonary tuberculosis and associated factors among adults living with HIV/AIDS attending public hospitals in Shashamene Town, Oromia Region, South Ethiopia . SAGE Open Med , 2022 . 10 : p. 20503121221122437 . OpenUrl PubMed 12. ↵ Belsti T , G.G. , Nakachew MA , Anmut A. , Incidence of tuberculosis among HIV-positive adults on antiretroviral therapy at Debre Markos Referral Hospital, Northwest Ethiopia: A retrospective record review . 13. ↵ Liao , K.M. , et al. , Using an Artificial Intelligence Approach to Predict the Adverse Effects and Prognosis of Tuberculosis . Diagnostics (Basel) , 2023 . 13 ( 6 ). 14. ↵ Anito , A.A. , et al. , Magnitude of Viral Load Suppression and Associated Factors among Clients on Antiretroviral Therapy in Public Hospitals of Hawassa City Administration, Ethiopia . HIV AIDS (Auckl) , 2022 . 14 : p. 529 – 538 . OpenUrl PubMed 15. ↵ Narwane , S.V. and S.D. Sawarkar , Machine Learning and Class Imbalance: A Literature Survey . Industrial Engineering Journal , 2019 . 16. ↵ Mulugeta , G. , et al. , Classification of imbalanced data using machine learning algorithms to predict the risk of renal graft failures in Ethiopia . BMC Medical Informatics and Decision Making , 2023 . 23 ( 1 ): p. 98 . OpenUrl 17. ↵ Mamo , D.N. , et al. , Machine learning to predict virological failure among HIV patients on antiretroviral therapy in the University of Gondar Comprehensive and Specialized Hospital, in Amhara Region, Ethiopia, 2022 . BMC Medical Informatics and Decision Making , 2023 . 23 ( 1 ): p. 75 . OpenUrl 18. ↵ Swana , E.F. , W. Doorsamy , and P. Bokoro Tomek Link and SMOTE Approaches for Machine Fault Classification with an Imbalanced Dataset . Sensors , 2022 . 22 , DOI: 10.3390/s22093246 . OpenUrl CrossRef PubMed 19. ↵ Gichuhi , H.W. , et al. , A machine learning approach to explore individual risk factors for tuberculosis treatment non-adherence in Mukono district . PLOS Glob Public Health , 2023 . 3 ( 7 ): p. e0001466 . OpenUrl 20. ↵ Joseph , V.R. and A. and Vakayil , SPlit: An Optimal Method for Data Splitting . Technometrics , 2022 . 64 ( 2 ): p. 166 – 176 . OpenUrl 21. ↵ Peng , A.Z. , et al. , Explainable machine learning for early predicting treatment failure risk among patients with TB-diabetes comorbidity . Sci Rep , 2024 . 14 ( 1 ): p. 6814 . OpenUrl PubMed 22. ↵ Balogun , O.S. , et al. , Investigating Machine Learning Methods for Tuberculosis Risk Factors Prediction - A Comparative Analysis and Evaluation . Proceedings of the 37th International Business Information Management Association (IBIMA). 1056-1070 . . 23. ↵ Getu , A. , et al. , Incidence and predictors of Tuberculosis among patients enrolled in Anti-Retroviral Therapy after universal test and treat program, Addis Ababa, Ethiopia. A retrospective follow -up study . PLOS ONE , 2022 . 17 ( 8 ): p. e0272358 . OpenUrl PubMed 24. ↵ Liu , E. , et al. , Tuberculosis incidence rate and risk factors among HIV-infected adults with access to antiretroviral therapy . Aids , 2015 . 29 ( 11 ): p. 1391 – 9 . OpenUrl CrossRef PubMed View the discussion thread. Back to top Previous Next Posted August 28, 2025. Download PDF Data/Code Email Thank you for your interest in spreading the word about medRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Predicting Tuberculosis Incidence in Adult HIV Patients on ART in Debre Markos, Ethiopia: A Machine Learning Approach Message Subject (Your Name) has forwarded a page to you from medRxiv Message Body (Your Name) thought you would like to see this page from the medRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Predicting Tuberculosis Incidence in Adult HIV Patients on ART in Debre Markos, Ethiopia: A Machine Learning Approach Desalegn Meseret Tadele , Getaye Tizazu Biwota , Lijalem Megibaru Enyew , Maru Meseret Tadele , Gizaw Hailiye Teferi medRxiv 2025.08.24.25334330; doi: https://doi.org/10.1101/2025.08.24.25334330 Share This Article: Copy Citation Tools Predicting Tuberculosis Incidence in Adult HIV Patients on ART in Debre Markos, Ethiopia: A Machine Learning Approach Desalegn Meseret Tadele , Getaye Tizazu Biwota , Lijalem Megibaru Enyew , Maru Meseret Tadele , Gizaw Hailiye Teferi medRxiv 2025.08.24.25334330; doi: https://doi.org/10.1101/2025.08.24.25334330 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Health Informatics Subject Areas All Articles Addiction Medicine (568) Allergy and Immunology (863) Anesthesia (299) Cardiovascular Medicine (4425) Dentistry and Oral Medicine (443) Dermatology (382) Emergency Medicine (607) Endocrinology (including Diabetes Mellitus and Metabolic Disease) (1507) Epidemiology (15221) Forensic Medicine (30) Gastroenterology (1123) Genetic and Genomic Medicine (6588) Geriatric Medicine (667) Health Economics (997) Health Informatics (4524) Health Policy (1368) Health Systems and Quality Improvement (1612) Hematology (540) HIV/AIDS (1264) Infectious Diseases (except HIV/AIDS) (15910) Intensive Care and Critical Care Medicine (1103) Medical Education (623) Medical Ethics (145) Nephrology (667) Neurology (6588) Nursing (346) Nutrition (998) Obstetrics and Gynecology (1143) Occupational and Environmental Health (956) Oncology (3331) Ophthalmology (970) Orthopedics (369) Otolaryngology (420) Pain Medicine (435) Palliative Medicine (129) Pathology (663) Pediatrics (1690) Pharmacology and Therapeutics (691) Primary Care Research (710) Psychiatry and Clinical Psychology (5440) Public and Global Health (9220) Radiology and Imaging (2195) Rehabilitation Medicine and Physical Therapy (1369) Respiratory Medicine (1196) Rheumatology (593) Sexual and Reproductive Health (710) Sports Medicine (529) Surgery (710) Toxicology (99) Transplantation (289) Urology (265) (function(){function c(){var b=a.contentDocument||a.contentWindow.document;if(b){var d=b.createElement('script');d.innerHTML="window.__CF$cv$params={r:'9ffd1ff00d450700',t:'MTc3OTQ2NjU4MA=='};var a=document.createElement('script');a.src='/cdn-cgi/challenge-platform/scripts/jsd/main.js';document.getElementsByTagName('head')[0].appendChild(a);";b.getElementsByTagName('head')[0].appendChild(d)}}if(document.body){var a=document.createElement('iframe');a.height=1;a.width=1;a.style.position='absolute';a.style.top=0;a.style.left=0;a.style.border='none';a.style.visibility='hidden';document.body.appendChild(a);if('loading'!==document.readyState)c();else if(window.addEventListener)document.addEventListener('DOMContentLoaded',c);else{var e=document.onreadystatechange||function(){};document.onreadystatechange=function(b){e(b);'loading'!==document.readyState&&(document.onreadystatechange=e,c())}}}})();

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00