A Study on the Identification of Risk Factors for Latent Tuberculosis Infection in Xinjiang Using Machine Learning

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background Latent tuberculosis infection (LTBI) is a significant reservoir foractive tuberculosis (TB) development. Identifying key risk factors for LTBI is crucialfor effective prevention and control strategies. Machine learning (ML) techniques can uncover complex relationships between risk factors and disease outcomes. Methods Data were collected from the "Tuberculosis Management Information System" in China. LTBI was defined by positive tuberculin skin tests. Four ML models—random forest, XGBoost, support vector machine, and neural network—were used for feature importance analysis, alongside LASSO and logistic regression to identify key risk factors. A risk nomogram was constructed based on selected variables. Results Key risk factors identified included age, body mass index (BMI), smoking status, occupational dust exposure, diabetes, and family history of TB. Logistic regression also highlighted medical insurance type, immunosuppressant use, education level, silicosis, anemia, mental health status, TB contact history, and insomnia. The risk nomogram showed good discrimination (AUC = 0.839). Conclusion This study identified several key risk factors for LTBI in a Chinese population using ML techniques. The developed risk nomogram can aid in targeted LTBI screening and prevention, emphasizing interventions like smoking cessation and occupational dust control to reduce LTBI and active TB disease burden.
Full text 103,231 characters · extracted from preprint-html · click to expand
A Study on the Identification of Risk Factors for Latent Tuberculosis Infection in Xinjiang Using Machine Learning | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article A Study on the Identification of Risk Factors for Latent Tuberculosis Infection in Xinjiang Using Machine Learning YanJie Wang, Zhen Luo, Mairihaba kamili, LiTing Yuan, Yang Xiang, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5951117/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 7 You are reading this latest preprint version Abstract Background Latent tuberculosis infection (LTBI) is a significant reservoir foractive tuberculosis (TB) development. Identifying key risk factors for LTBI is crucialfor effective prevention and control strategies. Machine learning (ML) techniques can uncover complex relationships between risk factors and disease outcomes. Methods Data were collected from the "Tuberculosis Management Information System" in China. LTBI was defined by positive tuberculin skin tests. Four ML models—random forest, XGBoost, support vector machine, and neural network—were used for feature importance analysis, alongside LASSO and logistic regression to identify key risk factors. A risk nomogram was constructed based on selected variables. Results Key risk factors identified included age, body mass index (BMI), smoking status, occupational dust exposure, diabetes, and family history of TB. Logistic regression also highlighted medical insurance type, immunosuppressant use, education level, silicosis, anemia, mental health status, TB contact history, and insomnia. The risk nomogram showed good discrimination (AUC = 0.839). Conclusion This study identified several key risk factors for LTBI in a Chinese population using ML techniques. The developed risk nomogram can aid in targeted LTBI screening and prevention, emphasizing interventions like smoking cessation and occupational dust control to reduce LTBI and active TB disease burden. Health sciences/Risk factors Health sciences/Medical research/Epidemiology Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Figure 13 Figure 14 Figure 15 Introduction Tuberculosis (TB) continues to pose a significant global health challenge, with an estimated 10.8 million new cases and an incidence rate of 134 per 100,000 population in 2023. China ranks third in the world in terms of TB burden, accounting for 11% of global cases [1]. Despite implementing various control measures, the incidence of TB in China remains high, posing a considerable threat to public health [2].Latent tuberculosis infection (LTBI) is a state of persistent immune response to stimulation by Mycobacterium tuberculosis antigens without evidence of clinically manifested active TB [3]. It is estimated that approximately one-quarter of the world's population has LTBI, and 5-10% of infected individuals will develop active TB disease over their lifetime [3]. The identification and treatment of LTBI is a critical component of TB control strategies, as it can prevent the development of active disease and reduce transmission [4].Latent tuberculosis infection (LTBI) is characterized by a persistent immune response to stimulation by Mycobacterium tuberculosis antigens, in the absence of clinically manifested active TB [3]. It is estimated that approximately one-quarter of the global population is affected by LTBI, with 5-10% of infected individuals likely to develop active TB disease at some point in their lifetime [3]. Identifying and treating LTBI is a crucial aspect of tuberculosis control strategies, as it can prevent the progression to active disease and reduce transmission rates [4]. Machine learning (ML) has emerged as a powerful tool for predicting disease risk and identifying key risk factors [5]. ML algorithms excel at managing complex, high-dimensional data and capturing non-linear relationships among variables, which makes them particularly well-suited for modeling the multifactorial nature of tuberculosis (TB) [6]. Several studies have utilized ML techniques to predict TB risk and identify significant predictors, demonstrating their potential to enhance TB control efforts [7-9].However, the majority of existing studies have concentrated on active tuberculosis (TB), with limited research on the application of machine learning (ML) for predicting latent TB infection (LTBI) risk and its progression. Furthermore, there are few studies that have compared the performance of various ML algorithms and evaluated their clinical utility through decision curve analysis. Recent advances in ensemble deep learning frameworks have shown promising results in detecting various diseases, demonstrating how adaptive integration of multiple models can improve diagnostic accuracy [10]. Furthermore, research on optimized information fusion techniques has highlighted the importance of selective network combination for enhancing detection capabilities in respiratory conditions[11].Consequently, this study aims to develop and validate ML models for predicting LTBI risk and for identifying key risk factors within a Chinese population, as well as to assess the clinical usefulness of these models using a risk nomogram and decision curve analysis. Methods Study Population This study is a case-control design, with subjects into two distinct groups: active tuberculosis (ATB) patients and with latent tuberculosis infection (LTBI).The data for ATB patients were obtained from a public health surveillance cohort, specifically the Tuberculosis Management Information System (TBMIS) maintained by a disease prevention and control center in Xinjiang. These cases were diagnosed through laboratory tests from January to December 2022. In contrast, the data for LTBI individuals were derived from a community-based active screening cohort. Volunteers in this cohort underwent tuberculin skin testing (TST) as part of a regional tuberculosis prevention and control program, with positive results recorded from May to December 2022. The study design integrates both active and passive monitoring data to provide a comprehensive analysis. Justification for Data Integration and Machine Learning Approach While this study utilizes data from different sources (TBMIS surveillance data and community screening data), rigorous propensity score matching was implemented to minimize selection bias and ensure comparability between groups. This approach is increasingly recognized in epidemiological research as a valid method to leverage existing datasets when randomized controlled trials are not feasible or ethical. The complex, multifactorial nature of tuberculosis risk necessitates analytical approaches that can capture non-linear relationships and interactions among risk factors. Traditional statistical methods often assume linear relationships and independence among variables, which may not reflect the true complexity of TB risk factors. Machine learning algorithms, particularly ensemble methods like Random Forest and XGBoost, can model these complex relationships without requiring pre-specified assumptions about the functional form of relationships between predictors and outcomes. To ensure the validity of findings, multiple complementary analyses were implemented: traditional statistical methods (univariate analysis and logistic regression), variable selection techniques (LASSO), and interpretable machine learning approaches (SHAP analysis). Diagnostic Criteria ATB diagnostic criteria followed The diagnostic criteria for tuberculosis patients refer to the "Diagnostic Criteria for Pulmonary Tuberculosis (WS288-2017) / Health Industry Standard of the People's Republic of China." There are no clinical symptoms of active tuberculosis, such as persistent cough, sputum production, fever, night sweats, or weight loss. Imaging examinations show no abnormalities: chest X-rays or CT scans do not reveal any tuberculosis lesions.Microbiological tests are negative: Mycobacterium tuberculosis is not detected in samples such as sputum or bronchial secretions. The inclusion criteria are as follows: (1)Sputum smear and/or culture positive; (2) Chest X-ray: Presence of exudative lesions, caseous lesions, cavities, proliferative lesions, or hematogenous disseminated tuberculosis lesions in the lungs. The exclusion criteria include: (1)Tuberculosis of hilar lymph nodes; (2)Tuberculous pleurisy; (3)Extrapulmonary tuberculosis with intrapulmonary lesions; (4)HIV infection; (5)Hematologic malignancies LTBI diagnostic criteria followed The tuberculin skin test (TST) screening follows the "Diagnostic Criteria for Pulmonary Tuberculosis (WS288-2017)/Health Industry Standard of the People's Republic of China." The inclusion criteria are as follows: (1) TST positive (induration diameter ≥ 10 mm in BCG-vaccinated areas or ≥ 5 mm in non-BCG-vaccinated areas); (2) No clinical respiratory or systemic manifestations such as cough, sputum production, hemoptysis, or fever; (3) No history of mental illness; (4) Participation is voluntary. The exclusion criteria include: (1) Previously diagnosed with tuberculosis; (2) HIV infection; (3) Hematologic malignancies. Sample Size Calculation The sample size was calculated using the formula: In this study, N represents the sample size, Z denotes the statistic (1.96 for a 95% confidence level), p indicates the probability of individuals with latent tuberculosis infection (LTBI) developing active tuberculosis (0.03375 based on previous studies), and 1-p equals 0.96625. The margin of error, d, is set at 1.5%. The calculated sample size (N) was 557. Taking into account a 20% loss to follow-up rate, the final sample size for each group was adjusted to 669. The study received approval from the Ethics Committee of the First Affiliated Hospital of Xinjiang Medical University, and informed consent was obtained from all participants. Statistical Analysis Statistical analysis was performed using SPSS version 26.0. Continuous variables were compared using the rank-sum test, while categorical variables were analyzed with the chi-square test to determine if there were statistically significant differences between the groups. To minimize confounding bias and improve the reliability of causal inferences, a 1:1 nearest neighbor matching approach was employed for propensity score matching. The matching variables included potential confounders such as age, gender, and education level.Four machine learning models were developed, specifically random forest, XGBoost, support vector machine, and neural network. Grid search and 5-fold cross-validation were utilized to assess the performance of these models. Additionally, the Bootstrap method was applied to further evaluate the performance of the four machine learning models. Calibration curves were plotted for each model to assess the reliability of the predicted probabilities. To understand both the importance and direction of influence for each risk factor, the Shapley Additive Explanations (SHAP) methodology was applied to the machine learning models. SHAP values provide a unified measure of feature importance while also indicating the directional impact of each feature on model predictions. SHAP values for both XGBoost and Random Forest models were calculated to compare their feature importance rankings and interpretations.For the XGBoost model, the model was first trained using optimal hyperparameters identified through grid search and cross-validation. Then SHAP values for each instance in the dataset were calculated using the SHAPforxgboost package. This produced both global importance measures across the entire dataset and individual importance measures for each prediction. SHAP summary plots were created to visualize the overall feature importance and impact direction. Additionally, SHAP dependence plots were generated for the top features to examine how the relationship between feature values and their impact on predictions might be moderated by other important features. These dependence plots reveal non-linear patterns and interactions between features that traditional statistical methods might not capture.To compare the consistency between models, feature importance values from both Random Forest (using Mean Decrease Gini) and XGBoost (using SHAP gain values) were normalized and a direct comparison of their feature rankings and importance distributions was conducted. Feature importance analysis was performed using the XGBoost model to identify key risk factors. Variables that showed statistical significance in the univariate analysis were included in the LASSO regression model to select the optimal predictors. The selected variables were then incorporated into a binary logistic regression model to explore the relationship between each factor and the risk of tuberculosis. A tuberculosis risk nomogram was developed based on the variables identified from the logistic regression analysis. The predictive performance of the nomogram model was evaluated using receiver operating characteristic (ROC) curves and decision curve analysis. Ethics approval and consent to participate The study was approved by the Ethics Committee of Xinjiang Medical University(ID:XJYKDXR20211015010). All methods were performed in accordance with the relevant guidelines and regulations. Results Univariate analysis of tuberculosis incidence The univariate analysis identified 31 factors, including age, body mass index (BMI), average monthly income over the past two years, type of medical insurance, and education level, which demonstrated statistically significant differences between the tuberculosis and non-tuberculosis groups (P < 0.05) (Table 1 and Table 2). Propensity Score Matching To mitigate confounding bias and enhance the reliability of causal inferences, a 1:1 nearest neighbor matching approach was employed for propensity score matching. The matching variables included potential confounders such as age, gender, and education level. Following the matching process, the propensity score distributions of the two groups became more comparable, thereby reducing systematic differences (Figure 1). The Love plot further illustrated that the standardized differences of most variables were maintained within 0.1 after matching, indicating a well-balanced distribution of characteristics between the two groups (Figure 2). Four machine learning models were developed: random forest, XGBoost, support vector machine, and neural network. Grid search and 5-fold cross-validation were utilized to assess the performance of each model. All models demonstrated strong predictive capabilities, with XGBoost achieving the highest area under the curve (AUC) of 0.898, followed closely by random forest (AUC = 0.895), support vector machine (AUC = 0.877), and neural network (AUC = 0.797). Additionally, XGBoost exhibited the highest accuracy (85.7%), sensitivity (84.2%), and specificity (86.9%) (Figure 3). Bootstrap Method The Bootstrap method was employed to assess the performance of the four machine learning models. The results indicated that the XGBoost model demonstrated superior overall predictive performance compared to the other models ( Figure 4 and Table 3). Calibration Curve Evaluation Calibration curves were generated for each model to evaluate the reliability of the predicted probabilities. The XGBoost model exhibited superior calibration performance compared to the other models, with its calibration curve closely aligning with the diagonal line (Figure 5). Feature importance analysis was conducted using the XGBoost model to pinpoint key risk factors. Variables that exhibited statistical significance in the univariate analysis were incorporated into the model. The analysis indicated that age, BMI, smoking status, occupational dust exposure, diabetes, and a family history of tuberculosis significantly impact the risk of developing tuberculosis (Figure 6). SHAP Analysis Results The SHAP analysis revealed nuanced insights into how each factor contributes to tuberculosis risk prediction. Figure 7 shows the SHAP summary plot for the XGBoost model, with features ordered by their overall importance. Age group emerged as the most influential predictor (mean |SHAP value| = 0.818), followed by type of medical insurance (0.599), income group (0.523), and education level (0.439). The SHAP values revealed not only feature importance but also the direction of impact. Higher age values consistently pushed predictions toward higher tuberculosis risk (positive SHAP values), while higher income generally reduced predicted risk (negative SHAP values). Education level showed a non-linear relationship with TB risk, where both very low and very high education levels decreased risk prediction, while mid-range education levels increased risk, as shown in Figure 7 . When comparing Random Forest and XGBoost models (Figure 8), substantial agreement in the top features identified was found. Both models ranked age group, type of medical insurance, and income group among their top five features. However, the models differed in the importance assigned to some factors: Random Forest emphasized marital status more heavily, while XGBoost gave greater weight to education level. These differences can be attributed to Random Forest's tendency to detect more complex feature interactions through its ensemble of decision trees, while XGBoost's gradient boosting approach may better capture the main effects of individual features. Individual feature contribution plots for representative cases (Figures 9-11) demonstrate how the models arrive at predictions for specific patients. In case #1, type of medical insurance and age group were the dominant factors pushing the prediction toward tuberculosis risk, while in case #10, age group and education level were most influential, with income group providing a protective effect. Models LASSO Regression The 31 variables that showed statistically significant differences in the univariate analysis were included in the LASSO regression model. When lambda (λ) was set to 0.0032, the model error was minimized, corresponding to 57 variables. After incorporating the dummy variables for the categorical variables, 28 optimal variables were identified.(Figure 12). The selected 28 variables were included in the binary logistic regression model. The results indicated that age, BMI, household monthly income, smoking status, exercise frequency, personality type, occupational exposure to dust and chemical fumes, history of tuberculosis contact, presence of anemia, insomnia, silicosis, use of immunosuppressants, awareness of tuberculosis transmission routes and free policies, type of medical insurance, and education level were significantly associated with the risk of tuberculosis.(Table 4). Nomogram A tuberculosis risk nomogram was developed based on the 11 variables selected from the logistic regression analysis (Figure 13). Age was standardized and categorized into three groups using tertiles: the low age group (lowest tertile), the middle age group (middle tertile), and the high age group (highest tertile). This three-level stratification of age was incorporated into the nomogram to enhance risk assessment based on different age categories. The most influential factor was the type of medical insurance, followed by the use of immunosuppressants, age group, education level, presence of silicosis, smoking status, income group, presence of anemia, mental health status, history of tuberculosis contact, and presence of insomnia. Each variable was assigned a score based on a predefined scoring scale, and the total score was calculated by summing the scores of all variables. A higher total score indicated an increased risk of tuberculosis. The area under the ROC curve (AUC) for the nomogram was 0.839, indicating good discrimination (Figure 14). Decision curve analysis (DCA) showed that the model performed optimally at a 20% risk threshold, achieving a net benefit of 0.269. At this threshold, the prediction model identified 1,283 high-risk individuals (64.44%) among the 1,991 study subjects, of whom 688 (53.62%) were confirmed as true positive cases. The model demonstrated high sensitivity and effectively identified potential high-risk populations ( Figure 15). Discussion This study investigated the risk factors associated with tuberculosis incidence using a large sample size and various statistical methods. The findings provide valuable insights into the epidemiological characteristics of tuberculosis and inform prevention strategies within the study population. Univariate analysis identified 31 factors significantly associated with tuberculosis incidence, encompassing demographic characteristics, socioeconomic status, living environment, lifestyle behaviors, medical history, and awareness of tuberculosis. Statistical differences were observed between the tuberculosis and non-tuberculosis groups for age, BMI, average monthly income over the past two years, type of medical insurance, and education level (P < 0.05). These findings are consistent with previous studies that have reported similar risk factors for tuberculosis [12-14]. The large sample size and comprehensive assessment in this study further validate and refine the associations between these factors and tuberculosis risk. Comparative analysis of Random Forest and XGBoost models offers insights into the relative strengths of these approaches for tuberculosis risk prediction. While both models achieved similar overall performance metrics (AUC of 0.895 and 0.898 respectively), their different learning mechanisms led to varying emphasis on risk factors. XGBoost's sequential tree-building process, which focuses on minimizing bias, resulted in more weight being placed on demographic factors like age and income. In contrast, Random Forest's multiple independent trees, which aim to minimize variance, highlighted factors related to medical history and exposure, such as TB contact history and marital status. These complementary perspectives strengthen confidence in the identified risk factors that were consistently important across both models. The SHAP analysis revealed complex, non-linear relationships between risk factors and tuberculosis outcomes that would have been difficult to detect using traditional statistical methods alone. For instance, the U-shaped relationship between education level and TB risk suggests that both very low and very high education levels may be protective, while intermediate education shows increased risk. This could potentially be explained by differential occupational exposures or healthcare utilization patterns across education levels. Similarly, the interaction between fitness activity level and medical insurance type indicates that the protective effect of regular exercise may be moderated by healthcare access, highlighting the multidimensional nature of tuberculosis risk.The analysis of individual SHAP values for case examples provides clinically relevant insights for personalized risk assessment. By quantifying how each factor contributes to an individual's predicted risk, healthcare providers can identify the most significant modifiable risk factors for targeted intervention. For example, in cases where occupational dust exposure substantially increases risk, workplace interventions might be prioritized, while in cases where low income is a major contributor, social support may be more beneficial. Based on the univariate analysis, both LASSO regression and logistic regression models were employed to identify key risk factors for tuberculosis incidence. The LASSO regression included 31 variables with statistically significant differences and ultimately selected 28 optimal variables, which encompassed age, BMI, average monthly household income, smoking status, frequency of exercise, personality type, exposure to occupational dust and chemical fumes, history of tuberculosis contact, anemia, insomnia, silicosis, use of immunosuppressants, routes of tuberculosis transmission, awareness of free healthcare policies, type of medical insurance, and education level. LASSO regression is commonly used to manage high-dimensional data and helps prevent overfitting [15], resulting in a set of selected variables that are more stable and interpretable. The 28 variables selected by LASSO regression were incorporated into a binary logistic regression model, revealing that multiple factors were significantly associated with the risk of tuberculosis. Age emerged as a crucial risk factor, with the risk of tuberculosis increasing progressively with age, likely due to a decline in immune function and a rise in underlying health conditions among the elderly [16]. Additionally, low BMI, low household income, smoking, infrequent exercise, introverted personality traits, exposure to occupational dust and chemical fumes, history of tuberculosis contact, anemia, insomnia, silicosis, and use of immunosuppressants were all linked to an elevated risk of tuberculosis. These findings align with existing epidemiological studies and clinical observations [17-21], providing a foundation for tuberculosis risk assessment and the development of intervention strategies. The type of medical insurance and education level had a particularly significant impact on the risk of tuberculosis. Compared to commercial medical insurance, self-pay, urban employee and resident basic medical insurance, and new rural cooperative medical insurance were associated with a significantly reduced risk of tuberculosis. This suggests that an effective medical security system is crucial for the prevention and control of tuberculosis, as it facilitates timely medical treatment and standardized care for patients, thereby helping to reduce the spread of the disease [22]. Furthermore, education level was negatively correlated with tuberculosis risk; individuals with a junior high school education or lower exhibited a significantly higher risk of illness compared to those with a college education or higher. This disparity may be related to how education influences personal hygiene practices, nutritional status, and awareness of diseases [23]. Enhancing health education and implementing behavioral interventions for individuals with lower education levels could help mitigate their risk of illness. The relationship between awareness of tuberculosis knowledge and the risk of illness was also analyzed. The risk of illness was significantly higher among individuals who understood that tuberculosis is primarily transmitted through the respiratory tract. Conversely, the risk of illness was notably lower for those who were aware of the free examination and treatment policies for tuberculosis. These findings suggest that enhancing public awareness regarding tuberculosis prevention and control, along with increasing outreach efforts to inform more people about and facilitate access to free tuberculosis prevention and control services, is crucial for managing the tuberculosis epidemic [19]. Several specific populations were examined concerning their risk of tuberculosis. The risk of illness was significantly higher in smokers and individuals exposed to secondhand smoke compared to non-smokers, likely due to the damage to the respiratory mucosa and the suppression of immune function caused by tobacco use [20]. Thus, tobacco control and the avoidance of secondhand smoke exposure are essential for tuberculosis prevention. Additionally, exposure to occupational dust and chemical fumes was identified as a risk factor for tuberculosis, potentially due to chronic lung inflammation and fibrosis associated with these exposures [21]. Strengthening occupational health protections and improving working conditions can effectively reduce the risk of tuberculosis among affected populations. Factors such as anemia, insomnia, silicosis, and the use of immunosuppressants are associated with an increased risk of tuberculosis. Anemia can impair immune function, making individuals more susceptible to tuberculosis infection [24]. Insomnia negatively impacts neuroendocrine function and cellular immunity, which diminishes the body's overall resistance [25]. In patients with silicosis, the destruction of lung tissue and the resulting fibrosis create an ideal environment for the proliferation of tuberculosis bacilli [27]. The use of immunosuppressants directly compromises the immune system, further elevating the risk of tuberculosis [25]. Therefore, it is crucial to enhance tuberculosis screening and prevention efforts for these high-risk populations to enable early detection and treatment. Comparative analysis of Random Forest and XGBoost models offers insights into the relative strengths of these approaches for tuberculosis risk prediction. While both models achieved similar overall performance metrics (AUC of 0.895 and 0.898 respectively), their different learning mechanisms led to varying emphasis on risk factors. XGBoost's sequential tree-building process, which focuses on minimizing bias, resulted in more weight being placed on demographic factors like age and income. In contrast, Random Forest's multiple independent trees, which aim to minimize variance, highlighted factors related to medical history and exposure, such as TB contact history and marital status. These complementary perspectives strengthen confidence in the identified risk factors that were consistently important across both models. The SHAP analysis revealed complex, non-linear relationships between risk factors and tuberculosis outcomes that would have been difficult to detect using traditional statistical methods alone. For instance, the U-shaped relationship between education level and TB risk suggests that both very low and very high education levels may be protective, while intermediate education shows increased risk. This could potentially be explained by differential occupational exposures or healthcare utilization patterns across education levels. Similarly, the interaction between fitness activity level and medical insurance type indicates that the protective effect of regular exercise may be moderated by healthcare access, highlighting the multidimensional nature of tuberculosis risk. The analysis of individual SHAP values for case examples provides clinically relevant insights for personalized risk assessment. By quantifying how each factor contributes to an individual's predicted risk, healthcare providers can identify the most significant modifiable risk factors for targeted intervention. For example, in cases where occupational dust exposure substantially increases risk, workplace interventions might be prioritized, while in cases where low income is a major contributor, social support may be more beneficial. In summary, this study utilized multiple statistical methods to conduct a thorough analysis of the risk factors associated with tuberculosis incidence in a large sample population. It identified several key factors, including age, Body Mass Index (BMI), household income, smoking status, occupational exposure, history of tuberculosis contact, anemia, insomnia, silicosis, use of immunosuppressants, type of medical insurance, and education level. These findings provide valuable insights for developing effective tuberculosis prevention and control strategies, which may include enhancing health education, improving medical coverage, implementing tobacco control measures, strengthening occupational protections, and identifying high-risk populations. Moving forward, it will be crucial to validate the results of this study across diverse populations and evaluate the impact of interventions tailored to these risk factors, thereby generating robust evidence to support targeted prevention and control measures for tuberculosis. Declarations Data availability The data underlying this article cannot be shared publicly due to the privacy of individuals that participated in the study. The data will be shared on reasonable request to the corresponding author. Relevant R code is available upon request to the corresponding author. Authors’ contributions Y X and M k designed the search strategy and searched the literature. Y X and Y W selected the studies and made the quality assessment of studies included. Z L and Y W extracted data and analyzed data.Y W and Z L edited the manuscript. L Y provided coordinate the scene and distribute the questionnaire.Y X and J W provided resources used in drafting the discussion. All authors read and approved the final manuscript. Conflict of Interest The authors declare no conflict of interest. Funding This work was supported by the Xinjiang Uygur Autonomous Region Science Foundation(No.2022D01C203), and the Xinjiang Uygur Autonomous Region's 14th Five-Year Plan for Higher Education: Featured Discipline - Public Health and Preventive Medicine,and the Xinjiang Uygur Autonomous Region Association for Science and Technology(No.XHXM000490). Author details 1 Epidemiology and Statistics, School of Public Health, Xinjiang Medical University, Urumqi 830017, China 2 Department of Health Service Management, School of Public Health, Xinjiang Medical University,Urumqi 830017, China 3 Department of Tuberculosis Prevention, Disease Control and Prevention Center of Wushi County, Aksu City, Xinjiang,Akesu 843400, China 4 Medical Affairs Department of the People's Hospital of Xinjiang Uygur Autonomous Region,Urumqi 830001, China References World Health Organization. Global tuberculosis report 2024 (World Health Organization, 2024). Cui, X., Gao, L. & Cao, B. Management of latent tuberculosis infection in China: Exploring solutions suitable for high-burden countries. Int. J. Infect. Dis. 92S , S37–S40 (2020). Getahun, H., Matteelli, A., Chaisson, R. E. & Raviglione, M. Latent Mycobacterium tuberculosis infection. N Engl. J. Med. 372 (22), 2127–2135 (2015). World Health Organization. Guidelines on the management of latent tuberculosis infection (World Health Organization, 2015). Deo, R. C. Machine learning in medicine. Circulation 132 (20), 1920–1930 (2015). Gao, J., Jiang, Q., Zhou, B. & Chen, D. Convolutional neural networks for computer-aided detection or diagnosis in medical image analysis: An overview. Math. BiosciEng . 16 (6), 6536–6561 (2019). Seixas, J. M. et al. Artificial neural network models to support the diagnosis of pleural tuberculosis in adult patients. Int. J. Tuberc Lung Dis. 17 (5), 682–686 (2013). Lakhani, P. & Sundaram, B. Deep learning at chest radiography: automated classification of pulmonary tuberculosis by using convolutional neural networks. Radiology 284 (2), 574–582 (2017). Pasa, F., Golkov, V., Pfeiffer, F., Cremers, D. & Pfeiffer, D. Efficient deep network architectures for fast chest X-ray tuberculosis screening and visualization. Sci. Rep. 9 (1), 6268 (2019). Iqbal, M. S. et al. An adaptive ensemble deep learning framework for reliable detection of pandemic patients. Comput. Biol. Med. 168 , 107836 (2024). Hamza, A. et al. COVID-19 classification using chest X-ray images based on fusion-assisted deep Bayesian optimization and Grad-CAM visualization. Front. Public. Health . 10 , 1046296 (2022). Narasimhan, P., Wood, J., MacIntyre, C. R. & Mathai, D. Risk factors for tuberculosis. Pulm Med. 2013 , 828939 (2013). Lönnroth, K. et al. Tuberculosis control and elimination 2010-50: cure, care, and social development. Lancet 375 (9728), 1814–1829 (2010). Ai, J. W., Ruan, Q. L., Liu, Q. H. & Zhang, W. H. Updates on the risk factors for latent tuberculosis reactivation and their managements. Emerg. Microbes Infect. 5 (2), e10 (2016). Tibshirani, R. Regression shrinkage and selection via the lasso. J. R Stat. Soc. Ser. B Stat. Methodol. 58 (1), 267–288 (1996). Byng-Maddick, R. & Noursadeghi, M. Does tuberculosis threaten our ageing populations? BMC Infect. Dis. 16 , 119 (2016). Marais, B. J. et al. Tuberculosis comorbidity with communicable and non-communicable diseases: integrating health services and control efforts. Lancet Infect. Dis. 13 (5), 436–448 (2013). Lin, H. H., Ezzati, M. & Murray, M. Tobacco smoke, indoor air pollution and tuberculosis: a systematic review and meta-analysis. PLoS Med. 4 (1), e20 (2007). Kamineni, V. V., Turk, T., Wilson, N., Satyanarayana, S. & Chauhan, L. S. A rapid assessment and response approach to review and enhance advocacy, communication and social mobilisation for tuberculosis control in Odisha state, India. BMC Public. Health . 11 , 463 (2011). Bates, M. N. et al. Risk of tuberculosis from exposure to tobacco smoke: a systematic review and meta-analysis. Arch. Intern. Med. 167 (4), 335–342 (2007). Yarahmadi, A. et al. Correlation between silica exposure and risk of tuberculosis in Lorestan Province of Iran. Tanaffos 12 (2), 34–40 (2013). Boccia, D. et al. Cash transfer and microfinance interventions for tuberculosis control: review of the impact evidence and policy implications. Int. J. TubercLung Dis. 15 (Suppl 2), S37–S49 (2011). Suk, M. H. et al. Educational attainment and differences in the risk of active tuberculosis: a cohort study in South Korea. Int. J. Tuberc Lung Dis. 24 (4), 425–431 (2020). Barzegari, S., Afshari, M., Movahednia, M. & Moosazadeh, M. Prevalence of anemia among patients with tuberculosis: A systematic review and meta-analysis. Indian J. Tuberc . 66 (2), 299–307 (2019). Jee, S. H. et al. Smoking and risk of tuberculosis incidence, mortality, and recurrence in South Korean men and women. Am. J. Epidemiol. 170 (12), 1478–1485 (2009). Leung, C. C., Yu, I. T. & Chen, W. Silicosis Lancet ; 379 (9830):2008–2018. (2012). Dobler, C. C., Cheung, K., Nguyen, J. & Martin, A. Risk of tuberculosis inpatients with solid cancers and haematological malignancies: a systematic review and meta-analysis. Eur. Respir J. 50 (2), 1700157 (2017). Tables Tables 1 to 4 are available in the Supplementary Files section. Additional Declarations No competing interests reported. Supplementary Files Table.docx Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 03 Apr, 2025 Reviewers agreed at journal 03 Apr, 2025 Reviews received at journal 03 Apr, 2025 Reviewers agreed at journal 03 Apr, 2025 Reviewers invited by journal 03 Apr, 2025 Submission checks completed at journal 03 Apr, 2025 First submitted to journal 30 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5951117","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":438190119,"identity":"af5297fb-b565-43e4-b912-df0c33039341","order_by":0,"name":"YanJie Wang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+UlEQVRIiWNgGAWjYBADZjYGxsYHHyoYeEjRwtxsOOMMCVqAgL1NmLONCHUGx88efs1Tc4edj72xjZlxXp2MOfsBxg8fc/BoOZOXZs1z7BkzG8/BtseF2w7zWPYkMEvO3IZbi9mBHDNjHrbDzGwSie3GM7cd4DE4kMDGzItPy/k3QC3/gFrkH7ZJ886p4zE4/4CAlhs5xo9520C2MAK1NDDzGNwgYIv9jTdmjHP7gFp4EoGBfOwwUMvDZrx+kezPMf7w5tvhZPn24w8ffKipszc4n3zww0c8WoCATQoYfclIAowNeNUDAfPHHwwMdoRUjYJRMApGwQgGAN2AUmCKB3JxAAAAAElFTkSuQmCC","orcid":"","institution":"Xinjiang Medical University","correspondingAuthor":true,"prefix":"","firstName":"YanJie","middleName":"","lastName":"Wang","suffix":""},{"id":438190120,"identity":"eca4ad35-bec7-4fe6-8b87-694895348b03","order_by":1,"name":"Zhen Luo","email":"","orcid":"","institution":"Xinjiang Medical University","correspondingAuthor":false,"prefix":"","firstName":"Zhen","middleName":"","lastName":"Luo","suffix":""},{"id":438190121,"identity":"4b2818f8-7a09-4bd6-952d-c63589fa50d3","order_by":2,"name":"Mairihaba kamili","email":"","orcid":"","institution":"Xinjiang Medical University","correspondingAuthor":false,"prefix":"","firstName":"Mairihaba","middleName":"","lastName":"kamili","suffix":""},{"id":438190122,"identity":"1e93b29d-d60c-4566-bb29-9ad758ae8267","order_by":3,"name":"LiTing Yuan","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"LiTing","middleName":"","lastName":"Yuan","suffix":""},{"id":438190123,"identity":"dc40ad41-42dc-4153-9b12-0b01973bc2e4","order_by":4,"name":"Yang Xiang","email":"","orcid":"","institution":"Xinjiang Medical University","correspondingAuthor":false,"prefix":"","firstName":"Yang","middleName":"","lastName":"Xiang","suffix":""},{"id":438190124,"identity":"2cc2fda1-ebbb-42ed-8c7f-75559349adb7","order_by":5,"name":"Yu Wu","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Yu","middleName":"","lastName":"Wu","suffix":""},{"id":438190125,"identity":"df775cfe-9758-45bf-a492-e2025d067daf","order_by":6,"name":"JingJing Wei","email":"","orcid":"","institution":"Xinjiang Medical University","correspondingAuthor":false,"prefix":"","firstName":"JingJing","middleName":"","lastName":"Wei","suffix":""}],"badges":[],"createdAt":"2025-02-03 12:53:14","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5951117/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5951117/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":79888745,"identity":"126ec5a5-a6de-43fe-8ee4-63c64daad3c0","added_by":"auto","created_at":"2025-04-04 06:46:23","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":26060,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePropensity score distributions before and after matching\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/a1640fb37497cb1e1b89ce98.jpeg"},{"id":79887856,"identity":"eca38e63-80a1-4e8e-9614-203aba7c31de","added_by":"auto","created_at":"2025-04-04 06:30:24","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":14606,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLove plot of standardized differences before and after propensity score matching\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/f734cb0bc7b9cf8190b372f8.jpeg"},{"id":79888238,"identity":"ee0f7cb1-a6a4-4862-a2d1-b7243c454c27","added_by":"auto","created_at":"2025-04-04 06:38:23","extension":"jpeg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":37227,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eMachine learning model performance comparison\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/7abc7272b86167d4e26baa1f.jpeg"},{"id":79887837,"identity":"feae3900-6965-44f6-8d77-e7047c84b277","added_by":"auto","created_at":"2025-04-04 06:30:23","extension":"jpeg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":32347,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eBootstrap evaluation of machine learning model performance\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"4.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/36544cf19dcc07073b1f3d06.jpeg"},{"id":79887863,"identity":"688b7c5f-924c-48f9-964a-a86876a1b2ba","added_by":"auto","created_at":"2025-04-04 06:30:24","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":81526,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eMachine learning model calibration curve evaluation\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/db96069a8ef6610220648310.jpeg"},{"id":79887839,"identity":"c511ec82-bd83-4bad-a487-042b64da2c7a","added_by":"auto","created_at":"2025-04-04 06:30:23","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":50910,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eXGBoost model feature importance analysis\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/f85fbed729156f5b7a90bdea.png"},{"id":79887844,"identity":"f2ef2ca1-1ccb-4e6f-8ae4-c5e0053828e8","added_by":"auto","created_at":"2025-04-04 06:30:23","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":250249,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSHAP Value Distribution of Key Features for Tuberculosis Risk Prediction\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/a5dcf5b04c0ef9cfd424ccdb.png"},{"id":79887842,"identity":"607eacb1-4b72-4145-834a-1f35b52072fd","added_by":"auto","created_at":"2025-04-04 06:30:23","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":146969,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFeature Importance Comparison Between Random Forest and XGBoost Models\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"8.png","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/3a3c240d3a6ae4771919bcb6.png"},{"id":79887864,"identity":"91545b34-f698-4cd4-b6fe-600d927330ad","added_by":"auto","created_at":"2025-04-04 06:30:24","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":51380,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSHAP Feature Contributions for Case #1: Individual Risk Factor Analysis\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"9.png","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/900e47d29172e213c684c28a.png"},{"id":79888239,"identity":"5d218fcc-89fa-42ff-a5ba-65ce0c931560","added_by":"auto","created_at":"2025-04-04 06:38:23","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":50215,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSHAP Feature Contributions for Case #2: Individual Risk Factor Analysis\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"10.png","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/6634adc682f749b2f07bbb50.png"},{"id":79887849,"identity":"ffeebccf-14c7-4ce1-893e-51fad0b788a1","added_by":"auto","created_at":"2025-04-04 06:30:23","extension":"png","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":40983,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSHAP Feature Contributions for Case #3: Individual Risk Factor Analysis\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"11.png","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/b507561e84ce2d356ad30295.png"},{"id":79889171,"identity":"50c0f7e4-ab02-4fe2-8b2b-7994c92916db","added_by":"auto","created_at":"2025-04-04 06:54:28","extension":"png","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":386883,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLASSO regression variable\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"12.png","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/61956b1c9e35b9f236bab5f8.png"},{"id":79887881,"identity":"f86f1e86-84a5-417a-a083-c89d461ad980","added_by":"auto","created_at":"2025-04-04 06:30:25","extension":"jpeg","order_by":13,"title":"Figure 13","display":"","copyAsset":false,"role":"figure","size":47007,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTuberculosis risk\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"13.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/b8f5f0b92ccec3c62bd67b62.jpeg"},{"id":79887851,"identity":"b2fad4ea-ce0e-4f58-b7e0-86f127ba0212","added_by":"auto","created_at":"2025-04-04 06:30:23","extension":"jpeg","order_by":14,"title":"Figure 14","display":"","copyAsset":false,"role":"figure","size":25148,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eNomogram model ROC curve\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"14.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/bcfec26285044d41470f1e0b.jpeg"},{"id":79887879,"identity":"f7b5f019-ef9f-4312-8a50-a0951ca03921","added_by":"auto","created_at":"2025-04-04 06:30:25","extension":"png","order_by":15,"title":"Figure 15","display":"","copyAsset":false,"role":"figure","size":351756,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eNomogram model decision curve analysis\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"15.png","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/e6a912a5caf1f77f4d565ef9.png"},{"id":79889487,"identity":"2175cc07-e5c8-492b-b58f-0db8044b13b2","added_by":"auto","created_at":"2025-04-04 07:02:27","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2564816,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/1be1c821-d491-4654-a350-ea50adef370b.pdf"},{"id":79887832,"identity":"384be6b9-a4a5-48d7-b20f-3cccea74011c","added_by":"auto","created_at":"2025-04-04 06:30:23","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":57164,"visible":true,"origin":"","legend":"","description":"","filename":"Table.docx","url":"https://assets-eu.researchsquare.com/files/rs-5951117/v1/ca50f0646dfef91c7e97104c.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Study on the Identification of Risk Factors for Latent Tuberculosis Infection in Xinjiang Using Machine Learning","fulltext":[{"header":"Introduction","content":"\u003cp\u003eTuberculosis (TB) continues to pose a significant global health challenge, with an estimated 10.8 million new cases and an incidence rate of 134 per 100,000 population in 2023. China ranks third in the world in terms of TB burden, accounting for 11% of global cases [1]. Despite implementing various control measures, the incidence of TB in China remains high, posing a considerable threat to public health [2].Latent tuberculosis infection (LTBI) is a state of persistent immune response to stimulation by Mycobacterium tuberculosis antigens without evidence of clinically manifested active TB [3]. It is estimated that approximately one-quarter of the world\u0026apos;s population has LTBI, and 5-10% of infected individuals will develop active TB disease over their lifetime [3]. The identification and treatment of LTBI is a critical component of TB control strategies, as it can prevent the development of active disease and reduce transmission [4].Latent tuberculosis infection (LTBI) is characterized by a persistent immune response to stimulation by Mycobacterium tuberculosis antigens, in the absence of clinically manifested active TB [3]. It is estimated that approximately one-quarter of the global population is affected by LTBI, with 5-10% of infected individuals likely to develop active TB disease at some point in their lifetime [3]. Identifying and treating LTBI is a crucial aspect of tuberculosis control strategies, as it can prevent the progression to active disease and reduce transmission rates [4].\u003c/p\u003e\n\u003cp\u003eMachine learning (ML) has emerged as a powerful tool for predicting disease risk and identifying key risk factors [5]. ML algorithms excel at managing complex, high-dimensional data and capturing non-linear relationships among variables, which makes them particularly well-suited for modeling the multifactorial nature of tuberculosis (TB) [6]. Several studies have utilized ML techniques to predict TB risk and identify significant predictors, demonstrating their potential to enhance TB control efforts [7-9].However, the majority of existing studies have concentrated on active tuberculosis (TB), with limited research on the application of machine learning (ML) for predicting latent TB infection (LTBI) risk and its progression. Furthermore, there are few studies that have compared the performance of various ML algorithms and evaluated their clinical utility through decision curve analysis. Recent advances in ensemble deep learning frameworks have shown promising results in detecting various diseases, demonstrating how adaptive integration of multiple models can improve diagnostic accuracy [10]. Furthermore, research on optimized information fusion techniques has highlighted the importance of selective network combination for enhancing detection capabilities in respiratory conditions[11].Consequently, this study aims to develop and validate ML models for predicting LTBI risk and for identifying key risk factors within a Chinese population, as well as to assess the clinical usefulness of these models using a risk nomogram and decision curve analysis.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e\u003cstrong\u003eStudy Population\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study is a case-control design, with subjects into two distinct groups: active tuberculosis (ATB) patients and with latent tuberculosis infection (LTBI).The data for ATB patients were obtained from a public health surveillance cohort, specifically the Tuberculosis Management Information System (TBMIS) maintained by a disease prevention and control center in Xinjiang. These cases were diagnosed through laboratory tests from January to December 2022. In contrast, the data for LTBI individuals were derived from a community-based active screening cohort. Volunteers in this cohort underwent tuberculin skin testing (TST) as part of a regional tuberculosis prevention and control program, with positive results recorded from May to December 2022. The study design integrates both active and passive monitoring data to provide a comprehensive analysis.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eJustification for Data Integration and Machine Learning Approach\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWhile this study utilizes data from different sources (TBMIS surveillance data and community screening data), rigorous propensity score matching was implemented to minimize selection bias and ensure comparability between groups. This approach is increasingly recognized in epidemiological research as a valid method to leverage existing datasets when randomized controlled trials are not feasible or ethical.\u003c/p\u003e\n\u003cp\u003eThe complex, multifactorial nature of tuberculosis risk necessitates analytical approaches that can capture non-linear relationships and interactions among risk factors. Traditional statistical methods often assume linear relationships and independence among variables, which may not reflect the true complexity of TB risk factors. Machine learning algorithms, particularly ensemble methods like Random Forest and XGBoost, can model these complex relationships without requiring pre-specified assumptions about the functional form of relationships between predictors and outcomes.\u003c/p\u003e\n\u003cp\u003eTo ensure the validity of findings, multiple complementary analyses were implemented: traditional statistical methods (univariate analysis and logistic regression), variable selection techniques (LASSO), and interpretable machine learning approaches (SHAP analysis).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDiagnostic Criteria\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eATB diagnostic criteria followed\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe diagnostic criteria for tuberculosis patients refer to the \u0026quot;Diagnostic Criteria for Pulmonary Tuberculosis (WS288-2017) / Health Industry Standard of the People\u0026apos;s Republic of China.\u0026quot; There are no clinical symptoms of active tuberculosis, such as persistent cough, sputum production, fever, night sweats, or weight loss. Imaging examinations show no abnormalities: chest X-rays or CT scans do not reveal any tuberculosis lesions.Microbiological tests are negative: Mycobacterium tuberculosis is not detected in samples such as sputum or bronchial secretions.\u003c/p\u003e\n\u003cp\u003eThe inclusion criteria are as follows:\u003c/p\u003e\n\u003cp\u003e(1)Sputum smear and/or culture positive;\u003c/p\u003e\n\u003cp\u003e(2) Chest X-ray: Presence of exudative lesions, caseous lesions, cavities, proliferative lesions, or hematogenous disseminated tuberculosis lesions in the lungs.\u003c/p\u003e\n\u003cp\u003eThe exclusion criteria include:\u003c/p\u003e\n\u003cp\u003e(1)Tuberculosis of hilar lymph nodes;\u003c/p\u003e\n\u003cp\u003e(2)Tuberculous pleurisy;\u003c/p\u003e\n\u003cp\u003e(3)Extrapulmonary tuberculosis with intrapulmonary lesions;\u003c/p\u003e\n\u003cp\u003e(4)HIV infection;\u003c/p\u003e\n\u003cp\u003e(5)Hematologic malignancies\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLTBI diagnostic criteria followed\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe tuberculin skin test (TST) screening follows the \u0026quot;Diagnostic Criteria for Pulmonary Tuberculosis (WS288-2017)/Health Industry Standard of the People\u0026apos;s Republic of China.\u0026quot;\u003c/p\u003e\n\u003cp\u003eThe inclusion criteria are as follows:\u003c/p\u003e\n\u003cp\u003e(1) TST positive (induration diameter \u0026ge; 10 mm in BCG-vaccinated areas or \u0026ge; 5 mm in non-BCG-vaccinated areas);\u003c/p\u003e\n\u003cp\u003e(2) No clinical respiratory or systemic manifestations such as cough, sputum production, hemoptysis, or fever;\u003c/p\u003e\n\u003cp\u003e(3) No history of mental illness;\u003c/p\u003e\n\u003cp\u003e(4) Participation is voluntary. The exclusion criteria include:\u003c/p\u003e\n\u003cp\u003e(1) Previously diagnosed with tuberculosis;\u003c/p\u003e\n\u003cp\u003e(2) HIV infection;\u003c/p\u003e\n\u003cp\u003e(3) Hematologic malignancies.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSample Size Calculation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe sample size was calculated using the formula:\u003c/p\u003e\n\u003cp\u003e\u003cimg src=\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAOIAAABYCAYAAADoQSvEAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAAJcEhZcwAAFiUAABYlAUlSJPAAAAg6SURBVHhe7d1La9RcGAfw0/cLSNWVCxF140rwtvAGLnSkgrgQq+BOQepCUMFbdeVdREGw6sJ1FYW6UWtd1pVVaVcuqiKua9VPMG/+T8/TZtIkM8nMmZyZ/H9wOJkk1clknjnXJD3VgCGiQv1ncyIqEAORyAMMRCIPMBCJPMBAJPIAA5HIAwxEIg8wEIk8wEAk8gADkcgDDEQiDzAQiTzAQCTyAAORyAMMRCIPMBCJPMBAJPIAA5HIAwxEIg8wEB37/fu3uXPnjtm8ebPp6emRdPLkSVlPNA83jyJ3KpVKtbe3tzo+Pi6vh4eHcbOu6sDAgLwmApaIbXDhwgWzfft2WT58+LDZtGmT+fHjh7zuBHivKMXXrl1r13Q3nCOkDx8+2DXuMRAzQBVTq5f1khodHTXnzp2zr+YsW7bMLvnv2bNn8sMBY2NjksfBlxbVb3xGne7hw4dmw4YNZv/+/WZwcNCudcyWjNSA27dvS7US1c2ZmRm7dsGaNWtkO/ZL8v37d9nn9evXdo2/9HhRnU6C4+nv75f96h17p5mcnJRmBc63awzEDPSLiS9flG7DiYsLUhWULh3RPsQPRVpg4RgvXbokx4tARN5tgQho2+O4cKwusWqaEappq1evtq/moAf01q1bsnzz5s3EqifaWai+DQ0N2TX+OnXqlAmCyxw7dsyuqfXy5UvJp6enpfq6ZcsWed1t0LYPSkRz48YNt+16G5DUBPxa4qNE1TQJSsFO6SnVnt0spQCqb/ibbisRQWsHLs8fS8Qm4VcSv5Zw7do1yaOePHliJiYmakrCvXv32iU30OunHUfagYIcPZ9YhxzvK86rV68kP3TokOSdDMe8dOlSOWb9zN+8eTM/rott6JBJG9ft6+uT2gFKfmdsQFJO2lGR1KDXBn+43Yg2Jta5piU1cm2bosTS0gvp8ePHdu8Fui0Ln0tELdHwHvEZIMf7xLIeK85jGj0+HQ9uNQZiE7Qhn3aCwl/6aHINXzb9v6I9n7ot+oOgx4T3nYUep4+BGD5P0epleBt+NJPo5+Xq+Fg1bYJWRYNf0/kB+yiMIwafc2zKCtVgVKVQrcwyRS748klVNUw7Yf78+dPWgesiBW34RR1lOG86Tvru3TvJi8BAzAntBT1x2kZ0DT2VCJygamu+fv1q19a3atUqu7QAPbto9/gC7Tdt02ZJWdraSTODEKBFYyDmdPnyZcmD9tei4QxXDh48KMGDL866devs2vx8GnJIqzmkJfxdszCLpmgMxBzQ24hSCUFx5swZu9Y9BPzs7Kz59u1bS6bJ8QqQOf/+/bNLxWEgZoQv78WLF2UZk7mjAYF2HKpMndDu+vTpk+QrVqyQPKxMQfr582fJV65cKXkc18HKQMzo6dOn0k5D9TA6mRs+fvxol/zx4sULu7QAY2lQqVRqqtba6aRB2k3Qpo/+wOA11qN2s3v3brt2MQ3WrVu3St5yQT2bGoTxP3xkSHETobEds2uwPWk4o53CwxcYL9SxTLw3DFsgxXXZ69hoo8eg/x7+BsMYaXNtixAeosDwBc4TIMf4KtanTWwH7INjdIUlYgZ37961S8YcOXJkUQ8eSkm0HX2DUm9kZMQsX75c3icu78FwBkq99evX270WHDhwQPK3b99KnkRn6ezYsUNqCYDSRf8fzK31Cc4P2tgYrtDzhRQE6qLhnTCtPaTt0zQbkNSFmhmERsmOEsC30i0PLRGzTlJQOllBS1IXWCJSrAcPHkgpd+/ePbumnFAaopR3PUzFQKRYmOiMLx8mKzid7Oyxqakpc/ToUanKXr9+3a51owfFol12Dl37qJOrycnJ2DYKZkvgVyioFnXUbSWI8mpriahFe9D2kPz48eOSJ2EQUlkUUjXFJGQEI3rtkq6JIyqTQgJxyZIlcksJwCwVTrWisiuss+bEiRPSCEbP3JUrV+zaOe2aRE3ki0J7Te/fvy/5o0ePauZmxl220ygM1OZJ3XA/TupchQYi5jWivQinT5+WvFnoBM6T4uaNErVLoYEIV69eZccNlV7hgYghCnbcUNkVHoiQ1nHji7h2JVO5Uyt5EYgQ7rj5+fOnLOcR94E1kup11sS1K5nKnVrJm0AMd9wgGPOK+8AaSeysoSJ5E4igHTdEZdPWQNSHeCRVPcMdN0Rl0rZAxHVde/bskWVUPZOudtaOG6IyaetlUEQUz6s2IlFZMRCJPMBAJPIAA5HIAwxEIg8wEIk8wEAk8gADkVLhnqZ4SrFOjo/CbCncWl9vv499ebeD7BiIlEqfkZEE96mdmJgwY2NjMnke+58/f760NyXOi4FIddW7mRcuYdN9MHEffv36JTk1hoFITUEpqM9UBN4UOh8GIrUUqqS4lA3P+6fGMRCpZbTjZmhoiPemzYiBSDVwf1l0uGgvKXpBG3ngKG761d/fL0Ho9IGeXYqBSPNQrcTTf/HUYyS0/9Bjqhd0J0EQ4gleZ8+eZRDmxOsRSSDY9ILs6enpmk4XbNPH6cV9XRB82B5+hiACc3R01L6ielgiknj//r3czhJBFe35TGvv4abQf//+rQlCfeY8NY6BSGJkZETyLM8dQZUUN4XGQ2W1TYm0b98+uwc1ioFIuaHknJ2dlepqNLFamg0DkcgDDEQSGzdulPzLly+Sh/F5JO4xEEls27ZN8ufPn5upqSlZVr4+j6SbMBBJ9PX1zQ9f7Nq1ywwODsrlTLi8aefOnfPDF9EgpRYJGtZEYmZmpjowMFDt7e3FYGG1UqlUx8fHZRteI2Hb8PCwrKPW4YA+kQdYNSXyAAORyAMMRCIPMBCJPMBAJCqcMf8DEtx/rS17LGYAAAAASUVORK5CYII=\" style=\"width: 167px;\"\u003e\u003c/p\u003e\n\u003cp\u003eIn this study, N represents the sample size, Z denotes the statistic (1.96 for a 95% confidence level), p indicates the probability of individuals with latent tuberculosis infection (LTBI) developing active tuberculosis (0.03375 based on previous studies), and 1-p equals 0.96625. The margin of error, d, is set at 1.5%. The calculated sample size (N) was 557. Taking into account a 20% loss to follow-up rate, the final sample size for each group was adjusted to 669. The study received approval from the Ethics Committee of the First Affiliated Hospital of Xinjiang Medical University, and informed consent was obtained from all participants.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStatistical Analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eStatistical analysis was performed using SPSS version 26.0. Continuous variables were compared using the rank-sum test, while categorical variables were analyzed with the chi-square test to determine if there were statistically significant differences between the groups. To minimize confounding bias and improve the reliability of causal inferences, a 1:1 nearest neighbor matching approach was employed for propensity score matching. The matching variables included potential confounders such as age, gender, and education level.Four machine learning models were developed, specifically random forest, XGBoost, support vector machine, and neural network. Grid search and 5-fold cross-validation were utilized to assess the performance of these models. Additionally, the Bootstrap method was applied to further evaluate the performance of the four machine learning models. Calibration curves were plotted for each model to assess the reliability of the predicted probabilities.\u003c/p\u003e\n\u003cp\u003eTo understand both the importance and direction of influence for each risk factor, the Shapley Additive Explanations (SHAP) methodology was applied to the machine learning models. SHAP values provide a unified measure of feature importance while also indicating the directional impact of each feature on model predictions. SHAP values for both XGBoost and Random Forest models were calculated to compare their feature importance rankings and interpretations.For the XGBoost model, the model was first trained using optimal hyperparameters identified through grid search and cross-validation. Then SHAP values for each instance in the dataset were calculated using the SHAPforxgboost package. This produced both global importance measures across the entire dataset and individual importance measures for each prediction. SHAP summary plots were created to visualize the overall feature importance and impact direction. Additionally, SHAP dependence plots were generated for the top features to examine how the relationship between feature values and their impact on predictions might be moderated by other important features. These dependence plots reveal non-linear patterns and interactions between features that traditional statistical methods might not capture.To compare the consistency between models, feature importance values from both Random Forest (using Mean Decrease Gini) and XGBoost (using SHAP gain values) were normalized and a direct comparison of their feature rankings and importance distributions was conducted.\u003c/p\u003e\n\u003cp\u003eFeature importance analysis was performed using the XGBoost model to identify key risk factors. Variables that showed statistical significance in the univariate\u003c/p\u003e\n\u003cp\u003eanalysis were included in the LASSO regression model to select the optimal predictors. The selected variables were then incorporated into a binary logistic regression model to explore the relationship between each factor and the risk of tuberculosis. A tuberculosis risk nomogram was developed based on the variables identified from the logistic regression analysis. The predictive performance of the nomogram model was evaluated using receiver operating characteristic (ROC) curves and decision curve analysis.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe study was approved by the Ethics Committee of Xinjiang Medical University(ID:XJYKDXR20211015010). All methods were performed in accordance with the relevant guidelines and regulations.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eUnivariate analysis of tuberculosis incidence\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe univariate analysis identified 31 factors, including age, body mass index (BMI), average monthly income over the past two years, type of medical insurance, and education level, which demonstrated statistically significant differences between the tuberculosis and non-tuberculosis groups (P \u0026lt; 0.05) (Table 1 and Table 2).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ePropensity Score Matching\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo mitigate confounding bias and enhance the reliability of causal inferences, a 1:1 nearest neighbor matching approach was employed for propensity score matching. The matching variables included potential confounders such as age, gender, and education level. Following the matching process, the propensity score distributions of the two groups became more comparable, thereby reducing systematic differences (Figure 1). The Love plot further illustrated that the standardized differences of most variables were maintained within 0.1 after matching, indicating a well-balanced distribution of characteristics between the two groups (Figure 2).\u003c/p\u003e\n\u003cp\u003eFour machine learning models were developed: random forest, XGBoost, support vector machine, and neural network. Grid search and 5-fold cross-validation were utilized to assess the performance of each model. All models demonstrated strong predictive capabilities, with XGBoost achieving the highest area under the curve (AUC) of 0.898, followed closely by random forest (AUC = 0.895), support vector machine (AUC = 0.877), and neural network (AUC = 0.797). Additionally, XGBoost exhibited the highest accuracy (85.7%), sensitivity (84.2%), and specificity (86.9%) (Figure 3).\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eBootstrap Method\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Bootstrap method was employed to assess the performance of the four machine learning models. The results indicated that the XGBoost model demonstrated superior overall predictive performance compared to the other models ( Figure 4 and Table 3).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCalibration Curve Evaluation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCalibration curves were generated for each model to evaluate the reliability of the predicted probabilities. The XGBoost model exhibited superior calibration performance compared to the other models, with its calibration curve closely aligning with the diagonal line (Figure 5).\u003c/p\u003e\n\u003cp\u003eFeature importance analysis was conducted using the XGBoost model to pinpoint key risk factors. Variables that exhibited statistical significance in the univariate analysis were incorporated into the model. The analysis indicated that age, BMI, smoking status, occupational dust exposure, diabetes, and a family history of tuberculosis significantly impact the risk of developing tuberculosis (Figure 6).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSHAP Analysis Results\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe SHAP analysis revealed nuanced insights into how each factor contributes to tuberculosis risk prediction. Figure 7 shows the SHAP summary plot for the XGBoost model, with features ordered by their overall importance. Age group emerged as the most influential predictor (mean |SHAP value| = 0.818), followed by type of medical insurance (0.599), income group (0.523), and education level (0.439).\u003c/p\u003e\n\u003cp\u003eThe SHAP values revealed not only feature importance but also the direction of impact. Higher age values consistently pushed predictions toward higher tuberculosis risk (positive SHAP values), while higher income generally reduced predicted risk (negative SHAP values). Education level showed a non-linear relationship with TB risk, where both very low and very high education levels decreased risk prediction, while mid-range education levels increased risk, as shown in Figure 7 .\u003c/p\u003e\n\u003cp\u003eWhen comparing Random Forest and XGBoost models (Figure 8), substantial agreement in the top features identified was found. Both models ranked age group, type of medical insurance, and income group among their top five features. However, the models differed in the importance assigned to some factors: Random Forest emphasized marital status more heavily, while XGBoost gave greater weight to education level. These differences can be attributed to Random Forest's tendency to detect more complex feature interactions through its ensemble of decision trees, while XGBoost's gradient boosting approach may better capture the main effects of individual features.\u003c/p\u003e\n\u003cp\u003eIndividual feature contribution plots for representative cases (Figures 9-11) demonstrate how the models arrive at predictions for specific patients. In case #1, type of medical insurance and age group were the dominant factors pushing the prediction toward tuberculosis risk, while in case #10, age group and education level were most influential, with income group providing a protective effect.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModels\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLASSO Regression\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe 31 variables that showed statistically significant differences in the univariate analysis were included in the LASSO regression model. When lambda (λ) was set to 0.0032, the model error was minimized, corresponding to 57 variables. After incorporating the dummy variables for the categorical variables, 28 optimal variables were identified.(Figure 12).\u003c/p\u003e\n\u003cp\u003eThe selected 28 variables were included in the binary logistic regression model. The results indicated that age, BMI, household monthly income, smoking status, exercise frequency, personality type, occupational exposure to dust and chemical fumes, history of tuberculosis contact, presence of anemia, insomnia, silicosis, use of immunosuppressants, awareness of tuberculosis transmission routes and free policies, type of medical insurance, and education level were significantly associated with the risk of tuberculosis.(Table 4).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eNomogram\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA tuberculosis risk nomogram was developed based on the 11 variables selected from the logistic regression analysis (Figure 13). Age was standardized and categorized into three groups using tertiles: the low age group (lowest tertile), the middle age group (middle tertile), and the high age group (highest tertile). This three-level stratification of age was incorporated into the nomogram to enhance risk assessment based on different age categories.\u003c/p\u003e\n\u003cp\u003eThe most influential factor was the type of medical insurance, followed by the use of immunosuppressants, age group, education level, presence of silicosis, smoking status, income group, presence of anemia, mental health status, history of tuberculosis contact, and presence of insomnia. Each variable was assigned a score based on a predefined scoring scale, and the total score was calculated by summing the scores of all variables. A higher total score indicated an increased risk of tuberculosis.\u003c/p\u003e\n\u003cp\u003eThe area under the ROC curve (AUC) for the nomogram was 0.839, indicating good discrimination (Figure 14). Decision curve analysis (DCA) showed that the model performed optimally at a 20% risk threshold, achieving a net benefit of 0.269. At this threshold, the prediction model identified 1,283 high-risk individuals (64.44%) among the 1,991 study subjects, of whom 688 (53.62%) were confirmed as true positive cases. The model demonstrated high sensitivity and effectively identified potential high-risk populations ( Figure 15).\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study investigated the risk factors associated with tuberculosis incidence using a large sample size and various statistical methods. The findings provide valuable insights into the epidemiological characteristics of tuberculosis and inform prevention strategies within the study population.\u003c/p\u003e\n\u003cp\u003eUnivariate analysis identified 31 factors significantly associated with tuberculosis incidence, encompassing demographic characteristics, socioeconomic status, living environment, lifestyle behaviors, medical history, and awareness of tuberculosis. Statistical differences were observed between the tuberculosis and non-tuberculosis groups for age, BMI, average monthly income over the past two years, type of medical insurance, and education level (P \u0026lt; 0.05). These findings are consistent with previous studies that have reported similar risk factors for tuberculosis [12-14]. The large sample size and comprehensive assessment in this study further validate and refine the associations between these factors and tuberculosis risk.\u003c/p\u003e\n\u003cp\u003eComparative analysis of Random Forest and XGBoost models offers insights into the relative strengths of these approaches for tuberculosis risk prediction. While both models achieved similar overall performance metrics (AUC of 0.895 and 0.898 respectively), their different learning mechanisms led to varying emphasis on risk factors. XGBoost's sequential tree-building process, which focuses on minimizing bias, resulted in more weight being placed on demographic factors like age and income. In contrast, Random Forest's multiple independent trees, which aim to minimize variance, highlighted factors related to medical history and exposure, such as TB contact history and marital status. These complementary perspectives strengthen confidence in the identified risk factors that were consistently important across both models.\u003c/p\u003e\n\u003cp\u003eThe SHAP analysis revealed complex, non-linear relationships between risk factors and tuberculosis outcomes that would have been difficult to detect using traditional statistical methods alone. For instance, the U-shaped relationship between education level and TB risk suggests that both very low and very high education levels may be protective, while intermediate education shows increased risk. This could potentially be explained by differential occupational exposures or healthcare utilization patterns across education levels. Similarly, the interaction between fitness activity level and medical insurance type indicates that the protective effect of regular exercise may be moderated by healthcare access, highlighting the multidimensional nature of tuberculosis risk.The analysis of individual SHAP values for case examples provides clinically relevant insights for personalized risk assessment. By quantifying how each factor contributes to an individual's predicted risk, healthcare providers can identify the most significant modifiable risk factors for targeted intervention. For example, in cases where occupational dust exposure substantially increases risk, workplace interventions might be prioritized, while in cases where low income is a major contributor, social support may be more beneficial.\u003c/p\u003e\n\u003cp\u003eBased on the univariate analysis, both LASSO regression and logistic regression models were employed to identify key risk factors for tuberculosis incidence. The LASSO regression included 31 variables with statistically significant differences and ultimately selected 28 optimal variables, which encompassed age, BMI, average\u003c/p\u003e\n\u003cp\u003emonthly household income, smoking status, frequency of exercise, personality type, exposure to occupational dust and chemical fumes, history of tuberculosis contact, anemia, insomnia, silicosis, use of immunosuppressants, routes of tuberculosis transmission, awareness of free healthcare policies, type of medical insurance, and education level. LASSO regression is commonly used to manage high-dimensional data and helps prevent overfitting [15], resulting in a set of selected variables that are more stable and interpretable.\u003c/p\u003e\n\u003cp\u003eThe 28 variables selected by LASSO regression were incorporated into a binary logistic regression model, revealing that multiple factors were significantly associated with the risk of tuberculosis. Age emerged as a crucial risk factor, with the risk of tuberculosis increasing progressively with age, likely due to a decline in immune function and a rise in underlying health conditions among the elderly [16]. Additionally, low BMI, low household income, smoking, infrequent exercise, introverted personality traits, exposure to occupational dust and chemical fumes, history of tuberculosis contact, anemia, insomnia, silicosis, and use of immunosuppressants were all linked to an elevated risk of tuberculosis. These findings align with existing epidemiological studies and clinical observations [17-21], providing a foundation for tuberculosis risk assessment and the development of intervention strategies.\u003c/p\u003e\n\u003cp\u003eThe type of medical insurance and education level had a particularly significant impact on the risk of tuberculosis. Compared to commercial medical insurance, self-pay, urban employee and resident basic medical insurance, and new rural cooperative medical insurance were associated with a significantly reduced risk of tuberculosis. This suggests that an effective medical security system is crucial for the prevention and control of tuberculosis, as it facilitates timely medical treatment and standardized care for patients, thereby helping to reduce the spread of the disease [22]. Furthermore, education level was negatively correlated with tuberculosis risk; individuals with a junior high school education or lower exhibited a significantly higher risk of illness compared to those with a college education or higher. This disparity may be related to how education influences personal hygiene practices, nutritional status, and awareness of diseases [23]. Enhancing health education and implementing behavioral interventions for individuals with lower education levels could help mitigate their risk of illness.\u003c/p\u003e\n\u003cp\u003eThe relationship between awareness of tuberculosis knowledge and the risk of illness was also analyzed. The risk of illness was significantly higher among individuals who understood that tuberculosis is primarily transmitted through the respiratory tract. Conversely, the risk of illness was notably lower for those who were aware of the free examination and treatment policies for tuberculosis. These findings suggest that enhancing public awareness regarding tuberculosis prevention and control, along with increasing outreach efforts to inform more people about and facilitate access to free tuberculosis prevention and control services, is crucial for managing the tuberculosis epidemic [19].\u003c/p\u003e\n\u003cp\u003eSeveral specific populations were examined concerning their risk of tuberculosis. The risk of illness was significantly higher in smokers and individuals exposed to secondhand smoke compared to non-smokers, likely due to the damage to the respiratory mucosa and the suppression of immune function caused by tobacco use [20]. Thus, tobacco control and the avoidance of secondhand smoke exposure are essential for tuberculosis prevention. Additionally, exposure to occupational dust and chemical fumes was identified as a risk factor for tuberculosis, potentially due to chronic lung inflammation and fibrosis associated with these exposures [21]. Strengthening occupational health protections and improving working conditions can effectively reduce the risk of tuberculosis among affected populations.\u003c/p\u003e\n\u003cp\u003eFactors such as anemia, insomnia, silicosis, and the use of immunosuppressants are associated with an increased risk of tuberculosis. Anemia can impair immune function, making individuals more susceptible to tuberculosis infection [24]. Insomnia negatively impacts neuroendocrine function and cellular immunity, which diminishes the body's overall resistance [25]. In patients with silicosis, the destruction of lung tissue and the resulting fibrosis create an ideal environment for the proliferation of tuberculosis bacilli [27]. The use of immunosuppressants directly compromises the immune system, further elevating the risk of tuberculosis [25]. Therefore, it is crucial to enhance tuberculosis screening and prevention efforts for these high-risk populations to enable early detection and treatment.\u003c/p\u003e\n\u003cp\u003eComparative analysis of Random Forest and XGBoost models offers insights into the relative strengths of these approaches for tuberculosis risk prediction. While both models achieved similar overall performance metrics (AUC of 0.895 and 0.898 respectively), their different learning mechanisms led to varying emphasis on risk factors. XGBoost's sequential tree-building process, which focuses on minimizing bias, resulted in more weight being placed on demographic factors like age and income. In contrast, Random Forest's multiple independent trees, which aim to minimize variance, highlighted factors related to medical history and exposure, such as TB contact history and marital status. These complementary perspectives strengthen confidence in the identified risk factors that were consistently important across both models.\u003c/p\u003e\n\u003cp\u003eThe SHAP analysis revealed complex, non-linear relationships between risk factors and tuberculosis outcomes that would have been difficult to detect using traditional statistical methods alone. For instance, the U-shaped relationship between education level and TB risk suggests that both very low and very high education levels may be protective, while intermediate education shows increased risk. This could potentially be explained by differential occupational exposures or healthcare utilization patterns across education levels. Similarly, the interaction between fitness activity level and medical insurance type indicates that the protective effect of regular exercise may be moderated by healthcare access, highlighting the multidimensional nature of tuberculosis risk.\u003c/p\u003e\n\u003cp\u003eThe analysis of individual SHAP values for case examples provides clinically relevant insights for personalized risk assessment. By quantifying how each factor contributes to an individual's predicted risk, healthcare providers can identify the most significant modifiable risk factors for targeted intervention. For example, in cases where occupational dust exposure substantially increases risk, workplace interventions might be prioritized, while in cases where low income is a major contributor, social support may be more beneficial.\u003c/p\u003e\n\u003cp\u003eIn summary, this study utilized multiple statistical methods to conduct a thorough analysis of the risk factors associated with tuberculosis incidence in a large sample population. It identified several key factors, including age, Body Mass Index (BMI), household income, smoking status, occupational exposure, history of tuberculosis contact, anemia, insomnia, silicosis, use of immunosuppressants, type of medical insurance, and education level. These findings provide valuable insights for developing effective tuberculosis prevention and control strategies, which may include enhancing health education, improving medical coverage, implementing tobacco control measures, strengthening occupational protections, and identifying high-risk populations. Moving forward, it will be crucial to validate the results of this study across diverse populations and evaluate the impact of interventions tailored to these risk factors, thereby generating robust evidence to support targeted prevention and control measures for tuberculosis.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data underlying this article cannot be shared publicly due to the privacy of individuals that participated in the study. The data will be shared on reasonable request to the corresponding author. Relevant R code is available upon request to the corresponding author.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors’ contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eY X and M k designed the search strategy and searched the literature. Y X and Y W selected the studies and made the quality assessment of studies included. Z L and Y W extracted data and analyzed data.Y W and Z L edited the manuscript. L Y provided coordinate the scene and distribute the questionnaire.Y X and J W provided resources used in drafting the discussion. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of Interest \u003c/strong\u003eThe authors declare no conflict of interest. \u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by the Xinjiang Uygur Autonomous Region Science Foundation(No.2022D01C203), and the Xinjiang Uygur Autonomous Region's 14th Five-Year Plan for Higher Education: Featured Discipline - Public Health and Preventive Medicine,and the Xinjiang Uygur Autonomous Region Association for Science and Technology(No.XHXM000490).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor details\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e1 Epidemiology and Statistics, School of Public Health, Xinjiang Medical University, Urumqi 830017, China\u003c/p\u003e\n\u003cp\u003e2 Department of Health Service Management, School of Public Health, Xinjiang Medical University,Urumqi 830017, China\u003c/p\u003e\n\u003cp\u003e3 Department of Tuberculosis Prevention, Disease Control and Prevention Center of Wushi County, Aksu City, Xinjiang,Akesu 843400, China\u003c/p\u003e\n\u003cp\u003e4 Medical Affairs Department of the People's Hospital of Xinjiang Uygur Autonomous Region,Urumqi 830001, China\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eWorld Health Organization. \u003cem\u003eGlobal tuberculosis report 2024\u003c/em\u003e (World Health Organization, 2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCui, X., Gao, L. \u0026amp; Cao, B. Management of latent tuberculosis infection in China: Exploring solutions suitable for high-burden countries. \u003cem\u003eInt. J. Infect. Dis.\u003c/em\u003e \u003cb\u003e92S\u003c/b\u003e, S37\u0026ndash;S40 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGetahun, H., Matteelli, A., Chaisson, R. E. \u0026amp; Raviglione, M. Latent Mycobacterium tuberculosis infection. \u003cem\u003eN Engl. J. Med.\u003c/em\u003e \u003cb\u003e372\u003c/b\u003e (22), 2127\u0026ndash;2135 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWorld Health Organization. \u003cem\u003eGuidelines on the management of latent tuberculosis infection\u003c/em\u003e (World Health Organization, 2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeo, R. C. Machine learning in medicine. \u003cem\u003eCirculation\u003c/em\u003e \u003cb\u003e132\u003c/b\u003e (20), 1920\u0026ndash;1930 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao, J., Jiang, Q., Zhou, B. \u0026amp; Chen, D. Convolutional neural networks for computer-aided detection or diagnosis in medical image analysis: An overview. \u003cem\u003eMath. BiosciEng\u003c/em\u003e. \u003cb\u003e16\u003c/b\u003e (6), 6536\u0026ndash;6561 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSeixas, J. M. et al. Artificial neural network models to support the diagnosis of pleural tuberculosis in adult patients. \u003cem\u003eInt. J. Tuberc Lung Dis.\u003c/em\u003e \u003cb\u003e17\u003c/b\u003e (5), 682\u0026ndash;686 (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLakhani, P. \u0026amp; Sundaram, B. Deep learning at chest radiography: automated classification of pulmonary tuberculosis by using convolutional neural networks. \u003cem\u003eRadiology\u003c/em\u003e \u003cb\u003e284\u003c/b\u003e (2), 574\u0026ndash;582 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePasa, F., Golkov, V., Pfeiffer, F., Cremers, D. \u0026amp; Pfeiffer, D. Efficient deep network architectures for fast chest X-ray tuberculosis screening and visualization. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e9\u003c/b\u003e (1), 6268 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIqbal, M. S. et al. An adaptive ensemble deep learning framework for reliable detection of pandemic patients. \u003cem\u003eComput. Biol. Med.\u003c/em\u003e \u003cb\u003e168\u003c/b\u003e, 107836 (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHamza, A. et al. COVID-19 classification using chest X-ray images based on fusion-assisted deep Bayesian optimization and Grad-CAM visualization. \u003cem\u003eFront. Public. Health\u003c/em\u003e. \u003cb\u003e10\u003c/b\u003e, 1046296 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNarasimhan, P., Wood, J., MacIntyre, C. R. \u0026amp; Mathai, D. Risk factors for tuberculosis. \u003cem\u003ePulm Med.\u003c/em\u003e \u003cb\u003e2013\u003c/b\u003e, 828939 (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eL\u0026ouml;nnroth, K. et al. Tuberculosis control and elimination 2010-50: cure, care, and social development. \u003cem\u003eLancet\u003c/em\u003e \u003cb\u003e375\u003c/b\u003e (9728), 1814\u0026ndash;1829 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAi, J. W., Ruan, Q. L., Liu, Q. H. \u0026amp; Zhang, W. H. Updates on the risk factors for latent tuberculosis reactivation and their managements. \u003cem\u003eEmerg. Microbes Infect.\u003c/em\u003e \u003cb\u003e5\u003c/b\u003e (2), e10 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTibshirani, R. Regression shrinkage and selection via the lasso. \u003cem\u003eJ. R Stat. Soc. Ser. B Stat. Methodol.\u003c/em\u003e \u003cb\u003e58\u003c/b\u003e (1), 267\u0026ndash;288 (1996).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eByng-Maddick, R. \u0026amp; Noursadeghi, M. Does tuberculosis threaten our ageing populations? \u003cem\u003eBMC Infect. Dis.\u003c/em\u003e \u003cb\u003e16\u003c/b\u003e, 119 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMarais, B. J. et al. Tuberculosis comorbidity with communicable and non-communicable diseases: integrating health services and control efforts. \u003cem\u003eLancet Infect. Dis.\u003c/em\u003e \u003cb\u003e13\u003c/b\u003e (5), 436\u0026ndash;448 (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin, H. H., Ezzati, M. \u0026amp; Murray, M. Tobacco smoke, indoor air pollution and tuberculosis: a systematic review and meta-analysis. \u003cem\u003ePLoS Med.\u003c/em\u003e \u003cb\u003e4\u003c/b\u003e (1), e20 (2007).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKamineni, V. V., Turk, T., Wilson, N., Satyanarayana, S. \u0026amp; Chauhan, L. S. A rapid assessment and response approach to review and enhance advocacy, communication and social mobilisation for tuberculosis control in Odisha state, India. \u003cem\u003eBMC Public. Health\u003c/em\u003e. \u003cb\u003e11\u003c/b\u003e, 463 (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBates, M. N. et al. Risk of tuberculosis from exposure to tobacco smoke: a systematic review and meta-analysis. \u003cem\u003eArch. Intern. Med.\u003c/em\u003e \u003cb\u003e167\u003c/b\u003e (4), 335\u0026ndash;342 (2007).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYarahmadi, A. et al. Correlation between silica exposure and risk of tuberculosis in Lorestan Province of Iran. \u003cem\u003eTanaffos\u003c/em\u003e \u003cb\u003e12\u003c/b\u003e (2), 34\u0026ndash;40 (2013).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoccia, D. et al. Cash transfer and microfinance interventions for tuberculosis control: review of the impact evidence and policy implications. \u003cem\u003eInt. J. TubercLung Dis.\u003c/em\u003e \u003cb\u003e15\u003c/b\u003e (Suppl 2), S37\u0026ndash;S49 (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSuk, M. H. et al. Educational attainment and differences in the risk of active tuberculosis: a cohort study in South Korea. \u003cem\u003eInt. J. Tuberc Lung Dis.\u003c/em\u003e \u003cb\u003e24\u003c/b\u003e (4), 425\u0026ndash;431 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarzegari, S., Afshari, M., Movahednia, M. \u0026amp; Moosazadeh, M. Prevalence of anemia among patients with tuberculosis: A systematic review and meta-analysis. \u003cem\u003eIndian J. Tuberc\u003c/em\u003e. \u003cb\u003e66\u003c/b\u003e (2), 299\u0026ndash;307 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJee, S. H. et al. Smoking and risk of tuberculosis incidence, mortality, and recurrence in South Korean men and women. \u003cem\u003eAm. J. Epidemiol.\u003c/em\u003e \u003cb\u003e170\u003c/b\u003e (12), 1478\u0026ndash;1485 (2009).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLeung, C. C., Yu, I. T. \u0026amp; Chen, W. \u003cem\u003eSilicosis Lancet\u003c/em\u003e ;\u003cb\u003e379\u003c/b\u003e(9830):2008\u0026ndash;2018. (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDobler, C. C., Cheung, K., Nguyen, J. \u0026amp; Martin, A. Risk of tuberculosis inpatients with solid cancers and haematological malignancies: a systematic review and meta-analysis. \u003cem\u003eEur. Respir J.\u003c/em\u003e \u003cb\u003e50\u003c/b\u003e (2), 1700157 (2017).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTables 1 to 4 are available in the Supplementary Files section.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-5951117/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5951117/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground \u003c/strong\u003eLatent tuberculosis infection (LTBI) is a significant reservoir foractive tuberculosis (TB) development. Identifying key risk factors for LTBI is crucialfor effective prevention and control strategies. Machine learning (ML) techniques can uncover complex relationships between risk factors \u0026nbsp;and disease outcomes.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods \u003c/strong\u003eData were collected from the \u0026nbsp;\"Tuberculosis Management Information System\" in China. LTBI was defined by positive tuberculin skin tests. Four ML models—random forest, XGBoost, support vector machine, and neural network—were used for feature importance analysis, alongside LASSO and logistic regression to identify key risk factors. A risk nomogram was constructed based on selected variables.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults \u003c/strong\u003eKey risk factors identified included age, body mass index (BMI), smoking status, occupational dust exposure, diabetes, and family history of TB. Logistic regression also highlighted medical insurance type, immunosuppressant use, education level, silicosis, anemia, mental health status, TB contact history, and insomnia. The risk nomogram showed good discrimination (AUC = 0.839).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion \u003c/strong\u003eThis study identified several key risk factors for LTBI in a Chinese population using ML techniques. The developed risk nomogram can aid in targeted LTBI screening and prevention, emphasizing interventions like smoking cessation and occupational dust control to reduce LTBI and active TB disease burden.\u003c/p\u003e","manuscriptTitle":"A Study on the Identification of Risk Factors for Latent Tuberculosis Infection in Xinjiang Using Machine Learning","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-04 06:30:17","doi":"10.21203/rs.3.rs-5951117/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2025-04-04T01:23:23+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"38204286898723532351502966720515116435","date":"2025-04-04T01:13:44+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-03T07:39:13+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"161542416399476270703151546044720175169","date":"2025-04-03T07:37:01+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-04-03T07:17:17+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-04-03T05:50:20+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2025-03-30T21:16:57+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"cdc10a4c-94c1-48dd-bd7d-bad6663cba7a","owner":[],"postedDate":"April 4th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":46661241,"name":"Health sciences/Risk factors"},{"id":46661242,"name":"Health sciences/Medical research/Epidemiology"}],"tags":[],"updatedAt":"2025-04-08T06:08:21+00:00","versionOfRecord":[],"versionCreatedAt":"2025-04-04 06:30:17","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5951117","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5951117","identity":"rs-5951117","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00