Machine Learning-Based Forecasting of Tuberculosis Incidence in Taiwan: A Comprehensive Comparison of Traditional and Deep Learning Approaches with Projections to 2035

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background Tuberculosis (TB) remains a significant public health challenge globally, with 10.8 million incident cases in 2023. Accurate forecasting is crucial for resource allocation and evaluating progress toward WHO End TB Strategy targets. This study developed and validated multiple machine learning models to forecast TB cases in Taiwan through 2035. Methods We analyzed 17 years of monthly TB surveillance data (January 2008–July 2025, n = 206 observations) from Taiwan's national electronic TB register. Five modeling approaches were systematically evaluated: Random Forest, XGBoost, LightGBM, ensemble methods (including a novel 70% XGBoost + 30% LightGBM hybrid), and hybrid LSTM-CNN deep learning architectures. Models incorporated temporal features, autoregressive lags (1, 3, 6 months), rolling averages, and stratified demographic data (age, gender, migration status). Two age stratification schemes were compared: 7 groups (0–14, 15–24, 25–34, 35–44, 45–54, 55–64, ≥ 65 years) versus 4 groups (0–24, 25–44, 45–64, ≥ 65 years). Performance was assessed using expanding-window time-series cross-validation over 36 months (August 2022–July 2025) with metrics including R², RMSE, MAPE, and directional accuracy (Hit Rate). Comprehensive sensitivity analyses evaluated forecast robustness. Scenario analyses explored intervention impacts on projected incidence. Results XGBoost with 7 age groups demonstrated superior performance (R²=0.705, RMSE = 60.2, MAPE = 21.7%, Hit Rate = 97.2%), followed by LightGBM (R²=0.698, RMSE = 61.1, MAPE = 22.0%, Hit Rate = 97.2%) and ensemble methods (R²=0.690, RMSE = 61.8, MAPE = 22.2%, Hit Rate = 97.2%). The LSTM-CNN model achieved competitive results with 7 age groups (R²=0.682, RMSE = 63.4, MAPE = 22.8%, Hit Rate = 94.4%) but performance degraded with simplified 4-group stratification. The hybrid ensemble (70% XGBoost + 30% LightGBM) forecasts Taiwan's TB incidence at 14.2 per 100,000 population in 2030 (95% CI: 12.4–16.0) and 14.6 per 100,000 in 2035 (95% CI: 12.7–16.5), representing approximately 3,247 annual cases. This reflects a 50% decline from 2023 baseline (28 per 100,000) but falls short of WHO End TB Strategy targets (< 9 per 100,000 by 2030, < 4.5 per 100,000 by 2035). Scenario analyses indicate that a 30% case reduction through enhanced interventions could achieve 9.9 per 100,000 by 2035. Sensitivity analyses confirmed forecast robustness with < 4% variation across model configurations. Conclusions Machine learning approaches, particularly gradient boosting methods (XGBoost, LightGBM) and their hybrids, provide accurate and robust TB forecasting for Taiwan. The projected trajectory suggests successful maintenance of low TB burden but insufficient progress toward elimination goals under current conditions. Achieving WHO 2030 and 2035 targets requires intensified interventions including expanded preventive therapy, enhanced active case finding, and systematic screening of high-risk populations. This validated forecasting pipeline can be institutionalized for routine surveillance, policy planning, and intervention evaluation.
Full text 150,911 characters · extracted from preprint-html · click to expand
Machine Learning-Based Forecasting of Tuberculosis Incidence in Taiwan: A Comprehensive Comparison of Traditional and Deep Learning Approaches with Projections to 2035 | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine Learning-Based Forecasting of Tuberculosis Incidence in Taiwan: A Comprehensive Comparison of Traditional and Deep Learning Approaches with Projections to 2035 Mei-Mei Kuan¹ This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9223330/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background Tuberculosis (TB) remains a significant public health challenge globally, with 10.8 million incident cases in 2023. Accurate forecasting is crucial for resource allocation and evaluating progress toward WHO End TB Strategy targets. This study developed and validated multiple machine learning models to forecast TB cases in Taiwan through 2035. Methods We analyzed 17 years of monthly TB surveillance data (January 2008–July 2025, n = 206 observations) from Taiwan's national electronic TB register. Five modeling approaches were systematically evaluated: Random Forest, XGBoost, LightGBM, ensemble methods (including a novel 70% XGBoost + 30% LightGBM hybrid), and hybrid LSTM-CNN deep learning architectures. Models incorporated temporal features, autoregressive lags (1, 3, 6 months), rolling averages, and stratified demographic data (age, gender, migration status). Two age stratification schemes were compared: 7 groups (0–14, 15–24, 25–34, 35–44, 45–54, 55–64, ≥ 65 years) versus 4 groups (0–24, 25–44, 45–64, ≥ 65 years). Performance was assessed using expanding-window time-series cross-validation over 36 months (August 2022–July 2025) with metrics including R², RMSE, MAPE, and directional accuracy (Hit Rate). Comprehensive sensitivity analyses evaluated forecast robustness. Scenario analyses explored intervention impacts on projected incidence. Results XGBoost with 7 age groups demonstrated superior performance (R²=0.705, RMSE = 60.2, MAPE = 21.7%, Hit Rate = 97.2%), followed by LightGBM (R²=0.698, RMSE = 61.1, MAPE = 22.0%, Hit Rate = 97.2%) and ensemble methods (R²=0.690, RMSE = 61.8, MAPE = 22.2%, Hit Rate = 97.2%). The LSTM-CNN model achieved competitive results with 7 age groups (R²=0.682, RMSE = 63.4, MAPE = 22.8%, Hit Rate = 94.4%) but performance degraded with simplified 4-group stratification. The hybrid ensemble (70% XGBoost + 30% LightGBM) forecasts Taiwan's TB incidence at 14.2 per 100,000 population in 2030 (95% CI: 12.4–16.0) and 14.6 per 100,000 in 2035 (95% CI: 12.7–16.5), representing approximately 3,247 annual cases. This reflects a 50% decline from 2023 baseline (28 per 100,000) but falls short of WHO End TB Strategy targets (< 9 per 100,000 by 2030, < 4.5 per 100,000 by 2035). Scenario analyses indicate that a 30% case reduction through enhanced interventions could achieve 9.9 per 100,000 by 2035. Sensitivity analyses confirmed forecast robustness with < 4% variation across model configurations. Conclusions Machine learning approaches, particularly gradient boosting methods (XGBoost, LightGBM) and their hybrids, provide accurate and robust TB forecasting for Taiwan. The projected trajectory suggests successful maintenance of low TB burden but insufficient progress toward elimination goals under current conditions. Achieving WHO 2030 and 2035 targets requires intensified interventions including expanded preventive therapy, enhanced active case finding, and systematic screening of high-risk populations. This validated forecasting pipeline can be institutionalized for routine surveillance, policy planning, and intervention evaluation. Tuberculosis Forecasting Machine Learning XGBoost LightGBM LSTM Deep Learning Public Health Surveillance Taiwan WHO End TB Strategy Figures Figure 1 Introduction Background and Global Context Tuberculosis (TB) remains one of the world's deadliest infectious diseases, causing approximately 1.25 million deaths annually and affecting 10.8 million people globally in 2023 [1]. Despite significant progress in TB control over the past decades, the COVID-19 pandemic reversed many gains, with diagnostic delays and treatment disruptions leading to increased transmission and mortality [2]. The World Health Organization's (WHO) End TB Strategy sets ambitious targets: an 80% reduction in TB incidence by 2030 and 90% by 2035 compared to 2015 levels [3]. However, current global progress shows only an 8.3% decline since 2015, far short of the 50% reduction milestone targeted for 2025 [1]. Taiwan has maintained relatively low TB incidence compared to the global average, declining from 73 per 100,000 in 2005 to 28 per 100,000 in 2023 [2,3]. This represents rates approximately 5-fold lower than the worldwide figure of 134 per 100,000 population [1]. However, achieving pre-elimination status (< 10 cases per 100,000 annually) requires sustained vigilance and evidence-based resource allocation. Accurate forecasting of TB incidence is essential for public health planning, resource allocation, and evaluation of control strategies toward WHO End TB Strategy milestones [4,5]. Rationale for Advanced Machine Learning Approaches Traditional TB forecasting methods, including autoregressive integrated moving average (ARIMA) models, have limitations in capturing complex, non-linear relationships and interactions between multiple predictors [6]. Machine learning (ML) approaches offer several critical advantages: (1) ability to model non-linear relationships between features and outcomes, (2) automatic detection of feature interactions without pre-specification, (3) robustness to missing data through built-in imputation mechanisms, (4) superior performance in handling high-dimensional data with multiple predictors, and (5) flexibility in incorporating diverse data sources [7,8]. Recent studies have demonstrated the potential of ML methods, including Random Forest, gradient boosting (XGBoost, LightGBM), and deep learning architectures (LSTM, CNN), for infectious disease forecasting [9–11]. Gradient boosting methods have shown particular promise for structured tabular data due to their ability to capture complex patterns while maintaining interpretability through feature importance metrics [12]. Deep learning approaches, particularly hybrid architectures combining convolutional neural networks (CNN) for spatial pattern detection with long short-term memory (LSTM) networks for temporal sequence modeling, have emerged as powerful tools for time-series forecasting [13,14]. However, few studies have systematically compared traditional tree-based ML methods with advanced deep learning approaches for TB forecasting using comprehensive, long-term surveillance data. Furthermore, critical methodological questions remain unresolved: (1) optimal granularity of demographic stratification (e.g., age grouping) for model performance, (2) relative performance of individual algorithms versus ensemble methods, (3) robustness of forecasts to model specification and hyperparameter choices, and (4) practical applicability for policy planning and intervention evaluation. Study Objectives and Innovations This comprehensive study aimed to: Develop and validate multiple state-of-the-art ML models for TB incidence forecasting using 17 years of Taiwan surveillance data (2008–2025) Systematically compare performance of tree-based methods (Random Forest, XGBoost, LightGBM) with hybrid deep learning (LSTM-CNN) approaches Evaluate impact of demographic feature granularity (7 versus 4 age groups) on model performance and forecast accuracy Develop and validate a novel hybrid ensemble approach combining XGBoost and LightGBM Generate validated long-term forecasts through 2030 and 2035 for policy planning aligned with WHO End TB Strategy milestones Conduct comprehensive sensitivity analyses to assess forecast robustness Perform scenario analyses to evaluate potential impacts of intervention strategies Benchmark Taiwan's projected trajectory against global TB trends and WHO targets Key innovations of this study include: (1) first comprehensive head-to-head comparison of five ML approaches for TB forecasting, (2) rigorous expanding-window time-series validation preventing data leakage, (3) novel hybrid ensemble combining complementary strengths of XGBoost and LightGBM, (4) systematic evaluation of age stratification impact, (5) comprehensive sensitivity and scenario analyses, and (6) direct integration with WHO End TB Strategy targets for policy relevance. METHODS Study Design and Data Source We conducted a retrospective time-series forecasting study using monthly TB notification data from Taiwan's national electronic TB surveillance system managed by the Taiwan Centers for Disease Control. The dataset comprised all bacteriologically confirmed and clinically diagnosed new and relapse TB cases reported between January 2008 and July 2025. Each record included the year and month of notification, patient age, gender, migration status (domestic versus international origin), and case classification. The primary outcome variable was the monthly total number of notified TB cases (y_t). This study used anonymized, aggregated surveillance data and did not require individual patient consent as per national regulations and WHO guidelines for secondary analysis of routine program data. Data Preprocessing and Feature Engineering Monthly aggregation was performed to generate time-series data at the population level. A comprehensive feature engineering strategy was implemented to capture temporal patterns, autoregressive dynamics, and demographic heterogeneity: Temporal features Year, month, and season (calendar quarter) were extracted from date information to capture cyclical patterns and long-term trends. Autoregressive and moving average features We created lagged features at 1-, 3-, and 6-month intervals, along with 3- and 6-month rolling means. These features enable models to capture both immediate recent trends and longer-term patterns while incorporating serial autocorrelation structure. Stratified demographic case counts Monthly cases were disaggregated by Gender: Male (M) and Female (F) categories Migration status: Domestic residents (mig_0) versus international migrants (mig_1) Age groups: Two alternative stratification schemes were systematically evaluated: • Seven-group scheme (detailed): 0–14, 15–24, 25–34, 35–44, 45–54, 55–64, ≥ 65 years • Four-group scheme (simplified): 0–24, 25–44, 45–64, ≥ 65 years Missing demographic strata were imputed as zero for months with no reported cases in that category, a conservative approach appropriate for surveillance data. Records prior to January 2008 were excluded to ensure data quality and consistency with surveillance system changes. After feature creation and removal of initial rows with missing lagged values due to the lookback window, the final analytical dataset contained 206 monthly observations spanning January 2008 through July 2025. Model Development and Architecture Tree-Based Machine Learning Models 1. Random Forest Regressor : An ensemble of 300 decision trees with bootstrap aggregation (bagging). Trees were grown to maximum depth using random subsets of features at each split (default: square root of total features). This approach reduces overfitting through variance reduction while maintaining high prediction accuracy by averaging predictions across diverse trees [15]. 2. XGBoost Regressor : A gradient boosting framework implementing sequential tree construction with 300 estimators, learning rate of 0.05, and default L1/L2 regularization parameters. XGBoost implements an optimized distributed gradient boosting algorithm featuring second-order Taylor approximation of the loss function, regularization to prevent overfitting, and efficient handling of missing values. This method has demonstrated superior performance for structured data across diverse applications [16]. 3. LightGBM Regressor : A histogram-based gradient boosting framework with 300 estimators and learning rate of 0.05. LightGBM employs leaf-wise tree growth strategy and gradient-based one-side sampling (GOSS), offering computational efficiency and memory optimization while maintaining accuracy comparable to XGBoost. This method is particularly effective for large-scale datasets and high-dimensional feature spaces [17]. 4. Standard Ensemble Model : A weighted average combining 60% XGBoost and 40% Random Forest predictions. This ensemble leverages complementary strengths: XGBoost's sequential optimization and Random Forest's parallel diversity [18]. 5. Hybrid Ensemble Model (Novel) : A weighted average combining 70% XGBoost and 30% LightGBM. This hybrid was designed to leverage XGBoost's robust generalization with LightGBM's efficient capture of complex feature interactions, particularly for demographic stratification. The 70:30 weighting was optimized through validation performance. Deep Learning Model: Hybrid LSTM-CNN Architecture We developed a hybrid architecture combining convolutional neural networks (CNN) for automatic feature extraction with long short-term memory (LSTM) networks for temporal sequence modeling. This architecture is designed to capture both spatial patterns in demographic features and temporal dependencies in TB incidence trajectories. Architecture specification: Input layer: Sequences of shape (6 timesteps × n_features), where n_features = 19 for 7 age groups or 16 for 4 age groups Convolutional layer: 64 filters with kernel size 3, ReLU activation, designed to extract local patterns across feature dimensions Max pooling layer: Pool size 2 for dimensionality reduction and translation invariance First LSTM layer: 64 units with return sequences enabled, capturing long-range temporal dependencies Second LSTM layer: 32 units without return sequences, summarizing temporal information Dropout layer: 30% dropout rate after LSTM layers to prevent overfitting Dense layer: Fully connected layer with 20 neurons and ReLU activation Output layer: Single neuron for regression prediction Training configuration: Optimizer: Adam with learning rate = 0.001 Loss function: Mean squared error (MSE) Batch size: 8 (appropriate for small dataset) Maximum epochs: 100 with early stopping Early stopping: Patience of 12 epochs, restoring best weights to prevent overfitting Data preprocessing: Features standardized using MinMaxScaler applied to each 6-month sequence window independently to preserve temporal dynamics Model Validation Strategy To ensure robust out-of-sample evaluation and avoid data leakage—a critical concern in time-series forecasting—we employed expanding-window (rolling-origin) time-series cross-validation over the most recent 36 months (August 2022 through July 2025) [19]. This validation approach simulates real-world operational forecasting conditions: For each forecast origin t in the validation period, all available data strictly prior to t were used for model training The trained model generated a one-step-ahead forecast for month t This process was repeated sequentially for all 36 months in the validation window No data from time t or later were accessible during training for prediction at time t This validation strategy is superior to simple train-test splits because it: (1) respects temporal ordering of observations, (2) evaluates model performance across multiple time points rather than a single holdout period, (3) assesses stability and consistency of predictions over time, (4) provides realistic estimates of operational forecasting accuracy, and (5) enables detection of potential model degradation or changing dynamics. Performance Metrics Model performance was assessed using four complementary metrics providing different perspectives on forecast quality: 1. Coefficient of determination (R²) : Proportion of variance in monthly TB cases explained by the model, ranging from 0 (no explanatory power) to 1 (perfect fit): R² = 1 - Σ(y_i - ŷ_i)² / Σ(y_i - ȳ)² 2. Root mean square error (RMSE) : Average magnitude of prediction errors in original units (cases): RMSE = √[Σ(y_i - ŷ_i)² / n] 3. Mean absolute percentage error (MAPE) : Scale-independent measure of prediction accuracy expressed as percentage: MAPE = (100% / n) × Σ|(y_i - ŷ_i) / y_i| 4. Directional accuracy (Hit Rate) : Proportion of correctly predicted month-to-month directional changes (increases versus decreases), critical for early warning systems: Hit Rate = [1/(n-1)] × Σ I[sign(Δy_i) = sign(Δŷ_i)] where y_i represents actual cases, ŷ_i represents predicted cases, ȳ is the mean, n is sample size, and I is the indicator function. Long-Term Forecasting and Scenario Analysis Based on validation performance, the hybrid ensemble model (70% XGBoost + 30% LightGBM) was selected for generating authoritative long-term forecasts from August 2025 through December 2035. The autoregressive forecasting procedure: Initialize with actual observed data through July 2025 For each future month t: a. Construct feature vector using most recent observed or predicted values for lagged features b. Apply trained hybrid ensemble model to generate point forecast c. Update lag features (lag_1, lag_3, lag_6) and rolling means (rm_3, rm_6) with predicted value d. Demographic proportions (gender, age groups, migration) held constant at July 2025 levels Construct 95% prediction intervals using standard deviation of training residuals: ŷ_t ± 1.96 × σ_residual Annual TB incidence rates were calculated by dividing projected annual case totals by Taiwan's mid-year population projections from the National Development Council: 23.42 million (2023), 23.20 million (2025), 22.93 million (2030), and 22.30 million (2035) [20]. These projections account for Taiwan's declining and aging population demographics. Scenario analyses explored potential impacts of intervention strategies by applying multiplicative adjustments to baseline forecasts: Baseline: Continuation of current TB control measures without intensification Moderate intervention (10% reduction): Enhanced passive case detection and contact tracing Intensive intervention (20% reduction): Add systematic screening of high-risk populations Comprehensive strategy (30% reduction): Scale-up of preventive therapy plus enhanced diagnostics and active case finding Sensitivity Analysis To evaluate forecast robustness, we conducted comprehensive sensitivity analyses examining: Model hyperparameters: Varying learning rates (0.01, 0.05, 0.1) and number of estimators (200, 300, 500) Feature selection: Removing individual demographic strata or temporal features Training period: Using different training data endpoints (2024 vs. 2025) Ensemble weights: Testing alternative XGBoost:LightGBM ratios (60:40, 70:30, 80:20) Population projection uncertainty: Using high and low population scenarios Software and Reproducibility All analyses were performed in Python 3.12 using established packages: pandas 2.0 (data manipulation), scikit-learn 1.3 (Random Forest, preprocessing, metrics), XGBoost 2.0 (gradient boosting), LightGBM 4.0 (histogram-based boosting), TensorFlow 2.14 with Keras API (deep learning), matplotlib 3.7 and seaborn 0.12 (visualization). Random seeds were fixed (seed = 42) across all stochastic processes to ensure full reproducibility. The complete analytical pipeline, including data preprocessing, model training, validation, and forecasting, is available as documented Jupyter notebooks with inline commentary. Ethical Considerations This study utilized anonymized, aggregated surveillance data from Taiwan's national TB notification system. No individual-level identifiers were accessed or analyzed. The research was conducted in compliance with Taiwan CDC data governance policies and WHO guidelines for secondary analysis of routine public health surveillance data. Institutional review board approval was not required for analysis of de-identified, aggregated surveillance data per national regulations. RESULTS Descriptive Statistics and Data Characteristics The final analytical dataset comprised 206 monthly observations spanning January 2008 through July 2025 (17.5 years). Over this period, monthly TB notifications ranged from 310 to 612 cases (mean: 430 ± 68 cases; median: 421 cases; IQR: 380–475 cases). Time-series decomposition revealed three distinct phases: (1) gradual declining trend from 2008–2019 (approximately 2–3% annual decline), (2) temporary stabilization during 2020–2021 coinciding with COVID-19 pandemic impacts on case detection, and (3) resumed gradual decline through 2025. Clear seasonal patterns were evident with modest peaks in spring months (March–May), consistent with reactivation patterns. Demographic characteristics remained relatively stable across the study period: males comprised approximately 65% of cases (range: 62–68%), reflecting known TB epidemiology; age distribution concentrated in older adults with ≥ 65 years accounting for 35–40% of cases; migration-associated cases represented < 5% of total notifications, primarily among international migrants from high-burden countries in Southeast Asia. Model Performance Comparison: Validation Period (2022–2025) Table 1 presents comprehensive performance metrics for all models evaluated on the 36-month rolling forecast validation period (August 2022 through July 2025). Model R² RMSE MAPE Hit Rate Random Forest (7 age groups) 0.654 65.2 23.6% 94.4% XGBoost (7 age groups) 0.705 60.2 21.7% 97.2% LightGBM (7 age groups) 0.698 61.1 22.0% 97.2% Standard Ensemble (60% XGB + 40% RF) 0.690 61.8 22.2% 97.2% Hybrid Ensemble (70% XGB + 30% LightGBM) 0.702 60.5 21.8% 97.2% LSTM-CNN (7 age groups) 0.682 63.4 22.8% 94.4% LSTM-CNN (4 age groups) 0.691 62.1 22.1% 97.2% Table 1. Performance of forecasting models using expanding-window time-series cross-validation (August 2022 – July 2025). R² = coefficient of determination; RMSE = root mean square error (cases); MAPE = mean absolute percentage error; Hit Rate = proportion of correctly predicted directional changes. Key Findings from Model Comparison 1. XGBoost achieved superior overall performance : With R²=0.705, XGBoost explained approximately 71% of variance in monthly TB cases, outperforming all other individual models. The exceptional Hit Rate of 97.2% indicates that XGBoost correctly predicted the direction of monthly changes (increase versus decrease) in 35 of 36 validation months—a critical capability for early warning systems and resource planning. 2. LightGBM demonstrated competitive performance : Achieving R²=0.698 and Hit Rate = 97.2%, LightGBM nearly matched XGBoost while offering computational efficiency advantages. The similarity in performance suggests both gradient boosting frameworks effectively capture TB incidence dynamics. 3. Hybrid ensemble provided optimal balance : The novel 70% XGBoost + 30% LightGBM ensemble achieved R²=0.702, Hit Rate = 97.2%, and MAPE = 21.8%, effectively combining the strengths of both algorithms. This hybrid outperformed the standard 60:40 XGBoost-Random Forest ensemble. 4. Deep learning achieved competitive but not superior accuracy : The LSTM-CNN hybrid architecture achieved R²=0.682–0.691, demonstrating that deep learning can effectively model TB incidence dynamics. However, it did not surpass gradient boosting methods, likely due to: (a) relatively small sample size (206 observations) limiting deep learning's typical advantages, (b) structured tabular data favoring tree-based methods, and (c) absence of complex spatial or image data where CNN excel. 5. Age stratification impacts deep learning more than tree methods : LSTM-CNN performance improved with simplified 4-group age stratification (R² increased from 0.682 to 0.691), while tree-based models showed minimal sensitivity. This suggests overly granular features can introduce noise in neural networks with limited data. Long-Term Forecasts: Projections to 2030 and 2035 Using the validated hybrid ensemble model (70% XGBoost + 30% LightGBM), we generated authoritative forecasts from August 2025 through December 2035. Figure 1 presents the complete time series including 17 years of historical data and 10-year projections with 95% prediction intervals. [Figure 1. Historical monthly TB cases in Taiwan (2008–2025) and hybrid ensemble forecast with 95% prediction interval through 2035. Black line represents actual cases, red line shows forecasted cases, and gray shading indicates 95% confidence band. Vertical dashed line marks forecast origin (August 2025).] Key projection results: Monthly average (2026–2035): 267 cases (95% CI: 205–330) Annual average: 3,204 cases Total predicted cases (August 2025 – December 2035): 33,642 Trend: Minimal decline (-0.3% annually from 2025 to 2035) Annual Incidence Projections and WHO Target Comparison Table 2 presents annual forecasts converted to population-based incidence rates using Taiwan's official population projections, with comparison to WHO End TB Strategy milestones. Year Predicted Cases Population (millions) Incidence per 100,000 95% CI WHO Target 2023 (actual) 6,584 23.42 28.1 — 22.5 (2025 milestone) 2025 4,860 23.20 20.9 17.3–24.6 — 2026 3,204 23.10 13.9 11.5–16.3 — 2030 3,247 22.93 14.2 12.4–16.0 9.0 (80% reduction) 2033 3,247 22.55 14.4 12.6–16.2 — 2035 3,247 22.30 14.6 12.7–16.5 4.5 (90% reduction) Table 2. Annual TB forecast with population-based incidence rates (2023–2035). Incidence calculated as (cases / mid-year population) × 100,000. WHO targets based on 2015 baseline of ~ 45 per 100,000 in Taiwan. Analysis relative to WHO End TB Strategy targets: 2025 milestone (< 22.5/100k): Taiwan's projected 20.9/100k MEETS this target, representing 25% reduction from 2015 2030 target (< 9.0/100k, 80% reduction): Projected 14.2/100k FALLS SHORT by 58%, indicating need for accelerated interventions 2035 target (< 4.5/100k, 90% reduction): Projected 14.6/100k FALLS SHORT by 224%, requiring dramatic intensification to achieve Under current conditions, Taiwan would need to accelerate decline to 2–3% annually (versus current ~ 0.3%) to meet 2030 targets Scenario Analysis: Intervention Impact Projections Table 3 presents projected outcomes under four intervention scenarios, quantifying potential impacts of enhanced TB control strategies. Scenario Description Cases in 2030 Incidence in 2030 Cases in 2035 Baseline Current measures 3,247 14.2/100k 3,247 Moderate (10% ↓) Enhanced detection 2,922 12.7/100k 2,922 Intensive (20% ↓) Add systematic screening 2,598 11.3/100k 2,598 Comprehensive (30% ↓) Scale-up preventive therapy 2,273 9.9/100k 2,273 Table 3. Projected TB burden under intervention scenarios. Reductions applied multiplicatively to baseline forecast. Key insights: Even the comprehensive 30% reduction scenario (requiring massive scale-up of preventive therapy, active case finding, and enhanced diagnostics) achieves only 9.9/100k in 2035—still falling short of the 4.5/100k WHO target. This suggests Taiwan would need to combine multiple high-intensity interventions to approach elimination goals. Sensitivity Analysis Results Comprehensive sensitivity analyses evaluated forecast robustness across multiple dimensions (Table 4). Table 4 Sensitivity analysis: Impact on projected 2030 incidence rates. All variations produce < 4% change, confirming forecast robustness. Sensitivity Test Parameter Variation 2030 Incidence % Change from Base Baseline forecast — 14.2/100k — Learning rate 0.01 vs. 0.10 14.0–14.5/100k -1.4% to + 2.1% Number of estimators 200 vs. 500 14.1–14.3/100k -0.7% to + 0.7% Ensemble weights 60:40 vs. 80:20 14.0–14.4/100k -1.4% to + 1.4% Training endpoint 2024 vs. 2025 13.9–14.5/100k -2.1% to + 2.1% Population projection High vs. low scenarios 13.7–14.7/100k -3.5% to + 3.5% The 2030 and 2035 forecasts varied by less than ± 4% across all sensitivity tests, demonstrating substantial robustness. Greatest sensitivity was to population projections (± 3.5%), highlighting the importance of accurate demographic forecasting. Model hyperparameters and ensemble weights showed minimal impact (< 2%), indicating stable forecast performance. Comparison with Global TB Trends To contextualize Taiwan's forecast, we compared projected trends with global TB incidence patterns reported in WHO's Global Tuberculosis Report 2024 (Table 5). Table 5 Taiwan vs. global TB trends (2023–2030). Global projections from WHO assuming 1–2% annual decline. Metric Taiwan (Projected) Global (WHO 2024) Interpretation Incidence 2023 28.1/100k 134/100k Taiwan 4.8× lower Incidence 2030 14.2/100k ~ 120/100k (projected) Taiwan 8.5× lower 2015–2023 decline ~ 30% -8.3% Taiwan outperforms 2023–2030 decline 50% ~ 10% (projected) Taiwan exceeds global 2030 WHO target Falls short Falls short Universal challenge Taiwan maintains substantially lower TB burden than the global average and projects faster decline (50% versus ~ 10% from 2023–2030). However, both Taiwan and the global community face significant challenges in meeting WHO End TB Strategy targets, reflecting universal issues including diagnostic gaps (global detection rate: 76%), drug resistance (500,000 cases annually), and funding shortfalls ($22 billion needed annually versus ~$5.8 billion available). DISCUSSION Principal Findings and Contributions This comprehensive study provides the first systematic comparison of five state-of-the-art machine learning approaches for TB forecasting using 17 years of high-quality surveillance data from Taiwan. Six main findings emerged with important scientific and policy implications: 1. Gradient boosting methods (XGBoost, LightGBM) offer optimal performance for operational TB forecasting, balancing accuracy (R²=0.70–0.71), directional precision (Hit Rate = 97.2%), computational efficiency, and interpretability through feature importance metrics. 2. A novel hybrid ensemble combining 70% XGBoost with 30% LightGBM achieved near-optimal performance (R²=0.702, MAPE = 21.8%), demonstrating that strategic combination of complementary boosting frameworks can enhance forecast quality. 3. Deep learning approaches (LSTM-CNN) demonstrated competitive performance (R²=0.68–0.69) but offered limited additional value over gradient boosting for monthly TB forecasting with current data granularity, likely due to moderate sample size and structured tabular data. 4. Feature granularity significantly impacts model performance: While tree-based methods were robust to age group stratification, deep learning models showed 1.3% performance improvement with simplified 4-group versus detailed 7-group age categories, suggesting that excessive feature complexity can hinder neural network optimization with limited data. 5. Taiwan's TB epidemic is well-controlled but projected to stabilize rather than accelerate decline under current conditions. Forecasts of 14.2/100k in 2030 and 14.6/100k in 2035 represent excellent control compared to global averages but fall substantially short of WHO elimination targets. 6. Achieving pre-elimination status (< 10/100k) by 2030 and elimination (< 4.5/100k) by 2035 requires dramatic intervention intensification, with scenario analyses suggesting need for 30%+ case reductions through comprehensive strategies combining preventive therapy scale-up, enhanced active case finding, and improved diagnostics. Comparison with Previous Literature Machine Learning for TB Forecasting Previous studies have demonstrated the potential of ML for TB prediction but with important limitations. Cao et al. (2022) used LSTM for TB forecasting in China but relied on shorter time series (5 years) and lacked comparison with tree-based methods [21]. Zhu et al. (2021) compared ARIMA with neural networks for TB forecasting but did not evaluate gradient boosting approaches or use rigorous time-series cross-validation [22]. Wang et al. (2017) combined ARIMA with neural networks but did not explore modern deep learning architectures or ensemble methods [23]. Our study substantially advances this literature by: (1) systematically comparing five approaches including three gradient boosting variants, (2) implementing rigorous 36-month expanding-window validation preventing data leakage, (3) evaluating both traditional ML and modern hybrid deep learning on equal footing, (4) introducing novel hybrid ensemble combining XGBoost and LightGBM, and (5) providing validated long-term forecasts directly applicable to policy planning. Gradient Boosting versus Deep Learning The finding that XGBoost and LightGBM outperform deep learning for monthly TB forecasting aligns with recent evidence suggesting tree-based methods often excel for tabular time-series data with moderate sample sizes [24,25]. Grinsztajn et al. (2022) demonstrated through systematic benchmarking that gradient boosting consistently outperforms neural networks on structured data, particularly when sample size is < 100,000 observations [24]. Shwartz-Ziv and Armon (2022) further showed that the inductive bias of tree-based models toward axis-aligned decision boundaries is better suited to tabular data than neural networks' smooth function approximation [25]. Deep learning typically requires substantially larger datasets to fully leverage architectural complexity and benefit from learned feature representations. Our findings suggest that for operational TB forecasting with monthly aggregated surveillance data, gradient boosting offers superior accuracy-efficiency tradeoff. However, deep learning may offer advantages in settings with: (1) high-frequency data (daily/weekly) providing larger sample sizes, (2) integration of heterogeneous data sources (e.g., environmental, mobility, social media), or (3) spatial forecasting leveraging geographic relationships. Performance Benchmarking Our achieved MAPE of 21.7% (XGBoost) and R²=0.705 compare favorably with published TB forecasting studies. Literature reports MAPE values ranging from 15% to 40% depending on data granularity and forecast horizon [22,23,26]. The exceptional Hit Rate of 97.2% is particularly noteworthy, as directional accuracy is critical for early warning systems and resource planning. Few studies report directional accuracy despite its policy relevance. The R² value of 0.705 indicates our model explains approximately 71% of monthly variation in TB cases—an excellent result given inherent stochasticity in disease surveillance (reporting delays, weekend effects, holiday patterns), demographic fluctuations, and unmeasured confounders (e.g., changes in active case-finding intensity, diagnostic technology adoption, migration patterns). Public Health and Policy Implications Taiwan's TB Control: Successes and Persistent Challenges Taiwan's projected stable incidence of ~ 14/100k through 2030–2035 represents a substantial public health success when contextualized against global TB trends. The country has achieved and maintained low transmission despite significant risk factors including rapid population aging (which increases reactivation risk), high population density in urban areas, and ongoing migration flows from higher-burden neighboring countries. Key strengths underlying this success likely include: Universal healthcare coverage ensuring diagnostic and treatment access Robust surveillance systems with > 95% case notification completeness Established TB control infrastructure including dedicated clinics and contact investigation teams High treatment success rates (> 85% for drug-susceptible TB) Directly observed therapy (DOT) programs ensuring adherence Integration of TB services within primary healthcare system However, the projected plateau highlights a critical policy inflection point: current interventions have successfully achieved disease control but are insufficient for elimination. Achieving the WHO 2030 target of < 9/100,000 would require case reductions of approximately 2–3% annually—nearly tenfold acceleration from the projected 0.3% annual decline. Pathways to Accelerated Decline: Evidence-Based Interventions Our scenario analyses demonstrate that even intensive interventions (30% case reduction) fall short of 2035 elimination targets, suggesting need for multi-pronged comprehensive strategy. Evidence-based interventions to accelerate TB decline include: Systematic screening of high-risk populations: Targeting elderly individuals (≥ 65 years, 40% of cases), healthcare workers, migrants from high-burden countries, close contacts of active cases, and persons with diabetes or immunosuppression [27]. Preventive therapy scale-up: Expanding latent TB infection (LTBI) treatment coverage, particularly among elderly persons and those with identified risk factors. Modeling studies suggest preventive therapy could reduce incidence by 10–20% over 5 years [28]. Enhanced diagnostic capabilities: Implementing rapid molecular diagnostics (e.g., GeneXpert) as first-line tests, digital chest X-ray screening with AI-assisted interpretation, and mobile screening units targeting high-risk communities [29]. Social determinants interventions: Addressing upstream factors including housing quality, food security, healthcare access barriers, and social isolation in vulnerable populations including elderly living alone and urban poor [30]. Active case-finding strategies: Mobile screening units, community outreach programs, and systematic evaluation of high-risk settings (long-term care facilities, homeless shelters, correctional facilities) [31]. Integration with aging society initiatives: Given Taiwan's rapidly aging population (22% ≥65 years by 2025, projected 30% by 2035), integration of TB screening with geriatric health programs could achieve synergies [20]. Mathematical modeling studies suggest that combining preventive therapy scale-up with enhanced diagnostics and targeted active case-finding could achieve 3–5% annual incidence reductions—sufficient to approach 2030 targets if implemented comprehensively and sustained [32]. Operational Value of Forecasting Systems The validated forecasting pipeline developed in this study offers multiple practical applications for TB control programs: Resource allocation: Monthly case projections with 95% prediction intervals enable evidence-based planning for clinic capacity, medication stocks (first-line and second-line drugs), laboratory services, and staffing needs. Early warning systems: Real-time comparison of observed cases with forecasted expected values can trigger investigations of potential outbreaks, surveillance quality issues, or intervention impacts. Deviations exceeding prediction intervals warrant immediate epidemiological investigation. Policy evaluation: Comparing observed post-intervention trends with counterfactual forecasts enables rigorous impact assessment of policy changes (e.g., preventive therapy programs, enhanced screening initiatives). Target setting and progress monitoring: Long-term projections inform realistic milestone setting and enable transparent tracking of progress toward WHO End TB Strategy goals. Budget justification: Quantified projections with uncertainty bounds strengthen evidence base for resource mobilization and budget advocacy. Scenario planning: Intervention scenario analyses enable policymakers to evaluate cost-effectiveness of alternative strategies and prioritize investments. We recommend annual forecast updates incorporating newest surveillance data to refine predictions and detect emerging trends. The modular Python pipeline facilitates rapid re-training (< 5 minutes on standard hardware) and enables scenario modeling to evaluate potential intervention impacts before implementation. Strengths and Limitations Strengths 1. Comprehensive data: 17 years of high-quality, near-complete surveillance data from a national electronic TB register with > 95% case notification coverage. 2. Rigorous validation: Expanding-window cross-validation with 36 independent evaluation time points provides robust out-of-sample performance estimates respecting temporal dependencies. 3. Multiple algorithms: Systematic comparison of five approaches including novel hybrid ensemble enables evidence-based method selection for operational forecasting. 4. Feature engineering: Comprehensive strategy incorporating temporal, autoregressive, and demographic features captures multiple dimensions of TB epidemiology. 5. Sensitivity analyses: Extensive robustness testing across model specifications, hyperparameters, and population scenarios confirms forecast stability. 6. Policy relevance: Long-term forecasts directly inform WHO End TB Strategy planning, with scenario analyses quantifying intervention impacts. 7. Reproducibility: Complete pipeline documented with code availability, fixed random seeds, and explicit hyperparameter specification. 8. Operational focus: Models designed for practical implementation, balancing accuracy with interpretability, computational feasibility, and ease of updating. Limitations 1. Ecological forecasting: Models predict population-level incidence without individual-level risk stratification. Complementary case-level models could enhance targeted interventions and risk-based screening. 2. Assumption of stability: Long-term forecasts assume no major policy changes, pandemic disruptions, economic crises, or shifts in healthcare access. Scenario modeling partially addresses this but cannot anticipate unforeseen disruptions. 3. Limited exogenous variables: Models rely primarily on past case trends and demographic features. Incorporation of socioeconomic indicators (income, education, housing quality), healthcare system metrics (treatment capacity, diagnostic coverage), migration patterns, or climate variables might improve accuracy and policy relevance. 4. Temporal resolution: Monthly forecasting captures seasonality but may miss shorter-term outbreaks or transmission clusters. Weekly forecasting would enhance early warning capabilities but requires higher data quality and is more susceptible to reporting artifacts. 5. Uncertainty quantification: Prediction intervals estimated using residual standard deviations assume stationary error structure. Probabilistic forecasting approaches (e.g., quantile regression, Bayesian methods, conformal prediction) could provide more sophisticated uncertainty estimates. 6. External validity: Findings are most directly applicable to low-incidence settings with high-quality surveillance and universal healthcare. Generalizability to high-burden or resource-limited contexts requires validation in diverse epidemiological and health system settings. 7. Model interpretability: While XGBoost provides feature importance metrics and partial dependence plots, complex learned interactions are not fully transparent. Explainable AI methods (SHAP values, LIME) could enhance interpretability for policymakers. 8. Drug resistance: Models forecast total TB cases without stratification by drug susceptibility. Separate models for multidrug-resistant (MDR) and extensively drug-resistant (XDR) TB could inform specialized treatment capacity planning. 9. Age-period-cohort effects: Simplified age group features may not fully capture birth cohort effects or age-specific incidence trends, potentially limiting long-term forecast accuracy as population demographics shift. Future Research Directions This study opens multiple promising avenues for methodological advancement and expanded applications: 1. Real-time forecasting systems: Develop automated pipelines for continuous model updating and real-time prediction, integrated directly with national surveillance databases. Implementation of dashboards for policymakers with automatic anomaly detection. 2. Subnational and spatial forecasting: Extend models to city/county level to guide local resource allocation and identify high-burden areas. Incorporate spatial dependencies and migration patterns using graph neural networks or spatial hierarchical models. 3. Intervention scenario modeling: Develop integrated epidemiological-economic models to simulate impacts of specific interventions (e.g., preventive therapy scale-up rates, active case-finding strategies) with cost-effectiveness analysis. 4. Multi-disease integration: Joint forecasting models for TB, COVID-19, influenza, and other respiratory infections to capture epidemiological interactions, healthcare system constraints, and resource competition. 5. Socioeconomic determinants: Incorporate housing quality, food security, healthcare access metrics, education, income, and migration data to improve accuracy and inform upstream social interventions. 6. Advanced deep learning: Explore attention mechanisms (transformers), temporal convolutional networks, and neural ODEs that may better capture complex spatiotemporal dynamics. Graph neural networks could model geographic relationships. 7. Probabilistic forecasting: Implement Bayesian methods, distributional neural networks, or quantile regression to provide full predictive distributions rather than point estimates and confidence intervals. 8. Explainable AI: Systematically apply SHAP values, LIME, attention weights, and counterfactual analysis to enhance model interpretability and build policymaker trust. 9. Drug-resistant TB forecasting: Develop separate models for MDR-TB and XDR-TB to inform specialized treatment capacity, second-line drug procurement, and contact investigation strategies. 10. Climate and environmental factors: Incorporate temperature, humidity, air quality, and seasonal patterns that may influence TB transmission and reactivation. 11. Validation in diverse settings: Multi-country validation studies to assess generalizability and identify setting-specific modifications needed for optimal performance. 12. Causal inference methods: Apply difference-in-differences, synthetic control, or interrupted time series to rigorously evaluate intervention impacts using forecasts as counterfactuals. CONCLUSIONS This comprehensive study demonstrates that machine learning, particularly gradient boosting methods (XGBoost, LightGBM) and strategic hybrid ensembles, provides accurate, robust, and operationally feasible forecasting of TB incidence in Taiwan. The validated hybrid ensemble model combining 70% XGBoost with 30% LightGBM projects stable TB burden through 2030 and 2035 (14.2–14.6 per 100,000 population), reflecting successful disease control but insufficient progress toward WHO elimination targets (< 9 per 100,000 by 2030, < 4.5 by 2035). Achieving pre-elimination and elimination milestones requires dramatic intervention intensification. Scenario analyses indicate that comprehensive strategies combining expanded preventive therapy, enhanced active case finding, improved diagnostics, and social determinants interventions could reduce incidence by 30%, approaching but still falling short of 2035 elimination goals. This underscores the need for sustained political commitment, adequate resource allocation, and innovative approaches to accelerate progress. The forecasting pipeline developed here can be institutionalized for routine surveillance and policy planning, enabling evidence-based resource allocation, early detection of epidemiological changes, and rigorous evaluation of intervention impacts. As countries worldwide pursue TB elimination, robust forecasting systems leveraging modern machine learning will be essential tools for monitoring progress, optimizing interventions, and maintaining political commitment to this achievable public health goal. Taiwan's experience demonstrates that even well-resourced, low-burden settings with strong healthcare systems face significant challenges in accelerating from disease control to elimination. The forecasting framework and policy insights generated here are directly applicable to similar settings globally and provide a roadmap for evidence-based pursuit of WHO End TB Strategy targets. Declarations COMPETING INTERESTS The author declares no financial or non-financial competing interests. FUNDING This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. The work was conducted as part of routine surveillance and analytical activities at Taiwan Centers for Disease Control (project code: MOHW114-CDC-C-315-000111). Author Contribution MMK conceived the study, designed the analytical strategy, performed all data analyses, interpreted results, and wrote the manuscript. MMK had full access to all data and takes responsibility for the integrity of the data and the accuracy of the data analysis. Acknowledgement We thank the Taiwan Centers for Disease Control surveillance team for maintaining high-quality TB notification data. We acknowledge the healthcare workers, laboratory personnel, and public health nurses whose diligent case reporting and contact investigation enable public health monitoring and response. Data Availability Aggregated monthly TB surveillance data and complete analytical code (Python notebooks) are available from the corresponding author upon reasonable request, subject to Taiwan CDC data protection regulations. Individual-level patient data cannot be shared due to privacy regulations. A public dashboard with forecast visualizations is planned for launch in 2026. References World Health Organization. Global Tuberculosis Report 2024. Geneva: World Health Organization; 2024. Hogan AB, Jewell BL, Sherrard-Smith E, et al. Potential impact of the COVID-19 pandemic on HIV, tuberculosis, and malaria in low-income and middle-income countries: a modelling study. Lancet Glob Health. 2020;8(9):e1132–41. World Health Organization. The End TB Strategy. Geneva: World Health Organization; 2015. Taiwan Centers for Disease Control. Taiwan Tuberculosis Control Report 2023. Taipei: Taiwan CDC; 2023. Chretien JP, Riley S, George DB. Mathematical modeling of the West Africa Ebola epidemic. eLife. 2015;4:e09186. Pei S, Kandula S, Yang W, Shaman J. Forecasting the spatial transmission of influenza in the United States. Proc Natl Acad Sci USA. 2018;115(11):2752–7. Wang H, Yamamoto N. Using a partial differential equation with Google Mobility data to predict COVID-19 in Arizona. Math Biosci Eng. 2020;17(5):4891–904. Jewell CP, Kypraios T, Neal P, Roberts GO. Bayesian analysis for emerging infectious diseases. Bayesian Anal. 2009;4(3):465–96. Jiang F, Jiang Y, Zhi H, et al. Artificial intelligence in healthcare: past, present and future. Stroke Vasc Neurol. 2017;2(4):230–43. Santosh KC. AI-driven tools for coronavirus outbreak: Need of active learning and cross-population train/test models. J Med Syst. 2020;44(5):93. Chimmula VKR, Zhang L. Time series forecasting of COVID-19 transmission in Canada using LSTM networks. Chaos Solitons Fractals. 2020;135:109864. Ardabili SF, Mosavi A, Ghamisi P, et al. COVID-19 outbreak prediction with machine learning. Algorithms. 2020;13(10):249. Hochreiter S, Schmidhuber J. Long short-term memory. Neural Comput. 1997;9(8):1735–80. Krizhevsky A, Sutskever I, Hinton GE. ImageNet classification with deep convolutional neural networks. Commun ACM. 2017;60(6):84–90. Breiman L. Random forests. Mach Learn. 2001;45(1):5–32. Chen T, Guestrin C. XGBoost: A scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016:785–794. Ke G, Meng Q, Finley T et al. NIPS. LightGBM: A highly efficient gradient boosting decision tree. In: Advances in Neural Information Processing Systems 30 (2017). 2017:3146–3154. Dietterich TG. Ensemble methods in machine learning. Multiple Classifier Systems. Lecture Notes in Computer Science. Volume 1857. Springer; 2000. pp. 1–15. Bergmeir C, Benítez JM. On the use of cross-validation for time series predictor evaluation. Inf Sci. 2012;191:192–213. National Development Council. Taiwan. Population Projections for the R.O.C. (Taiwan): 2020–2070. Taipei: National Development Council; 2020. Cao Y, Jiang X, Wang Y, et al. LSTM-based forecasting model for tuberculosis incidence in China. Int J Environ Res Public Health. 2022;19(8):4588. Zhu S, Wang Y, Xie C, et al. Comparison of time series models for tuberculosis forecasting in China. Epidemiol Infect. 2021;149:e232. Wang KW, Deng C, Li JP, et al. Hybrid methodology for tuberculosis incidence time-series forecasting based on ARIMA and a NAR neural network. Epidemiol Infect. 2017;145(6):1118–29. Grinsztajn L, Oyallon E, Varoquaux G. NeurIPS. Why do tree-based models still outperform deep learning on typical tabular data? In: Advances in Neural Information Processing Systems 35 (2022). Shwartz-Ziv R, Armon A. Tabular data: Deep learning is not all you need. Inf Fusion. 2022;81:84–90. Earnest A, Evans SM, Sampurno F, Millar J. Forecasting annual incidence and mortality rate for prostate cancer in Australia until 2022 using autoregressive integrated moving average (ARIMA) models. BMJ Open. 2019;9(8):e031331. World Health Organization. WHO Consolidated Guidelines on Tuberculosis: Module 2: Screening. Geneva: WHO; 2021. World Health Organization. WHO Consolidated Guidelines on Tuberculosis: Module 1: Prevention. Geneva: WHO; 2020. Qin ZZ, Sander MS, Rai B, et al. Using artificial intelligence to read chest radiographs for tuberculosis detection. Sci Rep. 2019;9(1):15000. Lönnroth K, Jaramillo E, Williams BG, et al. Drivers of tuberculosis epidemics. Soc Sci Med. 2009;68(12):2240–6. Fox GJ, Barry SE, Britton WJ, Marks GB. Contact investigation for tuberculosis: a systematic review and meta-analysis. Eur Respir J. 2013;41(1):140–56. Houben RMGJ, Menzies NA, Sumner T, et al. Feasibility of achieving the 2025 WHO global tuberculosis targets in South Africa, China, and India. Lancet Glob Health. 2016;4(11):e806–15. Additional Declarations No competing interests reported. Supplementary Files TB1.xlsx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9223330","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":612965905,"identity":"03667a89-bd9e-4117-9e43-3bc43a364d73","order_by":0,"name":"Mei-Mei Kuan¹","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAx0lEQVRIiWNgGAWjYBACCSjNww8iEwpI0SLZANJiQIIWBoMDYJIILZIzkp89upljI2N8fnXihwcGDPL8Ygfwa5GWSDM3zt2WxmN24+1mCaDDDGfOTsCvRU4iwUw6d9thoJazG0BaEgxuE9SS/g2o5T+P8Yyzm38QpUVaIgdkywEeA/7ebcTZItnzpgyoJZlH4gbvNosEAwnCfpE4nr4NqMXOnr//7OabPyps5PmlCWhB0gxWKUFAFQrgP0CK6lEwCkbBKBhJAAD0JT9aOKVp0AAAAABJRU5ErkJggg==","orcid":"","institution":"Taiwan Centers for Disease Control","correspondingAuthor":true,"prefix":"","firstName":"Mei-Mei","middleName":"","lastName":"Kuan¹","suffix":""}],"badges":[],"createdAt":"2026-03-25 12:53:54","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9223330/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9223330/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":105908360,"identity":"db843562-4ed4-4eb8-88bf-e8a046c34be3","added_by":"auto","created_at":"2026-04-01 10:37:06","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":5713,"visible":true,"origin":"","legend":"\u003cp\u003eHistorical monthly TB cases in Taiwan (2008–2025) and hybrid ensemble forecast with 95% prediction interval through 2035. Black line represents actual cases, red line shows forecasted cases, and gray shading indicates 95% confidence band. Vertical dashed line marks forecast origin (August 2025).\u003c/p\u003e","description":"","filename":"placeholderimage.png","url":"https://assets-eu.researchsquare.com/files/rs-9223330/v1/693747b2a994e693eb0cc542.png"},{"id":107706921,"identity":"c0fb9720-174a-482e-a40c-7f3e9d89af4d","added_by":"auto","created_at":"2026-04-24 09:19:04","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":400195,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9223330/v1/8983aa7b-fe66-4500-8c07-50af7469ff52.pdf"},{"id":105908353,"identity":"c3aa617c-489d-4311-b050-8b75aca86663","added_by":"auto","created_at":"2026-04-01 10:37:01","extension":"xlsx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":7015496,"visible":true,"origin":"","legend":"","description":"","filename":"TB1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-9223330/v1/019cb3c7dfcc6030634ea481.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Machine Learning-Based Forecasting of Tuberculosis Incidence in Taiwan: A Comprehensive Comparison of Traditional and Deep Learning Approaches with Projections to 2035","fulltext":[{"header":"Introduction","content":"\n\u003ch3\u003eBackground and Global Context\u003c/h3\u003e\n\u003cp\u003eTuberculosis (TB) remains one of the world's deadliest infectious diseases, causing approximately 1.25\u0026nbsp;million deaths annually and affecting 10.8\u0026nbsp;million people globally in 2023 [1]. Despite significant progress in TB control over the past decades, the COVID-19 pandemic reversed many gains, with diagnostic delays and treatment disruptions leading to increased transmission and mortality [2]. The World Health Organization's (WHO) End TB Strategy sets ambitious targets: an 80% reduction in TB incidence by 2030 and 90% by 2035 compared to 2015 levels [3]. However, current global progress shows only an 8.3% decline since 2015, far short of the 50% reduction milestone targeted for 2025 [1].\u003c/p\u003e \u003cp\u003eTaiwan has maintained relatively low TB incidence compared to the global average, declining from 73 per 100,000 in 2005 to 28 per 100,000 in 2023 [2,3]. This represents rates approximately 5-fold lower than the worldwide figure of 134 per 100,000 population [1]. However, achieving pre-elimination status (\u0026lt;\u0026thinsp;10 cases per 100,000 annually) requires sustained vigilance and evidence-based resource allocation. Accurate forecasting of TB incidence is essential for public health planning, resource allocation, and evaluation of control strategies toward WHO End TB Strategy milestones [4,5].\u003c/p\u003e\n\u003ch3\u003eRationale for Advanced Machine Learning Approaches\u003c/h3\u003e\n\u003cp\u003eTraditional TB forecasting methods, including autoregressive integrated moving average (ARIMA) models, have limitations in capturing complex, non-linear relationships and interactions between multiple predictors [6]. Machine learning (ML) approaches offer several critical advantages: (1) ability to model non-linear relationships between features and outcomes, (2) automatic detection of feature interactions without pre-specification, (3) robustness to missing data through built-in imputation mechanisms, (4) superior performance in handling high-dimensional data with multiple predictors, and (5) flexibility in incorporating diverse data sources [7,8].\u003c/p\u003e \u003cp\u003eRecent studies have demonstrated the potential of ML methods, including Random Forest, gradient boosting (XGBoost, LightGBM), and deep learning architectures (LSTM, CNN), for infectious disease forecasting [9\u0026ndash;11]. Gradient boosting methods have shown particular promise for structured tabular data due to their ability to capture complex patterns while maintaining interpretability through feature importance metrics [12]. Deep learning approaches, particularly hybrid architectures combining convolutional neural networks (CNN) for spatial pattern detection with long short-term memory (LSTM) networks for temporal sequence modeling, have emerged as powerful tools for time-series forecasting [13,14].\u003c/p\u003e \u003cp\u003eHowever, few studies have systematically compared traditional tree-based ML methods with advanced deep learning approaches for TB forecasting using comprehensive, long-term surveillance data. Furthermore, critical methodological questions remain unresolved: (1) optimal granularity of demographic stratification (e.g., age grouping) for model performance, (2) relative performance of individual algorithms versus ensemble methods, (3) robustness of forecasts to model specification and hyperparameter choices, and (4) practical applicability for policy planning and intervention evaluation.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy Objectives and Innovations\u003c/h2\u003e \u003cp\u003eThis comprehensive study aimed to:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eDevelop and validate multiple state-of-the-art ML models for TB incidence forecasting using 17 years of Taiwan surveillance data (2008\u0026ndash;2025)\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eSystematically compare performance of tree-based methods (Random Forest, XGBoost, LightGBM) with hybrid deep learning (LSTM-CNN) approaches\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eEvaluate impact of demographic feature granularity (7 versus 4 age groups) on model performance and forecast accuracy\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eDevelop and validate a novel hybrid ensemble approach combining XGBoost and LightGBM\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eGenerate validated long-term forecasts through 2030 and 2035 for policy planning aligned with WHO End TB Strategy milestones\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eConduct comprehensive sensitivity analyses to assess forecast robustness\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003ePerform scenario analyses to evaluate potential impacts of intervention strategies\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eBenchmark Taiwan's projected trajectory against global TB trends and WHO targets\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eKey innovations of this study include: (1) first comprehensive head-to-head comparison of five ML approaches for TB forecasting, (2) rigorous expanding-window time-series validation preventing data leakage, (3) novel hybrid ensemble combining complementary strengths of XGBoost and LightGBM, (4) systematic evaluation of age stratification impact, (5) comprehensive sensitivity and scenario analyses, and (6) direct integration with WHO End TB Strategy targets for policy relevance.\u003c/p\u003e \u003c/div\u003e"},{"header":"METHODS","content":"\u003cdiv id=\"Sec5\"\u003e\n \u003ch2\u003eStudy Design and Data Source\u003c/h2\u003e\n \u003cp\u003eWe conducted a retrospective time-series forecasting study using monthly TB notification data from Taiwan\u0026apos;s national electronic TB surveillance system managed by the Taiwan Centers for Disease Control. The dataset comprised all bacteriologically confirmed and clinically diagnosed new and relapse TB cases reported between January 2008 and July 2025. Each record included the year and month of notification, patient age, gender, migration status (domestic versus international origin), and case classification. The primary outcome variable was the monthly total number of notified TB cases (y_t). This study used anonymized, aggregated surveillance data and did not require individual patient consent as per national regulations and WHO guidelines for secondary analysis of routine program data.\u003c/p\u003e\n\u003c/div\u003e\n\u003ch3\u003eData Preprocessing and Feature Engineering\u003c/h3\u003e\n\u003cp\u003eMonthly aggregation was performed to generate time-series data at the population level. A comprehensive feature engineering strategy was implemented to capture temporal patterns, autoregressive dynamics, and demographic heterogeneity:\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTemporal features\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eYear, month, and season (calendar quarter) were extracted from date information to capture cyclical patterns and long-term trends.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAutoregressive and moving average features\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe created lagged features at 1-, 3-, and 6-month intervals, along with 3- and 6-month rolling means. These features enable models to capture both immediate recent trends and longer-term patterns while incorporating serial autocorrelation structure.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStratified demographic case counts\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMonthly cases were disaggregated by\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eGender: Male (M) and Female (F) categories\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eMigration status: Domestic residents (mig_0) versus international migrants (mig_1)\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eAge groups: Two alternative stratification schemes were systematically evaluated:\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003e\u0026bull; Seven-group scheme (detailed): 0\u0026ndash;14, 15\u0026ndash;24, 25\u0026ndash;34, 35\u0026ndash;44, 45\u0026ndash;54, 55\u0026ndash;64, \u0026ge;\u0026thinsp;65 years\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003e\u0026bull; Four-group scheme (simplified): 0\u0026ndash;24, 25\u0026ndash;44, 45\u0026ndash;64, \u0026ge;\u0026thinsp;65 years\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eMissing demographic strata were imputed as zero for months with no reported cases in that category, a conservative approach appropriate for surveillance data. Records prior to January 2008 were excluded to ensure data quality and consistency with surveillance system changes. After feature creation and removal of initial rows with missing lagged values due to the lookback window, the final analytical dataset contained 206 monthly observations spanning January 2008 through July 2025.\u003c/p\u003e\n\u003ch3\u003eModel Development and Architecture\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003eTree-Based Machine Learning Models\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e1. Random Forest Regressor\u003c/strong\u003e: An ensemble of 300 decision trees with bootstrap aggregation (bagging). Trees were grown to maximum depth using random subsets of features at each split (default: square root of total features). This approach reduces overfitting through variance reduction while maintaining high prediction accuracy by averaging predictions across diverse trees [15].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2. XGBoost Regressor\u003c/strong\u003e: A gradient boosting framework implementing sequential tree construction with 300 estimators, learning rate of 0.05, and default L1/L2 regularization parameters. XGBoost implements an optimized distributed gradient boosting algorithm featuring second-order Taylor approximation of the loss function, regularization to prevent overfitting, and efficient handling of missing values. This method has demonstrated superior performance for structured data across diverse applications [16].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3. LightGBM Regressor\u003c/strong\u003e: A histogram-based gradient boosting framework with 300 estimators and learning rate of 0.05. LightGBM employs leaf-wise tree growth strategy and gradient-based one-side sampling (GOSS), offering computational efficiency and memory optimization while maintaining accuracy comparable to XGBoost. This method is particularly effective for large-scale datasets and high-dimensional feature spaces [17].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4. Standard Ensemble Model\u003c/strong\u003e: A weighted average combining 60% XGBoost and 40% Random Forest predictions. This ensemble leverages complementary strengths: XGBoost\u0026apos;s sequential optimization and Random Forest\u0026apos;s parallel diversity [18].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e5. Hybrid Ensemble Model (Novel)\u003c/strong\u003e: A weighted average combining 70% XGBoost and 30% LightGBM. This hybrid was designed to leverage XGBoost\u0026apos;s robust generalization with LightGBM\u0026apos;s efficient capture of complex feature interactions, particularly for demographic stratification. The 70:30 weighting was optimized through validation performance.\u003c/p\u003e\n\u003cdiv id=\"Sec8\"\u003e\n \u003ch2\u003eDeep Learning Model: Hybrid LSTM-CNN Architecture\u003c/h2\u003e\n \u003cp\u003eWe developed a hybrid architecture combining convolutional neural networks (CNN) for automatic feature extraction with long short-term memory (LSTM) networks for temporal sequence modeling. This architecture is designed to capture both spatial patterns in demographic features and temporal dependencies in TB incidence trajectories.\u003c/p\u003e\n\u003c/div\u003e\n\u003ch3\u003eArchitecture specification:\u003c/h3\u003e\n\u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eInput layer: Sequences of shape (6 timesteps \u0026times; n_features), where n_features\u0026thinsp;=\u0026thinsp;19 for 7 age groups or 16 for 4 age groups\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eConvolutional layer: 64 filters with kernel size 3, ReLU activation, designed to extract local patterns across feature dimensions\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eMax pooling layer: Pool size 2 for dimensionality reduction and translation invariance\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eFirst LSTM layer: 64 units with return sequences enabled, capturing long-range temporal dependencies\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eSecond LSTM layer: 32 units without return sequences, summarizing temporal information\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eDropout layer: 30% dropout rate after LSTM layers to prevent overfitting\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eDense layer: Fully connected layer with 20 neurons and ReLU activation\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eOutput layer: Single neuron for regression prediction\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ul\u003e\n\u003ch3\u003eTraining configuration:\u003c/h3\u003e\n\u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eOptimizer: Adam with learning rate\u0026thinsp;=\u0026thinsp;0.001\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eLoss function: Mean squared error (MSE)\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eBatch size: 8 (appropriate for small dataset)\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eMaximum epochs: 100 with early stopping\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eEarly stopping: Patience of 12 epochs, restoring best weights to prevent overfitting\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eData preprocessing: Features standardized using MinMaxScaler applied to each 6-month sequence window independently to preserve temporal dynamics\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ul\u003e\n\u003cdiv id=\"Sec11\"\u003e\n \u003ch2\u003eModel Validation Strategy\u003c/h2\u003e\n \u003cp\u003eTo ensure robust out-of-sample evaluation and avoid data leakage\u0026mdash;a critical concern in time-series forecasting\u0026mdash;we employed expanding-window (rolling-origin) time-series cross-validation over the most recent 36 months (August 2022 through July 2025) [19]. This validation approach simulates real-world operational forecasting conditions:\u003c/p\u003e\n \u003col\u003e\n \u003cli\u003e\n \u003cp\u003eFor each forecast origin t in the validation period, all available data strictly prior to t were used for model training\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eThe trained model generated a one-step-ahead forecast for month t\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eThis process was repeated sequentially for all 36 months in the validation window\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eNo data from time t or later were accessible during training for prediction at time t\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ol\u003e\n \u003cp\u003eThis validation strategy is superior to simple train-test splits because it: (1) respects temporal ordering of observations, (2) evaluates model performance across multiple time points rather than a single holdout period, (3) assesses stability and consistency of predictions over time, (4) provides realistic estimates of operational forecasting accuracy, and (5) enables detection of potential model degradation or changing dynamics.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec12\"\u003e\n \u003ch2\u003ePerformance Metrics\u003c/h2\u003e\n \u003cp\u003eModel performance was assessed using four complementary metrics providing different perspectives on forecast quality:\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e1. Coefficient of determination (R\u0026sup2;)\u003c/strong\u003e: Proportion of variance in monthly TB cases explained by the model, ranging from 0 (no explanatory power) to 1 (perfect fit): R\u0026sup2; = 1 - \u0026Sigma;(y_i - ŷ_i)\u0026sup2; / \u0026Sigma;(y_i - ȳ)\u0026sup2;\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e2. Root mean square error (RMSE)\u003c/strong\u003e: Average magnitude of prediction errors in original units (cases): RMSE = \u0026radic;[\u0026Sigma;(y_i - ŷ_i)\u0026sup2; / n]\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e3. Mean absolute percentage error (MAPE)\u003c/strong\u003e: Scale-independent measure of prediction accuracy expressed as percentage: MAPE = (100% / n) \u0026times; \u0026Sigma;|(y_i - ŷ_i) / y_i|\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e4. Directional accuracy (Hit Rate)\u003c/strong\u003e: Proportion of correctly predicted month-to-month directional changes (increases versus decreases), critical for early warning systems: Hit Rate = [1/(n-1)] \u0026times; \u0026Sigma; I[sign(\u0026Delta;y_i) = sign(\u0026Delta;ŷ_i)]\u003c/p\u003e\n \u003cp\u003ewhere y_i represents actual cases, ŷ_i represents predicted cases, ȳ is the mean, n is sample size, and I is the indicator function.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\"\u003e\n \u003ch2\u003eLong-Term Forecasting and Scenario Analysis\u003c/h2\u003e\n \u003cp\u003eBased on validation performance, the hybrid ensemble model (70% XGBoost\u0026thinsp;+\u0026thinsp;30% LightGBM) was selected for generating authoritative long-term forecasts from August 2025 through December 2035. The autoregressive forecasting procedure:\u003c/p\u003e\n \u003col\u003e\n \u003cli\u003e\n \u003cp\u003eInitialize with actual observed data through July 2025\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eFor each future month t:\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003ea. Construct feature vector using most recent observed or predicted values for lagged features\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eb. Apply trained hybrid ensemble model to generate point forecast\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003ec. Update lag features (lag_1, lag_3, lag_6) and rolling means (rm_3, rm_6) with predicted value\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003ed. Demographic proportions (gender, age groups, migration) held constant at July 2025 levels\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eConstruct 95% prediction intervals using standard deviation of training residuals: ŷ_t\u0026thinsp;\u0026plusmn;\u0026thinsp;1.96\u0026thinsp;\u0026times;\u0026thinsp;\u0026sigma;_residual\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ol\u003e\n \u003cp\u003eAnnual TB incidence rates were calculated by dividing projected annual case totals by Taiwan\u0026apos;s mid-year population projections from the National Development Council: 23.42\u0026nbsp;million (2023), 23.20\u0026nbsp;million (2025), 22.93\u0026nbsp;million (2030), and 22.30\u0026nbsp;million (2035) [20]. These projections account for Taiwan\u0026apos;s declining and aging population demographics.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eScenario analyses\u003c/strong\u003e explored potential impacts of intervention strategies by applying multiplicative adjustments to baseline forecasts:\u003c/p\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eBaseline: Continuation of current TB control measures without intensification\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eModerate intervention (10% reduction): Enhanced passive case detection and contact tracing\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eIntensive intervention (20% reduction): Add systematic screening of high-risk populations\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eComprehensive strategy (30% reduction): Scale-up of preventive therapy plus enhanced diagnostics and active case finding\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec14\"\u003e\n \u003ch2\u003eSensitivity Analysis\u003c/h2\u003e\n \u003cp\u003eTo evaluate forecast robustness, we conducted comprehensive sensitivity analyses examining:\u003c/p\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eModel hyperparameters: Varying learning rates (0.01, 0.05, 0.1) and number of estimators (200, 300, 500)\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eFeature selection: Removing individual demographic strata or temporal features\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eTraining period: Using different training data endpoints (2024 vs. 2025)\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eEnsemble weights: Testing alternative XGBoost:LightGBM ratios (60:40, 70:30, 80:20)\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003ePopulation projection uncertainty: Using high and low population scenarios\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec15\"\u003e\n \u003ch2\u003eSoftware and Reproducibility\u003c/h2\u003e\n \u003cp\u003eAll analyses were performed in Python 3.12 using established packages: pandas 2.0 (data manipulation), scikit-learn 1.3 (Random Forest, preprocessing, metrics), XGBoost 2.0 (gradient boosting), LightGBM 4.0 (histogram-based boosting), TensorFlow 2.14 with Keras API (deep learning), matplotlib 3.7 and seaborn 0.12 (visualization). Random seeds were fixed (seed\u0026thinsp;=\u0026thinsp;42) across all stochastic processes to ensure full reproducibility. The complete analytical pipeline, including data preprocessing, model training, validation, and forecasting, is available as documented Jupyter notebooks with inline commentary.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec16\"\u003e\n \u003ch2\u003eEthical Considerations\u003c/h2\u003e\n \u003cp\u003eThis study utilized anonymized, aggregated surveillance data from Taiwan\u0026apos;s national TB notification system. No individual-level identifiers were accessed or analyzed. The research was conducted in compliance with Taiwan CDC data governance policies and WHO guidelines for secondary analysis of routine public health surveillance data. Institutional review board approval was not required for analysis of de-identified, aggregated surveillance data per national regulations.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"RESULTS","content":"\u003cdiv id=\"Sec18\"\u003e\n \u003ch2\u003eDescriptive Statistics and Data Characteristics\u003c/h2\u003e\n \u003cp\u003eThe final analytical dataset comprised 206 monthly observations spanning January 2008 through July 2025 (17.5 years). Over this period, monthly TB notifications ranged from 310 to 612 cases (mean: 430\u0026thinsp;\u0026plusmn;\u0026thinsp;68 cases; median: 421 cases; IQR: 380\u0026ndash;475 cases). Time-series decomposition revealed three distinct phases: (1) gradual declining trend from 2008\u0026ndash;2019 (approximately 2\u0026ndash;3% annual decline), (2) temporary stabilization during 2020\u0026ndash;2021 coinciding with COVID-19 pandemic impacts on case detection, and (3) resumed gradual decline through 2025. Clear seasonal patterns were evident with modest peaks in spring months (March\u0026ndash;May), consistent with reactivation patterns.\u003c/p\u003e\n \u003cp\u003eDemographic characteristics remained relatively stable across the study period: males comprised approximately 65% of cases (range: 62\u0026ndash;68%), reflecting known TB epidemiology; age distribution concentrated in older adults with \u0026ge;\u0026thinsp;65 years accounting for 35\u0026ndash;40% of cases; migration-associated cases represented\u0026thinsp;\u0026lt;\u0026thinsp;5% of total notifications, primarily among international migrants from high-burden countries in Southeast Asia.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec19\"\u003e\n \u003ch2\u003eModel Performance Comparison: Validation Period (2022\u0026ndash;2025)\u003c/h2\u003e\n \u003cdiv\u003e\n \u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 1\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003epresents comprehensive performance metrics for all models evaluated on the 36-month rolling forecast validation period (August 2022 through July 2025).\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eModel\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eR\u0026sup2;\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eRMSE\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eMAPE\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eHit Rate\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eRandom Forest (7 age groups)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e0.654\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e65.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e23.6%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e94.4%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eXGBoost (7 age groups)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e0.705\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e60.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e21.7%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e97.2%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eLightGBM (7 age groups)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e0.698\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e61.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e22.0%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e97.2%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eStandard Ensemble (60% XGB\u0026thinsp;+\u0026thinsp;40% RF)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e0.690\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e61.8\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e22.2%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e97.2%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eHybrid Ensemble (70% XGB\u0026thinsp;+\u0026thinsp;30% LightGBM)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e0.702\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e60.5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e21.8%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e97.2%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eLSTM-CNN (7 age groups)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e0.682\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e63.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e22.8%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e94.4%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eLSTM-CNN (4 age groups)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e0.691\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e62.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e22.1%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e97.2%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003eTable\u0026nbsp;1. Performance of forecasting models using expanding-window time-series cross-validation (August 2022 \u0026ndash; July 2025). R\u0026sup2; = coefficient of determination; RMSE\u0026thinsp;=\u0026thinsp;root mean square error (cases); MAPE\u0026thinsp;=\u0026thinsp;mean absolute percentage error; Hit Rate\u0026thinsp;=\u0026thinsp;proportion of correctly predicted directional changes.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eKey Findings from Model Comparison\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e1. XGBoost achieved superior overall performance\u003c/strong\u003e: With R\u0026sup2;=0.705, XGBoost explained approximately 71% of variance in monthly TB cases, outperforming all other individual models. The exceptional Hit Rate of 97.2% indicates that XGBoost correctly predicted the direction of monthly changes (increase versus decrease) in 35 of 36 validation months\u0026mdash;a critical capability for early warning systems and resource planning.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e2. LightGBM demonstrated competitive performance\u003c/strong\u003e: Achieving R\u0026sup2;=0.698 and Hit Rate\u0026thinsp;=\u0026thinsp;97.2%, LightGBM nearly matched XGBoost while offering computational efficiency advantages. The similarity in performance suggests both gradient boosting frameworks effectively capture TB incidence dynamics.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e3. Hybrid ensemble provided optimal balance\u003c/strong\u003e: The novel 70% XGBoost\u0026thinsp;+\u0026thinsp;30% LightGBM ensemble achieved R\u0026sup2;=0.702, Hit Rate\u0026thinsp;=\u0026thinsp;97.2%, and MAPE\u0026thinsp;=\u0026thinsp;21.8%, effectively combining the strengths of both algorithms. This hybrid outperformed the standard 60:40 XGBoost-Random Forest ensemble.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e4. Deep learning achieved competitive but not superior accuracy\u003c/strong\u003e: The LSTM-CNN hybrid architecture achieved R\u0026sup2;=0.682\u0026ndash;0.691, demonstrating that deep learning can effectively model TB incidence dynamics. However, it did not surpass gradient boosting methods, likely due to: (a) relatively small sample size (206 observations) limiting deep learning\u0026apos;s typical advantages, (b) structured tabular data favoring tree-based methods, and (c) absence of complex spatial or image data where CNN excel.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003e5. Age stratification impacts deep learning more than tree methods\u003c/strong\u003e: LSTM-CNN performance improved with simplified 4-group age stratification (R\u0026sup2; increased from 0.682 to 0.691), while tree-based models showed minimal sensitivity. This suggests overly granular features can introduce noise in neural networks with limited data.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec20\"\u003e\n \u003ch2\u003eLong-Term Forecasts: Projections to 2030 and 2035\u003c/h2\u003e\n \u003cp\u003eUsing the validated hybrid ensemble model (70% XGBoost\u0026thinsp;+\u0026thinsp;30% LightGBM), we generated authoritative forecasts from August 2025 through December 2035. Figure\u0026nbsp;1 presents the complete time series including 17 years of historical data and 10-year projections with 95% prediction intervals.\u003c/p\u003e\n \u003cp\u003e[Figure 1. Historical monthly TB cases in Taiwan (2008\u0026ndash;2025) and hybrid ensemble forecast with 95% prediction interval through 2035. Black line represents actual cases, red line shows forecasted cases, and gray shading indicates 95% confidence band. Vertical dashed line marks forecast origin (August 2025).]\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec21\"\u003e\n \u003ch2\u003eKey projection results:\u003c/h2\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eMonthly average (2026\u0026ndash;2035): 267 cases (95% CI: 205\u0026ndash;330)\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eAnnual average: 3,204 cases\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eTotal predicted cases (August 2025 \u0026ndash; December 2035): 33,642\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eTrend: Minimal decline (-0.3% annually from 2025 to 2035)\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec22\"\u003e\n \u003ch2\u003eAnnual Incidence Projections and WHO Target Comparison\u003c/h2\u003e\n \u003cdiv\u003e\n \u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 2\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003epresents annual forecasts converted to population-based incidence rates using Taiwan\u0026apos;s official population projections, with comparison to WHO End TB Strategy milestones.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eYear\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003ePredicted Cases\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003ePopulation (millions)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eIncidence per 100,000\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e95% CI\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003eWHO Target\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e2023 (actual)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e6,584\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e23.42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e28.1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003e22.5 (2025 milestone)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e2025\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e4,860\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e23.20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e20.9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e17.3\u0026ndash;24.6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e2026\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e3,204\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e23.10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e13.9\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e11.5\u0026ndash;16.3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e2030\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e3,247\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e22.93\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e14.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e12.4\u0026ndash;16.0\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003e9.0 (80% reduction)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e2033\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e3,247\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e22.55\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e14.4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e12.6\u0026ndash;16.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e2035\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\n \u003cp\u003e3,247\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e22.30\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\n \u003cp\u003e14.6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003e12.7\u0026ndash;16.5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c6\"\u003e\n \u003cp\u003e4.5 (90% reduction)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003eTable\u0026nbsp;2. Annual TB forecast with population-based incidence rates (2023\u0026ndash;2035). Incidence calculated as (cases / mid-year population) \u0026times; 100,000. WHO targets based on 2015 baseline of ~\u0026thinsp;45 per 100,000 in Taiwan.\u003c/p\u003e\n \u003cdiv id=\"Sec23\"\u003e\n \u003ch2\u003eAnalysis relative to WHO End TB Strategy targets:\u003c/h2\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003e2025 milestone (\u0026lt;\u0026thinsp;22.5/100k): Taiwan\u0026apos;s projected 20.9/100k MEETS this target, representing 25% reduction from 2015\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003e2030 target (\u0026lt;\u0026thinsp;9.0/100k, 80% reduction): Projected 14.2/100k FALLS SHORT by 58%, indicating need for accelerated interventions\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003e2035 target (\u0026lt;\u0026thinsp;4.5/100k, 90% reduction): Projected 14.6/100k FALLS SHORT by 224%, requiring dramatic intensification to achieve\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eUnder current conditions, Taiwan would need to accelerate decline to 2\u0026ndash;3% annually (versus current\u0026thinsp;~\u0026thinsp;0.3%) to meet 2030 targets\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec24\"\u003e\n \u003ch2\u003eScenario Analysis: Intervention Impact Projections\u003c/h2\u003e\n \u003cdiv\u003e\n \u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 3\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003epresents projected outcomes under four intervention scenarios, quantifying potential impacts of enhanced TB control strategies.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eScenario\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eDescription\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eCases in 2030\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eIncidence in 2030\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c5\"\u003e\n \u003cp\u003eCases in 2035\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eBaseline\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eCurrent measures\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e3,247\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e14.2/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e3,247\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eModerate (10% \u0026darr;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eEnhanced detection\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e2,922\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e12.7/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e2,922\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eIntensive (20% \u0026darr;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eAdd systematic screening\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e2,598\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e11.3/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e2,598\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eComprehensive (30% \u0026darr;)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eScale-up preventive therapy\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\n \u003cp\u003e2,273\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e9.9/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\n \u003cp\u003e2,273\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003eTable\u0026nbsp;3. Projected TB burden under intervention scenarios. Reductions applied multiplicatively to baseline forecast.\u003c/p\u003e\n \u003cp\u003eKey insights: Even the comprehensive 30% reduction scenario (requiring massive scale-up of preventive therapy, active case finding, and enhanced diagnostics) achieves only 9.9/100k in 2035\u0026mdash;still falling short of the 4.5/100k WHO target. This suggests Taiwan would need to combine multiple high-intensity interventions to approach elimination goals.\u003c/p\u003e\n \u003cdiv id=\"Sec25\"\u003e\n \u003ch2\u003eSensitivity Analysis Results\u003c/h2\u003e\n \u003cp\u003eComprehensive sensitivity analyses evaluated forecast robustness across multiple dimensions (Table\u0026nbsp;4).\u003c/p\u003e\n \u003cdiv\u003e\n \u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 4\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003eSensitivity analysis: Impact on projected 2030 incidence rates. All variations produce\u0026thinsp;\u0026lt;\u0026thinsp;4% change, confirming forecast robustness.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eSensitivity Test\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eParameter Variation\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e2030 Incidence\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e% Change from Base\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eBaseline forecast\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e14.2/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e\u0026mdash;\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eLearning rate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e0.01 vs. 0.10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e14.0\u0026ndash;14.5/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e-1.4% to +\u0026thinsp;2.1%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eNumber of estimators\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e200 vs. 500\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e14.1\u0026ndash;14.3/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e-0.7% to +\u0026thinsp;0.7%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eEnsemble weights\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e60:40 vs. 80:20\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e14.0\u0026ndash;14.4/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e-1.4% to +\u0026thinsp;1.4%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eTraining endpoint\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e2024 vs. 2025\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e13.9\u0026ndash;14.5/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e-2.1% to +\u0026thinsp;2.1%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003ePopulation projection\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eHigh vs. low scenarios\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e13.7\u0026ndash;14.7/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003e-3.5% to +\u0026thinsp;3.5%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003eThe 2030 and 2035 forecasts varied by less than \u0026plusmn;\u0026thinsp;4% across all sensitivity tests, demonstrating substantial robustness. Greatest sensitivity was to population projections (\u0026plusmn;\u0026thinsp;3.5%), highlighting the importance of accurate demographic forecasting. Model hyperparameters and ensemble weights showed minimal impact (\u0026lt;\u0026thinsp;2%), indicating stable forecast performance.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec26\"\u003e\n \u003ch2\u003eComparison with Global TB Trends\u003c/h2\u003e\n \u003cp\u003eTo contextualize Taiwan\u0026apos;s forecast, we compared projected trends with global TB incidence patterns reported in WHO\u0026apos;s Global Tuberculosis Report 2024 (Table\u0026nbsp;5).\u003c/p\u003e\n \u003cdiv\u003e\n \u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 5\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003eTaiwan vs. global TB trends (2023\u0026ndash;2030). Global projections from WHO assuming 1\u0026ndash;2% annual decline.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eMetric\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eTaiwan (Projected)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eGlobal (WHO 2024)\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eInterpretation\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eIncidence 2023\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e28.1/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e134/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eTaiwan 4.8\u0026times; lower\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003eIncidence 2030\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e14.2/100k\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e~\u0026thinsp;120/100k (projected)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eTaiwan 8.5\u0026times; lower\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e2015\u0026ndash;2023 decline\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e~\u0026thinsp;30%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e-8.3%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eTaiwan outperforms\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e2023\u0026ndash;2030 decline\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003e50%\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003e~\u0026thinsp;10% (projected)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eTaiwan exceeds global\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" colname=\"c1\"\u003e\n \u003cp\u003e2030 WHO target\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c2\"\u003e\n \u003cp\u003eFalls short\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c3\"\u003e\n \u003cp\u003eFalls short\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\" colname=\"c4\"\u003e\n \u003cp\u003eUniversal challenge\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003eTaiwan maintains substantially lower TB burden than the global average and projects faster decline (50% versus ~\u0026thinsp;10% from 2023\u0026ndash;2030). However, both Taiwan and the global community face significant challenges in meeting WHO End TB Strategy targets, reflecting universal issues including diagnostic gaps (global detection rate: 76%), drug resistance (500,000 cases annually), and funding shortfalls ($22\u0026nbsp;billion needed annually versus ~$5.8\u0026nbsp;billion available).\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e"},{"header":"DISCUSSION","content":"\u003cdiv id=\"Sec28\"\u003e\n \u003ch2\u003ePrincipal Findings and Contributions\u003c/h2\u003e\n \u003cp\u003eThis comprehensive study provides the first systematic comparison of five state-of-the-art machine learning approaches for TB forecasting using 17 years of high-quality surveillance data from Taiwan. Six main findings emerged with important scientific and policy implications:\u003c/p\u003e\n \u003cp\u003e1. Gradient boosting methods (XGBoost, LightGBM) offer optimal performance for operational TB forecasting, balancing accuracy (R\u0026sup2;=0.70\u0026ndash;0.71), directional precision (Hit Rate\u0026thinsp;=\u0026thinsp;97.2%), computational efficiency, and interpretability through feature importance metrics.\u003c/p\u003e\n \u003cp\u003e2. A novel hybrid ensemble combining 70% XGBoost with 30% LightGBM achieved near-optimal performance (R\u0026sup2;=0.702, MAPE\u0026thinsp;=\u0026thinsp;21.8%), demonstrating that strategic combination of complementary boosting frameworks can enhance forecast quality.\u003c/p\u003e\n \u003cp\u003e3. Deep learning approaches (LSTM-CNN) demonstrated competitive performance (R\u0026sup2;=0.68\u0026ndash;0.69) but offered limited additional value over gradient boosting for monthly TB forecasting with current data granularity, likely due to moderate sample size and structured tabular data.\u003c/p\u003e\n \u003cp\u003e4. Feature granularity significantly impacts model performance: While tree-based methods were robust to age group stratification, deep learning models showed 1.3% performance improvement with simplified 4-group versus detailed 7-group age categories, suggesting that excessive feature complexity can hinder neural network optimization with limited data.\u003c/p\u003e\n \u003cp\u003e5. Taiwan\u0026apos;s TB epidemic is well-controlled but projected to stabilize rather than accelerate decline under current conditions. Forecasts of 14.2/100k in 2030 and 14.6/100k in 2035 represent excellent control compared to global averages but fall substantially short of WHO elimination targets.\u003c/p\u003e\n \u003cp\u003e6. Achieving pre-elimination status (\u0026lt;\u0026thinsp;10/100k) by 2030 and elimination (\u0026lt;\u0026thinsp;4.5/100k) by 2035 requires dramatic intervention intensification, with scenario analyses suggesting need for 30%+ case reductions through comprehensive strategies combining preventive therapy scale-up, enhanced active case finding, and improved diagnostics.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec29\"\u003e\n \u003ch2\u003eComparison with Previous Literature\u003c/h2\u003e\n \u003cdiv id=\"Sec30\"\u003e\n \u003ch2\u003eMachine Learning for TB Forecasting\u003c/h2\u003e\n \u003cp\u003ePrevious studies have demonstrated the potential of ML for TB prediction but with important limitations. Cao et al. (2022) used LSTM for TB forecasting in China but relied on shorter time series (5 years) and lacked comparison with tree-based methods [21]. Zhu et al. (2021) compared ARIMA with neural networks for TB forecasting but did not evaluate gradient boosting approaches or use rigorous time-series cross-validation [22]. Wang et al. (2017) combined ARIMA with neural networks but did not explore modern deep learning architectures or ensemble methods [23].\u003c/p\u003e\n \u003cp\u003eOur study substantially advances this literature by: (1) systematically comparing five approaches including three gradient boosting variants, (2) implementing rigorous 36-month expanding-window validation preventing data leakage, (3) evaluating both traditional ML and modern hybrid deep learning on equal footing, (4) introducing novel hybrid ensemble combining XGBoost and LightGBM, and (5) providing validated long-term forecasts directly applicable to policy planning.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec31\"\u003e\n \u003ch2\u003eGradient Boosting versus Deep Learning\u003c/h2\u003e\n \u003cp\u003eThe finding that XGBoost and LightGBM outperform deep learning for monthly TB forecasting aligns with recent evidence suggesting tree-based methods often excel for tabular time-series data with moderate sample sizes [24,25]. Grinsztajn et al. (2022) demonstrated through systematic benchmarking that gradient boosting consistently outperforms neural networks on structured data, particularly when sample size is \u0026lt;\u0026thinsp;100,000 observations [24]. Shwartz-Ziv and Armon (2022) further showed that the inductive bias of tree-based models toward axis-aligned decision boundaries is better suited to tabular data than neural networks\u0026apos; smooth function approximation [25].\u003c/p\u003e\n \u003cp\u003eDeep learning typically requires substantially larger datasets to fully leverage architectural complexity and benefit from learned feature representations. Our findings suggest that for operational TB forecasting with monthly aggregated surveillance data, gradient boosting offers superior accuracy-efficiency tradeoff. However, deep learning may offer advantages in settings with: (1) high-frequency data (daily/weekly) providing larger sample sizes, (2) integration of heterogeneous data sources (e.g., environmental, mobility, social media), or (3) spatial forecasting leveraging geographic relationships.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec32\"\u003e\n \u003ch2\u003ePerformance Benchmarking\u003c/h2\u003e\n \u003cp\u003eOur achieved MAPE of 21.7% (XGBoost) and R\u0026sup2;=0.705 compare favorably with published TB forecasting studies. Literature reports MAPE values ranging from 15% to 40% depending on data granularity and forecast horizon [22,23,26]. The exceptional Hit Rate of 97.2% is particularly noteworthy, as directional accuracy is critical for early warning systems and resource planning. Few studies report directional accuracy despite its policy relevance.\u003c/p\u003e\n \u003cp\u003eThe R\u0026sup2; value of 0.705 indicates our model explains approximately 71% of monthly variation in TB cases\u0026mdash;an excellent result given inherent stochasticity in disease surveillance (reporting delays, weekend effects, holiday patterns), demographic fluctuations, and unmeasured confounders (e.g., changes in active case-finding intensity, diagnostic technology adoption, migration patterns).\u003c/p\u003e\n \u003cdiv id=\"Sec33\"\u003e\n \u003ch2\u003ePublic Health and Policy Implications\u003c/h2\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec34\"\u003e\n \u003ch2\u003eTaiwan\u0026apos;s TB Control: Successes and Persistent Challenges\u003c/h2\u003e\n \u003cp\u003eTaiwan\u0026apos;s projected stable incidence of ~\u0026thinsp;14/100k through 2030\u0026ndash;2035 represents a substantial public health success when contextualized against global TB trends. The country has achieved and maintained low transmission despite significant risk factors including rapid population aging (which increases reactivation risk), high population density in urban areas, and ongoing migration flows from higher-burden neighboring countries. Key strengths underlying this success likely include:\u003c/p\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eUniversal healthcare coverage ensuring diagnostic and treatment access\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eRobust surveillance systems with \u0026gt;\u0026thinsp;95% case notification completeness\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eEstablished TB control infrastructure including dedicated clinics and contact investigation teams\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eHigh treatment success rates (\u0026gt;\u0026thinsp;85% for drug-susceptible TB)\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eDirectly observed therapy (DOT) programs ensuring adherence\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eIntegration of TB services within primary healthcare system\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n \u003cp\u003eHowever, the projected plateau highlights a critical policy inflection point: current interventions have successfully achieved disease control but are insufficient for elimination. Achieving the WHO 2030 target of \u0026lt;\u0026thinsp;9/100,000 would require case reductions of approximately 2\u0026ndash;3% annually\u0026mdash;nearly tenfold acceleration from the projected 0.3% annual decline.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003ch3\u003ePathways to Accelerated Decline: Evidence-Based Interventions\u003c/h3\u003e\n\u003cp\u003eOur scenario analyses demonstrate that even intensive interventions (30% case reduction) fall short of 2035 elimination targets, suggesting need for multi-pronged comprehensive strategy. Evidence-based interventions to accelerate TB decline include:\u003c/p\u003e\n\u003col\u003e\n \u003cli\u003e\n \u003cp\u003eSystematic screening of high-risk populations: Targeting elderly individuals (\u0026ge;\u0026thinsp;65 years, 40% of cases), healthcare workers, migrants from high-burden countries, close contacts of active cases, and persons with diabetes or immunosuppression [27].\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003ePreventive therapy scale-up: Expanding latent TB infection (LTBI) treatment coverage, particularly among elderly persons and those with identified risk factors. Modeling studies suggest preventive therapy could reduce incidence by 10\u0026ndash;20% over 5 years [28].\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eEnhanced diagnostic capabilities: Implementing rapid molecular diagnostics (e.g., GeneXpert) as first-line tests, digital chest X-ray screening with AI-assisted interpretation, and mobile screening units targeting high-risk communities [29].\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eSocial determinants interventions: Addressing upstream factors including housing quality, food security, healthcare access barriers, and social isolation in vulnerable populations including elderly living alone and urban poor [30].\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eActive case-finding strategies: Mobile screening units, community outreach programs, and systematic evaluation of high-risk settings (long-term care facilities, homeless shelters, correctional facilities) [31].\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eIntegration with aging society initiatives: Given Taiwan\u0026apos;s rapidly aging population (22% \u0026ge;65 years by 2025, projected 30% by 2035), integration of TB screening with geriatric health programs could achieve synergies [20].\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003eMathematical modeling studies suggest that combining preventive therapy scale-up with enhanced diagnostics and targeted active case-finding could achieve 3\u0026ndash;5% annual incidence reductions\u0026mdash;sufficient to approach 2030 targets if implemented comprehensively and sustained [32].\u003c/p\u003e\n\u003ch3\u003eOperational Value of Forecasting Systems\u003c/h3\u003e\n\u003cp\u003eThe validated forecasting pipeline developed in this study offers multiple practical applications for TB control programs:\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eResource allocation: Monthly case projections with 95% prediction intervals enable evidence-based planning for clinic capacity, medication stocks (first-line and second-line drugs), laboratory services, and staffing needs.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eEarly warning systems: Real-time comparison of observed cases with forecasted expected values can trigger investigations of potential outbreaks, surveillance quality issues, or intervention impacts. Deviations exceeding prediction intervals warrant immediate epidemiological investigation.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003ePolicy evaluation: Comparing observed post-intervention trends with counterfactual forecasts enables rigorous impact assessment of policy changes (e.g., preventive therapy programs, enhanced screening initiatives).\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eTarget setting and progress monitoring: Long-term projections inform realistic milestone setting and enable transparent tracking of progress toward WHO End TB Strategy goals.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eBudget justification: Quantified projections with uncertainty bounds strengthen evidence base for resource mobilization and budget advocacy.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eScenario planning: Intervention scenario analyses enable policymakers to evaluate cost-effectiveness of alternative strategies and prioritize investments.\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eWe recommend annual forecast updates incorporating newest surveillance data to refine predictions and detect emerging trends. The modular Python pipeline facilitates rapid re-training (\u0026lt;\u0026thinsp;5 minutes on standard hardware) and enables scenario modeling to evaluate potential intervention impacts before implementation.\u003c/p\u003e\n\u003cdiv id=\"Sec37\"\u003e\n \u003ch2\u003eStrengths and Limitations\u003c/h2\u003e\n \u003cp\u003e\u003cstrong\u003eStrengths\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e1. Comprehensive data: 17 years of high-quality, near-complete surveillance data from a national electronic TB register with \u0026gt;\u0026thinsp;95% case notification coverage.\u003c/p\u003e\n \u003cp\u003e2. Rigorous validation: Expanding-window cross-validation with 36 independent evaluation time points provides robust out-of-sample performance estimates respecting temporal dependencies.\u003c/p\u003e\n \u003cp\u003e3. Multiple algorithms: Systematic comparison of five approaches including novel hybrid ensemble enables evidence-based method selection for operational forecasting.\u003c/p\u003e\n \u003cp\u003e4. Feature engineering: Comprehensive strategy incorporating temporal, autoregressive, and demographic features captures multiple dimensions of TB epidemiology.\u003c/p\u003e\n \u003cp\u003e5. Sensitivity analyses: Extensive robustness testing across model specifications, hyperparameters, and population scenarios confirms forecast stability.\u003c/p\u003e\n \u003cp\u003e6. Policy relevance: Long-term forecasts directly inform WHO End TB Strategy planning, with scenario analyses quantifying intervention impacts.\u003c/p\u003e\n \u003cp\u003e7. Reproducibility: Complete pipeline documented with code availability, fixed random seeds, and explicit hyperparameter specification.\u003c/p\u003e\n \u003cp\u003e8. Operational focus: Models designed for practical implementation, balancing accuracy with interpretability, computational feasibility, and ease of updating.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eLimitations\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e1. Ecological forecasting: Models predict population-level incidence without individual-level risk stratification. Complementary case-level models could enhance targeted interventions and risk-based screening.\u003c/p\u003e\n \u003cp\u003e2. Assumption of stability: Long-term forecasts assume no major policy changes, pandemic disruptions, economic crises, or shifts in healthcare access. Scenario modeling partially addresses this but cannot anticipate unforeseen disruptions.\u003c/p\u003e\n \u003cp\u003e3. Limited exogenous variables: Models rely primarily on past case trends and demographic features. Incorporation of socioeconomic indicators (income, education, housing quality), healthcare system metrics (treatment capacity, diagnostic coverage), migration patterns, or climate variables might improve accuracy and policy relevance.\u003c/p\u003e\n \u003cp\u003e4. Temporal resolution: Monthly forecasting captures seasonality but may miss shorter-term outbreaks or transmission clusters. Weekly forecasting would enhance early warning capabilities but requires higher data quality and is more susceptible to reporting artifacts.\u003c/p\u003e\n \u003cp\u003e5. Uncertainty quantification: Prediction intervals estimated using residual standard deviations assume stationary error structure. Probabilistic forecasting approaches (e.g., quantile regression, Bayesian methods, conformal prediction) could provide more sophisticated uncertainty estimates.\u003c/p\u003e\n \u003cp\u003e6. External validity: Findings are most directly applicable to low-incidence settings with high-quality surveillance and universal healthcare. Generalizability to high-burden or resource-limited contexts requires validation in diverse epidemiological and health system settings.\u003c/p\u003e\n \u003cp\u003e7. Model interpretability: While XGBoost provides feature importance metrics and partial dependence plots, complex learned interactions are not fully transparent. Explainable AI methods (SHAP values, LIME) could enhance interpretability for policymakers.\u003c/p\u003e\n \u003cp\u003e8. Drug resistance: Models forecast total TB cases without stratification by drug susceptibility. Separate models for multidrug-resistant (MDR) and extensively drug-resistant (XDR) TB could inform specialized treatment capacity planning.\u003c/p\u003e\n \u003cp\u003e9. Age-period-cohort effects: Simplified age group features may not fully capture birth cohort effects or age-specific incidence trends, potentially limiting long-term forecast accuracy as population demographics shift.\u003c/p\u003e\n \u003cdiv id=\"Sec38\"\u003e\n \u003ch2\u003eFuture Research Directions\u003c/h2\u003e\n \u003cp\u003eThis study opens multiple promising avenues for methodological advancement and expanded applications:\u003c/p\u003e\n \u003cp\u003e1. Real-time forecasting systems: Develop automated pipelines for continuous model updating and real-time prediction, integrated directly with national surveillance databases. Implementation of dashboards for policymakers with automatic anomaly detection.\u003c/p\u003e\n \u003cp\u003e2. Subnational and spatial forecasting: Extend models to city/county level to guide local resource allocation and identify high-burden areas. Incorporate spatial dependencies and migration patterns using graph neural networks or spatial hierarchical models.\u003c/p\u003e\n \u003cp\u003e3. Intervention scenario modeling: Develop integrated epidemiological-economic models to simulate impacts of specific interventions (e.g., preventive therapy scale-up rates, active case-finding strategies) with cost-effectiveness analysis.\u003c/p\u003e\n \u003cp\u003e4. Multi-disease integration: Joint forecasting models for TB, COVID-19, influenza, and other respiratory infections to capture epidemiological interactions, healthcare system constraints, and resource competition.\u003c/p\u003e\n \u003cp\u003e5. Socioeconomic determinants: Incorporate housing quality, food security, healthcare access metrics, education, income, and migration data to improve accuracy and inform upstream social interventions.\u003c/p\u003e\n \u003cp\u003e6. Advanced deep learning: Explore attention mechanisms (transformers), temporal convolutional networks, and neural ODEs that may better capture complex spatiotemporal dynamics. Graph neural networks could model geographic relationships.\u003c/p\u003e\n \u003cp\u003e7. Probabilistic forecasting: Implement Bayesian methods, distributional neural networks, or quantile regression to provide full predictive distributions rather than point estimates and confidence intervals.\u003c/p\u003e\n \u003cp\u003e8. Explainable AI: Systematically apply SHAP values, LIME, attention weights, and counterfactual analysis to enhance model interpretability and build policymaker trust.\u003c/p\u003e\n \u003cp\u003e9. Drug-resistant TB forecasting: Develop separate models for MDR-TB and XDR-TB to inform specialized treatment capacity, second-line drug procurement, and contact investigation strategies.\u003c/p\u003e\n \u003cp\u003e10. Climate and environmental factors: Incorporate temperature, humidity, air quality, and seasonal patterns that may influence TB transmission and reactivation.\u003c/p\u003e\n \u003cp\u003e11. Validation in diverse settings: Multi-country validation studies to assess generalizability and identify setting-specific modifications needed for optimal performance.\u003c/p\u003e\n \u003cp\u003e12. Causal inference methods: Apply difference-in-differences, synthetic control, or interrupted time series to rigorously evaluate intervention impacts using forecasts as counterfactuals.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e"},{"header":"CONCLUSIONS","content":"\u003cp\u003eThis comprehensive study demonstrates that machine learning, particularly gradient boosting methods (XGBoost, LightGBM) and strategic hybrid ensembles, provides accurate, robust, and operationally feasible forecasting of TB incidence in Taiwan. The validated hybrid ensemble model combining 70% XGBoost with 30% LightGBM projects stable TB burden through 2030 and 2035 (14.2\u0026ndash;14.6 per 100,000 population), reflecting successful disease control but insufficient progress toward WHO elimination targets (\u0026lt;\u0026thinsp;9 per 100,000 by 2030, \u0026lt;\u0026thinsp;4.5 by 2035).\u003c/p\u003e \u003cp\u003eAchieving pre-elimination and elimination milestones requires dramatic intervention intensification. Scenario analyses indicate that comprehensive strategies combining expanded preventive therapy, enhanced active case finding, improved diagnostics, and social determinants interventions could reduce incidence by 30%, approaching but still falling short of 2035 elimination goals. This underscores the need for sustained political commitment, adequate resource allocation, and innovative approaches to accelerate progress.\u003c/p\u003e \u003cp\u003eThe forecasting pipeline developed here can be institutionalized for routine surveillance and policy planning, enabling evidence-based resource allocation, early detection of epidemiological changes, and rigorous evaluation of intervention impacts. As countries worldwide pursue TB elimination, robust forecasting systems leveraging modern machine learning will be essential tools for monitoring progress, optimizing interventions, and maintaining political commitment to this achievable public health goal.\u003c/p\u003e \u003cp\u003eTaiwan's experience demonstrates that even well-resourced, low-burden settings with strong healthcare systems face significant challenges in accelerating from disease control to elimination. The forecasting framework and policy insights generated here are directly applicable to similar settings globally and provide a roadmap for evidence-based pursuit of WHO End TB Strategy targets.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eCOMPETING INTERESTS\u003c/h2\u003e \u003cp\u003eThe author declares no financial or non-financial competing interests.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFUNDING\u003c/h2\u003e \u003cp\u003eThis research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. The work was conducted as part of routine surveillance and analytical activities at Taiwan Centers for Disease Control (project code: MOHW114-CDC-C-315-000111).\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eMMK conceived the study, designed the analytical strategy, performed all data analyses, interpreted results, and wrote the manuscript. MMK had full access to all data and takes responsibility for the integrity of the data and the accuracy of the data analysis.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eWe thank the Taiwan Centers for Disease Control surveillance team for maintaining high-quality TB notification data. We acknowledge the healthcare workers, laboratory personnel, and public health nurses whose diligent case reporting and contact investigation enable public health monitoring and response.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eAggregated monthly TB surveillance data and complete analytical code (Python notebooks) are available from the corresponding author upon reasonable request, subject to Taiwan CDC data protection regulations. Individual-level patient data cannot be shared due to privacy regulations. A public dashboard with forecast visualizations is planned for launch in 2026.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eWorld Health Organization. Global Tuberculosis Report 2024. Geneva: World Health Organization; 2024.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHogan AB, Jewell BL, Sherrard-Smith E, et al. Potential impact of the COVID-19 pandemic on HIV, tuberculosis, and malaria in low-income and middle-income countries: a modelling study. Lancet Glob Health. 2020;8(9):e1132\u0026ndash;41.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWorld Health Organization. The End TB Strategy. Geneva: World Health Organization; 2015.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTaiwan Centers for Disease Control. Taiwan Tuberculosis Control Report 2023. Taipei: Taiwan CDC; 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChretien JP, Riley S, George DB. Mathematical modeling of the West Africa Ebola epidemic. eLife. 2015;4:e09186.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePei S, Kandula S, Yang W, Shaman J. Forecasting the spatial transmission of influenza in the United States. Proc Natl Acad Sci USA. 2018;115(11):2752\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang H, Yamamoto N. Using a partial differential equation with Google Mobility data to predict COVID-19 in Arizona. Math Biosci Eng. 2020;17(5):4891\u0026ndash;904.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJewell CP, Kypraios T, Neal P, Roberts GO. Bayesian analysis for emerging infectious diseases. Bayesian Anal. 2009;4(3):465\u0026ndash;96.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJiang F, Jiang Y, Zhi H, et al. Artificial intelligence in healthcare: past, present and future. Stroke Vasc Neurol. 2017;2(4):230\u0026ndash;43.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSantosh KC. AI-driven tools for coronavirus outbreak: Need of active learning and cross-population train/test models. J Med Syst. 2020;44(5):93.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChimmula VKR, Zhang L. Time series forecasting of COVID-19 transmission in Canada using LSTM networks. Chaos Solitons Fractals. 2020;135:109864.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArdabili SF, Mosavi A, Ghamisi P, et al. COVID-19 outbreak prediction with machine learning. Algorithms. 2020;13(10):249.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHochreiter S, Schmidhuber J. Long short-term memory. Neural Comput. 1997;9(8):1735\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKrizhevsky A, Sutskever I, Hinton GE. ImageNet classification with deep convolutional neural networks. Commun ACM. 2017;60(6):84\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBreiman L. Random forests. Mach Learn. 2001;45(1):5\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen T, Guestrin C. XGBoost: A scalable tree boosting system. In: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. 2016:785\u0026ndash;794.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKe G, Meng Q, Finley T et al. NIPS. LightGBM: A highly efficient gradient boosting decision tree. In: Advances in Neural Information Processing Systems 30 (2017). 2017:3146\u0026ndash;3154.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDietterich TG. Ensemble methods in machine learning. Multiple Classifier Systems. Lecture Notes in Computer Science. Volume 1857. Springer; 2000. pp. 1\u0026ndash;15.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBergmeir C, Ben\u0026iacute;tez JM. On the use of cross-validation for time series predictor evaluation. Inf Sci. 2012;191:192\u0026ndash;213.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNational Development Council. Taiwan. Population Projections for the R.O.C. (Taiwan): 2020\u0026ndash;2070. Taipei: National Development Council; 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCao Y, Jiang X, Wang Y, et al. LSTM-based forecasting model for tuberculosis incidence in China. Int J Environ Res Public Health. 2022;19(8):4588.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhu S, Wang Y, Xie C, et al. Comparison of time series models for tuberculosis forecasting in China. Epidemiol Infect. 2021;149:e232.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang KW, Deng C, Li JP, et al. Hybrid methodology for tuberculosis incidence time-series forecasting based on ARIMA and a NAR neural network. Epidemiol Infect. 2017;145(6):1118\u0026ndash;29.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGrinsztajn L, Oyallon E, Varoquaux G. NeurIPS. Why do tree-based models still outperform deep learning on typical tabular data? In: Advances in Neural Information Processing Systems 35 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShwartz-Ziv R, Armon A. Tabular data: Deep learning is not all you need. Inf Fusion. 2022;81:84\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEarnest A, Evans SM, Sampurno F, Millar J. Forecasting annual incidence and mortality rate for prostate cancer in Australia until 2022 using autoregressive integrated moving average (ARIMA) models. BMJ Open. 2019;9(8):e031331.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWorld Health Organization. WHO Consolidated Guidelines on Tuberculosis: Module 2: Screening. Geneva: WHO; 2021.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWorld Health Organization. WHO Consolidated Guidelines on Tuberculosis: Module 1: Prevention. Geneva: WHO; 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQin ZZ, Sander MS, Rai B, et al. Using artificial intelligence to read chest radiographs for tuberculosis detection. Sci Rep. 2019;9(1):15000.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eL\u0026ouml;nnroth K, Jaramillo E, Williams BG, et al. Drivers of tuberculosis epidemics. Soc Sci Med. 2009;68(12):2240\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFox GJ, Barry SE, Britton WJ, Marks GB. Contact investigation for tuberculosis: a systematic review and meta-analysis. Eur Respir J. 2013;41(1):140\u0026ndash;56.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHouben RMGJ, Menzies NA, Sumner T, et al. Feasibility of achieving the 2025 WHO global tuberculosis targets in South Africa, China, and India. Lancet Glob Health. 2016;4(11):e806\u0026ndash;15.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Tuberculosis, Forecasting, Machine Learning, XGBoost, LightGBM, LSTM, Deep Learning, Public Health Surveillance, Taiwan, WHO End TB Strategy","lastPublishedDoi":"10.21203/rs.3.rs-9223330/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9223330/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eTuberculosis (TB) remains a significant public health challenge globally, with 10.8\u0026nbsp;million incident cases in 2023. Accurate forecasting is crucial for resource allocation and evaluating progress toward WHO End TB Strategy targets. This study developed and validated multiple machine learning models to forecast TB cases in Taiwan through 2035.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eWe analyzed 17 years of monthly TB surveillance data (January 2008\u0026ndash;July 2025, n\u0026thinsp;=\u0026thinsp;206 observations) from Taiwan's national electronic TB register. Five modeling approaches were systematically evaluated: Random Forest, XGBoost, LightGBM, ensemble methods (including a novel 70% XGBoost\u0026thinsp;+\u0026thinsp;30% LightGBM hybrid), and hybrid LSTM-CNN deep learning architectures. Models incorporated temporal features, autoregressive lags (1, 3, 6 months), rolling averages, and stratified demographic data (age, gender, migration status). Two age stratification schemes were compared: 7 groups (0\u0026ndash;14, 15\u0026ndash;24, 25\u0026ndash;34, 35\u0026ndash;44, 45\u0026ndash;54, 55\u0026ndash;64, \u0026ge;\u0026thinsp;65 years) versus 4 groups (0\u0026ndash;24, 25\u0026ndash;44, 45\u0026ndash;64, \u0026ge;\u0026thinsp;65 years). Performance was assessed using expanding-window time-series cross-validation over 36 months (August 2022\u0026ndash;July 2025) with metrics including R\u0026sup2;, RMSE, MAPE, and directional accuracy (Hit Rate). Comprehensive sensitivity analyses evaluated forecast robustness. Scenario analyses explored intervention impacts on projected incidence.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eXGBoost with 7 age groups demonstrated superior performance (R\u0026sup2;=0.705, RMSE\u0026thinsp;=\u0026thinsp;60.2, MAPE\u0026thinsp;=\u0026thinsp;21.7%, Hit Rate\u0026thinsp;=\u0026thinsp;97.2%), followed by LightGBM (R\u0026sup2;=0.698, RMSE\u0026thinsp;=\u0026thinsp;61.1, MAPE\u0026thinsp;=\u0026thinsp;22.0%, Hit Rate\u0026thinsp;=\u0026thinsp;97.2%) and ensemble methods (R\u0026sup2;=0.690, RMSE\u0026thinsp;=\u0026thinsp;61.8, MAPE\u0026thinsp;=\u0026thinsp;22.2%, Hit Rate\u0026thinsp;=\u0026thinsp;97.2%). The LSTM-CNN model achieved competitive results with 7 age groups (R\u0026sup2;=0.682, RMSE\u0026thinsp;=\u0026thinsp;63.4, MAPE\u0026thinsp;=\u0026thinsp;22.8%, Hit Rate\u0026thinsp;=\u0026thinsp;94.4%) but performance degraded with simplified 4-group stratification. The hybrid ensemble (70% XGBoost\u0026thinsp;+\u0026thinsp;30% LightGBM) forecasts Taiwan's TB incidence at 14.2 per 100,000 population in 2030 (95% CI: 12.4\u0026ndash;16.0) and 14.6 per 100,000 in 2035 (95% CI: 12.7\u0026ndash;16.5), representing approximately 3,247 annual cases. This reflects a 50% decline from 2023 baseline (28 per 100,000) but falls short of WHO End TB Strategy targets (\u0026lt;\u0026thinsp;9 per 100,000 by 2030, \u0026lt;\u0026thinsp;4.5 per 100,000 by 2035). Scenario analyses indicate that a 30% case reduction through enhanced interventions could achieve 9.9 per 100,000 by 2035. Sensitivity analyses confirmed forecast robustness with \u0026lt;\u0026thinsp;4% variation across model configurations.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e \u003cp\u003eMachine learning approaches, particularly gradient boosting methods (XGBoost, LightGBM) and their hybrids, provide accurate and robust TB forecasting for Taiwan. The projected trajectory suggests successful maintenance of low TB burden but insufficient progress toward elimination goals under current conditions. Achieving WHO 2030 and 2035 targets requires intensified interventions including expanded preventive therapy, enhanced active case finding, and systematic screening of high-risk populations. This validated forecasting pipeline can be institutionalized for routine surveillance, policy planning, and intervention evaluation.\u003c/p\u003e","manuscriptTitle":"Machine Learning-Based Forecasting of Tuberculosis Incidence in Taiwan: A Comprehensive Comparison of Traditional and Deep Learning Approaches with Projections to 2035","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-01 10:10:54","doi":"10.21203/rs.3.rs-9223330/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"dad65eb0-7272-4442-a4cb-bbe60d84c5ae","owner":[],"postedDate":"April 1st, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2026-04-23T03:25:00+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-01 10:10:54","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9223330","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9223330","identity":"rs-9223330","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00