Developing and validating clinical features-based machine learning algorithms to predict influenza infection in influenza-like illness patients.

OA: gold CC-BY-NC-ND-4.0
⚙ AI-generated summary by qwen3.7-flash, 2026-08-21 ⓘ

A prospective cohort study validated an eXtreme Gradient Boosting model using clinical features to predict influenza infection in ILI patients, demonstrating superior performance over conventional models.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

⚙ AI-generated deep summary by qwen3.7-flash, 2026-08-21 · read from full text ⓘ

This prospective cohort study developed and validated machine learning algorithms to predict influenza infection among patients presenting with influenza-like illness in emergency departments across the US and Taiwan. The researchers utilized clinical features such as demographics, vital signs, and symptoms to train seven different ML models, comparing their performance against existing clinical prediction rules using a testing dataset. The results indicated that these objective, data-driven approaches could effectively identify influenza cases, potentially addressing the undertreatment issues associated with lower sensitivity in traditional clinical gestalt or rapid diagnostic tests. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

BackgroundSeasonal influenza poses a significant risk, and patients can benefit from early diagnosis and treatment. However, underdiagnosis and undertreatment remain widespread. We developed and compared clinical feature-based machine learning (ML) algorithms that can accurately predict influenza infection in emergency departments (EDs) among patients with influenza-like illness (ILI).Material and methodsWe conducted a prospective cohort study in five EDs in the US and Taiwan from 2015 to 2020. Adult patients visiting the EDs with symptoms of ILI were recruited and tested by real-time RT-PCR for influenza. We evaluated seven ML algorithms and compared their results with previously developed clinical prediction models.ResultsOut of the 2189 enrolled patients, 1104 tested positive for influenza. The eXtreme Gradient Boosting achieved superior performance with an area under the receiver operating characteristic curve of 0.82 (95% confidence interval [CI] = 0.79-0.85), with a sensitivity of 0.92 (95% CI = 0.88-0.95), specificity of 0.89 (95% CI = 0.86-0.92), and accuracy of 0.72 (95% CI = 0.69-0.76) in the testing set over cut-offs of 0.4, 0.6 and 0.5, respectively. These results were superior to those of previously proposed clinical prediction models. The model interpretation revealed that body temperature, cough, rhinorrhea, and exposure history were positively associated with and the days of illness and influenza vaccine were negatively associated with influenza infection. We also found the week of the influenza season, pulse rate, and oxygen saturation to be associated with influenza infection.ConclusionsThe clinical feature-based ML model outperformed conventional models for predicting influenza infection.
Full text 30,719 characters · extracted from pmc-nxml · 10 sections · click to expand

Goals

The primary aim of this study was to develop and compare clinical feature-based ML algorithms that can predict influenza infection among patients with ILI in EDs. The secondary aim was to compare the performance of clinical feature-based ML algorithms with previously developed clinical prediction models.

Funding

This study was supported by the Ministry of Science and Technology and Chang Gung Memorial Hospital in Taiwan (MOST 109-2327-B-182-002, 109-2314-B-182-036, CMRPG2K0211, CMRPG2K0241 and CMRPG2L0281) and National Institutes of Health in the US (HHSN272201400007C and 75N93021C00045). The funder of this study had no role in study design, data collection, data analysis, data interpretation, writing of the report, or in the decision to submit for publication. The authors had full access to all the data, and the corresponding author was responsible for the.

Methods

We conducted a prospective cohort study in two hospitals in the US and three hospitals in Taiwan at the Centers of Excellence for Influenza Research and Surveillance. The Institutional Review Board of Johns Hopkins University (IRB00135664, IRB00041233, IRB00141101, IRB00052743, and IRB00091667) and the Chang Gung Medical Foundation (201406930B0) approved the study. The hospitals where the patients were enrolled ranged from community hospitals to medical centers in both the suburbs and the metropolitan areas. Appendix 1 provides a detailed description of these hospitals. To report our study, we followed the transparent reporting of a multivariable prediction model for individual prognosis or diagnosis [ 29 ]. Adult patients were presented to the study EDs hospitals between 2015 and 2020 with ILI symptoms modified from the WHO were eligible for this study. ILI was defined as a documented or reported fever and any of the three respiratory symptoms, including cough, headache, or sore throat, within the past seven days before the ED visit [ 30 ]. Patients who could not provide written informed consent, were currently incarcerated, or were previously enrolled in the study during the same influenza season were excluded. After informed consent, patient demographic information, comorbidities, history of exposure to a confirmed influenza infection patient in the past five days, travel, vaccination, and medication history, initial vital signs at triage, date of ED visit, and clinical symptoms were collected by research coordinators as the feature candidates. The coordinators also reviewed the medical history in the electronic medical record and obtained a nasopharyngeal (NP) swab for the influenza PCR test. The PCR was performed using a next-generation, fully automated, and integrated system, Cepheid® Xpert Flu Assay multiplex real-time PCR (Cepheid), with an overall sensitivity and specificity of 98–100% [ 31 ]. Additionally, external features, including the dominant viral subtype of the season, the week of the influenza season, and visits during the influenza season, were also collected for model development. We then randomly partitioned our dataset into a 75% training dataset and a 25% testing dataset, stratified by influenza infection status. The training dataset was used for feature selection and model development. We performed ten-fold cross-validations on the training dataset to fine-tune the hyperparameters. The testing dataset was kept aside for performance validation. To facilitate model training, we normalized the continuous features. To extract important candidate features and remove noisy or redundant features that may result in an inefficient, impractical, or overfitting model, we adopted a wrapper method with the Boruta algorithm to rank the predictive influenza infection features. We repeated the Boruta algorithm 300 times for practical reasons and chose the top 20 ranked features. For model development, we evaluated the seven ML algorithms, including tree-and ensemble-based models: eXtreme Gradient Boosting (XGBoost), conditional random forest (CForest), random forest (RF), and RANdom forest GEneRator (RANGER); a distance-based model: support vector machine (SVM); neural network-based models: artificial neural networks (ANN) and deep learning models with two, three, and four layers. Appendix 2 provides a detailed description and the specific hyperparameters for each algorithm. After model development, the candidate models and the previously developed clinical prediction models developed by Zimmerman et al., Anderson et al., and Dugas et al. were compared using the testing dataset [ 19 , 20 , 32 ]. Performance was evaluated on the area under the receiver-operating characteristic curves (AUROC). We further adopted modifiable cut-offs for the candidate ML algorithm to evaluate corresponding sensitivities and specificities. To interpret the ML models, we used the Shapley additive extension (SHAP) values to illustrate the direction and strength of the selected features in the final model. Additionally, we plotted a calibration curve and estimated the calibration by the Pearson Correlation for goodness-of-fit diagnosis, which enabled us to qualitatively compare the predicted probability of influenza infection to the empirical probability. Demographics and clinical characteristics of recruited patients are presented as mean (standard deviation) or median (interquartile range) for continuous features and counts and percentages for discrete features. ML modeling was performed using R (version 4.0.1, R Foundation for Statistical Computing, Vienna, Austria) with the Boruta [ 33 ] and Caret packages [ 34 ]. Deep learning modeling was performed using Python (Python Software Foundation, Wilmington, DE) with sequential, dense, and dropout packages. Prediction models were compared using R with the DeLong method of the pROC package [ 35 ]. All statistical tests were two-sided, and statistical significance was defined as a p -value of <0.05. We evaluated the final model in the subgroups of different countries (Taiwan and the US) and different subtype-dominant seasons (H1N1 and H3N2). The subtype-dominant season was defined by surveillance reports from the Centers for Disease Control and Prevention in the United States and reports from Taiwan Centers for Disease Control in Taiwan [ 36 , 37 ]. To compare model performance in different subgroups, we applied the following two methods: first, we applied the model developed from the overall training dataset to the testing set of each subgroup; then retrained the model in different subgroups of the training dataset and then evaluated model performance in different subgroups of the testing set.

Results

Among 2189 patients with ILI recruited from 2015 to 2020, 1104 patients tested positive for influenza infection (50%). These patients were over 40 years (median; 50 and IQR: 29–54) and were women (55%, Table 1 and Appendix 3). Patients with influenza infection were older and more likely to have a history of exposure to a confirmed influenza infection patient in the past five days and receive an antiviral agent within 30 days. They were also less likely to receive an influenza vaccine that year and to have taken antibiotics within the past 30 days ( Table 1 ), but were more likely to visit the ED earlier and visit during the influenza season. Patients with influenza infection had higher body temperatures, pulse rates, respiratory rates, and lower oxygen saturation at triage. They had significantly more respiratory symptoms, including a cough, with or without sputum, a sore throat, rhinorrhea, shortness of breath, and wheezing, but less nausea and stomach pain. Table 1 Characteristics of enrolled patients with influenza-like illness. Table 1 Mean/N (SD/%) Overall Influenza Negative Influenza Positive (n = 2189) (n = 1085) (n = 1104) Demographics and related medical history  Age, years (median, (IQR)) 40 (29–54) 40 (29–53) 40.5 (30–55)  Female 1203 (55.0) 610 (56.2) 593 (53.7)  BMI, kg/m 2 27.30 (7.29) 27.33 (7.33) 27.27 (7.26)  Received influenza vaccination of the year 686 (31.3) 385 (35.5) 301 (27.3)  Exposed to influenza a 292 (13.3) 83 (7.6) 209 (18.9)  Travel history within the past 30 days 568 (25.9) 264 (24.3) 304 (27.5)  Taken antibiotics within the past 30 days 525 (24.0) 296 (27.3) 229 (20.7)  Taken influenza antiviral agent within the past 30 days 217 (9.9) 86 (7.9) 131 (11.9)  Day of illness of ILI while visiting the EDs 3 (2–5) 4 (2–5) 3 (2–4)  Visiting during influenza season 1372 (62.7) 633 (58.3) 739 (66.9)  Recruited during H1N1 dominant season 1261 (57.6) 537 (49.5) 724 (65.6)  Recruited in the US 1103 (50.4) 575 (53.0) 528 (47.8) Triage vital sign  Body temperature, °C 37.69 (1.10) 37.49 (1.11) 37.92 (1.05)  Pulse rate, bpm 103.94 (19.06) 102.78 (19.31) 105.25 (18.70)  Respiratory rate, bpm 18.68 (3.38) 18.52 (2.64) 18.87 (4.05)  Oxygen saturation, % 96.64 (2.88) 96.84 (2.80) 96.39 (2.95) Past medical history  Chronic lung disease 572 (26.1) 295 (27.2) 277 (25.1)  HIV 140 (6.4) 71 (6.5) 69 (6.2)  Autoimmune disease 91 (4.2) 51 (4.7) 40 (3.6) Clinical symptoms  Cough 1908 (87.2) 849 (78.2) 1059 (95.9)  Cough with sputum 1410 (64.4) 613 (56.5) 797 (72.2)  Headache 1768 (80.8) 890 (82.0) 878 (79.5)  Sore throat 1462 (66.8) 689 (63.5) 773 (70.0)  Body aches 1816 (83.0) 886 (81.7) 930 (84.2)  Rhinorrhea 1541 (70.4) 696 (64.1) 845 (76.5)  Shortness of breath 1469 (67.1) 698 (64.3) 771 (69.8)  Sinus pain 520 (23.8) 267 (24.6) 253 (22.9)  Wheezing 968 (44.2) 452 (41.7) 516 (46.7)  Fatigue 1943 (88.8) 951 (87.6) 992 (89.9)  Nausea 1110 (50.7) 574 (52.9) 536 (48.6)  Diarrhea 638 (29.1) 332 (30.6) 306 (27.7)  Stomach pain 763 (34.9) 410 (37.8) 353 (32.0) a Been exposed to human with confirmed influenza infection within the past 5 days. Characteristics of enrolled patients with influenza-like illness. Been exposed to human with confirmed influenza infection within the past 5 days. We selected the top 20 most important clinical features from 115 features using the wrapper with 300 iterations ( Fig. 1 and Appendix 4), including patient demographics (body height and weight), past medical history (chronic lung disease and influenza vaccination), influenza infection-related history (travel history, exposed to human with confirmed influenza infection within the past five days, antiviral agent administered within 30 days before the hospital visit, the week of the influenza season, and visit during influenza season), symptoms (day of illness, cough, cough with sputum, rhinorrhea, sore throat, and sinus pain), and signs (body temperature, pulse rate, oxygen saturation, systolic blood pressure, and respiratory rate). Fig. 1 Shapley Additive exPlanations (SHAP) summary plot to explain the feature importance obtained by the XGBoost algorithm. Features with higher feature value (red) and positive SHAP value (on the right side) illustrated in the figure show a positive association, while features with higher feature value (blue) and negative SHAP value (on the left side) show a negative association. Human exposure: State of having been exposed to people with confirmed influenza infection within the past 5 days; Influenza season: Period during which there is a prevalent outbreak of influenza; Influenza vaccine: Received influenza vaccination of the year; Travel history: Travel history within the past 30 days; Influenza antivirals: State of having taken an influenza antiviral agent within the past 30 days. Fig. 1 Shapley Additive exPlanations (SHAP) summary plot to explain the feature importance obtained by the XGBoost algorithm. Features with higher feature value (red) and positive SHAP value (on the right side) illustrated in the figure show a positive association, while features with higher feature value (blue) and negative SHAP value (on the left side) show a negative association. Human exposure: State of having been exposed to people with confirmed influenza infection within the past 5 days; Influenza season: Period during which there is a prevalent outbreak of influenza; Influenza vaccine: Received influenza vaccination of the year; Travel history: Travel history within the past 30 days; Influenza antivirals: State of having taken an influenza antiviral agent within the past 30 days. We then used these clinical features to develop and compare seven ML algorithms ( Table 2 ). The XGBoost model performed the best in the testing set without the overfitting issues noted in other models (accuracy: 0.72 (95% confidence interval [CI] = 0.68–0.76); AUROC, 0.82 (95% CI = 0.79–0.85); final hyperparameter setting is reported in Appendix 2). Among different cut-offs, XGBoost had a sensitivity of 0.92 (95% CI = 0.88–0.96), specificity of 0.89 (95% CI = 0.86–0.92) and accuracy of 0.72 (95% CI = 0.69–0.76) in the testing set over cut-offs of 0.4, 0.6 and 0.5, respectively (Appendix 5). The calibration plot of the XGBoost algorithm revealed a good correlation (Pearson Correlation coefficients = 0.97) between the predicted and observed probabilities of influenza infection in the testing dataset (Appendix 6). Table 2 Performance of machine learning algorithms based on top 20 features. Table 2 Algorithm Dataset Accuracy (95% C.I.) AUROC (95% C.I.) XGBoost Training 0.74 (0.71–0.76) 0.82 (0.80–0.84) Testing 0.72 (0.68–0.76) 0.82 (0.79–0.85) Ranger Training 0.97 (0.96–0.98) 0.99 (0.99–1.00) Testing 0.71 (0.67–0.74) 0.79 (0.75–0.82) Random Forest Training 1.00 (0.99–1.00) 1.00 (1.00–1.00) Testing 0.70 (0.66–0.73) 0.78 (0.75–0.82) Cforest Training 0.83 (0.81–0.85) 0.92 (0.90–0.93) Testing 0.69 (0.66–0.73) 0.77 (0.73–0.80) SVM Training 0.75 (0.72–0.77) 0.84 (0.82–0.86) Testing 0.70 (0.66–0.73) 0.77 (0.73–0.80) Artificial Neural Network Training 0.74 (0.72–0.76) 0.82 (0.80–0.84) Testing 0.70 (0.67–0.74) 0.76 (0.72–0.79) Deep Learning (2-layers) Training 0.70 (0.67–0.72) 0.79 (0.77–0.81) Testing 0.67 (0.63–0.70) 0.75 (0.71–0.79) Deep Learning (3-layers) Training 0.69 (0.67–0.71) 0.76 (0.73–0.78) Testing 0.68 (0.64–0.72) 0.73 (0.69–0.77) Deep Learning (4-layers) Training 0.67 (0.64–0.69) 0.73 (0.71–0.76) Testing 0.65 (0.61–0.68) 0.71 (0.67–0.75) Performance of machine learning algorithms based on top 20 features. The SHAP summary plot in Fig. 1 shows that body temperature, cough, rhinorrhea, and exposure history were highly positively associated with influenza infection in the final XGBoost model (illustrated as higher features value with positive SHAP values). In contrast, the day of illness and a history of influenza vaccination were negatively associated with influenza infection (illustrated as higher feature values with negative SHAP values). In addition, the week of the influenza season, pulse rate, and oxygen saturation were strongly associated with influenza infection. We then compared the performance of the XGBoost model with that of previously proposed models for predicting influenza infection among patients with ILI. Our XGBoost model performed significantly better than the other models proposed by Zimmerman et al., Anderson et al., and Dugas et al. in the testing dataset (AUROC, 0.60, 0.61, and 0.65, respectively; all p values < 0.001; Table 3 and Fig. 2 ). We also found that the model proposed by Zimmerman et al. has the best sensitivity of 0.96 (95% CI = 0.93–0.98) but the worst specificity (0.23 [95% CI = 0.19–0.28]) and AUROC (0.60 [95% CI = 0.56–0.64]). Table 3 Comparison of performance between the XGBoost model and other clinical prediction models. Table 3 Models Dataset Accuracy (95% C.I.) Sensitivity (95% C.I.) Specificity (95% C.I.) AUROC (95% C.I.) XGboost Training 0.74 (0.71–0.76) a 0.91 (0.89–0.93) a 0.88 (0.85–0.90) a 0.82 (0.80–0.84) Testing 0.72 (0.69–0.76) a 0.92 (0.88–0.95) a 0.89 (0.86–0.92) a 0.82 (0.79–0.85) Zimmerman Training 0.59 (0.56–0.61) 0.96 (0.94–0.97) 0.21 (0.18–0.24) 0.59 (0.56–0.62) Testing 0.60 (0.56–0.64) 0.96 (0.93–0.98) 0.23 (0.19–0.28) 0.60 (0.56–0.64)∗∗∗ Anderson Training 0.64 (0.61–0.66) 0.55 (0.51–0.58) 0.72 (0.69–0.75) 0.67 (0.64–0.70) Testing 0.57 (0.53–0.61) 0.52 (0.46–0.58) 0.63 (0.57–0.68) 0.61 (0.56–0.66)∗∗∗ Dugas Training 0.59 (0.56–0.61) 0.84 (0.82–0.87) 0.33 (0.30–0.36) 0.61 (0.55–0.67) Testing 0.59 (0.55–0.64) 0.84 (0.80–0.88) 0.34 (0.28–0.40) 0.65 (0.62–0.68)∗∗∗ ∗∗∗: p -value< 0.001 represents the difference between XGBoost, Zimmerman, Anderson, and Dugas compared to the corresponding testing dataset, respectively. a Accuracy, sensitivity and specificity of the XGBoost model were measured under the cutoffs of 0.5, 0.4 and 0.6, respectively. Fig. 2 Receiver-operating characteristic curves of the XGBoost model and the other clinical prediction models. Fig. 2 Comparison of performance between the XGBoost model and other clinical prediction models. ∗∗∗: p -value< 0.001 represents the difference between XGBoost, Zimmerman, Anderson, and Dugas compared to the corresponding testing dataset, respectively. Accuracy, sensitivity and specificity of the XGBoost model were measured under the cutoffs of 0.5, 0.4 and 0.6, respectively. Receiver-operating characteristic curves of the XGBoost model and the other clinical prediction models. Applying the developed model from the overall training dataset to each subgroup, we found that the XGBoost model performed better in the Taiwan subgroup than in the US (AUROC, 0.78 (95% CI = 0.72–0.84) vs. 0.72 (95% CI = 0.59–0.85), p  = 0.004) but similarly between H1N1 and H3N2 dominant seasons (0.84 [95% CI = 0.79–0.90] vs. 0.81 (95% CI = 0.75–0.87), p  = 0.062). Additionally, we retrained the model separately for different subgroups. The performance of the XGBoost models improved slightly in the US subgroup (0.83 [95% CI = 0.79–0.88]) but not in the Taiwan subgroup and different subtype-dominant seasons (Appendix 7). We also developed an applet that could be easily linked with electrical medical records to enhance clinical utility, available at the following hyperlink ( https://cgmher.shinyapps.io/shinyapp_for_flu_prediction/ , Appendix 8).

Background

Seasonal influenza poses a significant risk. In 2017, over 300,000 deaths resulted from influenza-associated respiratory diseases, and nearly 10 million patients were hospitalized because of influenza-associated lower respiratory tract infections worldwide [ 1 , 2 ]. Influenza was also the principal infectious disease in Europe between 2009 and 2013 [ 3 ]. Furthermore, influenza poses an even more significant threat to low-income countries in sub-Saharan Africa, western Pacific, and Southeast Asia [ 2 ]. The CDC recommends early antiviral medication, especially for patients with complications or those requiring hospitalization for influenza infection [ 4 ]. Early (≤48 h after illness onset) antiviral treatment should also be considered for otherwise healthy symptomatic outpatients to reduce the duration of symptoms. Timely diagnosis and early administration of antiviral agents decrease the illness duration, viral transmission, symptom severity, hospitalization risk, antibiotic usage, and mortality rate [ 5 , 6 ]. However, according to a 2015 study, only 29% of the laboratory-confirmed influenza infection patients who were hospitalized or presented to emergency departments (EDs) were clinically diagnosed with fever and respiratory symptoms [ [7] , [8] , [9] , [10] ]. Another study published in 2019 found that over 60% of influenza-like illness (ILI) patients who were presented to primary care settings with posthoc laboratory-confirmed influenza infection did not initially receive an antiviral agent [ 6 ]. This discrepancy is even more widespread in high-risk groups that would benefit significantly from prompt treatment [ 5 , 6 , 11 ]. Therefore, influenza's undertreatment is common in EDs and primary care settings.

Discussion

This is the first prospective binational multicenter study to use clinical feature-based ML algorithms to predict influenza infection in patients with ILI. The XGBoost model outperformed the other seven ML algorithms and three previously developed clinical prediction models with an AUROC of 0.82 (95% CI = 0.77–0.87) and achieved a sensitivity of 0.92 (95% CI = 0.88–0.95) and specificity of 0.89 (95% CI = 0.86–0.92) with different cut-offs. Body temperature, cough, rhinorrhea, and exposure history were positively associated with influenza infection, whereas days of illness and history of influenza vaccine were negatively associated with influenza infection. Furthermore, the infection was found to be strongly associated with the week of the influenza season, pulse rate, and oxygen saturation. The major strength of our study is that we conducted a prospective multicenter study with a comprehensive collection of data from patients with ILI in front-line health care settings. The first crucial challenge in developing machine learning models is assembling a representative and diverse dataset. Previous efforts that relied on retrospective extraction from medical records may have missed essential features that were not routinely recorded, lowering the model's quality [ 44 ]. In our study, well-trained research coordinators used a predefined and comprehensive questionnaire for every enrolled patient to minimize potential recall bias and maximize model reliability. Another advantage is that it spanned four influenza seasons in multiple scaled EDs in two countries. This reduced the potential spectrum bias related to different age distributions, clinical presentations, and time distributions between viral subtypes [ 45 , 46 ]. We also compared state-of-the-art ML algorithms with available clinical prediction models designed to predict influenza infection, eliminating the need for a historical control that may represent an invalid comparison. Two recent studies attempted to predict influenza infections using the classification and regression tree methods with the rpart and ptree package in R or the Classification and Regression Trees (CART) software; however, their results were suboptimal, and they did not provide detail of the hyperparameter fine-tuning process. The first study conducted in 2016 by Zimmerman et al. included three features (fever, cough, and fatigue) in their prediction model and found a sensitivity of 84%, specificity of 49%, and AUROC of 0.69 [ 20 ]. Anderson et al. (2018) incorporated four features, including cough, rhinorrhea, chills, and body aches, to achieve a sensitivity of 52.1%, specificity of 82.9%, and AUROC of 0.689 [ 19 ]. However, in our testing data set, our XGBoost significantly outperformed these two models (AUROC, 0.82, 0.60, and 0.61, respectively; both p  < 0.001). In contrast, we found that fatigue, chills, and body aches were similar between patients with ILI, and with and without influenza infection. We hypothesized that the subjective definitions of these symptoms might contribute to the limited results of these studies. However, the usefulness of external features and the advantage of our XGBoost model in dealing with complex cases could be attributed to improved performance. In agreement with previous studies, our findings showed that cough, rhinorrhea, and sore throat were positively associated with influenza infection. This supports the current understanding of cough as an important feature in influenza patients [ 30 ]. The influenza virus is thought to infect epithelial cells in the airway, causing inflammation and cytokine release from the host immune system. Furthermore, sore throat was previously included in the WHO ILI definition but was removed in 2011 owing to conflicting evidence [ [47] , [48] , [49] ]. A sore throat in young children is challenging to diagnose [ 11 ]. However, we demonstrated that sore throat has a significant predictive value, implying that further revision of the definition based on age groups is warranted and could facilitate a more prompt and precise prediction of influenza infection [ 26 , 27 ]. We also evaluated another clinical prediction model developed by Dugas et al., composed of only three clinical symptoms, including fever, cough, and headache [ 32 ]. Similarly to the original study, our findings indicated that sensitivity was high (84%) but at the expense of low specificity (34%). This clinical prediction model was developed using data from a single influenza season and focused solely on high-risk patients, which could explain its low specificity. Furthermore, the presence or absence of a headache shows questionable value in distinguishing adult patients with influenza and other pathogens [ 32 ]. Our data revealed that headaches did not differ significantly between ILI patients with and without influenza infection. Another explanation is that simple prediction models frequently disregard more complex interactions between clinical features. The ability of ML algorithms to handle high–order interactions and nonlinear relationships is a significant advantage that can enhance prediction models. Similar studies used natural language processing (NLP) and free text reports to detect influenza, identical to our effort to use clinical features to predict influenza infection. Pineda and Tsui demonstrated, in two retrospective studies, respectively, that NLP and ML algorithms could be used to extract features from electronic health records to detect influenza infection [ 26 , 27 ]. However, the model developed by Pineda et al. was trained in a single health system without external validation. Furthermore, the two studies used non-ILI patients as control groups, which may have overestimated model performance [ 28 ]. In contrast to previous clinical prediction models that relied solely on patient signs and symptoms, we discovered that external features such as the week of the influenza season, the season of the year, and the day of illness have a significant role to play in our XGBoost model. These factors are not intuitive, and their nonlinear relationship with other features makes embedding them into conventional prediction models difficult. However, modern ML algorithms and text-mining techniques make this process possible and therefore warrant further investigation. Given the current availability of rapid influenza diagnostic tools such as point-of-care testing and RT-PCR, the need for a model to effectively predict influenza infection arises. However, clinical gestalt is still required for a sustainable health care environment to increase pre-test probability and direct testing toward specific pathogens [ 17 ]. Our model only included clinical and external features that could be obtained easily in front-line health care settings and used in resource-limited regions and pandemic preparedness in the future. In conclusion, we created and validated a clinical feature-based ML algorithm to predict influenza infections among ILI patients in the ED. We demonstrated that the XGBoost model outperformed the other seven selected ML algorithms and surpassed previously designed models. We created an applet that can be easily linked with electronic medical records to improve clinical utility based on our findings. However, in the new era of COVID-19, our model needs further validation and modification, which can serve as future research directions.

Importance

Diagnosing influenza is clinically challenging. The general symptoms often overlap with other respiratory virus infections. The presentations vary among patients of different age groups and in various clinical settings [ 12 , 13 ]. Therefore, many clinicians rely on the rapid influenza diagnostic test (RIDT) to screen for potential influenza infection in EDs and other primary care settings [ 14 , 15 ]. However, because its low sensitivity, RIDT is not recommended in the updated CDC guidelines anymore [ 16 ]. Although other laboratory methods with higher accuracy, namely RT-PCR assays, rapid molecular assays other nucleic acid amplification tests, and viral cultures, are available, clinical gestalt is still required to direct the diagnostic pathway for a sustainable health care environment [ 17 ]. However, previous studies demonstrated that sole reliance on clinical diagnosis had poor sensitivity [ [18] , [19] , [20] ]. Clinical prediction models that incorporate several different symptoms may be helpful but are still insufficient with a wide range of sensitivities (36–80%) and specificities (78–98%) [ [18] , [19] , [20] ]. Therefore, researchers are turning to machine learning (ML) algorithms, which are objective and have replicable approaches, for integrating multiple features to improve diagnostic ability in conditions such as sepsis, urinary tract infection, acute myocardial infarction, and congestive heart failure [ [21] , [22] , [23] , [24] , [25] ]. Previous studies have shown that ML algorithms outperform expert-built classifiers in predicting the influenza infection [ 26 ]. However, these studies were either conducted retrospectively or only used restrictive feature sets with no reliable validations [ 19 , 20 , [26] , [27] , [28] ].

Limitations

Here are some limitations of our research. First, the model performance based on the AUROC differed between the two subgroups of recruiting countries. The differences in the retrained models for these subgroups remained. We hypothesized that the disparity was because of different dominant season-specific influenza viral subtypes or different respiratory pathogen compositions between the enrolled sites [ [38] , [39] , [40] ]. However, the subgroup analysis of varying subtype-dominant seasons was almost similar. Models must be retrained or updated in different subgroups with larger sample sizes. Second, previous research has shown that up to 40% of influenza infection patients, particularly the elderly, may be afebrile [ 41 , 42 ]. The inclusion criteria derived from the modified WHO ILI definition would invariably result in a different study population. We, therefore, caution readers about selecting patient groups before applying our findings. Third, when it was compared with previously proposed clinical prediction models, our model used up to 20 features for prediction. The more information requires manual data entry hampers the clinical practicality. By increasing global usage of electronic medical records and their integration with a diagnostic decision support system (DDSS), the obstacles to clinical practicality can be narrowed down, but further investigations are needed [ 43 ]. Finally, we conducted the surveillance before the COVID-19 outbreak, which may impact applicability and should be used cautiously. The performance of the clinical features warrants further investigation in the new era of COVID-19.

Coi Statement

There are no conflicts of interest to declare.

Data Availability

The datasets generated and analyzed during the current study are available at https://github.com/wujinja-cgu/FLu_ML_prediction .

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-09-27T09:11:36.575535+00:00
unpaywall
last seen: 2026-08-12T06:43:03.944938+00:00
License: CC-BY-NC-ND-4.0