Staged Identification of CAP in Fever Patients Across Epidemic Environments: Modeling &Validation

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background Diagnosing community-acquired pneumonia (CAP) relies on costly imaging, posing challenges in resource-limited settings. Traditional tools focus on diagnostic tests for clinicians rather than patient use. Additionally, classification of subtypes in traditional Chinese medicine (TCM) lacks criteria. Materials and Methods We developed a multimodal fusion model using machine learning algorithms and clinical variables from basic information, medical records, and lab tests to assess CAP risk in fever patients. The model integrates top-performing models via ensemble learning to predict pneumonia probability. We trained on 2,193 visits at Beijing Traditional Chinese Medicine Hospital’s fever clinic from Dec 2021 to Dec 2022, and validated on 300 visits from Jan to July 2024. Use unsupervised learning to classify subtypes. Results The training cohort included 1,781 CAP and similar patients, with 210 in the external validation cohort. CAPs were diagnosed via chest CT. The α model, based on pre-visit medical records, performed well (AUC internal =0.80, 95%CI 0.77–0.83; AUC external =0.80, 95%CI 0.71–0.87). The β model added four lab indicators, optimizing performance (AUC internal =0.93, 95%CI 0.92–0.95; AUC external =0.81, 95%CI 0.70–0.90). Two models were developed into online calculators. Latent class analysis distinguished Cold/Heat syndrome as subtypes. Conclusion Two models performed good across epidemic environments. We provided low-cost, and accurate tools for staged identification of CAP.
Full text 174,351 characters · extracted from preprint-html · click to expand
Staged Identification of CAP in Fever Patients Across Epidemic Environments: Modeling &Validation | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Staged Identification of CAP in Fever Patients Across Epidemic Environments: Modeling &Validation Gao Ziheng, Chen Tengfei, Ha Yanxiang, Shi Yifan, Xu Xiaolong, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6768762/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 18 Dec, 2025 Read the published version in Scientific Reports → Version 1 posted 10 You are reading this latest preprint version Abstract Background Diagnosing community-acquired pneumonia (CAP) relies on costly imaging, posing challenges in resource-limited settings. Traditional tools focus on diagnostic tests for clinicians rather than patient use. Additionally, classification of subtypes in traditional Chinese medicine (TCM) lacks criteria. Materials and Methods We developed a multimodal fusion model using machine learning algorithms and clinical variables from basic information, medical records, and lab tests to assess CAP risk in fever patients. The model integrates top-performing models via ensemble learning to predict pneumonia probability. We trained on 2,193 visits at Beijing Traditional Chinese Medicine Hospital’s fever clinic from Dec 2021 to Dec 2022, and validated on 300 visits from Jan to July 2024. Use unsupervised learning to classify subtypes. Results The training cohort included 1,781 CAP and similar patients, with 210 in the external validation cohort. CAPs were diagnosed via chest CT. The α model, based on pre-visit medical records, performed well (AUC internal =0.80, 95%CI 0.77–0.83; AUC external =0.80, 95%CI 0.71–0.87). The β model added four lab indicators, optimizing performance (AUC internal =0.93, 95%CI 0.92–0.95; AUC external =0.81, 95%CI 0.70–0.90). Two models were developed into online calculators. Latent class analysis distinguished Cold/Heat syndrome as subtypes. Conclusion Two models performed good across epidemic environments. We provided low-cost, and accurate tools for staged identification of CAP. Health sciences/Medical research/Experimental models of disease Health sciences/Risk factors Health sciences/Diseases/Infectious diseases/Influenza virus Health sciences/Diseases/Respiratory tract diseases Community-Acquired Pneumonia Machine Learning Risk Prediction Model Traditional Chinese Medicine Epidemic Environment Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction Community-acquired pneumonia (CAP) is a common and frequently occurring respiratory infectious disease with high incidence and mortality rates[ 1 ][ 2 ]. Early identification of CAP is challenging, as its confirmation relies on expensive imaging equipment, such as chest CT[ 3 ][ 4 ]. But even X-ray examinations can only be interpreted by medical professionals. During sudden outbreaks, when medical resources are strained, medical congestion is particularly prominent in primary care clinics, emergency departments, and fever clinics. Moreover, pneumonia in the elderly population is clinically insidious and more difficult to identify early, leading to delayed treatment[ 5 ][ 6 ]. The Pneumonia Severity Index (PSI) is a commonly used tool in clinical practice for clinicians and researchers to assess the severity of illness in hospitalized patients[ 7 ][ 8 ]. It includes many precise ancillary tests, making it difficult for patients to use and challenging to implement in rudimentary settings. Outpatients also struggle to complete all the required ancillary tests to calculate the score. Other scoring systems like CURB-65, A-DROP, and SMART-COP can rapidly assess pneumonia severity using a limited number of variables[ 8 ][ 9 ]. However, these assessments rely on significantly abnormal vital signssuch as respiratory rate, level of consciousness, blood pressure, and heart rateand specific ancillary tests (such as blood gas analysis, albumin, or blood urea nitrogen testing). They are primarily designed to quickly identify critically ill patients and provide emergency care, rather than to differentiate between pneumonia and non-pneumonia patients. Existing pneumonia prediction models often incorporate a wide range of clinical characteristic variables[ 10 ][ 11 ], including high-cost chest CT imaging and specific biomarkers (such as IL-6, IL-10), aiming to accurately identify CAP patients, improve detection accuracy, and precisely identify pathogens. However, compared to further improving diagnostic accuracy, the ability to quickly, accurately, and cost-effectively identify pneumonia patients and enable stratified diagnosis and treatment during sudden outbreak responses is equally of great clinical value. Additionally, the TCM treatment of CAP primarily focuses more on classifying syndromes by clinical symptoms than emphasizing etiology, a feature that has proven uniquely advantageous in the COVID-19 pandemic response[ 12 ][ 13 ]. Clinicians empirically classify CAP into Cold/Heat syndrome and select corresponding TCM interventions. Current TCM guideline only mentions the coexistence of “External Cold”“Internal Heat” without differentiation[ 14 ]. Therefore, the clinical classification of the two syndromes still lacks a standard. During the COVID-19 pandemic, China implemented long-term and strict prevention and control policies from December 2019 to December 2022, creating a unique epidemiological environment. In this specialized social environment that effectively blocked the spread of the coronavirus, influenza viruses and other pathogens became the mainstream of respiratory viral infections[ 15 ][ 16 ]. After the end of these policies in December 2022, the Omicron variant of the coronavirus, rather than the Alpha or Delta variants, became widely prevalent. It has a lower proportion of severe cases and mortality rate, but stronger transmissibility and immune escape ability. The prevalent respiratory pathogens gradually shifted to mainly include the coronavirus, influenza virus, rhinovirus, and others[ 17 ][ 18 ]. It has become a new challenge for models to accurately identify pneumonia patients in different epidemiological environments,and even in the presence of new pathogens. This study develops and validates a step-by-step CAP risk prediction model using multimodal data and machine learning across different epidemic settings. It aims to provide an easy-to-use, sensitive, and robust tool for early CAP detection. Additionally, it classifies CAP subtypes based on clinical symptoms to support TCM standardization. Materials and Methods Materials Data Source The data used in this study were derived from two temporally and spatially independent cohorts of patients who visited the fever clinic of Beijing Hospital of Traditional Chinese Medicine (BHTCM), representing different epidemic environments. All data were anonymously extracted from the hospital information management system. From December 2021 to December 2022, 2,193 visits were selected from a total of 7,203 visits throughout the year based on disease diagnosis. After inclusion and exclusion criteria were applied, these visits formed the training cohort of 1,781 patients used for model development (also used for internal validation). From January to July 2024, 300 visits were randomly selected from a total of 2,771 visits over the six-month period (without disease diagnosis selection) to form the external validation cohort of 210 patients. The data from the two cohorts were kept separate, and the external validation cohort was used only after the model was completed (Fig. 1 ). Inclusion Criteria We included patients with community-acquired pneumonia (CAP) and similar conditions (age over 18). The diagnoses of similar conditions needed to be related to respiratory tract infections, including acute upper respiratory tract infection, influenza, common cold, acute tracheitis/bronchitis, acute tonsillitis, acute pharyngitis, etc. All disease diagnoses were obtained from the outpatient information system. The clinical diagnosis and the CAP diagnostic criteria referenced in this study were as follows[ 3 ][ 4 ]: I. Onset in the community. II. Presence of any one of the following manifestations related to pneumonia: a. New onset of cough, sputum production, or worsening of pre-existing respiratory symptoms, with or without purulent sputum, chest pain, dyspnea, and hemoptysis; b. Fever; c. Physical signs of lung consolidation and/or auscultation of moist rales; d. Peripheral blood leukocyte count > 10 × 10^9/L or < 4 × 10^9/L, with or without left shift. III. Chest imaging showing new pulmonary infiltrates, lobar or segmental consolidation, ground-glass opacity, or interstitial changes, with or without pleural effusion. Exclusion Criteria All patients included in the study were assessed against the following exclusion criteria: I. Exclusion of follow-up medical records, retaining only the records of patients' initial visits; II. Exclusion of patients diagnosed with pneumonia who lack chest imaging evidence; III. Exclusion of patients with conditions similar to severe CAP, such as pulmonary edema, diffuse alveolar hemorrhage, interstitial lung disease, etc.; IV. Exclusion of patients who were transferred to another hospital or died on the day of their visit; V. Exclusion of patients with missing important clinical variables (basic information, Tmax, number of days since onset of illness, disease diagnosis, etc.). Methods Disease Diagnostic Criteria Community-acquired pneumonia (CAP) was diagnosed according to the latest relevant guideline criteria, and all patients with pneumonia in the study were required to have chest CT imaging to support the diagnosis. For practicality, we established the training cohort from patients with CAP and similar conditions rather than all patients who visited the clinic. Due to the different diagnostic perspectives, conditions such as "pulmonary infection" "pneumonia" "pneumonia caused by known viruses" "viral pneumonia" "bacterial pneumonia" and "fungal pneumonia" were considered clinically synonymous with CAP. Non-pneumonia patients were diagnosed with conditions such as "acute upper respiratory tract infection" "influenza caused by known viruses" "common cold" and "acute tonsillitis", all of which were determined based on the standard ICD-10 disease names. Because of the differences in the epidemic environment, there were no patients with COVID-19 infection in the training cohort, and therefore no patients with CAP caused by COVID-19. Thus, diagnoses as "COVID-19 infection" and "COVID-19 pneumonia" were only present in the external validation cohort. Data Collection & Management All patient data were anonymously extracted from the hospital information system,including basic information, medical records, laboratory test results, and chest CT examination results. Basic information included sex age, and date of visit, while medical records included chief complaints, medical history, past medical history, physical examination, and disease diagnosis. All patient data were anonymized (identifiable patient information was removed) and aggregated into an Excel database. Modeling & Validation In predictor selection, three strategies were compared comprehensively: I. Directly using all extracted clinical feature variables; II. Using only variables with significant differences between the pneumonia and non-pneumonia groups; III. Using only important variables selected by Lasso regularization regression. Randomforest was used to model the training cohort, comparing the variable importance rankings among the three strategies and comparing the AUC values of the three models using 10-fold cross-validation. The strategy with the fewest variables was chosen as the predictor set. The same method was used to compare the variable importance rankings with interaction terms added, and the gain and necessity of adding interaction terms were assessed. For modeling algorithms, traditional Logistic regression and five common machine learning algorithms were used (Logisticnet, Randomforest, XGBoost, AdaBoost, and CatBoost). To fully utilize the data, instead of using traditional random splitting methods(e.g., 7:3 training and validation set split), 10-fold cross-validation was conducted on the complete training set to obtain the best models and parameters for the six different algorithms. The six models were then internally validated and compared in the entire patient population. The temporally and spatially independent external validation set was used to evaluate the model performance in real clinical settings. The best-performing model among the six was determined as the final α model,with all predictors obtainable by patients before their visit. The Shapley Additive Explanations (SHAP) method was used to assess the impact of original variables on the output. For multimodal clinical data, a fusion modeling strategy was adopted by adding laboratory test-related variables to develop the β model. The β model was built on the probability obtained from the α model and included additional variables such as the neutrophil-to-lymphocyte ratio (NLR = NEU/LYM) and the C-reactive protein-to-platelet count ratio (CRP/PLT). These variables can be obtained through routine blood tests. The same six algorithms were used with 10-fold cross-validation for parameter tuning,modeling,and internal and external validation, with the best-performing β model selected. Both α/β models were developed into Shiny online web calculators for patients to assess pneumonia risk probability at different stages of their visit. Statistical Analysis Methods This predictive modeling study employed a cross-sectional research design. Continuous variables were presented as medians and interquartile ranges, while categorical variables were described using counts and percentages. Normally distributed variables were compared using the two-sample t-test, non-normally distributed variables were compared using the Mann-Whitney U test, and categorical variables were compared using chi-square tests, Fisher's exact tests, or Kruskal-Wallis tests as appropriate. Missing values (with the proportion of missing values in each variable being less than 10%) were handled using multiple imputation by chained equations based on the random forest algorithm, with 10 imputations performed and the results combined according to Rubin's rules. During model training, 10-fold cross-validation was utilized. The entire training set was randomly divided into 10 equal subsets, with one subset (10%) serving as the validation set and the remaining nine subsets (90%) as the training set for each iteration. This process was repeated 10 times, and the parameters that performed best in the validation set were selected as the optimal parameters for the model based on that algorithm.All p-values and confidence intervals were derived from two-sided tests, with a significance level set at 0.05. Bonferroni correction was applied to multiple testing to control the family-wise error rate (FWER) by using a more stringent significance level. Data cleaning, organization, and statistical analysis were conducted using RStudio (R 4.4.3) and Excel software. All R packages used have been described, and the main R codes had been provided also ( Supplement ). Methods’ Technical Review Our study was based on past cross-sectional data from the fever clinic. We downloaded all the anonymized data from the Beijing Hospital of Traditional Chinese Medicine’s YiduCloud HIS database after obtaining approval from the hospital’s clinical research department, which reviewed our protocol and waived informed consent from the original patients. Results are reported according to the TRIPOD statement ( Supplement Table 2 ). Research Support and Funding This study was funded by the National Natural Science Foundation of China (NNSFC 81774146) through the Beijing Hospital of Traditional Chinese Medicine. The sponsor had no role in the study. As its retrospective, anonymized, and non-intervention nature, the hospital’s ethics committee waived ethical approval during the model’s development, following the Declaration of Helsinki. Results Patient Clinical Characteristics We compared the clinical characteristics of patients with pneumonia and those without pneumonia in the training cohort (N = 362 vs.N = 1419) and calculated the odds ratios (OR) (Table 1 ). Similar comparison also conducted in the external validation cohort after modeling ( Supplementary Table 1 ). Among the 33 variables derived from general information and medical records, seven variables—age, Tmax, days (since onset of illness), presence or worsening of pharyngeal discomfort, presence or worsening of cough, presence or worsening of dyspnea, and presence or worsening of altered mental status—remained significantly different between the groups after adjustment (all P < 0.001). The presence or worsening of headache and abdominal pain were only significantly different before adjustment (OR hb_pain =0.49, 95%CI 0.29–0.78, P = 0.002; OR ab_pain =3.41, 95%CI 1.07–10.5, P = 0.039). Regarding underlying diseases, we examined the differences between groups for cardiovascular disease, respiratory disease, neurological disease, kidney disease, blood disorders, and oncological/immune diseases. Only cardiovascular disease and respiratory disease showed significant differences before adjustment (OR card =1.38, 95%CI 1.04–1.82, P = 0.025 ;OR resp =1.83, 95%CI 1.15–2.83, P = 0.011), but these differences were not significant after adjustment. In terms of medication history, we compared the use of medications since the onset of illness. There were no significant differences in the use of antipyretics, antibiotics, or antiviral drugs. For traditional Chinese medicine (TCM), the use of TCM_XFJB was only significantly different before adjustment (OR tcm_xfjb =0.74, 95%CI 0.57–0.95, P = 0.019). In laboratory tests, we extracted six key clinical variables related to infection: neutrophil count (NEU), lymphocyte count (LYM), platelet count (PLT), C-reactive protein (CRP), neutrophil-to-lymphocyte ratio (NLR), and C-reactive protein-to-platelet count ratio (CRP/PLT). All these variables showed significant differences between groups after adjustment (all P < 0.001). In the hypothesis testing above, the Bonferroni correction was used to adjust the significance level to 0.05/33 + 6 ≈ 0.001. Among the three predictor selection strategies for the random forest model, the third strategy performed the best (AUC I =0.78; AUC II =0.78; AUC III =0.80). After comparing the variable importance rankings, the final α model used only the seven variables that remained significantly different between groups after multiple testing as predictors (Fig. 2 ). Table 1 Comparison of Characteristic between Pneumonia and Non-Pneumonia in the Training Cohort Characteristic Non-pneumonia Pneumonia Odds Ratio 95%CI P value (N = 1419) (N = 362) Basic Profile Sex Female = 0 713 (50.2%) 172 (47.5%) 1.12 [0.89–1.41] 0.354 Male = 1 706 (49.8%) 190 (52.5%) Age(y) 39.2 (17.2) 62.2 (22.2) 1.06 [1.05–1.06] < 0.001 *** Days(d) 2.58 (4.05) 3.50 (4.49) 1.04 [1.02–1.07] < 0.001 *** Tmax(℃) 38.1 (0.64) 38.3 (0.71) 1.47 [1.24–1.75] < 0.001 *** Past Medication History Medication history No = 0 430 (30.3%) 112 (30.9%) 0.97 [0.76–1.25] 0.864 Yes = 1 989 (69.7%) 250 (69.1%) Used antipyretics No = 0 980 (69.1%) 263 (72.7%) 0.84 [0.65–1.08] 0.206 Yes = 1 439 (30.9%) 99 (27.3%) Used antibiotics No = 0 1109 (78.2%) 287 (79.3%) 0.94 [0.70–1.24] 0.694 Yes = 1 310 (21.8%) 75 (20.7%) Used antivirals No = 0 1416 (99.8%) 360 (99.4%) 2.68 [0.31–17.7] 0.269 Yes = 1 3 (0.21%) 2 (0.55%) TCM_XFJB No = 0 923 (65.0%) 259 (71.5%) 0.74 [0.57–0.95] 0.023 * Yes = 1 496 (35.0%) 103 (28.5%) TCM_XRJD No = 0 1161 (81.8%) 303 (83.7%) 0.88 [0.64–1.19] 0.448 Yes = 1 258 (18.2%) 59 (16.3%) TCM_WLSH No = 0 1418 (99.9%) 362 (100%) 1 1 Yes = 1 1 (0.07%) 0 (0.00%) Clinical Manifestation Heat feeling No = 0 153 (10.8%) 46 (12.7%) 0.83 [0.59–1.19] 0.345 Yes = 1 1266 (89.2%) 316 (87.3%) Cold feeling No = 0 1279 (90.1%) 326 (90.1%) 1.01 [0.68–1.47] 1 Yes = 1 140 (9.87%) 36 (9.94%) Headache༆body pain No = 0 1273 (89.7%) 343 (94.8%) 0.49 [0.29–0.78] 0.004 ** Yes = 1 146 (10.3%) 19 (5.25%) Chest pain No = 0 1418 (99.9%) 361 (99.7%) 3.92 [0.10–153] 0.365 Yes = 1 1 (0.07%) 1 (0.28%) Abdominal pain No = 0 1412 (99.5%) 356 (98.3%) 3.41 [1.07–10.5] 0.032 * Yes = 1 7 (0.49%) 6 (1.66%) Nasal congestion No = 0 1340 (94.4%) 349 (96.4%) 0.64 [0.33–1.12] 0.167 Yes = 1 79 (5.57%) 13 (3.59%) Pharyngeal discomfort No = 0 1040 (73.3%) 330 (91.2%) 0.27 [0.18–0.39] < 0.001 *** Yes = 1 379 (26.7%) 32 (8.84%) Cough No = 0 1241 (87.5%) 290 (80.1%) 1.73 [1.27–2.34] < 0.001 *** Yes = 1 178 (12.5%) 72 (19.9%) Dyspnea No = 0 1402 (98.8%) 343 (94.8%) 4.56 [2.33–8.99] < 0.001 *** Yes = 1 17 (1.20%) 19 (5.25%) Chest tightness No = 0 1412 (99.5%) 360 (99.4%) 1.18 [0.16–5.06] 1 Yes = 1 7 (0.49%) 2 (0.55%) Nausea༆vomiting No = 0 1396 (98.4%) 352 (97.2%) 1.74 [0.78–3.61] 0.223 Yes = 1 23 (1.62%) 10 (2.76%) Diarrhea No = 0 1393 (98.2%) 349 (96.4%) 2.01 [0.99–3.89] 0.066 Yes = 1 26 (1.83%) 13 (3.59%) Fatigue No = 0 827 (58.3%) 230 (63.5%) 0.80 [0.63–1.02] 0.079 Yes = 1 592 (41.7%) 132 (36.5%) Poor appetite No = 0 1406 (99.1%) 362 (100%) 1 1 Yes = 1 13 (0.92%) 0 (0.00%) Altered mental status No = 0 1408 (99.2%) 325 (89.8%) 14.4 [7.49–30.1] < 0.001 *** Yes = 1 11 (0.78%) 37 (10.2%) Underlying Disease Cardiovascular disease No = 0 1161 (81.8%) 277 (76.5%) 1.38 [1.04–1.82] 0.027 * Yes = 1 258 (18.2%) 85 (23.5%) Respiratory disease No = 0 1352 (95.3%) 332 (91.7%) 1.83 [1.15–2.83] 0.011 * Yes = 1 67 (4.72%) 30 (8.29%) Neurological disease No = 0 1324 (93.3%) 330 (91.2%) 1.36 [0.88–2.04] 0.193 Yes = 1 95 (6.69%) 32 (8.84%) Endocrine disease No = 0 1295 (91.3%) 321 (88.7%) 1.34 [0.91–1.93] 0.157 Yes = 1 124 (8.74%) 41 (11.3%) Kidney disease No = 0 1386 (97.7%) 354 (97.8%) 0.96 [0.41–2.01] 1 Yes = 1 33 (2.33%) 8 (2.21%) Blood disorders No = 0 1398 (98.5%) 354 (97.8%) 1.52 [0.62–3.36] 0.455 Yes = 1 21 (1.48%) 8 (2.21%) Oncological/immune diseases No = 0 1354 (95.4%) 344 (95.0%) 1.10 [0.62–1.84] 0.86 Yes = 1 65 (4.58%) 18 (4.97%) Blood Test Results NEU(*10 9 /L) 6.88 (3.62) 9.17 (5.00) 1.14 [1.11–1.17] < 0.001 *** LYM(*10 9 /L) 1.35 (0.76) 1.12 (0.60) 0.54 [0.43–0.67] < 0.001 *** NLR 6.43 (5.38) 11.9 (12.7) 1.09 [1.07–1.10] < 0.001 *** PLT(*10 9 /L) 220 (62.7) 213 (78.2) 1.00 [1.00–1.00] 0.115 CRP(mg/L) 22.3 (31.5) 66.5 (68.8) 1.02 [1.02–1.02] < 0.001 *** CRP/PLT 0.11 (0.17) 0.36 (0.42) 28.7 [16.8–49.1] < 0.001 *** Mean(SD) ; n(%), *P < 0.05, **P < 0.01, ***P < 0.001. Interaction Analysis Noticing that there was a significant age difference between pneumonia and non-pneumonia patients in the training cohort (mean age 62.2 vs.39.2), which is consistent with clinical practice, we introduced interaction terms of age with six other variables that showed significant differences between the groups. This is because elderly CAP patients often have insidious onset of illness and are prone to neglecting their condition. We used the random forest algorithm to model and evaluate this strategy, and found that the model performance was quite satisfactory (AUC interaction =0.80). The six interaction terms with age that we constructed all had higher importance rankings than the original features, indicating the presence of interactions (Fig. 2 ). Given the vigilance required for the elderly, when integrating the β model, in addition to using the probability calculated from the α model as a new factor and incorporating laboratory test-derived variables, we once again included age as a predictor. Model Performance Evaluation Discrimination Performance “Whether diagnosed with pneumonia” as a binary classification prediction variable, its discrimination performance of the model was measured using the receiver operating characteristic curve (ROC) and the area under the curve (AUC), also known as the C-index. In the performance evaluation of the α model, comparison of six algorithms using 10-fold cross-validation suggested that five machine learning algorithms were superior to traditional Logistic regression, indicating practical value in selecting complex machine learning algorithms (Fig. 3 ). In the comparison of the optimal models of the six algorithms using the entire dataset as the internal validation set, the random forest model had the highest AUC (AUC rf =0.94, 95%CI 0.93–0.95), followed by the AdaBoost, XGBoost, and CatBoost models (AUC ada =0.85, 95%CI 0.83–0.87; AUC xgb =0.82, 95%CI 0.80–0.85; AUC cat =0.80, 95%CI 0.77–0.83). In the external validation set that was temporally and spatially independent, the six models were compared again, with the CatBoost model performing the best (AUC cat =0.80, 95%CI 0.71–0.87). The difference in AUC (|ΔAUC|) between internal and external validation was used to measure the robustness of the models, with CatBoost being the best (|ΔAUC| cat <0.01), while randomforest was the worst (|ΔAUC| rf =0.19). The Integrated Discrimination Improvement (IDI) was used to measure the improvement in the prediction probability of actual pneumonia patients by selecting the CatBoost algorithm ( Figure S1 ). After a comprehensive comparison, the CatBoost model was selected as the final α model. The β model also employed the same six algorithms and conducted internal/external validation, among other processes. In the comparison of models built with different algorithms, the final CatBoost model emerged as the winner (AUC in =0.93, 95%CI 0.92–0.95; AUC ex =0.81, 95%CI 0.70–0.90). When integrating the β model, the α model was incorporated into the modeling in the form of predicted probabilities as a new variable. As a result, the β model utilized the most clinical raw variables, covering multimodal clinical information including general information, medical records, and laboratory tests. The ROC curves of the α/β models were also compared in internal/external validation sets, with AUC calculated and DeLong tests conducted. It was found that the β model improved predictive performance in the training cohort compared to the α model (Z in =-13.91, P < 0.001), while the difference was not significant in the external validation cohort (Z ex =-0.32, P = 0.748) ( Figure S2 ). We further calculated the optimal cutoff and confusion matrices for the α/β models during internal and external validation. The cutoffs were selected based on the ROC curve, corresponding to the threshold that maximizes the sum of sensitivity and specificity. Compared to the α model, the β model had lower optimal cutoff during internal/external validation. However, the optimal cutoffs for the same model were similar across internal/external validation (Cutoff α_in =0.246, Cutoff β_in =0.145; Cutoff α_ex =0.222, Cutoff β_ex =0.151), reflecting the robustness of the α/β models under different epidemiological conditions. The final model is intended for screening pneumonia patients, and both the training and external validation cohorts were imbalanced datasets (i.e., the number of non-pneumonia cases was several times higher than that of pneumonia cases). Therefore, we focused on sensitivity, the Matthews Correlation Coefficient (MCC), and other indicators related to the confusion matrix. The Net Reclassification Improvement (NRI) analysis also indicated that the β model was superior (NRI α/β_in =0.20; NRI α/β_ex =0.33) (Table 2 ). Table 2 Comparison of Performance based on the Confusion Matrix for the α/β Models Comparison Internal Validation External Validation α_Model β_Model α_Model β_Model True Positives 246 329 19 17 True Negatives 1163 1099 120 147 False Positives 256 320 66 39 False Negatives 116 33 5 7 NRIα/β(Cutoffβ) \ 0.20 \ 0.33 IDIα/β \ 0.26 \ 0.12 MCC 0.45 0.57 0.28 0.36 Sensitivity 0.68 0.82 0.79 0.71 Specificity 0.82 0.79 0.65 0.79 Accuracy 0.79 0.79 0.66 0.78 Balanced Accuracy 0.75 0.80 0.72 0.75 Kappa 0.44 0.49 0.21 0.32 Calibration Performance Calibration curves were drawn to assess the calibration performance.The closer the curve is to the diagonal line (from the bottom left to the top right), the better the calibration. The Brier score, which ranges from 0 to 1, is used to evaluate the accuracy of probability predictions, with lower values indicating better performance. During internal validation, the modeling processes of the α and β models using various algorithms were relatively close to the diagonal line. However, the random forest model in the β model exhibited overfitting (AUC rf =1), making it impossible to draw a calibration curveas it would overlap with the diagonal line. During external validation, the α/β models built using the CatBoost algorithm performed the best. Comparing the α and β models, all curves were found to be closer to the diagonal line, indicating that the integrated modeling strategy indeed provided added value. Comparing the Brier scores of internal/external validation also suggested that the β model was superior (Brier α_in =0.13, Brier β_in =0.08; Brier α_ex =0.11, Brier β_ex =0.08) (Fig. 4 ). Clinical Utility Clinical decision curves were plotted to assess the clinical utility. By comparing the intersection relationships between the curves and two reference lines (treating all patients and treating no patients), the clinical benefit of applying the models was evaluated. During internal validation, the α/β models built by various algorithms were all above the reference lines across the entire probability range, indicating the potential clinical benefit of the models. The clinical net benefit value was obtained at the leftmost intersection point between the curve and the reference line (Net Benefit = 0.200), which means that approximately 20 more patients out of every 100 in the training cohort could receive correct diagnosis and treatment. During external validation, the α/β models built using the CatBoost algorithm still maintained curves above the two reference lines for most of the probability range. The β models built by all algorithms improved the clinical decision curves compared to the α models. Based on the final α/β models, approximately 11.3 more patients out of every 100 in the external validation cohort could receive correct diagnosis and treatment (Net Benefit = 0.113) (Fig. 4 ). Model Interpretation and Practical Application Prior to modeling, we had already used the randomforest to calculate the importance ranking of variables. After determining the final model, we conducted SHAP interpretability analysis on the predictors of the α model again, finding that the SHAP value of the interaction term between age and Tmax was the highest, indicating that it is the most important for the α model. The β model incorporates additional test results and adjusts the predicted probability of the known α model. Therefore, we first calculated the difference in predicted probability between the β model and the α model (Probability Delta). We then used restricted cubic spline regression (RCS) on the predictors of the β model to elucidate the nonlinear relationships. Since age was highly collinear with the predicted probability of the α model, it was not included in the regression (Fig. 7). From the RCS regression plot, it can be seen that NLR (neutrophil-to-lymphocyte ratio) and CRP/PLT (C-reactive protein to platelet count ratio) are positively correlated with the probability difference (F NLR =11.61, P < 0.001; F CRP/PLT =-3.99, P = 0.008). An increase in these two factors indicates an upregulation of CAP risk. Regression analysis of the predicted probabilities of the α model shows that the β model makes smaller adjustments to extreme predicted probabilities (close to 0%or 100%) but downregulates most of the predicted probabilities. Both α/β models have been transformed into online web-based calculators using the Shiny package in RStudio. By directly entering clinical features and calculating the probabilities before and during the fever clinical visit, patients can autonomously assess the CAP risk ( https://xjbqxmbt.shinyapps.io/shiny_app_alpha/ ) ( https://xjbqxmbt.shinyapps.io/shiny_app_beta/ ). Identification of CAP’s Clinical Subtypes After successfully developing the CAP risk prediction model, latent class analysis (LCA), an unsupervised machine learning method, was used to automatically distinguish clinical subtypes in 362 CAP patients only based on the probability distribution of clinical manifestations (Fig. 8). With setted latent classes number of three, LCA identified three clinically meaningful categories: I. Class 1 (N = 48) with more severe illness, presenting with altered mental status and dyspnea, which are easily detectable and clearly related to pulmonary infection. These patients could not be further classified by symptoms due to the mental problem. II. Class 2 (N = 22) presenting with cold feeling, headach༆body pain, nausea༆vomiting, abdominal pain, and diarrhea, consistent with the TCM concept of Cold Syndrome. III. Class 3 (N = 292) presenting with heat feeling, fatigue, cough, pharyngeal discomfort, and nasal congestion, consistent with the TCM concept of Heat Syndrome. The predicted CAP risk probabilities were compared across the three classes. Class 1 had higher risk probabilities (Pred_α 1 = 0.57 ± 0.09; Pred_β 1 = 0.82 ± 0.18; W α =2851.5, P < 0.001; W β =2931, P < 0.001). Class 2/3 had similar risk (Pred_α 3 = 0.38 ± 0.18, Pred_β 3 = 0.52 ± 0.28; Pred_α2 = 0.33 ± 0.20, Pred_β 2 = 0.49 ± 0.34; W α =2685, P = 0.199; W β =3078, P = 0.745). This confirms that the model can not only identify CAP patients, but also distinguish disease severity based on risk probabilities. Discussion We explored the risk factors for CAP within the training cohort based on a cross-sectional study design framework. Some variables that were significant before adjustment (TCM_XFJB, Headache & body pain, Abdominal pain, Cardiovascular disease, Respiratory disease) were not included in the α model, but somes may still be potential risk factors. For example, underlying cardiovascular and respiratory diseases imply poorer cardiopulmonary functional reserve, which is associated with worse prognosis when facing respiratory infections. Guidelines for heart failure and COPD also indicate that infection is an important trigger for the acute exacerbation of these underlying diseases, and the two may interact causally in the overall progression[ 19 ][ 20 ]. Our model is designed to screen high-risk CAP patients in large cohorts of respiratory/febrile patients, not for disease endpoints. Thus, assessing comorbidities after initial imaging confirmation is feasible. In this study, we aimed to develop a screening tool with as few features as possible, hence only seven clinical features with significant differences between groups were ultimately selected as predictors for the α model. Among them, altered mental status (OR = 14, 95%CI 7.49–30.1) and dyspnea (OR = 4.56, 95%CI 2.33–8.99) had very high odds ratios (OR). However, in the variable importance assessment based on randomforest and the SHAP analysis of the CatBoost model, the rankings of these two variables and their interaction terms with age were relatively low. In contrast, age (OR = 1.06, 95%CI 1.05–1.06), Tmax (OR = 1.47, 95%CI 1.24–1.75), and their interaction term were ranked higher. This indicates that machine learning algorithms do not rely on extreme clinical feature differences between groups, but can effectively capture the impact of common clinical features and their interactions on CAP risk. Therefore, even in external validation across different epidemic environments, the model still has good generalization ability. Clinically, altered mental status and dyspnea are key signs of severe CAP, easily alerting patients and doctors. However, mild symptoms make CAP risk hard to estimate, often delaying diagnosis and treatment, especially in elderly patients. Our model highlights the interaction between age and other factors, emphasizing the predictive value of mild symptoms like Tmax, days since onset of illness, cough, and pharyngeal discomfort in elderly patients. This helps dynamically assess and detect CAP early in this high-risk group. When considering the predictors such as the number of days since the onset of illness (OR = 1.04, 95% CI 1.02–1.07), pharyngeal discomfort (OR = 0.27, 95% CI 0.18–0.39), and cough (OR = 1.73, 95% CI 1.27–2.34), in combination with other clinical features like headache༆body pain (OR = 0.49, 95% CI 0.29–0.78), abdominal pain (OR = 3.41, 95% CI 1.07–10.5), and the use of TCM_XFJB (OR = 0.74, 95% CI 0.57–0.95), the differences in clinical features between groups align with the natural course of respiratory infectious diseases caused by pathogens. Referring to guidelines for common pathogens such as COVID-19 and influenza[ 21 ][ 22 ], the progression from mild to moderate/severe/critical conditions is often marked by the migration of the infection from the upper respiratory tract to the lower. Latent class analysis (LCA) was conducted to further explore whether these clinical symptoms can distinguish patients ( Supplement Fig. 3 ). This transition is frequently described in traditional Chinese medicine theory as the progression from "superficial syndromes" (headache༆body pain, pharyngeal discomfort) to "interior syndromes" (cough, abdominal pain, dyspnea, altered mental status). This may explain why the use of TCM_XFJB (which are indicated for superficial syndromes) could be a potential protective factor against CAP. In other medication comparisons,no significant differences were found in the use of antipyretics,antibiotics,and antiviral drugs.This may reflect potential misuse,as these drugs can be obtained without clear diagnoses in community settings.Their preemptive use without confirmed indications does not reduce CAP risk.The single-center study design may introduce selection bias,and future multicenter studies are needed. As an exploration of CAP diagnosis based on clinical symptoms, we identified three clinical subtypes, two of which align with TCM’s cold and heat syndromes. These syndromes determine different TCM treatments: “Warming the Cold” or “Clearing the Heat”. However, there are no standard criteria to classify the two. Therefore, we used unsupervised learning (latent class analysis) without adding labels manually. Class 1 represented severe cases with noticeable altered mental status and dyspnea. Class 2 and Class 3, which accounted for the majority (86.7%) of pneumonia patients, had CT evidence of pneumonia but lacked severe pulmonary infection symptoms. Compared to current assessment tools like PSI and CURB-65 that rely heavily on extensive tests or significant abnormal signs, our model (α/β) effectively identified these patients and reflected disease severity through predicted probabilities. In the β model, we adjusted the predicted probabilities of the α model using NLR and CRP/PLT based on clinical experience. The NLR, calculated as the ratio of NEU to LYM in peripheral blood, is a biomarker that reflects both the innate immune response (mediated by NEU) and adaptive immunity (supported by LYM)[ 23 ][ 24 ][ 25 ]. An elevated NLR is often associated with inflammation, tissue damage, and a systemic inflammatory response, as neutrophils increase and lymphocytes decrease during infection or immune suppression. CRP/PLT is a composite index proposed by us based on clinical practice. It serves as an initial diagnostic marker in studies on neonatal pneumonia and sepsis[ 26 ]. However, research on adult pneumonia related to this index is rather limited and not very satisfactory[ 27 ][ 28 ][ 29 ]. CRP levels acutely rise in bacterial infections and severe viral infections, while PLT can significantly decrease due to inflammation or coagulation consumption, leading to a higher CRP/PLT. These two indicators, were chosen because they reflect the interplay of inflammation, immunity, and coagulation, which are central to the progression of respiratory infections from mild to severe[ 30 ][ 31 ][ 32 ][ 33 ]. In the β model, increases in NLR and CRP/PLT both imply an upregulation of the predicted CAP risk. According to the RCS regression, these two factors can bring about a maximum increase of approximately 10% and 40% in the predicted probability. In this study, only initial ancillary tests (complete blood count + CRP, chest CT) were used. Future work could develop more comprehensive dynamic prediction models incorporating imaging data and multiple follow-up test results to cover the entire disease course of respiratory infections. The final α/β model, as a convenient online CAP screening tool, utilized a limited number of clinical features.The use of machine learning, multimodal information, and fusion strategies proved successful. Given the complexity of interpreting machine learning algorithms, this study supplemented the analysis with SHAP and RCS regression to evaluate the contribution of each predictor. Additionally, LCA was conducted to explore the natural classification of CAP clinical subtypes, laying the groundwork for future research. Conclusion This study developed and externally validated two prediction models calculating the probability of confirmed CAP in a phased and stepwise manner. The final α/β models demonstrated satisfactory predictive performance and represent a new tool for assessing CAP risk in the population of febrile patients. The classification of CAP clinical subtypes corroborates the existing TCM experience in distinguishing Cold/Heat syndromes, providing support for future standardization. Declarations Conflict of Interest All authors have no unreported potential conflicts of interest. Author Contribution All authors contributed to the study. G.Z. and C.T. wrote the manuscript. G.Z. and H.Y. collected clinical data, G.Z. and S.Y. conducted statistical analysis, X.X. and L.B. guided the methodology, and L.Q. explored clinical significance. All authors reviewed the manuscript. Data Availability The datasets generated and/or analyzed during the current study are available from the corresponding authors upon reasonable request. References Quan, T. P. et al. Increasing burden of community-acquired pneumonia leading to hospitalisation, 1998–2014. Thorax 71 (6), 535–542 (2016). Chalmers, J. D. et al. Severity assessment tools for predicting mortality in hospitalised patients with community-acquired pneumonia. Systematic review and meta-analysis. Thorax 65 (10), 878–883 (2010). Metlay, J. P. et al. Diagnosis and Treatment of Adults with Community-acquired Pneumonia. An Official Clinical Practice Guideline of the American Thoracic Society and Infectious Diseases Society of America. Am. J. Respir Crit. Care Med. 200 (7), e45–e67 (2019). Chinese Medical Doctor Association, Division of Emergency Physicians, Chinese Emergency Medicine Consortium. Beijing Association for Emergency Medicine. Clinical Practice Guidelines for Community-Acquired Pneumonia in Adult Emergency Patients 2024 Edition. Chinese Journal of Emergency Medicine,2025,3403300-317. Simonetti, A. F., Viasus, D., Garcia-Vidal, C. & Carratalà, J. Management of community-acquired pneumonia in older adults. Ther. Adv. Infect. Dis. 2 (1), 3–16 (2014). Cillóniz, C., Rodríguez-Hurtado, D. & Torres, A. Characteristics and Management of Community-Acquired Pneumonia in the Era of Global Aging. Med. Sci. (Basel) . 6 (2), 35 (2018). Published 2018 Apr 30. Bradley, J. et al. Pneumonia Severity Index and CURB-65 Score Are Good Predictors of Mortality in Hospitalized Patients With SARS-CoV-2 Community-Acquired Pneumonia. Chest 161 (4), 927–936 (2022). Tuta-Quintero, E. et al. Comparison of performances between risk scores for predicting mortality at 30 days in patients with community acquired pneumonia. BMC Infect. Dis. 24 (1), 912 (2024). Published 2024 Sep 3. SHINDO, Y. et al. Comparison of severity scoring systems A-DROP and CURB-65 for community-acquired pneumonia. Respirology 13 , 731–735 (2008). Shao, J. et al. A multimodal integration pipeline for accurate diagnosis, pathogen identification, and prognosis prediction of pulmonary infections. Innov. (Camb) . 5 (4), 100648 (2024). Published 2024 May 22. Yang, Z. et al. Development and validation of machine learning-based prediction model for severe pneumonia: A multicenter cohort study. Heliyon 10 (17), e37367 (2024). Published 2024 Sep 3. Tong et al. Strategic Considerations for Strengthening the Construction of the Traditional Chinese Medicine Emergency Prevention and Control System for Emerging and Sudden Infectious Diseases in ChinaJ.Bulletin of the Chinese Academy of Sciences,2020,3509,1087–1095 . Liu, J. et al. Combination of Hua Shi Bai Du granule (Q-14) and standard care in the treatment of patients with coronavirus disease 2019 (COVID-19): A single-center, open-label, randomized controlled trial. Phytomedicine 91 , 153671 (2021). Department of Internal Medicine,China Association of Chinese Medicine;Department of Pulmonary Diseases,China Association of Chinese Medicine. Department of Pulmonary Diseases,China Nationality Medicine Association;Yu Xueqing,Xie Yang,Li Jiansheng. Guidelines for the Diagnosis and Treatment of Community-Acquired Pneumonia in Traditional Chinese Medicine, Revised Edition 2018[J]. Journal of Traditional Chinese Medicine,2019,04350 – 360. Shasha et al. The Incoming Influenza Season — China, the United Kingdom, and the United States, 2021–2022[J]. China CDC Wkly. 3 (49), 1039–1045 (2021). Ge, J. The COVID-19 pandemic in China: from dynamic zero-COVID to current policy. Herz 48 , 226–228 (2023). Jue et al. Trends of SARS-CoV-2 Infection in Sentinel Community-Based Surveillance After the Optimization of Prevention and Control Measures — China, December 2022–January 2023[J]. China CDC Wkly. 5 (7), 159–164 (2023). China, C. D. C. Reported Cases and Deaths of National Notifiable Infectious Diseases — China, April 2023*[J]. China CDC Wkly. 5 (33), 742–743 (2023). Chinese Geriatrics Society,Division of Electrocardiology and Cardiac Function,Chinese Medical Doctor Association,Division of Cardiology,Expert Committee of the Heart Failure Center Alliance,Yang Jiefu. Expert Consensus on the Comprehensive Management of Patients with Worsening Chronic Heart Failure in China,2022. J. Chin. Circulation J. 373 , 215–225 (2022). Expert Group on the Diagnosis and Treatment of Acute Exacerbation of Chronic Obstructive Pulmonary Disease. Chinese Expert Consensus on the Diagnosis and Treatment of Acute Exacerbation of Chronic Obstructive Pulmonary DiseaseRevised Edition 2023[J]. International Journal of Respiratory Diseases,2023,432132-149. General Office of the National Health Commission of the Peoples Republic of China&General Office of the National Administration of Traditional Chinese. Medicine of the Peoples Republic of China.2023. Diagnosis and Treatment Protocol for COVID-19. Infect. Trial Version 10 Chin. Med. , 1802 ,161–166 . National Health Commission of the People's Republic of China & National Administration of Traditional Chinese Medicine. Diagnosis and Treatment Protocol for Influenza (2025 Edition). Chin. J. Ration. Drug Use . 22 (02), 1–7 (2025). Cataudella, E. et al. Neutrophil-To-Lymphocyte Ratio: An Emerging Marker Predicting Prognosis in Elderly Adults with Community-Acquired Pneumonia. J. Am. Geriatr. Soc. 65 (8), 1796–1801 (2017). Drăgoescu, A. N. et al. Neutrophil to Lymphocyte Ratio (NLR)-A Useful Tool for the Prognosis of Sepsis in the ICU. Biomedicines 10 (1), 75 (2021). Published 2021 Dec 30. Li, X. et al. Predictive values of neutrophil-to-lymphocyte ratio on disease severity and mortality in COVID-19 patients: a systematic review and meta-analysis. Crit. Care . 24 , 647 (2020). Li, X. et al. C-reactive protein to platelet ratio as an early biomarker in differentiating neonatal late-onset sepsis in neonates with pneumonia. Sci. Rep. 15 (1), 10760 (2025). Published 2025 Mar 28. Ge, S., Ma, Y., Xie, M., Qiao, T. & Zhou, J. The role of platelet to mean platelet volume ratio in the identification of adult-onset still's disease from sepsis. Clin. (Sao Paulo) . 76 , e2307 (2021). Published 2021 Apr 16. Zhang, Y. et al. Diagnostic Value and Prognostic Significance of Procalcitonin Combined with C-Reactive Protein in Patients with Bacterial Bloodstream Infection. Comput. Math. Methods Med. 2022 , 6989229 (2022). Published 2022 Aug 11. Mirsaeidi, M. et al. Thrombocytopenia and thrombocytosis at time of hospitalization predict mortality in patients with community-acquired pneumonia. Chest 137 (2), 416–420 (2010). Kombe Kombe, A. J. et al. The Role of Inflammation in the Pathogenesis of Viral Respiratory Infections. Microorganisms 12 (12), 2526 (2024). Published 2024 Dec 7. Yang, Y. & Tang, H. Aberrant coagulation causes a hyper-inflammatory response in severe influenza pneumonia. Cell. Mol. Immunol. 13 (4), 432–442 (2016). Tay, M. Z., Poh, C. M., Rénia, L., MacAry, P. A. & Ng, L. F. P. The trinity of COVID-19: immunity, inflammation and intervention. Nat. Rev. Immunol. 20 (6), 363–374 (2020). Bouwman, J. J., Visseren, F. L., Bosch, M. C., Bouter, K. P. & Diepersloot, R. J. Procoagulant and inflammatory response of virus-infected monocytes. Eur. J. Clin. Invest. 32 (10), 759–766 (2002). Additional Declarations No competing interests reported. Supplementary Files StagedIdentificationofCAPinFeverPatientsAcrossEpidemicEnvironmentsModelingValidationsupplement.pdf Cite Share Download PDF Status: Published Journal Publication published 18 Dec, 2025 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Revision requested 29 Jul, 2025 Reviews received at journal 18 Jul, 2025 Reviewers agreed at journal 29 Jun, 2025 Reviews received at journal 22 Jun, 2025 Reviewers agreed at journal 21 Jun, 2025 Reviewers invited by journal 20 Jun, 2025 Editor assigned by journal 20 Jun, 2025 Editor invited by journal 04 Jun, 2025 Submission checks completed at journal 31 May, 2025 First submitted to journal 31 May, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6768762","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":474758345,"identity":"d9ff9bf8-bb96-4c88-a591-90839334eb1e","order_by":0,"name":"Gao Ziheng","email":"","orcid":"","institution":"Bejing University of Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Gao","middleName":"","lastName":"Ziheng","suffix":""},{"id":474758346,"identity":"ce948a38-bcab-4948-9c1d-36593ac5524c","order_by":1,"name":"Chen Tengfei","email":"","orcid":"","institution":"Beijing Traditional Chinese Medicine Hospital","correspondingAuthor":false,"prefix":"","firstName":"Chen","middleName":"","lastName":"Tengfei","suffix":""},{"id":474758347,"identity":"fe945eeb-3d65-43de-8af4-cdfffd956f9c","order_by":2,"name":"Ha Yanxiang","email":"","orcid":"","institution":"Beijing Traditional Chinese Medicine Hospital","correspondingAuthor":false,"prefix":"","firstName":"Ha","middleName":"","lastName":"Yanxiang","suffix":""},{"id":474758348,"identity":"f4e9b453-1e93-4701-8a92-4c152f0ffbee","order_by":3,"name":"Shi Yifan","email":"","orcid":"","institution":"Bejing University of Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Shi","middleName":"","lastName":"Yifan","suffix":""},{"id":474758349,"identity":"ab27486d-553d-4a7b-b39d-e143df2ba98c","order_by":4,"name":"Xu Xiaolong","email":"","orcid":"","institution":"Beijing Traditional Chinese Medicine Hospital","correspondingAuthor":false,"prefix":"","firstName":"Xu","middleName":"","lastName":"Xiaolong","suffix":""},{"id":474758350,"identity":"a2bf8a3d-da0b-481c-bdf6-4c597047edd3","order_by":5,"name":"Li Bo","email":"","orcid":"","institution":"Beijing Traditional Chinese Medicine Hospital","correspondingAuthor":false,"prefix":"","firstName":"Li","middleName":"","lastName":"Bo","suffix":""},{"id":474758351,"identity":"8e86134c-b17e-4ea0-a5b3-cf5cc1910413","order_by":6,"name":"Liu Qingquan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+0lEQVRIiWNgGAWjYBACPmYgkQBhMz5IMKiRY2NvPoBXCxuSFmaDDxXHjPl4jiXg14LMlpxxhjlxnkSOAn4t7DyGNx7uqJXdcPzsAWneNrb0NoYcBoYfFdvwOIzH2CLxzHHjDWfyEox522Ry2xjOHmDsOXMbnxYzicS2Y4kbDuQYJANtyW1j7EtgZmwjRsv5NwaHeduY04EiBsRoqUnccCPHsBHo/QQ2NoJa2IotEtsOGM+88caYARjIhm08bAkH8fmFn//wxps/2+pk+87nmP8ARqW8/PzHBx/8qMCtBQQkGBgOMzYgixzAqx6ipQ5VyygYBaNgFIwCZAAAunVWmsoAozgAAAAASUVORK5CYII=","orcid":"","institution":"Beijing Traditional Chinese Medicine Hospital","correspondingAuthor":true,"prefix":"","firstName":"Liu","middleName":"","lastName":"Qingquan","suffix":""}],"badges":[],"createdAt":"2025-05-28 13:53:23","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6768762/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6768762/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-025-29689-6","type":"published","date":"2025-12-18T15:58:40+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":85616943,"identity":"ddd64f22-81e8-405f-8474-33f0cb87ba1e","added_by":"auto","created_at":"2025-06-29 14:40:25","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":176476,"visible":true,"origin":"","legend":"\u003cp\u003eFlowchart of Inclusion/Exclusion of Training/External Validation Cohort and Composition of 10-Fold Cross/Internal/External Validation Datasets\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-6768762/v1/9715f25d1e010e72256fbf97.png"},{"id":85615408,"identity":"0978d99f-a06c-4051-9407-57d0a9672bc4","added_by":"auto","created_at":"2025-06-29 14:24:25","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":213263,"visible":true,"origin":"","legend":"\u003cp\u003eEvaluation Plots for Predictor Selection Strategies and Interaction Construction (\u003cstrong\u003eA.\u003c/strong\u003eForest plot of clinical variables; \u003cstrong\u003eB.\u003c/strong\u003e Comparison of variable importance for three strategies selecting predictors; \u003cstrong\u003eC\u0026amp;D.\u003c/strong\u003e Lasso regression plots; \u003cstrong\u003eE\u0026amp;G\u0026amp;H.\u003c/strong\u003e Performance comparison of Randomforest models using three strategies; \u003cstrong\u003eF.\u003c/strong\u003e Comparison of variable importance in Randomforest models after incorporating interaction terms)\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-6768762/v1/b54dd7c35725a7fe4cf7d934.png"},{"id":85616944,"identity":"a1cd56ed-5133-4156-8717-c7022041589e","added_by":"auto","created_at":"2025-06-29 14:40:25","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":246845,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of ROC Curves for 2×6 Models (\u003cstrong\u003eA\u0026amp;C\u003c/strong\u003e. Internal validation; \u003cstrong\u003eB\u0026amp;D\u003c/strong\u003e. External validation)\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-6768762/v1/6b3bf107dcb7b0111f10c92b.png"},{"id":85615727,"identity":"e476568d-8997-44c6-a333-9b69559569e2","added_by":"auto","created_at":"2025-06-29 14:32:25","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":216512,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of Calibration and Decision Curves for 2×6 Models (\u003cstrong\u003eA\u0026amp;C\u0026amp;E\u0026amp;G\u003c/strong\u003e. Internal validation; \u003cstrong\u003eB\u0026amp;D\u0026amp;F\u0026amp;H\u003c/strong\u003e. External validation)\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-6768762/v1/56baa1d934a787fb01b6419b.png"},{"id":85615416,"identity":"65747513-5364-4cec-9f8f-cd45078310d6","added_by":"auto","created_at":"2025-06-29 14:24:25","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":219718,"visible":true,"origin":"","legend":"\u003cp\u003eSeries of Interpretability Analysis Plots for the Two Best α/β Models (\u003cstrong\u003eA&B.\u003c/strong\u003e SHAP analysis for the α model; \u003cstrong\u003eC.\u003c/strong\u003e RCS regression of the β model on the delta of predicted probability between the α and β)\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-6768762/v1/b6ae02d4e6bd938e990a6a92.png"},{"id":85615725,"identity":"1ad9b657-71f9-4ee4-86ba-0e534b13b4b3","added_by":"auto","created_at":"2025-06-29 14:32:25","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":122563,"visible":true,"origin":"","legend":"\u003cp\u003eSeries of Latent Class Analysis Charts for Automatic Subtype Classification of CAP (\u003cstrong\u003eA. \u003c/strong\u003eLCA plot of three latent classes containing the Cold/Heat syndrome; \u003cstrong\u003eB.\u003c/strong\u003e LCA’s Probability differences between Cold/Heat syndrome)\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-6768762/v1/d9f115ce36086a5a87002fbc.png"},{"id":98814347,"identity":"45352df1-d7ec-4769-8449-4b5e405501d7","added_by":"auto","created_at":"2025-12-22 16:12:24","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2319455,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6768762/v1/fb8f2b70-39dc-480b-a7a8-6aa16891f958.pdf"},{"id":85615723,"identity":"d98a60e7-57c6-4094-9069-0abaa3d7ab5e","added_by":"auto","created_at":"2025-06-29 14:32:25","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":586963,"visible":true,"origin":"","legend":"","description":"","filename":"StagedIdentificationofCAPinFeverPatientsAcrossEpidemicEnvironmentsModelingValidationsupplement.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6768762/v1/3a8c994b344d10c60db08cb5.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Staged Identification of CAP in Fever Patients Across Epidemic Environments: Modeling \u0026Validation","fulltext":[{"header":"Introduction","content":"\u003cp\u003eCommunity-acquired pneumonia (CAP) is a common and frequently occurring respiratory infectious disease with high incidence and mortality rates[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e][\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Early identification of CAP is challenging, as its confirmation relies on expensive imaging equipment, such as chest CT[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e][\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. But even X-ray examinations can only be interpreted by medical professionals. During sudden outbreaks, when medical resources are strained, medical congestion is particularly prominent in primary care clinics, emergency departments, and fever clinics. Moreover, pneumonia in the elderly population is clinically insidious and more difficult to identify early, leading to delayed treatment[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e][\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe Pneumonia Severity Index (PSI) is a commonly used tool in clinical practice for clinicians and researchers to assess the severity of illness in hospitalized patients[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e][\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. It includes many precise ancillary tests, making it difficult for patients to use and challenging to implement in rudimentary settings. Outpatients also struggle to complete all the required ancillary tests to calculate the score. Other scoring systems like CURB-65, A-DROP, and SMART-COP can rapidly assess pneumonia severity using a limited number of variables[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e][\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. However, these assessments rely on significantly abnormal vital signssuch as respiratory rate, level of consciousness, blood pressure, and heart rateand specific ancillary tests (such as blood gas analysis, albumin, or blood urea nitrogen testing). They are primarily designed to quickly identify critically ill patients and provide emergency care, rather than to differentiate between pneumonia and non-pneumonia patients. Existing pneumonia prediction models often incorporate a wide range of clinical characteristic variables[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e][\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e], including high-cost chest CT imaging and specific biomarkers (such as IL-6, IL-10), aiming to accurately identify CAP patients, improve detection accuracy, and precisely identify pathogens. However, compared to further improving diagnostic accuracy, the ability to quickly, accurately, and cost-effectively identify pneumonia patients and enable stratified diagnosis and treatment during sudden outbreak responses is equally of great clinical value.\u003c/p\u003e \u003cp\u003eAdditionally, the TCM treatment of CAP primarily focuses more on classifying syndromes by clinical symptoms than emphasizing etiology, a feature that has proven uniquely advantageous in the COVID-19 pandemic response[\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e][\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Clinicians empirically classify CAP into Cold/Heat syndrome and select corresponding TCM interventions. Current TCM guideline only mentions the coexistence of \u0026ldquo;External Cold\u0026rdquo;\u0026ldquo;Internal Heat\u0026rdquo; without differentiation[\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Therefore, the clinical classification of the two syndromes still lacks a standard.\u003c/p\u003e \u003cp\u003eDuring the COVID-19 pandemic, China implemented long-term and strict prevention and control policies from December 2019 to December 2022, creating a unique epidemiological environment. In this specialized social environment that effectively blocked the spread of the coronavirus, influenza viruses and other pathogens became the mainstream of respiratory viral infections[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e][\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. After the end of these policies in December 2022, the Omicron variant of the coronavirus, rather than the Alpha or Delta variants, became widely prevalent. It has a lower proportion of severe cases and mortality rate, but stronger transmissibility and immune escape ability. The prevalent respiratory pathogens gradually shifted to mainly include the coronavirus, influenza virus, rhinovirus, and others[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e][\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. It has become a new challenge for models to accurately identify pneumonia patients in different epidemiological environments,and even in the presence of new pathogens.\u003c/p\u003e \u003cp\u003eThis study develops and validates a step-by-step CAP risk prediction model using multimodal data and machine learning across different epidemic settings. It aims to provide an easy-to-use, sensitive, and robust tool for early CAP detection. Additionally, it classifies CAP subtypes based on clinical symptoms to support TCM standardization.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eMaterials\u003c/h2\u003e \u003cdiv id=\"Sec4\" class=\"Section3\"\u003e \u003ch2\u003eData Source\u003c/h2\u003e \u003cp\u003eThe data used in this study were derived from two temporally and spatially independent cohorts of patients who visited the fever clinic of Beijing Hospital of Traditional Chinese Medicine (BHTCM), representing different epidemic environments. All data were anonymously extracted from the hospital information management system. From December 2021 to December 2022, 2,193 visits were selected from a total of 7,203 visits throughout the year based on disease diagnosis. After inclusion and exclusion criteria were applied, these visits formed the training cohort of 1,781 patients used for model development (also used for internal validation). From January to July 2024, 300 visits were randomly selected from a total of 2,771 visits over the six-month period (without disease diagnosis selection) to form the external validation cohort of 210 patients. The data from the two cohorts were kept separate, and the external validation cohort was used only after the model was completed (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e\n\u003ch3\u003eInclusion Criteria\u003c/h3\u003e\n\u003cp\u003eWe included patients with community-acquired pneumonia (CAP) and similar conditions (age over 18). The diagnoses of similar conditions needed to be related to respiratory tract infections, including acute upper respiratory tract infection, influenza, common cold, acute tracheitis/bronchitis, acute tonsillitis, acute pharyngitis, etc. All disease diagnoses were obtained from the outpatient information system. The clinical diagnosis and the CAP diagnostic criteria referenced in this study were as follows[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e][\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]:\u003c/p\u003e \u003cp\u003e \u003cb\u003eI.\u003c/b\u003e Onset in the community. \u003cb\u003eII.\u003c/b\u003e Presence of any one of the following manifestations related to pneumonia: \u003cb\u003ea.\u003c/b\u003e New onset of cough, sputum production, or worsening of pre-existing respiratory symptoms, with or without purulent sputum, chest pain, dyspnea, and hemoptysis; \u003cb\u003eb.\u003c/b\u003e Fever; \u003cb\u003ec.\u003c/b\u003e Physical signs of lung consolidation and/or auscultation of moist rales; \u003cb\u003ed.\u003c/b\u003e Peripheral blood leukocyte count\u0026thinsp;\u0026gt;\u0026thinsp;10 \u0026times; 10^9/L or \u0026lt;\u0026thinsp;4 \u0026times; 10^9/L, with or without left shift. \u003cb\u003eIII.\u003c/b\u003e Chest imaging showing new pulmonary infiltrates, lobar or segmental consolidation, ground-glass opacity, or interstitial changes, with or without pleural effusion.\u003c/p\u003e\n\u003ch3\u003eExclusion Criteria\u003c/h3\u003e\n\u003cp\u003eAll patients included in the study were assessed against the following exclusion criteria: \u003cb\u003eI.\u003c/b\u003e Exclusion of follow-up medical records, retaining only the records of patients' initial visits; \u003cb\u003eII.\u003c/b\u003e Exclusion of patients diagnosed with pneumonia who lack chest imaging evidence; \u003cb\u003eIII.\u003c/b\u003e Exclusion of patients with conditions similar to severe CAP, such as pulmonary edema, diffuse alveolar hemorrhage, interstitial lung disease, etc.; \u003cb\u003eIV.\u003c/b\u003e Exclusion of patients who were transferred to another hospital or died on the day of their visit; \u003cb\u003eV.\u003c/b\u003e Exclusion of patients with missing important clinical variables (basic information, Tmax, number of days since onset of illness, disease diagnosis, etc.).\u003c/p\u003e\n\u003ch3\u003eMethods\u003c/h3\u003e\n\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eDisease Diagnostic Criteria\u003c/h2\u003e \u003cp\u003e Community-acquired pneumonia (CAP) was diagnosed according to the latest relevant guideline criteria, and all patients with pneumonia in the study were required to have chest CT imaging to support the diagnosis. For practicality, we established the training cohort from patients with CAP and similar conditions rather than all patients who visited the clinic. Due to the different diagnostic perspectives, conditions such as \"pulmonary infection\" \"pneumonia\" \"pneumonia caused by known viruses\" \"viral pneumonia\" \"bacterial pneumonia\" and \"fungal pneumonia\" were considered clinically synonymous with CAP. Non-pneumonia patients were diagnosed with conditions such as \"acute upper respiratory tract infection\" \"influenza caused by known viruses\" \"common cold\" and \"acute tonsillitis\", all of which were determined based on the standard ICD-10 disease names. Because of the differences in the epidemic environment, there were no patients with COVID-19 infection in the training cohort, and therefore no patients with CAP caused by COVID-19. Thus, diagnoses as \"COVID-19 infection\" and \"COVID-19 pneumonia\" were only present in the external validation cohort.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eData Collection \u0026 Management\u003c/h3\u003e\n\u003cp\u003eAll patient data were anonymously extracted from the hospital information system,including basic information, medical records, laboratory test results, and chest CT examination results. Basic information included sex age, and date of visit, while medical records included chief complaints, medical history, past medical history, physical examination, and disease diagnosis. All patient data were anonymized (identifiable patient information was removed) and aggregated into an Excel database.\u003c/p\u003e\n\u003ch3\u003eModeling \u0026 Validation\u003c/h3\u003e\n\u003cp\u003eIn predictor selection, three strategies were compared comprehensively: \u003cb\u003eI.\u003c/b\u003e Directly using all extracted clinical feature variables; \u003cb\u003eII.\u003c/b\u003e Using only variables with significant differences between the pneumonia and non-pneumonia groups; \u003cb\u003eIII.\u003c/b\u003e Using only important variables selected by Lasso regularization regression. Randomforest was used to model the training cohort, comparing the variable importance rankings among the three strategies and comparing the AUC values of the three models using 10-fold cross-validation. The strategy with the fewest variables was chosen as the predictor set. The same method was used to compare the variable importance rankings with interaction terms added, and the gain and necessity of adding interaction terms were assessed.\u003c/p\u003e \u003cp\u003eFor modeling algorithms, traditional Logistic regression and five common machine learning algorithms were used (Logisticnet, Randomforest, XGBoost, AdaBoost, and CatBoost). To fully utilize the data, instead of using traditional random splitting methods(e.g., 7:3 training and validation set split), 10-fold cross-validation was conducted on the complete training set to obtain the best models and parameters for the six different algorithms. The six models were then internally validated and compared in the entire patient population. The temporally and spatially independent external validation set was used to evaluate the model performance in real clinical settings. The best-performing model among the six was determined as the final α model,with all predictors obtainable by patients before their visit. The Shapley Additive Explanations (SHAP) method was used to assess the impact of original variables on the output. For multimodal clinical data, a fusion modeling strategy was adopted by adding laboratory test-related variables to develop the β model. The β model was built on the probability obtained from the α model and included additional variables such as the neutrophil-to-lymphocyte ratio (NLR\u0026thinsp;=\u0026thinsp;NEU/LYM) and the C-reactive protein-to-platelet count ratio (CRP/PLT). These variables can be obtained through routine blood tests. The same six algorithms were used with 10-fold cross-validation for parameter tuning,modeling,and internal and external validation, with the best-performing β model selected. Both α/β models were developed into Shiny online web calculators for patients to assess pneumonia risk probability at different stages of their visit.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis Methods\u003c/h2\u003e \u003cp\u003eThis predictive modeling study employed a cross-sectional research design. Continuous variables were presented as medians and interquartile ranges, while categorical variables were described using counts and percentages. Normally distributed variables were compared using the two-sample t-test, non-normally distributed variables were compared using the Mann-Whitney U test, and categorical variables were compared using chi-square tests, Fisher's exact tests, or Kruskal-Wallis tests as appropriate.\u003c/p\u003e \u003cp\u003eMissing values (with the proportion of missing values in each variable being less than 10%) were handled using multiple imputation by chained equations based on the random forest algorithm, with 10 imputations performed and the results combined according to Rubin's rules. During model training, 10-fold cross-validation was utilized. The entire training set was randomly divided into 10 equal subsets, with one subset (10%) serving as the validation set and the remaining nine subsets (90%) as the training set for each iteration. This process was repeated 10 times, and the parameters that performed best in the validation set were selected as the optimal parameters for the model based on that algorithm.All p-values and confidence intervals were derived from two-sided tests, with a significance level set at 0.05. Bonferroni correction was applied to multiple testing to control the family-wise error rate (FWER) by using a more stringent significance level. Data cleaning, organization, and statistical analysis were conducted using RStudio (R 4.4.3) and Excel software. All R packages used have been described, and the main R codes had been provided also (\u003cb\u003eSupplement\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eMethods\u0026rsquo; Technical Review\u003c/h2\u003e \u003cp\u003eOur study was based on past cross-sectional data from the fever clinic. We downloaded all the anonymized data from the Beijing Hospital of Traditional Chinese Medicine\u0026rsquo;s YiduCloud HIS database after obtaining approval from the hospital\u0026rsquo;s clinical research department, which reviewed our protocol and waived informed consent from the original patients. Results are reported according to the TRIPOD statement (\u003cb\u003eSupplement\u003c/b\u003e Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eResearch Support and Funding\u003c/h2\u003e \u003cp\u003eThis study was funded by the National Natural Science Foundation of China (NNSFC 81774146) through the Beijing Hospital of Traditional Chinese Medicine. The sponsor had no role in the study. As its retrospective, anonymized, and non-intervention nature, the hospital\u0026rsquo;s ethics committee waived ethical approval during the model\u0026rsquo;s development, following the Declaration of Helsinki.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003ePatient Clinical Characteristics\u003c/h2\u003e \u003cp\u003eWe compared the clinical characteristics of patients with pneumonia and those without pneumonia in the training cohort (N\u0026thinsp;=\u0026thinsp;362 vs.N\u0026thinsp;=\u0026thinsp;1419) and calculated the odds ratios (OR) (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Similar comparison also conducted in the external validation cohort after modeling (\u003cb\u003eSupplementary Table\u0026nbsp;1\u003c/b\u003e). Among the 33 variables derived from general information and medical records, seven variables\u0026mdash;age, Tmax, days (since onset of illness), presence or worsening of pharyngeal discomfort, presence or worsening of cough, presence or worsening of dyspnea, and presence or worsening of altered mental status\u0026mdash;remained significantly different between the groups after adjustment (all P\u0026thinsp;\u0026lt;\u0026thinsp;0.001). The presence or worsening of headache and abdominal pain were only significantly different before adjustment (OR\u003csub\u003ehb_pain\u003c/sub\u003e=0.49, 95%CI 0.29\u0026ndash;0.78, P\u0026thinsp;=\u0026thinsp;0.002; OR\u003csub\u003eab_pain\u003c/sub\u003e=3.41, 95%CI 1.07\u0026ndash;10.5, P\u0026thinsp;=\u0026thinsp;0.039). Regarding underlying diseases, we examined the differences between groups for cardiovascular disease, respiratory disease, neurological disease, kidney disease, blood disorders, and oncological/immune diseases. Only cardiovascular disease and respiratory disease showed significant differences before adjustment (OR\u003csub\u003ecard\u003c/sub\u003e=1.38, 95%CI 1.04\u0026ndash;1.82, P\u0026thinsp;=\u0026thinsp;0.025 ;OR\u003csub\u003eresp\u003c/sub\u003e=1.83, 95%CI 1.15\u0026ndash;2.83, P\u0026thinsp;=\u0026thinsp;0.011), but these differences were not significant after adjustment. In terms of medication history, we compared the use of medications since the onset of illness. There were no significant differences in the use of antipyretics, antibiotics, or antiviral drugs. For traditional Chinese medicine (TCM), the use of TCM_XFJB was only significantly different before adjustment (OR\u003csub\u003etcm_xfjb\u003c/sub\u003e=0.74, 95%CI 0.57\u0026ndash;0.95, P\u0026thinsp;=\u0026thinsp;0.019). In laboratory tests, we extracted six key clinical variables related to infection: neutrophil count (NEU), lymphocyte count (LYM), platelet count (PLT), C-reactive protein (CRP), neutrophil-to-lymphocyte ratio (NLR), and C-reactive protein-to-platelet count ratio (CRP/PLT). All these variables showed significant differences between groups after adjustment (all P\u0026thinsp;\u0026lt;\u0026thinsp;0.001). In the hypothesis testing above, the Bonferroni correction was used to adjust the significance level to 0.05/33\u0026thinsp;+\u0026thinsp;6\u0026thinsp;\u0026asymp;\u0026thinsp;0.001. Among the three predictor selection strategies for the random forest model, the third strategy performed the best (AUC\u003csub\u003eI\u003c/sub\u003e=0.78; AUC\u003csub\u003eII\u003c/sub\u003e=0.78; AUC\u003csub\u003eIII\u003c/sub\u003e=0.80). After comparing the variable importance rankings, the final α model used only the seven variables that remained significantly different between groups after multiple testing as predictors (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of Characteristic between Pneumonia and Non-Pneumonia in the Training Cohort\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"3\" morerows=\"1\" nameend=\"c3\" namest=\"c1\" rowspan=\"2\"\u003e \u003cp\u003eCharacteristic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNon-pneumonia\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePneumonia\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eOdds Ratio 95%CI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eP value\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e(N\u0026thinsp;=\u0026thinsp;1419)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e(N\u0026thinsp;=\u0026thinsp;362)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"4\" rowspan=\"5\"\u003e \u003cp\u003eBasic Profile\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSex\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFemale\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e713 (50.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e172 (47.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1.12 [0.89\u0026ndash;1.41]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.354\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMale\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e706 (49.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e190 (52.5%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eAge(y)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e39.2 (17.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e62.2 (22.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.06 [1.05\u0026ndash;1.06]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003csup\u003e***\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eDays(d)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2.58 (4.05)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3.50 (4.49)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.04 [1.02\u0026ndash;1.07]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003csup\u003e***\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eTmax(℃)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e38.1 (0.64)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e38.3 (0.71)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.47 [1.24\u0026ndash;1.75]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003csup\u003e***\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"13\" rowspan=\"14\"\u003e \u003cp\u003ePast Medication History\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eMedication history\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e430 (30.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e112 (30.9%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.97 [0.76\u0026ndash;1.25]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.864\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e989 (69.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e250 (69.1%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eUsed antipyretics\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e980 (69.1%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e263 (72.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.84 [0.65\u0026ndash;1.08]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.206\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e439 (30.9%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e99 (27.3%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eUsed antibiotics\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1109 (78.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e287 (79.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.94 [0.70\u0026ndash;1.24]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.694\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e310 (21.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e75 (20.7%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eUsed antivirals\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1416 (99.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e360 (99.4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e2.68 [0.31\u0026ndash;17.7]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.269\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3 (0.21%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2 (0.55%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eTCM_XFJB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e923 (65.0%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e259 (71.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.74 [0.57\u0026ndash;0.95]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.023\u003csup\u003e*\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e496 (35.0%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e103 (28.5%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eTCM_XRJD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1161 (81.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e303 (83.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.88 [0.64\u0026ndash;1.19]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.448\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e258 (18.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e59 (16.3%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eTCM_WLSH\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1418 (99.9%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e362 (100%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1 (0.07%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0 (0.00%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"29\" rowspan=\"30\"\u003e \u003cp\u003eClinical Manifestation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eHeat feeling\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e153 (10.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e46 (12.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.83 [0.59\u0026ndash;1.19]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.345\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1266 (89.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e316 (87.3%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eCold feeling\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1279 (90.1%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e326 (90.1%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1.01 [0.68\u0026ndash;1.47]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e140 (9.87%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e36 (9.94%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eHeadache༆body pain\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1273 (89.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e343 (94.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.49 [0.29\u0026ndash;0.78]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.004\u003csup\u003e**\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e146 (10.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e19 (5.25%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eChest pain\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1418 (99.9%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e361 (99.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e3.92 [0.10\u0026ndash;153]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.365\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1 (0.07%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1 (0.28%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eAbdominal pain\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1412 (99.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e356 (98.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e3.41 [1.07\u0026ndash;10.5]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.032\u003csup\u003e*\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e7 (0.49%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e6 (1.66%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eNasal congestion\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1340 (94.4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e349 (96.4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.64 [0.33\u0026ndash;1.12]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.167\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e79 (5.57%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e13 (3.59%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003ePharyngeal discomfort\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1040 (73.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e330 (91.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.27 [0.18\u0026ndash;0.39]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003csup\u003e***\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e379 (26.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e32 (8.84%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eCough\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1241 (87.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e290 (80.1%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1.73 [1.27\u0026ndash;2.34]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003csup\u003e***\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e178 (12.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e72 (19.9%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eDyspnea\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1402 (98.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e343 (94.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e4.56 [2.33\u0026ndash;8.99]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003csup\u003e***\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e17 (1.20%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e19 (5.25%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eChest tightness\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1412 (99.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e360 (99.4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1.18 [0.16\u0026ndash;5.06]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e7 (0.49%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2 (0.55%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eNausea༆vomiting\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1396 (98.4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e352 (97.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1.74 [0.78\u0026ndash;3.61]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.223\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e23 (1.62%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10 (2.76%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eDiarrhea\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1393 (98.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e349 (96.4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e2.01 [0.99\u0026ndash;3.89]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.066\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e26 (1.83%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e13 (3.59%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eFatigue\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e827 (58.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e230 (63.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.80 [0.63\u0026ndash;1.02]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.079\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e592 (41.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e132 (36.5%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003ePoor appetite\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1406 (99.1%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e362 (100%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e13 (0.92%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0 (0.00%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eAltered mental status\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1408 (99.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e325 (89.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e14.4 [7.49\u0026ndash;30.1]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003csup\u003e***\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e11 (0.78%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e37 (10.2%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"13\" rowspan=\"14\"\u003e \u003cp\u003eUnderlying Disease\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eCardiovascular disease\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1161 (81.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e277 (76.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1.38 [1.04\u0026ndash;1.82]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.027\u003csup\u003e*\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e258 (18.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e85 (23.5%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eRespiratory disease\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1352 (95.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e332 (91.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1.83 [1.15\u0026ndash;2.83]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.011\u003csup\u003e*\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e67 (4.72%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e30 (8.29%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eNeurological disease\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1324 (93.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e330 (91.2%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1.36 [0.88\u0026ndash;2.04]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.193\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e95 (6.69%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e32 (8.84%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eEndocrine disease\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1295 (91.3%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e321 (88.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1.34 [0.91\u0026ndash;1.93]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.157\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e124 (8.74%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e41 (11.3%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eKidney disease\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1386 (97.7%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e354 (97.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.96 [0.41\u0026ndash;2.01]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e33 (2.33%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e8 (2.21%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eBlood disorders\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1398 (98.5%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e354 (97.8%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1.52 [0.62\u0026ndash;3.36]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.455\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e21 (1.48%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e8 (2.21%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eOncological/immune diseases\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNo\u0026thinsp;=\u0026thinsp;0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1354 (95.4%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e344 (95.0%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e1.10 [0.62\u0026ndash;1.84]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e0.86\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eYes\u0026thinsp;=\u0026thinsp;1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e65 (4.58%)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e18 (4.97%)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"5\" rowspan=\"6\"\u003e \u003cp\u003eBlood Test Results\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eNEU(*10\u003csup\u003e9\u003c/sup\u003e/L)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6.88 (3.62)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e9.17 (5.00)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.14 [1.11\u0026ndash;1.17]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003csup\u003e***\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eLYM(*10\u003csup\u003e9\u003c/sup\u003e/L)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.35 (0.76)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.12 (0.60)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.54 [0.43\u0026ndash;0.67]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003csup\u003e***\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eNLR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6.43 (5.38)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e11.9 (12.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.09 [1.07\u0026ndash;1.10]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003csup\u003e***\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003ePLT(*10\u003csup\u003e9\u003c/sup\u003e/L)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e220 (62.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e213 (78.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.00 [1.00\u0026ndash;1.00]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.115\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eCRP(mg/L)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e22.3 (31.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e66.5 (68.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.02 [1.02\u0026ndash;1.02]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003csup\u003e***\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eCRP/PLT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.11 (0.17)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.36 (0.42)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e28.7 [16.8\u0026ndash;49.1]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003csup\u003e***\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eMean(SD) ; n(%), *P\u0026thinsp;\u0026lt;\u0026thinsp;0.05, **P\u0026thinsp;\u0026lt;\u0026thinsp;0.01, ***P\u0026thinsp;\u0026lt;\u0026thinsp;0.001.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eInteraction Analysis\u003c/h2\u003e \u003cp\u003eNoticing that there was a significant age difference between pneumonia and non-pneumonia patients in the training cohort (mean age 62.2 vs.39.2), which is consistent with clinical practice, we introduced interaction terms of age with six other variables that showed significant differences between the groups. This is because elderly CAP patients often have insidious onset of illness and are prone to neglecting their condition. We used the random forest algorithm to model and evaluate this strategy, and found that the model performance was quite satisfactory (AUC\u003csub\u003einteraction\u003c/sub\u003e=0.80). The six interaction terms with age that we constructed all had higher importance rankings than the original features, indicating the presence of interactions (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Given the vigilance required for the elderly, when integrating the β model, in addition to using the probability calculated from the α model as a new factor and incorporating laboratory test-derived variables, we once again included age as a predictor.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eModel Performance Evaluation\u003c/h2\u003e \u003cdiv id=\"Sec18\" class=\"Section3\"\u003e \u003ch2\u003eDiscrimination Performance\u003c/h2\u003e \u003cp\u003e\u0026ldquo;Whether diagnosed with pneumonia\u0026rdquo; as a binary classification prediction variable, its discrimination performance of the model was measured using the receiver operating characteristic curve (ROC) and the area under the curve (AUC), also known as the C-index. In the performance evaluation of the α model, comparison of six algorithms using 10-fold cross-validation suggested that five machine learning algorithms were superior to traditional Logistic regression, indicating practical value in selecting complex machine learning algorithms (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). In the comparison of the optimal models of the six algorithms using the entire dataset as the internal validation set, the random forest model had the highest AUC (AUC\u003csub\u003erf\u003c/sub\u003e=0.94, 95%CI 0.93\u0026ndash;0.95), followed by the AdaBoost, XGBoost, and CatBoost models (AUC\u003csub\u003eada\u003c/sub\u003e=0.85, 95%CI 0.83\u0026ndash;0.87; AUC\u003csub\u003exgb\u003c/sub\u003e=0.82, 95%CI 0.80\u0026ndash;0.85; AUC\u003csub\u003ecat\u003c/sub\u003e=0.80, 95%CI 0.77\u0026ndash;0.83). In the external validation set that was temporally and spatially independent, the six models were compared again, with the CatBoost model performing the best (AUC\u003csub\u003ecat\u003c/sub\u003e=0.80, 95%CI 0.71\u0026ndash;0.87). The difference in AUC (|ΔAUC|) between internal and external validation was used to measure the robustness of the models, with CatBoost being the best (|ΔAUC|\u003csub\u003ecat\u003c/sub\u003e\u0026lt;0.01), while randomforest was the worst (|ΔAUC|\u003csub\u003erf\u003c/sub\u003e=0.19). The Integrated Discrimination Improvement (IDI) was used to measure the improvement in the prediction probability of actual pneumonia patients by selecting the CatBoost algorithm (\u003cb\u003eFigure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e). After a comprehensive comparison, the CatBoost model was selected as the final α model.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe β model also employed the same six algorithms and conducted internal/external validation, among other processes. In the comparison of models built with different algorithms, the final CatBoost model emerged as the winner (AUC\u003csub\u003ein\u003c/sub\u003e=0.93, 95%CI 0.92\u0026ndash;0.95; AUC\u003csub\u003eex\u003c/sub\u003e=0.81, 95%CI 0.70\u0026ndash;0.90). When integrating the β model, the α model was incorporated into the modeling in the form of predicted probabilities as a new variable. As a result, the β model utilized the most clinical raw variables, covering multimodal clinical information including general information, medical records, and laboratory tests. The ROC curves of the α/β models were also compared in internal/external validation sets, with AUC calculated and DeLong tests conducted. It was found that the β model improved predictive performance in the training cohort compared to the α model (Z\u003csub\u003ein\u003c/sub\u003e=-13.91, P\u0026thinsp;\u0026lt;\u0026thinsp;0.001), while the difference was not significant in the external validation cohort (Z\u003csub\u003eex\u003c/sub\u003e=-0.32, P\u0026thinsp;=\u0026thinsp;0.748) (\u003cb\u003eFigure S2\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eWe further calculated the optimal cutoff and confusion matrices for the α/β models during internal and external validation. The cutoffs were selected based on the ROC curve, corresponding to the threshold that maximizes the sum of sensitivity and specificity. Compared to the α model, the β model had lower optimal cutoff during internal/external validation. However, the optimal cutoffs for the same model were similar across internal/external validation (Cutoff\u003csub\u003eα_in\u003c/sub\u003e=0.246, Cutoff\u003csub\u003eβ_in\u003c/sub\u003e=0.145; Cutoff\u003csub\u003eα_ex\u003c/sub\u003e=0.222, Cutoff\u003csub\u003eβ_ex\u003c/sub\u003e=0.151), reflecting the robustness of the α/β models under different epidemiological conditions. The final model is intended for screening pneumonia patients, and both the training and external validation cohorts were imbalanced datasets (i.e., the number of non-pneumonia cases was several times higher than that of pneumonia cases). Therefore, we focused on sensitivity, the Matthews Correlation Coefficient (MCC), and other indicators related to the confusion matrix. The Net Reclassification Improvement (NRI) analysis also indicated that the β model was superior (NRI\u003csub\u003eα/β_in\u003c/sub\u003e=0.20; NRI\u003csub\u003eα/β_ex\u003c/sub\u003e=0.33) (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of Performance based on the Confusion Matrix for the α/β Models\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eComparison\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eInternal Validation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e \u003cp\u003eExternal Validation\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eα_Model\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eβ_Model\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eα_Model\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eβ_Model\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTrue Positives\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e246\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e329\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTrue Negatives\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1163\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1099\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e120\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e147\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFalse Positives\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e256\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e320\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e39\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFalse Negatives\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e116\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNRIα/β(Cutoffβ)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e\\\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\\\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.33\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIDIα/β\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e\\\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\\\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.12\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMCC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.57\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.36\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSensitivity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSpecificity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.79\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.78\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBalanced Accuracy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKappa\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.44\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.32\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eCalibration Performance\u003c/h2\u003e \u003cp\u003eCalibration curves were drawn to assess the calibration performance.The closer the curve is to the diagonal line (from the bottom left to the top right), the better the calibration. The Brier score, which ranges from 0 to 1, is used to evaluate the accuracy of probability predictions, with lower values indicating better performance. During internal validation, the modeling processes of the α and β models using various algorithms were relatively close to the diagonal line. However, the random forest model in the β model exhibited overfitting (AUC\u003csub\u003erf\u003c/sub\u003e=1), making it impossible to draw a calibration curveas it would overlap with the diagonal line. During external validation, the α/β models built using the CatBoost algorithm performed the best. Comparing the α and β models, all curves were found to be closer to the diagonal line, indicating that the integrated modeling strategy indeed provided added value. Comparing the Brier scores of internal/external validation also suggested that the β model was superior (Brier\u003csub\u003eα_in\u003c/sub\u003e=0.13, Brier\u003csub\u003eβ_in\u003c/sub\u003e=0.08; Brier\u003csub\u003eα_ex\u003c/sub\u003e=0.11, Brier\u003csub\u003eβ_ex\u003c/sub\u003e=0.08) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eClinical Utility\u003c/h2\u003e \u003cp\u003eClinical decision curves were plotted to assess the clinical utility. By comparing the intersection relationships between the curves and two reference lines (treating all patients and treating no patients), the clinical benefit of applying the models was evaluated. During internal validation, the α/β models built by various algorithms were all above the reference lines across the entire probability range, indicating the potential clinical benefit of the models. The clinical net benefit value was obtained at the leftmost intersection point between the curve and the reference line (Net Benefit\u0026thinsp;=\u0026thinsp;0.200), which means that approximately 20 more patients out of every 100 in the training cohort could receive correct diagnosis and treatment. During external validation, the α/β models built using the CatBoost algorithm still maintained curves above the two reference lines for most of the probability range. The β models built by all algorithms improved the clinical decision curves compared to the α models. Based on the final α/β models, approximately 11.3 more patients out of every 100 in the external validation cohort could receive correct diagnosis and treatment (Net Benefit\u0026thinsp;=\u0026thinsp;0.113) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eModel Interpretation and Practical Application\u003c/h2\u003e \u003cp\u003ePrior to modeling, we had already used the randomforest to calculate the importance ranking of variables. After determining the final model, we conducted SHAP interpretability analysis on the predictors of the α model again, finding that the SHAP value of the interaction term between age and Tmax was the highest, indicating that it is the most important for the α model. The β model incorporates additional test results and adjusts the predicted probability of the known α model. Therefore, we first calculated the difference in predicted probability between the β model and the α model (Probability Delta). We then used restricted cubic spline regression (RCS) on the predictors of the β model to elucidate the nonlinear relationships. Since age was highly collinear with the predicted probability of the α model, it was not included in the regression (Fig.\u0026nbsp;7). From the RCS regression plot, it can be seen that NLR (neutrophil-to-lymphocyte ratio) and CRP/PLT (C-reactive protein to platelet count ratio) are positively correlated with the probability difference (F\u003csub\u003eNLR\u003c/sub\u003e=11.61, P\u0026thinsp;\u0026lt;\u0026thinsp;0.001; F\u003csub\u003eCRP/PLT\u003c/sub\u003e=-3.99, P\u0026thinsp;=\u0026thinsp;0.008). An increase in these two factors indicates an upregulation of CAP risk. Regression analysis of the predicted probabilities of the α model shows that the β model makes smaller adjustments to extreme predicted probabilities (close to 0%or 100%) but downregulates most of the predicted probabilities. Both α/β models have been transformed into online web-based calculators using the Shiny package in RStudio. By directly entering clinical features and calculating the probabilities before and during the fever clinical visit, patients can autonomously assess the CAP risk (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://xjbqxmbt.shinyapps.io/shiny_app_alpha/\u003c/span\u003e\u003cspan address=\"https://xjbqxmbt.shinyapps.io/shiny_app_alpha/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://xjbqxmbt.shinyapps.io/shiny_app_beta/\u003c/span\u003e\u003cspan address=\"https://xjbqxmbt.shinyapps.io/shiny_app_beta/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e).\u003c/span\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003eIdentification of CAP\u0026rsquo;s Clinical Subtypes\u003c/h2\u003e \u003cp\u003eAfter successfully developing the CAP risk prediction model, latent class analysis (LCA), an unsupervised machine learning method, was used to automatically distinguish clinical subtypes in 362 CAP patients only based on the probability distribution of clinical manifestations (Fig.\u0026nbsp;8). With setted latent classes number of three, LCA identified three clinically meaningful categories: \u003cb\u003eI.\u003c/b\u003e Class 1 (N\u0026thinsp;=\u0026thinsp;48) with more severe illness, presenting with altered mental status and dyspnea, which are easily detectable and clearly related to pulmonary infection. These patients could not be further classified by symptoms due to the mental problem. \u003cb\u003eII.\u003c/b\u003e Class 2 (N\u0026thinsp;=\u0026thinsp;22) presenting with cold feeling, headach༆body pain, nausea༆vomiting, abdominal pain, and diarrhea, consistent with the TCM concept of Cold Syndrome. \u003cb\u003eIII.\u003c/b\u003e Class 3 (N\u0026thinsp;=\u0026thinsp;292) presenting with heat feeling, fatigue, cough, pharyngeal discomfort, and nasal congestion, consistent with the TCM concept of Heat Syndrome. The predicted CAP risk probabilities were compared across the three classes. Class 1 had higher risk probabilities (Pred_α\u003csub\u003e1\u003c/sub\u003e\u0026thinsp;=\u0026thinsp;0.57\u0026thinsp;\u0026plusmn;\u0026thinsp;0.09; Pred_β\u003csub\u003e1\u003c/sub\u003e\u0026thinsp;=\u0026thinsp;0.82\u0026thinsp;\u0026plusmn;\u0026thinsp;0.18; W\u003csub\u003eα\u003c/sub\u003e=2851.5, P\u0026thinsp;\u0026lt;\u0026thinsp;0.001; W\u003csub\u003eβ\u003c/sub\u003e=2931, P\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Class 2/3 had similar risk (Pred_α\u003csub\u003e3\u003c/sub\u003e\u0026thinsp;=\u0026thinsp;0.38\u0026thinsp;\u0026plusmn;\u0026thinsp;0.18, Pred_β\u003csub\u003e3\u003c/sub\u003e\u0026thinsp;=\u0026thinsp;0.52\u0026thinsp;\u0026plusmn;\u0026thinsp;0.28; Pred_α2\u0026thinsp;=\u0026thinsp;0.33\u0026thinsp;\u0026plusmn;\u0026thinsp;0.20, Pred_β\u003csub\u003e2\u003c/sub\u003e\u0026thinsp;=\u0026thinsp;0.49\u0026thinsp;\u0026plusmn;\u0026thinsp;0.34; W\u003csub\u003eα\u003c/sub\u003e=2685, P\u0026thinsp;=\u0026thinsp;0.199; W\u003csub\u003eβ\u003c/sub\u003e=3078, P\u0026thinsp;=\u0026thinsp;0.745). This confirms that the model can not only identify CAP patients, but also distinguish disease severity based on risk probabilities.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eWe explored the risk factors for CAP within the training cohort based on a cross-sectional study design framework. Some variables that were significant before adjustment (TCM_XFJB, Headache \u0026amp; body pain, Abdominal pain, Cardiovascular disease, Respiratory disease) were not included in the α model, but somes may still be potential risk factors. For example, underlying cardiovascular and respiratory diseases imply poorer cardiopulmonary functional reserve, which is associated with worse prognosis when facing respiratory infections. Guidelines for heart failure and COPD also indicate that infection is an important trigger for the acute exacerbation of these underlying diseases, and the two may interact causally in the overall progression[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e][\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Our model is designed to screen high-risk CAP patients in large cohorts of respiratory/febrile patients, not for disease endpoints. Thus, assessing comorbidities after initial imaging confirmation is feasible.\u003c/p\u003e \u003cp\u003eIn this study, we aimed to develop a screening tool with as few features as possible, hence only seven clinical features with significant differences between groups were ultimately selected as predictors for the α model. Among them, altered mental status (OR\u0026thinsp;=\u0026thinsp;14, 95%CI 7.49\u0026ndash;30.1) and dyspnea (OR\u0026thinsp;=\u0026thinsp;4.56, 95%CI 2.33\u0026ndash;8.99) had very high odds ratios (OR). However, in the variable importance assessment based on randomforest and the SHAP analysis of the CatBoost model, the rankings of these two variables and their interaction terms with age were relatively low. In contrast, age (OR\u0026thinsp;=\u0026thinsp;1.06, 95%CI 1.05\u0026ndash;1.06), Tmax (OR\u0026thinsp;=\u0026thinsp;1.47, 95%CI 1.24\u0026ndash;1.75), and their interaction term were ranked higher. This indicates that machine learning algorithms do not rely on extreme clinical feature differences between groups, but can effectively capture the impact of common clinical features and their interactions on CAP risk. Therefore, even in external validation across different epidemic environments, the model still has good generalization ability. Clinically, altered mental status and dyspnea are key signs of severe CAP, easily alerting patients and doctors. However, mild symptoms make CAP risk hard to estimate, often delaying diagnosis and treatment, especially in elderly patients. Our model highlights the interaction between age and other factors, emphasizing the predictive value of mild symptoms like Tmax, days since onset of illness, cough, and pharyngeal discomfort in elderly patients. This helps dynamically assess and detect CAP early in this high-risk group.\u003c/p\u003e \u003cp\u003eWhen considering the predictors such as the number of days since the onset of illness (OR\u0026thinsp;=\u0026thinsp;1.04, 95% CI 1.02\u0026ndash;1.07), pharyngeal discomfort (OR\u0026thinsp;=\u0026thinsp;0.27, 95% CI 0.18\u0026ndash;0.39), and cough (OR\u0026thinsp;=\u0026thinsp;1.73, 95% CI 1.27\u0026ndash;2.34), in combination with other clinical features like headache༆body pain (OR\u0026thinsp;=\u0026thinsp;0.49, 95% CI 0.29\u0026ndash;0.78), abdominal pain (OR\u0026thinsp;=\u0026thinsp;3.41, 95% CI 1.07\u0026ndash;10.5), and the use of TCM_XFJB (OR\u0026thinsp;=\u0026thinsp;0.74, 95% CI 0.57\u0026ndash;0.95), the differences in clinical features between groups align with the natural course of respiratory infectious diseases caused by pathogens. Referring to guidelines for common pathogens such as COVID-19 and influenza[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e][\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], the progression from mild to moderate/severe/critical conditions is often marked by the migration of the infection from the upper respiratory tract to the lower. Latent class analysis (LCA) was conducted to further explore whether these clinical symptoms can distinguish patients (\u003cb\u003eSupplement\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). This transition is frequently described in traditional Chinese medicine theory as the progression from \"superficial syndromes\" (headache༆body pain, pharyngeal discomfort) to \"interior syndromes\" (cough, abdominal pain, dyspnea, altered mental status). This may explain why the use of TCM_XFJB (which are indicated for superficial syndromes) could be a potential protective factor against CAP. In other medication comparisons,no significant differences were found in the use of antipyretics,antibiotics,and antiviral drugs.This may reflect potential misuse,as these drugs can be obtained without clear diagnoses in community settings.Their preemptive use without confirmed indications does not reduce CAP risk.The single-center study design may introduce selection bias,and future multicenter studies are needed.\u003c/p\u003e \u003cp\u003eAs an exploration of CAP diagnosis based on clinical symptoms, we identified three clinical subtypes, two of which align with TCM\u0026rsquo;s cold and heat syndromes. These syndromes determine different TCM treatments: \u0026ldquo;Warming the Cold\u0026rdquo; or \u0026ldquo;Clearing the Heat\u0026rdquo;. However, there are no standard criteria to classify the two. Therefore, we used unsupervised learning (latent class analysis) without adding labels manually. Class 1 represented severe cases with noticeable altered mental status and dyspnea. Class 2 and Class 3, which accounted for the majority (86.7%) of pneumonia patients, had CT evidence of pneumonia but lacked severe pulmonary infection symptoms. Compared to current assessment tools like PSI and CURB-65 that rely heavily on extensive tests or significant abnormal signs, our model (α/β) effectively identified these patients and reflected disease severity through predicted probabilities.\u003c/p\u003e \u003cp\u003eIn the β model, we adjusted the predicted probabilities of the α model using NLR and CRP/PLT based on clinical experience. The NLR, calculated as the ratio of NEU to LYM in peripheral blood, is a biomarker that reflects both the innate immune response (mediated by NEU) and adaptive immunity (supported by LYM)[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e][\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e][\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. An elevated NLR is often associated with inflammation, tissue damage, and a systemic inflammatory response, as neutrophils increase and lymphocytes decrease during infection or immune suppression. CRP/PLT is a composite index proposed by us based on clinical practice. It serves as an initial diagnostic marker in studies on neonatal pneumonia and sepsis[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. However, research on adult pneumonia related to this index is rather limited and not very satisfactory[\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e][\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e][\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. CRP levels acutely rise in bacterial infections and severe viral infections, while PLT can significantly decrease due to inflammation or coagulation consumption, leading to a higher CRP/PLT. These two indicators, were chosen because they reflect the interplay of inflammation, immunity, and coagulation, which are central to the progression of respiratory infections from mild to severe[\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e][\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e][\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e][\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. In the β model, increases in NLR and CRP/PLT both imply an upregulation of the predicted CAP risk. According to the RCS regression, these two factors can bring about a maximum increase of approximately 10% and 40% in the predicted probability. In this study, only initial ancillary tests (complete blood count\u0026thinsp;+\u0026thinsp;CRP, chest CT) were used. Future work could develop more comprehensive dynamic prediction models incorporating imaging data and multiple follow-up test results to cover the entire disease course of respiratory infections.\u003c/p\u003e \u003cp\u003eThe final α/β model, as a convenient online CAP screening tool, utilized a limited number of clinical features.The use of machine learning, multimodal information, and fusion strategies proved successful. Given the complexity of interpreting machine learning algorithms, this study supplemented the analysis with SHAP and RCS regression to evaluate the contribution of each predictor. Additionally, LCA was conducted to explore the natural classification of CAP clinical subtypes, laying the groundwork for future research.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis study developed and externally validated two prediction models calculating the probability of confirmed CAP in a phased and stepwise manner. The final α/β models demonstrated satisfactory predictive performance and represent a new tool for assessing CAP risk in the population of febrile patients. The classification of CAP clinical subtypes corroborates the existing TCM experience in distinguishing Cold/Heat syndromes, providing support for future standardization.\u003c/p\u003e "},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eConflict of Interest\u003c/h2\u003e \u003cp\u003eAll authors have no unreported potential conflicts of interest.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eAll authors contributed to the study. G.Z. and C.T. wrote the manuscript. G.Z. and H.Y. collected clinical data, G.Z. and S.Y. conducted statistical analysis, X.X. and L.B. guided the methodology, and L.Q. explored clinical significance. All authors reviewed the manuscript.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets generated and/or analyzed during the current study are available from the corresponding authors upon reasonable request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eQuan, T. P. et al. Increasing burden of community-acquired pneumonia leading to hospitalisation, 1998\u0026ndash;2014. \u003cem\u003eThorax\u003c/em\u003e \u003cb\u003e71\u003c/b\u003e (6), 535\u0026ndash;542 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChalmers, J. D. et al. Severity assessment tools for predicting mortality in hospitalised patients with community-acquired pneumonia. Systematic review and meta-analysis. \u003cem\u003eThorax\u003c/em\u003e \u003cb\u003e65\u003c/b\u003e (10), 878\u0026ndash;883 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMetlay, J. P. et al. Diagnosis and Treatment of Adults with Community-acquired Pneumonia. An Official Clinical Practice Guideline of the American Thoracic Society and Infectious Diseases Society of America. \u003cem\u003eAm. J. Respir Crit. Care Med.\u003c/em\u003e \u003cb\u003e200\u003c/b\u003e (7), e45\u0026ndash;e67 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChinese Medical Doctor Association, Division of Emergency Physicians, Chinese Emergency Medicine Consortium. Beijing Association for Emergency Medicine. Clinical Practice Guidelines for Community-Acquired Pneumonia in Adult Emergency Patients 2024 Edition. Chinese Journal of Emergency Medicine,2025,3403300-317.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSimonetti, A. F., Viasus, D., Garcia-Vidal, C. \u0026amp; Carratal\u0026agrave;, J. Management of community-acquired pneumonia in older adults. \u003cem\u003eTher. Adv. Infect. Dis.\u003c/em\u003e \u003cb\u003e2\u003c/b\u003e (1), 3\u0026ndash;16 (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCill\u0026oacute;niz, C., Rodr\u0026iacute;guez-Hurtado, D. \u0026amp; Torres, A. Characteristics and Management of Community-Acquired Pneumonia in the Era of Global Aging. \u003cem\u003eMed. Sci. (Basel)\u003c/em\u003e. \u003cb\u003e6\u003c/b\u003e (2), 35 (2018). Published 2018 Apr 30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBradley, J. et al. Pneumonia Severity Index and CURB-65 Score Are Good Predictors of Mortality in Hospitalized Patients With SARS-CoV-2 Community-Acquired Pneumonia. \u003cem\u003eChest\u003c/em\u003e \u003cb\u003e161\u003c/b\u003e (4), 927\u0026ndash;936 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTuta-Quintero, E. et al. Comparison of performances between risk scores for predicting mortality at 30 days in patients with community acquired pneumonia. \u003cem\u003eBMC Infect. Dis.\u003c/em\u003e \u003cb\u003e24\u003c/b\u003e (1), 912 (2024). Published 2024 Sep 3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSHINDO, Y. et al. Comparison of severity scoring systems A-DROP and CURB-65 for community-acquired pneumonia. \u003cem\u003eRespirology\u003c/em\u003e \u003cb\u003e13\u003c/b\u003e, 731\u0026ndash;735 (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShao, J. et al. A multimodal integration pipeline for accurate diagnosis, pathogen identification, and prognosis prediction of pulmonary infections. \u003cem\u003eInnov. (Camb)\u003c/em\u003e. \u003cb\u003e5\u003c/b\u003e (4), 100648 (2024). Published 2024 May 22.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, Z. et al. Development and validation of machine learning-based prediction model for severe pneumonia: A multicenter cohort study. \u003cem\u003eHeliyon\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e (17), e37367 (2024). Published 2024 Sep 3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTong et al. Strategic Considerations for Strengthening the Construction of the Traditional Chinese Medicine Emergency Prevention and Control System for Emerging and Sudden Infectious Diseases in ChinaJ.Bulletin of the Chinese Academy of Sciences,2020,3509,1087\u0026ndash;1095 .\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, J. et al. Combination of Hua Shi Bai Du granule (Q-14) and standard care in the treatment of patients with coronavirus disease 2019 (COVID-19): A single-center, open-label, randomized controlled trial. \u003cem\u003ePhytomedicine\u003c/em\u003e \u003cb\u003e91\u003c/b\u003e, 153671 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDepartment of Internal Medicine,China Association of Chinese Medicine;Department of Pulmonary Diseases,China Association of Chinese Medicine. Department of Pulmonary Diseases,China Nationality Medicine Association;Yu Xueqing,Xie Yang,Li Jiansheng. Guidelines for the Diagnosis and Treatment of Community-Acquired Pneumonia in Traditional Chinese Medicine, Revised Edition 2018[J]. Journal of Traditional Chinese Medicine,2019,04350\u0026thinsp;\u0026ndash;\u0026thinsp;360.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShasha et al. The Incoming Influenza Season \u0026mdash; China, the United Kingdom, and the United States, 2021\u0026ndash;2022[J]. \u003cem\u003eChina CDC Wkly.\u003c/em\u003e \u003cb\u003e3\u003c/b\u003e (49), 1039\u0026ndash;1045 (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGe, J. The COVID-19 pandemic in China: from dynamic zero-COVID to current policy. \u003cem\u003eHerz\u003c/em\u003e \u003cb\u003e48\u003c/b\u003e, 226\u0026ndash;228 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJue et al. Trends of SARS-CoV-2 Infection in Sentinel Community-Based Surveillance After the Optimization of Prevention and Control Measures \u0026mdash; China, December 2022\u0026ndash;January 2023[J]. \u003cem\u003eChina CDC Wkly.\u003c/em\u003e \u003cb\u003e5\u003c/b\u003e (7), 159\u0026ndash;164 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChina, C. D. C. Reported Cases and Deaths of National Notifiable Infectious Diseases \u0026mdash; China, April 2023*[J]. \u003cem\u003eChina CDC Wkly.\u003c/em\u003e \u003cb\u003e5\u003c/b\u003e (33), 742\u0026ndash;743 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChinese Geriatrics Society,Division of Electrocardiology and Cardiac Function,Chinese Medical Doctor Association,Division of Cardiology,Expert Committee of the Heart Failure Center Alliance,Yang Jiefu. Expert Consensus on the Comprehensive Management of Patients with Worsening Chronic Heart Failure in China,2022. \u003cem\u003eJ. Chin. Circulation J.\u003c/em\u003e \u003cb\u003e373\u003c/b\u003e, 215\u0026ndash;225 (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eExpert Group on the Diagnosis and Treatment of Acute Exacerbation of Chronic Obstructive Pulmonary Disease. Chinese Expert Consensus on the Diagnosis and Treatment of Acute Exacerbation of Chronic Obstructive Pulmonary DiseaseRevised Edition 2023[J]. International Journal of Respiratory Diseases,2023,432132-149.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGeneral Office of the National Health Commission of the Peoples Republic of China\u0026amp;General Office of the National Administration of Traditional Chinese. Medicine of the Peoples Republic of China.2023. Diagnosis and Treatment Protocol for COVID-19. \u003cem\u003eInfect. Trial Version 10 Chin. Med.\u003c/em\u003e, \u003cb\u003e1802\u003c/b\u003e,161\u0026ndash;166 .\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNational Health Commission of the People's Republic of China \u0026amp; National Administration of Traditional Chinese Medicine. Diagnosis and Treatment Protocol for Influenza (2025 Edition). \u003cem\u003eChin. J. Ration. Drug Use\u003c/em\u003e. \u003cb\u003e22\u003c/b\u003e (02), 1\u0026ndash;7 (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCataudella, E. et al. Neutrophil-To-Lymphocyte Ratio: An Emerging Marker Predicting Prognosis in Elderly Adults with Community-Acquired Pneumonia. \u003cem\u003eJ. Am. Geriatr. Soc.\u003c/em\u003e \u003cb\u003e65\u003c/b\u003e (8), 1796\u0026ndash;1801 (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDrăgoescu, A. N. et al. Neutrophil to Lymphocyte Ratio (NLR)-A Useful Tool for the Prognosis of Sepsis in the ICU. \u003cem\u003eBiomedicines\u003c/em\u003e \u003cb\u003e10\u003c/b\u003e (1), 75 (2021). Published 2021 Dec 30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, X. et al. Predictive values of neutrophil-to-lymphocyte ratio on disease severity and mortality in COVID-19 patients: a systematic review and meta-analysis. \u003cem\u003eCrit. Care\u003c/em\u003e. \u003cb\u003e24\u003c/b\u003e, 647 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, X. et al. C-reactive protein to platelet ratio as an early biomarker in differentiating neonatal late-onset sepsis in neonates with pneumonia. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e15\u003c/b\u003e (1), 10760 (2025). Published 2025 Mar 28.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGe, S., Ma, Y., Xie, M., Qiao, T. \u0026amp; Zhou, J. The role of platelet to mean platelet volume ratio in the identification of adult-onset still's disease from sepsis. \u003cem\u003eClin. (Sao Paulo)\u003c/em\u003e. \u003cb\u003e76\u003c/b\u003e, e2307 (2021). Published 2021 Apr 16.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, Y. et al. Diagnostic Value and Prognostic Significance of Procalcitonin Combined with C-Reactive Protein in Patients with Bacterial Bloodstream Infection. \u003cem\u003eComput. Math. Methods Med.\u003c/em\u003e \u003cb\u003e2022\u003c/b\u003e, 6989229 (2022). Published 2022 Aug 11.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMirsaeidi, M. et al. Thrombocytopenia and thrombocytosis at time of hospitalization predict mortality in patients with community-acquired pneumonia. \u003cem\u003eChest\u003c/em\u003e \u003cb\u003e137\u003c/b\u003e (2), 416\u0026ndash;420 (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKombe Kombe, A. J. et al. The Role of Inflammation in the Pathogenesis of Viral Respiratory Infections. \u003cem\u003eMicroorganisms\u003c/em\u003e \u003cb\u003e12\u003c/b\u003e (12), 2526 (2024). Published 2024 Dec 7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, Y. \u0026amp; Tang, H. Aberrant coagulation causes a hyper-inflammatory response in severe influenza pneumonia. \u003cem\u003eCell. Mol. Immunol.\u003c/em\u003e \u003cb\u003e13\u003c/b\u003e (4), 432\u0026ndash;442 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTay, M. Z., Poh, C. M., R\u0026eacute;nia, L., MacAry, P. A. \u0026amp; Ng, L. F. P. The trinity of COVID-19: immunity, inflammation and intervention. \u003cem\u003eNat. Rev. Immunol.\u003c/em\u003e \u003cb\u003e20\u003c/b\u003e (6), 363\u0026ndash;374 (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBouwman, J. J., Visseren, F. L., Bosch, M. C., Bouter, K. P. \u0026amp; Diepersloot, R. J. Procoagulant and inflammatory response of virus-infected monocytes. \u003cem\u003eEur. J. Clin. Invest.\u003c/em\u003e \u003cb\u003e32\u003c/b\u003e (10), 759\u0026ndash;766 (2002).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Community-Acquired Pneumonia, Machine Learning, Risk Prediction Model, Traditional Chinese Medicine, Epidemic Environment","lastPublishedDoi":"10.21203/rs.3.rs-6768762/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6768762/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e \u003cp\u003eDiagnosing community-acquired pneumonia (CAP) relies on costly imaging, posing challenges in resource-limited settings. Traditional tools focus on diagnostic tests for clinicians rather than patient use. Additionally, classification of subtypes in traditional Chinese medicine (TCM) lacks criteria.\u003c/p\u003e\u003ch2\u003eMaterials and Methods\u003c/h2\u003e \u003cp\u003eWe developed a multimodal fusion model using machine learning algorithms and clinical variables from basic information, medical records, and lab tests to assess CAP risk in fever patients. The model integrates top-performing models via ensemble learning to predict pneumonia probability. We trained on 2,193 visits at Beijing Traditional Chinese Medicine Hospital\u0026rsquo;s fever clinic from Dec 2021 to Dec 2022, and validated on 300 visits from Jan to July 2024. Use unsupervised learning to classify subtypes.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eThe training cohort included 1,781 CAP and similar patients, with 210 in the external validation cohort. CAPs were diagnosed via chest CT. The α model, based on pre-visit medical records, performed well (AUC\u003csub\u003einternal\u003c/sub\u003e=0.80, 95%CI 0.77\u0026ndash;0.83; AUC\u003csub\u003eexternal\u003c/sub\u003e=0.80, 95%CI 0.71\u0026ndash;0.87). The β model added four lab indicators, optimizing performance (AUC\u003csub\u003einternal\u003c/sub\u003e=0.93, 95%CI 0.92\u0026ndash;0.95; AUC\u003csub\u003eexternal\u003c/sub\u003e=0.81, 95%CI 0.70\u0026ndash;0.90). Two models were developed into online calculators. Latent class analysis distinguished Cold/Heat syndrome as subtypes.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eTwo models performed good across epidemic environments. We provided low-cost, and accurate tools for staged identification of CAP.\u003c/p\u003e","manuscriptTitle":"Staged Identification of CAP in Fever Patients Across Epidemic Environments: Modeling \u0026amp;Validation","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-29 14:24:20","doi":"10.21203/rs.3.rs-6768762/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-07-29T16:04:56+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-07-19T03:26:23+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"127948247798475900397622438605166175346","date":"2025-06-30T01:35:39+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-06-22T06:46:27+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"126898687109376953178439260853731304786","date":"2025-06-21T14:08:32+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-06-20T07:30:55+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-06-20T07:28:49+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-06-04T07:58:00+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-05-31T11:26:46+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2025-05-31T11:23:09+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"979d752c-0777-4c58-adda-dd900485cd81","owner":[],"postedDate":"June 29th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":50410919,"name":"Health sciences/Medical research/Experimental models of disease"},{"id":50410920,"name":"Health sciences/Risk factors"},{"id":50410921,"name":"Health sciences/Diseases/Infectious diseases/Influenza virus"},{"id":50410922,"name":"Health sciences/Diseases/Respiratory tract diseases"}],"tags":[],"updatedAt":"2025-12-22T16:07:21+00:00","versionOfRecord":{"articleIdentity":"rs-6768762","link":"https://doi.org/10.1038/s41598-025-29689-6","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2025-12-18 15:58:40","publishedOnDateReadable":"December 18th, 2025"},"versionCreatedAt":"2025-06-29 14:24:20","video":"","vorDoi":"10.1038/s41598-025-29689-6","vorDoiUrl":"https://doi.org/10.1038/s41598-025-29689-6","workflowStages":[]},"version":"v1","identity":"rs-6768762","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6768762","identity":"rs-6768762","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00