Machine Learning-Based Cancer Prediction Using Complete Blood Count: A Retrospective Study on the Diagnostic Potential of Hematological Parameters in Lung Cancer Screening

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background Early detection of lung cancer is crucial for improving outcomes, yet existing screening methods are costly and limited in accessibility. This study evaluated the diagnostic potential of routine complete blood count (CBC) parameters combined with machine learning (ML) for lung cancer prediction. Design and methods : Data from 12,964 lung cancer patients and 169,703 healthy controls were retrospectively collected, including CBC, coagulation, and tumor marker results. After rigorous data preprocessing, multiple ML models were developed and validated, with CatBoost showing the best performance (AUC 95.84%, precision 92.53%, accuracy 89.81%, recall 86.60%, F1-score 89.46%). Results Feature importance analysis identified platelet distribution width (PDW), age, neutrophil percentage (NE%), and red blood cell count (RBC) as the most significant predictors. A reduced model using these four features retained high accuracy (AUC 94.93%), indicating their strong discriminative value. Compared to tumor markers and coagulation data, CBC-derived features alone were robust for lung cancer prediction. Conclusions Routine CBC parameters, paired with ML, may enable accurate and cost-effective lung cancer screening in retrospective, single-center data. Key features such as PDW, NE%, and RBC may serve as early diagnostic indicators. This approach offers a scalable solution for early cancer detection, particularly in resource-limited settings, and requires prospective, multi-center validation prior to clinical implementation.
Full text 155,730 characters · extracted from preprint-html · click to expand
Machine Learning-Based Cancer Prediction Using Complete Blood Count: A Retrospective Study on the Diagnostic Potential of Hematological Parameters in Lung Cancer Screening | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine Learning-Based Cancer Prediction Using Complete Blood Count: A Retrospective Study on the Diagnostic Potential of Hematological Parameters in Lung Cancer Screening Ting Liu, Yingxin Wang, Lin Wang, Hao Wang, Chunhui Yang This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7811809/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 8 You are reading this latest preprint version Abstract Background Early detection of lung cancer is crucial for improving outcomes, yet existing screening methods are costly and limited in accessibility. This study evaluated the diagnostic potential of routine complete blood count (CBC) parameters combined with machine learning (ML) for lung cancer prediction. Design and methods : Data from 12,964 lung cancer patients and 169,703 healthy controls were retrospectively collected, including CBC, coagulation, and tumor marker results. After rigorous data preprocessing, multiple ML models were developed and validated, with CatBoost showing the best performance (AUC 95.84%, precision 92.53%, accuracy 89.81%, recall 86.60%, F1-score 89.46%). Results Feature importance analysis identified platelet distribution width (PDW), age, neutrophil percentage (NE%), and red blood cell count (RBC) as the most significant predictors. A reduced model using these four features retained high accuracy (AUC 94.93%), indicating their strong discriminative value. Compared to tumor markers and coagulation data, CBC-derived features alone were robust for lung cancer prediction. Conclusions Routine CBC parameters, paired with ML, may enable accurate and cost-effective lung cancer screening in retrospective, single-center data. Key features such as PDW, NE%, and RBC may serve as early diagnostic indicators. This approach offers a scalable solution for early cancer detection, particularly in resource-limited settings, and requires prospective, multi-center validation prior to clinical implementation. Complete blood count Lung cancer Machine learning Early diagnosis Platelet distribution width Neutrophil percentage Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1. Introduction The Complete Blood Count (CBC) is a crucial diagnostic tool widely used in medical practice for the quantitative assessment of blood cellular components [ 1 ]. This analysis aids in diagnosing infections, anemia, and blood disorders while providing a comprehensive overview of a patient's health status [ 2 ]. Cancer poses a significant public health challenge globally, particularly in China, where both incidence and mortality rates are rising, threatening public health. According to the "China Cancer Report," China records over 4 million new cancer cases annually, with approximately 3 million cancer-related deaths, making cancer a leading cause of mortality. High-incidence cancers in China, such as lung, gastric, liver, and colorectal cancers, rank among the highest worldwide in terms of incidence and mortality rates [ 3 , 4 ]. Early diagnosis is pivotal for enhancing cancer patient survival rates [ 5 ]. Current screening methods, primarily imaging tests like computed tomography (CT) and magnetic resonance imaging (MRI) and tumor marker detection such as carcinoembryonic antigen (CEA) and neuron-specific enolase‌ (NSE), are limited [6,7]. These tumor markers are often specific to certain cancer types, reducing their predictive utility. Additionally, the high cost and limited accessibility of these methods restrict their use in primary healthcare. As societal advancements, shifts in disease patterns, and heightened health awareness evolve, traditional biomarkers fall short of meeting clinical demands [8]. Consequently, there is an urgent need to develop new, cost-effective, and broadly applicable cancer prediction tools to address the complex challenges of clinical practice. CBC testing offers notable advantages, including cost-effectiveness, non-invasiveness, and broad applicability. As a routine diagnostic tool, its data coverage surpasses that of other screening methods [ 9 ]. Fully leveraging latent information in CBC data to develop a cancer prediction model could significantly enhance large-scale cancer screening [ 10 ]. Recent studies indicate that certain CBC parameters not only reflect basic physiological status but also correlate with the development and progression of various chronic diseases and cancers. For instance, CBC-derived inflammatory biomarkers have been linked to increased all-cause and respiratory disease mortality in adults with asthma [ 11 ]. Furthermore, CBC data can distinguish between idiopathic pulmonary fibrosis and connective tissue-related diseases and predict their severity with accuracy comparable to or exceeding current clinical practices [ 12 ]. An elevated platelet count has been associated with increased cancer risk at multiple sites. PDW indicates platelet heterogeneity and is closely related to angiogenesis and inflammation, while RBC distribution width (RDW), reflecting RBC volume variability, correlates with chronic inflammation, cardiovascular diseases, and the prognosis of certain malignancies [ 13 – 15 ]. Historically, CBC testing has concentrated on numerical changes in individual parameters, frequently neglecting the relationships among these indicators. Advances in machine learning now facilitate multi-dimensional data analysis, allowing for deeper insights from CBC parameters and offering novel perspectives for early disease screening. In clinical practice, efficient, cost-effective, and reliable diagnostic tools are essential for accurate disease diagnosis. Models should utilize readily available clinical indicators that are both quick and inexpensive, enabling rapid and precise patient diagnosis. Machine learning models can process multidimensional data, identify both linear and nonlinear relationships, and self-optimize through training to deliver accurate predictions, thereby aiding clinicians in making more informed decisions. Moreover, the interpretability of machine learning models helps identify critical information within the data, facilitating the discovery of key diagnostic indicators and enabling the extraction of precise diagnostic insights from clinical data [ 2 , 16 , 17 ]. Consequently, machine learning technology is increasingly adopted in the medical field [ 18 , 19 ]. Data quality is vital for the credibility of machine learning models. Reliable foundational data enhances training effectiveness and prediction accuracy. Clinical data often suffers from missing values, duplicate records, incorrect annotations, and formatting inconsistencies. Data cleaning is essential in machine learning applications for clinical data analysis, involving a thorough review to resolve these issues and ensure data accuracy and consistency [ 20 ]. This process establishes a solid foundation for machine learning applications, enabling the extraction and utilization of valuable information from large, complex clinical datasets. Ensuring model accuracy and credibility allows for feature ranking, exploration of core disease indicators, and the development of efficient diagnostic approaches. 2. Materials and methods 2.1. Study design and patient selection This retrospective cohort study, approved by the Ethics Committee of the Second Affiliated Hospital of Dalian Medical University (IRB No. KY2025-137-01), received an IRB-granted waiver of written informed consent because all analyses used de-identified historical data and posed no more than minimal risk to participants, in accordance with the Declaration of Helsinki and the Measures for the Ethical Review of Biomedical Research Involving Humans (NHFPC, 2016). The study involves both outpatients and inpatients, CBC, coagulation tests, and tumor marker test data from the Laboratory Information System of the Second Affiliated Hospital of Dalian Medical University from 2015 to 2024. Disease diagnoses are coded according to the National Clinical Edition 2.0 of the International Classification of Diseases. The initial laboratory test results for each patient are matched with their disease diagnoses to develop a data model for analysis. The study seeks to identify correlations between abnormal laboratory results and disease diagnoses, thereby exploring the potential utility of laboratory data in early disease diagnosis. 2.2. Identification of research variables and data collection The baseline model incorporated gender, age, and 24 CBC parameters. Correlation analysis identified strong associations among several feature groups, including white blood cell count (WBC) with absolute neutrophil count (NE#), NE% with lymphocyte percentage (LY%), and eosinophil percentage (EO%) with absolute eosinophil count (EO#). Based on the research significance of specific features in similar studies, 11 variables were excluded: hematocrit (HCT), LY%, mean platelet volume (MPV), plateletcrit (PCT), hemoglobin (HGB), WBC, mean corpuscular hemoglobin (MCH), absolute basophil count (BASO#), EO#, NE#, and RDW-coefficient of variation (RDW-CV). Statistical verification confirmed correlations among the remaining features. While some CBC test results had missing values, missing data were more common in coagulation tests and tumor marker tests. Samples with over half of their features missing were excluded. For tumor marker data, initial feature selection was based on the distribution of missing values. For lung cancer patients, NSE and cytokeratin-19 fragment (CYFRA 21 − 1) were retained. Missing data were imputed with random values within the medical reference range. Subsequently, an equal number of samples were randomly selected from the health check-up group to match the lung cancer cases, forming the final dataset for analysis. 2.3. Data transformation and normalization The dataset presented three main issues: missing data, duplicate records, and inconsistent formatting. As machine learning models require complete and reliable input, these problems were addressed systematically. For missing data, common interpolation techniques (e.g., multiple imputation, expectation-maximization) assume completely random missingness, which is often not the case in medical datasets and may introduce error. In this study, if more than 80% of values for a feature within a disease group were missing, or if more than half of the features for an individual sample were absent, the feature or sample was excluded. Remaining missing values were imputed with random values within the corresponding clinical reference ranges. Duplicate records mainly originated from repeated testing of hospitalized patients; in such cases, only the first test after admission was retained. Formatting issues included special characters in test results, which were standardized to boundary values, and instrument-generated invalid entries, which were converted into missing values and imputed. For unstructured diagnostic text, irregular formatting, ambiguous diagnoses, and uncoded diseases were addressed by a unified preprocessing step, with diagnoses mapped to the National Clinical Version 2.0 Disease Codes. All preprocessing was implemented in SAS, which enabled batch data cleaning, efficient storage, and optimized extraction, ensuring reliable datasets for subsequent analysis. 2.4. Model construction and feature selection To reduce potential overfitting caused by limited sample size, the RF model was first tested to evaluate the predictive value of CBC data, followed by the addition of coagulation data and then tumor markers. Tumor markers, as important indicators for cancer diagnosis, were also used to validate predictive accuracy. Feature importance analysis further identified variables with greater predictive power than tumor markers, providing insight into valuable predictors.Six models—LR, SVM, XGB, LGBM, CART, and CatBoost—were compared. Model evaluation used a 70:30 train-test split with control samples matched to the experimental group and 10-fold cross-validation. To quantify the role of individual features, permutation feature importance (PFI) was applied, ranking features by their contribution to model performance. The final feature subset was determined by aggregating importance rankings across models.All procedures were implemented using the ranger package in R and Scikit-Learn in Python. 2.5. Statistical analysis Continuous variables were tested for normality using the Shapiro test and described as mean ± standard deviation or median with interquartile range, as appropriate. Categorical variables were expressed as percentages. Group differences were assessed using Chi-square tests for categorical variables and t-tests or Kruskal–Wallis tests for continuous variables, depending on distribution. To account for covariates in model development, correlation analyses were performed: Spearman correlation for continuous variables, Cramer’s V for categorical variables, and point-biserial correlation for continuous–categorical pairs. Statistical analyses were conducted in RStudio version 4.2.3, with p < 0.05 considered statistically significant. 3. Result 3.1. Clinical characteristics of patients The dataset included 791,867 visits, comprising 169,703 physical examination cases and 622,164 disease diagnoses. For modeling, 12,964 lung cancer patients (malignant tumors of the bronchi and lungs) were selected, with 2,757 breast cancer patients used for robustness validation. Based on physical examination data (Table 1 ), participants were divided into a healthy group (169,703) and a lung cancer group (12,964), including 88,417 males (48.40%) and 94,250 females (51.60%). Lung cancer patients were older, and significant differences were observed in most CBC, coagulation, and tumor marker parameters compared with the healthy group (Fig. 1 ). Table 1 Comparison of Features between C34 Group and Healthy Group Feature Name Level Abbreviation N C34 Mean ± Standard Deviation Healthy Group Mean ± Standard Deviation p-value Gender Male Gender 88417 (48.40%) 6098 (47.04%) 82319 (48.51%) 0 Female Gender 94250 (51.60%) 6866 (52.96%) 87384 (51.49%) 0 Age [0,10) Age 11 (0.01%) 0 (0.00%) 11 (0.01%) 0 [10,20) Age 3518 (1.93%) 9 (0.07%) 3509 (2.07%) 0 [20,30) Age 21555 (11.80%) 83 (0.64%) 21472 (12.65%) 0 [30,40) Age 43331 (23.72%) 336 (2.59%) 42995 (25.34%) 0 [40,50) Age 37911 (20.76%) 917 (7.07%) 36994 (21.80%) 0 [50,60) Age 39551 (21.65%) 3186 (24.58%) 36365 (21.43%) 0 [60,70) Age 23307 (12.76%) 5065 (39.07%) 18242 (10.75%) 0 [70,80) Age 10311 (5.64%) 2978 (22.97%) 7333 (4.32%) 0.008 [80,90) Age 2878 (1.58%) 380 (2.93%) 2498 (1.47%) 0 [90,100] Age 286 (0.16%) 10 (0.08%) 276 (0.16%) 0 White Blood Cell Count WBC 182667 8.51 (3.99) 6.12 (1.81) 0 Neutrophil Percentage NE% 182667 72.79 (11.76) 57.03 (8.53) 0 Lymphocyte Percentage LY% 182667 19.01 (10.18) 33.33 (7.92) 0 Monocyte Percentage MO% 182667 6.37 (2.21) 7.02 (1.84) 0 Eosinophil Percentage EO% 182667 1.52 (2.14) 2.05 (1.78) 0 Basophil Percentage BA% 182667 0.31 (0.26) 0.57 (0.29) 0 Absolute Neutrophil Count NE# 182667 6.46 (3.83) 3.54 (1.41) 0 Absolute Lymphocyte Count LY# 182667 1.39 (0.60) 2.00 (0.71) 0 Absolute Monocyte Count MO# 182667 0.52 (0.25) 0.42 (0.14) 0 Absolute Eosinophil Count EO# 182667 0.12 (0.41) 0.12 (0.12) 0 Absolute Basophil Count BA# 182667 0.02 (0.02) 0.03 (0.04) 0 Red Blood Cell Count RBC 182667 4.24 (0.58) 4.76 (0.52) 0 Hemoglobin HGB 182667 128.36 (18.51) 143.10 (17.08) 0 Hematocrit PCV 182667 38.68 (5.18) 42.80 (4.50) 0.003 Mean Corpuscular Volume MCV 182667 91.47 (4.75) 90.14 (4.68) 0.746 Mean Corpuscular Hemoglobin MCH 182667 30.33 (1.90) 30.11 (2.02) 0.382 Mean Corpuscular Hemoglobin Concentration MCHC 182667 331.44 (9.36) 333.89 (11.05) 0 Red Cell Distribution Width-Coefficient of Variation RDW-CV 182667 13.09 (1.26) 12.68 (1.12) 0 Red Cell Distribution Width-Standard Deviation RDW-SD 182667 43.53 (4.03) 41.60 (3.40) 0 Platelet Count PLT 182667 230.55 (79.52) 245.46 (59.14) 0.035 Mean Platelet Volume MPV 182667 14.92 (2.25) 11.61 (1.94) 0.781 Plateletcrit PCT 182667 9.88 (1.06) 10.10 (0.87) 0.278 Platelet Distribution Width PDW 182667 0.22 (0.07) 0.25 (0.06) 0 Large Platelet Ratio P-LCR 182667 24.94 (7.47) 25.77 (6.95) 0 Thrombin Time TT 32655 12.83 (0.91) 12.74 (1.35) p_value Prothrombin Time Activity Percentage PTActivity% 32655 111.92 (16.74) 111.03 (14.52) 0 International Normalized Ratio INR 32655 0.96 (0.07) 0.97 (0.10) 0 International Normalized Ratio INR2 32655 0.96 (0.09) 0.96 (0.15) 0 Fibrinogen FIB 37001 3.78 (1.31) 3.41 (0.91) 0 Activated Partial Thromboplastin Time APTT 32529 35.36 (4.31) 35.07 (4.47) 0 Thrombin Time TT 32529 17.80 (1.03) 17.23 (5.07) 0 D-Dimer D-Di 9319 0.97 (1.74) 0.89 (1.69) 0 Fibrin Degradation Products FDP 151 5.15 (2.80) 5.56 (14.61) 0 Antithrombin III AT III 154 105.54 (25.07) 97.56 (12.77) 0 Neuron-specific enolase NSE 5017 19.44 (28.91) 15.02 (3.73) 0.008 Cytokeratin-19 fragment CYFRA 21 − 1 5073 4.77 (13.19) 1.85 (0.86) 0 3.2. Removing covariates and interpolating missing values The baseline model included gender, age, and 24 CBC parameters. Correlation analysis revealed strong associations among several features, leading to the exclusion of 11 variables (HCT, LY%, MPV, PCT, HGB, WBC, MCH, BASO#, EO#, NE#, RDW-CV). Correlation patterns among the remaining variables are shown in Fig. 2 . CBC, coagulation, and tumor marker data exhibited varying degrees of missingness. Samples with more than half of their features missing were excluded. For tumor markers, features were screened according to missingness, and NSE and CYFRA 21 − 1 were retained for lung cancer analysis. Remaining missing values were imputed within clinical reference ranges. Finally, an equal number of samples was randomly drawn from the physical examination group to match the lung cancer group, forming the dataset for subsequent analysis. 3.3. Model performance Seven machine learning models (RF, LR, SVM, XGB, LGBM, CART, and CatBoost) were trained in Python. Model performance was sequentially evaluated using CBC data alone, followed by the addition of coagulation parameters, tumor markers, and finally the full set of variables. Except for CART (Table 2 ), all models demonstrated favorable predictive ability; thus, CART was not subjected to further optimization. In addition to Table 2 , we visualized part of these results using a forest plot of AUC values with 95% confidence intervals (Fig. 3 ). This plot shows that the ensemble methods—XGBoost, LGBM, and RF—consistently achieved the highest AUC values (> 0.95) with narrow confidence intervals, indicating superior accuracy and stable performance. In contrast, CART yielded the lowest discriminative ability (AUC = 0.841, 95% CI: 0.836–0.847), whereas LR and SVM achieved moderately high AUC values (~ 0.94–0.95) but still lagged behind the ensemble models. As shown subsequently in Fig. 4 , the PR curves further indicate stable predictive performance without evidence of overfitting. Table 2 The evaluation indicators for each model Model Data Type ROC_AUC Accuracy Precision Recall F1_score CART Blood routine 84.07 (83.58–84.79) 84.07 (83.58–84.79) 83.50 (82.10-84.89) 84.89 (84.09–85.78) 84.18 (83.67–84.81) CART Blood routine + Coagulation + Tumor markers 88.16 (87.50-89.42) 88.16 (87.50-89.42) 88.08 (86.84–89.09) 88.46 (87.45–89.85) 88.27 (87.67–89.47) CatBoost Blood routine 95.94 (95.49–96.20) 89.81 (89.44–90.17) 92.53 (90.95–93.58) 86.60 (85.94–87.76) 89.46 (89.05–89.78) CatBoost Blood routine + Coagulation + Tumor markers 98.28 (97.97–98.45) 93.62 (93.27–94.05) 95.49 (94.87–96.05) 91.67 (90.81–92.56) 93.54 (93.20-93.97) LGBM Blood routine 95.79 (95.36–96.11) 89.60 (88.99–90.13) 91.98 (90.33–93.09) 86.75 (85.67–87.57) 89.29 (88.55–89.74) LGBM Blood routine + Coagulation + Tumor markers 98.11 (97.77–98.26) 93.36 (92.89–93.67) 94.87 (94.02–95.57) 91.78 (91.12–92.21) 93.30 (92.82–93.65) LR Blood routine 94.34 (94.00-94.62) 88.02 (87.72–88.38) 90.84 (89.99–91.88) 84.55 (83.96–85.35) 87.58 (87.18–87.91) LR Blood routine + Coagulation + Tumor markers 96.00 (95.65–96.25) 90.11 (89.51–90.86) 92.59 (92.00-93.28) 87.35 (86.16–88.93) 89.89 (89.22–90.67) RF Blood routine 95.52 (95.02–95.85) 89.37 (88.75–89.58) 92.23 (90.85–93.71) 85.97 (85.07–87.15) 88.98 (88.19–89.26) RF Blood routine + Coagulation + Tumor markers 97.46 (96.95–97.65) 92.55 (92.09–93.08) 94.54 (93.59–95.37) 90.43 (89.29–91.20) 92.43 (91.97–92.95) SVM Blood routine 94.96 (94.53–95.27) 89.49 (89.04–90.01) 93.86 (92.74–95.15) 84.50 (83.55–85.54) 88.93 (88.65–89.30) SVM Blood routine + Coagulation + Tumor markers 96.88 (96.40–97.10) 91.63 (91.14–92.02) 94.66 (93.94–95.31) 88.36 (87.10-89.36) 91.40 (91.02–91.79) XGB Blood routine 95.47 (95.05–95.72) 89.25 (88.66–89.70) 91.39 (90.10-92.91) 86.66 (85.76–88.04) 88.96 (88.19–89.37) XGB Blood routine + Coagulation + Tumor markers 98.11 (97.11–98.32) 93.52 (93.05–94.07) 95.04 (94.26–95.73) 91.94 (89.39–92.53) 93.46 (92.53–93.97) Note: The model results obtained when the Data Type is 'Blood routine + Coagulation' or 'Blood routine + Tumor markers' are not presented in the table. 3.4. Construct and validate the feature subsets PFI analysis (Fig. 5 ) showed that PDW, age, NE%, and RBC were the most important features for lung cancer prediction. Using only these four features, CatBoost still achieved strong performance (AUC 94.9%; Table 3 , Fig. 6 ). These features also have clear clinical relevance: PDW reflects platelet activation and tumor-associated inflammation; age is a well-recognized non-modifiable risk factor; NE% indicates systemic inflammatory response; and RBC changes may reflect tumor-related hypoxia or altered erythropoiesis. Their consistent importance across models highlights biological plausibility and supports their integration into simplified, cost-effective tools for early lung cancer screening in clinical practice. Table 3 The evaluation metrics for each model using top four PFI features Model ROC_AUC Accuracy Precision Recall F1_score CatBoost 94.93 (94.55–95.38) 88.05 (87.90-88.19) 89.64 (89.51–89.77) 86.29 (86.17–86.40) 87.93 (87.81–88.06) LGBM 94.79 (94.47–95.20) 87.71 (87.22–88.24) 89.32 (89.05–89.74) 86.25 (86.10-86.54) 87.79 (87.62–88.11) LR 93.10 (92.59-94.00) 86.09 (85.85–86.25) 87.43 (86.91–87.81) 84.52 (84.44–84.63) 85.95 (85.68–86.09) RF 94.28 (93.94–94.67) 87.57 (87.05–87.93) 89.29 (88.83–90.01) 85.58 (85.09–86.26) 87.39 (86.92–87.64) SVM 93.30 (92.68–93.75) 87.98 (87.79–88.28) 90.50 (90.21–90.71) 85.05 (84.25–85.65) 87.69 (87.36–88.05) XGB 94.32 (93.97–94.66) 87.65 (87.22–88.31) 88.98 (88.38–90.01) 86.16 (86.03–86.23) 87.54 (87.19–88.08) 4. Discussion In this study, we developed a machine learning framework for lung cancer prediction using routine laboratory data. Among the tested models, CatBoost achieved the best performance, with high accuracy in both internal validation and external testing. Using PFI, four key features (PDW, age, NE%, and RBC) were identified, and a simplified model with only these variables still maintained high predictive accuracy (AUC 94.9%). These findings suggest that routine and low-cost laboratory indicators can provide valuable support for early cancer screening. In clinical research, encountering missing and imbalanced data is inevitable in real-world studies. Missing data, if not properly addressed, can lead to information loss, biased estimation, and impaired model performance. For instance, a study on patient-specific MACE risk prediction highlighted that different imputation methods can significantly affect predictive accuracy, and improper handling of missing values undermines the model’s reliability [ 21 ]. Furthermore, the effectiveness of commonly used imputation techniques varies across different real-world scenarios, and selecting an appropriate method requires careful consideration of the dataset’s specific characteristics [ 22 ]. A systematic review examining machine learning-based clinical prediction models revealed that missing data is often poorly handled and inadequately reported [ 23 ]. Many studies either failed to disclose how missing values were treated or adopted simplistic strategies such as complete case analysis or mean imputation, both of which can lead to substantial information loss and reduced model accuracy. In this study, we explored the underlying causes of data missingness and implemented a feature-informed imputation strategy. By leveraging available clinical test information, we aimed to restore missing values in a manner that closely approximates their true distribution, thereby preserving the integrity and representativeness of the dataset. Similarly, class imbalance is a common phenomenon in medical diagnostic datasets. When not properly handled, it introduces systematic bias into model training, as machine learning algorithms tend to favor the majority class. This results in overly optimistic performance metrics that fail to generalize to minority populations. This challenge has been widely acknowledged in previous literature. For example, Tasci et al emphasized the impact of class imbalance in oncologic datasets and advocated for inclusive and bias-aware frameworks to ensure equitable predictive performance across all subgroups [ 24 ]. In a recent review, Wang et al demonstrated that the performance and reliability of deep learning models are significantly affected by imbalanced data distributions in medical imaging tasks [ 25 ]. Similarly, Zhang et al reviewed a decade of progress in addressing imbalance in medical datasets and highlighted how conventional machine learning methods often fail to detect rare or minority-class cases, thereby limiting diagnostic accuracy [ 26 ].] To mitigate this issue, we applied a random under-sampling approach to achieve class balance and implemented k-fold cross-validation to validate consistency and reliability. The predictive pipeline developed in this study also demonstrated scalability, supporting integration with external data sources such as genomic data, which could enable adaptive, multimodal models. Model interpretability is essential for promoting clinical trust and ensuring responsible deployment of AI systems in healthcare. Complex models such as ensemble methods often lack transparency, which can hinder their acceptance in clinical practice [ 27 ]. To address this, we applied PFI to identify key predictive variables in our CatBoost model. This method quantitatively ranks feature contributions and enhances the interpretability of model outputs [ 28 ]. Our analysis revealed that PDW, age, NE%, and RBC were the most influential features, supporting both the model’s clinical relevance and its practical applicability. Our interpretability analysis further reinforces the clinical utility of the model by highlighting four key features—PDW, age, NE%, and RBC—which not only demonstrated strong statistical importance but also have solid biological plausibility in the context of lung cancer. PDW reflects the variability in platelet size, which increases during platelet activation, a process linked to tumor-driven inflammation and angiogenesis. Elevated PDW levels have been reported in various malignancies and are associated with poor prognosis, likely due to the role of activated platelets in promoting tumor cell proliferation, metastasis, and immune evasion. The consistently highest importance score of PDW across multiple models suggests that it may serve as a sensitive hematologic biomarker for early cancer detection. Age is a well-recognized independent risk factor for lung cancer. As individuals age, cumulative exposure to carcinogens, decline in immune function, and accumulation of genetic mutations contribute to increased cancer susceptibility. Its strong predictive contribution in all models highlights its essential role in risk stratification and supports its integration into AI-based clinical tools. NE% is a marker of systemic inflammation and immune status. Elevated NE% is frequently observed in cancer patients and reflects tumor-induced neutrophilia, which can suppress anti-tumor immune responses and support tumor progression through the secretion of growth factors and proteases. Its inclusion in the top features aligns with current evidence linking systemic inflammatory markers with cancer prognosis. RBC count, although less commonly emphasized in oncology, may indicate chronic hypoxia, nutritional deficiency, or impaired erythropoiesis—conditions often associated with tumor burden and cancer-related metabolic changes. In this study, RBC demonstrated consistent relevance across models, suggesting that subtle hematologic shifts detectable via routine CBC may contribute valuable diagnostic signals. Collectively, these four features allow for the construction of a highly interpretable and simplified diagnostic model that maintains excellent predictive performance (AUC 94.93%, Fig. 5 ), thus offering a practical solution for early lung cancer detection. Importantly, these parameters are readily available, cost-effective, and routinely measured in clinical practice, supporting their potential integration into large-scale screening frameworks, particularly in resource-limited healthcare settings. Additionally, because consent was waived and only de-identified historical records were used, we could not perform patient-level re-contact or adjudicate outcomes beyond the available electronic records, which may leave residual misclassification unaddressed. This study has several limitations. First, only conventional machine learning models were evaluated; future work could explore deep learning approaches. Second, comorbidities may influence model performance in real-world settings, increasing the risk of misclassification. Third, although commonly available laboratory indicators provided strong predictive accuracy, the additional value of tumor markers was limited. Future research could incorporate broader clinical or molecular features to further enhance prediction. To mitigate the risk of spectrum and selection biases inherent to single-center retrospective datasets, future work should include pre-registered, prospective, multi-center cohorts with standardized CBC acquisition protocols and blinded adjudication. External validation across diverse clinical settings and laboratory analyzers will be essential to confirm transportability and support any downstream clinical deployment. Declarations Ethics approval and consent to participate This study was reviewed and approved by the Ethics Committee of the Second Affiliated Hospital of Dalian Medical University (IRB approval No. KY2025-137-01). Given the retrospective design and the use of fully de-identified data, the Institutional Review Board (IRB) waived the requirement for obtaining individual informed consent in accordance with the Declaration of Helsinki and the Measures for the Ethical Review of Biomedical Research Involving Humans (National Health and Family Planning Commission of the People’s Republic of China, 2016). The study posed no more than minimal risk to participants, and obtaining consent was impracticable in this context; therefore, the IRB determined that a consent waiver met the applicable regulatory criteria. Consent for publication Not applicable. Competing interests All authors have no additional conflicts of interest to declare. Author details 1 Department of Clinical Laboratory, The Second Hospital of Dalian Medical University, Dalian 116023, China. 2 School of statistics, Dongbei University of Finance and Economics, Dalian 116025, China Funding This work was supported by a grant from the Dalian Science and Technology Innovation Fund Program (No. 2024JJ13PT070) and United Foundation for Dalian Institute of Chemical Physics, Chinese Academy of Sciences and the Second Hospital of Dalian Medical University (No. DMU-2&DICP UN202410), Dalian Life and Health Field Guidance Program Project (No. 2024ZDJH01PT084). Author Contribution TL and YW conceived the study and performed initial discussions with LW, HW and CY. Afterward, TL, YW and LW designed the research protocols and coordinated the data collection and analyses. TL and YW carried out the data curation, preprocessing, and machine learning analyses. LW and HW supervised the statistical analysis and model validation. LW, HW and CY oversaw the clinical interpretation of the results. TL and YW wrote the first draft of the manuscript. LW, HW and CY critically reviewed and revised the manuscript for important intellectual content. All authors read and approved the final version of the manuscript. Acknowledgements Not applicable. Data Availability Data are available upon reasonable request to corresponding author. References Agarwal R, Sarkar A, Bhowmik A, Mukherjee D, Chakraborty S. A portable spinning disc for complete blood count (CBC). Biosens Bioelectron. 2020;150:111935. Luo G. MLBCD: a machine learning tool for big clinical data. Health Inf Sci Syst. 2015;3:3. Cao M, Li H, Sun D, Chen W. Cancer burden of major cancers in China: A need for sustainable actions. Cancer Commun (Lond). 2020;40(5):205–10. GBD. 2019 Colorectal Cancer Collaborators, Global, regional, and national burden of colorectal cancer and its risk factors, 1990–2019: a systematic analysis for the Global Burden of Disease Study 2019, Lancet Gastroenterol Hepatol 7(7) (2022) 627–647. Mazzone PJ, Silvestri GA, Souter LH, Caverly TJ, Kanne JP, Katki HA, et al. Screening for Lung Cancer: CHEST Guideline and Expert Panel Report. Chest. 2021;160(5):e427–94. Wang Y, Zhao Y, Li M, Hou H, Jian Z, Li W, et al. Conversion of primary liver cancer after targeted therapy for liver cancer combined with AFP-targeted CAR T-cell therapy: a case report. Front Immunol. 2023;14:1180001. Grunnet M, Sorensen JB. Carcinoembryonic antigen (CEA) as tumor marker in lung cancer. Lung Cancer 76(2) (2012) 138 – 43. 2019 Cancer Global Burden of Disease, Collaboration JM, Kocarnik K, Compton FE, Dean W, Fu BL, Gaw et al. Cancer Incidence, Mortality, Years of Life Lost, Years Lived With Disability, and Disability-Adjusted Life Years for 29 Cancer Groups From 2010 to 2019: A Systematic Analysis for the Global Burden of Disease Study 2019, JAMA Oncol 8(3) (2022) 420–444. Raess PW, van de Geijn GJ, Njo TL, Klop B, Sukhachev D, Wertheim G, et al. Automated screening for myelodysplastic syndromes through analysis of complete blood count and cell population data parameters. Am J Hematol. 2014;89(4):369–74. Boutault R, Peterlin P, Boubaya M, Sockel K, Chevallier P, Garnier A, et al. A novel complete blood count-based score to screen for myelodysplastic syndrome in cytopenic patients. Br J Haematol. 2018;183(5):736–46. Ke J, Qiu F, Fan W, Wei S. Associations of complete blood cell count-derived inflammatory biomarkers with asthma and mortality in adults: a population-based study. Front Immunol. 2023;14:1205687. Mueller AN, Miller HA, Taylor MJ, Suliman SA, Frieboes HB. Identification of Idiopathic Pulmonary Fibrosis and Prediction of Disease Severity via Machine Learning Analysis of Comprehensive Metabolic Panel and Complete Blood Count Data. Lung. 2024;202(2):139–50. Bongiovanni D, Han J, Klug M, Kirmes K, Viggiani G, von Scheidt M, et al. Role of Reticulated Platelets in Cardiovascular Disease. Arterioscler Thromb Vasc Biol. 2022;42(5):527–39. Kwon O, Ahn JH, Koh JS, Park Y, Hwang SJ, Tantry US, et al. Platelet-fibrin clot strength and platelet reactivity predicting cardiovascular events after percutaneous coronary interventions. Eur Heart J. 2024;45(25):2217–31. Giannakeas V, Kotsopoulos J, Cheung MC, Rosella L, Brooks JD, Lipscombe L, et al. Analysis of Platelet Count and New Cancer Diagnosis Over a 10-Year Period. JAMA Netw Open. 2022;5(1):e2141633. Sanchez-Pinto LN, Bennett TD. Evaluation of Machine Learning Models for Clinical Prediction Problems. Pediatr Crit Care Med. 2022;23(5):405–8. Stevens L, Kao D, Hall J, Görg C, Abdo K, Linstead E. A Preliminary Study of an Interactive Visual Analysis Tool Facilitating Clinical Applications of Machine Learning for Precision Medicine. Appl Sci (Basel). 2020;10(9):3309. Haymond S, Master SR. How Can We Ensure Reproducibility and Clinical Translation of Machine Learning Applications in Laboratory Medicine? Clin Chem. 2022;68(3):392–5. Tan YY, Rim TH, Ting DSJ, Hsieh YT, Kim TI. Editorial: Big data and artificial intelligence in ophthalmology - clinical application and future exploration. Front Med (Lausanne). 2023;10:1339280. Chirumbolo S, Berretta M, Tirelli U. Trust, trustworthiness and acceptability of a machine learning adoption in data-driven clinical decision support system. Some comments. Int J Med Inf. 2024;184:105374. Rios R, Miller RJH, Manral N, Sharir T, Einstein AJ, Fish MB, et al. Handling missing values in machine learning to predict patient-specific risk of adverse cardiac events: Insights from REFINE SPECT registry. Comput Biol Med. 2022;145:105449. Li J, Guo S, Ma R, He J, Zhang X, Rui D, et al. Comparison of the effects of imputation methods for missing data in predictive modelling of cohort study datasets. BMC Med Res Methodol. 2024;24(1):41. Nijman S, Leeuwenberg AM, Beekers I, Verkouter I, Jacobs J, Bots ML, et al. Missing data is poorly handled and reported in prediction model studies using machine learning: a literature review. J Clin Epidemiol. 2022;142:218–29. Tasci E, Zhuge Y, Camphausen K, Krauze AV. Cancers (Basel). 2022;14(12):2897. Bias and Class Imbalance in Oncologic Data-Towards Inclusive and Transferrable AI in Large Scale Oncology Data Sets. Zhang J, Xie Y, Wu Q, Xia Y. Medical image classification using synergic deep learning. Med Image Anal. 2019;54:10–9. Gao L, Zhang L, Liu C, Wu S. Handling imbalanced medical image data: A deep-learning-based one-class classification approach. Artif Intell Med. 2020;108:101935. Rudin C. Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead. Nat Mach Intell. 2017;1(5):206–15. Wang S, Liu Y, Wang W, Zhao G, Liang H. Interpretable machine learning guided by physical mechanisms reveals drivers of runoff under dynamic land use changes. J Environ Manage. 2024;367:121978. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 21 Dec, 2025 Reviewers agreed at journal 30 Nov, 2025 Reviewers agreed at journal 29 Oct, 2025 Reviewers invited by journal 29 Oct, 2025 Editor assigned by journal 19 Oct, 2025 Editor invited by journal 17 Oct, 2025 Submission checks completed at journal 17 Oct, 2025 First submitted to journal 17 Oct, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7811809","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":541876092,"identity":"d8f56027-9190-4042-9848-70e02df8e1b6","order_by":0,"name":"Ting Liu","email":"","orcid":"","institution":"The Second Hospital of Dalian Medical University","correspondingAuthor":false,"prefix":"","firstName":"Ting","middleName":"","lastName":"Liu","suffix":""},{"id":541876093,"identity":"1e0ae410-0ed0-40e9-9509-5640687ba956","order_by":1,"name":"Yingxin Wang","email":"","orcid":"","institution":"The Second Hospital of Dalian Medical University","correspondingAuthor":false,"prefix":"","firstName":"Yingxin","middleName":"","lastName":"Wang","suffix":""},{"id":541876094,"identity":"051c8a68-4065-4bfb-9623-bdd08ed7415d","order_by":2,"name":"Lin Wang","email":"","orcid":"","institution":"Dongbei University of Finance and Economics","correspondingAuthor":false,"prefix":"","firstName":"Lin","middleName":"","lastName":"Wang","suffix":""},{"id":541876095,"identity":"ebf3cf64-555a-48f5-9a9e-dcd0ab77b259","order_by":3,"name":"Hao Wang","email":"","orcid":"","institution":"Dongbei University of Finance and Economics","correspondingAuthor":false,"prefix":"","firstName":"Hao","middleName":"","lastName":"Wang","suffix":""},{"id":541876096,"identity":"c2a98d8b-73f5-4cd9-9d9b-70fb56afbc6d","order_by":4,"name":"Chunhui Yang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA10lEQVRIie3RsQrCMBCA4YRAu6S6xkH0EQKFIujDXBcdbDs7CGaQOrrat1CE4lgo2OXc+xJCRye1XRxj3Bzyww0H9y0JITbbX8YJgdVssyfQbcyQNDinmfqF0Gxb0mNhSmR1K5inGPOr5VyQ1TRU7q3QE0yAeRfHCfCeC4KLUPEE9KTmknnIeVDHuaBpGSrBpQFJhfAPHXkaEpqlUkrREWVABhjJ9pEBBN7PE7gu/JRHetKrULZf+YL+Lj7VzXo63LuoJ+OCuI/PBu042vu2kfp2YbPZbLY3z0FFqIgz33cAAAAASUVORK5CYII=","orcid":"","institution":"The Second Hospital of Dalian Medical University","correspondingAuthor":true,"prefix":"","firstName":"Chunhui","middleName":"","lastName":"Yang","suffix":""}],"badges":[],"createdAt":"2025-10-09 02:34:41","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7811809/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7811809/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":95655817,"identity":"eb59fbe8-5087-4558-aa33-9914120a69b0","added_by":"auto","created_at":"2025-11-11 16:17:00","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1479438,"visible":true,"origin":"","legend":"","description":"","filename":"Manuscript.docx","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/cebf1563c9f24f0351e33405.docx"},{"id":95655140,"identity":"1f2a080a-0f38-44a4-ae05-3d351bf726df","added_by":"auto","created_at":"2025-11-11 16:14:21","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6952,"visible":true,"origin":"","legend":"","description":"","filename":"b52d22e9b2a2450eb2a30c91b52160c3.json","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/bd2ab0dfeb38bae7494f019f.json"},{"id":95565211,"identity":"82ba0e44-02be-4cb3-9be3-37f8684ce8d5","added_by":"auto","created_at":"2025-11-10 16:14:32","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":126296,"visible":true,"origin":"","legend":"","description":"","filename":"b52d22e9b2a2450eb2a30c91b52160c31enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/0876d28fab25635d1b2746be.xml"},{"id":95654831,"identity":"3d8a7d4e-e16c-464f-9aae-b07c9c71bdee","added_by":"auto","created_at":"2025-11-11 16:13:17","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":85108,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/2ceddd0b31a53b4077cf0056.png"},{"id":95565214,"identity":"f80dc304-1e9e-4359-a36a-cab1253da43e","added_by":"auto","created_at":"2025-11-10 16:14:32","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":34952,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/14478ca8d824620d36abdc41.png"},{"id":95655568,"identity":"795fb33b-b5cc-4f7a-959b-c05f333928ee","added_by":"auto","created_at":"2025-11-11 16:16:29","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":18159,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/22b2062fbb9d6256ee1bae01.png"},{"id":95565215,"identity":"026e8360-4516-468c-82ed-cea3471ed821","added_by":"auto","created_at":"2025-11-10 16:14:32","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":48557,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/0f6cf7f70f53096dac2a6978.png"},{"id":95655259,"identity":"4ec122f5-4d24-401b-9309-48308041a73a","added_by":"auto","created_at":"2025-11-11 16:14:53","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":33115,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/3ca5da2b6a2485e56888dc08.png"},{"id":95655823,"identity":"0b48183d-9671-4f7d-89e7-fbe6ad631d39","added_by":"auto","created_at":"2025-11-11 16:17:01","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":51676,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/4925ec83bb23a289666a695f.png"},{"id":95655441,"identity":"ae0ab7f3-ab30-4c18-a1e7-d54170e21cd3","added_by":"auto","created_at":"2025-11-11 16:16:08","extension":"xml","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":124023,"visible":true,"origin":"","legend":"","description":"","filename":"b52d22e9b2a2450eb2a30c91b52160c31structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/8a4b72cbe7be8470d1e0a312.xml"},{"id":95655479,"identity":"e8beb6cd-f405-4b8f-8808-d91f911e0c5e","added_by":"auto","created_at":"2025-11-11 16:16:18","extension":"html","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":131269,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/5609843a4d12d266b75775f9.html"},{"id":95565210,"identity":"cc767178-cada-4d75-83ac-e9d188115e28","added_by":"auto","created_at":"2025-11-10 16:14:32","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":720837,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStudy workflow for machine learning-based lung cancer prediction using clinical laboratory data.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/f217027d383f00ab60cff5a9.png"},{"id":95565212,"identity":"131849b3-3978-404b-8dc3-48c3e354ed93","added_by":"auto","created_at":"2025-11-10 16:14:32","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":139906,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCorrelation heatmaps of CBC parameters before and after feature reduction.\u003c/strong\u003e\u003cbr\u003e\n(A) Heatmap and hierarchical clustering of correlations among all initial CBC features derived from 182,667 patients, showing multiple groups of highly correlated variables (e.g., WBC with NE#, NE% with LY%). (B) The Correlation Heatmap between features derived from the data of 14,108 patients.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/7437af18d636771fb4fce757.png"},{"id":95565209,"identity":"935228ea-d40e-4023-a3ba-87d039fade78","added_by":"auto","created_at":"2025-11-10 16:14:32","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":62228,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eForest plot of AUC scores with 95% confidence intervals (95% CI) for different machine learning models based on blood routine data. \u003c/strong\u003eEnsemble methods (XGBoost, LGBM, RF) achieved the highest AUC values (\u0026gt;0.95) with narrow confidence intervals, indicating superior and stable performance.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/51697b66a134cb14d30618ce.png"},{"id":95565222,"identity":"e51748d4-495c-403f-beac-5d12c9add86d","added_by":"auto","created_at":"2025-11-10 16:14:32","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":183726,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePrecision-recall curves for six machine learning models in distinguishing lung cancer from healthy controls based on CBC features.\u003c/strong\u003e\u003cbr\u003e\n(A–F) show PR curves for CatBoost, LGBM, LR, RF, SVM, and XGB models, respectively. For each model, performance is evaluated on training and validation datasets using area under the precision-recall curve (AUPRC) as a metric. All models demonstrated high precision and recall, with CatBoost and RF achieving the best overall AUPRC values. Minimal performance drop between training and validation curves suggests limited overfitting and robust model generalization.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/70dffbb8f073a8637d3730d0.png"},{"id":95565216,"identity":"0936375c-90d3-4cbf-99bc-6f9b7bf23053","added_by":"auto","created_at":"2025-11-10 16:14:32","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":121646,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFeature importance ranking for six machine learning models based on PFI.\u003c/strong\u003e\u003cbr\u003e\nBar plots display the mean PFI values for the top 10 features contributing to prediction accuracy for each model: (A) CatBoost, (B) LGBM, (C) LR, (D) RF, (E) SVM, and (F) XGB. Across all models, PDW, age, NE%, and RBC consistently rank as the most important features, highlighting their central role in lung cancer prediction using CBC data.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/117483f719969e28a39123ed.png"},{"id":95565223,"identity":"0623bb77-a675-4425-9377-91601614a4cc","added_by":"auto","created_at":"2025-11-10 16:14:32","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":191165,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePrecision-recall curves for six machine learning models using only the top four CBC features (PDW, age, NE%, and RBC).\u003c/strong\u003e(A–F) show PR curves for CatBoost, LGBM, LR, RF, SVM, and XGB, trained using only the four most important predictors. Despite substantial reduction in input features, models maintain high precision and recall, with CatBoost achieving an AUPRC of 0.95 on the validation set. This demonstrates that a simplified model based on routinely available CBC parameters can effectively identify lung cancer, supporting the feasibility of cost-effective, population-level screening.\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/6e41ccf0e8375a6b5642748a.png"},{"id":95660213,"identity":"3199a2e3-0354-4356-a6b2-65deacc8076c","added_by":"auto","created_at":"2025-11-11 16:31:08","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2702580,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7811809/v1/8de03843-dfa9-42ef-b42e-4d01f7725141.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Machine Learning-Based Cancer Prediction Using Complete Blood Count: A Retrospective Study on the Diagnostic Potential of Hematological Parameters in Lung Cancer Screening","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eThe Complete Blood Count (CBC) is a crucial diagnostic tool widely used in medical practice for the quantitative assessment of blood cellular components [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. This analysis aids in diagnosing infections, anemia, and blood disorders while providing a comprehensive overview of a patient's health status [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Cancer poses a significant public health challenge globally, particularly in China, where both incidence and mortality rates are rising, threatening public health. According to the \"China Cancer Report,\" China records over 4\u0026nbsp;million new cancer cases annually, with approximately 3\u0026nbsp;million cancer-related deaths, making cancer a leading cause of mortality. High-incidence cancers in China, such as lung, gastric, liver, and colorectal cancers, rank among the highest worldwide in terms of incidence and mortality rates [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Early diagnosis is pivotal for enhancing cancer patient survival rates [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Current screening methods, primarily imaging tests like computed tomography (CT) and magnetic resonance imaging (MRI) and tumor marker detection such as carcinoembryonic antigen (CEA) and neuron-specific enolase\u0026zwnj; (NSE), are limited [6,7]. These tumor markers are often specific to certain cancer types, reducing their predictive utility. Additionally, the high cost and limited accessibility of these methods restrict their use in primary healthcare. As societal advancements, shifts in disease patterns, and heightened health awareness evolve, traditional biomarkers fall short of meeting clinical demands [8]. Consequently, there is an urgent need to develop new, cost-effective, and broadly applicable cancer prediction tools to address the complex challenges of clinical practice.\u003c/p\u003e\u003cp\u003eCBC testing offers notable advantages, including cost-effectiveness, non-invasiveness, and broad applicability. As a routine diagnostic tool, its data coverage surpasses that of other screening methods [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Fully leveraging latent information in CBC data to develop a cancer prediction model could significantly enhance large-scale cancer screening [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Recent studies indicate that certain CBC parameters not only reflect basic physiological status but also correlate with the development and progression of various chronic diseases and cancers. For instance, CBC-derived inflammatory biomarkers have been linked to increased all-cause and respiratory disease mortality in adults with asthma [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. Furthermore, CBC data can distinguish between idiopathic pulmonary fibrosis and connective tissue-related diseases and predict their severity with accuracy comparable to or exceeding current clinical practices [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. An elevated platelet count has been associated with increased cancer risk at multiple sites. PDW indicates platelet heterogeneity and is closely related to angiogenesis and inflammation, while RBC distribution width (RDW), reflecting RBC volume variability, correlates with chronic inflammation, cardiovascular diseases, and the prognosis of certain malignancies [\u003cspan additionalcitationids=\"CR14\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eHistorically, CBC testing has concentrated on numerical changes in individual parameters, frequently neglecting the relationships among these indicators. Advances in machine learning now facilitate multi-dimensional data analysis, allowing for deeper insights from CBC parameters and offering novel perspectives for early disease screening. In clinical practice, efficient, cost-effective, and reliable diagnostic tools are essential for accurate disease diagnosis. Models should utilize readily available clinical indicators that are both quick and inexpensive, enabling rapid and precise patient diagnosis. Machine learning models can process multidimensional data, identify both linear and nonlinear relationships, and self-optimize through training to deliver accurate predictions, thereby aiding clinicians in making more informed decisions. Moreover, the interpretability of machine learning models helps identify critical information within the data, facilitating the discovery of key diagnostic indicators and enabling the extraction of precise diagnostic insights from clinical data [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. Consequently, machine learning technology is increasingly adopted in the medical field [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Data quality is vital for the credibility of machine learning models. Reliable foundational data enhances training effectiveness and prediction accuracy. Clinical data often suffers from missing values, duplicate records, incorrect annotations, and formatting inconsistencies. Data cleaning is essential in machine learning applications for clinical data analysis, involving a thorough review to resolve these issues and ensure data accuracy and consistency [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. This process establishes a solid foundation for machine learning applications, enabling the extraction and utilization of valuable information from large, complex clinical datasets. Ensuring model accuracy and credibility allows for feature ranking, exploration of core disease indicators, and the development of efficient diagnostic approaches.\u003c/p\u003e"},{"header":"2. Materials and methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.1. Study design and patient selection\u003c/h2\u003e\u003cp\u003e This retrospective cohort study, approved by the Ethics Committee of the Second Affiliated Hospital of Dalian Medical University (IRB No. KY2025-137-01), received an IRB-granted waiver of written informed consent because all analyses used de-identified historical data and posed no more than minimal risk to participants, in accordance with the Declaration of Helsinki and the Measures for the Ethical Review of Biomedical Research Involving Humans (NHFPC, 2016). The study involves both outpatients and inpatients, CBC, coagulation tests, and tumor marker test data from the Laboratory Information System of the Second Affiliated Hospital of Dalian Medical University from 2015 to 2024. Disease diagnoses are coded according to the National Clinical Edition 2.0 of the International Classification of Diseases. The initial laboratory test results for each patient are matched with their disease diagnoses to develop a data model for analysis. The study seeks to identify correlations between abnormal laboratory results and disease diagnoses, thereby exploring the potential utility of laboratory data in early disease diagnosis.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003ch2\u003e2.2. Identification of research variables and data collection\u003c/h2\u003e\u003cp\u003eThe baseline model incorporated gender, age, and 24 CBC parameters. Correlation analysis identified strong associations among several feature groups, including white blood cell count (WBC) with absolute neutrophil count (NE#), NE% with lymphocyte percentage (LY%), and eosinophil percentage (EO%) with absolute eosinophil count (EO#). Based on the research significance of specific features in similar studies, 11 variables were excluded: hematocrit (HCT), LY%, mean platelet volume (MPV), plateletcrit (PCT), hemoglobin (HGB), WBC, mean corpuscular hemoglobin (MCH), absolute basophil count (BASO#), EO#, NE#, and RDW-coefficient of variation (RDW-CV). Statistical verification confirmed correlations among the remaining features. While some CBC test results had missing values, missing data were more common in coagulation tests and tumor marker tests. Samples with over half of their features missing were excluded. For tumor marker data, initial feature selection was based on the distribution of missing values. For lung cancer patients, NSE and cytokeratin-19 fragment (CYFRA 21\u0026thinsp;\u0026minus;\u0026thinsp;1) were retained. Missing data were imputed with random values within the medical reference range. Subsequently, an equal number of samples were randomly selected from the health check-up group to match the lung cancer cases, forming the final dataset for analysis.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\u003ch2\u003e2.3. Data transformation and normalization\u003c/h2\u003e\u003cp\u003eThe dataset presented three main issues: missing data, duplicate records, and inconsistent formatting. As machine learning models require complete and reliable input, these problems were addressed systematically. For missing data, common interpolation techniques (e.g., multiple imputation, expectation-maximization) assume completely random missingness, which is often not the case in medical datasets and may introduce error. In this study, if more than 80% of values for a feature within a disease group were missing, or if more than half of the features for an individual sample were absent, the feature or sample was excluded. Remaining missing values were imputed with random values within the corresponding clinical reference ranges.\u003c/p\u003e\u003cp\u003eDuplicate records mainly originated from repeated testing of hospitalized patients; in such cases, only the first test after admission was retained. Formatting issues included special characters in test results, which were standardized to boundary values, and instrument-generated invalid entries, which were converted into missing values and imputed. For unstructured diagnostic text, irregular formatting, ambiguous diagnoses, and uncoded diseases were addressed by a unified preprocessing step, with diagnoses mapped to the National Clinical Version 2.0 Disease Codes. All preprocessing was implemented in SAS, which enabled batch data cleaning, efficient storage, and optimized extraction, ensuring reliable datasets for subsequent analysis.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\u003ch2\u003e2.4. Model construction and feature selection\u003c/h2\u003e\u003cp\u003e To reduce potential overfitting caused by limited sample size, the RF model was first tested to evaluate the predictive value of CBC data, followed by the addition of coagulation data and then tumor markers. Tumor markers, as important indicators for cancer diagnosis, were also used to validate predictive accuracy. Feature importance analysis further identified variables with greater predictive power than tumor markers, providing insight into valuable predictors.Six models\u0026mdash;LR, SVM, XGB, LGBM, CART, and CatBoost\u0026mdash;were compared. Model evaluation used a 70:30 train-test split with control samples matched to the experimental group and 10-fold cross-validation. To quantify the role of individual features, permutation feature importance (PFI) was applied, ranking features by their contribution to model performance. The final feature subset was determined by aggregating importance rankings across models.All procedures were implemented using the ranger package in R and Scikit-Learn in Python.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\u003ch2\u003e2.5. Statistical analysis\u003c/h2\u003e\u003cp\u003eContinuous variables were tested for normality using the Shapiro test and described as mean\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation or median with interquartile range, as appropriate. Categorical variables were expressed as percentages. Group differences were assessed using Chi-square tests for categorical variables and t-tests or Kruskal\u0026ndash;Wallis tests for continuous variables, depending on distribution.\u003c/p\u003e\u003cp\u003eTo account for covariates in model development, correlation analyses were performed: Spearman correlation for continuous variables, Cramer\u0026rsquo;s V for categorical variables, and point-biserial correlation for continuous\u0026ndash;categorical pairs. Statistical analyses were conducted in RStudio version 4.2.3, with p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 considered statistically significant.\u003c/p\u003e\u003c/div\u003e"},{"header":"3. Result","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003e3.1. Clinical characteristics of patients\u003c/h2\u003e\u003cp\u003eThe dataset included 791,867 visits, comprising 169,703 physical examination cases and 622,164 disease diagnoses. For modeling, 12,964 lung cancer patients (malignant tumors of the bronchi and lungs) were selected, with 2,757 breast cancer patients used for robustness validation. Based on physical examination data (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), participants were divided into a healthy group (169,703) and a lung cancer group (12,964), including 88,417 males (48.40%) and 94,250 females (51.60%). Lung cancer patients were older, and significant differences were observed in most CBC, coagulation, and tumor marker parameters compared with the healthy group (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eComparison of Features between C34 Group and Healthy Group\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"7\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFeature Name\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLevel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAbbreviation\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eN\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eC34 Mean\u0026thinsp;\u0026plusmn;\u0026thinsp;Standard Deviation\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eHealthy Group Mean\u0026thinsp;\u0026plusmn;\u0026thinsp;Standard Deviation\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003ep-value\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eGender\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGender\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e88417\u003c/p\u003e\u003cp\u003e(48.40%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e6098\u003c/p\u003e\u003cp\u003e(47.04%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e82319\u003c/p\u003e\u003cp\u003e(48.51%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eFemale\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGender\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e94250\u003c/p\u003e\u003cp\u003e(51.60%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e6866\u003c/p\u003e\u003cp\u003e(52.96%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e87384\u003c/p\u003e\u003cp\u003e(51.49%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"9\" rowspan=\"10\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e[0,10)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e11\u003c/p\u003e\u003cp\u003e(0.01%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0\u003c/p\u003e\u003cp\u003e(0.00%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e11\u003c/p\u003e\u003cp\u003e(0.01%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e[10,20)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e3518\u003c/p\u003e\u003cp\u003e(1.93%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e9\u003c/p\u003e\u003cp\u003e(0.07%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e3509\u003c/p\u003e\u003cp\u003e(2.07%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e[20,30)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e21555\u003c/p\u003e\u003cp\u003e(11.80%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e83\u003c/p\u003e\u003cp\u003e(0.64%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e21472\u003c/p\u003e\u003cp\u003e(12.65%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e[30,40)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e43331\u003c/p\u003e\u003cp\u003e(23.72%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e336\u003c/p\u003e\u003cp\u003e(2.59%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e42995\u003c/p\u003e\u003cp\u003e(25.34%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e[40,50)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e37911\u003c/p\u003e\u003cp\u003e(20.76%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e917\u003c/p\u003e\u003cp\u003e(7.07%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e36994\u003c/p\u003e\u003cp\u003e(21.80%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e[50,60)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e39551\u003c/p\u003e\u003cp\u003e(21.65%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e3186\u003c/p\u003e\u003cp\u003e(24.58%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e36365\u003c/p\u003e\u003cp\u003e(21.43%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e[60,70)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e23307\u003c/p\u003e\u003cp\u003e(12.76%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e5065\u003c/p\u003e\u003cp\u003e(39.07%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e18242\u003c/p\u003e\u003cp\u003e(10.75%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e[70,80)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e10311\u003c/p\u003e\u003cp\u003e(5.64%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2978\u003c/p\u003e\u003cp\u003e(22.97%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e7333\u003c/p\u003e\u003cp\u003e(4.32%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.008\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e[80,90)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e2878\u003c/p\u003e\u003cp\u003e(1.58%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e380\u003c/p\u003e\u003cp\u003e(2.93%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e2498\u003c/p\u003e\u003cp\u003e(1.47%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e[90,100]\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e286\u003c/p\u003e\u003cp\u003e(0.16%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e10\u003c/p\u003e\u003cp\u003e(0.08%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e276\u003c/p\u003e\u003cp\u003e(0.16%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eWhite Blood Cell Count\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eWBC\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e8.51\u003c/p\u003e\u003cp\u003e(3.99)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e6.12\u003c/p\u003e\u003cp\u003e(1.81)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNeutrophil Percentage\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNE%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e72.79\u003c/p\u003e\u003cp\u003e(11.76)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e57.03\u003c/p\u003e\u003cp\u003e(8.53)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLymphocyte Percentage\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLY%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e19.01\u003c/p\u003e\u003cp\u003e(10.18)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e33.33\u003c/p\u003e\u003cp\u003e(7.92)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMonocyte Percentage\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMO%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e6.37\u003c/p\u003e\u003cp\u003e(2.21)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e7.02\u003c/p\u003e\u003cp\u003e(1.84)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEosinophil Percentage\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eEO%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1.52\u003c/p\u003e\u003cp\u003e(2.14)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e2.05\u003c/p\u003e\u003cp\u003e(1.78)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBasophil Percentage\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eBA%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.31\u003c/p\u003e\u003cp\u003e(0.26)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.57\u003c/p\u003e\u003cp\u003e(0.29)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAbsolute Neutrophil Count\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNE#\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e6.46\u003c/p\u003e\u003cp\u003e(3.83)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e3.54\u003c/p\u003e\u003cp\u003e(1.41)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAbsolute Lymphocyte Count\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLY#\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1.39\u003c/p\u003e\u003cp\u003e(0.60)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e2.00\u003c/p\u003e\u003cp\u003e(0.71)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAbsolute Monocyte Count\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMO#\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.52\u003c/p\u003e\u003cp\u003e(0.25)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.42\u003c/p\u003e\u003cp\u003e(0.14)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAbsolute Eosinophil Count\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eEO#\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.12\u003c/p\u003e\u003cp\u003e(0.41)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.12\u003c/p\u003e\u003cp\u003e(0.12)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAbsolute Basophil Count\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eBA#\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.02\u003c/p\u003e\u003cp\u003e(0.02)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.03\u003c/p\u003e\u003cp\u003e(0.04)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRed Blood Cell Count\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRBC\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e4.24\u003c/p\u003e\u003cp\u003e(0.58)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e4.76\u003c/p\u003e\u003cp\u003e(0.52)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHemoglobin\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eHGB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e128.36\u003c/p\u003e\u003cp\u003e(18.51)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e143.10\u003c/p\u003e\u003cp\u003e(17.08)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHematocrit\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePCV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e38.68\u003c/p\u003e\u003cp\u003e(5.18)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e42.80\u003c/p\u003e\u003cp\u003e(4.50)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.003\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMean Corpuscular Volume\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMCV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e91.47\u003c/p\u003e\u003cp\u003e(4.75)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e90.14\u003c/p\u003e\u003cp\u003e(4.68)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.746\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMean Corpuscular Hemoglobin\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMCH\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e30.33\u003c/p\u003e\u003cp\u003e(1.90)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e30.11\u003c/p\u003e\u003cp\u003e(2.02)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.382\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMean Corpuscular Hemoglobin Concentration\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMCHC\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e331.44\u003c/p\u003e\u003cp\u003e(9.36)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e333.89\u003c/p\u003e\u003cp\u003e(11.05)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRed Cell Distribution Width-Coefficient of Variation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRDW-CV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e13.09\u003c/p\u003e\u003cp\u003e(1.26)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12.68\u003c/p\u003e\u003cp\u003e(1.12)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRed Cell Distribution Width-Standard Deviation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRDW-SD\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e43.53\u003c/p\u003e\u003cp\u003e(4.03)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e41.60\u003c/p\u003e\u003cp\u003e(3.40)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePlatelet Count\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePLT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e230.55\u003c/p\u003e\u003cp\u003e(79.52)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e245.46\u003c/p\u003e\u003cp\u003e(59.14)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.035\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMean Platelet Volume\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMPV\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e14.92\u003c/p\u003e\u003cp\u003e(2.25)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e11.61\u003c/p\u003e\u003cp\u003e(1.94)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.781\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePlateletcrit\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePCT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e9.88\u003c/p\u003e\u003cp\u003e(1.06)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e10.10\u003c/p\u003e\u003cp\u003e(0.87)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.278\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePlatelet Distribution Width\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePDW\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.22\u003c/p\u003e\u003cp\u003e(0.07)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.25\u003c/p\u003e\u003cp\u003e(0.06)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLarge Platelet Ratio\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eP-LCR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e182667\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e24.94\u003c/p\u003e\u003cp\u003e(7.47)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e25.77\u003c/p\u003e\u003cp\u003e(6.95)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eThrombin Time\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e32655\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e12.83\u003c/p\u003e\u003cp\u003e(0.91)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e12.74\u003c/p\u003e\u003cp\u003e(1.35)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003ep_value\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eProthrombin Time Activity Percentage\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePTActivity%\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e32655\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e111.92\u003c/p\u003e\u003cp\u003e(16.74)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e111.03\u003c/p\u003e\u003cp\u003e(14.52)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eInternational Normalized Ratio\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eINR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e32655\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.96\u003c/p\u003e\u003cp\u003e(0.07)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.97\u003c/p\u003e\u003cp\u003e(0.10)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eInternational Normalized Ratio\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eINR2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e32655\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.96\u003c/p\u003e\u003cp\u003e(0.09)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.96\u003c/p\u003e\u003cp\u003e(0.15)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFibrinogen\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eFIB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e37001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e3.78\u003c/p\u003e\u003cp\u003e(1.31)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e3.41\u003c/p\u003e\u003cp\u003e(0.91)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eActivated Partial Thromboplastin Time\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAPTT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e32529\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e35.36\u003c/p\u003e\u003cp\u003e(4.31)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e35.07\u003c/p\u003e\u003cp\u003e(4.47)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eThrombin Time\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e32529\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e17.80\u003c/p\u003e\u003cp\u003e(1.03)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e17.23\u003c/p\u003e\u003cp\u003e(5.07)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eD-Dimer\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eD-Di\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e9319\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.97\u003c/p\u003e\u003cp\u003e(1.74)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.89\u003c/p\u003e\u003cp\u003e(1.69)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFibrin Degradation Products\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eFDP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e151\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e5.15\u003c/p\u003e\u003cp\u003e(2.80)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e5.56\u003c/p\u003e\u003cp\u003e(14.61)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAntithrombin III\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAT III\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e154\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e105.54\u003c/p\u003e\u003cp\u003e(25.07)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e97.56\u003c/p\u003e\u003cp\u003e(12.77)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNeuron-specific enolase\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eNSE\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e5017\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e19.44\u003c/p\u003e\u003cp\u003e(28.91)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e15.02\u003c/p\u003e\u003cp\u003e(3.73)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.008\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCytokeratin-19 fragment\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eCYFRA 21\u0026thinsp;\u0026minus;\u0026thinsp;1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e5073\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e4.77\u003c/p\u003e\u003cp\u003e(13.19)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e1.85\u003c/p\u003e\u003cp\u003e(0.86)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\u003ch2\u003e3.2. Removing covariates and interpolating missing values\u003c/h2\u003e\u003cp\u003eThe baseline model included gender, age, and 24 CBC parameters. Correlation analysis revealed strong associations among several features, leading to the exclusion of 11 variables (HCT, LY%, MPV, PCT, HGB, WBC, MCH, BASO#, EO#, NE#, RDW-CV). Correlation patterns among the remaining variables are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. CBC, coagulation, and tumor marker data exhibited varying degrees of missingness. Samples with more than half of their features missing were excluded. For tumor markers, features were screened according to missingness, and NSE and CYFRA 21\u0026thinsp;\u0026minus;\u0026thinsp;1 were retained for lung cancer analysis. Remaining missing values were imputed within clinical reference ranges. Finally, an equal number of samples was randomly drawn from the physical examination group to match the lung cancer group, forming the dataset for subsequent analysis.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003e3.3. Model performance\u003c/h2\u003e\u003cp\u003eSeven machine learning models (RF, LR, SVM, XGB, LGBM, CART, and CatBoost) were trained in Python. Model performance was sequentially evaluated using CBC data alone, followed by the addition of coagulation parameters, tumor markers, and finally the full set of variables. Except for CART (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e), all models demonstrated favorable predictive ability; thus, CART was not subjected to further optimization. In addition to Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, we visualized part of these results using a forest plot of AUC values with 95% confidence intervals (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). This plot shows that the ensemble methods\u0026mdash;XGBoost, LGBM, and RF\u0026mdash;consistently achieved the highest AUC values (\u0026gt;\u0026thinsp;0.95) with narrow confidence intervals, indicating superior accuracy and stable performance. In contrast, CART yielded the lowest discriminative ability (AUC\u0026thinsp;=\u0026thinsp;0.841, 95% CI: 0.836\u0026ndash;0.847), whereas LR and SVM achieved moderately high AUC values (~\u0026thinsp;0.94\u0026ndash;0.95) but still lagged behind the ensemble models. As shown subsequently in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, the PR curves further indicate stable predictive performance without evidence of overfitting.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eThe evaluation indicators for each model\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"7\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eData Type\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eROC_AUC\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eAccuracy\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003ePrecision\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eRecall\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eF1_score\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCART\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e84.07\u003c/p\u003e\u003cp\u003e(83.58\u0026ndash;84.79)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e84.07\u003c/p\u003e\u003cp\u003e(83.58\u0026ndash;84.79)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e83.50\u003c/p\u003e\u003cp\u003e(82.10-84.89)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e84.89\u003c/p\u003e\u003cp\u003e(84.09\u0026ndash;85.78)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e84.18\u003c/p\u003e\u003cp\u003e(83.67\u0026ndash;84.81)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCART\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u0026thinsp;+\u0026thinsp;Coagulation\u0026thinsp;+\u0026thinsp;Tumor markers\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e88.16\u003c/p\u003e\u003cp\u003e(87.50-89.42)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e88.16\u003c/p\u003e\u003cp\u003e(87.50-89.42)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e88.08\u003c/p\u003e\u003cp\u003e(86.84\u0026ndash;89.09)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e88.46\u003c/p\u003e\u003cp\u003e(87.45\u0026ndash;89.85)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e88.27\u003c/p\u003e\u003cp\u003e(87.67\u0026ndash;89.47)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCatBoost\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e95.94\u003c/p\u003e\u003cp\u003e(95.49\u0026ndash;96.20)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e89.81\u003c/p\u003e\u003cp\u003e(89.44\u0026ndash;90.17)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e92.53\u003c/p\u003e\u003cp\u003e(90.95\u0026ndash;93.58)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e86.60\u003c/p\u003e\u003cp\u003e(85.94\u0026ndash;87.76)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e89.46\u003c/p\u003e\u003cp\u003e(89.05\u0026ndash;89.78)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCatBoost\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u0026thinsp;+\u0026thinsp;Coagulation\u0026thinsp;+\u0026thinsp;Tumor markers\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e98.28\u003c/p\u003e\u003cp\u003e(97.97\u0026ndash;98.45)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e93.62\u003c/p\u003e\u003cp\u003e(93.27\u0026ndash;94.05)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e95.49\u003c/p\u003e\u003cp\u003e(94.87\u0026ndash;96.05)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e91.67\u003c/p\u003e\u003cp\u003e(90.81\u0026ndash;92.56)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e93.54\u003c/p\u003e\u003cp\u003e(93.20-93.97)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLGBM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e95.79\u003c/p\u003e\u003cp\u003e(95.36\u0026ndash;96.11)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e89.60\u003c/p\u003e\u003cp\u003e(88.99\u0026ndash;90.13)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e91.98\u003c/p\u003e\u003cp\u003e(90.33\u0026ndash;93.09)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e86.75\u003c/p\u003e\u003cp\u003e(85.67\u0026ndash;87.57)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e89.29\u003c/p\u003e\u003cp\u003e(88.55\u0026ndash;89.74)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLGBM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u0026thinsp;+\u0026thinsp;Coagulation\u0026thinsp;+\u0026thinsp;Tumor markers\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e98.11\u003c/p\u003e\u003cp\u003e(97.77\u0026ndash;98.26)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e93.36\u003c/p\u003e\u003cp\u003e(92.89\u0026ndash;93.67)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e94.87\u003c/p\u003e\u003cp\u003e(94.02\u0026ndash;95.57)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e91.78\u003c/p\u003e\u003cp\u003e(91.12\u0026ndash;92.21)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e93.30\u003c/p\u003e\u003cp\u003e(92.82\u0026ndash;93.65)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e94.34\u003c/p\u003e\u003cp\u003e(94.00-94.62)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e88.02\u003c/p\u003e\u003cp\u003e(87.72\u0026ndash;88.38)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e90.84\u003c/p\u003e\u003cp\u003e(89.99\u0026ndash;91.88)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e84.55\u003c/p\u003e\u003cp\u003e(83.96\u0026ndash;85.35)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e87.58\u003c/p\u003e\u003cp\u003e(87.18\u0026ndash;87.91)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u0026thinsp;+\u0026thinsp;Coagulation\u0026thinsp;+\u0026thinsp;Tumor markers\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e96.00\u003c/p\u003e\u003cp\u003e(95.65\u0026ndash;96.25)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e90.11\u003c/p\u003e\u003cp\u003e(89.51\u0026ndash;90.86)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e92.59\u003c/p\u003e\u003cp\u003e(92.00-93.28)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e87.35\u003c/p\u003e\u003cp\u003e(86.16\u0026ndash;88.93)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e89.89\u003c/p\u003e\u003cp\u003e(89.22\u0026ndash;90.67)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRF\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e95.52\u003c/p\u003e\u003cp\u003e(95.02\u0026ndash;95.85)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e89.37\u003c/p\u003e\u003cp\u003e(88.75\u0026ndash;89.58)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e92.23\u003c/p\u003e\u003cp\u003e(90.85\u0026ndash;93.71)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e85.97\u003c/p\u003e\u003cp\u003e(85.07\u0026ndash;87.15)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e88.98\u003c/p\u003e\u003cp\u003e(88.19\u0026ndash;89.26)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRF\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u0026thinsp;+\u0026thinsp;Coagulation\u0026thinsp;+\u0026thinsp;Tumor markers\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e97.46\u003c/p\u003e\u003cp\u003e(96.95\u0026ndash;97.65)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e92.55\u003c/p\u003e\u003cp\u003e(92.09\u0026ndash;93.08)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e94.54\u003c/p\u003e\u003cp\u003e(93.59\u0026ndash;95.37)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e90.43\u003c/p\u003e\u003cp\u003e(89.29\u0026ndash;91.20)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e92.43\u003c/p\u003e\u003cp\u003e(91.97\u0026ndash;92.95)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSVM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e94.96\u003c/p\u003e\u003cp\u003e(94.53\u0026ndash;95.27)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e89.49\u003c/p\u003e\u003cp\u003e(89.04\u0026ndash;90.01)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e93.86\u003c/p\u003e\u003cp\u003e(92.74\u0026ndash;95.15)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e84.50\u003c/p\u003e\u003cp\u003e(83.55\u0026ndash;85.54)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e88.93\u003c/p\u003e\u003cp\u003e(88.65\u0026ndash;89.30)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSVM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u0026thinsp;+\u0026thinsp;Coagulation\u0026thinsp;+\u0026thinsp;Tumor markers\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e96.88\u003c/p\u003e\u003cp\u003e(96.40\u0026ndash;97.10)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e91.63\u003c/p\u003e\u003cp\u003e(91.14\u0026ndash;92.02)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e94.66\u003c/p\u003e\u003cp\u003e(93.94\u0026ndash;95.31)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e88.36\u003c/p\u003e\u003cp\u003e(87.10-89.36)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e91.40\u003c/p\u003e\u003cp\u003e(91.02\u0026ndash;91.79)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eXGB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e95.47\u003c/p\u003e\u003cp\u003e(95.05\u0026ndash;95.72)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e89.25\u003c/p\u003e\u003cp\u003e(88.66\u0026ndash;89.70)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e91.39\u003c/p\u003e\u003cp\u003e(90.10-92.91)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e86.66\u003c/p\u003e\u003cp\u003e(85.76\u0026ndash;88.04)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e88.96\u003c/p\u003e\u003cp\u003e(88.19\u0026ndash;89.37)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eXGB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eBlood routine\u0026thinsp;+\u0026thinsp;Coagulation\u0026thinsp;+\u0026thinsp;Tumor markers\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e98.11\u003c/p\u003e\u003cp\u003e(97.11\u0026ndash;98.32)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e93.52\u003c/p\u003e\u003cp\u003e(93.05\u0026ndash;94.07)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e95.04\u003c/p\u003e\u003cp\u003e(94.26\u0026ndash;95.73)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e91.94\u003c/p\u003e\u003cp\u003e(89.39\u0026ndash;92.53)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e93.46\u003c/p\u003e\u003cp\u003e(92.53\u0026ndash;93.97)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"7\"\u003eNote: The model results obtained when the Data Type is 'Blood routine\u0026thinsp;+\u0026thinsp;Coagulation' or 'Blood routine\u0026thinsp;+\u0026thinsp;Tumor markers' are not presented in the table.\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\u003ch2\u003e3.4. Construct and validate the feature subsets\u003c/h2\u003e\u003cp\u003ePFI analysis (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e) showed that PDW, age, NE%, and RBC were the most important features for lung cancer prediction. Using only these four features, CatBoost still achieved strong performance (AUC 94.9%; Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e). These features also have clear clinical relevance: PDW reflects platelet activation and tumor-associated inflammation; age is a well-recognized non-modifiable risk factor; NE% indicates systemic inflammatory response; and RBC changes may reflect tumor-related hypoxia or altered erythropoiesis. Their consistent importance across models highlights biological plausibility and supports their integration into simplified, cost-effective tools for early lung cancer screening in clinical practice.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eThe evaluation metrics for each model using top four PFI features\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eROC_AUC\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAccuracy\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003ePrecision\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eRecall\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eF1_score\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCatBoost\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e94.93 (94.55\u0026ndash;95.38)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e88.05 (87.90-88.19)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e89.64 (89.51\u0026ndash;89.77)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e86.29 (86.17\u0026ndash;86.40)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e87.93 (87.81\u0026ndash;88.06)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLGBM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e94.79 (94.47\u0026ndash;95.20)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e87.71 (87.22\u0026ndash;88.24)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e89.32 (89.05\u0026ndash;89.74)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e86.25 (86.10-86.54)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e87.79 (87.62\u0026ndash;88.11)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e93.10 (92.59-94.00)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e86.09 (85.85\u0026ndash;86.25)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e87.43 (86.91\u0026ndash;87.81)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e84.52 (84.44\u0026ndash;84.63)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e85.95 (85.68\u0026ndash;86.09)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRF\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e94.28 (93.94\u0026ndash;94.67)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e87.57 (87.05\u0026ndash;87.93)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e89.29 (88.83\u0026ndash;90.01)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e85.58 (85.09\u0026ndash;86.26)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e87.39 (86.92\u0026ndash;87.64)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSVM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e93.30 (92.68\u0026ndash;93.75)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e87.98 (87.79\u0026ndash;88.28)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e90.50 (90.21\u0026ndash;90.71)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e85.05 (84.25\u0026ndash;85.65)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e87.69 (87.36\u0026ndash;88.05)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eXGB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e94.32 (93.97\u0026ndash;94.66)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e87.65 (87.22\u0026ndash;88.31)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e88.98 (88.38\u0026ndash;90.01)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e86.16 (86.03\u0026ndash;86.23)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e87.54 (87.19\u0026ndash;88.08)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eIn this study, we developed a machine learning framework for lung cancer prediction using routine laboratory data. Among the tested models, CatBoost achieved the best performance, with high accuracy in both internal validation and external testing. Using PFI, four key features (PDW, age, NE%, and RBC) were identified, and a simplified model with only these variables still maintained high predictive accuracy (AUC 94.9%). These findings suggest that routine and low-cost laboratory indicators can provide valuable support for early cancer screening.\u003c/p\u003e\u003cp\u003eIn clinical research, encountering missing and imbalanced data is inevitable in real-world studies. Missing data, if not properly addressed, can lead to information loss, biased estimation, and impaired model performance. For instance, a study on patient-specific MACE risk prediction highlighted that different imputation methods can significantly affect predictive accuracy, and improper handling of missing values undermines the model\u0026rsquo;s reliability [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Furthermore, the effectiveness of commonly used imputation techniques varies across different real-world scenarios, and selecting an appropriate method requires careful consideration of the dataset\u0026rsquo;s specific characteristics [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eA systematic review examining machine learning-based clinical prediction models revealed that missing data is often poorly handled and inadequately reported [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. Many studies either failed to disclose how missing values were treated or adopted simplistic strategies such as complete case analysis or mean imputation, both of which can lead to substantial information loss and reduced model accuracy. In this study, we explored the underlying causes of data missingness and implemented a feature-informed imputation strategy. By leveraging available clinical test information, we aimed to restore missing values in a manner that closely approximates their true distribution, thereby preserving the integrity and representativeness of the dataset.\u003c/p\u003e\u003cp\u003eSimilarly, class imbalance is a common phenomenon in medical diagnostic datasets. When not properly handled, it introduces systematic bias into model training, as machine learning algorithms tend to favor the majority class. This results in overly optimistic performance metrics that fail to generalize to minority populations. This challenge has been widely acknowledged in previous literature. For example, Tasci et al emphasized the impact of class imbalance in oncologic datasets and advocated for inclusive and bias-aware frameworks to ensure equitable predictive performance across all subgroups [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. In a recent review, Wang et al demonstrated that the performance and reliability of deep learning models are significantly affected by imbalanced data distributions in medical imaging tasks [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Similarly, Zhang et al reviewed a decade of progress in addressing imbalance in medical datasets and highlighted how conventional machine learning methods often fail to detect rare or minority-class cases, thereby limiting diagnostic accuracy [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e].]\u003c/p\u003e\u003cp\u003eTo mitigate this issue, we applied a random under-sampling approach to achieve class balance and implemented k-fold cross-validation to validate consistency and reliability. The predictive pipeline developed in this study also demonstrated scalability, supporting integration with external data sources such as genomic data, which could enable adaptive, multimodal models. Model interpretability is essential for promoting clinical trust and ensuring responsible deployment of AI systems in healthcare. Complex models such as ensemble methods often lack transparency, which can hinder their acceptance in clinical practice [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. To address this, we applied PFI to identify key predictive variables in our CatBoost model. This method quantitatively ranks feature contributions and enhances the interpretability of model outputs [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. Our analysis revealed that PDW, age, NE%, and RBC were the most influential features, supporting both the model\u0026rsquo;s clinical relevance and its practical applicability.\u003c/p\u003e\u003cp\u003eOur interpretability analysis further reinforces the clinical utility of the model by highlighting four key features\u0026mdash;PDW, age, NE%, and RBC\u0026mdash;which not only demonstrated strong statistical importance but also have solid biological plausibility in the context of lung cancer. PDW reflects the variability in platelet size, which increases during platelet activation, a process linked to tumor-driven inflammation and angiogenesis. Elevated PDW levels have been reported in various malignancies and are associated with poor prognosis, likely due to the role of activated platelets in promoting tumor cell proliferation, metastasis, and immune evasion. The consistently highest importance score of PDW across multiple models suggests that it may serve as a sensitive hematologic biomarker for early cancer detection. Age is a well-recognized independent risk factor for lung cancer. As individuals age, cumulative exposure to carcinogens, decline in immune function, and accumulation of genetic mutations contribute to increased cancer susceptibility. Its strong predictive contribution in all models highlights its essential role in risk stratification and supports its integration into AI-based clinical tools. NE% is a marker of systemic inflammation and immune status. Elevated NE% is frequently observed in cancer patients and reflects tumor-induced neutrophilia, which can suppress anti-tumor immune responses and support tumor progression through the secretion of growth factors and proteases. Its inclusion in the top features aligns with current evidence linking systemic inflammatory markers with cancer prognosis. RBC count, although less commonly emphasized in oncology, may indicate chronic hypoxia, nutritional deficiency, or impaired erythropoiesis\u0026mdash;conditions often associated with tumor burden and cancer-related metabolic changes. In this study, RBC demonstrated consistent relevance across models, suggesting that subtle hematologic shifts detectable via routine CBC may contribute valuable diagnostic signals. Collectively, these four features allow for the construction of a highly interpretable and simplified diagnostic model that maintains excellent predictive performance (AUC 94.93%, Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e), thus offering a practical solution for early lung cancer detection. Importantly, these parameters are readily available, cost-effective, and routinely measured in clinical practice, supporting their potential integration into large-scale screening frameworks, particularly in resource-limited healthcare settings. Additionally, because consent was waived and only de-identified historical records were used, we could not perform patient-level re-contact or adjudicate outcomes beyond the available electronic records, which may leave residual misclassification unaddressed.\u003c/p\u003e\u003cp\u003eThis study has several limitations. First, only conventional machine learning models were evaluated; future work could explore deep learning approaches. Second, comorbidities may influence model performance in real-world settings, increasing the risk of misclassification. Third, although commonly available laboratory indicators provided strong predictive accuracy, the additional value of tumor markers was limited. Future research could incorporate broader clinical or molecular features to further enhance prediction. To mitigate the risk of spectrum and selection biases inherent to single-center retrospective datasets, future work should include pre-registered, prospective, multi-center cohorts with standardized CBC acquisition protocols and blinded adjudication. External validation across diverse clinical settings and laboratory analyzers will be essential to confirm transportability and support any downstream clinical deployment.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003cp\u003e This study was reviewed and approved by the Ethics Committee of the Second Affiliated Hospital of Dalian Medical University (IRB approval No. KY2025-137-01). Given the retrospective design and the use of fully de-identified data, the Institutional Review Board (IRB) waived the requirement for obtaining individual informed consent in accordance with the Declaration of Helsinki and the Measures for the Ethical Review of Biomedical Research Involving Humans (National Health and Family Planning Commission of the People\u0026rsquo;s Republic of China, 2016). The study posed no more than minimal risk to participants, and obtaining consent was impracticable in this context; therefore, the IRB determined that a consent waiver met the applicable regulatory criteria.\u003c/p\u003e\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003cp\u003eNot applicable.\u003c/p\u003e\u003c/p\u003e\u003cp\u003e\u003ch2\u003eCompeting interests\u003c/h2\u003e\u003cp\u003eAll authors have no additional conflicts of interest to declare.\u003c/p\u003e\u003c/p\u003e\u003cp\u003e\u003ch2\u003eAuthor details\u003c/h2\u003e\u003cp\u003e\u003csup\u003e1\u003c/sup\u003eDepartment of Clinical Laboratory, The Second Hospital of Dalian Medical University, Dalian 116023, China.\u003c/p\u003e\u003cp\u003e\u003csup\u003e2\u003c/sup\u003eSchool of statistics, Dongbei University of Finance and Economics, Dalian 116025, China\u003c/p\u003e\u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e\u003cp\u003eThis work was supported by a grant from the Dalian Science and Technology Innovation Fund Program (No. 2024JJ13PT070) and United Foundation for Dalian Institute of Chemical Physics, Chinese Academy of Sciences and the Second Hospital of Dalian Medical University (No. DMU-2\u0026amp;DICP UN202410), Dalian Life and Health Field Guidance Program Project (No. 2024ZDJH01PT084).\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eTL and YW conceived the study and performed initial discussions with LW, HW and CY. Afterward, TL, YW and LW designed the research protocols and coordinated the data collection and analyses. TL and YW carried out the data curation, preprocessing, and machine learning analyses. LW and HW supervised the statistical analysis and model validation. LW, HW and CY oversaw the clinical interpretation of the results. TL and YW wrote the first draft of the manuscript. LW, HW and CY critically reviewed and revised the manuscript for important intellectual content. All authors read and approved the final version of the manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgements\u003c/h2\u003e\u003cp\u003eNot applicable.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eData are available upon reasonable request to corresponding author.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAgarwal R, Sarkar A, Bhowmik A, Mukherjee D, Chakraborty S. A portable spinning disc for complete blood count (CBC). Biosens Bioelectron. 2020;150:111935.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLuo G. MLBCD: a machine learning tool for big clinical data. Health Inf Sci Syst. 2015;3:3.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCao M, Li H, Sun D, Chen W. Cancer burden of major cancers in China: A need for sustainable actions. Cancer Commun (Lond). 2020;40(5):205\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGBD. 2019 Colorectal Cancer Collaborators, Global, regional, and national burden of colorectal cancer and its risk factors, 1990\u0026ndash;2019: a systematic analysis for the Global Burden of Disease Study 2019, Lancet Gastroenterol Hepatol 7(7) (2022) 627\u0026ndash;647.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMazzone PJ, Silvestri GA, Souter LH, Caverly TJ, Kanne JP, Katki HA, et al. Screening for Lung Cancer: CHEST Guideline and Expert Panel Report. Chest. 2021;160(5):e427\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWang Y, Zhao Y, Li M, Hou H, Jian Z, Li W, et al. Conversion of primary liver cancer after targeted therapy for liver cancer combined with AFP-targeted CAR T-cell therapy: a case report. Front Immunol. 2023;14:1180001.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGrunnet M, Sorensen JB. Carcinoembryonic antigen (CEA) as tumor marker in lung cancer. Lung Cancer 76(2) (2012) 138\u0026thinsp;\u0026ndash;\u0026thinsp;43.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003e2019 Cancer Global Burden of Disease, Collaboration JM, Kocarnik K, Compton FE, Dean W, Fu BL, Gaw et al. Cancer Incidence, Mortality, Years of Life Lost, Years Lived With Disability, and Disability-Adjusted Life Years for 29 Cancer Groups From 2010 to 2019: A Systematic Analysis for the Global Burden of Disease Study 2019, JAMA Oncol 8(3) (2022) 420\u0026ndash;444.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRaess PW, van de Geijn GJ, Njo TL, Klop B, Sukhachev D, Wertheim G, et al. Automated screening for myelodysplastic syndromes through analysis of complete blood count and cell population data parameters. Am J Hematol. 2014;89(4):369\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBoutault R, Peterlin P, Boubaya M, Sockel K, Chevallier P, Garnier A, et al. A novel complete blood count-based score to screen for myelodysplastic syndrome in cytopenic patients. Br J Haematol. 2018;183(5):736\u0026ndash;46.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKe J, Qiu F, Fan W, Wei S. Associations of complete blood cell count-derived inflammatory biomarkers with asthma and mortality in adults: a population-based study. Front Immunol. 2023;14:1205687.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMueller AN, Miller HA, Taylor MJ, Suliman SA, Frieboes HB. Identification of Idiopathic Pulmonary Fibrosis and Prediction of Disease Severity via Machine Learning Analysis of Comprehensive Metabolic Panel and Complete Blood Count Data. Lung. 2024;202(2):139\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBongiovanni D, Han J, Klug M, Kirmes K, Viggiani G, von Scheidt M, et al. Role of Reticulated Platelets in Cardiovascular Disease. Arterioscler Thromb Vasc Biol. 2022;42(5):527\u0026ndash;39.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKwon O, Ahn JH, Koh JS, Park Y, Hwang SJ, Tantry US, et al. Platelet-fibrin clot strength and platelet reactivity predicting cardiovascular events after percutaneous coronary interventions. Eur Heart J. 2024;45(25):2217\u0026ndash;31.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGiannakeas V, Kotsopoulos J, Cheung MC, Rosella L, Brooks JD, Lipscombe L, et al. Analysis of Platelet Count and New Cancer Diagnosis Over a 10-Year Period. JAMA Netw Open. 2022;5(1):e2141633.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSanchez-Pinto LN, Bennett TD. Evaluation of Machine Learning Models for Clinical Prediction Problems. Pediatr Crit Care Med. 2022;23(5):405\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eStevens L, Kao D, Hall J, G\u0026ouml;rg C, Abdo K, Linstead E. A Preliminary Study of an Interactive Visual Analysis Tool Facilitating Clinical Applications of Machine Learning for Precision Medicine. Appl Sci (Basel). 2020;10(9):3309.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHaymond S, Master SR. How Can We Ensure Reproducibility and Clinical Translation of Machine Learning Applications in Laboratory Medicine? Clin Chem. 2022;68(3):392\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTan YY, Rim TH, Ting DSJ, Hsieh YT, Kim TI. Editorial: Big data and artificial intelligence in ophthalmology - clinical application and future exploration. Front Med (Lausanne). 2023;10:1339280.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChirumbolo S, Berretta M, Tirelli U. Trust, trustworthiness and acceptability of a machine learning adoption in data-driven clinical decision support system. Some comments. Int J Med Inf. 2024;184:105374.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRios R, Miller RJH, Manral N, Sharir T, Einstein AJ, Fish MB, et al. Handling missing values in machine learning to predict patient-specific risk of adverse cardiac events: Insights from REFINE SPECT registry. Comput Biol Med. 2022;145:105449.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi J, Guo S, Ma R, He J, Zhang X, Rui D, et al. Comparison of the effects of imputation methods for missing data in predictive modelling of cohort study datasets. BMC Med Res Methodol. 2024;24(1):41.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNijman S, Leeuwenberg AM, Beekers I, Verkouter I, Jacobs J, Bots ML, et al. Missing data is poorly handled and reported in prediction model studies using machine learning: a literature review. J Clin Epidemiol. 2022;142:218\u0026ndash;29.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTasci E, Zhuge Y, Camphausen K, Krauze AV. Cancers (Basel). 2022;14(12):2897. Bias and Class Imbalance in Oncologic Data-Towards Inclusive and Transferrable AI in Large Scale Oncology Data Sets.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhang J, Xie Y, Wu Q, Xia Y. Medical image classification using synergic deep learning. Med Image Anal. 2019;54:10\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGao L, Zhang L, Liu C, Wu S. Handling imbalanced medical image data: A deep-learning-based one-class classification approach. Artif Intell Med. 2020;108:101935.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRudin C. Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead. Nat Mach Intell. 2017;1(5):206\u0026ndash;15.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWang S, Liu Y, Wang W, Zhao G, Liang H. Interpretable machine learning guided by physical mechanisms reveals drivers of runoff under dynamic land use changes. J Environ Manage. 2024;367:121978.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-informatics-and-decision-making","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"midm","sideBox":"Learn more about [BMC Medical Informatics and Decision Making](http://bmcmedinformdecismak.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/midm/default.aspx","title":"BMC Medical Informatics and Decision Making","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Complete blood count, Lung cancer, Machine learning, Early diagnosis, Platelet distribution width, Neutrophil percentage","lastPublishedDoi":"10.21203/rs.3.rs-7811809/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7811809/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground\u003c/h2\u003e\u003cp\u003eEarly detection of lung cancer is crucial for improving outcomes, yet existing screening methods are costly and limited in accessibility. This study evaluated the diagnostic potential of routine complete blood count (CBC) parameters combined with machine learning (ML) for lung cancer prediction.\u003c/p\u003e\u003ch2\u003eDesign and methods\u003c/h2\u003e\u003cp\u003e: Data from 12,964 lung cancer patients and 169,703 healthy controls were retrospectively collected, including CBC, coagulation, and tumor marker results. After rigorous data preprocessing, multiple ML models were developed and validated, with CatBoost showing the best performance (AUC 95.84%, precision 92.53%, accuracy 89.81%, recall 86.60%, F1-score 89.46%).\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e\u003cp\u003eFeature importance analysis identified platelet distribution width (PDW), age, neutrophil percentage (NE%), and red blood cell count (RBC) as the most significant predictors. A reduced model using these four features retained high accuracy (AUC 94.93%), indicating their strong discriminative value. Compared to tumor markers and coagulation data, CBC-derived features alone were robust for lung cancer prediction.\u003c/p\u003e\u003ch2\u003eConclusions\u003c/h2\u003e\u003cp\u003eRoutine CBC parameters, paired with ML, may enable accurate and cost-effective lung cancer screening in retrospective, single-center data. Key features such as PDW, NE%, and RBC may serve as early diagnostic indicators. This approach offers a scalable solution for early cancer detection, particularly in resource-limited settings, and requires prospective, multi-center validation prior to clinical implementation.\u003c/p\u003e","manuscriptTitle":"Machine Learning-Based Cancer Prediction Using Complete Blood Count: A Retrospective Study on the Diagnostic Potential of Hematological Parameters in Lung Cancer Screening","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-11-10 16:14:27","doi":"10.21203/rs.3.rs-7811809/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2025-12-21T15:21:19+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"74610114164845137301853597823765466893","date":"2025-11-30T21:48:00+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"101264630982859466889327300380772650803","date":"2025-10-29T08:50:26+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-10-29T08:38:25+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-10-19T07:00:48+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-10-17T06:42:42+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-10-17T06:27:16+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Medical Informatics and Decision Making","date":"2025-10-17T06:24:12+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-medical-informatics-and-decision-making","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"midm","sideBox":"Learn more about [BMC Medical Informatics and Decision Making](http://bmcmedinformdecismak.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/midm/default.aspx","title":"BMC Medical Informatics and Decision Making","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"37153ee4-6a7a-4df3-94dc-b392f854a1e6","owner":[],"postedDate":"November 10th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2025-11-10T16:14:27+00:00","versionOfRecord":[],"versionCreatedAt":"2025-11-10 16:14:27","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7811809","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7811809","identity":"rs-7811809","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00