Cervical cancer prediction using machine learning models based on blood routine analysis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Cervical cancer prediction using machine learning models based on blood routine analysis Jie Su, Hui Lu, RuiHuan Zhang, Na Cui, Chao Chen, Qin Si, Biao Song This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4761322/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background and objective: Cervical cancer is the fourth most common cancer among women globally. The key of prevention and treatment of cervical cancer is early detection, diagnosis and treatment. We aimed to develop an interpretable model to predict the risk for patients with cervical cancer based on blood routine data and used the Shapley additive interpretation (SHAP) method to explain the model and explore factors for cervical cancer. Methods In this paper, medical records of patients from 2013 to 2023 were collected for retrospective study. 2533 patients with cervical cancer were used as the case group, and 9879 patients with apparent healthy subjects were used as the control group. Using age, clinical diagnosis information and 22 blood cell analysis results, four different algorithm were used to construct cervical cancer prediction model. Results Using lasso regression and random forest method, 15 important blood routine features were finally selected from 23 features for model training. Comparatively, the XGBoost model had the highest predictive performance among four models with an area under the curve (AUC) of 0.964, whereas RF had the poorest generalization ability (AUC = 0.907). The SHAP method reveals the top 6 predictors of cervical cancer according to the importance ranking, and the average of the PDW was recognized as the most important predictor variable. Conclusion In conclusion, we select the best ML based on performance and rank the importance of features according to Shapley Additive Explanation (SHAP) values. Compared to the other 4 algorithms, the results showed that the XGB had the best prediction performance for successfully predicting cervical cancer recurrence and was adopted in the establishment of the prediction model. Blood routine Cervical cancer Machine learning Shapley additive interpretation Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Cervical Cancer Epidemiology and Risk Factors Cervical cancer (CC) is one of the leading causes of death among women worldwide[ 1 ]. Cervical cancer accounts for 6.5 per cent of all cancers in women. In China, cervical cancer ranks the sixth among female malignant tumors and the eighth among the causes of cancer death [ 1 ], with up to 130,000 new cases per year, accounting for about 1/4 of the total number of new cases in the world[ 2 , 3 ]. Cervical cancer destroys deep tissues of the cervix and can gradually reach other areas of the human body, such as the lungs, liver, and vagina. Early cervical cancer may not have obvious signs and symptoms, which is the main reason for late detection of cervical cancer in patients. The key of prevention and treatment of cervical cancer is early detection, diagnosis and treatment. Cervical cancer screening can effectively detect precancerous lesions early and intervene in time. Current screening techniques include Papanicolaou (Pap) cytology-based test, Acetic Acid Papanicolaou Stain, liquid-based cytology, and HPV testing techniques [ 4 – 6 ]. Recently, many studies have been conducted on cervical cancer using modern techniques that provide prediction in the early stage. Blood routine is a common hematological index used in bedside included Red Blood Cell Count (RBC), White Blood Cell Count (WBC), Hemoglobin (HGB), Platelet Count (PLT) and so on. In the last several years, the Hemoglobin, Albumin, Lymphocyte, Platelet Score (HALP) has emerged in the literature as a new prognostic biomarker that has been used to predict a number of clinical outcomes in the context of various neoplasms [ 7 ]. Machine learning is to learn our existing data through certain logic, understand its internal connection and distribution, mine hidden information that we can't see by the naked eye, and predict, classify and evaluate the things we study to a certain extent through the result output, so as to give full play to the maximum value of the data and realize the leap from data to value. Various machine learning (ML) techniques have been used to detect various cancer types including lung, brain, breast, skin, blood, hippocampus, pancreas, cardiac, colon cancer, hepatic vessel, spleen etc [ 8 – 13 ]. Bychkov et al. combine Surface-enhanced laser desorption ionization-time of flight mass spectrometry (SELDI-TOF) mass spectrometry and artificial neural networks to analyze serum protein with high production sensitivity and specificity values for detecting and diagnosing colorectal cancer [ 14 ]. Chaudhary et al. used a novel approach using Deep Learning to identify multi-omics features of survival in patients with liver cancer and apply to future prognostic predictions of liver cancer [ 15 ]. Machine learning can classify clinical images of skin diseases, including intraepithelial carcinoma, actinic keratosis, seborrheic keratosis, lentigo, pyogenic granuloma et al. [ 16 , 17 ]. A number of mechanistic algorithms have been applied to the mental screening and diagnosis of cervical cancer, including the K Nearest Neighbors (KNN), logistic regression (LR), decision tree (DT), supporting vector machine (SVM), Artificial Neural Network (ANN) and related derivation algorithm. Mehmood et al. proposes an approach named CervDetect that uses machine learning algorithms to evaluate the risk elements of malignant cervical formation. CervDetect select significant features including Age, IUD (intrauterine derice), smokes, STDs (sexually transmitted disease), and so on. CervDetect achieved an accuracy of 93.6%, mean squared error (MSE) error of 0.07111, false-positive rate (FPR) of 6.4%, and false-negative rate (FNR) of 100%[ 18 ]. Devi et al. used Decision Tree (DT), Random Forest (RF), Logistic Regression (LR), Multi-layer Perceptron (MLP) and Naive Bayes (NB) to find out and analyze the predictors of Cervical Malignancy and evaluate the unbalanced input and target datasets. The predictors used in the current study included age, smokes, Infection of the cervix et al. Logistic Regression and Decision Tree algorithms were identified as the best-performed ML (Machine Learning) algorithm classifiers to detect the significant predictors[ 19 ]. Mudawi and Alazeb utilized random forest (RF), decision tree (DT), adaptive boosting, and gradient boosting algorithms to predict cervical cancer. The cancer detection dataset includes 32 variables, such as age, IUD, smokes, STDs, and so on[ 20 ]. Shapley additive interpretation (SHAP) is a model interpretation tool based on Shapley values, which is a post-interpretable scheme. SHAP can explain predictions by computing the contributions of individual variables, accounting for local accuracy, missingness, and consistency, to formulate absolute magnitude and directionality of impact in the prediction of the desired outcome. The purpose of our study is to develop an interpretable model to predict the risk for patients with cervical cancer based on the blood routine. In addition, the SHAP method is used to explain the odel and explore factors for cervical cancer. Materials and methods Patients A retrospective study was conducted by collecting medical records of patients from hospitals from 2013 to 2023, with cervical cancer data as the case group and noncervical cancer patient data as the control group. According to the inclusion and exclusion criteria, 2533 cases were included in the case group, and 5011 cases were included in the control group. The dataset mainly collects demographics, clinical characteristics, and baseline data related to cervical cancer, including age, pathological type, clinical diagnosis, and routine test data. The inclusion criteria were as follows: (1) Patients who meet the diagnostic criteria for cervical cancer in the "Guidelines for Standardized Treatment of Cervical Cancer and Precancerous Lesions (Trial)"; (2) The clinical staging of the patient is stages I to II, with concurrent radical hysterectomy and bilateral lymph node dissection; (3) Patients with complete clinical and pathological data, and included for the first time; (4) By persistently accepting medical advice and conducting regular follow-up. The exclusion criteria were as follows: (1) Pregnant or lactating women; (2) Patients with gynecological malignancies or a history of malignancies in other areas; (3) Patients with coagulation dysfunction; (4) Patients with a history of anti-tumor treatment such as radiotherapy and chemotherapy. Ethical Considerations Ethical approval and individual patient consent was not necessary because all the protected health information was anonymized. Data preprocessing Organize and preprocess the collected patient dataset, including data cleaning, feature selection, and data normalization. Due to the fact that this study was conducted on cervical cancer patients, the gender was set to female and the age range was set to 18–90 years old. The purpose of data cleaning is to improve the quality of the dataset by removing duplicate, invalid, and abnormal data, deleting data with missing or duplicate values, and ensuring the consistency of data units. Feature selection is the process of selecting the subset of features with the most predictive and imaging power from a dataset. The specific operation is to select features based on age and conventional test data, mainly using lassoregression and random forest methods. Data normalization is the process of converting raw data into data with unified standards, thereby eliminating dimensional differences between different features, enhancing the stability of the model training process, and improving the model's generalization ability. The Z-Score normalization method is mainly used, which converts the data into a standard normal distribution with a mean of 0 and a standard deviation of 1. After Z-score standardization, all data is converted into dimensionless standardized values, so that even if the original data has different scales or units, it can be compared and analyzed on the same scale. Research and Design of Machine Learning Analysis Firstly, divide the dataset into cervical cancer and noncervical cancer groups based on clinical diagnosis, and then screen the obtained features, which mainly include demographic, clinical, and baseline data information. After extracting 23 features from the preprocessed dataset in the feature selection stage, the optimal feature combination is selected through lassoregression and random forest model. In machine learning models, this study uses Extreme Gradient Boosting (XGB), Support Vector Machine (SVM), Random Forest (RF), and Neural Network (NN) to train classification from selected features. Evaluation process Cross validation is an evaluation method that prevents overflow and improves accuracy when evaluating model performance. This study used tenfold cross validation to verify the classification performance of the algorithm. To ensure comparability of results, each method used the same training and testing sets. Usually, binary classification algorithms are evaluated based on basic decision scores. The scores include True Positive (TP) as the number of positive samples correctly classified as positive, True Negative (TN) as the number of negative samples correctly classified as negative, False Positive (FP) as the number of negative samples misclassified as positive, and False Negative (FN) as the number of positive samples incorrectly classifed as negative. Using the above scores, evaluated each model using the following metrics. Recall rate, also known as sensitivity, is a measure of the model's ability to detect patients with illnesses. Specificity is the ability of a model to detect disease-free individuals. The F1 score is the harmonic mean of precision and recall. Accuracy refers to the proportion of total predicted misclassifications that are true, and the proportion of total predicted misclassifications that are incorrect. Recall = TP /(TP + FN) (1) Specificity = TN/ (TN + FP) (2) F1score = (2 *precision*recall)/(precision + recall) (3) Accuracy = (TP + TN)/(TP + FN + TN + FP) (4) Model explainability SHapley Additive exPlanations (SHAP) was used as explainability model, “ shap 0.39.0 ” package, to give local explanations and allow computation of the contribution of each variable to the prediction for each individual [ 21 ]. SHAP is an interpretable method based on game theory, drawing on the concept of Shapley values, aimed at quantifying the contribution of each feature to a specific predictive output. By analyzing the Shapley values of each feature, we can determine how each part of the machine learning model's prediction results affects the final prediction. A positive SHAP value indicates that the feature improves the result value and has a positive effect, while a negative SHAP value indicates that the feature reduces the result value and has a negative impact. This method can output the importance ranking of features and the relationship between features and results. Result Patient characteristics Among 6297 patients with cervical cancer, a total of 2503 adult patients diagnosed with cervical cancer were included in the final cohort for this study. The patient screening process is shown in Fig. 1 . Table 1 showed 6297 cervical cancer patients (average age, 52.81 years; range,18–90 years) from 2013 to 2023 were identified, and 2503 (39.7%) diagnosed with cervical cancer and 3794 (60.3%) non cervical cancer patients. Table 1 Statistical Table of the Dataset feature Total number average Standard error median standard deviation minimum value Maximum value AGE 6297 52.818 0.177 53 14.071 18 90 RBC 6297 4.023 0.008 4.14 0.670 0.81 7.23 MCV 6297 90.262 0.093 90.66 7.389 56.11 139.9 PDW 6297 14.142 0.037 15.7 2.908 0 23.44 WBC 6297 6.424 0.044 5.81 3.475 0.3 60.38 NEUT% 6297 64.174 0.172 64.3 13.642 0 98.7 LYMPH% 6297 25.999 0.148 25.8 11.722 0 90 EO% 6297 1.941 0.035 1.3 2.756 0 65.4 BASO% 6297 0.328 0.005 0.3 0.432 0 13.2 NEUT 6297 4.311 0.039 3.59 3.058 0 53.76 LYMPH 6297 1.535 0.013 1.45 0.998 0 36.8 BASO 6297 0.019 0.000 0.01 0.039 0 1.59 HGB 6297 117.534 0.284 122 22.558 0 178 HCT 6297 0.361 0.001 0.371 0.058 0.081 0.556 MCH 6297 29.551 0.037 29.8 2.965 15.1 79.5 MCHC 6297 327.044 0.168 328 13.309 249 697 R-CV 6297 14.029 0.031 13.3 2.470 0.117 34.93 PLT 6297 222.597 1.069 215 84.816 1 1547 MPV 6297 9.331 0.019 9.4 1.491 0 15.7 PCT 6297 0.205 0.001 0.2 0.077 0 1.11 MONO 6297 0.440 0.006 0.39 0.468 0 32.18 MONO% 6297 7.219 0.043 6.7 3.426 0 61.5 EO 6297 0.114 0.003 0.07 0.219 0 8.14 RBC: Red Blood Cell Count; MCV: Mean Corpuscular; PDW: Platelet Distribution Width; WBC: White Blood Cell Count; NEUT%:Neutrophil Percentage; LYMPH: Lymphocyte; EO%: Eosinophil Percentage; BASO: Basophil; NEUT: Neutrophil; HGB: Hemoglobin; HCT: Hematocrit; MCH: Mean Corpuscular Hemoglobin; MCHC: Mean Corpuscular Hemoglobin Concentration ; R-CV: Red Cell Distribution Width; PLT: Platelet Count; MPV: Mean Platelet Volume; PCT: platelet count; MONO: Monocyte Absolute Value ; EO: Eosinophil Absolute Value; Feature selection This study used lasso regression and random forest methods to select features from the dataset. Use lasso regression to perform feature selection on the dataset, as shown in Fig. 2 A. Features with values greater than zero in the graph have a positive linear relationship, while features with values less than zero have a negative linear relationship. Finally, 19 features were selected from the 23 features. The specific positive linear relationships include: EO, MONO%, HGB, BASO, LYMPH, NEUT, LYMPH%, NEUT%, WBC, RBC, AGE; Negative linear relationships include: PCT, MPV, R-CV, MCHC, MCH, BASO%, EO%, PDW. Figure 23B shows the use of random forest for feature selection on the dataset. The top 19 features were selected based on their importance, including LYMPH, PDW, MPV, BASO%, MONO, RBC, R-CV, WBC, MCHC, AGE, HCT, MONO%, MCV, BASO, HGB, PCT, LYMPH%, PLT and NEUT%. Based on the 19 features selected by each of the two methods, a common feature set was extracted,and finally 15 features were selected to participate in subsequent model training, as shown in Fig. 2 C. Machine learning classification architectures For machine learning classification, this study used the Extreme Gradient Boosting (XGB), Support Vector Machine (SVM), Random Forest (RF), and Neural Network (NN) architectures. XGB creates a powerful predictive model by iteratively constructing and combining multiple weak learners. As shown in Figure.3A, the pre pruning method is used to compensate for the error of the previous tree and create the next tree. The RF model as shown in Figure.3B is a collection of multiple decision trees, each of which predicts the data and ultimately integrates the results of all decision trees through voting to obtain classification results. As shown in Figure.3c, when two types of data are classified, SVM attempts to find a hyperplane. This hyperplane not only correctly divides samples of different categories, but also maximizes the distance from the closest sample point to the hyperplane. Neural network is a computational model that simulates the neural structure of the human brain, consisting of multiple nodes or neurons that transmit and process data through weighted connections, as shown in Fig. 3 D.With all of these models, a default value is designated for each parameter. Default values for the main parameters of the XGB model are: booster = gbtree, n_estimators = 10, learning rate eta = 0.2, class _weights= {0:0.06,1:0.9},and max default = 6. The SVM’s main parameter default values are: kernel = RBF, probability = True, class_weight={0:0.1,1:0.25}, and the kernel factor gamma = 1. Default values for the main parameters of the RF model are: criterion="gini", max_depth = 25, nestimators = 80, min _samples _split = 2,class_wei-ght={0:0.01,1:3.1}.The NN’s main parameter default values are: 2 hidden layers with 128 and 32 neurons respectively, the first hidden layer uses L1 regularization and regulation_weight = 0.01,activation function is relu, Optimization function is Adam, learning rate is 0.1,loss function is binary_crossentropy, batch_size = 64,epochs = 50, class_weight={0:0.1,1:0.7}. Model evaluation Use the test set to evaluate the performance of two models, with specific evaluation indicators including recall, specificity, F1-score, accuracy and AUC (area under the subject working characteristic curve) (Table 2 ). Among models, XGB has the highest AUC of 0.964, followed closely by NN with an AUC of 0.954. The difference between RF and SVM is not significant, with values of 0.907 and 0.924, respectively. The AUC of XGB is 0.036 points higher than the average level of the three algorithms (0.928). Figure 4 visualizes the comparison of four models. Table 2 Performance evaluation table of the model model Number of features Recall Specificity F1-Score Accuracy AUC NN 15 0.888 0.910 0.827 0.905 0.954 RF 15 0.836 0.840 0.727 0.839 0.907 SVM 15 0.844 0.847 0.738 0.846 0.924 XGB 15 0.912 0.910 0.840 0.911 0.964 Explanation of XGB Model with the SHAP method The SHAP algorithm was used to obtain the importance of each predictor variable to the outcome predicted by the XGBoost model. The variable importance plot lists the most significant variables in a descending order (Fig. 5 A). The PDW had the strongest predictive value for all prediction horizons, followed quite closely by the LYMPH, BASO%, MPV, AGE, and WBC. Furthermore, to detect the positive and negative relationships of the predictors with the target result, SHAP values were applied to uncover the risk factors. As presented in Fig. 5 B, the horizontal location shows whether the effect of that value is associated with a higher or lower prediction and the color shows whether that variable is high (in red) or low (in blue) for that observation; we can see that increases in the average PDW has a positive impact and push the prediction toward cervical cancer, whereas increases in BASO% has a negative impact and push the prediction toward noncervical cancer. Discussion Diagnosing of cervical cancer is crucial because it does not exhibit specific early symptoms. Most women seek medication at an advanced stage of this cancer, which makes treatment more complicated and places a significant financial and psychological burden on patients[ 22 ]. The present study investigated the important predictors and the most popular algorithms for predicting cervical cancer. Arif et al established Artificial Neural Network (ANN), Decision Tree(DT), Random Forest(RF) and Logistic Model Tree (LMT) model based on 858 subjects in which the RF model was the best classifier for detecting cervical cancer models [ 23 ]. Bogani et al. established a Cox regression model based on 1503 subjects, and its performance in predicting CIN2 + lesions reached more than 0.7[ 24 ]. Weegar et al. established support vector (SVM), naive Bayes (NB), Random Forest (RF) and other machine learning cervical cancer diagnosis models based on hospital electronic medical data based on 17533 subjects (including 1723 patients with cervical cancer). The AUC value of the optimal model was 0.91 and the accuracy was 0.71[ 25 ]. Based on 1523 research subjects and the established logistic regression prediction model, Lee et al. achieved an accuracy of 0.875 in predicting CIN2+[ 26 ].Rothberg et al. established a generalized linear regression model based on 99,319 subjects to predict the risk of cervical cancer. The AUC value predicted by this model was 0.81[ 27 ]. The above study shows that different machine learning models are strongly influenced by the quantity and quality of data when predicting cervical cancer related risks. In this study, we attempted for the first time to bulid a cervical cancer prediction model based on blood routine. The results showed that the XGB model performed the best in assessing the risk of cervical cancer, with AUC, accuracy, sensitivity, specificity and F1-Score above 0.9. This may be because the sample size and feature number of the data used in this study are both high, which fully shows the performance advantage of XGB model in processing high-dimensional and large-sample data. At present, human papillomavirus (HPV), especially high-risk HPV virus infection, is known to be the main factor inducing cervical cancer [ 28 , 29 ]. At the same time, many factors such as age, early marriage and early pregnancy, premature birth and fertility, premature, smoking, personal health level, social status, education level, long-term use of oral contraceptives, cytological examination, pathological examination are also closely related to the incidence of cervical cancer[ 30 – 33 ]. Studies have shown a correlation between cervical cancer and blood tests. Mota et al validated parameters of blood counts and fasting glucose levels prior to cervical cancer treatment as associated with prognosis and survival (IIB-IVB staging)[ 34 ]. Studies at home and abroad have shown that there is a close relationship between anemia and the prognosis of cervical cancer[ 35 ]. Deng et al showed that the mean platelet volume/platelet count ratio (MPV/PC) based on preoperative peripheral MPV and PC can be used to predict the prognosis of multiple malignant tumors[ 36 ]. In this study, the explanatory analysis of SHAP model and the weight importance analysis of lasso feature showed that PDW、LYMPH、BASO%、MPV、AGE and WBC had the greatest influence on the prediction of cervical cancer. The next step is to increase the sample size and enrich the feature dimensions to improve the predictive performance of the model. In future research, we will consider integrating more types of data to better adapt the model to real clinical scenarios. At the same time, we will conduct an in-depth exploration of potential features related to the risk of cervical cancer, providing new opportunities for further improvement of model performance. Conclusions This study described the application of blood routine data based ML in patients with cervical cancer, generating algorithm models that reliably predicts the probability of cervical cancer. In our study, ML combined with the explainability method of SHAP makes the black box model of ML explainable, which is more suitable for predicting cervical cancer. The use of the SHAP value provides transparency and interpretability to ML models. Abbreviations CC: Cervical cancer; ML: Machine Learning; KNN: K Nearest Neighbors, LR: logistic regression, DT: decision tree, SVM: supporting vector machine, ANN: Artificial Neural Network; STDs: sexually transmitted disease; MSE: mean squared error; FPR: false-positive rate; FNR: false-negative rate; DT: Decision Tree; RF: Random Forest; LR: Logistic Regression; MLP: Multi-layer Perceptron; NB: Naive Bayes; XGB :Extreme Gradient Boosting; TP: True Positive; TN: True Negative; FP: False Positive; RBC: Red Blood Cell Count; MCV: Mean Corpuscular; PDW: Platelet Distribution Width; WBC: White Blood Cell Count; NEUT%:Neutrophil Percentage; LYMPH: Lymphocyte; EO%: Eosinophil Percentage; BASO: Basophil; NEUT: Neutrophil; HGB: Hemoglobin; HCT: Hematocrit; MCH: Mean Corpuscular Hemoglobin; MCHC: Mean Corpuscular Hemoglobin Concentration ; R-CV: Red Cell Distribution Width; PLT: Platelet Count; MPV: Mean Platelet Volume; PCT: platelet count; MONO: Monocyte Absolute Value ; EO: Eosinophil Absolute Value; Shapley additive interpretation (SHAP) Declarations Ethics approval and consent to participate Approval for the study was given from the Peking University Cancer Hospital-Inner Mongolia Campus/Affiliated Cancer Hospital of Inner Mongolia Medical University/Inner Mongolia Autonomous Region Cancer Center Gynecological oncology review board. All patients included in the study signed an informed consent form, authorizing the collection and use of their data, at enrollment time. All the authors have approved the manuscript and agree with submission to your esteemed journal. Consent for publication Not applicable. Competing interests The authors declare that they have no competing interests. Funding This study was supported by the Project of Revitalizing Mongolia through Science and Technology (2021- Revitalizing Mongolia through Science and Technology - Independent innovation demonstration zone − 01). Author Contribution Jie Su performed the research and wrote the original manuscript; Hui Lu is responsible for data collection and interpretation; Huanrui Zhang analysed the data; Na Cui is responsible for data collection; Chao Chen prepared figures 1-5; Qin Si is responsible for data collection and interpretation; Biao Song designed the research. The authors read and approved the final manuscript. Data Availability The datasets analyzed during the current study are not publicly available due to privacy but are available from the corresponding author on reasonablerequest. References Gaffney DK, Hashibe M, Kepka D, Maurer KA, Werner TL. Too many women are dying from cervix cancer: Problems and solutions. Gynecol Oncol. 2018;151(3):547–54. Yang X, Li Y, Tang Y, Li Z, Wang S, Luo X, He T, Yin A, Luo M. Cervical HPV infect cxion in Guangzhou, China: an epidemiological study of 198,111 women from 2015 to 2021. Emerg Microbes Infect. 2023;12(1):e2176009. Yuan M, Zhao X, Wang H, Hu S, Zhao F. Trend in Cervical Cancer Incidence and Mortality Rates in China, 2006–2030: A Bayesian Age-Period-Cohort Modeling Study. Cancer Epidemiol Biomarkers Prev. 2023;32(6):825–33. Koliopoulos G, Nyaga VN, Santesso N, Bryant A, Martin-Hirsch PP, Mustafa RA, Schünemann H, Paraskevaidis E, Arbyn M. Cytology versus HPV testing for cervical cancer screening in the general population. Cochrane Database Syst Rev. 2017;8(8):Cd008587. Goel G, Halder A, Joshi D, Anil AC, Kapoor N. Rapid, Economic, Acetic Acid Papanicolaou Stain (REAP): An Economical, Rapid, and Appropriate Substitute to Conventional Pap Stain for Staining Cervical Smears. J Cytol. 2020;37(4):170–3. Yang CM, Sung FC, Hsue CS, Muo CH, Wang SW, Shieh SH. Comparisons of Papanicolaou Utilization and Cervical Cancer Detection between Rural and Urban Women in Taiwan. Int J Environ Res Public Health. 2020;18(1). Mehmood M, Rizwan M, Gregus Ml M, Abbas S. Machine Learning Assisted Cervical Cancer Detection. Front Public Health. 2021;9:788376. Huang S, Yang J, Shen N, Xu Q, Zhao Q. Artificial intelligence in lung cancer diagnosis and prognosis: Current application and future perspective. Semin Cancer Biol. 2023;89:30–7. Sahu A, Das PK, Meher S. Recent advancements in machine learning and deep learning-based breast cancer detection using mammograms. Phys Med. 2023;114:103138. Tharwat M, Sakr NA, El-Sappagh S, Soliman H, Kwak KS, Elmogy M. Colon Cancer Diagnosis Based on Machine Learning and Deep Learning: Modalities and Analysis Techniques. Sens (Basel). 2022;22(23):9250. Lynch CM, Abdollahi B, Fuqua JD, de Carlo AR, Bartholomai JA, Balgemann RN, van Berkel VH, Frieboes HB. Prediction of lung cancer patient survival via supervised machine learning classification techniques. Int J Med Inf. 2017;108:1–8. Majumder A, Sen D. Artificial intelligence in cancer diagnostics and therapy: current perspectives. Indian J Cancer. 2021;58(4):481–92. Mäkitie AA, Alabi RO, Ng SP, Takes RP, Robbins KT, Ronen O, Shaha AR, Bradley PJ, Saba NF, Nuyts S, Triantafyllou A, Piazza C, Rinaldo A, Ferlito A. Artificial Intelligence in Head and Neck Cancer: A Systematic Review of Systematic Reviews. Adv Ther. 2023;40(8):3360–80. Bychkov D, Linder N, Turkki R, Nordling S, Kovanen PE, Verrill C, Walliander M, Lundin M, Haglund C, Lundin J. Deep learning based tissue analysis predicts outcome in colorectal cancer. Sci Rep. 2018;8(1):3395. Chaudhary K, Poirion OB, Lu L, Garmire LX. Deep Learning-Based Multi-Omics Integration Robustly Predicts Survival in Liver Cancer. Clin Cancer Res. 2018;24(6):1248–59. Narayanan DL, Saladi RN, Fox JL. Ultraviolet radiation and skin cancer. Int J Dermatol. 2010;49(9):978–86. Esteva A, Kuprel B, Novoa RA, Ko J, Swetter SM, Blau HM, Thrun S. Dermatologist-level classification of skin cancer with deep neural networks. Nature. 2017;542:115–8. Mehmood M, Rizwan M, Gregus Ml M, Abbas S. Machine Learning Assisted Cervical Cancer Detection. Front Public Health. 2021;9:788376. Devi S, Gaikwad SR. Prediction and Detection of Cervical Malignancy Using Machine Learning Models. Asian Pac J Cancer Prev. 2023;24(4):1419–33. Al Mudawi N, Alazeb A. A Model for Predicting Cervical Cancer Using Machine Learning Algorithms. Sens (Basel). 2022; 22(11). Lundberg SM, Lee SI. A unified approach to interpreting model predictions. Neural Information Processing Systems. NIPS 2017: Proceedings of the 31st International Conference on Neural Information Processing Systems; 2017; Long Beach, CA, USA. Red Hook: Curran Associates; 2017;4768–4777. Schiffman M, Castle PE, Jeronimo J, Rodriguez AC, Wacholder S. Human papillomavirus and cervical cancer. Lancet. 2007;370(9590):890–907. Arif -Ul-Islam A, Ripon SH, Bhuiyan NQ. Cervical Cancer Risk Factors: Classification and Mining Associations. APTIKOM J Comput Sci Inform Technol. 2019;4(1):8–18. Bogani G, Tagliabue E, Ferla S, Martinelli F, Ditto A, Chiappa V, Leone Roberti Maggiore U, Taverna F, Lombardo C, Lorusso D, et al. Nomogram-based prediction of cervical dysplasia persistence/recurrence. Eur J Cancer Prev. 2019;28(5):435–40. Weegar R, Sundström K. Using machine learning for predicting cervical cancer from Swedish electronic health records by mining hierarchical representations. PLoS ONE. 2020;15(8):e0237911. Lee CH, Peng CY, Li RN, Chen YC, Tsai HT, Hung YH, Chan TF, Huang HL, Lai TC, Wu MT. Risk evaluation for the development of cervical intraepithelial neoplasia: development and validation of risk-scoring schemes. Int J Cancer. 2015;136(2):340–9. Rothberg MB, Hu B, Lipold L, Schramm S, Jin XW, Sikon A, Taksler GB. A risk prediction model to allow personalized screening for cervical cancer. Cancer Causes Control. 2018;29(3):297–304. zur Hausen H. Papillomaviruses and cancer: from basic studies to clinical application. Nat Rev Cancer. 2002;2(5):342–50. Muñoz N, Bosch FX, de Sanjosé S, Herrero R, Castellsagué X, Shah KV, Snijders PJ, Meijer CJ. Epidemiologic classification of human papillomavirus types associated with cervical cancer. N Engl J Med. 2003;348(6):518–27. C FA. Supervised Algorithms of Machine Learning for the Prediction of Cervical Cancer. J Biomed Phys Eng. 2020;10(4):513–22. Nithya B, Ilango V. Evaluation of machine learning based optimized feature selection approaches and classification methods for cervical cancer prediction. SN Appl Sci. 2019;1(6):1–16. Ouahab IBA, Khamlichi SE, Bouhorma M, Sedqui A. PREDICTION OF CERVICAL CANCER RISK USING MACHINE LEARNING. 20w21. Geetha R, Sivasubramanian S, Kaliappan M, Vimal S, Annamalai S. Cervical Cancer Identification with Synthetic Minority Oversampling Technique and PCA Analysis using Random Forest Classifier. J Med Syst. 2019;43(9):286. Mota SDS, Otaño SS, Murta EFC, Nomelini R. Blood count and fasting blood glucose level in the assessment of prognosis and survival in advanced cervical cancer. Rev Assoc Med Bras. 2022;68(2):234–8. Wassie M, Aemro A, Fentie B. Prevalence and associated factors of baseline anemia among cervical cancer patients in Tikur Anbesa Specialized Hospital, Ethiopia. BMC Womens Health. 2021;21(1):36. Deng Q, Long Q, Liu Y, Yang Z, Du Y, Chen X. Prognostic value of preoperative peripheral blood mean platelet volume/platelet count ratio (MPV/PC) in patients with resectable cervical cancer. BMC Cancer. 2021;21(1):1282. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4761322","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":336044393,"identity":"0fa0b249-8c7e-4587-b595-a134fc94d671","order_by":0,"name":"Jie Su","email":"","orcid":"","institution":"Inner Mongolia Medical University","correspondingAuthor":false,"prefix":"","firstName":"Jie","middleName":"","lastName":"Su","suffix":""},{"id":336044395,"identity":"12d75988-20b1-4c55-a3e0-a7de2ae1ad02","order_by":1,"name":"Hui Lu","email":"","orcid":"","institution":"Inner Mongolia University","correspondingAuthor":false,"prefix":"","firstName":"Hui","middleName":"","lastName":"Lu","suffix":""},{"id":336044397,"identity":"09887dcf-4d2d-4b52-9360-6e09c38faf45","order_by":2,"name":"RuiHuan Zhang","email":"","orcid":"","institution":"Medical Intelligent Diagnostics Big Data Research Institute","correspondingAuthor":false,"prefix":"","firstName":"RuiHuan","middleName":"","lastName":"Zhang","suffix":""},{"id":336044398,"identity":"5d4f7335-fdb2-4e64-9885-3c823d1e4992","order_by":3,"name":"Na Cui","email":"","orcid":"","institution":"Peking University Cancer Hospital (Inner Mongolia Campus/Affiliated Cancer Hospital of Inner Mongolia Medical University, Inner Mongolia Autonomous Region Cancer Center Gynecological oncology)","correspondingAuthor":false,"prefix":"","firstName":"Na","middleName":"","lastName":"Cui","suffix":""},{"id":336044399,"identity":"913602d8-a066-42ce-9176-f69a5c33750f","order_by":4,"name":"Chao Chen","email":"","orcid":"","institution":"Medical Intelligent Diagnostics Big Data Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Chao","middleName":"","lastName":"Chen","suffix":""},{"id":336044401,"identity":"de3218b1-0fb0-4914-bff9-3d23153c4a31","order_by":5,"name":"Qin Si","email":"","orcid":"","institution":"Peking University Cancer Hospital (Inner Mongolia Campus/Affiliated Cancer Hospital of Inner Mongolia Medical University, Inner Mongolia Autonomous Region Cancer Center Gynecological oncology)","correspondingAuthor":false,"prefix":"","firstName":"Qin","middleName":"","lastName":"Si","suffix":""},{"id":336044402,"identity":"7a13bd20-4d35-41a0-9a0f-2f3d08714035","order_by":6,"name":"Biao Song","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5ElEQVRIie3PsQrCMBCA4SsBuwRdA4p9hYNAESz1WUohU4fuDkYcuujex/ARGoO6FOcOboKLS8HFwUGruLZ1E8w/HATugwuAyfSjEcDndJTMSvT8Lwgjc5XGImxJqpi90LTcWLJp3UlW+hLH/hD6c6k9zAjYeruuI5gfxDjFkMNASR3hsQtUiKKWsMjlFEkgIajImQCjbi1x0heZvckItSWbCBQRP1HUgWRPAm0I5rlLUtxzoEqqJYqw0/QXJ1nya3yfDsFOTuXt7vk9W+/qDwPosGpOss+zYb2KlC2WTCaT6Z97AJu2SyLLdbOhAAAAAElFTkSuQmCC","orcid":"","institution":"Medical Intelligent Diagnostics Big Data Research Institute","correspondingAuthor":true,"prefix":"","firstName":"Biao","middleName":"","lastName":"Song","suffix":""}],"badges":[],"createdAt":"2024-07-18 09:11:51","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4761322/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4761322/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":63024130,"identity":"ed6856a9-50c3-4538-9c2a-6a2de655362d","added_by":"auto","created_at":"2024-08-22 08:06:30","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":107529,"visible":true,"origin":"","legend":"\u003cp\u003eML model training process for cervical cancer classification.\u003c/p\u003e","description":"","filename":"Figure1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4761322/v1/906dc0a17b807bdec654835d.jpg"},{"id":63024131,"identity":"87249d05-c62e-45f7-9add-4482e92ba29f","added_by":"auto","created_at":"2024-08-22 08:06:30","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":408648,"visible":true,"origin":"","legend":"\u003cp\u003eFeature selection for cervical cancer dataset. A. Lasso regression for feature selection on datasets; B. Random Forest for feature selection of Datasets; C. Venn diagram shows the feature selection of Lasso and RF, identifying the intersection of the two methods.\u003c/p\u003e","description":"","filename":"Figure2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4761322/v1/819fff6c8fb39fb4c8d99bdf.jpg"},{"id":63024784,"identity":"26f42543-3901-45bc-bcc3-d26d65140a80","added_by":"auto","created_at":"2024-08-22 08:14:30","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":83950,"visible":true,"origin":"","legend":"\u003cp\u003eStructure diagram of the machine model for cervical cancer classification. (A) In XGB, the y-value corresponding to one feature is used as the input value to predict the next feature, and the last y-value in the process is combined with the weight. (B) RF creates multiple decision trees, each resulting in a category, and synthesizes the results of each tree to determine the final category. (C) The SVM looks for a hyperplane with maximum interval in the feature space to perform linear classification of the data, and (D) the neural network learns the complex mapping relationship between the input and output through the connected layers and nodes, and outputs the final category.\u003c/p\u003e","description":"","filename":"Figure3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4761322/v1/f72ffdba24802be4b7011456.jpg"},{"id":63024133,"identity":"ff8a1824-bb75-4e34-8cbe-9333b30fc6d1","added_by":"auto","created_at":"2024-08-22 08:06:30","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":165222,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of model performance. a. Comparison of Accuracy. b. Comparison of Recall. c. Comparison of Specificity. d. Comparison of AUC. e. Comparison of F1-Score\u003c/p\u003e","description":"","filename":"Figure4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4761322/v1/43b59fa9227d56aa395075b2.jpg"},{"id":63024146,"identity":"1cab57a1-e36b-46a8-a276-e612c6f75085","added_by":"auto","created_at":"2024-08-22 08:06:32","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":58410,"visible":true,"origin":"","legend":"\u003cp\u003eSHAP explains the results of the XGB model. A. Bar chart of the mean absolute SHAP value for each predictor; B. Summarize the overall distribution of SHAP values for each feature.\u003c/p\u003e","description":"","filename":"Figure5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4761322/v1/be92e18b19bf65711dca063a.jpg"},{"id":65581460,"identity":"1daf9fd5-97fd-4def-9165-13b3f8e63cb8","added_by":"auto","created_at":"2024-09-30 08:24:27","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1399137,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4761322/v1/69ecf38b-cadb-4733-b7f9-16ba3d49601c.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Cervical cancer prediction using machine learning models based on blood routine analysis","fulltext":[{"header":"Introduction","content":"\u003cp\u003eCervical Cancer Epidemiology and Risk Factors Cervical cancer (CC) is one of the leading causes of death among women worldwide[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Cervical cancer accounts for 6.5 per cent of all cancers in women. In China, cervical cancer ranks the sixth among female malignant tumors and the eighth among the causes of cancer death [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e], with up to 130,000 new cases per year, accounting for about 1/4 of the total number of new cases in the world[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Cervical cancer destroys deep tissues of the cervix and can gradually reach other areas of the human body, such as the lungs, liver, and vagina. Early cervical cancer may not have obvious signs and symptoms, which is the main reason for late detection of cervical cancer in patients. The key of prevention and treatment of cervical cancer is early detection, diagnosis and treatment. Cervical cancer screening can effectively detect precancerous lesions early and intervene in time. Current screening techniques include Papanicolaou (Pap) cytology-based test, Acetic Acid Papanicolaou Stain, liquid-based cytology, and HPV testing techniques [\u003cspan additionalcitationids=\"CR5\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Recently, many studies have been conducted on cervical cancer using modern techniques that provide prediction in the early stage.\u003c/p\u003e \u003cp\u003eBlood routine is a common hematological index used in bedside included Red Blood Cell Count (RBC), White Blood Cell Count (WBC), Hemoglobin (HGB), Platelet Count (PLT) and so on. In the last several years, the Hemoglobin, Albumin, Lymphocyte, Platelet Score (HALP) has emerged in the literature as a new prognostic biomarker that has been used to predict a number of clinical outcomes in the context of various neoplasms [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eMachine learning is to learn our existing data through certain logic, understand its internal connection and distribution, mine hidden information that we can't see by the naked eye, and predict, classify and evaluate the things we study to a certain extent through the result output, so as to give full play to the maximum value of the data and realize the leap from data to value. Various machine learning (ML) techniques have been used to detect various cancer types including lung, brain, breast, skin, blood, hippocampus, pancreas, cardiac, colon cancer, hepatic vessel, spleen etc [\u003cspan additionalcitationids=\"CR9 CR10 CR11 CR12\" citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Bychkov et al. combine Surface-enhanced laser desorption ionization-time of flight mass spectrometry (SELDI-TOF) mass spectrometry and artificial neural networks to analyze serum protein with high production sensitivity and specificity values for detecting and diagnosing colorectal cancer [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Chaudhary et al. used a novel approach using Deep Learning to identify multi-omics features of survival in patients with liver cancer and apply to future prognostic predictions of liver cancer [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. Machine learning can classify clinical images of skin diseases, including intraepithelial carcinoma, actinic keratosis, seborrheic keratosis, lentigo, pyogenic granuloma et al. [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eA number of mechanistic algorithms have been applied to the mental screening and diagnosis of cervical cancer, including the K Nearest Neighbors (KNN), logistic regression (LR), decision tree (DT), supporting vector machine (SVM), Artificial Neural Network (ANN) and related derivation algorithm. Mehmood et al. proposes an approach named CervDetect that uses machine learning algorithms to evaluate the risk elements of malignant cervical formation. CervDetect select significant features including Age, IUD (intrauterine derice), smokes, STDs (sexually transmitted disease), and so on. CervDetect achieved an accuracy of 93.6%, mean squared error (MSE) error of 0.07111, false-positive rate (FPR) of 6.4%, and false-negative rate (FNR) of 100%[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Devi et al. used Decision Tree (DT), Random Forest (RF), Logistic Regression (LR), Multi-layer Perceptron (MLP) and Naive Bayes (NB) to find out and analyze the predictors of Cervical Malignancy and evaluate the unbalanced input and target datasets. The predictors used in the current study included age, smokes, Infection of the cervix et al. Logistic Regression and Decision Tree algorithms were identified as the best-performed ML (Machine Learning) algorithm classifiers to detect the significant predictors[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Mudawi and Alazeb utilized random forest (RF), decision tree (DT), adaptive boosting, and gradient boosting algorithms to predict cervical cancer. The cancer detection dataset includes 32 variables, such as age, IUD, smokes, STDs, and so on[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eShapley additive interpretation (SHAP) is a model interpretation tool based on Shapley values, which is a post-interpretable scheme. SHAP can explain predictions by computing the contributions of individual variables, accounting for local accuracy, missingness, and consistency, to formulate absolute magnitude and directionality of impact in the prediction of the desired outcome.\u003c/p\u003e \u003cp\u003eThe purpose of our study is to develop an interpretable model to predict the risk for patients with cervical cancer based on the blood routine. In addition, the SHAP method is used to explain the odel and explore factors for cervical cancer.\u003c/p\u003e"},{"header":"Materials and methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003ePatients\u003c/h2\u003e \u003cp\u003eA retrospective study was conducted by collecting medical records of patients from hospitals from 2013 to 2023, with cervical cancer data as the case group and noncervical cancer patient data as the control group. According to the inclusion and exclusion criteria, 2533 cases were included in the case group, and 5011 cases were included in the control group. The dataset mainly collects demographics, clinical characteristics, and baseline data related to cervical cancer, including age, pathological type, clinical diagnosis, and routine test data.\u003c/p\u003e \u003cp\u003e The inclusion criteria were as follows: (1) Patients who meet the diagnostic criteria for cervical cancer in the \"Guidelines for Standardized Treatment of Cervical Cancer and Precancerous Lesions (Trial)\"; (2) The clinical staging of the patient is stages I to II, with concurrent radical hysterectomy and bilateral lymph node dissection; (3) Patients with complete clinical and pathological data, and included for the first time;\u003c/p\u003e\u003cp\u003e(4) By persistently accepting medical advice and conducting regular follow-up.\u003c/p\u003e \u003cp\u003eThe exclusion criteria were as follows: (1) Pregnant or lactating women; (2) Patients with gynecological malignancies or a history of malignancies in other areas; (3) Patients with coagulation dysfunction; (4) Patients with a history of anti-tumor treatment such as radiotherapy and chemotherapy.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eEthical Considerations\u003c/h2\u003e \u003cp\u003eEthical approval and individual patient consent was not necessary because all the protected health information was anonymized.\u003c/p\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eData preprocessing\u003c/h2\u003e \u003cp\u003eOrganize and preprocess the collected patient dataset, including data cleaning, feature selection, and data normalization. Due to the fact that this study was conducted on cervical cancer patients, the gender was set to female and the age range was set to 18\u0026ndash;90 years old.\u003c/p\u003e \u003cp\u003eThe purpose of data cleaning is to improve the quality of the dataset by removing duplicate, invalid, and abnormal data, deleting data with missing or duplicate values, and ensuring the consistency of data units.\u003c/p\u003e \u003cp\u003eFeature selection is the process of selecting the subset of features with the most predictive and imaging power from a dataset. The specific operation is to select features based on age and conventional test data, mainly using lassoregression and random forest methods.\u003c/p\u003e \u003cp\u003eData normalization is the process of converting raw data into data with unified standards, thereby eliminating dimensional differences between different features, enhancing the stability of the model training process, and improving the model's generalization ability. The Z-Score normalization method is mainly used, which converts the data into a standard normal distribution with a mean of 0 and a standard deviation of 1. After Z-score standardization, all data is converted into dimensionless standardized values, so that even if the original data has different scales or units, it can be compared and analyzed on the same scale.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eResearch and Design of Machine Learning Analysis\u003c/h2\u003e \u003cp\u003eFirstly, divide the dataset into cervical cancer and noncervical cancer groups based on clinical diagnosis, and then screen the obtained features, which mainly include demographic, clinical, and baseline data information. After extracting 23 features from the preprocessed dataset in the feature selection stage, the optimal feature combination is selected through lassoregression and random forest model. In machine learning models, this study uses Extreme Gradient Boosting (XGB), Support Vector Machine (SVM), Random Forest (RF), and Neural Network (NN) to train classification from selected features.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eEvaluation process\u003c/h2\u003e \u003cp\u003eCross validation is an evaluation method that prevents overflow and improves accuracy when evaluating model performance. This study used tenfold cross validation to verify the classification performance of the algorithm. To ensure comparability of results, each method used the same training and testing sets.\u003c/p\u003e \u003cp\u003eUsually, binary classification algorithms are evaluated based on basic decision scores. The scores include True Positive (TP) as the number of positive samples correctly classified as positive, True Negative (TN) as the number of negative samples correctly classified as negative, False Positive (FP) as the number of negative samples misclassified as positive, and False Negative (FN) as the number of positive samples incorrectly classifed as negative.\u003c/p\u003e \u003cp\u003eUsing the above scores, evaluated each model using the following metrics. Recall rate, also known as sensitivity, is a measure of the model's ability to detect patients with illnesses. Specificity is the ability of a model to detect disease-free individuals. The F1 score is the harmonic mean of precision and recall. Accuracy refers to the proportion of total predicted misclassifications that are true, and the proportion of total predicted misclassifications that are incorrect.\u003c/p\u003e \u003cp\u003eRecall\u0026thinsp;=\u0026thinsp;TP /(TP\u0026thinsp;+\u0026thinsp;FN) (1)\u003c/p\u003e \u003cp\u003eSpecificity\u0026thinsp;=\u0026thinsp;TN/ (TN\u0026thinsp;+\u0026thinsp;FP) (2)\u003c/p\u003e \u003cp\u003eF1score = (2 *precision*recall)/(precision\u0026thinsp;+\u0026thinsp;recall) (3)\u003c/p\u003e \u003cp\u003eAccuracy = (TP\u0026thinsp;+\u0026thinsp;TN)/(TP\u0026thinsp;+\u0026thinsp;FN\u0026thinsp;+\u0026thinsp;TN\u0026thinsp;+\u0026thinsp;FP) (4)\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eModel explainability\u003c/h2\u003e \u003cp\u003eSHapley Additive exPlanations (SHAP) was used as explainability model, \u0026ldquo;\u003cem\u003eshap 0.39.0\u003c/em\u003e\u0026rdquo; package, to give local explanations and allow computation of the contribution of each variable to the prediction for each individual [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. SHAP is an interpretable method based on game theory, drawing on the concept of Shapley values, aimed at quantifying the contribution of each feature to a specific predictive output. By analyzing the Shapley values of each feature, we can determine how each part of the machine learning model's prediction results affects the final prediction. A positive SHAP value indicates that the feature improves the result value and has a positive effect, while a negative SHAP value indicates that the feature reduces the result value and has a negative impact. This method can output the importance ranking of features and the relationship between features and results.\u003c/p\u003e \u003c/div\u003e"},{"header":"Result","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003ePatient characteristics\u003c/h2\u003e \u003cp\u003eAmong 6297 patients with cervical cancer, a total of 2503 adult patients diagnosed with cervical cancer were included in the final cohort for this study. The patient screening process is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e showed 6297 cervical cancer patients (average age, 52.81 years; range,18\u0026ndash;90 years) from 2013 to 2023 were identified, and 2503 (39.7%) diagnosed with cervical cancer and 3794 (60.3%) non cervical cancer patients.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eStatistical Table of the Dataset\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003efeature\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTotal number\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eaverage\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eStandard error\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003emedian\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003estandard deviation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eminimum value\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eMaximum value\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAGE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e52.818\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.177\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e14.071\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRBC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.023\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.008\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4.14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.670\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e7.23\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMCV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e90.262\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.093\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e90.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e7.389\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e56.11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e139.9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePDW\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e14.142\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.037\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e15.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.908\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e23.44\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWBC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.424\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.044\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5.81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e3.475\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e60.38\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNEUT%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e64.174\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.172\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e64.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e13.642\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e98.7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLYMPH%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e25.999\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.148\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e25.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e11.722\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e90\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEO%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.941\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.035\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.756\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e65.4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBASO%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.328\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.005\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.432\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e13.2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNEUT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.311\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.039\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3.59\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e3.058\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e53.76\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLYMPH\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.535\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.013\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.998\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e36.8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBASO\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.019\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.039\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e1.59\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHGB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e117.534\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.284\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e122\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e22.558\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e178\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHCT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.361\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.371\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.058\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.081\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e0.556\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMCH\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e29.551\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.037\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e29.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.965\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e15.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e79.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMCHC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e327.044\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.168\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e328\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e13.309\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e249\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e697\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eR-CV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e14.029\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.031\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e13.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.470\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.117\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e34.93\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePLT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e222.597\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1.069\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e215\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e84.816\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e1547\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMPV\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e9.331\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.019\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e9.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1.491\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e15.7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePCT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.205\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.077\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e1.11\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMONO\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.440\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.006\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.468\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e32.18\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMONO%\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e7.219\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.043\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e6.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e3.426\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e61.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEO\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.114\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.003\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.219\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e8.14\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eRBC: Red Blood Cell Count; MCV: Mean Corpuscular; PDW: Platelet Distribution Width; WBC: White Blood Cell Count; NEUT%:Neutrophil Percentage; LYMPH: Lymphocyte; EO%: Eosinophil Percentage; BASO: Basophil; NEUT: Neutrophil; HGB: Hemoglobin; HCT: Hematocrit; MCH: Mean Corpuscular Hemoglobin; MCHC: Mean Corpuscular Hemoglobin Concentration ; R-CV: Red Cell Distribution Width; PLT: Platelet Count; MPV: Mean Platelet Volume; PCT: platelet count; MONO: Monocyte Absolute Value ; EO: Eosinophil Absolute Value;\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eFeature selection\u003c/h2\u003e \u003cp\u003eThis study used lasso regression and random forest methods to select features from the dataset. Use lasso regression to perform feature selection on the dataset, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA. Features with values greater than zero in the graph have a positive linear relationship, while features with values less than zero have a negative linear relationship. Finally, 19 features were selected from the 23 features. The specific positive linear relationships include: EO, MONO%, HGB, BASO, LYMPH, NEUT, LYMPH%, NEUT%, WBC, RBC, AGE; Negative linear relationships include: PCT, MPV, R-CV, MCHC, MCH, BASO%, EO%, PDW. Figure\u0026nbsp;23B shows the use of random forest for feature selection on the dataset. The top 19 features were selected based on their importance, including LYMPH, PDW, MPV, BASO%, MONO, RBC, R-CV, WBC, MCHC, AGE, HCT, MONO%, MCV, BASO, HGB, PCT, LYMPH%, PLT and NEUT%. Based on the 19 features selected by each of the two methods, a common feature set was extracted,and finally 15 features were selected to participate in subsequent model training, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eMachine learning classification architectures\u003c/h2\u003e \u003cp\u003eFor machine learning classification, this study used the Extreme Gradient Boosting (XGB), Support Vector Machine (SVM), Random Forest (RF), and Neural Network (NN) architectures. XGB creates a powerful predictive model by iteratively constructing and combining multiple weak learners. As shown in Figure.3A, the pre pruning method is used to compensate for the error of the previous tree and create the next tree. The RF model as shown in Figure.3B is a collection of multiple decision trees, each of which predicts the data and ultimately integrates the results of all decision trees through voting to obtain classification results. As shown in Figure.3c, when two types of data are classified, SVM attempts to find a hyperplane. This hyperplane not only correctly divides samples of different categories, but also maximizes the distance from the closest sample point to the hyperplane. Neural network is a computational model that simulates the neural structure of the human brain, consisting of multiple nodes or neurons that transmit and process data through weighted connections, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD.With all of these models, a default value is designated for each parameter. Default values for the main parameters of the XGB model are: booster\u0026thinsp;=\u0026thinsp;gbtree, n_estimators\u0026thinsp;=\u0026thinsp;10, learning rate eta\u0026thinsp;=\u0026thinsp;0.2, class _weights= {0:0.06,1:0.9},and max default\u0026thinsp;=\u0026thinsp;6. The SVM\u0026rsquo;s main parameter default values are: kernel\u0026thinsp;=\u0026thinsp;RBF, probability\u0026thinsp;=\u0026thinsp;True, class_weight={0:0.1,1:0.25}, and the kernel factor gamma\u0026thinsp;=\u0026thinsp;1. Default values for the main parameters of the RF model are: criterion=\"gini\", max_depth\u0026thinsp;=\u0026thinsp;25, nestimators\u0026thinsp;=\u0026thinsp;80, min _samples _split\u0026thinsp;=\u0026thinsp;2,class_wei-ght={0:0.01,1:3.1}.The NN\u0026rsquo;s main parameter default values are: 2 hidden layers with 128 and 32 neurons respectively, the first hidden layer uses L1 regularization and regulation_weight\u0026thinsp;=\u0026thinsp;0.01,activation function is relu, Optimization function is Adam, learning rate is 0.1,loss function is binary_crossentropy, batch_size\u0026thinsp;=\u0026thinsp;64,epochs\u0026thinsp;=\u0026thinsp;50, class_weight={0:0.1,1:0.7}.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eModel evaluation\u003c/h2\u003e \u003cp\u003eUse the test set to evaluate the performance of two models, with specific evaluation indicators including recall, specificity, F1-score, accuracy and AUC (area under the subject working characteristic curve) (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Among models, XGB has the highest AUC of 0.964, followed closely by NN with an AUC of 0.954. The difference between RF and SVM is not significant, with values of 0.907 and 0.924, respectively. The AUC of XGB is 0.036 points higher than the average level of the three algorithms (0.928). Figure\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e visualizes the comparison of four models.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePerformance evaluation table of the model\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003emodel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNumber of features\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eSpecificity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eF1-Score\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eAccuracy\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.888\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.910\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.827\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.905\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.954\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.836\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.840\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.727\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.839\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.907\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSVM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.844\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.847\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.738\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.846\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.924\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXGB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.912\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.910\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.840\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.911\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.964\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eExplanation of XGB Model with the SHAP method\u003c/h2\u003e \u003cp\u003eThe SHAP algorithm was used to obtain the importance of each predictor variable to the outcome predicted by the XGBoost model. The variable importance plot lists the most significant variables in a descending order (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA). The PDW had the strongest predictive value for all prediction horizons, followed quite closely by the LYMPH, BASO%, MPV, AGE, and WBC. Furthermore, to detect the positive and negative relationships of the predictors with the target result, SHAP values were applied to uncover the risk factors. As presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB, the horizontal location shows whether the effect of that value is associated with a higher or lower prediction and the color shows whether that variable is high (in red) or low (in blue) for that observation; we can see that increases in the average PDW has a positive impact and push the prediction toward cervical cancer, whereas increases in BASO% has a negative impact and push the prediction toward noncervical cancer.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eDiagnosing of cervical cancer is crucial because it does not exhibit specific early symptoms. Most women seek medication at an advanced stage of this cancer, which makes treatment more complicated and places a significant financial and psychological burden on patients[\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. The present study investigated the important predictors and the most popular algorithms for predicting cervical cancer. Arif et al established Artificial Neural Network (ANN), Decision Tree(DT), Random Forest(RF) and Logistic Model Tree (LMT) model based on 858 subjects in which the RF model was the best classifier for detecting cervical cancer models [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. Bogani et al. established a Cox regression model based on 1503 subjects, and its performance in predicting CIN2\u0026thinsp;+\u0026thinsp;lesions reached more than 0.7[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Weegar et al. established support vector (SVM), naive Bayes (NB), Random Forest (RF) and other machine learning cervical cancer diagnosis models based on hospital electronic medical data based on 17533 subjects (including 1723 patients with cervical cancer). The AUC value of the optimal model was 0.91 and the accuracy was 0.71[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Based on 1523 research subjects and the established logistic regression prediction model, Lee et al. achieved an accuracy of 0.875 in predicting CIN2+[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e].Rothberg et al. established a generalized linear regression model based on 99,319 subjects to predict the risk of cervical cancer. The AUC value predicted by this model was 0.81[\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. The above study shows that different machine learning models are strongly influenced by the quantity and quality of data when predicting cervical cancer related risks. In this study, we attempted for the first time to bulid a cervical cancer prediction model based on blood routine. The results showed that the XGB model performed the best in assessing the risk of cervical cancer, with AUC, accuracy, sensitivity, specificity and F1-Score above 0.9. This may be because the sample size and feature number of the data used in this study are both high, which fully shows the performance advantage of XGB model in processing high-dimensional and large-sample data.\u003c/p\u003e \u003cp\u003eAt present, human papillomavirus (HPV), especially high-risk HPV virus infection, is known to be the main factor inducing cervical cancer [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. At the same time, many factors such as age, early marriage and early pregnancy, premature birth and fertility, premature, smoking, personal health level, social status, education level, long-term use of oral contraceptives, cytological examination, pathological examination are also closely related to the incidence of cervical cancer[\u003cspan additionalcitationids=\"CR31 CR32\" citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Studies have shown a correlation between cervical cancer and blood tests. Mota et al validated parameters of blood counts and fasting glucose levels prior to cervical cancer treatment as associated with prognosis and survival (IIB-IVB staging)[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. Studies at home and abroad have shown that there is a close relationship between anemia and the prognosis of cervical cancer[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. Deng et al showed that the mean platelet volume/platelet count ratio (MPV/PC) based on preoperative peripheral MPV and PC can be used to predict the prognosis of multiple malignant tumors[\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. In this study, the explanatory analysis of SHAP model and the weight importance analysis of lasso feature showed that PDW、LYMPH、BASO%、MPV、AGE and WBC had the greatest influence on the prediction of cervical cancer.\u003c/p\u003e \u003cp\u003eThe next step is to increase the sample size and enrich the feature dimensions to improve the predictive performance of the model. In future research, we will consider integrating more types of data to better adapt the model to real clinical scenarios. At the same time, we will conduct an in-depth exploration of potential features related to the risk of cervical cancer, providing new opportunities for further improvement of model performance.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eThis study described the application of blood routine data based ML in patients with cervical cancer, generating algorithm models that reliably predicts the probability of cervical cancer. In our study, ML combined with the explainability method of SHAP makes the black box model of ML explainable, which is more suitable for predicting cervical cancer. The use of the SHAP value provides transparency and interpretability to ML models.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eCC: Cervical cancer; ML: Machine Learning; KNN: K Nearest Neighbors, LR: logistic regression, DT: decision tree, SVM: supporting vector machine, ANN: Artificial Neural Network; STDs: sexually transmitted disease; MSE: mean squared error; FPR: false-positive rate; FNR: false-negative rate; DT: Decision Tree; RF: Random Forest; LR: Logistic Regression; MLP: Multi-layer Perceptron; NB: Naive Bayes; XGB :Extreme Gradient Boosting; TP: True Positive; TN: True Negative; FP: False Positive; RBC: Red Blood Cell Count; MCV: Mean Corpuscular; PDW: Platelet Distribution Width; WBC: White Blood Cell Count; NEUT%:Neutrophil Percentage; LYMPH: Lymphocyte; EO%: Eosinophil Percentage; BASO: Basophil; NEUT: Neutrophil; HGB: Hemoglobin; HCT: Hematocrit; MCH: Mean Corpuscular Hemoglobin; MCHC: Mean Corpuscular Hemoglobin Concentration ; R-CV: Red Cell Distribution Width; PLT: Platelet Count; MPV: Mean Platelet Volume; PCT: platelet count; MONO: Monocyte Absolute Value ; EO: Eosinophil Absolute Value; Shapley additive interpretation (SHAP)\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e \u003cp\u003e Approval for the study was given from the Peking University Cancer Hospital-Inner Mongolia Campus/Affiliated Cancer Hospital of Inner Mongolia Medical University/Inner Mongolia Autonomous Region Cancer Center Gynecological oncology review board. All patients included in the study signed an informed consent form, authorizing the collection and use of their data, at enrollment time. All the authors have approved the manuscript and agree with submission to your esteemed journal.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eConsent for publication\u003c/strong\u003e \u003cp\u003eNot applicable.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eCompeting interests\u003c/strong\u003e \u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eThis study was supported by the Project of Revitalizing Mongolia through Science and Technology (2021- Revitalizing Mongolia through Science and Technology - Independent innovation demonstration zone \u0026minus;\u0026thinsp;01).\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eJie Su performed the research and wrote the original manuscript; Hui Lu is responsible for data collection and interpretation; Huanrui Zhang analysed the data; Na Cui is responsible for data collection; Chao Chen prepared figures 1-5; Qin Si is responsible for data collection and interpretation; Biao Song designed the research. The authors read and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets analyzed during the current study are not publicly available due to privacy but are available from the corresponding author on reasonablerequest.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eGaffney DK, Hashibe M, Kepka D, Maurer KA, Werner TL. Too many women are dying from cervix cancer: Problems and solutions. Gynecol Oncol. 2018;151(3):547\u0026ndash;54.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang X, Li Y, Tang Y, Li Z, Wang S, Luo X, He T, Yin A, Luo M. Cervical HPV infect cxion in Guangzhou, China: an epidemiological study of 198,111 women from 2015 to 2021. Emerg Microbes Infect. 2023;12(1):e2176009.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYuan M, Zhao X, Wang H, Hu S, Zhao F. Trend in Cervical Cancer Incidence and Mortality Rates in China, 2006\u0026ndash;2030: A Bayesian Age-Period-Cohort Modeling Study. Cancer Epidemiol Biomarkers Prev. 2023;32(6):825\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKoliopoulos G, Nyaga VN, Santesso N, Bryant A, Martin-Hirsch PP, Mustafa RA, Sch\u0026uuml;nemann H, Paraskevaidis E, Arbyn M. Cytology versus HPV testing for cervical cancer screening in the general population. Cochrane Database Syst Rev. 2017;8(8):Cd008587.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoel G, Halder A, Joshi D, Anil AC, Kapoor N. Rapid, Economic, Acetic Acid Papanicolaou Stain (REAP): An Economical, Rapid, and Appropriate Substitute to Conventional Pap Stain for Staining Cervical Smears. J Cytol. 2020;37(4):170\u0026ndash;3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang CM, Sung FC, Hsue CS, Muo CH, Wang SW, Shieh SH. Comparisons of Papanicolaou Utilization and Cervical Cancer Detection between Rural and Urban Women in Taiwan. Int J Environ Res Public Health. 2020;18(1).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMehmood M, Rizwan M, Gregus Ml M, Abbas S. Machine Learning Assisted Cervical Cancer Detection. Front Public Health. 2021;9:788376.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang S, Yang J, Shen N, Xu Q, Zhao Q. Artificial intelligence in lung cancer diagnosis and prognosis: Current application and future perspective. Semin Cancer Biol. 2023;89:30\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSahu A, Das PK, Meher S. Recent advancements in machine learning and deep learning-based breast cancer detection using mammograms. Phys Med. 2023;114:103138.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTharwat M, Sakr NA, El-Sappagh S, Soliman H, Kwak KS, Elmogy M. Colon Cancer Diagnosis Based on Machine Learning and Deep Learning: Modalities and Analysis Techniques. Sens (Basel). 2022;22(23):9250.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLynch CM, Abdollahi B, Fuqua JD, de Carlo AR, Bartholomai JA, Balgemann RN, van Berkel VH, Frieboes HB. Prediction of lung cancer patient survival via supervised machine learning classification techniques. Int J Med Inf. 2017;108:1\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMajumder A, Sen D. Artificial intelligence in cancer diagnostics and therapy: current perspectives. Indian J Cancer. 2021;58(4):481\u0026ndash;92.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eM\u0026auml;kitie AA, Alabi RO, Ng SP, Takes RP, Robbins KT, Ronen O, Shaha AR, Bradley PJ, Saba NF, Nuyts S, Triantafyllou A, Piazza C, Rinaldo A, Ferlito A. Artificial Intelligence in Head and Neck Cancer: A Systematic Review of Systematic Reviews. Adv Ther. 2023;40(8):3360\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBychkov D, Linder N, Turkki R, Nordling S, Kovanen PE, Verrill C, Walliander M, Lundin M, Haglund C, Lundin J. Deep learning based tissue analysis predicts outcome in colorectal cancer. Sci Rep. 2018;8(1):3395.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChaudhary K, Poirion OB, Lu L, Garmire LX. Deep Learning-Based Multi-Omics Integration Robustly Predicts Survival in Liver Cancer. Clin Cancer Res. 2018;24(6):1248\u0026ndash;59.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNarayanan DL, Saladi RN, Fox JL. Ultraviolet radiation and skin cancer. Int J Dermatol. 2010;49(9):978\u0026ndash;86.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEsteva A, Kuprel B, Novoa RA, Ko J, Swetter SM, Blau HM, Thrun S. Dermatologist-level classification of skin cancer with deep neural networks. Nature. 2017;542:115\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMehmood M, Rizwan M, Gregus Ml M, Abbas S. Machine Learning Assisted Cervical Cancer Detection. Front Public Health. 2021;9:788376.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDevi S, Gaikwad SR. Prediction and Detection of Cervical Malignancy Using Machine Learning Models. Asian Pac J Cancer Prev. 2023;24(4):1419\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAl Mudawi N, Alazeb A. A Model for Predicting Cervical Cancer Using Machine Learning Algorithms. Sens (Basel). 2022; 22(11).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLundberg SM, Lee SI. A unified approach to interpreting model predictions. Neural Information Processing Systems. NIPS 2017: Proceedings of the 31st International Conference on Neural Information Processing Systems; 2017; Long Beach, CA, USA. Red Hook: Curran Associates; 2017;4768\u0026ndash;4777.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchiffman M, Castle PE, Jeronimo J, Rodriguez AC, Wacholder S. Human papillomavirus and cervical cancer. Lancet. 2007;370(9590):890\u0026ndash;907.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArif -Ul-Islam A, Ripon SH, Bhuiyan NQ. Cervical Cancer Risk Factors: Classification and Mining Associations. APTIKOM J Comput Sci Inform Technol. 2019;4(1):8\u0026ndash;18.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBogani G, Tagliabue E, Ferla S, Martinelli F, Ditto A, Chiappa V, Leone Roberti Maggiore U, Taverna F, Lombardo C, Lorusso D, et al. Nomogram-based prediction of cervical dysplasia persistence/recurrence. Eur J Cancer Prev. 2019;28(5):435\u0026ndash;40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeegar R, Sundstr\u0026ouml;m K. Using machine learning for predicting cervical cancer from Swedish electronic health records by mining hierarchical representations. PLoS ONE. 2020;15(8):e0237911.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee CH, Peng CY, Li RN, Chen YC, Tsai HT, Hung YH, Chan TF, Huang HL, Lai TC, Wu MT. Risk evaluation for the development of cervical intraepithelial neoplasia: development and validation of risk-scoring schemes. Int J Cancer. 2015;136(2):340\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRothberg MB, Hu B, Lipold L, Schramm S, Jin XW, Sikon A, Taksler GB. A risk prediction model to allow personalized screening for cervical cancer. Cancer Causes Control. 2018;29(3):297\u0026ndash;304.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ezur Hausen H. Papillomaviruses and cancer: from basic studies to clinical application. Nat Rev Cancer. 2002;2(5):342\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMu\u0026ntilde;oz N, Bosch FX, de Sanjos\u0026eacute; S, Herrero R, Castellsagu\u0026eacute; X, Shah KV, Snijders PJ, Meijer CJ. Epidemiologic classification of human papillomavirus types associated with cervical cancer. N Engl J Med. 2003;348(6):518\u0026ndash;27.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eC FA. Supervised Algorithms of Machine Learning for the Prediction of Cervical Cancer. J Biomed Phys Eng. 2020;10(4):513\u0026ndash;22.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNithya B, Ilango V. Evaluation of machine learning based optimized feature selection approaches and classification methods for cervical cancer prediction. SN Appl Sci. 2019;1(6):1\u0026ndash;16.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOuahab IBA, Khamlichi SE, Bouhorma M, Sedqui A. PREDICTION OF CERVICAL CANCER RISK USING MACHINE LEARNING. 20w21.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGeetha R, Sivasubramanian S, Kaliappan M, Vimal S, Annamalai S. Cervical Cancer Identification with Synthetic Minority Oversampling Technique and PCA Analysis using Random Forest Classifier. J Med Syst. 2019;43(9):286.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMota SDS, Ota\u0026ntilde;o SS, Murta EFC, Nomelini R. Blood count and fasting blood glucose level in the assessment of prognosis and survival in advanced cervical cancer. Rev Assoc Med Bras. 2022;68(2):234\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWassie M, Aemro A, Fentie B. Prevalence and associated factors of baseline anemia among cervical cancer patients in Tikur Anbesa Specialized Hospital, Ethiopia. BMC Womens Health. 2021;21(1):36.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeng Q, Long Q, Liu Y, Yang Z, Du Y, Chen X. Prognostic value of preoperative peripheral blood mean platelet volume/platelet count ratio (MPV/PC) in patients with resectable cervical cancer. BMC Cancer. 2021;21(1):1282.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Blood routine, Cervical cancer, Machine learning, Shapley additive interpretation","lastPublishedDoi":"10.21203/rs.3.rs-4761322/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4761322/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground and objective:\u003c/h2\u003e \u003cp\u003eCervical cancer is the fourth most common cancer among women globally. The key of prevention and treatment of cervical cancer is early detection, diagnosis and treatment. We aimed to develop an interpretable model to predict the risk for patients with cervical cancer based on blood routine data and used the Shapley additive interpretation (SHAP) method to explain the model and explore factors for cervical cancer.\u003c/p\u003e\u003ch2\u003eMethods\u003c/h2\u003e \u003cp\u003eIn this paper, medical records of patients from 2013 to 2023 were collected for retrospective study. 2533 patients with cervical cancer were used as the case group, and 9879 patients with apparent healthy subjects were used as the control group. Using age, clinical diagnosis information and 22 blood cell analysis results, four different algorithm were used to construct cervical cancer prediction model.\u003c/p\u003e\u003ch2\u003eResults\u003c/h2\u003e \u003cp\u003eUsing lasso regression and random forest method, 15 important blood routine features were finally selected from 23 features for model training. Comparatively, the XGBoost model had the highest predictive performance among four models with an area under the curve (AUC) of 0.964, whereas RF had the poorest generalization ability (AUC\u0026thinsp;=\u0026thinsp;0.907). The SHAP method reveals the top 6 predictors of cervical cancer according to the importance ranking, and the average of the PDW was recognized as the most important predictor variable.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003eIn conclusion, we select the best ML based on performance and rank the importance of features according to Shapley Additive Explanation (SHAP) values. Compared to the other 4 algorithms, the results showed that the XGB had the best prediction performance for successfully predicting cervical cancer recurrence and was adopted in the establishment of the prediction model.\u003c/p\u003e","manuscriptTitle":"Cervical cancer prediction using machine learning models based on blood routine analysis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-08-22 08:06:25","doi":"10.21203/rs.3.rs-4761322/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"84f113da-a5a0-4085-8b08-73f9349f4555","owner":[],"postedDate":"August 22nd, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-10-23T05:08:05+00:00","versionOfRecord":[],"versionCreatedAt":"2024-08-22 08:06:25","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4761322","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4761322","identity":"rs-4761322","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.