A Predictive Model for Diabetic Retinopathy Based on Ensemble Learning

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Objective The purpose of this passage is to predict the risk of type 2 daibetes complicated with retinopathy,we evaluated 14 commonly used models and fusion them in a glm stacking classifier. Methods The Clinical data of this passage comes from National Health Science Data Center (Diabetic complications early warning dataset), all the statistical analysis were finished in R-4.2.1, Rstudio. We used recursive feature elimination to get variables we need, we create models in caret package, and computed Accuracy, Precision, Sensitivity, Specificity, F1-score of every models,choose the better models to the stacking classifier. Results REF feature screening shows that the accuracy of the models improve with the number of variables, and tends to be flat after more than 30 variables, in order to prevent overfitting, combined with the literature, a total of 45 variables are selected into the model, and the evaluation indicators show that the support vector machine, AdaBoost, XGBoost, rotating forest, are excellent in the first-stage modeling. The fusion of stacking models of generalized linear models is better than stage one models. Conclusion The stacking fusion model can improve the performance of the model on the basis of a single model, and can play a certain role in the screening and prediction of high-risk groups with type 2 diabetes complicated by retinopathy in the clinic.
Full text 125,664 characters · extracted from preprint-html · click to expand
A Predictive Model for Diabetic Retinopathy Based on Ensemble Learning | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Predictive Model for Diabetic Retinopathy Based on Ensemble Learning Jin feng Wei, Xiang lin Yin, Ze min Huang, Jia rui Zheng, Shi jie Deng, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-2832556/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Objective The purpose of this passage is to predict the risk of type 2 daibetes complicated with retinopathy,we evaluated 14 commonly used models and fusion them in a glm stacking classifier. Methods The Clinical data of this passage comes from National Health Science Data Center (Diabetic complications early warning dataset), all the statistical analysis were finished in R-4.2.1, Rstudio. We used recursive feature elimination to get variables we need, we create models in caret package, and computed Accuracy, Precision, Sensitivity, Specificity, F1-score of every models,choose the better models to the stacking classifier. Results REF feature screening shows that the accuracy of the models improve with the number of variables, and tends to be flat after more than 30 variables, in order to prevent overfitting, combined with the literature, a total of 45 variables are selected into the model, and the evaluation indicators show that the support vector machine, AdaBoost, XGBoost, rotating forest, are excellent in the first-stage modeling. The fusion of stacking models of generalized linear models is better than stage one models. Conclusion The stacking fusion model can improve the performance of the model on the basis of a single model, and can play a certain role in the screening and prediction of high-risk groups with type 2 diabetes complicated by retinopathy in the clinic. Machine learning diabetes Vascular lesions Retinopathy Predictive Model Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 1. Introduction With the advancement of modern computing and technology, the significance of big data applications in the medical field has increased significantly. As such, with the advent of machine learning and other cutting-edge methods, their application has been gaining immense significance in recent times. Clinical prediction models, utilizing multi-factor models to estimate the likelihood of developing a particular disease or outcome, have become a cutting-edge area of research [ 1 ]. Type 2 diabetes (T2DM), a prevalent non-communicable disease, has become a leading cause of death globally due to the growing incidence of obesity and overweight [ 2 – 9 ]. In China, the disease burden attributed to diabetes is mounting. A majority of T2DM patients suffer from at least one complication, with cardiovascular disease being the primary cause of morbidity and mortality [ 9 – 11 ]. Diabetic retinopathy (DR), a common complication of T2DM, encompasses severe DR stages such as proliferative DR (characterized by anomalous growth of new retinal blood vessels) and diabetic macular edema (DME), where the central part of the retina exhibits exudation and edema. DR is directly linked to the long-term duration of diabetes, hyperglycemia, and hypertension. Although it is conventionally considered a microvascular disease, retinal neurodegeneration is also involved [ 12 , 13 ]. The research aimed to comprehensively apply machine learning to the clinical prediction model of diabetes. This was achieved by integrating various complications, physiological and biochemical indicators to develop a highly reliable clinical model that could guide future clinical diagnosis. 2. Data Preprocessing 2.1 Data description The data used in this study was obtained from the National Population Health Science Data Center's "Diabetes Complications Warning Dataset". The dataset consists of 1500 cases each in the T2DM group and the T2DM concurrent DR group. It contains a patient number, an outcome indicator for diabetic retinopathy, four basic information fields such as age and gender, 33 types of complications (binary data), and 50 types of physiological and biochemical indicators (continuous data). 2.2 Data cleaning 2.2.1 Initial cleaning To ensure data quality, missing values in the data features were identified and addressed. The missing values were primarily concentrated in the physiological and biochemical indicators, with most column variables having missing values below 30%. Consequently, variables with missing values greater than 40% were removed from the dataset. Additionally, missing row observations were checked, and observations with missing ratios greater than 10% were removed to ensure that only complete data was used for the analysis. 2.2.2 Missing value imputation Multiple imputation was used as the method to address missing values in the dataset. Two imputation kernels, pmm (predictive mean matching) and rf (random forest imputation), both included in the mice package, were utilized in the imputation process. Five imputations were performed, and the average of 10 results was taken as the final result. It is worth noting that after deleting missing values initially, some indicators, such as TP (total protein), had only one missing value. In the random forest imputation, an algorithm error in the "Ranger" package occurred when subsetting one of the matrices. Therefore, the issue was resolved by manually adjusting to the "randomForest" package. 2.2.3 Feature selection The dataset included indicators of height and weight, which allowed us to calculate the BMI index separately. We recalculated the imputed BMI (BMI = weight/height2) and removed the original height and weight indicators from the dataset. To improve model performance and reduce computational costs, we performed RFE feature selection to obtain the optimal subset of variables. Figure 1 shows the results of this analysis. In evaluating the random forest and bagged trees models, we found that the accuracy and Kappa coefficient tended to flatten out when the number of variables exceeded 30. However, it is important to note that the optimal subset may not always be the best solution, and there is a risk of overfitting, which can reduce clinical interpretability. After reviewing previous research in the field, and considering a 2% model loss range, we determined that keeping the variables within a 40–50 range was an appropriate range. 2.2.4 Feature Normalization In order to account for the diverse range of models that will be built in the next stage and the significant differences between various indicators, we chose to normalize the data. The goal of normalization is to ensure a proportional distribution of data while eliminating differences in data dimensionality. We chose to use the Min-Max normalization method, which maps the data to a range of [0, 1] and is expressed by the following formula: \({x}^{{\prime }}=\frac{x-\text{m}\text{i}\text{n}\left(x\right)}{\text{max}\left(x\right)-\text{m}\text{i}\text{n}\left(x\right)}\) 。 2.3 Variable Selection Prior to incorporating the variables into the model, we conducted a thorough check of the variables. We removed variables that had correlation coefficients greater than 0.7 and those with zero or near-zero variance and examined the issue of variable collinearity. After this process, we retained the following variables, arranged in descending order of variable importance based on RFE feature selection: Binary variable: Nephropathy、Arteriopathy of lower extremity、Coronary heart disease、Other tumor、Other Endocrine disease、Hematopathy、Digestive carcinoma、Renal failure、Fatty liver、Hyperlipidemia、hypertension、Atherosclerosis、Sex、Myocardial infarction、Respiratory diseases、Cerebral apoplexy、Biliary tract disease、Nation. Numerical variables: Glycosylated hemoglobin、Blood urea、Red blood cell accumulation、Total protein、Direct bilirubin、Glycosylated serum protein、Glutamine transferase、Age、Systolic blood pressure、Fibrinogen、Fasting blood glucose、Lactate dehydrogenase、Diastolic blood pressure、Prothrombin time、Serum uric acid、Indirect bilirubin、Glutamic oxalacetic transaminase、Globulin,Globulin、Tumor markers、Prothrombin activity、Part Activated prothrombin time、Triglycerides、HDL cholesterol、Alkaline phosphatase、Body index、LDL cholesterol、Platelets. There are 46 variables in total, including 18 binary variables and 28 numerical variables. Taking diabetes complicated with retina as the index, the patients were divided into 10 parts according to 2:8, which were respectively used as the test set and the training set. There were 2352 training set samples and 587 test set samples. 3. Phase 1 Modeling In this phase, we evaluated 14 machine learning methods. We began by narrowing down the options through random parameter tuning, and then determined the hyperparameters through grid tuning. To select appropriate submodels for stacking phase 1, we used ROC curves and confusion matrices as evaluation metrics. To ensure the reproducibility of our results, we used a random seed of "135" and conducted ten-fold cross-validation to estimate accuracy throughout the article. 3.1 XGBoost Algorithm The XGBoost algorithm is a variant of the gradient boosting method that employs an ensemble of decision trees. There are various versions of XGBoost, including XGBlinear, XGBtree, XGBDART, and others. For this study, we utilized the xgboost package as the algorithm source, with XGBDART as the operational model.The hyperparameter optimization is shown in Table 1 . Table 1 Model parameter tuning results of XGBoost Model algorithm Adjust parameter Final parameter XGBoost nround(Maximum number of iterations) 720 eta(Model learning rate) 0.0243 gamma(Minimum loss required to add a single child node branch in the tree model) 1.96 max_depth(Modeling process, model individual tree depth) 4 colsample_bytree(The feature ratio is randomly selected when the tree is built) 0.48 rate_drop(The probability of a single discard process) 0.088 skip_drop(The probability of skipping the discard process) 0.583 min_child_weigh(The minimum weight to use when lifting the tree) 4 subsample(The proportion of subsample data to the total observation) 0.723 3.2 GBM Algorithm: The GBM algorithm is a popular Boosting ensemble learning algorithm that uses multiple tree models to make overall predictions. The later models are used to correct the errors of the previous models, leading to improved accuracy. In this study, the gbm package was used to build the GBM model.The hyperparameter optimization is shown in Table 2 . Table 2 Model parameter tuning results of GBM Model algorithm Adjust parameter Final parameter GBM n.trees(Recursive regression tree number) 2000 Interaction.depth(Depth of the decision tree) 9 Shrinkage(Learning rate) 0.036 n.minobsinnode(The minimum number of outcomes at the end points of the decision tree model) 12 3.3 AdaBoost Algorithm: The AdaBoost algorithm is another classic Boosting algorithm that uses multiple weak classifiers to build a strong classifier. The ada package was used as the algorithm source to construct the AdaBoost model.The hyperparameter optimization is shown in Table 3 . Table 3 Model parameter tuning results of AdaBoost Model algorithm Adjust parameter Final parameter AdaBoost iter(The number of iterations performed was raised) 680 maxdepth(Maximum depth of a single tree model) 10 nu(Learning rate) 0.017 3.4 CatBoost Algorithm: The CatBoost algorithm is a Boosting algorithm known for its fast convergence speed. We used the catboost package as the algorithm source to construct the CatBoost model.The hyperparameter optimization is shown in Table 4 . Table 4 Model parameter tuning results of CatBoost Model algorithm Adjust parameter Final parameter CatBoost depth(The depth of a single tree) 2 learning_rate(Learning rate) 0.17 Interations(Number of tree models) 100 l2_leaf_reg(L2 regularization penalty) 0.1 Rsm(Random subspace) 0.9 border_count(Segmentation number of numerical features) 13 3.5 C5.0 Algorithm: The C5.0 algorithm is a decision tree algorithm that uses the maximum information gain ratio principle. It is an improved version of the C4.5 algorithm, with enhanced execution efficiency and memory usage. In this study, the C50 package was used to construct the C5.0 model.The hyperparameter optimization is shown in Table 5 . Table 5 Model parameter tuning results of C5.0 Model algorithm Adjust parameter Final parameter C5.0 trails(Number of loop iterations) 99 models(Use a tree model or a rule dependency model) rules winnow(Whether features are filtered) FALSE 3.6 MARS Algorithm: The MARS algorithm is an adaptive process for regression, ideal for high-dimensional problems, and can be considered an improvement to enhance the effect of CART in regression.The hyperparameter optimization is shown in Table 6 . Table 6 Model parameter tuning results of MARS Model algorithm Adjust parameter Final parameter MARS Nprune(The maximum number of items after pruning) 21 Degree(Maximum degree of interaction) 1 3.7 Random Forest Algorithm: The Random Forest algorithm is a commonly used Bagging algorithm. It constructs multiple tree models that are unrelated and then integrates the classification results to obtain the final result. The randomForest package was used as the algorithm source to construct the Random Forest model.The hyperparameter optimization is shown in Table 7 . Table 7 Model parameter tuning results of RanomForest Model algorithm Adjust parameter Final parameter RandomForest Mtry(Default number of hyperparameters) 2 n_tree(Number of tree models) 500 3.8 Rotation Forest Algorithm: The Rotation Forest algorithm is a variant of the Random Forest algorithm. It divides attributes into K non-overlapping subsets of equal size and constructs the model by combining PCA and rotation matrix. In this study, the rotationForest package was used to calculate the Random Forest model.The hyperparameter optimization is shown in Table 8 . Table 8 Model parameter tuning results of Rotation Forest Model algorithm Adjust parameter Final parameter Rotation Forest K(Number of variable subsets) 1 L(Number of base classifiers) 96 Cp(Complexity limiting parameter) 0.00036 3.9 Naive Bayes Algorithm: The Naive Bayes algorithm is a classification method based on the Bayesian theorem and the feature condition independence assumption. It uses probability statistics to classify the sample dataset. The klaR package was used as the algorithm source to construct the Naive Bayes classifier.The hyperparameter optimization is shown in Table 9 . Table 9 Model parameter tuning results of Bayes Model algorithm Adjust parameter Final parameter Naive Bayes fl (Laplace Correction) 0 Usekernel (normal density F or kernel density T) TRUE adjust(Bandwidth adjustment) 1 3.10 Support-Vector Machine Algorithm: The support vector machine (SVM) is a supervised learning algorithm with multiple derivative variants. In this article, we used two different kernel functions of SVM, namely, the radial basis function kernel support vector machine (SVM1) and the least squares support vector machine with a polynomial kernel (SVM2). We constructed these models using the kernlab package as the algorithm source.The hyperparameter optimization is shown in Table 10 – 11 . Table 10 Model parameter tuning results of lssvmRadialWight Model algorithm Adjust parameter Final parameter lssvmRadialWight(svm1) sigma(kernel function coefficient) 0.003 C((Penalty coefficient of objective function) 14.8 wight(weight ratio) 2.15 Table 11 Model parameter tuning results of lssvmPoly Model algorithm Adjust parameter Final parameter lssvmPoly(svm2) degree(degree of polynomial) 3 Scale(scale) 0.0115 Tau(Regularized parameter) 0.521 3.11 Discriminant analysis algorithm: Discriminant analysis (LDA) is a classic supervised learning algorithm. The common distance discriminant method is to use the centroid. We classify the sample dataset by comparing which category is closer to the centroid. In this paper, we used the polynomial kernel distance weighted discriminant analysis. We constructed the discriminant analysis model using the kerndwd package as the algorithm source.The hyperparameter optimization is shown in Table 12 . Table 12 Model parameter tuning results of dwdPoly Model algorithm Adjust parameter Final parameter dwdPoly lambda(System of polynomials) 0.0007 qval(dwd index) 0.29 degree(Degree of polynomial) 3 scale(scale) 0.15 3.12 Neural network algorithm: The neural network (nnet) is a machine learning technique that simulates the human brain to achieve artificial intelligence similar to humans. In this research, we adopted a two-layer neural network with a single hidden layer and constructed the neural network model using the nnet package as the algorithm source.The hyperparameter optimization is shown in Table 13 . Table 13 Model parameter tuning results of nnet Model algorithm Adjust parameter Final parameter nnet size(Number of hidden layer neurons) 15 decay(Learning rate attenuation factor) 0.761 3.13 KNN algorithm: The K-nearest neighbor algorithm (KNN) is a non-parametric and lazy algorithm model. We only set the value of the nearest neighbor. We constructed the KNN model using the knn package as the algorithm source.The hyperparameter optimization is shown in Table 14 . Table 14 Model parameter tuning results of knn Model algorithm Adjust parameter Final parameter knn K(Nearest neighbor number) 9 4. Model Selection And Interpretation 4.1 Model Selection: We evaluated 14 machine learning models by predicting the test set and calculating the confusion matrix to obtain five performance metrics: accuracy, precision, sensitivity, specificity, and F1-score. We constructed a Receiver Operating Characteristic (ROC) curve, calculated the Area Under the Curve (AUC), and generated a histogram. The results are presented below.The results are shown in Table 15 . Table 15 Model evaluates metrics of Confusion Matrix Algorithm name Accuracy rate Precision sensitivity specificity F1-score XGBoost 0.761 0.766 0.756 0.767 0.761 AdaBoost 0.77 0.77 0.773 0.767 0.772 CatBoost 0.761 0.772 0.746 0.777 0.746 GBM 0.739 0.75 0.722 0.757 0.736 C5.0 0.744 0.76 0.719 0.771 0.739 MARS 0.738 0.744 0.729 0.747 0.736 svm1 0.716 0.833 0.542 0.89 0.657 svm2 0.772 0.784 0.753 0.791 0.768 NB 0.658 0.826 0.403 0.914 0.543 KNN 0.693 0.687 0.715 0.671 0.701 nnet 0.751 0.756 0.746 0.757 0.751 Rotating forest 0.765 0.763 0.773 0.757 0.768 Random forest 0.758 0.767 0.746 0.771 0.756 LDA 0.751 0.756 0.746 0.757 0.751 As shown in Figs. 2 , 3 and 4 , most models have an AUC value above 0.8, indicating some clinical value. Among them, the polynomial kernel least squares support vector machine achieved the highest accuracy of 0.772, the radial basis function kernel support vector machine achieved the highest precision of 0.833, the rotated forest achieved the highest sensitivity of 0.773, the Bayesian model achieved the highest specificity of 0.914, and Adaboost achieved the highest F1-score and AUC of 0.772. After considering the model metrics, we removed models with poor performance, such as MARS, discriminant analysis, KNN, and C5.0, in preparation for the second-stage modeling. 4.2 Model Interpretation: After selecting the models, we further analyzed the boosting and bagging models (CatBoost, RandomForest, and Adaboost) for variable interpretation, with the aim of identifying comorbidities and hidden variables strongly associated with diabetic retinopathy (DR), and quantifying the impact of each variable to better serve clinical practice. We believe that DR in the T2DM population is not the result of a single factor, but rather the cumulative effect of multiple variables with interactive effects. To calculate the shap values of the samples and plot the scatter diagram as an attribution analysis method, we used the iml package. For important variables, the final attribution interpretation of the two methods is similar, as the shapley value is the average of different order attributions [ 17 ]. As the outcome variable is a binary factor, we transformed it into the maximum likelihood estimate probability and used RMSE as the loss function to calculate the importance of the top 10 variables, as shown in Fig. 5 . Figures 6 – 8 illustrate the positive and negative shape contributions of the factors in the entire valid sample. Figure 5 indicates that HBA1C, kidney disease, lower limb arterial disease, coronary heart disease, uric acid, age, erythrocyte squeeze, indirect hemoglobin, and glutamate dehydrogenase are significant factors. Figures 6–8 suggest that lower limb vascular disease, kidney disease, and endocrine diseases are important positive factors to consider as comorbidities that require attention in DR patients. Age, GSP, HBA1C, TP, PT, GGT, and other biochemical indicators are significant related indicators that have clinical significance in detection. 5. Second-stage Modeling In the stacking approach, the selected models were used as the first-level sub-models, and the caretEnsemble package was used for modeling [ 18 ]. However, since the efficiency of the stacking fusion method is related to the strength (strong or weak) of correlation between models, we could not include all the models in the second-stage modeling. After considering the AUC values mentioned above, we chose to use the nnet, xgb, lsvm, and nb models to construct the glm generalized linear model as the second-stage model to obtain the stacking fusion model.The results are shown in Table 16 . Table 16 Stacking Model evaluates metrics of Confusion Matrix Algorithm name Accuracy rate Precision sensitivity specificity F1-score Stack.glm 0.781 0.782 0.780 0.780 0.780 6. Discussion According to our model, LEADDP, vascular disease, stroke, kidney disease, and endocrine system diseases are risk factors for DR. Blood indicators such as HBAIC, PCV, DBILI, GGT, and PT are closely related to the occurrence of DR. It has been established that glucose metabolism is closely related to kidney function [ 19 , 20 ]. Our model also suggests that kidney disease and kidney failure are risk factors for DR. As kidney function declines, the capacity for glucose reabsorption is affected, leading to a vicious cycle [ 21 ]. Furthermore, pathological enhancement of GLUT1 protein activity can increase the metabolic burden of the eye, leading to DR [ 22 ]. It is noteworthy that CHD, as a vascular disease, appears to act as a protective factor, meaning that having coronary heart disease can reduce the incidence of DR. This conclusion is consistent with some previous studies [ 23 , 24 ], but it contradicts the conclusion regarding lower limb vascular disease, which is also a cardiovascular disease in this study. We are not entirely certain about the clinical causal basis for this relationship, but we propose two possible explanations. Firstly, patients diagnosed with coronary heart disease often take medication that slows down the progression of microvascular disease, reducing the possibility of retinopathy. Secondly, since the data comes from hospitals, lower limb vascular disease is not a high mortality disease, while CHD is a disease with a high mortality rate and a long latency period. Therefore, selection bias may occur in hospital data, excluding those patients who died quickly outside the hospital due to CHD, and this cannot be avoided in cross-sectional studies. Glucose metabolism is closely related to liver function, and T2DM patients experience functional compensation of liver function due to insulin resistance and progressive impairment of secretion, resulting in an increasing trend in liver function indicators. Our SHAP chart suggests that fatty liver disease is a risk factor, so paying attention to trends in liver function indicators during blood tests is an important way to predict the occurrence of DR [ 25 , 26 ]. At the same time, the coagulation factor of DR patients shows a decreasing trend. Stabilizing the coagulation ability of diabetic patients is an important step in preventing DR, microvascular disease, and thrombosis. 7. Conclusion Clinical decision-making recommendations for individuals with high blood sugar include controlling their diet, protecting their liver and kidneys, measuring their blood glucose levels regularly, and paying attention to the condition of varicose veins in their lower extremities. In Fig. 5 and Figs. 6 – 8 , we found that the feature importance of some minority features was inconsistent due to the different interpretation methods of the black box model. To compare these methods and derive better clinical interpretations is a promising area for future clinical prediction models. We observed that the stacking model did not significantly improve our model's accuracy, and we believe there are three reasons for this phenomenon. First, the sample size of 3000 training samples may not be sufficient to accurately represent the entire population. Second, there may be insufficient accuracy or overlap in the first-layer models. Third, we could not include all 14 models in the first-layer due to model correlations, resulting in too few base models. As shown in Fig. 9 , our model generally has an AUC exceeding 0.8, indicating good internal consistency in terms of prediction efficiency. However, external testing with a larger sample size is still required to validate the model's effectiveness. 8. App Production As shown in Fig. 10 ,based on our research results, we developed a retinopathy prediction shinyapp for diabetic patients [ 27 , 28 ], The following is the app ICONS.This is a part of the web site ( https://r-test-model2.shinyapps.io/appsvm/ ). Considering computing time, server costs, etc. As shown in Fig. 11 , we only uploaded lightweight models. Declarations Acknowledgements Thanks to Population Health Data sharing Platform and 301 Hospital of National Clinical Science Data Center for data support of this study. The data for this study were obtained from the National Population Health Sciences Data Center Data Warehouse PHDA (http://www.ncmi.cn). Authors’ contributions Wei Jin feng,Yin Xiang lin and Huang Ze min drafted the manuscript, Wei Jin feng, Huang Ze min and Qiu Hong bin revised the manuscript. Zheng Jia rui,Deng Shi jie,Yu Yang and Xu Wei jing drew Figures and tables.All authors read and approved the final manuscript. Funding No funding. Ethics approval and consent to participate The data analyzed in this study were from public databases, so ethical approval and consent participation were not required. All methods in this study were performed in accordance with the relevant guidelines and regulations. Availability of data and materials The datasets used and analyzed during the current study are available were obtained from the National Population Health Sciences Data Center Data Warehouse PHDA (http://www.ncmi.cn). The datasets used and/or analysed during the current study available from the corresponding author on reasonable request. Consent for publication Not applicable. Competing interests The authors declare that they have no competing interests. References The dataset for diabetes complications early warning system is from the General Hospital of the Chinese People's Liberation Army. National Population Health Science Data Center Data Warehouse PHDA, 2022. https://doi.org/. WU J H, LIU T, HSU W T, et al. Performance and Limitation of Machine Learning Algorithms for Diabetic Retinopathy Screening: Meta-analysis[J]. J Med Internet Res, 2021,23(7): e23863. ZHANG X X, KONG J, YUN K. Prevalence of Diabetic Nephropathy among Patients with Type 2 Diabetes Mellitus in China: A Meta-Analysis of Observational Studies[J]. J Diabetes Res, 2020,2020: 2315607. Worldwide trends in diabetes since 1980: a pooled analysis of 751 population-based studies with 4.4 million participants[J]. Lancet, 2016,387(10027): 1513-1530. MOK C H, KWOK H, NG C S, et al. Health State Utility Values for Type 2 Diabetes and Related Complications in East and Southeast Asia: A Systematic Review and Meta-Analysis[J]. Value Health, 2021,24(7): 1059-1067. CLOETE L. Diabetes mellitus: an overview of the types, symptoms, complications and management[J]. Nurs Stand, 2022,37(1): 61-66. TINAJERO M G, MALIK V S. An Update on the Epidemiology of Type 2 Diabetes: A Global Perspective[J]. Endocrinol Metab Clin North Am, 2021,50(3): 337-355. RANASINGHE P,JAYAWARDENA R, GAMAGE N, et al. Prevalence and trends of the diabetes epidemic in urban and rural India: A pooled systematic review and meta-analysis of 1.7 million adults[J]. Ann Epidemiol, 2021,58: 128-148. ZHENG Y, LEY S H, HU F B. Global aetiology and epidemiology of type 2 diabetes mellitus and its complications[J]. Nat Rev Endocrinol, 2018,14(2): 88-98. Ning G.Current situation and prospect of diabetes prevention and treatment in China[J].Science in China: Life Sciences,2018,48(08):810-811. WONG Y H, WONG S H, WONG X T, et al. Genetic associated complications of type 2 diabetes mellitus[J]. Panminerva Med, 2022,64(2): 274-288. TEO Z L, THAM Y C, YU M, et al. Global Prevalence of Diabetic Retinopathy and Projection of Burden through 2045: Systematic Review and Meta-analysis[J]. Ophthalmology, 2021,128(11): 1580-1591. LI J Q, WELCHOWSKI T, SCHMID M, et al. Prevalence, incidence and future projection of diabetic eye disease in Europe: a systematic review and meta-analysis[J]. Eur J Epidemiol, 2020,35(1): 11-23. GOSIEWSKA A, BIECEK P. Do Not Trust Additive Explanations[J]. arXiv: Learning, 2019. RAFFERTY A, NENUTIL R, RAJAN A. Explainable Artificial Intelligence for Breast Tumour Classification: Helpful or Harmful, Cham: Springer Nature Switzerland, 2022. MOLNAR C, CASALICCHIO G, BISCHL B. iml: An R package for Interpretable Machine Learning[J]. J. Open Source Softw., 2018,3: 786. TRUMBELJ E, KONONENKO I. Explaining prediction models and individual predictions with feature contributions[J]. Knowledge and Information Systems, 2013,41: 647-665.[11] WONG Y H, WONG S H, WONG X T, et al. Genetic associated complications of type 2 diabetes mellitus[J]. Panminerva Med, 2022,64(2): 274-288. Xu, H. (2018). Research and Improvement of Stacking Algorithm [D]. South China University of Technology. SHAHWAN M, HASSAN N, SHAHEEN R A, et al. Diabetes Mellitus and Renal Function: Current Medical Research and Opinion[J]. Curr Diabetes Rev, 2021,17(9): e1763711712. JEPSON C, HSU J Y, FISCHER M J, et al. Incident Type 2 Diabetes Among Individuals With CKD: Findings From the Chronic Renal Insufficiency Cohort (CRIC) Study[J]. Am J Kidney Dis, 2019,73(1): 72-81. FENG X, ZHENG Y, GUAN H, et al. The Association between Urinary Glucose and Renal Uric Acid Excretion in Non-diabetic Patients with Stage 1-2 Chronic Kidney Disease[J]. Endocr Res, 2021,46(1): 28-36. TAI G J, YU Q Q, LI J P, et al. NLRP3 inflammasome links vascular senescence to diabetic vascular lesions[J]. Pharmacol Res, 2022,178: 106143. Cao, W., Ying, J., Chen, G., & Zhou, D. (2016). Comparative study on the prediction of type 2 diabetes complications - retinopathy risk based on logistic regression and random forest algorithms [J]. Chinese Medical Equipment, 31(03), 33-38+69. Zhang, H., Qiu, H., & Zhang, Y. (2022). Risk factors for diabetic retinopathy [J]. Journal of Mudanjiang Medical University, 43(03), 64-68. DOI: 10.13799/j.cnki.mdjyxyxb.2022.03.009. DILLMANN W H. Diabetic Cardiomyopathy[J]. Circ Res, 2019,124(8): 1160-1162. TANASE D M, GOSAV E M, COSTEA C F, et al. The Intricate Relationship between Type 2 Diabetes Mellitus (T2DM), Insulin Resistance (IR), and Nonalcoholic Fatty Liver Disease (NAFLD)[J]. J Diabetes Res, 2020,2020: 3920196. DUNNING M J, VOWLER S L, LALONDE E, et al. Mining Human Prostate Cancer Datasets: The "camcAPP" Shiny App[J]. EBioMedicine, 2017,17: 5-6. HE X, JIAO Y, YANG X, et al. A Novel Prediction Tool for Overall Survival of Patients Living with Spinal Metastatic Disease[J]. World Neurosurgery, 2020,144: e824-e836. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-2832556","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":194344873,"identity":"ece15f0d-55fe-4049-b0e0-8dbb21ace795","order_by":0,"name":"Jin feng Wei","email":"","orcid":"","institution":"Jiamusi University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jin","middleName":"feng","lastName":"Wei","suffix":""},{"id":194344876,"identity":"737fbc93-709a-4a62-b259-b2a299f225de","order_by":1,"name":"Xiang lin Yin","email":"","orcid":"","institution":"Jiamusi University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Xiang","middleName":"lin","lastName":"Yin","suffix":""},{"id":194344879,"identity":"8876412a-44fd-4200-880b-d10e510d4d0d","order_by":2,"name":"Ze min Huang","email":"","orcid":"","institution":"Jiamusi University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ze","middleName":"min","lastName":"Huang","suffix":""},{"id":194344881,"identity":"a1b46314-25d0-4c6e-b460-df7bb3b5a4e8","order_by":3,"name":"Jia rui Zheng","email":"","orcid":"","institution":"First Affiliated Hospital of Jiamusi University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jia","middleName":"rui","lastName":"Zheng","suffix":""},{"id":194344884,"identity":"480aca5c-f67f-49ac-9b05-95d25fbec5ea","order_by":4,"name":"Shi jie Deng","email":"","orcid":"","institution":"Jiamusi University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Shi","middleName":"jie","lastName":"Deng","suffix":""},{"id":194344887,"identity":"a9fcca32-5c5a-4b15-90b2-c15792881a48","order_by":5,"name":"Yang Yu","email":"","orcid":"","institution":"Jiamusi University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yang","middleName":"","lastName":"Yu","suffix":""},{"id":194344890,"identity":"f5ed685e-50d9-4eb8-b009-16bf33022aa2","order_by":6,"name":"Wei jing Xu","email":"","orcid":"","institution":"Jiamusi University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wei","middleName":"jing","lastName":"Xu","suffix":""},{"id":194344892,"identity":"d722c935-8924-4573-8ca5-1e09e9d3982e","order_by":7,"name":"Hong bin Qiu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA60lEQVRIiWNgGAWjYDACZgY2EJXAxsB8wAAilEC0FrYEIrUwQLUwMPBAdRDSwnecx+zBxx21eXzSPR+Kbvw5zMDPnmPA8HMHbi2Sh3nMDWeeOV7MJnN2g3Fu22EGyZ43Boy9Z3BrMTjMYybN23YssU0iF6il4TCDwY0cA2bGNqK05DwwzgE6zJ5ILTUgLQzGOWxAWyQIaJE8zFYmObPtQDGbRJoB0C/pPBJnnhUc7MWjhe/84W0SH9vq8uRnJD8DOsxajr89eeODn3i0MBwAk4dBBBsoYngQgvi11IEI5gd4VY6CUTAKRsGIBQBVfE7vjuEXUwAAAABJRU5ErkJggg==","orcid":"","institution":"Jiamusi University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Hong","middleName":"bin","lastName":"Qiu","suffix":""}],"badges":[],"createdAt":"2023-04-18 14:29:42","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-2832556/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-2832556/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":36289734,"identity":"68da126c-cf72-47d1-96b9-de23059af883","added_by":"auto","created_at":"2023-04-25 18:09:17","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":60452,"visible":true,"origin":"","legend":"\u003cp\u003eRFE Feature filter results\u003c/p\u003e","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2832556/v1/35ad9ec1321ef78b35cb2dc9.jpg"},{"id":36289735,"identity":"960f9794-7610-49e9-b795-b9243cb427f0","added_by":"auto","created_at":"2023-04-25 18:09:17","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":83940,"visible":true,"origin":"","legend":"\u003cp\u003eHistograms of Model metric\u003c/p\u003e","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2832556/v1/909d120ab224feb29716e35c.jpg"},{"id":36290899,"identity":"616397c5-be34-46d1-a1dc-e0ac3bf464a6","added_by":"auto","created_at":"2023-04-25 18:25:17","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":50266,"visible":true,"origin":"","legend":"\u003cp\u003eROC curve of 14 Models\u003c/p\u003e","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2832556/v1/ced572a89260c8ce5b8f0a6b.jpg"},{"id":36290898,"identity":"37613463-3b0b-40f4-9d39-a2ba9199252b","added_by":"auto","created_at":"2023-04-25 18:25:17","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":66767,"visible":true,"origin":"","legend":"\u003cp\u003eAUC score of 14 Models\u003c/p\u003e","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2832556/v1/a00b5fb63601fc43fa0b1cbb.jpg"},{"id":36290538,"identity":"ee46ad5d-f934-4e1c-bd0b-7492afdba50f","added_by":"auto","created_at":"2023-04-25 18:17:17","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":51464,"visible":true,"origin":"","legend":"\u003cp\u003eTop ten Variable’s RMSE loss importance of Boosting Models\u003c/p\u003e","description":"","filename":"5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2832556/v1/0ed51855f9c7809b116f582b.jpg"},{"id":36289737,"identity":"abac35f0-3de3-4dc8-9c16-0332e7dc51d0","added_by":"auto","created_at":"2023-04-25 18:09:17","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":110833,"visible":true,"origin":"","legend":"\u003cp\u003eVariable shap effect of Randomforest Model\u003c/p\u003e","description":"","filename":"6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2832556/v1/2a4936045db94d1c3a9ed8cb.jpg"},{"id":36289741,"identity":"a7ea9787-7305-46ac-a6a3-dbae78d06e7a","added_by":"auto","created_at":"2023-04-25 18:09:17","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":90444,"visible":true,"origin":"","legend":"\u003cp\u003eVariable shap value effect of CatBoost Model\u003c/p\u003e","description":"","filename":"7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2832556/v1/61239c9221273846ac1c3e65.jpg"},{"id":36290541,"identity":"f65ebacd-4639-4b8c-a2ff-c3159ab18411","added_by":"auto","created_at":"2023-04-25 18:17:18","extension":"jpg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":103333,"visible":true,"origin":"","legend":"\u003cp\u003eVariable shap effect of XGBoost Model\u003c/p\u003e","description":"","filename":"8.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2832556/v1/b6d9a0e10fd4f26ed1735054.jpg"},{"id":36289742,"identity":"7bcd3ecf-3295-481b-a508-2734d3ce4f73","added_by":"auto","created_at":"2023-04-25 18:09:17","extension":"jpg","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":41453,"visible":true,"origin":"","legend":"\u003cp\u003eHeatmap of Models correlation\u003c/p\u003e","description":"","filename":"9.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2832556/v1/1daa25bac4810de308e740e2.jpg"},{"id":36290539,"identity":"fb122a55-5995-41df-bffb-d4a8df208043","added_by":"auto","created_at":"2023-04-25 18:17:17","extension":"jpg","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":20888,"visible":true,"origin":"","legend":"\u003cp\u003eSymbol of APP\u003c/p\u003e","description":"","filename":"10.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2832556/v1/a24a52dde955a1c450da7f95.jpg"},{"id":36289744,"identity":"82391cae-5fae-4d95-a09c-f72fa3ef6ba7","added_by":"auto","created_at":"2023-04-25 18:09:18","extension":"jpg","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":43256,"visible":true,"origin":"","legend":"\u003cp\u003ePicture of Web-app\u003c/p\u003e","description":"","filename":"11.jpg","url":"https://assets-eu.researchsquare.com/files/rs-2832556/v1/c0ec39799d0eaad82bb9f344.jpg"},{"id":40457171,"identity":"27fa91c7-7afb-492d-9242-8918a7ec1264","added_by":"auto","created_at":"2023-07-24 08:59:35","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1077556,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-2832556/v1/ea3e2647-3cea-4132-86dd-142c905b5f38.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Predictive Model for Diabetic Retinopathy Based on Ensemble Learning","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eWith the advancement of modern computing and technology, the significance of big data applications in the medical field has increased significantly. As such, with the advent of machine learning and other cutting-edge methods, their application has been gaining immense significance in recent times. Clinical prediction models, utilizing multi-factor models to estimate the likelihood of developing a particular disease or outcome, have become a cutting-edge area of research [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eType 2 diabetes (T2DM), a prevalent non-communicable disease, has become a leading cause of death globally due to the growing incidence of obesity and overweight [\u003cspan additionalcitationids=\"CR3 CR4 CR5 CR6 CR7 CR8\" citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. In China, the disease burden attributed to diabetes is mounting. A majority of T2DM patients suffer from at least one complication, with cardiovascular disease being the primary cause of morbidity and mortality [\u003cspan additionalcitationids=\"CR10\" citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eDiabetic retinopathy (DR), a common complication of T2DM, encompasses severe DR stages such as proliferative DR (characterized by anomalous growth of new retinal blood vessels) and diabetic macular edema (DME), where the central part of the retina exhibits exudation and edema. DR is directly linked to the long-term duration of diabetes, hyperglycemia, and hypertension. Although it is conventionally considered a microvascular disease, retinal neurodegeneration is also involved [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe research aimed to comprehensively apply machine learning to the clinical prediction model of diabetes. This was achieved by integrating various complications, physiological and biochemical indicators to develop a highly reliable clinical model that could guide future clinical diagnosis.\u003c/p\u003e"},{"header":"2. Data Preprocessing","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Data description\u003c/h2\u003e \u003cp\u003eThe data used in this study was obtained from the National Population Health Science Data Center's \"Diabetes Complications Warning Dataset\". The dataset consists of 1500 cases each in the T2DM group and the T2DM concurrent DR group. It contains a patient number, an outcome indicator for diabetic retinopathy, four basic information fields such as age and gender, 33 types of complications (binary data), and 50 types of physiological and biochemical indicators (continuous data).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Data cleaning\u003c/h2\u003e \u003cdiv id=\"Sec5\" class=\"Section3\"\u003e \u003ch2\u003e2.2.1 Initial cleaning\u003c/h2\u003e \u003cp\u003eTo ensure data quality, missing values in the data features were identified and addressed. The missing values were primarily concentrated in the physiological and biochemical indicators, with most column variables having missing values below 30%. Consequently, variables with missing values greater than 40% were removed from the dataset. Additionally, missing row observations were checked, and observations with missing ratios greater than 10% were removed to ensure that only complete data was used for the analysis.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section3\"\u003e \u003ch2\u003e2.2.2 Missing value imputation\u003c/h2\u003e \u003cp\u003eMultiple imputation was used as the method to address missing values in the dataset. Two imputation kernels, pmm (predictive mean matching) and rf (random forest imputation), both included in the mice package, were utilized in the imputation process. Five imputations were performed, and the average of 10 results was taken as the final result. It is worth noting that after deleting missing values initially, some indicators, such as TP (total protein), had only one missing value. In the random forest imputation, an algorithm error in the \"Ranger\" package occurred when subsetting one of the matrices. Therefore, the issue was resolved by manually adjusting to the \"randomForest\" package.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e \u003ch2\u003e2.2.3 Feature selection\u003c/h2\u003e \u003cp\u003eThe dataset included indicators of height and weight, which allowed us to calculate the BMI index separately. We recalculated the imputed BMI (BMI\u0026thinsp;=\u0026thinsp;weight/height2) and removed the original height and weight indicators from the dataset. To improve model performance and reduce computational costs, we performed RFE feature selection to obtain the optimal subset of variables. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the results of this analysis. In evaluating the random forest and bagged trees models, we found that the accuracy and Kappa coefficient tended to flatten out when the number of variables exceeded 30. However, it is important to note that the optimal subset may not always be the best solution, and there is a risk of overfitting, which can reduce clinical interpretability. After reviewing previous research in the field, and considering a 2% model loss range, we determined that keeping the variables within a 40\u0026ndash;50 range was an appropriate range.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e \u003ch2\u003e2.2.4 Feature Normalization\u003c/h2\u003e \u003cp\u003eIn order to account for the diverse range of models that will be built in the next stage and the significant differences between various indicators, we chose to normalize the data. The goal of normalization is to ensure a proportional distribution of data while eliminating differences in data dimensionality. We chose to use the Min-Max normalization method, which maps the data to a range of [0, 1] and is expressed by the following formula:\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\({x}^{{\\prime }}=\\frac{x-\\text{m}\\text{i}\\text{n}\\left(x\\right)}{\\text{max}\\left(x\\right)-\\text{m}\\text{i}\\text{n}\\left(x\\right)}\\)\u003c/span\u003e \u003c/span\u003e。\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Variable Selection\u003c/h2\u003e \u003cp\u003ePrior to incorporating the variables into the model, we conducted a thorough check of the variables. We removed variables that had correlation coefficients greater than 0.7 and those with zero or near-zero variance and examined the issue of variable collinearity. After this process, we retained the following variables, arranged in descending order of variable importance based on RFE feature selection:\u003c/p\u003e \u003cp\u003eBinary variable: Nephropathy、Arteriopathy of lower extremity、Coronary heart disease、Other tumor、Other Endocrine disease、Hematopathy、Digestive carcinoma、Renal failure、Fatty liver、Hyperlipidemia、hypertension、Atherosclerosis、Sex、Myocardial infarction、Respiratory diseases、Cerebral apoplexy、Biliary tract disease、Nation.\u003c/p\u003e \u003cp\u003eNumerical variables: Glycosylated hemoglobin、Blood urea、Red blood cell accumulation、Total protein、Direct bilirubin、Glycosylated serum protein、Glutamine transferase、Age、Systolic blood pressure、Fibrinogen、Fasting blood glucose、Lactate dehydrogenase、Diastolic blood pressure、Prothrombin time、Serum uric acid、Indirect bilirubin、Glutamic oxalacetic transaminase、Globulin,Globulin、Tumor markers、Prothrombin activity、Part Activated prothrombin time、Triglycerides、HDL cholesterol、Alkaline phosphatase、Body index、LDL cholesterol、Platelets.\u003c/p\u003e \u003cp\u003eThere are 46 variables in total, including 18 binary variables and 28 numerical variables. Taking diabetes complicated with retina as the index, the patients were divided into 10 parts according to 2:8, which were respectively used as the test set and the training set. There were 2352 training set samples and 587 test set samples.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Phase 1 Modeling","content":"\u003cp\u003eIn this phase, we evaluated 14 machine learning methods. We began by narrowing down the options through random parameter tuning, and then determined the hyperparameters through grid tuning. To select appropriate submodels for stacking phase 1, we used ROC curves and confusion matrices as evaluation metrics. To ensure the reproducibility of our results, we used a random seed of \"135\" and conducted ten-fold cross-validation to estimate accuracy throughout the article.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e3.1 XGBoost Algorithm\u003c/h2\u003e \u003cp\u003eThe XGBoost algorithm is a variant of the gradient boosting method that employs an ensemble of decision trees. There are various versions of XGBoost, including XGBlinear, XGBtree, XGBDART, and others. For this study, we utilized the xgboost package as the algorithm source, with XGBDART as the operational model.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of XGBoost\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"8\" rowspan=\"9\"\u003e \u003cp\u003eXGBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003enround(Maximum number of iterations)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e720\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eeta(Model learning rate)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.0243\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003egamma(Minimum loss required to add a single child node branch in the tree model)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.96\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emax_depth(Modeling process, model individual tree depth)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ecolsample_bytree(The feature ratio is randomly selected when the tree is built)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.48\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003erate_drop(The probability of a single discard process)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.088\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eskip_drop(The probability of skipping the discard process)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.583\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emin_child_weigh(The minimum weight to use when lifting the tree)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003esubsample(The proportion of subsample data to the total observation)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.723\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e3.2 GBM Algorithm:\u003c/h2\u003e \u003cp\u003eThe GBM algorithm is a popular Boosting ensemble learning algorithm that uses multiple tree models to make overall predictions. The later models are used to correct the errors of the previous models, leading to improved accuracy. In this study, the gbm package was used to build the GBM model.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of GBM\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003eGBM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003en.trees(Recursive regression tree number)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2000\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eInteraction.depth(Depth of the decision tree)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eShrinkage(Learning rate)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.036\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003en.minobsinnode(The minimum number of outcomes at the end points of the decision tree model)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e3.3 AdaBoost Algorithm:\u003c/h2\u003e \u003cp\u003eThe AdaBoost algorithm is another classic Boosting algorithm that uses multiple weak classifiers to build a strong classifier. The ada package was used as the algorithm source to construct the AdaBoost model.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of AdaBoost\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eAdaBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eiter(The number of iterations performed was raised)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e680\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emaxdepth(Maximum depth of a single tree model)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003enu(Learning rate)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.017\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e3.4 CatBoost Algorithm:\u003c/h2\u003e \u003cp\u003eThe CatBoost algorithm is a Boosting algorithm known for its fast convergence speed. We used the catboost package as the algorithm source to construct the CatBoost model.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of CatBoost\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"5\" rowspan=\"6\"\u003e \u003cp\u003eCatBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003edepth(The depth of a single tree)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003elearning_rate(Learning rate)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.17\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eInterations(Number of tree models)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e100\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003el2_leaf_reg(L2 regularization penalty)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRsm(Random subspace)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eborder_count(Segmentation number of numerical features)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e13\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e3.5 C5.0 Algorithm:\u003c/h2\u003e \u003cp\u003eThe C5.0 algorithm is a decision tree algorithm that uses the maximum information gain ratio principle. It is an improved version of the C4.5 algorithm, with enhanced execution efficiency and memory usage. In this study, the C50 package was used to construct the C5.0 model.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of C5.0\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eC5.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003etrails(Number of loop iterations)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e99\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emodels(Use a tree model or a rule dependency model)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003erules\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ewinnow(Whether features are filtered)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFALSE\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e3.6 MARS Algorithm:\u003c/h2\u003e \u003cp\u003eThe MARS algorithm is an adaptive process for regression, ideal for high-dimensional problems, and can be considered an improvement to enhance the effect of CART in regression.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of MARS\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eMARS\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNprune(The maximum number of items after pruning)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDegree(Maximum degree of interaction)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e3.7 Random Forest Algorithm:\u003c/h2\u003e \u003cp\u003eThe Random Forest algorithm is a commonly used Bagging algorithm. It constructs multiple tree models that are unrelated and then integrates the classification results to obtain the final result. The randomForest package was used as the algorithm source to construct the Random Forest model.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab7\" class=\"InternalRef\"\u003e7\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of RanomForest\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eRandomForest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMtry(Default number of hyperparameters)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003en_tree(Number of tree models)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e500\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003e3.8 Rotation Forest Algorithm:\u003c/h2\u003e \u003cp\u003eThe Rotation Forest algorithm is a variant of the Random Forest algorithm. It divides attributes into K non-overlapping subsets of equal size and constructs the model by combining PCA and rotation matrix. In this study, the rotationForest package was used to calculate the Random Forest model.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab8\" class=\"InternalRef\"\u003e8\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab8\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 8\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of Rotation Forest\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eRotation Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eK(Number of variable subsets)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eL(Number of base classifiers)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e96\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCp(Complexity limiting parameter)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.00036\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003e3.9 Naive Bayes Algorithm:\u003c/h2\u003e \u003cp\u003eThe Naive Bayes algorithm is a classification method based on the Bayesian theorem and the feature condition independence assumption. It uses probability statistics to classify the sample dataset. The klaR package was used as the algorithm source to construct the Naive Bayes classifier.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab9\" class=\"InternalRef\"\u003e9\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab9\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 9\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of Bayes\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eNaive Bayes\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003efl (Laplace Correction)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eUsekernel (normal density F or kernel density T)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eTRUE\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eadjust(Bandwidth adjustment)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003e3.10 Support-Vector Machine Algorithm:\u003c/h2\u003e \u003cp\u003eThe support vector machine (SVM) is a supervised learning algorithm with multiple derivative variants. In this article, we used two different kernel functions of SVM, namely, the radial basis function kernel support vector machine (SVM1) and the least squares support vector machine with a polynomial kernel (SVM2). We constructed these models using the kernlab package as the algorithm source.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab10\" class=\"InternalRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Tab11\" class=\"InternalRef\"\u003e11\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab10\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 10\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of lssvmRadialWight\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003elssvmRadialWight(svm1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003esigma(kernel function coefficient)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.003\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eC((Penalty coefficient of objective function)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e14.8\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ewight(weight ratio)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.15\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab11\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 11\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of lssvmPoly\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003elssvmPoly(svm2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003edegree(degree of polynomial)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eScale(scale)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.0115\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTau(Regularized parameter)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.521\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003e3.11 Discriminant analysis algorithm:\u003c/h2\u003e \u003cp\u003eDiscriminant analysis (LDA) is a classic supervised learning algorithm. The common distance discriminant method is to use the centroid. We classify the sample dataset by comparing which category is closer to the centroid. In this paper, we used the polynomial kernel distance weighted discriminant analysis. We constructed the discriminant analysis model using the kerndwd package as the algorithm source.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab12\" class=\"InternalRef\"\u003e12\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab12\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 12\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of dwdPoly\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e \u003cp\u003edwdPoly\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003elambda(System of polynomials)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.0007\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eqval(dwd index)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.29\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003edegree(Degree of polynomial)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003escale(scale)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.15\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003e3.12 Neural network algorithm:\u003c/h2\u003e \u003cp\u003eThe neural network (nnet) is a machine learning technique that simulates the human brain to achieve artificial intelligence similar to humans. In this research, we adopted a two-layer neural network with a single hidden layer and constructed the neural network model using the nnet package as the algorithm source.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab13\" class=\"InternalRef\"\u003e13\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab13\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 13\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of nnet\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003ennet\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003esize(Number of hidden layer neurons)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003edecay(Learning rate attenuation factor)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.761\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec23\" class=\"Section2\"\u003e \u003ch2\u003e3.13 KNN algorithm:\u003c/h2\u003e \u003cp\u003eThe K-nearest neighbor algorithm (KNN) is a non-parametric and lazy algorithm model. We only set the value of the nearest neighbor. We constructed the KNN model using the knn package as the algorithm source.The hyperparameter optimization is shown in Table\u0026nbsp;\u003cspan refid=\"Tab14\" class=\"InternalRef\"\u003e14\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab14\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 14\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel parameter tuning results of knn\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel algorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdjust parameter\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFinal parameter\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eknn\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eK(Nearest neighbor number)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"4. Model Selection And Interpretation","content":"\u003cdiv id=\"Sec25\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Model Selection:\u003c/h2\u003e \u003cp\u003eWe evaluated 14 machine learning models by predicting the test set and calculating the confusion matrix to obtain five performance metrics: accuracy, precision, sensitivity, specificity, and F1-score. We constructed a Receiver Operating Characteristic (ROC) curve, calculated the Area Under the Curve (AUC), and generated a histogram. The results are presented below.The results are shown in Table\u0026nbsp;\u003cspan refid=\"Tab15\" class=\"InternalRef\"\u003e15\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab15\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 15\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel evaluates metrics of Confusion Matrix\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlgorithm name\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy rate\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003esensitivity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003especificity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eF1-score\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXGBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.761\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.766\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.756\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.767\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.761\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAdaBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.77\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.77\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.773\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.767\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.772\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCatBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.761\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.772\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.746\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.777\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.746\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGBM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.739\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.722\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.757\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.736\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC5.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.744\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.76\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.719\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.771\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.739\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMARS\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.738\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.744\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.729\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.747\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.736\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003esvm1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.716\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.833\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.542\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.89\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.657\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003esvm2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.772\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.784\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.753\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.791\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.768\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNB\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.658\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.826\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.403\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.914\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.543\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.693\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.687\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.715\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.671\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.701\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ennet\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.751\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.756\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.746\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.757\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.751\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRotating forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.765\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.763\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.773\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.757\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.768\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.758\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.767\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.746\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.771\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.756\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLDA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.751\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.756\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.746\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.757\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.751\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eAs shown in Figs.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, \u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e and \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, most models have an AUC value above 0.8, indicating some clinical value. Among them, the polynomial kernel least squares support vector machine achieved the highest accuracy of 0.772, the radial basis function kernel support vector machine achieved the highest precision of 0.833, the rotated forest achieved the highest sensitivity of 0.773, the Bayesian model achieved the highest specificity of 0.914, and Adaboost achieved the highest F1-score and AUC of 0.772.\u003c/p\u003e \u003cp\u003eAfter considering the model metrics, we removed models with poor performance, such as MARS, discriminant analysis, KNN, and C5.0, in preparation for the second-stage modeling.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec26\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Model Interpretation:\u003c/h2\u003e \u003cp\u003eAfter selecting the models, we further analyzed the boosting and bagging models (CatBoost, RandomForest, and Adaboost) for variable interpretation, with the aim of identifying comorbidities and hidden variables strongly associated with diabetic retinopathy (DR), and quantifying the impact of each variable to better serve clinical practice.\u003c/p\u003e \u003cp\u003eWe believe that DR in the T2DM population is not the result of a single factor, but rather the cumulative effect of multiple variables with interactive effects. To calculate the shap values of the samples and plot the scatter diagram as an attribution analysis method, we used the iml package. For important variables, the final attribution interpretation of the two methods is similar, as the shapley value is the average of different order attributions [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAs the outcome variable is a binary factor, we transformed it into the maximum likelihood estimate probability and used RMSE as the loss function to calculate the importance of the top 10 variables, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e5\u003c/span\u003e. Figures\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e illustrate the positive and negative shape contributions of the factors in the entire valid sample.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e5\u003c/span\u003e indicates that HBA1C, kidney disease, lower limb arterial disease, coronary heart disease, uric acid, age, erythrocyte squeeze, indirect hemoglobin, and glutamate dehydrogenase are significant factors.\u003c/p\u003e \u003cp\u003eFigures 6\u0026ndash;8 suggest that lower limb vascular disease, kidney disease, and endocrine diseases are important positive factors to consider as comorbidities that require attention in DR patients. Age, GSP, HBA1C, TP, PT, GGT, and other biochemical indicators are significant related indicators that have clinical significance in detection.\u003c/p\u003e"},{"header":"5. Second-stage Modeling","content":"\u003cp\u003eIn the stacking approach, the selected models were used as the first-level sub-models, and the caretEnsemble package was used for modeling [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. However, since the efficiency of the stacking fusion method is related to the strength (strong or weak) of correlation between models, we could not include all the models in the second-stage modeling. After considering the AUC values mentioned above, we chose to use the nnet, xgb, lsvm, and nb models to construct the glm generalized linear model as the second-stage model to obtain the stacking fusion model.The results are shown in Table\u0026nbsp;\u003cspan refid=\"Tab16\" class=\"InternalRef\"\u003e16\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab16\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 16\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eStacking Model evaluates metrics of Confusion Matrix\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlgorithm name\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAccuracy rate\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003esensitivity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003especificity\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eF1-score\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStack.glm\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.781\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.782\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.780\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.780\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.780\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"6. Discussion","content":"\u003cp\u003eAccording to our model, LEADDP, vascular disease, stroke, kidney disease, and endocrine system diseases are risk factors for DR. Blood indicators such as HBAIC, PCV, DBILI, GGT, and PT are closely related to the occurrence of DR. It has been established that glucose metabolism is closely related to kidney function [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Our model also suggests that kidney disease and kidney failure are risk factors for DR. As kidney function declines, the capacity for glucose reabsorption is affected, leading to a vicious cycle [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Furthermore, pathological enhancement of GLUT1 protein activity can increase the metabolic burden of the eye, leading to DR [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIt is noteworthy that CHD, as a vascular disease, appears to act as a protective factor, meaning that having coronary heart disease can reduce the incidence of DR. This conclusion is consistent with some previous studies [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e], but it contradicts the conclusion regarding lower limb vascular disease, which is also a cardiovascular disease in this study. We are not entirely certain about the clinical causal basis for this relationship, but we propose two possible explanations. Firstly, patients diagnosed with coronary heart disease often take medication that slows down the progression of microvascular disease, reducing the possibility of retinopathy. Secondly, since the data comes from hospitals, lower limb vascular disease is not a high mortality disease, while CHD is a disease with a high mortality rate and a long latency period. Therefore, selection bias may occur in hospital data, excluding those patients who died quickly outside the hospital due to CHD, and this cannot be avoided in cross-sectional studies.\u003c/p\u003e \u003cp\u003eGlucose metabolism is closely related to liver function, and T2DM patients experience functional compensation of liver function due to insulin resistance and progressive impairment of secretion, resulting in an increasing trend in liver function indicators. Our SHAP chart suggests that fatty liver disease is a risk factor, so paying attention to trends in liver function indicators during blood tests is an important way to predict the occurrence of DR [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. At the same time, the coagulation factor of DR patients shows a decreasing trend. Stabilizing the coagulation ability of diabetic patients is an important step in preventing DR, microvascular disease, and thrombosis.\u003c/p\u003e"},{"header":"7. Conclusion","content":"\u003cp\u003eClinical decision-making recommendations for individuals with high blood sugar include controlling their diet, protecting their liver and kidneys, measuring their blood glucose levels regularly, and paying attention to the condition of varicose veins in their lower extremities. In Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e5\u003c/span\u003e and Figs.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Fig8\" class=\"InternalRef\"\u003e8\u003c/span\u003e, we found that the feature importance of some minority features was inconsistent due to the different interpretation methods of the black box model. To compare these methods and derive better clinical interpretations is a promising area for future clinical prediction models. We observed that the stacking model did not significantly improve our model's accuracy, and we believe there are three reasons for this phenomenon. First, the sample size of 3000 training samples may not be sufficient to accurately represent the entire population. Second, there may be insufficient accuracy or overlap in the first-layer models. Third, we could not include all 14 models in the first-layer due to model correlations, resulting in too few base models. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e, our model generally has an AUC exceeding 0.8, indicating good internal consistency in terms of prediction efficiency. However, external testing with a larger sample size is still required to validate the model's effectiveness.\u003c/p\u003e"},{"header":"8. App Production","content":"\u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003e,based on our research results, we developed a retinopathy prediction shinyapp for diabetic patients [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], The following is the app ICONS.This is a part of the web site (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://r-test-model2.shinyapps.io/appsvm/\u003c/span\u003e\u003cspan address=\"https://r-test-model2.shinyapps.io/appsvm/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Considering computing time, server costs, etc. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig11\" class=\"InternalRef\"\u003e11\u003c/span\u003e, we only uploaded lightweight models.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThanks to Population Health Data sharing Platform and 301 Hospital of National Clinical Science Data Center for data support of this study. The data for this study were obtained from the National Population Health Sciences Data Center Data Warehouse PHDA (http://www.ncmi.cn).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWei Jin feng,Yin Xiang lin and Huang Ze min drafted the manuscript, Wei Jin feng, Huang Ze min and Qiu Hong bin revised the manuscript. Zheng Jia rui,Deng Shi jie,Yu Yang and Xu Wei jing drew Figures and tables.All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNo funding.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe data analyzed in this study were from public databases, so ethical approval and consent participation were not required. All methods in this study were performed in accordance with the relevant guidelines and regulations.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets used and analyzed during the current study are available were obtained from the National Population Health Sciences Data Center Data Warehouse PHDA (http://www.ncmi.cn). The datasets used and/or analysed during the current study available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eThe dataset for diabetes complications early warning system is from the General Hospital of the Chinese People\u0026apos;s Liberation Army. National Population Health Science Data Center Data Warehouse PHDA, 2022. https://doi.org/.\u003c/li\u003e\n\u003cli\u003eWU J H, LIU T, HSU W T, et al. Performance and Limitation of Machine Learning Algorithms for Diabetic Retinopathy Screening: Meta-analysis[J]. J Med Internet Res, 2021,23(7): e23863.\u003c/li\u003e\n\u003cli\u003eZHANG X X, KONG J, YUN K. Prevalence of Diabetic Nephropathy among Patients with Type 2 Diabetes Mellitus in China: A Meta-Analysis of Observational Studies[J]. J Diabetes Res, 2020,2020: 2315607.\u003c/li\u003e\n\u003cli\u003eWorldwide trends in diabetes since 1980: a pooled analysis of 751 population-based studies with 4.4 million participants[J]. Lancet, 2016,387(10027): 1513-1530.\u003c/li\u003e\n\u003cli\u003eMOK C H, KWOK H, NG C S, et al. Health State Utility Values for Type 2 Diabetes and Related Complications in East and Southeast Asia: A Systematic Review and Meta-Analysis[J]. Value Health, 2021,24(7): 1059-1067.\u003c/li\u003e\n\u003cli\u003eCLOETE L. Diabetes mellitus: an overview of the types, symptoms, complications and management[J]. Nurs Stand, 2022,37(1): 61-66.\u003c/li\u003e\n\u003cli\u003eTINAJERO M G, MALIK V S. An Update on the Epidemiology of Type 2 Diabetes: A Global Perspective[J]. Endocrinol Metab Clin North Am, 2021,50(3): 337-355.\u003c/li\u003e\n\u003cli\u003eRANASINGHE P,JAYAWARDENA R, GAMAGE N, et al. Prevalence and trends of the diabetes epidemic in urban and rural India: A pooled systematic review and meta-analysis of 1.7 million adults[J]. Ann Epidemiol, 2021,58: 128-148.\u003c/li\u003e\n\u003cli\u003eZHENG Y, LEY S H, HU F B. Global aetiology and epidemiology of type 2 diabetes mellitus and its complications[J]. Nat Rev Endocrinol, 2018,14(2): 88-98.\u003c/li\u003e\n\u003cli\u003eNing G.Current situation and prospect of diabetes prevention and treatment in China[J].Science in China: Life Sciences,2018,48(08):810-811.\u003c/li\u003e\n\u003cli\u003eWONG Y H, WONG S H, WONG X T, et al. Genetic associated complications of type 2 diabetes mellitus[J]. Panminerva Med, 2022,64(2): 274-288.\u003c/li\u003e\n\u003cli\u003eTEO Z L, THAM Y C, YU M, et al. Global Prevalence of Diabetic Retinopathy and Projection of Burden through 2045: Systematic Review and Meta-analysis[J]. Ophthalmology, 2021,128(11): 1580-1591.\u003c/li\u003e\n\u003cli\u003eLI J Q, WELCHOWSKI T, SCHMID M, et al. Prevalence, incidence and future projection of diabetic eye disease in Europe: a systematic review and meta-analysis[J]. Eur J Epidemiol, 2020,35(1): 11-23.\u003c/li\u003e\n\u003cli\u003eGOSIEWSKA A, BIECEK P. Do Not Trust Additive Explanations[J]. arXiv: Learning, 2019.\u003c/li\u003e\n\u003cli\u003eRAFFERTY A, NENUTIL R, RAJAN A. Explainable Artificial Intelligence for Breast Tumour Classification: Helpful or Harmful, Cham: Springer Nature Switzerland, 2022.\u003c/li\u003e\n\u003cli\u003eMOLNAR C, CASALICCHIO G, BISCHL B. iml: An R package for Interpretable Machine Learning[J]. J. Open Source Softw., 2018,3: 786.\u003c/li\u003e\n\u003cli\u003eTRUMBELJ E, KONONENKO I. Explaining prediction models and individual predictions with feature contributions[J]. Knowledge and Information Systems, 2013,41: 647-665.[11] WONG Y H, WONG S H, WONG X T, et al. Genetic associated complications of type 2 diabetes mellitus[J]. Panminerva Med, 2022,64(2): 274-288.\u003c/li\u003e\n\u003cli\u003eXu, H. (2018). Research and Improvement of Stacking Algorithm [D]. South China University of Technology.\u003c/li\u003e\n\u003cli\u003eSHAHWAN M, HASSAN N, SHAHEEN R A, et al. Diabetes Mellitus and Renal Function: Current Medical Research and Opinion[J]. Curr Diabetes Rev, 2021,17(9): e1763711712.\u003c/li\u003e\n\u003cli\u003eJEPSON C, HSU J Y, FISCHER M J, et al. Incident Type 2 Diabetes Among Individuals With CKD: Findings From the Chronic Renal Insufficiency Cohort (CRIC) Study[J]. Am J Kidney Dis, 2019,73(1): 72-81.\u003c/li\u003e\n\u003cli\u003eFENG X, ZHENG Y, GUAN H, et al. The Association between Urinary Glucose and Renal Uric Acid Excretion in Non-diabetic Patients with Stage 1-2 Chronic Kidney Disease[J]. Endocr Res, 2021,46(1): 28-36.\u003c/li\u003e\n\u003cli\u003eTAI G J, YU Q Q, LI J P, et al. NLRP3 inflammasome links vascular senescence to diabetic vascular lesions[J]. Pharmacol Res, 2022,178: 106143.\u003c/li\u003e\n\u003cli\u003eCao, W., Ying, J., Chen, G., \u0026amp; Zhou, D. (2016). Comparative study on the prediction of type 2 diabetes complications - retinopathy risk based on logistic regression and random forest algorithms [J]. Chinese Medical Equipment, 31(03), 33-38+69.\u003c/li\u003e\n\u003cli\u003eZhang, H., Qiu, H., \u0026amp; Zhang, Y. (2022). Risk factors for diabetic retinopathy [J]. Journal of Mudanjiang Medical University, 43(03), 64-68. DOI: 10.13799/j.cnki.mdjyxyxb.2022.03.009.\u003c/li\u003e\n\u003cli\u003eDILLMANN W H. Diabetic Cardiomyopathy[J]. Circ Res, 2019,124(8): 1160-1162.\u003c/li\u003e\n\u003cli\u003eTANASE D M, GOSAV E M, COSTEA C F, et al. The Intricate Relationship between Type 2 Diabetes Mellitus (T2DM), Insulin Resistance (IR), and Nonalcoholic Fatty Liver Disease (NAFLD)[J]. J Diabetes Res, 2020,2020: 3920196.\u003c/li\u003e\n\u003cli\u003eDUNNING M J, VOWLER S L, LALONDE E, et al. Mining Human Prostate Cancer Datasets: The \u0026quot;camcAPP\u0026quot; Shiny App[J]. EBioMedicine, 2017,17: 5-6.\u003c/li\u003e\n\u003cli\u003eHE X, JIAO Y, YANG X, et al. A Novel Prediction Tool for Overall Survival of Patients Living with Spinal Metastatic Disease[J]. World Neurosurgery, 2020,144: e824-e836.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Machine learning, diabetes, Vascular lesions, Retinopathy, Predictive Model","lastPublishedDoi":"10.21203/rs.3.rs-2832556/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-2832556/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eObjective\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe purpose of this passage is to predict the risk of type 2 daibetes complicated with retinopathy,we evaluated 14 commonly used models and fusion them in a glm stacking classifier.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Clinical data of this passage comes from National Health Science Data Center (Diabetic complications early warning dataset), all the statistical analysis were finished in R-4.2.1, Rstudio. We used recursive feature elimination to get variables we need, we create models in caret package, and computed Accuracy, Precision, Sensitivity, Specificity, F1-score of every models,choose the better models to the stacking classifier.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eREF feature screening shows that the accuracy of the models improve with the number of variables, and tends to be flat after more than 30 variables, in order to prevent overfitting, combined with the literature, a total of 45 variables are selected into the model, and the evaluation indicators show that the support vector machine, AdaBoost, XGBoost, rotating forest, are excellent in the first-stage modeling. The fusion of stacking models of generalized linear models is better than stage one models.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe stacking fusion model can improve the performance of the model on the basis of a single model, and can play a certain role in the screening and prediction of high-risk groups with type 2 diabetes complicated by retinopathy in the clinic.\u003c/p\u003e","manuscriptTitle":"A Predictive Model for Diabetic Retinopathy Based on Ensemble Learning","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-04-25 18:09:12","doi":"10.21203/rs.3.rs-2832556/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"1169d771-9eb5-4f36-8ae3-02fad18e7a4e","owner":[],"postedDate":"April 25th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2023-07-24T08:59:22+00:00","versionOfRecord":[],"versionCreatedAt":"2023-04-25 18:09:12","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-2832556","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-2832556","identity":"rs-2832556","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00