Predicting hypoproteinemia among patients undergoing maintenance hemodialysis: A development and validation study based on machine learning algorithms | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Predicting hypoproteinemia among patients undergoing maintenance hemodialysis: A development and validation study based on machine learning algorithms Wang Yao, Yang Jingshu, Wang Haiyan, Zhang Huiru, Duan Xiaotian, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3219283/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 10 You are reading this latest preprint version Abstract Purpose Maintenance hemodialysis (MHD), which can cause various complications, is a common alternative therapy for patients with ESRD. This research built a prediction model of hypoproteinemia among ESRD patients based on machine learning algorithms. Method A total of 468 patients were selected as subjects. The “hypoproteinemia risk factor data extraction table” was drawn up after a literature review. Univariate analysis was used to screen independent risk factors as prediction variables. After hyper parameter adjustment by k-fold (k = 5) cross-validation and grid search, random forest (RF), support vector machine (SVM), back propagation (BP) neural network and logistic regression (LR) prediction models were developed. The model was evaluated by 6 dimensions, including AUROC, accuracy, precision, sensitivity, specificity and F1 score, and an importance matrix diagram was used to describe the importance. Result The incidence of hypoproteinemia in total was 30.8%. According to univariate analysis, the difference between the hypoproteinemia and nonhypoproteinemia groups was significant in 18 aspects, including age, weight, dialysis duration, and dialysis frequency. In the training set, the AUROC values of the RF, SVM, and LR models were all greater than 0.8 unlike the BP neural network (0.798). The RF model had the highest AUC value (0.924). The specificities of the LR and RF models were similar (0.846 and 0.839, respectively), while the RF model had the best accuracy (0.924) and balanced F1 score (0.751). The models had higher performance indexes in the test set than in the training set, with the RF and BP models performing better in AUROC (0.981, 0.948) and the RF model being better in accuracy, specificity balanced F1 score and precision. The top 5 prediction variables were hypersensitivity C reactive protein, age, weight, usage of high-throughput dialyzers, and dialysis age. Conclusion The RF model performed best. The model could help recognize characteristics related to hypoproteinemia during clinical practice, thereby enhancing nurses’ risk perception and improving accurate screening, primary prevention and early intervention. Maintenance Hemodialysis Hypoproteinemia Machine Learning Algorithms Prediction Model Random Forest Back Propagation Neural Network Figures Figure 1 Figure 2 Figure 3 Figure 4 1. Introduction ESRD is the final stage of various renal diseases when renal function is irreversibly declining. Kidney transplantation, hemodialysis and peritoneal dialysis are the main treatment strategies[ 1 ]. Recently, because a shortage of kidney transplant donors has resulted in numerous ESRD patients not receiving transplantation in time, maintenance hemodialysis (MHD) has become an important alternate treatment for ESRD patients thanks to the maturity of dialysis techniques and the strong support of medical insurance. By 2020, the number of MHD patients in China was as high as 690,000[ 2 ]. Long-term hemodialysis may lead to a variety of complications, of which 20ཞ50% of MHD patients may suffer from hypoproteinemia[ 3 – 6 ]. In the laboratory, hypoproteinemia can be diagnosed by total plasma protein < 60 g/L or albumin < 35 g/L. In terms of clinical manifestations, hypoproteinemia, which is one of the risk factors for cardiovascular and cerebrovascular diseases, infection, readmission and even death, could present as malnutrition and be prevented if corrected in time[ 7 – 9 ]. Early screening and identification is the key to the prevention and treatment of hypoproteinemia, which can effectively prevent the occurrence and development of hypoproteinemia in MHD patients and reduce the social and family medical economic burden. It was reported that microinflammatory status, age, weight, diabetes, and dialysis frequency are early risk factors for hypoproteinemia[ 7 , 10 – 12 ]; however, there is no research on the building of an early risk prediction model for hypoproteinemia. Machine learning, as a hybrid of artificial intelligence-computer and statistics, can predict the occurrence of the disease by defining data attributes, using computers to mine the laws existing in data, and applying different classifiers[ 13 ]. Compared with traditional statistical methods, it can capture the nonlinear relationship in the data and the complexity between multiple predictive variables, explain the patient's specific level of prediction, accurately identify the risk factors for the disease, and make early diagnosis of the disease, which better solves the complexity and unpredictable nature of human physiology[ 14 ]. In view of all of the above, the purpose of this study is to mine the target data set and establish a risk prediction model of hypoproteinemia in MHD patients based on 4 classical algorithms, including RF, SVM, LR and BP neural networks, and evaluate the model prediction efficiency using different algorithms at the same time to build a model that functions best. The clinical importance of this model is that it can help nurses identify changes in patient-related risk factors as early as possible in clinical practice, enhance nurses' risk perception ability, improve the early accurate screening and primary prevention of hypoproteinemia, and provide a basis for accurate intervention and treatment. 2. Object and method 2.1 Objects This study is a single-center retrospective study. A total of 468 MHD patients who were admitted to the Department of Nephrology, First Bethune Hospital of Jilin University, Changchun City, Jilin Province, from January 1, 2021, to December 1, 2021, were selected as the subjects. A total of 43 variables were included in this study as candidate risk factors. According to the empirical method, the sample size was estimated to be 10ཞ15 times the variables[ 15 ]. Therefore, the sample size required for this study was 430 cases in theory. Considering the 10% loss of data, we needed to collect at least 468 MHD patients. 2.1.1 Inclusion criteria: Age ≥ 18; Consistent with the diagnostic criteria of ESRD in the International Classification of Diseases (ICD)[ 16 ]; Applying arteriovenous fistula for dialysis treatment; Continuous dialysis time; Dialysis frequency: 2ཞ3 times/week; Dialysis blood flow rate: 250ཞ300 mL/min; Dialysis fluid flow rate: 500 mL/min. 2.1.2 Exclusion criteria: Diagnosed with malignant tumor; Suffering from nutritional intake/absorption disorders caused by digestive diseases; Albumin synthesis disorder caused by decompensated liver function; With residual urine volume; Incomplete medical records. 2.2 Data collection Using the self-designed data collection tool Data Extraction Table of Hypoproteinemia Risk Factors in Maintenance Hemodialysis Patients', two researchers conducted a literature review of Web of Science, PubMed, CNKI, Wanfang and other databases. In the process of screening literature, if the opinions were not uniform, they were to be discussed with high-level researchers. The search strategies were as follows: Retrieval time: 2001.01.01ཞ2021.01.01; Database: English databases include Web of Science, Embase, PubMed, and Cochrane libraries, Chinese databases include China National Knowledge Infrastructure (CNKI), Wanfang Data Knowledge Service Platform (Wanfang), VIP Journal Resource Integration Service Platform (VIP), and China Biomedical Literature Database (CBM); Search terms: Chinese search terms included "maintenance hemodialysis/hemodialysis/MHD and hypoproteinemia/albumin levels and influencing factors/risk factors/predictive factors/risk factors"; English search terms included "maintenance hemodialysis/hematodialysis/MHD and hypoproteinemia/albumin levels and influencing factors/risk factors/predictors/risk factors". The data of MHD patients with hypoproteinemia mainly included the following 4 aspects: ① basic information; ② comorbidity data; ③ dialysis-related data; and ④ biochemical examination data (see Table 1 ). Table 1 Data of MHD Patients with Hypoproteinemia Data Type Factors Basic information Sex, age, height, body weight, and BMI Comorbidity Hepatitis B, pneumonia, tuberculosis, diabetes, hypertension and associated diseases Dialysis Dialysis age, dialysis frequency, hemodialysis treatment mode, whether to use high-throughput dialysis Laboratory examination Asparagine amino acid transferase, alkaline phosphatase, cholinesterase, prealbumin, total protein, albumin, globulin, hemoglobin, total bilirubin, total cholesterol, triglycerides, high-density lipoprotein, retinol-binding protein, urea nitrogen, creatinine, uric acid, cystatin, blood calcium, serum phosphorus, serum iron, ferritin, hypersensitivity C reactive protein, erythrocyte count, parathyroid hormone, procalcitonin, cardiac troponin, myoglobin, CO 2 binding force 2.3 Data preprocessing Data preprocessing is one of the important procedures of machine learning and data mining and includes the treatment of missing and abnormal values. Following the law of missing value replacement in Python, we use mode to fill missing values in discrete characteristics and means and medians to fill continuous characteristics; isolated forest algorithm was applied to remove outliers; data Z score normalization was used to standardize the data, and label coding and one-hot transformation were used for discrete variables. 2.4 Model development Considering that this study was a small sample study and collected multiple binary variables, four algorithms of RF, SVM, LR and BP neural network were selected to develop the model. The process is shown in Fig. 1 . Feature selection. In this data modeling analysis, the incidence of hypoproteinemia in MHD patients was used as the classification target. Classification prediction was performed according to the first serum albumin and total protein results of objects. Serum albumin 60 g/L was selected as the output variable of the model. Through statistical analysis, the characteristic variables with P < 0.05 in “the risk factor data extraction table” were used as the predictive variables of the model. Model training. The optimal parameters of the four models were determined by the k-fold (k = 5) cross-validation algorithm and grid search method, including max _ depth, n _ estimators, max _ features, C, gamma, penalty, dropout, epoch, batch size and detail. Finally, 4 ML algorithm models were trained by the calculated hyperparameters. In addition, the robustness of the data set was verified and divided into 5 groups. Each group of 4/5 data points was used to train the model, and the other 1/5 data points were used to verify the model performance. The above process was repeated 5 times, and the average value of the five results was taken to generate an estimate for model comparison and evaluation. Model evaluation. We analyzed the receiver operating characteristic curve (ROC), confusion matrix, accuracy, balance F1 score, sensitivity, specificity and other indicators to evaluate the performance of different models. Model validation. We internally validated the trained model in the test set to determine whether the model is universal. Similarly, we followed the method of step 3 to calculate the performance index of the model in the test set. 2.5 Statistical analysis The original data collected were analyzed by SPSS 25.0 software. The measurement data were described by (x ± s) or M (P25, P75). Independent sample t tests or Mann‒Whitney U tests were used for comparisons between groups. We used the number of cases (%) for numeration data description and the X 2 test for comparison between groups of binary classification variables. Variables with P < 0.05 were imported into Python3.10 for model development and evaluation. The importance of input variables was analyzed, and the risk factors for prehypoproteinemia in MHD patients in the optimal model were listed. 3. Results 3.1 Basic information Of all the 468 objectives included, 144 had hypoproteinemia, accounting for approximately 30.8%. The cohort of patients was randomly divided into a training set and a test set at a ratio of 8:2. A total of 374 subjects were included in the training set, of which 118 suffered from hypoproteinemia, accounting for approximately 31.6%. Other patients of 94 subjects were included in the test set, among whom 26 patients had hypoproteinemia, accounting for approximately 27.7%. 3.2 Characteristic variable selection A total of 43 research variables were included for data preprocessing. After removing variables with data missing greater than 30% (myoglobin (47.2%), cardiac troponin (43.6%), cystatin C (37.8%), uric acid (31.8%)), we included a total of 39 characteristic variables. There were significant differences in age, weight, high-sensitivity C-reactive protein, dialysis age, diabetes, retinol-binding protein and 18 other indicators between the 2 groups (P < 0.05, see Table 1 ). Table 1 Comparison of baseline data between patients with hypoalbuminemia and those without hypoalbuminemia Feature Nonhypoproteinemia group(n = 324) Hypoproteinemia group(n = 144) t / Z / X 2 P Age (years) b 52.82(42.00, 64.00) 57.49(48.00.69.00) -2.602 0.009 Sex (male) a 206(63.5) 76(52.7) 2.317 0.128 Weight(kg) b 67.60(58.25, 74.75) 58.57(48.75, 67.25) -3.633 < 0.001 Height(cm) b 166.72(160.00.173.50) 164.88(160.00, 170.00) -0.776 0.983 BMI(kg/m 2 ) b 23.99(20.81, 25.95) 23.32(20.86, 25.38) -3.763 0.398 Dialysis age (years) b 4.00(1.77, 7.94) 2.00(0.50, 5.00) -5.784 < 0.001 Dialysis frequency (times/week) b 2.87(3.00, 3.00) 2.83(3.00, 3.00) -2.705 0.005 High-throughput dialysis a 50(15.4) 57(39.5) 22.362 < 0.001 Applying hemodialysis mode a 302(93.2) 124(86.1) 0.440 0.507 Diabetes a 98(30.2) 71(49.3) 10.014 0.002 Hypertension a 284(87.6) 124(86.1) 3.463 0.963 Hepatitis B a 20(6.1) 41(28.4) 8.065 < 0.001 Hepatitis C a 5(1.5) 8(5.5) 3.128 0.077 Pulmonary infection a 63(19.4) 45(31.2) 6.194 0.013 Pulmonary tuberculosis a 1(0.3) 8(5.5) 12.537 < 0.001 Number of basic diseases (≥ 3) a 202(62.35) 95(65.97) 0.064 0.801 AST (U/L) b 14.00(11.00, 18.00) 14.00(11.00, 21.00) -1.466 0.012 ALP (U/L) b 85.00(67.00, 120.25) 86.00(65.75, 111.75) -0.714 0.938 CHE (U/L) c 7327.49 ± 365.54 5320.19 ± 276.187 1.201 < 0.001 PAB (mg/L) c 377.36 ± 108.54 261.96 ± 105.19 1.173 < 0.001 Globulin (g/L) c 31.09 ± 5.07 29.19 ± 7.56 1.972 < 0.001 Hemoglobin (g/L) c 112.15 ± 31.80 92.24 ± 27.39 2.822 < 0.001 Hs-CRP (mg/dL) b 4.02(3.00, 6.00) 50.51(13.07, 110.00) -15.382 < 0.001 RBC (10 ~ 12/L) c 6.66 ± 1.64 4.32 ± 0.71 3.071 0.379 STB (umol/L) b 7.2(5.58, 9.10) 5.80(4.20, 8.03) -5.013 0.010 TC (mmol/L) c 4.42 ± 1.14 4.28 ± 1.50 -7.744 0.452 TG (mmol/L) b 1.56(1.00, 2.52) 1.18(0.87, 1.60) -3.577 0.199 HDL (mmol/L) b 1.19(0.94, 1.37) 1.17(0.86, 1.40) -1.442 0.947 LDL (mmol/L) b 2.43 (2.00, 3.00) 2.40(1.83, 3.03) -1.072 0.828 RBP (mg/L) b 97.55(78.00, 113.00) 100.54(75.50, 121.00) -5.439 < 0.001 BUN (mmol/L) c 18.88 ± 7.49 19.01 ± 10.72 2.330 0.808 Creatinine (umol/L) c 836.74 ± 305.07 659.63 ± 107.23 1.638 < 0.001 PTH (pg/ml) b 279.50(149.00, 579.25) 189.00(93.54, 314.00) -11.581 0.218 Procalcitonin (µg/L) b 0.66(0.40, 0.84) 0.44(0.24, 0.65) -15.727 0.740 SF (ug/L) b 75.00(32.00, 210.00) 103.50(39.75, 335.75) -1.374 0.724 CO 2 CP (mmol/L) c 25.52 ± 9.76 23.39 ± 4.65 2.947 0.215 Fe 2+ (µmol/L) c 12.63 ± 4.00 10.16 ± 5.97 1.166 0.165 Ca 2+ (mmol/L) c 2.27 ± 0.23 2.19 ± 0.72 1.783 0.092 P (mmol/L) c 1.86 ± 1.41 2.88 ± 0.87 7.496 0.238 Note: a: n(%), b: M༈P25, P75༉,c: ‾X ± S Abbreviations: BMI, body mass index; AST, aspartate aminotransferase; ALP, alkaline phosphatase; CHE, choline esterase; PAB, serum prealbumin; Hs-CRP, hypersensitive C-reaction protein; RBC, red blood cell; STB, total bilirubin; TC, total cholesterol; TG, triglyceride; HDL, high density lipoprotein; LDL, low density lipoprotein; RBP, retinol-binding protein; BUN, B-type natriuretic peptide; PTH, parathyroid hormone; SF, serum ferroprotein; CO 2 CP, carbon dioxide combining power; Fe 2+ , serum iron; Ca 2+ , serum calcium; P, serum phosphorus 3.3 Model building The 18 characteristic variables in Table 1 were used as the input variables of the model, and hypoproteinemia was the outcome variable. After optimizing the parameters by k-fold (k = 5) cross-validation and the grid search method, it was found that the RF algorithm has the best overall performance under the premise of max _ depth = 7, n _ estimators = 170, and max _ features = 2. The SVM kernel algorithm was applied to complete the prediction of large-scale problems through iterative solution of subproblems. The results showed that the model was more stable when C = 21 and gamma = 0.01. The LR algorithm performed better if C = 0.184 and penalty = 12. In the BP neural network, the input layer was used to collect information, and the hidden layer was used for analysis and processing. The model has the best performance under the conditions of dropout = 0.2, epochs = 100, batch size = 20, and verbose = 0. To ensure the stability of the model, the average value of the k-fold (k = 5) cross-validation of each algorithm was selected as the result of the algorithm in the accuracy, accuracy, balanced F score, AUC, sensitivity, and specificity indicators (Table 1 ). In the training set, the 4 machine learning algorithms showed significant differences in 6 dimensions (P < 0.05). Taking stability as the primary indicator of modeling, the mean values of each algorithm evaluation index were used to evaluate the effectiveness of the model. The differences among the 4 machine learning algorithms in all 6 evaluation dimensions were significant (P < 0.05). The RF algorithm had the best overall performance with the largest ROC curve area and AUC average value (0.924 ± 0.024) (see Table 2 ). Table 2 Mean of k-fold cross-validation results of 4 machine learning algorithms, Mean(95%CI) Model Precision (95%CI) Accuracy (95%CI) F1 score (95%CI) AUC (95%CI) Sensitivity (95%CI) Specificity (95%CI) RF 0.843(0.804–0.88) 0.924(0.896–0.954) 0.751(0.712–0.791) 0.924(0.891–0.956) 0.657(0.603–0.711) 0.839(0.805–0.874) SVM 0.846(0.805–0.886) 0.78(0.693–0.870) 0.736(0.637–0.836) 0.888(0.867–0.908) 0.815(0.727–0.902) 0.704(0.593–0.815) LR 0.740(0.628–0.852) 0.757(0.663–0.852) 0.696(0.585–0.808) 0.838(0.734–0.945) 0.639(0.584–0.695) 0.846(0.774–0.917) BP 0.753(0.662–0.843) 0.683(0.553–0.812) 0.704(0.659–0.748) 0.798(0.666–0.929) 0.756(0.657–0.835) 0.792(0.728–0.856) P value P < 0.05 P < 0.05 P < 0.05 P < 0.05 P < 0.05 P < 0.05 Note: RF, random forest; SVM, support vector machine; LR, logistic regression; BP: back propagation 3.4 Model evaluation The test set was used to verify the effectiveness of the model. The results showed that the RF model was superior to the other 3 models in the 6 evaluation dimensions, but the sensitivity of the SVM and BP neural networks was higher than that of the RF model (Fig. 2 , Table 3 ). The results of the confusion matrix (Fig. 3 ) showed that the SVM and BP neural networks produced many false negatives (FN) and false positives (FP) (SVM: FN = 9, FP = 3; BP: FN = 9, FP = 6) but fewer in RF and LR (RF: FN = 5, FP = 1; LR: FN = 3, FP = 7). Table 3 Comparison of model prediction results Model Precision Accuracy F1 score Sensitivity Specificity RF 0.955 0.936 0.875 0.808 0.985 SVM 0.719 0.872 0.793 0.885 0.868 LR 0.864 0.894 0.792 0.731 0.956 BP 0.759 0.883 0.800 0.846 0.897 3.5 Ranking of feature importance in RFs The RF model was used as the optimal prediction model to rank the importance of the characteristic variables, with ordering as follows: hs-CRP, age, weight, high-throughput dialysis, dialysis age, pulmonary infection, retinol-binding protein, diabetes, hemoglobin, and dialysis frequency (Fig. 4 ). 4. Discussion In this study, the number of patients with hypoproteinemia accounted for approximately 30.8%, of which the training set accounted for approximately 31.6%, and the test set accounted for approximately 27.7%. In the study of Tianming He et al. [ 17 ], the incidence of hypoproteinemia in patients during peridialysis was 17.78%. Tang X et al.[ 18 ] showed that the incidence of hypoproteinemia in MHD patients was 20.55% (188/915), which was higher than that in MHD patients. It is inferred that the incidence of hypoproteinemia in patients undergoing MHD may be high throughout the peridialysis period. The reason for the results of this study being slightly higher than those of similar studies may be due to the small sample size and the differences in dietary habits between regions. To our knowledge, this is the first study to develop and validate a risk prediction model for hypoproteinemia in MHD patients. Many risk prediction models have been developed to predict the related complications of hemodialysis patients, and the research methods were mostly single factor analysis and multiple linear regression[ 19 ]. Although the algorithm can be used to check the influence of input variables on output variables, it cannot provide a clear explanation of collinear variables, and as a consequence, the prediction accuracy of the established models is limited. Some scholars have also tried to apply machine learning algorithms to predict the mortality[ 20 ] and complications[ 21 , 22 ] of hemodialysis patients, but hypoproteinemia has not been analyzed as an independent outcome indicator. Therefore, this study is the first attempt to use machine learning algorithms to predict hypoproteinemia in MHD patients. We input all alternative predictors into the model for model development through a literature review rather than statistical numerical criteria, that is, to reduce the bias generated in the process of screening predictors ending up with the least model overfitting. According to the results of ROC analysis, the prediction accuracy of the RF algorithm model was higher than that of the other three models in both the validation set and the test set. In this study, the RF algorithm had the best potential to predict the occurrence of hypoproteinemia in MHD patients. The RF algorithm predicts large-scale data by introducing random samples and random features. It does not need to make strict assumptions on the original data. It can process thousands of input variables at the same time and can cluster and locate outliers. In the absence of many data features, it can still maintain high accuracy, which makes the random forest algorithm widely used in clinical research. A study enrolled 858 patients with CKD and predicted progression to renal failure and important features, which revealed that a random forest (RF) classifier with a synthetic minority oversampling technique (SMOTE) had the best predictive performance among patients with early-stage CKD who progressed within 3 and 5 years and among patients with advanced-stage CKD who progressed within 1 and 3 years[ 23 ]. The RF algorithm also has many applications in other disease prediction models. For example, Chen et al.[ 24 ] verified nine machine learning algorithms to predict the occurrence of acute kidney injury (AKI) in elderly orthopedic patients, and the results showed that the RF algorithm could better predict AKI. In the study of Kalafi et al.[ 25 ] predicting the survival rate of 4902 breast cancer patients, the results showed that the RF algorithm had a higher accuracy of 83.3%, while the SVM algorithm had a lower accuracy of 80.5%. The conclusions of the above studies are consistent with the optimal prediction model in this study. In addition, in this study, MHD patients had significant differences in most baseline factors between the hypoproteinemia group and the normal albumin group. Therefore, the baseline factors at the time of hypoproteinemia can be used to predict the occurrence of hypoproteinemia in other patients, which is replicable. In summary, the advantages of the ML algorithm lie in the following 3 aspects: First, ML algorithms can learn from past data or experience without explicit programming. The data used in this study were the relevant information in the electronic medical record system. Compared with the traditional modeling method, the ML algorithm can more effectively use all valid data and provide a clear explanation. Second, in terms of variable selection, traditional methods only include or eliminate variables based on significant differences between groups, which may lose some potential predictors. Traditional methods can only judge whether the variables affect the outcome indicators and cannot evaluate the weight of each variable in the model. However, practically, there are many factors affecting the occurrence of hypoproteinemia in MHD patients, and the factors are complicated and intertwined. As mentioned above, traditional algorithms are prone to bias and deviate from clinical reality. The ML algorithm could effectively compensate for the disadvantages above. The ML algorithm enabled the variables to be visualized by finding logical relationships. SHAP values could show the weight of predictors in the model and could also visually display risk or protective factors. Therefore, the prediction model in this study can fully reflect the influence of all predictors on hypoproteinemia and can predict the occurrence of hypoproteinemia before the change in serum albumin level, and the results are more accurate and effective. Third, ML is a subset of artificial intelligence that does not need to be calculated by assignment, omits the process of manual construction and evaluation, and has higher modeling efficiency, which greatly saves the time of clinical workers and improves work efficiency. In this study, many relevant data recorded in the hospital information system were used as data sources. With the support of big data analysis technology, the information was fully utilized and mined to develop a reliable and good predictive model for hypoproteinemia risk in MHD patients. Using this model enhanced the screening efficiency and accuracy of assessing hypoproteinemia in MHD patients, which could help doctors and nurses make more instructive clinical decisions and had far-reaching importance for the prevention, intervention and management of hypoproteinemia among MHD patients. In the future, we should combine information automation technology to further validate the RF model with good performance and develop software or apps to screen patients who may suffer from hypoproteinemia in a specified period. Through the system’s automatic warning, professionals could be prompted to perform further assessment and intervention. 5. Limitations This was a retrospective study. The selected subjects were patients who underwent regular dialysis in the hospital. There may be bias. Prospective medical records should be collected in future studies to further verify the model. This study was limited by region and sample size and did not obtain relevant information such as protein intake, which may limit the application of the model. In the future, comprehensive data on patient nutrition and etiology should be collected. This prediction model lacks external verification. In the future, it will be extrapolated to other research centers to verify the generalization ability of the model. 6. Conclusion Hs-CRP, age, weight, application of high-flux dialyzer, and diabetes are risk factors for hypoproteinemia in MHD patients. The RF model has a better effect on predicting the risk of hypoproteinemia in MHD patients. The algorithm has high classification accuracy and prediction accuracy, and its theory and method research are relatively mature. Therefore, using the RF algorithm to predict the risk of hypoproteinemia in MHD patients in future research and applications is recommended. Declarations Ethical approval and consent to participate The study has been performed in accordance with the Declaration of Helsinki and was approved by the Clinical Ethics Committee of the First Bethune Hospital of Jilin University with approval No. (2023-255). Informed consent was obtained from all subjects and/or their legal guardian(s). Consent to publish not applicable. Funding No funding. Conflict of interest All authors disclosed no relevant relationships. Author contribution This article is written by "Wang yao, Yang Jingshu", "Yang Jingshu, Wang yao, Wang haiyan, Wang Song Yu" independent research design and data analysis, "Cao Hongshi" led to write the article, "Duan Xiaotian" proofread all drafts, "Cao Hongshi" to guide the paper for the first time, "Wang Haiyan" objective review of the article. All of the authors have contributed to the further revision of this article. Finally, all the authors have contributed to the writing and revision of the article, and have some constructive suggestions for this article, ensuring that the article can accurately express complex research results. In addition, all study members received consistent study participant information and outcome reports based on evidence. Availability of data and materials see appendix 2 in the Related Files on the upload section. References Cui Jinrui, X.Q.K.J., The shared decision-making experience of elderly patients with end-stage renal disease:a systematic review and meta synthesis of qualitative studies. Chinese Journal of Nursing, 2022. 7(57): p. 863-871. Canaud, B., et al., Global prevalent use, trends and practices in hemodiafiltration. Nephrol Dial Transplant, 2020. 35(3): p. 398-407. Tang Xingming, Z.H.H.J., Current situation and risk factors analysis for hypoalbuminemia of maintenance hemodialysis patients: a multiple centers experience. Chinese Journal of Postgraduates of Medicine, 2021. 5(44): p. 411-415. Jin Jiansheng, Z.X.Z.M., Relationship between Microinflammatory State Related-Factors and Hypoalbuminemia in Maintenance Hemodialysis Patients. Journal of Fujian Medical University, 2006. 2(40): p. 161-163. Li Jie, W.Y.W.Q., Relationship between serum albumin level and prognosis of maintenance hemodialysis patients. Internal Medicine, 2012. 6(7): p. 586-588. Liang Weifeng, L.Y.X.S., Predictive value of bioelectrical impedance phase angle for hypoalbuminemia in maintenance hemodial-ysis patients. Chinese Journal of Blood Purification, 2022. 7(21): p. 478-482. Tang, J., et al., Early albumin level and mortality in hemodialysis patients: a retrospective study. Ann Palliat Med, 2021. 10(10): p. 10697-10705. Cheng, Y.J., et al., Effect of Intradialytic Exercise on Physical Performance and Cardiovascular Risk Factors in Patients Receiving Maintenance Hemodialysis: A Pilot and Feasibility Study. Blood Purif, 2020. 49(4): p. 409-418. Ma Li, F.S.Z.Y., The value of albumin to C-reactive protein ratio, red blood cell distribution width, and serum uric acid in evaluating the prognosis of maintenance hemodialysis patients. Chinese Journal of Blood Purification, 2021. 6(20): p. 373-377. Kim, Y., et al., Relative contributions of inflammation and inadequate protein intake to hypoalbuminemia in patients on maintenance hemodialysis. Int Urol Nephrol, 2013. 45(1): p. 215-27. Mohajerani, F., et al., Mass Transport in High-Flux Hemodialysis: Application of Engineering Principles to Clinical Prescription. Clin J Am Soc Nephrol, 2022. 17(5): p. 749-756. Adanan, N., et al., Investigating Physical and Nutritional Changes During Prolonged Intermittent Fasting in Hemodialysis Patients: A Prospective Cohort Study. J Ren Nutr, 2020. 30(2): p. e15-e26. Nwanosike, E.M., et al., Potential applications and performance of machine learning techniques and algorithms in clinical practice: A systematic review. Int J Med Inform, 2022. 159(5): p. 104679. Bacchi, S., et al., Machine learning in the prediction of medical inpatient length of stay. Intern Med J, 2022. 52(2): p. 176-185. Gao Yongxiang, Z.J., Determination of Sample Size in Logistic Regression Analysis. The Journal of Evidence-Based Medicine, 2018. 2(18): p. 122-124. Lambert, K., et al., Commentary on the 2020 update of the KDOQI clinical practice guideline for nutrition in chronic kidney disease. Nephrology (Carlton), 2022. 27(6): p. 537-540. He, T., et al., Risk factors for infection-related hospitalization in end-stage renal disease patients during peri-dialysis period. Ther Apher Dial, 2022. 26(4): p. 717-725. Tang Xingming, Z.H.H.J., Current situation and risk factors analysis for hypoalbuminemia of maintenance hemodialysis patients: a multiple centers experience. Chin J Postgrad Med, 2023. 44(5): p. 411-415. Wang, Y., et al., Clinical Prediction of Heart Failure in Hemodialysis Patients: Based on the Extreme Gradient Boosting Method. Front Genet, 2022. 13(4): p. 889378. Radovic, N., et al., Machine learning approach in mortality rate prediction for hemodialysis patients. Comput Methods Biomech Biomed Engin, 2022. 25(1): p. 111-122. Othman, M., et al., Early prediction of hemodialysis complications employing ensemble techniques. Biomed Eng Online, 2022. 21(1): p. 74. Zhang, H., et al., Real-time prediction of intradialytic hypotension using machine learning and cloud computing infrastructure. Nephrol Dial Transplant, 2023. 15(4): p. 1-28. Su, C., et al., Machine Learning Models for the Prediction of Renal Failure in Chronic Kidney Disease: A Retrospective Cohort Study. Diagnostics, 2022. 12(10): p. 2454. Chen, Q., et al., Application of Machine Learning Algorithms to Predict Acute Kidney Injury in Elderly Orthopedic Postoperative Patients. Clin Interv Aging, 2022. 17(5): p. 317-330. Kalafi, E.Y., et al., Machine Learning and Deep Learning Approaches in Breast Cancer Survival Prediction Using Clinical Data. Folia Biol (Praha), 2019. 65(5-6): p. 212-220. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 05 Apr, 2024 Reviewers agreed at journal 31 Mar, 2024 Reviewers agreed at journal 29 Feb, 2024 Reviews received at journal 17 Dec, 2023 Reviewers agreed at journal 16 Dec, 2023 Reviewers invited by journal 29 Nov, 2023 Editor assigned by journal 29 Nov, 2023 Editor invited by journal 30 Aug, 2023 Submission checks completed at journal 30 Aug, 2023 First submitted to journal 31 Jul, 2023 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3219283","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":229916627,"identity":"b6b95684-2667-4575-a2a7-60dc7c089617","order_by":0,"name":"Wang Yao","email":"","orcid":"","institution":"First Hospital of Jilin University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wang","middleName":"","lastName":"Yao","suffix":""},{"id":229916628,"identity":"ce727f9d-96d3-422b-a28e-f0153f423390","order_by":1,"name":"Yang Jingshu","email":"","orcid":"","institution":"First Hospital of Jilin University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yang","middleName":"","lastName":"Jingshu","suffix":""},{"id":229916629,"identity":"0e4be2ed-c9f7-46c9-9142-929701495ab3","order_by":2,"name":"Wang Haiyan","email":"","orcid":"","institution":"First Hospital of Jilin University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wang","middleName":"","lastName":"Haiyan","suffix":""},{"id":229916630,"identity":"56d50708-ed8b-40bb-88a5-e2d9e145b908","order_by":3,"name":"Zhang Huiru","email":"","orcid":"","institution":"First Hospital of Jilin University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Zhang","middleName":"","lastName":"Huiru","suffix":""},{"id":229916631,"identity":"476109ca-eb62-4e7c-8d79-ad327ffad13f","order_by":4,"name":"Duan Xiaotian","email":"","orcid":"","institution":"First Hospital of Jilin University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Duan","middleName":"","lastName":"Xiaotian","suffix":""},{"id":229916632,"identity":"bc132750-1fc7-4342-8bda-e9ca039df1f9","order_by":5,"name":"Wang Songyu","email":"","orcid":"","institution":"First Hospital of Jilin University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wang","middleName":"","lastName":"Songyu","suffix":""},{"id":229916633,"identity":"ce25597b-ba75-40ac-bf5b-1c7cb94909c2","order_by":6,"name":"Cao Hongshi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAwUlEQVRIiWNgGAWjYBAC/gYGNgaGCjYZEEeCKC0SB0BazrDxEK/FwAGohbGNgRQt7IePPfw5j4/H4ADzwds8DHZ5hLXwpKUb825jA2phS7bmYUguJqyFIcdMmhGshcdMmofhQGIDQS38b8wkf84BaeH/RqSWiBwzCd4GsC1sxGmRuPEsTZrnGBuP5GE2Y8s5BsmEtfD3Jx+T/FFzTI7vePPDG28q7AhrgYJjDAzMYHcSqR4IaohXOgpGwSgYBSMPAABYBzJwFKIN/QAAAABJRU5ErkJggg==","orcid":"","institution":"First Hospital of Jilin University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Cao","middleName":"","lastName":"Hongshi","suffix":""}],"badges":[],"createdAt":"2023-07-31 05:29:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3219283/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3219283/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":42664107,"identity":"88f6c215-d0bb-470e-9ccc-643273f67190","added_by":"auto","created_at":"2023-09-05 19:52:21","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":26403,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDiagram of the model\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-3219283/v1/d451bd07c6da06d46c88fa74.png"},{"id":42664109,"identity":"c9dd2aff-8d8a-45d3-9abf-1b362697f171","added_by":"auto","created_at":"2023-09-05 19:52:21","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":115024,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe ROC curves of the RF, SVM, LR and BP algorithms\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3219283/v1/25229cb334ce031ccc326ae2.jpg"},{"id":42664108,"identity":"bf6dc9ef-042f-4d86-8020-bc08286f14c7","added_by":"auto","created_at":"2023-09-05 19:52:21","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":147738,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eConfusion matrix of four machine learning models in the test set\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(A: RF; B: SVM; C: LR; D: BP neural networks)\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3219283/v1/5b7d3b082abf029f74adf546.jpg"},{"id":42664110,"identity":"9221d3ea-55d2-4458-a28b-4e2a79ce9168","added_by":"auto","created_at":"2023-09-05 19:52:21","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":221307,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eImportance of the features in the random forest algorithm\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3219283/v1/2dc435105c68fbe31ef5006a.jpg"},{"id":42664459,"identity":"ef4c4ae3-1c8c-4283-a7cc-a31a2a2e864a","added_by":"auto","created_at":"2023-09-05 20:00:21","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":625329,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3219283/v1/d5d16fe8-c9fb-4fc2-8e61-cf1248055849.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Predicting hypoproteinemia among patients undergoing maintenance hemodialysis: A development and validation study based on machine learning algorithms","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eESRD is the final stage of various renal diseases when renal function is irreversibly declining. Kidney transplantation, hemodialysis and peritoneal dialysis are the main treatment strategies[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Recently, because a shortage of kidney transplant donors has resulted in numerous ESRD patients not receiving transplantation in time, maintenance hemodialysis (MHD) has become an important alternate treatment for ESRD patients thanks to the maturity of dialysis techniques and the strong support of medical insurance.\u003c/p\u003e \u003cp\u003eBy 2020, the number of MHD patients in China was as high as 690,000[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Long-term hemodialysis may lead to a variety of complications, of which 20ཞ50% of MHD patients may suffer from hypoproteinemia[\u003cspan additionalcitationids=\"CR4 CR5\" citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. In the laboratory, hypoproteinemia can be diagnosed by total plasma protein\u0026thinsp;\u0026lt;\u0026thinsp;60 g/L or albumin\u0026thinsp;\u0026lt;\u0026thinsp;35 g/L. In terms of clinical manifestations, hypoproteinemia, which is one of the risk factors for cardiovascular and cerebrovascular diseases, infection, readmission and even death, could present as malnutrition and be prevented if corrected in time[\u003cspan additionalcitationids=\"CR8\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Early screening and identification is the key to the prevention and treatment of hypoproteinemia, which can effectively prevent the occurrence and development of hypoproteinemia in MHD patients and reduce the social and family medical economic burden. It was reported that microinflammatory status, age, weight, diabetes, and dialysis frequency are early risk factors for hypoproteinemia[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan additionalcitationids=\"CR11\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]; however, there is no research on the building of an early risk prediction model for hypoproteinemia.\u003c/p\u003e \u003cp\u003eMachine learning, as a hybrid of artificial intelligence-computer and statistics, can predict the occurrence of the disease by defining data attributes, using computers to mine the laws existing in data, and applying different classifiers[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Compared with traditional statistical methods, it can capture the nonlinear relationship in the data and the complexity between multiple predictive variables, explain the patient's specific level of prediction, accurately identify the risk factors for the disease, and make early diagnosis of the disease, which better solves the complexity and unpredictable nature of human physiology[\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn view of all of the above, the purpose of this study is to mine the target data set and establish a risk prediction model of hypoproteinemia in MHD patients based on 4 classical algorithms, including RF, SVM, LR and BP neural networks, and evaluate the model prediction efficiency using different algorithms at the same time to build a model that functions best. The clinical importance of this model is that it can help nurses identify changes in patient-related risk factors as early as possible in clinical practice, enhance nurses' risk perception ability, improve the early accurate screening and primary prevention of hypoproteinemia, and provide a basis for accurate intervention and treatment.\u003c/p\u003e"},{"header":"2. Object and method","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n\u003ch2\u003e2.1 Objects\u003c/h2\u003e\n\u003cp\u003eThis study is a single-center retrospective study. A total of 468 MHD patients who were admitted to the Department of Nephrology, First Bethune Hospital of Jilin University, Changchun City, Jilin Province, from January 1, 2021, to December 1, 2021, were selected as the subjects. A total of 43 variables were included in this study as candidate risk factors. According to the empirical method, the sample size was estimated to be 10ཞ15 times the variables[\u003cspan class=\"CitationRef\"\u003e15\u003c/span\u003e]. Therefore, the sample size required for this study was 430 cases in theory. Considering the 10% loss of data, we needed to collect at least 468 MHD patients.\u003c/p\u003e\n\u003cdiv id=\"Sec4\" class=\"Section3\"\u003e\n\u003ch2\u003e2.1.1 Inclusion criteria:\u003c/h2\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003eAge\u0026thinsp;\u0026ge;\u0026thinsp;18;\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eConsistent with the diagnostic criteria of ESRD in the International Classification of Diseases (ICD)[\u003cspan class=\"CitationRef\"\u003e16\u003c/span\u003e];\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eApplying arteriovenous fistula for dialysis treatment;\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eContinuous dialysis time;\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eDialysis frequency: 2ཞ3 times/week;\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eDialysis blood flow rate: 250ཞ300 mL/min;\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eDialysis fluid flow rate: 500 mL/min.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ol\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec5\" class=\"Section3\"\u003e\n\u003ch2\u003e2.1.2 Exclusion criteria:\u003c/h2\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003eDiagnosed with malignant tumor;\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eSuffering from nutritional intake/absorption disorders caused by digestive diseases;\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAlbumin synthesis disorder caused by decompensated liver function;\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eWith residual urine volume;\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eIncomplete medical records.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ol\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\n\u003ch2\u003e2.2 Data collection\u003c/h2\u003e\n\u003cp\u003eUsing the self-designed data collection tool Data Extraction Table of Hypoproteinemia Risk Factors in Maintenance Hemodialysis Patients', two researchers conducted a literature review of Web of Science, PubMed, CNKI, Wanfang and other databases. In the process of screening literature, if the opinions were not uniform, they were to be discussed with high-level researchers. The search strategies were as follows:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003eRetrieval time: 2001.01.01ཞ2021.01.01;\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eDatabase: English databases include Web of Science, Embase, PubMed, and Cochrane libraries, Chinese databases include China National Knowledge Infrastructure (CNKI), Wanfang Data Knowledge Service Platform (Wanfang), VIP Journal Resource Integration Service Platform (VIP), and China Biomedical Literature Database (CBM);\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eSearch terms: Chinese search terms included \"maintenance hemodialysis/hemodialysis/MHD and hypoproteinemia/albumin levels and influencing factors/risk factors/predictive factors/risk factors\"; English search terms included \"maintenance hemodialysis/hematodialysis/MHD and hypoproteinemia/albumin levels and influencing factors/risk factors/predictors/risk factors\".\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003eThe data of MHD patients with hypoproteinemia mainly included the following 4 aspects: ① basic information; ② comorbidity data; ③ dialysis-related data; and ④ biochemical examination data (see Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab1\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eData of MHD Patients with Hypoproteinemia\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eData Type\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eFactors\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eBasic information\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eSex, age, height, body weight, and BMI\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eComorbidity\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eHepatitis B, pneumonia, tuberculosis, diabetes, hypertension and associated diseases\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eDialysis\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eDialysis age, dialysis frequency, hemodialysis treatment mode, whether to use high-throughput dialysis\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eLaboratory examination\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAsparagine amino acid transferase, alkaline phosphatase, cholinesterase, prealbumin, total protein, albumin, globulin, hemoglobin, total bilirubin, total cholesterol, triglycerides, high-density lipoprotein, retinol-binding protein, urea nitrogen, creatinine, uric acid, cystatin, blood calcium, serum phosphorus, serum iron, ferritin, hypersensitivity C reactive protein, erythrocyte count, parathyroid hormone, procalcitonin, cardiac troponin, myoglobin, CO\u003csub\u003e2\u003c/sub\u003e binding force\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\n\u003ch2\u003e2.3 Data preprocessing\u003c/h2\u003e\n\u003cp\u003eData preprocessing is one of the important procedures of machine learning and data mining and includes the treatment of missing and abnormal values. Following the law of missing value replacement in Python, we use mode to fill missing values in discrete characteristics and means and medians to fill continuous characteristics; isolated forest algorithm was applied to remove outliers; data Z score normalization was used to standardize the data, and label coding and one-hot transformation were used for discrete variables.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\n\u003ch2\u003e2.4 Model development\u003c/h2\u003e\n\u003cp\u003eConsidering that this study was a small sample study and collected multiple binary variables, four algorithms of RF, SVM, LR and BP neural network were selected to develop the model. The process is shown in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003eFeature selection. In this data modeling analysis, the incidence of hypoproteinemia in MHD patients was used as the classification target. Classification prediction was performed according to the first serum albumin and total protein results of objects. Serum albumin 60 g/L was selected as the output variable of the model. Through statistical analysis, the characteristic variables with P\u0026thinsp;\u0026lt;\u0026thinsp;0.05 in \u0026ldquo;the risk factor data extraction table\u0026rdquo; were used as the predictive variables of the model.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eModel training. The optimal parameters of the four models were determined by the k-fold (k\u0026thinsp;=\u0026thinsp;5) cross-validation algorithm and grid search method, including max _ depth, n _ estimators, max _ features, C, gamma, penalty, dropout, epoch, batch size and detail. Finally, 4 ML algorithm models were trained by the calculated hyperparameters. In addition, the robustness of the data set was verified and divided into 5 groups. Each group of 4/5 data points was used to train the model, and the other 1/5 data points were used to verify the model performance. The above process was repeated 5 times, and the average value of the five results was taken to generate an estimate for model comparison and evaluation.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eModel evaluation. We analyzed the receiver operating characteristic curve (ROC), confusion matrix, accuracy, balance F1 score, sensitivity, specificity and other indicators to evaluate the performance of different models.\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eModel validation. We internally validated the trained model in the test set to determine whether the model is universal. Similarly, we followed the method of step 3 to calculate the performance index of the model in the test set.\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ol\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\n\u003ch2\u003e2.5 Statistical analysis\u003c/h2\u003e\n\u003cp\u003eThe original data collected were analyzed by SPSS 25.0 software. The measurement data were described by (x\u0026thinsp;\u0026plusmn;\u0026thinsp;s) or M (P25, P75). Independent sample t tests or Mann‒Whitney U tests were used for comparisons between groups. We used the number of cases (%) for numeration data description and the X\u003csup\u003e2\u003c/sup\u003e test for comparison between groups of binary classification variables. Variables with P\u0026thinsp;\u0026lt;\u0026thinsp;0.05 were imported into Python3.10 for model development and evaluation. The importance of input variables was analyzed, and the risk factors for prehypoproteinemia in MHD patients in the optimal model were listed.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n\u003ch2\u003e3.1 Basic information\u003c/h2\u003e\n\u003cp\u003eOf all the 468 objectives included, 144 had hypoproteinemia, accounting for approximately 30.8%. The cohort of patients was randomly divided into a training set and a test set at a ratio of 8:2. A total of 374 subjects were included in the training set, of which 118 suffered from hypoproteinemia, accounting for approximately 31.6%. Other patients of 94 subjects were included in the test set, among whom 26 patients had hypoproteinemia, accounting for approximately 27.7%.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\n\u003ch2\u003e3.2 Characteristic variable selection\u003c/h2\u003e\n\u003cp\u003eA total of 43 research variables were included for data preprocessing. After removing variables with data missing greater than 30% (myoglobin (47.2%), cardiac troponin (43.6%), cystatin C (37.8%), uric acid (31.8%)), we included a total of 39 characteristic variables. There were significant differences in age, weight, high-sensitivity C-reactive protein, dialysis age, diabetes, retinol-binding protein and 18 other indicators between the 2 groups (P\u0026thinsp;\u0026lt;\u0026thinsp;0.05, see Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab2\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eComparison of baseline data between patients with hypoalbuminemia and those without hypoalbuminemia\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eFeature\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eNonhypoproteinemia group(n\u0026thinsp;=\u0026thinsp;324)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eHypoproteinemia group(n\u0026thinsp;=\u0026thinsp;144)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003et\u003c/em\u003e/\u003cem\u003eZ\u003c/em\u003e/\u003cem\u003eX\u003c/em\u003e\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eP\u003c/em\u003e\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAge (years) \u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e52.82(42.00, 64.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e57.49(48.00.69.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-2.602\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.009\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eSex (male)\u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e206(63.5)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e76(52.7)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.317\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.128\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eWeight(kg)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e67.60(58.25, 74.75)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e58.57(48.75, 67.25)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-3.633\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eHeight(cm)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e166.72(160.00.173.50)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e164.88(160.00, 170.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-0.776\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.983\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eBMI(kg/m\u003csup\u003e2\u003c/sup\u003e)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e23.99(20.81, 25.95)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e23.32(20.86, 25.38)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-3.763\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.398\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eDialysis age (years)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e4.00(1.77, 7.94)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.00(0.50, 5.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-5.784\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eDialysis frequency (times/week)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.87(3.00, 3.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.83(3.00, 3.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-2.705\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.005\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eHigh-throughput dialysis \u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e50(15.4)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e57(39.5)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e22.362\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eApplying hemodialysis mode \u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e302(93.2)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e124(86.1)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.440\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.507\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eDiabetes \u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e98(30.2)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e71(49.3)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e10.014\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.002\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eHypertension \u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e284(87.6)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e124(86.1)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3.463\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.963\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eHepatitis B \u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e20(6.1)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e41(28.4)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8.065\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eHepatitis C \u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5(1.5)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8(5.5)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3.128\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.077\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePulmonary infection \u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e63(19.4)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e45(31.2)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6.194\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.013\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePulmonary tuberculosis \u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1(0.3)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8(5.5)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e12.537\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eNumber of basic diseases (\u0026ge;\u0026thinsp;3) \u003csup\u003ea\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e202(62.35)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e95(65.97)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.064\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.801\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAST (U/L)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e14.00(11.00, 18.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e14.00(11.00, 21.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-1.466\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.012\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eALP (U/L)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e85.00(67.00, 120.25)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e86.00(65.75, 111.75)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-0.714\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.938\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCHE (U/L)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7327.49\u0026thinsp;\u0026plusmn;\u0026thinsp;365.54\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5320.19\u0026thinsp;\u0026plusmn;\u0026thinsp;276.187\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.201\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePAB (mg/L)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e377.36\u0026thinsp;\u0026plusmn;\u0026thinsp;108.54\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e261.96\u0026thinsp;\u0026plusmn;\u0026thinsp;105.19\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.173\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eGlobulin (g/L)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e31.09\u0026thinsp;\u0026plusmn;\u0026thinsp;5.07\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e29.19\u0026thinsp;\u0026plusmn;\u0026thinsp;7.56\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.972\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eHemoglobin (g/L)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e112.15\u0026thinsp;\u0026plusmn;\u0026thinsp;31.80\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e92.24\u0026thinsp;\u0026plusmn;\u0026thinsp;27.39\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.822\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eHs-CRP (mg/dL)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e4.02(3.00, 6.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e50.51(13.07, 110.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-15.382\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eRBC (10\u0026thinsp;~\u0026thinsp;12/L) \u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e6.66\u0026thinsp;\u0026plusmn;\u0026thinsp;1.64\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e4.32\u0026thinsp;\u0026plusmn;\u0026thinsp;0.71\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3.071\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.379\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eSTB (umol/L)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7.2(5.58, 9.10)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e5.80(4.20, 8.03)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-5.013\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.010\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eTC (mmol/L)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e4.42\u0026thinsp;\u0026plusmn;\u0026thinsp;1.14\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e4.28\u0026thinsp;\u0026plusmn;\u0026thinsp;1.50\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-7.744\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.452\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eTG (mmol/L)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.56(1.00, 2.52)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.18(0.87, 1.60)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-3.577\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.199\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eHDL (mmol/L)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.19(0.94, 1.37)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.17(0.86, 1.40)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-1.442\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.947\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eLDL (mmol/L)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.43 (2.00, 3.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.40(1.83, 3.03)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-1.072\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.828\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eRBP (mg/L)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e97.55(78.00, 113.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e100.54(75.50, 121.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-5.439\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eBUN (mmol/L)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e18.88\u0026thinsp;\u0026plusmn;\u0026thinsp;7.49\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e19.01\u0026thinsp;\u0026plusmn;\u0026thinsp;10.72\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.330\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.808\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCreatinine (umol/L)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e836.74\u0026thinsp;\u0026plusmn;\u0026thinsp;305.07\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e659.63\u0026thinsp;\u0026plusmn;\u0026thinsp;107.23\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.638\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePTH (pg/ml)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e279.50(149.00, 579.25)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e189.00(93.54, 314.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-11.581\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.218\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eProcalcitonin (\u0026micro;g/L)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.66(0.40, 0.84)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.44(0.24, 0.65)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-15.727\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.740\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eSF (ug/L)\u003csup\u003eb\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e75.00(32.00, 210.00)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e103.50(39.75, 335.75)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-1.374\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.724\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCO\u003csub\u003e2\u003c/sub\u003eCP (mmol/L)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e25.52\u0026thinsp;\u0026plusmn;\u0026thinsp;9.76\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e23.39\u0026thinsp;\u0026plusmn;\u0026thinsp;4.65\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.947\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.215\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eFe\u003csup\u003e2+\u003c/sup\u003e(\u0026micro;mol/L)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e12.63\u0026thinsp;\u0026plusmn;\u0026thinsp;4.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e10.16\u0026thinsp;\u0026plusmn;\u0026thinsp;5.97\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.166\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.165\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eCa\u003csup\u003e2+\u003c/sup\u003e(mmol/L)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.27\u0026thinsp;\u0026plusmn;\u0026thinsp;0.23\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.19\u0026thinsp;\u0026plusmn;\u0026thinsp;0.72\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.783\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.092\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eP (mmol/L)\u003csup\u003ec\u003c/sup\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.86\u0026thinsp;\u0026plusmn;\u0026thinsp;1.41\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.88\u0026thinsp;\u0026plusmn;\u0026thinsp;0.87\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e7.496\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.238\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"5\" align=\"left\"\u003e\n\u003cp\u003eNote: a: n(%), b: M༈P25, P75༉,c: \u0026oline;X\u0026thinsp;\u0026plusmn;\u0026thinsp;S\u003c/p\u003e\n\u003cp\u003eAbbreviations: BMI, body mass index; AST, aspartate aminotransferase; ALP, alkaline phosphatase; CHE, choline esterase; PAB, serum prealbumin; Hs-CRP, hypersensitive C-reaction protein; RBC, red blood cell; STB, total bilirubin; TC, total cholesterol; TG, triglyceride; HDL, high density lipoprotein; LDL, low density lipoprotein; RBP, retinol-binding protein; BUN, B-type natriuretic peptide; PTH, parathyroid hormone; SF, serum ferroprotein; CO\u003csub\u003e2\u003c/sub\u003eCP, carbon dioxide combining power; Fe\u003csup\u003e2+\u003c/sup\u003e, serum iron; Ca\u003csup\u003e2+\u003c/sup\u003e, serum calcium; P, serum phosphorus\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\n\u003ch2\u003e3.3 Model building\u003c/h2\u003e\n\u003cp\u003eThe 18 characteristic variables in Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e were used as the input variables of the model, and hypoproteinemia was the outcome variable. After optimizing the parameters by k-fold (k\u0026thinsp;=\u0026thinsp;5) cross-validation and the grid search method, it was found that the RF algorithm has the best overall performance under the premise of max _ depth\u0026thinsp;=\u0026thinsp;7, n _ estimators\u0026thinsp;=\u0026thinsp;170, and max _ features\u0026thinsp;=\u0026thinsp;2.\u003c/p\u003e\n\u003cp\u003eThe SVM kernel algorithm was applied to complete the prediction of large-scale problems through iterative solution of subproblems. The results showed that the model was more stable when C\u0026thinsp;=\u0026thinsp;21 and gamma\u0026thinsp;=\u0026thinsp;0.01. The LR algorithm performed better if C\u0026thinsp;=\u0026thinsp;0.184 and penalty\u0026thinsp;=\u0026thinsp;12. In the BP neural network, the input layer was used to collect information, and the hidden layer was used for analysis and processing. The model has the best performance under the conditions of dropout\u0026thinsp;=\u0026thinsp;0.2, epochs\u0026thinsp;=\u0026thinsp;100, batch size\u0026thinsp;=\u0026thinsp;20, and verbose\u0026thinsp;=\u0026thinsp;0. To ensure the stability of the model, the average value of the k-fold (k\u0026thinsp;=\u0026thinsp;5) cross-validation of each algorithm was selected as the result of the algorithm in the accuracy, accuracy, balanced F score, AUC, sensitivity, and specificity indicators (Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). In the training set, the 4 machine learning algorithms showed significant differences in 6 dimensions (P\u0026thinsp;\u0026lt;\u0026thinsp;0.05). Taking stability as the primary indicator of modeling, the mean values of each algorithm evaluation index were used to evaluate the effectiveness of the model. The differences among the 4 machine learning algorithms in all 6 evaluation dimensions were significant (P\u0026thinsp;\u0026lt;\u0026thinsp;0.05). The RF algorithm had the best overall performance with the largest ROC curve area and AUC average value (0.924\u0026thinsp;\u0026plusmn;\u0026thinsp;0.024) (see Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab3\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eMean of k-fold cross-validation results of 4 machine learning algorithms, Mean(95%CI)\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eModel\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003ePrecision\u003c/p\u003e\n\u003cp\u003e(95%CI)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eAccuracy\u003c/p\u003e\n\u003cp\u003e(95%CI)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eF1 score\u003c/p\u003e\n\u003cp\u003e(95%CI)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eAUC\u003c/p\u003e\n\u003cp\u003e(95%CI)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eSensitivity\u003c/p\u003e\n\u003cp\u003e(95%CI)\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eSpecificity\u003c/p\u003e\n\u003cp\u003e(95%CI)\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eRF\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.843(0.804\u0026ndash;0.88)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.924(0.896\u0026ndash;0.954)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.751(0.712\u0026ndash;0.791)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.924(0.891\u0026ndash;0.956)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.657(0.603\u0026ndash;0.711)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.839(0.805\u0026ndash;0.874)\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eSVM\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.846(0.805\u0026ndash;0.886)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.78(0.693\u0026ndash;0.870)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.736(0.637\u0026ndash;0.836)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.888(0.867\u0026ndash;0.908)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.815(0.727\u0026ndash;0.902)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.704(0.593\u0026ndash;0.815)\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eLR\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.740(0.628\u0026ndash;0.852)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.757(0.663\u0026ndash;0.852)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.696(0.585\u0026ndash;0.808)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.838(0.734\u0026ndash;0.945)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.639(0.584\u0026ndash;0.695)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.846(0.774\u0026ndash;0.917)\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eBP\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.753(0.662\u0026ndash;0.843)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.683(0.553\u0026ndash;0.812)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.704(0.659\u0026ndash;0.748)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.798(0.666\u0026ndash;0.929)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.756(0.657\u0026ndash;0.835)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.792(0.728\u0026ndash;0.856)\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eP value\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003ctfoot\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"7\"\u003eNote: RF, random forest; SVM, support vector machine; LR, logistic regression; BP: back propagation\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tfoot\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\n\u003ch2\u003e3.4 Model evaluation\u003c/h2\u003e\n\u003cp\u003eThe test set was used to verify the effectiveness of the model. The results showed that the RF model was superior to the other 3 models in the 6 evaluation dimensions, but the sensitivity of the SVM and BP neural networks was higher than that of the RF model (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e, Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e). The results of the confusion matrix (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e) showed that the SVM and BP neural networks produced many false negatives (FN) and false positives (FP) (SVM: FN\u0026thinsp;=\u0026thinsp;9, FP\u0026thinsp;=\u0026thinsp;3; BP: FN\u0026thinsp;=\u0026thinsp;9, FP\u0026thinsp;=\u0026thinsp;6) but fewer in RF and LR (RF: FN\u0026thinsp;=\u0026thinsp;5, FP\u0026thinsp;=\u0026thinsp;1; LR: FN\u0026thinsp;=\u0026thinsp;3, FP\u0026thinsp;=\u0026thinsp;7).\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab4\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eComparison of model prediction results\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eModel\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003ePrecision\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eAccuracy\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eF1 score\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eSensitivity\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eSpecificity\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eRF\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.955\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.936\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.875\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.808\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.985\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eSVM\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.719\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.872\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.793\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.885\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.868\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eLR\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.864\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.894\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.792\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.731\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.956\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eBP\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.759\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.883\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.800\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.846\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.897\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\n\u003ch2\u003e3.5 Ranking of feature importance in RFs\u003c/h2\u003e\n\u003cp\u003eThe RF model was used as the optimal prediction model to rank the importance of the characteristic variables, with ordering as follows: hs-CRP, age, weight, high-throughput dialysis, dialysis age, pulmonary infection, retinol-binding protein, diabetes, hemoglobin, and dialysis frequency (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e\n\u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eIn this study, the number of patients with hypoproteinemia accounted for approximately 30.8%, of which the training set accounted for approximately 31.6%, and the test set accounted for approximately 27.7%. In the study of Tianming He et al. [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], the incidence of hypoproteinemia in patients during peridialysis was 17.78%. Tang X et al.[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] showed that the incidence of hypoproteinemia in MHD patients was 20.55% (188/915), which was higher than that in MHD patients. It is inferred that the incidence of hypoproteinemia in patients undergoing MHD may be high throughout the peridialysis period. The reason for the results of this study being slightly higher than those of similar studies may be due to the small sample size and the differences in dietary habits between regions.\u003c/p\u003e \u003cp\u003eTo our knowledge, this is the first study to develop and validate a risk prediction model for hypoproteinemia in MHD patients. Many risk prediction models have been developed to predict the related complications of hemodialysis patients, and the research methods were mostly single factor analysis and multiple linear regression[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Although the algorithm can be used to check the influence of input variables on output variables, it cannot provide a clear explanation of collinear variables, and as a consequence, the prediction accuracy of the established models is limited. Some scholars have also tried to apply machine learning algorithms to predict the mortality[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] and complications[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e] of hemodialysis patients, but hypoproteinemia has not been analyzed as an independent outcome indicator.\u003c/p\u003e \u003cp\u003eTherefore, this study is the first attempt to use machine learning algorithms to predict hypoproteinemia in MHD patients. We input all alternative predictors into the model for model development through a literature review rather than statistical numerical criteria, that is, to reduce the bias generated in the process of screening predictors ending up with the least model overfitting.\u003c/p\u003e \u003cp\u003eAccording to the results of ROC analysis, the prediction accuracy of the RF algorithm model was higher than that of the other three models in both the validation set and the test set. In this study, the RF algorithm had the best potential to predict the occurrence of hypoproteinemia in MHD patients. The RF algorithm predicts large-scale data by introducing random samples and random features. It does not need to make strict assumptions on the original data. It can process thousands of input variables at the same time and can cluster and locate outliers. In the absence of many data features, it can still maintain high accuracy, which makes the random forest algorithm widely used in clinical research.\u003c/p\u003e \u003cp\u003eA study enrolled 858 patients with CKD and predicted progression to renal failure and important features, which revealed that a random forest (RF) classifier with a synthetic minority oversampling technique (SMOTE) had the best predictive performance among patients with early-stage CKD who progressed within 3 and 5 years and among patients with advanced-stage CKD who progressed within 1 and 3 years[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. The RF algorithm also has many applications in other disease prediction models. For example, Chen et al.[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e] verified nine machine learning algorithms to predict the occurrence of acute kidney injury (AKI) in elderly orthopedic patients, and the results showed that the RF algorithm could better predict AKI. In the study of Kalafi et al.[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] predicting the survival rate of 4902 breast cancer patients, the results showed that the RF algorithm had a higher accuracy of 83.3%, while the SVM algorithm had a lower accuracy of 80.5%. The conclusions of the above studies are consistent with the optimal prediction model in this study.\u003c/p\u003e \u003cp\u003eIn addition, in this study, MHD patients had significant differences in most baseline factors between the hypoproteinemia group and the normal albumin group. Therefore, the baseline factors at the time of hypoproteinemia can be used to predict the occurrence of hypoproteinemia in other patients, which is replicable.\u003c/p\u003e \u003cp\u003eIn summary, the advantages of the ML algorithm lie in the following 3 aspects:\u003c/p\u003e \u003cp\u003eFirst, ML algorithms can learn from past data or experience without explicit programming. The data used in this study were the relevant information in the electronic medical record system. Compared with the traditional modeling method, the ML algorithm can more effectively use all valid data and provide a clear explanation.\u003c/p\u003e \u003cp\u003eSecond, in terms of variable selection, traditional methods only include or eliminate variables based on significant differences between groups, which may lose some potential predictors. Traditional methods can only judge whether the variables affect the outcome indicators and cannot evaluate the weight of each variable in the model. However, practically, there are many factors affecting the occurrence of hypoproteinemia in MHD patients, and the factors are complicated and intertwined. As mentioned above, traditional algorithms are prone to bias and deviate from clinical reality. The ML algorithm could effectively compensate for the disadvantages above. The ML algorithm enabled the variables to be visualized by finding logical relationships. SHAP values could show the weight of predictors in the model and could also visually display risk or protective factors. Therefore, the prediction model in this study can fully reflect the influence of all predictors on hypoproteinemia and can predict the occurrence of hypoproteinemia before the change in serum albumin level, and the results are more accurate and effective.\u003c/p\u003e \u003cp\u003eThird, ML is a subset of artificial intelligence that does not need to be calculated by assignment, omits the process of manual construction and evaluation, and has higher modeling efficiency, which greatly saves the time of clinical workers and improves work efficiency.\u003c/p\u003e \u003cp\u003eIn this study, many relevant data recorded in the hospital information system were used as data sources. With the support of big data analysis technology, the information was fully utilized and mined to develop a reliable and good predictive model for hypoproteinemia risk in MHD patients. Using this model enhanced the screening efficiency and accuracy of assessing hypoproteinemia in MHD patients, which could help doctors and nurses make more instructive clinical decisions and had far-reaching importance for the prevention, intervention and management of hypoproteinemia among MHD patients. In the future, we should combine information automation technology to further validate the RF model with good performance and develop software or apps to screen patients who may suffer from hypoproteinemia in a specified period. Through the system\u0026rsquo;s automatic warning, professionals could be prompted to perform further assessment and intervention.\u003c/p\u003e"},{"header":"5. Limitations","content":"\u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eThis was a retrospective study. The selected subjects were patients who underwent regular dialysis in the hospital. There may be bias. Prospective medical records should be collected in future studies to further verify the model.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eThis study was limited by region and sample size and did not obtain relevant information such as protein intake, which may limit the application of the model. In the future, comprehensive data on patient nutrition and etiology should be collected.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eThis prediction model lacks external verification. In the future, it will be extrapolated to other research centers to verify the generalization ability of the model.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e"},{"header":"6. Conclusion","content":"\u003cp\u003eHs-CRP, age, weight, application of high-flux dialyzer, and diabetes are risk factors for hypoproteinemia in MHD patients.\u003c/p\u003e\n\u003cp\u003eThe RF model has a better effect on predicting the risk of hypoproteinemia in MHD patients. The algorithm has high classification accuracy and prediction accuracy, and its theory and method research are relatively mature. Therefore, using the RF algorithm to predict the risk of hypoproteinemia in MHD patients in future research and applications is recommended.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthical approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe study has been performed in accordance with the Declaration of Helsinki and was approved by the Clinical Ethics Committee of the First Bethune Hospital of Jilin University with approval No. (2023-255). Informed consent was obtained from all subjects and/or their legal guardian(s).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to publish\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003enot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNo funding.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll authors disclosed no relevant relationships.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contribution\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis article is written by \u0026quot;Wang yao, Yang Jingshu\u0026quot;, \u0026quot;Yang Jingshu, Wang yao, Wang haiyan, Wang Song Yu\u0026quot; independent research design and data analysis, \u0026quot;Cao Hongshi\u0026quot; led to write the article, \u0026quot;Duan Xiaotian\u0026quot; proofread all drafts, \u0026quot;Cao Hongshi\u0026quot; to guide the paper for the first time, \u0026quot;Wang Haiyan\u0026quot; objective review of the article. All of the authors have contributed to the further revision of this article. Finally, all the authors have contributed to the writing and revision of the article, and have some constructive suggestions for this article, ensuring that the article can accurately express complex research results. In addition, all study members received consistent study participant information and outcome reports based on evidence.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003esee appendix 2 in the Related Files on the upload section.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eCui Jinrui, X.Q.K.J., The shared decision-making experience of elderly patients with end-stage renal disease:a systematic review and meta synthesis of qualitative studies. Chinese Journal of Nursing, 2022. 7(57): p. 863-871.\u003c/li\u003e\n\u003cli\u003eCanaud, B., et al., Global prevalent use, trends and practices in hemodiafiltration. Nephrol Dial Transplant, 2020. 35(3): p. 398-407.\u003c/li\u003e\n\u003cli\u003eTang Xingming, Z.H.H.J., Current situation and risk factors analysis for hypoalbuminemia of maintenance hemodialysis patients: a multiple centers experience. Chinese Journal of Postgraduates of Medicine, 2021. 5(44): p. 411-415.\u003c/li\u003e\n\u003cli\u003eJin Jiansheng, Z.X.Z.M., Relationship between Microinflammatory State Related-Factors and Hypoalbuminemia in Maintenance Hemodialysis Patients. Journal of Fujian Medical University, 2006. 2(40): p. 161-163.\u003c/li\u003e\n\u003cli\u003eLi Jie, W.Y.W.Q., Relationship between serum albumin level and prognosis of maintenance hemodialysis patients. Internal Medicine, 2012. 6(7): p. 586-588.\u003c/li\u003e\n\u003cli\u003eLiang Weifeng, L.Y.X.S., Predictive value of bioelectrical impedance phase angle for hypoalbuminemia in maintenance hemodial-ysis patients. Chinese Journal of Blood Purification, 2022. 7(21): p. 478-482.\u003c/li\u003e\n\u003cli\u003eTang, J., et al., Early albumin level and mortality in hemodialysis patients: a retrospective study. Ann Palliat Med, 2021. 10(10): p. 10697-10705.\u003c/li\u003e\n\u003cli\u003eCheng, Y.J., et al., Effect of Intradialytic Exercise on Physical Performance and Cardiovascular Risk Factors in Patients Receiving Maintenance Hemodialysis: A Pilot and Feasibility Study. Blood Purif, 2020. 49(4): p. 409-418.\u003c/li\u003e\n\u003cli\u003eMa Li, F.S.Z.Y., The value of albumin to C-reactive protein ratio, red blood cell distribution width, and serum uric acid in evaluating the prognosis of maintenance hemodialysis patients. Chinese Journal of Blood Purification, 2021. 6(20): p. 373-377.\u003c/li\u003e\n\u003cli\u003eKim, Y., et al., Relative contributions of inflammation and inadequate protein intake to hypoalbuminemia in patients on maintenance hemodialysis. Int Urol Nephrol, 2013. 45(1): p. 215-27.\u003c/li\u003e\n\u003cli\u003eMohajerani, F., et al., Mass Transport in High-Flux Hemodialysis: Application of Engineering Principles to Clinical Prescription. Clin J Am Soc Nephrol, 2022. 17(5): p. 749-756.\u003c/li\u003e\n\u003cli\u003eAdanan, N., et al., Investigating Physical and Nutritional Changes During Prolonged Intermittent Fasting in Hemodialysis Patients: A Prospective Cohort Study. J Ren Nutr, 2020. 30(2): p. e15-e26.\u003c/li\u003e\n\u003cli\u003eNwanosike, E.M., et al., Potential applications and performance of machine learning techniques and algorithms in clinical practice: A systematic review. Int J Med Inform, 2022. 159(5): p. 104679.\u003c/li\u003e\n\u003cli\u003eBacchi, S., et al., Machine learning in the prediction of medical inpatient length of stay. Intern Med J, 2022. 52(2): p. 176-185.\u003c/li\u003e\n\u003cli\u003eGao Yongxiang, Z.J., Determination of Sample Size in Logistic Regression Analysis. The Journal of Evidence-Based Medicine, 2018. 2(18): p. 122-124.\u003c/li\u003e\n\u003cli\u003eLambert, K., et al., Commentary on the 2020 update of the KDOQI clinical practice guideline for nutrition in chronic kidney disease. Nephrology (Carlton), 2022. 27(6): p. 537-540.\u003c/li\u003e\n\u003cli\u003eHe, T., et al., Risk factors for infection-related hospitalization in end-stage renal disease patients during peri-dialysis period. Ther Apher Dial, 2022. 26(4): p. 717-725.\u003c/li\u003e\n\u003cli\u003eTang Xingming, Z.H.H.J., Current situation and risk factors analysis for hypoalbuminemia of maintenance hemodialysis patients: a multiple centers experience. Chin J Postgrad Med, 2023. 44(5): p. 411-415.\u003c/li\u003e\n\u003cli\u003eWang, Y., et al., Clinical Prediction of Heart Failure in Hemodialysis Patients: Based on the Extreme Gradient Boosting Method. Front Genet, 2022. 13(4): p. 889378.\u003c/li\u003e\n\u003cli\u003eRadovic, N., et al., Machine learning approach in mortality rate prediction for hemodialysis patients. Comput Methods Biomech Biomed Engin, 2022. 25(1): p. 111-122.\u003c/li\u003e\n\u003cli\u003eOthman, M., et al., Early prediction of hemodialysis complications employing ensemble techniques. Biomed Eng Online, 2022. 21(1): p. 74.\u003c/li\u003e\n\u003cli\u003eZhang, H., et al., Real-time prediction of intradialytic hypotension using machine learning and cloud computing infrastructure. Nephrol Dial Transplant, 2023. 15(4): p. 1-28.\u003c/li\u003e\n\u003cli\u003eSu, C., et al., Machine Learning Models for the Prediction of Renal Failure in Chronic Kidney Disease: A Retrospective Cohort Study. Diagnostics, 2022. 12(10): p. 2454.\u003c/li\u003e\n\u003cli\u003eChen, Q., et al., Application of Machine Learning Algorithms to Predict Acute Kidney Injury in Elderly Orthopedic Postoperative Patients. Clin Interv Aging, 2022. 17(5): p. 317-330.\u003c/li\u003e\n\u003cli\u003eKalafi, E.Y., et al., Machine Learning and Deep Learning Approaches in Breast Cancer Survival Prediction Using Clinical Data. Folia Biol (Praha), 2019. 65(5-6): p. 212-220.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-nephrology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bnep","sideBox":"Learn more about [BMC Nephrology](http://bmcnephrol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bnep/default.aspx","title":"BMC Nephrology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Maintenance Hemodialysis, Hypoproteinemia, Machine Learning Algorithms, Prediction Model, Random Forest, Back Propagation Neural Network","lastPublishedDoi":"10.21203/rs.3.rs-3219283/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3219283/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003ePurpose\u003c/h2\u003e \u003cp\u003eMaintenance hemodialysis (MHD), which can cause various complications, is a common alternative therapy for patients with ESRD. This research built a prediction model of hypoproteinemia among ESRD patients based on machine learning algorithms.\u003c/p\u003e\u003ch2\u003eMethod\u003c/h2\u003e \u003cp\u003eA total of 468 patients were selected as subjects. The \u0026ldquo;hypoproteinemia risk factor data extraction table\u0026rdquo; was drawn up after a literature review. Univariate analysis was used to screen independent risk factors as prediction variables. After hyper parameter adjustment by k-fold (k\u0026thinsp;=\u0026thinsp;5) cross-validation and grid search, random forest (RF), support vector machine (SVM), back propagation (BP) neural network and logistic regression (LR) prediction models were developed. The model was evaluated by 6 dimensions, including AUROC, accuracy, precision, sensitivity, specificity and F1 score, and an importance matrix diagram was used to describe the importance.\u003c/p\u003e\u003ch2\u003eResult\u003c/h2\u003e \u003cp\u003eThe incidence of hypoproteinemia in total was 30.8%. According to univariate analysis, the difference between the hypoproteinemia and nonhypoproteinemia groups was significant in 18 aspects, including age, weight, dialysis duration, and dialysis frequency. In the training set, the AUROC values of the RF, SVM, and LR models were all greater than 0.8 unlike the BP neural network (0.798). The RF model had the highest AUC value (0.924). The specificities of the LR and RF models were similar (0.846 and 0.839, respectively), while the RF model had the best accuracy (0.924) and balanced F1 score (0.751). The models had higher performance indexes in the test set than in the training set, with the RF and BP models performing better in AUROC (0.981, 0.948) and the RF model being better in accuracy, specificity balanced F1 score and precision. The top 5 prediction variables were hypersensitivity C reactive protein, age, weight, usage of high-throughput dialyzers, and dialysis age.\u003c/p\u003e\u003ch2\u003eConclusion\u003c/h2\u003e \u003cp\u003e \u003cb\u003eThe\u003c/b\u003e RF model performed best. The model could help recognize characteristics related to hypoproteinemia during clinical practice, thereby enhancing nurses\u0026rsquo; risk perception and improving accurate screening, primary prevention and early intervention.\u003c/p\u003e","manuscriptTitle":"Predicting hypoproteinemia among patients undergoing maintenance hemodialysis: A development and validation study based on machine learning algorithms","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-09-05 19:52:16","doi":"10.21203/rs.3.rs-3219283/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2024-04-05T11:58:39+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"a465b9b2-3594-4ca0-a6a8-49f6414c6a9f_SNPRID","date":"2024-03-31T07:19:44+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"dd5492ff-dc1b-46a9-9d21-7f5d76bf28fc","date":"2024-02-29T20:31:33+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2023-12-17T08:04:30+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"1ab0b36b-47af-44de-b9be-6613e536fe42","date":"2023-12-16T08:30:42+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2023-11-29T10:43:00+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2023-11-29T05:55:05+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2023-08-30T15:05:22+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2023-08-30T15:04:18+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Nephrology","date":"2023-07-31T05:14:12+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-nephrology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bnep","sideBox":"Learn more about [BMC Nephrology](http://bmcnephrol.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bnep/default.aspx","title":"BMC Nephrology","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"579a9ccd-7001-411e-8e9f-10b4c0debd9e","owner":[],"postedDate":"September 5th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2024-05-10T08:58:00+00:00","versionOfRecord":[],"versionCreatedAt":"2023-09-05 19:52:16","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3219283","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3219283","identity":"rs-3219283","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.