An interpretable machine learning model for predicting chronic kidney disease risk among obese adults: a nationwide population-based study | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article An interpretable machine learning model for predicting chronic kidney disease risk among obese adults: a nationwide population-based study HyeokJun Yang, Young Gyun Seo, Nayoung Han This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9308066/v1 This work is licensed under a CC BY 4.0 License Status: Under Revision Version 1 posted 10 You are reading this latest preprint version Abstract Chronic kidney disease (CKD) is a major public health concern, particularly among individuals with obesity; however, early identification of high-risk individuals remains challenging. This study aimed to develop an interpretable machine learning model to predict CKD risk and identify key contributing factors using nationwide data. We analyzed 37,240 obese adults aged 19–79 years from the Korea National Health and Nutrition Examination Survey (2007–2023). Feature selection was performed using recursive feature elimination with cross-validation (RFECV), and an eXtreme Gradient Boosting (XGBoost) model was developed. Model performance was evaluated using cross-validation, calibration, and decision curve analysis, and interpretability was assessed using Shapley Additive exPlanations (SHAP) and individual conditional expectation analyses. Seventeen predictors were selected, with age, metabolic comorbidities, and central obesity measures identified as the most important factors. The final model showed moderate discriminative performance (AUROC 0.756) and probability calibration improved reliability. Lifestyle-related variables, including water and sodium intake, and frequency of eating out, showed modest contributions but were retained as modifiable predictors. Decision curve analysis indicated modest clinical utility, with the model primarily helping reduce unnecessary interventions. This interpretable model highlights the importance of central obesity and metabolic comorbidities while supporting the role of modifiable factors in CKD risk stratification. Health sciences/Diseases Health sciences/Health care Health sciences/Medical research Health sciences/Nephrology Health sciences/Risk factors chronic kidney disease obesity machine learning risk prevention Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Obesity is an independent risk factor for development and progression of chronic kidney disease (CKD). 1 A high body mass index (BMI) increases the prevalence of proteinuria by approximately 21%, 2 and an elevated waist-to-hip ratio increased the risk of CKD by 22%. 3 In addition, the increased prevalence of obesity is a major public health concern, owing to its significant association with chronic diseases, such as hypertension (HTN), diabetes mellitus (DM), and dyslipidemia (DL). 4 , 5 Metabolic syndrome contributes to a decline in renal function and the development of proteinuria. 6 Measures of central obesity, such as waist circumference (WC) or fat mass, may predict CKD incidence more accurately than BMI. 7 , 8 This is particularly relevant because patients with CKD are prone to malnutrition and muscle wasting, which can lead to an underestimation of adiposity when BMI alone is used. 9 , 10 In addition to obesity-related factors, demographic and physiological characteristics including age, sex differences, and muscle strength also influence CKD risk. 11 , 12 CKD often progresses silently, with minimal or no symptoms in its early stages, resulting in delayed diagnosis until significant renal impairment has occurred. 13 Patients with advanced CKD frequently require high-cost treatments, such as dialysis or kidney transplantation, owning to progression to end-stage kidney disease. Therefore, early identification of individuals at high risk for CKD is crucial, as timely intervention may delay disease progression and improve long-term outcomes and quality of life. Several risk prediction models have been developed to facilitate early detection of CKD. 14 , 15 However, previous studies have reported inconsistent findings regarding the demographic and clinical predictors of CKD, depending on the population studied. Moreover, most conventional approaches have relied on linear regression models, which are limited in their ability to capture complex nonlinear relationships and interactions among risk factors. Therefore, in this study, we aimed to identify key predictors of CKD among obese adults using a machine learning-based approach and to develop a robust and interpretable prediction model that accounts for nonlinear effects and feature interactions. Results Demographic characteristics Among the 133,375 participants, 37,240 were included in the final eligible population (Fig. 1 ). Participants with CKD were significantly older than those without CKD (63.5 ± 12.9 vs. 52.2 ± 15.2 years, p < 0.001) (Table 1 ). Although BMI did not significantly differ between the two groups, measures of central obesity, including WC and conicity index, were significantly higher in the CKD group than in the non-CKD group. Comorbid conditions, including HTN, DM, DL, metabolic syndrome, and cardiovascular disease, were significantly more common in the CKD group than in the non-CKD group (all p < 0.001). Regarding lifestyle factors, current smoking exposure was greater in the CKD group than in the non-CKD group, as reflected by longer smoking duration. Individuals with CKD reported lower daily water intake, while also exhibiting a higher frequency of eating out. In contrast, sodium intake and physical activity levels, including walking time and resistance exercise, were generally higher in the non-CKD group than in the CKD group, although some variables showed modest differences. Table 1 Demographic characteristics of study population. Variables CKD Non-CKD P-value N, % 2,966 34,274 Age 63.5 ± 12.9 52.2 ± 15.2 < 0.001 Male 1,507 (50.8) 16,739 (48.8) 0.039 BMI (kg/m 2 ) 27.1 ± 2.8 27.2 ± 2.7 0.123 Waist circumference (cm) 93.5 ± 7.2 91.5 ± 7.3 < 0.001 Conicity index 1.3 ± 0.1 1.26 ± 0.1 < 0.001 Comorbid diseases Hypertension 2,158 (72.8) 14,763 (43.1) < 0.001 Diabetes mellitus 1,245 (42.0) 6,247 (18.2) < 0.001 Dyslipidemia 2,124 (71.6) 19,565 (57.1) < 0.001 Metabolic syndrome 2,247 (75.8) 18,795 (54.8) < 0.001 Stroke 195 (6.6) 673 (2.0) < 0.001 Myocardial infarction 109 (3.7) 302 (0.9) < 0.001 Angina pectoris 176 (5.9) 652 (1.9) < 0.001 NAFLD 892 (30.1) 9,429 (27.5) Alcohol intake(g/day) 0.8 ± 1.7 1.0 ± 1.7 < 0.001 Smoking duration (pack per year) 10.5 ± 18.7 7.5 ± 14.5 < 0.001 Sleep duration 5 and < 9 hours 1,938 (65.3) 25,071 (73.1) ≥ 9 hours 380 (12.8) 3,150 (9.2) Working hours (h/week) 37.8 ± 20.7 42.1 ± 18.4 < 0.001 Walking time (min/week) 341.9 ± 1,105.8 367.3 ± 1,044.0 0.302 Weekly strength training days 0.015 Never 2,202 (74.2) 24,775 (72.3) 1 day 89 (3.0) 1,369 (4.0) 2 days 122 (4.1) 1,617 (4.7) 3 days 119 (4.0) 1,659 (4.8) 4 days 66 (2.2) 815 (2.4) 5 or more days 212 (7.1) 2,382 (6.9) Annual household income (per 10,000/year) 3,441.8 ± 3,473.5 4,582.3 ± 3,694.8 < 0.001 Household income quartile < 0.001 Low 1,070 (36.1) 6,638 (19.4) Lower-middle 773 (26.1) 8,949 (26.1) Upper-middle 559 (18.8) 9,364 (27.3) High 521 (17.6) 9,044 (26.4) Water intake (g/day) 1,960.3 ± 3,957.6 2,272.1 ± 5,824.7 < 0.001 Sodium intake (mg/day) 3,512.1 ± 2,476.9 4,055.8 ± 2,819.5 < 0.001 Frequency of eating out < 0.001 Twice a day or more 79 (2.7) 1,912 (5.6) Once a day 229 (7.7) 4,348 (12.7) 5–6 times a week 329 (11.1) 5,211 (15.2) 3–4 times a week 387 (13.0) 4,001 (11.7) 1–2 times a week 620 (20.9) 6,513 (19.0) 1–3 times a month 635 (21.4) 5,419 (15.8) a mean (IQR); b n (%) BMI, body mass index; CKD, chronic kidney disease; NAFLD, non-alcohol fatty liver disease Feature selection Recursive feature elimination with cross-validation (RFECV) process demonstrated a gradual improvement in model performance as the number of selected features increased. The mean area under the precision-recall curve (AUPRC) steadily increased from approximately 0.221 with a smaller subset of features and reached its peak at 17 features: age; sex; BMI; WC; comorbid diseases, including HTN, DM, and DL; concomitant metabolic syndrome; conicity index; daily alcohol intake; smoking duration; daily water intake; daily sodium intake; frequency of eating out; weekly walking time; weekly working hours; and annual total household income. Beyond this point, the addition of further variables did not result in meaningful performance gains and instead showed a slight decline in AUPRC, indicating diminishing returns from additional features. Based on these results, the optimal feature subset included 17 variables, which provided the best balance between predictive performance and model complexity. This pattern suggests that removing less informative variables improves model generalizability while reducing redundancy among predictors. Model development and evaluation Based on the confusion matrix of the model constructed using eXtreme Gradient Boosting (XGBoost), the model showed an accuracy of 84.51%, a precision of 0.244, a recall of 0.449, and an F1-score of 0.316. The final XGBoost model demonstrated acceptable discriminative performance, with an area under the receiver operating characteristic curve (AUROC) of 0.756 and an AUPRC of 0.240 (Fig. 2 ). The ROC curve indicated stable discrimination across thresholds, whereas the PR curve reflected the expected trade-off between precision and recall under class imbalance. Permutation importance analysis of the final XGBoost model indicated that age, comorbidities, and conicity index were major contributors to CKD prediction (Table 2 ). Feature importance ranking using Shapley Additive exPlanations (SHAP) analysis showed that age was the most strongly associated variable with CKD, followed by HTN, DM, DL, and conicity index (Fig. 3 ). Higher age and the presence of metabolic comorbidities were associated with increased CKD risk, whereas measures of central obesity showed stronger contributions than BMI. Lifestyle-related variables, including sodium intake, water intake, and frequency of eating out, demonstrated relatively low feature importance in the SHAP analysis. Although their individual contributions to model prediction were modest and showed heterogeneous patterns, these variables were consistently retained in the final model. Individual conditional expectation (ICE) plots further demonstrated that age had a nonlinear association with CKD risk, with a more pronounced increase at older ages (Fig. 4 ). In contrast, HTN and DM showed consistent positive effects across individuals, indicating more uniform contributions to risk. Table 2 Permutation feature importance of predictors in the XGBoost model Feature Importance Age 0.0735 ± 0.0061 Diabetes mellitus 0.0315 ± 0.0054 Hypertension 0.0309 ± 0.0055 Sex 0.0117 ± 0.0024 Smoking duration 0.0091 ± 0.0024 Dyslipidemia 0.0088 ± 0.0032 Conicity index 0.0067 ± 0.0036 Average weekly working hours 0.0060 ± 0.0022 Alcohol intake 0.0055 ± 0.0027 Sodium intake 0.0040 ± 0.0013 Household income quartile 0.0036 ± 0.0018 Water intake 0.0033 ± 0.0019 Metabolic syndrome 0.0032 ± 0.0023 Waist circumference 0.0029 ± 0.0024 Weekly walking time 0.0017 ± 0.0014 Frequency of eating out 0.0011 ± 0.0013 Body mass index -0.0003 ± 0.0010 Model calibration and clinical utility Probability calibration using Platt scaling substantially improved the agreement between predicted probabilities and observed outcomes. The calibrated model showed a markedly lower Brier score compared with the uncalibrated model (0.067 vs. 0.193), and the calibration curve was more closely aligned with the ideal reference line, indicating improved reliability of probability estimates. The optimal classification threshold notably differed between models. Although the uncalibrated model achieved the maximum F1-score at a higher threshold (0.656), the calibrated model achieved a comparable F1-score (0.319) at a substantially lower threshold (0.169). This shift suggests that probability calibration redistributes predicted risks, enabling more sensitive identification of CKD without compromising overall performance. Figure 5 presents the decision curve analysis (DCA) for the raw and calibrated XGBoost models. Across a wide range of threshold probabilities (0.0–1.0), both models showed similar net benefits (NBs) and largely overlapped with each other. Across the full range of threshold probabilities, the NB of both the raw and calibrated models remained close to zero and largely overlapped with the “treat none” strategy. However, the “treat all” strategy showed a marked decline in NB at higher threshold probabilities, whereas the model-based strategies remained relatively stable. Discussion In this study, we developed a machine learning model to predict CKD risk among obese adults using nationwide population-based data. The model demonstrated acceptable discriminative performance and identified key predictors, including age, metabolic comorbidities, and measures of central obesity. Through the integration of explainable artificial intelligence technology, we were able to gain insights into the relative importance and clinical effects of these predictor variables. Age was the most influential factor, showing a nonlinear association with CKD risk. The ICE analysis indicated that the predicted probability of CKD increased more significantly at older ages, suggesting an accelerated risk pattern beyond 40 years. This finding is consistent with the results of previous epidemiological studies demonstrating age-related decline in renal function and increased susceptibility to kidney damage. 16 , 17 The nonlinear pattern further highlights the advantage of machine learning approaches in capturing complex risk relationships that may be inadequately modeled using conventional linear methods. The identified comorbidities of this study are not only primary causes of CKD but also can be exacerbated by declining renal function, reflecting a bidirectional and synergistic relationship. 18 , 19 As most obese individuals have these clustered comorbidities, the cumulative risk of CKD significantly increases. 20 Among the factors related to obesity, the conicity index demonstrated a greater influence on CKD prediction than BMI or WC, suggesting that renal impairment is more closely associated with central adiposity than with overall body mass. 21 The conicity index is an indicator that incorporates WC, body weight, and height to determine the extent to which the body shape approximates a double cone, 22 measuring the degree of abdominal obesity. 23 These findings emphasize that managing abdominal fat distribution, rather than focusing solely on weight loss, is a more important clinical strategy for delaying the onset and progression of CKD. 24 Unlike non-modifiable factors, such as age, or partially modifiable conditions, such as established comorbidities, lifestyle-related variables, such as sodium intake, water intake, and frequency of eating out, showed relatively low but non-negligible contributions to CKD prediction. Variables related to physical activity and working conditions, including weekly working time and walking time, demonstrated minimal and non-directional contributions in the SHAP analysis. This may reflect the complex and multifactorial nature of lifestyle influences and potential measurement variability inherent in self-reported data. Nevertheless, the inclusion of these factors in the final model suggests that lifestyle factors represent modifiable risk components that can be targeted through behavioral interventions. 25 – 27 The final model demonstrated moderate discriminative performance. Although Platt calibration improved probability reliability, it did not lead to a meaningful change in decision-level NB; however, discrimination metrics, such as AUROC, remained largely unchanged. This finding is expected, as probability calibration improves the accuracy of predicted probabilities rather than the model’s ability to distinguish between classes. 28 Notably, the Brier score substantially decreased from 0.193 in the uncalibrated model, indicating poor reliability of predicted probabilities, to 0.067 after calibration, reflecting a marked improvement in probabilistic accuracy. Given that CKD screening relies on accurate estimation of individual risk rather than binary classification alone, this improvement in calibration is clinically meaningful and may enhance the applicability of the model in risk-based decision-making. 29 Despite these improvements, DCA demonstrated that the overall NB of the model remained modest and close to that of the “treat-none” strategy across most threshold probabilities. 30 Several factors may explain this observation. First, CKD in population-based cohorts is characterized by relatively low short-term event rates, which limits the absolute number of true positives and constrains achievable NB. 31 Second, the threshold for intervention in CKD is typically low, favoring early detection strategies, thereby reducing the relative advantage of model-based decision-making compared with default approaches. Third, many of the key predictors identified in this study, such as age and metabolic comorbidities, are already incorporated into routine clinical judgment, limiting the incremental value of the model. Nevertheless, the model consistently outperformed the “treat-all” strategy at moderate to high threshold probabilities, suggesting a potential role in reducing unnecessary interventions. 32 In clinical practice, indiscriminate application of screening or treatment may lead to increased healthcare burden and potential harm. In this context, the model may be useful for selectively identifying individuals who are less likely to benefit from intervention, thereby improving resource allocation. 33 This study has several strengths. It used a large, nationally representative dataset, enhancing generalizability within the target population. The use of a rigorous feature selection framework and nested cross-validation minimized overfitting and information leakage. Furthermore, the incorporation of SHAP and ICE analyses allowed for both global and individual-level interpretation, improving the transparency and clinical interpretability of the model. However, this study has some limitations. First, owing to the cross-sectional design, causal relationships cannot be established, and temporal prediction of CKD development is limited. In addition, the use of questionnaire-based variables may introduce measurement error and recall bias. Second, the proportion of CKD-positive cases was relatively low (approximately 8%), resulting in substantial class imbalance. In such settings, a model may achieve high overall accuracy by predominantly predicting the negative class, whereas its ability to correctly identify positive cases (e.g., precision and recall) remains limited. Because the AUPRC directly reflects class imbalance, the relatively low AUPRC observed in this study should be interpreted in this context. Third, the model was developed using data from a single country, and external validation in independent populations is required to confirm its generalizability. Finally, the modest NB observed in DCA suggests that further improvements in clinical utility may be achieved by incorporating longitudinal data or additional biomarkers in future studies. In conclusion, this study developed an interpretable machine learning model for CKD risk prediction among obese adults, highlighting the dominant roles of age, metabolic comorbidities, and central obesity. Although the model demonstrated acceptable performance and provided meaningful insights into risk factors, its clinical utility for guiding intervention remains modest. Future studie should focus on improving risk stratification through longitudinal modeling and integration of additional clinical and biological data. Methods Study design and data collection This study was a cross-sectional analysis using publicly available data from the Korea National Health and Nutrition Examination Survey (KNHANES) conducted between 2007 and 2023 (n = 133,375). As KNHANES data are publicly available and de-identified, institutional review board approval and informed consent were waived. Prior to analysis, inclusion and exclusion criteria were predefined: Adults aged 19–79 years were included, resulting in a final eligible population of 100,835 participants after applying the age criterion. Participants with missing values that precluded assessment of obesity status or CKD classification were excluded from the analysis. Obesity-related variables were defined and used as analytical features rather than inclusion criteria. Obesity was defined as a BMI ≥ 25 kg/m² and/or abdominal obesity, defined as WC ≥ 90 cm for men and ≥ 85 cm for women, according to Korean guidelines. CKD was defined based on the criteria of the Korean Society of Nephrology as any of the following: estimated glomerular filtration rate < 60 mL/min/1.73 m² calculated using the CKD-Epidemiology (CKD-EPI) or Modification of Diet in Renal Disease (MDRD) equation, presence of proteinuria on urinalysis, or self-reported physician-diagnosed CKD. The final analytic dataset was subsequently used for feature selection, model development, and evaluation. Complete-case analysis was applied for all subsequent modeling procedures. Data preprocessing Based on literature review, we selected 24 factors that had been significantly associated with CKD incidence in existing literature and could be directly identified or calculated from the KNHANES raw data: age; sex; comorbidities, such as DM, HTN, DL, metabolic syndrome, non-alcoholic fatty liver disease, myocardial infarction, and angina; history of cerebrovascular disease; BMI; WC; conicity index; alcohol intake; smoking duration; sleep time; water intake; sodium intake; frequency of eating out per week; weekly walking time; average weekly working hours; frequency of resistance training per week; and household income quartile (Table 3 ). 34 – 42 Table 3. Definition of variables used in this study Variables Definition Obesity BMI ≥ 25(kg/m 2 ) or WC ≥ 90cm (if male), ≥ 85 (if female) Body mass index weight(kg) / height(m) 2 Waist circumference ≥ 90 cm (if male), ≥ 85 cm (if female) Conicity index 0.109 −1 x WC(m) x [weight(kg)/height(m)] −1/2 ] Hypertension diagnosed with hypertension, taking antihypertensive medicines, SBP ≥140 mmHg or DBP ≥90 mmHg Diabetes mellitus diagnosed with diabetes mellitus, taking antidiabetic medicines, random glucose ≥126 mg/dL, or HbA1c ≥6.5% Dyslipidemia diagnosed with dyslipidemia, taking lipid-lowering medications, total cholesterol ≥240 mg/dL, LDL cholesterol ≥160 mg/dL, HDL cholesterol <40 mg/dL, or triglycerides ≥200 mg/dL Non-alcoholic fatty liver disease at least one of the three criteria applies: daily alcohol intake criteria (WHO standard; 10 g per drink) men: < 30 g/day women: < 20 g/day Hepatic Steatosis Index (HSI) cutoff: ≥ 36 HSI = 8 × (ALT/AST ratio) + BMI + 2 (if diabetes) + 2 (if female) Framingham Steatosis Index (FSI) cutoff: ≥ 23 FSI = -7.981 + 0.011 × age (years) + 0.146 × sex (female = 1, male = 0) + 0.173 × BMI (kg/m²) + 0.007 × triglycerides (mg/dL) + 0.593 × hypertension (yes = 1, no = 0) + 0.789 × diabetes (yes = 1, no = 0) + 1.1 × I(ALT:AST ratio ≥ 1.33) (yes = 1, no = 0) Metabolic syndrome having at least 3 out of 5 health conditions: abdominal obesity: waist circumference for men ≥90 cm; for women ≥80 cm high triglycerides: blood triglyceride levels ≥150 mg/dL low HDL cholesterol: blood HDL levels for men <40 mg/dL; for women <50 mg/dL high blood pressure: SBP/DBP ≥130/85 mmHg high fasting blood sugar: blood glucose levels ≥100 mg/dL History of stroke diagnosed with stroke or currently under treatment History of ischemic heart disease diagnosed with angina pectoris or myocardial infarction Alcohol intake average daily alcohol intake (g) Smoking duration (number of packs smoked per day) × (years of smoking) Sleep duration group less or equal to 5 hours = 1; more than 5 and less than 9 hours = 2; more or equal to 9 hours = 3 Water intake water intake from foods (g) + drinking water intake (cups) × 200 mL Sodium intake average daily sodium intake (mg) Frequency of eating out twice a day or more = 1; once a day = 2; 5-6 times a week = 3; 3-4 times a week = 4; 1-2 times a week = 5; 1-3 times a month = 6 Walking time daily total walking time × number of walking days per week Resistance training time none = 0; 1 or 2 days per week = 1; 3 or 4 days per week = 2; 5 days or more per week = 3 Total household income quartile low = 1; low-middle = 2; upper-middle = 3; high = 4 ALT, alanine aminotransferase; AST, aspartate transaminase; BMI, body mass index; DBP, diastolic blood pressure; HDL, high density lipoprotein; LDL, low density lipoprotein; SBP, systolic blood pressure; WC, waist circumference; WHO, World Health Organization To address missing values in all variables included in the analysis, a data imputation pipeline was constructed using Python (version 3.11.7). In the raw KNHANES dataset, variables were then categorized into continuous and categorical types, and missing values were handled separately for each group. Continuous variables were normalized using robust scaling to reduce the influence of outliers. Missing values in continuous variables were subsequently imputed using k-nearest neighbors (KNN) imputation with the number of neighbors set to k = 5, based on proximity in the scaled feature space. For categorical variables, missing values were imputed using a KNN-based approach adapted for categorical data, following appropriate numerical encoding of categorical levels. After imputation, the separately processed continuous and categorical datasets were merged to restore the original variable order for subsequent analyses. Feature selection To identify the most informative predictors for CKD classification while minimizing overfitting and information leakage, we performed RFECV within a nested cross-validation framework. 43 A Decision Tree classifier was used as the internal estimator for recursive feature elimination (RFE), as it can capture nonlinear relationships and interactions between clinical variables. At each elimination step, feature importance was computed using the impurity-based importance scores derived from the Decision Tree model, and the feature with the lowest importance was iteratively removed. Hyperparameter optimization of the internal estimator was conducted using Bayesian optimization with Optuna. At each feature elimination step, 40 Optuna trials were performed under a stratified 5-fold cross-validation framework to identify the hyperparameter combination that maximized the mean AUPRC. Class distribution was preserved across all folds. Given the class imbalance inherent in CKD prediction tasks, AUPRC was selected as the primary optimization metric, as it provides a more informative assessment of model performance in imbalanced datasets. The AUROC was additionally reported as a complementary metric. 44 Feature selection performance was evaluated using the cross-validated mean AUPRC derived from out-of-fold (OOF) predicted probabilities. The optimal number of features was defined as the subset yielding the highest mean AUPRC across folds. Model training and calibration using XGBoost Model training was performed using the XGBoost algorithm. To mitigate overfitting, both L1 (α) and L2 (λ) regularization terms on leaf weights were incorporated during model training. 45 Hyperparameter optimization was conducted using a Bayesian optimization framework. The Expected Improvement (EI) acquisition function was used with an exploration–exploitation trade-off parameter set to ξ = 0.01. The optimization process consisted of three initial random evaluations followed by 50 iterative optimization steps. Model performance at each evaluation point was assessed using stratified 5-fold cross-validation, with AUPRC selected as the primary optimization metric to appropriately account for class imbalance. 46 A Gaussian Process Regressor was used as the surrogate model, combining a Matérn kernel with a WhiteKernel to balance smoothness assumptions and noise tolerance within the hyperparameter search space. 47 The hyperparameter configuration yielding the highest cross-validated AUPRC was selected as the optimal set. Using the optimized hyperparameters, the final XGBoost model was retrained on the selected feature set. To improve the reliability of predicted probabilities, Platt scaling (sigmoid calibration) was applied. Calibration parameters were estimated via internal stratified 5-fold CV using the training data only. Model calibration performance before and after calibration was evaluated using the Brier score and calibration curves (reliability diagrams). 46 Model evaluation: Threshold determination and performance evaluation The optimal classification threshold was determined by maximizing the F1-score, following a commonly adopted approach for imbalanced binary classification tasks. Specifically, stratified 5-fold CV was performed on the training dataset, and OOF predicted probabilities were aggregated to construct a precision-recall curve. The probability threshold corresponding to the maximum F1-score was selected as the final cutoff and subsequently fixed for evaluation on the independent test set. $$\:F1\:score=\frac{2\times\:Precision\times\:Recall}{Precision+Recall}$$ Final model performance was assessed on the test set by computing accuracy, precision, recall, F1-score, confusion matrix, AUROC, and AUPRC. AUROC was used to evaluate overall discriminative ability, whereas AUPRC was used to assess performance for the positive class. Clinical utility was further evaluated using DCA, which quantifies NB across a range of threshold probabilities. 30 Net benefit was calculated as follows: $$\:NB\left(pt\right)=\frac{TP}{N}-\frac{FP}{N}\times\:\frac{pt}{1-pt}$$ where pt denotes the threshold probability. Model performance was compared against treat-all and treat-none strategies to assess potential clinical value. Model explanation To enhance model interpretability, two complementary approaches were used. First, permutation importance was calculated on the test set by randomly shuffling each feature and measuring the resulting decrease in AUPRC, thereby quantifying each feature’s contribution to predictive performance under unseen data conditions. Second, to enhance the interpretability of the final CKD prediction model, SHAP analysis was performed using TreeExplainer to leverage the tree-based structure of XGBoost. Feature importance was quantified using the mean absolute SHAP value across samples, and the most influential predictors were visualized using bar plots. In addition, SHAP summary (beeswarm) plots were generated to examine the directionality and distribution of each feature’s contribution to model predictions. Finally, ICE plots were generated to further investigate nonlinear feature effects. 48 ICE curves were constructed for the top three features identified by permutation importance, using both training and test sets, to examine individual-level prediction patterns as feature values varied. Probability calibration was performed using Platt scaling (sigmoid calibration). Specifically, the model’s raw output scores were mapped to calibrated probabilities via a logistic function \(\:P(y=1\mid\:s)=1/(1+\text{e}\text{x}\text{p}(As+B\left)\right)\) , where \(\:A\) and \(\:B\) were estimated using cross-validated predictions within the training data to prevent information leakage. Calibration performance was assessed using the Brier score and calibration curves (reliability diagrams). Statistical analyses Continuous variables are summarized as means ± standard deviations for normally distributed data or median with interquartile range (IQR) for non-normally distributed data. Normality was assessed using the Shapiro–Wilk test. Between-group comparisons were performed using the independent t -test for normally distributed variables and the Mann-Whitney U test for non-normally distributed variables. Categorical variables are presented as frequencies (n) and percentages (%), and group differences were assessed using the chi-squared (χ²) test. All statistical tests were two-tailed, and a p-value < 0.05 was considered statistically significant. Conventional statistical analyses were performed to compare baseline characteristics, whereas machine learning-based modeling and evaluation were performed separately as described above. All model development and analyses were performed using Python (version 3.11, 64-bit). Key libraries included pandas, numpy, matplotlib, scikit-learn, xgboost, shap, and scipy. To ensure reproducibility, random seeds were fixed (random_state = 42) throughout the analysis. Data preprocessing, model training, hyperparameter optimization, evaluation, and visualization were performed as an end-to-end automated pipeline, allowing the entire analysis to be reproducibly executed by specifying only the data paths and variable definitions. Declarations Competing interests The authors declare no competing interests. Ethics declarations This study was conducted in accordance with the Declaration of Helsinki and was approved by the Institutional Review Board of Jeju National University (IRB No. 2026-###-###). The data used in this study were obtained from the Korea National Health and Nutrition Examination Survey (KNHANES), which is publicly available. The requirement for informed consent was waived due to the use of anonymized secondary data. Funding Declaration This work was supported by the research grant of Jeju National University in 2023. Author Contribution N.H. conceived and designed the study. H.J.Y. and Y.G.S. performed data curation, formal analysis, and methodology development. H.J.Y. and developed the machine learning model and conducted statistical analyses. Y.G.S. drafted the manuscript and H.J.Y. contributed to data interpretation and critically revised the manuscript. All authors reviewed and approved the final manuscript. Data Availability The datasets used and/or analyzed during the current study are available from the corresponding author upon reasonable request. References Chang, A. R. et al. Adiposity and risk of decline in glomerular filtration rate: meta-analysis of individual participant data in a global consortium. BMJ 364 , k5301. 10.1136/bmj.k5301 (2019). Rosenstock, J. L., Pommier, M., Stoffels, G., Patel, S. & Michelis, M. F. Prevalence of Proteinuria and Albuminuria in an Obese Population and Associated Risk Factors. Front. Med. (Lausanne) . 5 , 122. 10.3389/fmed.2018.00122 (2018). Elsayed, E. F. et al. Waist-to-hip ratio, body mass index, and subsequent kidney disease and death. Am. J. Kidney Dis. 52 , 29–38. 10.1053/j.ajkd.2008.02.363 (2008). Jafari-Adli, S. et al. Prevalence of obesity and overweight in adults and children in Iran; a systematic review. J. Diabetes Metab. Disord . 13 , 121. 10.1186/s40200-014-0121-2 (2014). Despres, J. P. et al. Abdominal obesity and the metabolic syndrome: contribution to global cardiometabolic risk. Arterioscler. Thromb. Vasc Biol. 28 , 1039–1049. 10.1161/ATVBAHA.107.159228 (2008). Thomas, G. et al. Metabolic syndrome and kidney disease: a systematic review and meta-analysis. Clin. J. Am. Soc. Nephrol. 6 , 2364–2373. 10.2215/CJN.02180311 (2011). Jaroszynski, A. et al. Association of anthropometric measures of obesity and chronic kidney disease in elderly women. Ann. Agric. Environ. Med. 23 , 636–640. 10.5604/12321966.1226859 (2016). Evangelista, L. S., Cho, W. K. & Kim, Y. Obesity and chronic kidney disease: A population-based study among South Koreans. PLoS One . 13 , e0193559. 10.1371/journal.pone.0193559 (2018). Reis, J. P. et al. Comparison of overall obesity and body fat distribution in predicting risk of mortality. Obes. (Silver Spring) . 17 , 1232–1239. 10.1038/oby.2008.664 (2009). Kim, D. W. et al. Reproducibility and validity of an FFQ developed for the Korea National Health and Nutrition Examination Survey (KNHANES). Public. Health Nutr. 18 , 1369–1377. 10.1017/S1368980014001712 (2015). Carrero, J. J. Gender differences in chronic kidney disease: underpinnings and therapeutic implications. Kidney Blood Press. Res. 33 , 383–392. 10.1159/000320389 (2010). Chen, C. C. et al. Impact of resistance exercise on patients with chronic kidney disease. BMC Nephrol. 25 , 115. 10.1186/s12882-024-03547-5 (2024). Evans, M. et al. A Narrative Review of Chronic Kidney Disease in Clinical Practice: Current Challenges and Future Perspectives. Adv. Ther. 39 , 33–43. 10.1007/s12325-021-01927-z (2022). Lin, C. C. et al. Development and validation of a risk prediction model for chronic kidney disease among individuals with type 2 diabetes. Sci. Rep. 12 , 4794. 10.1038/s41598-022-08284-z (2022). Colli, V. A. et al. Chronic kidney disease risk prediction scores assessment and development in Mexican adult population. Front. Med. (Lausanne) . 9 , 903090. 10.3389/fmed.2022.903090 (2022). Glassock, R. J. & Rule, A. D. The implications of anatomical and functional changes of the aging kidney: with an emphasis on the glomeruli. Kidney Int. 82 , 270–277. 10.1038/ki.2012.65 (2012). Mallappallil, M., Friedman, E. A., Delano, B. G., McFarlane, S. I. & Salifu, M. O. Chronic kidney disease in the elderly: evaluation and management. Clin. Pract. (Lond) . 11 , 525–535. 10.2217/cpr.14.46 (2014). Kurella, M., Lo, J. C. & Chertow, G. M. Metabolic syndrome and the risk for chronic kidney disease among nondiabetic adults. J. Am. Soc. Nephrol. 16 , 2134–2140. 10.1681/ASN.2005010106 (2005). Sowers, J. R., Epstein, M. & Frohlich, E. D. Diabetes, hypertension, and cardiovascular disease: an update. Hypertension 37 , 1053–1059. 10.1161/01.hyp.37.4.1053 (2001). Kovesdy, C. P., Furth, S. L. & Zoccali, C. World Kidney Day Steering, C. Obesity and Kidney Disease: Hidden Consequences of the Epidemic. Can. J. Kidney Health Dis. 4 , 2054358117698669. 10.1177/2054358117698669 (2017). Wahba, I. M. & Mak, R. H. Obesity and obesity-initiated metabolic syndrome: mechanistic links to chronic kidney disease. Clin. J. Am. Soc. Nephrol. 2 , 550–562. 10.2215/CJN.04071206 (2007). Martins, C. A. et al. Conicity index as an indicator of abdominal obesity in individuals with chronic kidney disease on hemodialysis. PLoS One . 18 , e0284059. 10.1371/journal.pone.0284059 (2023). Valdez, R. A simple model-based index of abdominal adiposity. J. Clin. Epidemiol. 44 , 955–956. 10.1016/0895-4356(91)90059-i (1991). Postorino, M., Marino, C., Tripepi, G., Zoccali, C. & Group, C. W. Abdominal obesity and all-cause and cardiovascular mortality in end-stage renal disease. J. Am. Coll. Cardiol. 53 , 1265–1272. 10.1016/j.jacc.2008.12.040 (2009). McMahon, E. J., Campbell, K. L., Bauer, J. D., Mudge, D. W. & Kelly, J. T. Altered dietary salt intake for people with chronic kidney disease. Cochrane Database Syst. Rev. 6 , CD010070. 10.1002/14651858.CD010070.pub3 (2021). Clark, W. F. et al. Effect of Coaching to Increase Water Intake on Kidney Function Decline in Adults With Chronic Kidney Disease: The CKD WIT Randomized Clinical Trial. JAMA 319 , 1870–1879. 10.1001/jama.2018.4930 (2018). Kelly, J. T., Su, G. & Carrero, J. J. Lifestyle interventions for preventing and ameliorating CKD in primary and secondary care. Curr. Opin. Nephrol. Hypertens. 30 , 538–546. 10.1097/MNH.0000000000000745 (2021). Steyerberg, E. W. et al. Assessing the performance of prediction models: a framework for traditional and novel measures. Epidemiology 21 , 128–138. 10.1097/EDE.0b013e3181c30fb2 (2010). Van Calster, B., McLernon, D. J., Van Smeden, M., Wynants, L. & Steyerberg, E. W. Calibration: the Achilles heel of predictive analytics. BMC Med. 17 , 230 (2019). Vickers, A. J. & Elkin, E. B. Decision curve analysis: a novel method for evaluating prediction models. Med. Decis. Mak. 26 , 565–574. 10.1177/0272989X06295361 (2006). Cook, N. R. Use and misuse of the receiver operating characteristic curve in risk prediction. Circulation 115 , 928–935. 10.1161/CIRCULATIONAHA.106.672402 (2007). Moynihan, R., Doust, J. & Henry, D. Preventing overdiagnosis: how to stop harming the healthy. BMJ 344 , e3502. 10.1136/bmj.e3502 (2012). Komenda, P. et al. Cost-effectiveness of primary screening for CKD: a systematic review. Am. J. Kidney Dis. 63 , 789–797. 10.1053/j.ajkd.2013.12.012 (2014). Kang, S. H. et al. HbA1c Levels Are Associated with Chronic Kidney Disease in a Non-Diabetic Adult Population: A Nationwide Survey (KNHANES 2011–2013). PLoS One . 10 , e0145827. 10.1371/journal.pone.0145827 (2015). Jang, S. Y. et al. Chronic kidney disease and metabolic syndrome in a general Korean population: the Third Korea National Health and Nutrition Examination Survey (KNHANES III) Study. J. Public. Health (Oxf) . 32 , 538–546. 10.1093/pubmed/fdp127 (2010). Son, Y. B. et al. Smoking amplifies the risk of albuminuria in individuals with high sodium intake: the Korea National Health and Nutrition Examination Survey (KNHANES) 2008–2011 and 2014–2018. Kidney Res. Clin. Pract. 44 , 452–460. 10.23876/j.krcp.22.133 (2025). Han, E., Kim, M. K., Im, S. S., Jang, B. K. & Kim, H. S. Non-alcoholic fatty liver disease and sarcopenia is associated with the risk of albuminuria independent of insulin resistance, and obesity. J. Diabetes Complications . 36 , 108253. 10.1016/j.jdiacomp.2022.108253 (2022). Lo, J. A. et al. Impact of water consumption on renal function in the general population: a cross-sectional analysis of KNHANES data (2008–2017). Clin. Exp. Nephrol. 25 , 376–384. 10.1007/s10157-020-01997-3 (2021). Yu, J. H. et al. U-shaped association between sleep duration and urinary albumin excretion in Korean adults: 2011–2014 Korea National Health and Nutrition Examination Survey. PLoS One . 13 , e0192980. 10.1371/journal.pone.0192980 (2018). Weon, B. et al. Association between dyslipidemia and the risk of incident chronic kidney disease affected by genetic susceptibility: Polygenic risk score analysis. PLoS One . 19 , e0299605. 10.1371/journal.pone.0299605 (2024). Lee, Y., Seo, E., Mun, E. & Lee, W. A longitudinal study of working hours and chronic kidney disease in healthy workers: The Kangbuk Samsung Health Study. J. Occup. Health . 63 , e12266. 10.1002/1348-9585.12266 (2021). Lee, D. Y. & Shin, S. Association between Chronic Kidney Disease and Dynapenia in Elderly Koreans. Healthc. (Basel) . 11 10.3390/healthcare11222976 (2023). Guyon, I., Weston, J., Barnhill, S. & Vapnik, V. Gene selection for cancer classification using support vector machines. Mach. Learn. 46 , 389–422 (2002). Saito, T. & Rehmsmeier, M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. PloS one . 10 , e0118432 (2015). Chen, T. & Guestrin, C. in Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining. 785–794 (2016). Lee, W. V. et al. Development of a machine learning model for precision prognosis of rapid kidney function decline in people with diabetes and chronic kidney disease. Diabetes Res. Clin. Pract. 217 , 111897. 10.1016/j.diabres.2024.111897 (2024). Snoek, J., Larochelle, H. & Adams, R. P. Practical bayesian optimization of machine learning algorithms. Advances neural Inform. Process. systems 25 (2012). Goldstein, A., Kapelner, A., Bleich, J. & Pitkin, E. Peeking inside the black box: Visualizing statistical learning with plots of individual conditional expectation. J. Comput. Graphical Stat. 24 , 44–65 (2015). Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Revision Version 1 posted Editorial decision: Revision requested 08 May, 2026 Reviews received at journal 07 May, 2026 Reviews received at journal 15 Apr, 2026 Reviewers agreed at journal 15 Apr, 2026 Reviewers agreed at journal 15 Apr, 2026 Reviewers invited by journal 15 Apr, 2026 Editor invited by journal 08 Apr, 2026 Editor assigned by journal 06 Apr, 2026 Submission checks completed at journal 06 Apr, 2026 First submitted to journal 02 Apr, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9308066","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":626430900,"identity":"89ab7c45-a75d-4d98-aea5-a660f6c11cce","order_by":0,"name":"HyeokJun Yang","email":"","orcid":"","institution":"Jeju National University","correspondingAuthor":false,"prefix":"","firstName":"HyeokJun","middleName":"","lastName":"Yang","suffix":""},{"id":626430901,"identity":"9d7d635d-6707-48e2-83c9-d13f4eb11896","order_by":1,"name":"Young Gyun Seo","email":"","orcid":"","institution":"Jeju National University","correspondingAuthor":false,"prefix":"","firstName":"Young","middleName":"Gyun","lastName":"Seo","suffix":""},{"id":626430902,"identity":"17c30e90-8901-48af-9590-be0941df95b2","order_by":2,"name":"Nayoung Han","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA1UlEQVRIiWNgGAWjYBACAzBZwcbA2AAVkSBOyxmStTC2IYkQ1GLOfvbYY955fHLMM3KPSTDU2DFIzj6AX4tlT166Me82NmPGGXlpEgzHkhmk+RIIOOxAjpk0UEti44wcMwkGtgMMcjwEHGZw/g1Qyxy2eoiWf8RouQGypYEtgRGkhbHtAIM0YS1vzCTnHGMzbOx5Y2yR2JfMI9lD0GFAw9/UHJM3bM8xvPHhm52cxBkCWkCAiYfhGINhA5CVwMBAyFkQwPiDoYZBniilo2AUjIJRMCIBAGbYOKoXbZmjAAAAAElFTkSuQmCC","orcid":"","institution":"Jeju National University","correspondingAuthor":true,"prefix":"","firstName":"Nayoung","middleName":"","lastName":"Han","suffix":""}],"badges":[],"createdAt":"2026-04-03 02:53:36","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9308066/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9308066/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":107627061,"identity":"d5138276-3215-4412-8f1a-3cbde96d118c","added_by":"auto","created_at":"2026-04-23 10:48:54","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":61654,"visible":true,"origin":"","legend":"\u003cp\u003eFlow diagram for the inclusion and exclusion in our study.\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-9308066/v1/8c7aab5b338296601a3f4b51.png"},{"id":107627062,"identity":"dafaf4a1-80b2-4420-a05d-f5ba99a59a0c","added_by":"auto","created_at":"2026-04-23 10:48:54","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":257825,"visible":true,"origin":"","legend":"\u003cp\u003eROC and PR curves for the final XGBoost model.\u003cbr\u003e\nROC, receiver operating characteristic; PR, precision-recall curve; XGBoost, eXtreme Gradient Boosting\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-9308066/v1/d531804607d048ebe486a3ca.png"},{"id":107627063,"identity":"7ff40843-c9e3-4593-8b8c-867708f2b225","added_by":"auto","created_at":"2026-04-23 10:48:54","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":461101,"visible":true,"origin":"","legend":"\u003cp\u003eTreeSHAP feature importance and value summary plots. (A) SHAP global bar plot and (B) SHAP beeswarm plot of the top features.\u003cbr\u003e\nSHAP, SHapley Additive exPlanations\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-9308066/v1/8c4f7d238e57c7751254c1e0.png"},{"id":107705966,"identity":"78c80d83-c26b-42d2-80ff-96e2fe4e0351","added_by":"auto","created_at":"2026-04-24 09:16:46","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":3676057,"visible":true,"origin":"","legend":"\u003cp\u003eICE plots for top 3 features: (A) age, (B) hypertension, and (C) diabetes mellitus.\u003cbr\u003e\nICE, individual conditional expectation\u003c/p\u003e","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-9308066/v1/006f4cf16483546844060170.png"},{"id":107627065,"identity":"d0bbecab-402f-40bb-81f2-858cb635f034","added_by":"auto","created_at":"2026-04-23 10:48:54","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":189918,"visible":true,"origin":"","legend":"\u003cp\u003eDecision curve analysis plot\u003c/p\u003e","description":"","filename":"Figure5.png","url":"https://assets-eu.researchsquare.com/files/rs-9308066/v1/4d70d6e193c32c60465838e6.png"},{"id":107708949,"identity":"a893de71-8ddc-4824-9921-49a3b57ad883","added_by":"auto","created_at":"2026-04-24 09:33:56","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5119030,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9308066/v1/3e7a3eb2-d1cb-4d17-bec7-81dbff188a7c.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"An interpretable machine learning model for predicting chronic kidney disease risk among obese adults: a nationwide population-based study","fulltext":[{"header":"Introduction","content":"\u003cp\u003eObesity is an independent risk factor for development and progression of chronic kidney disease (CKD).\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e A high body mass index (BMI) increases the prevalence of proteinuria by approximately 21%,\u003csup\u003e2\u003c/sup\u003e and an elevated waist-to-hip ratio increased the risk of CKD by 22%.\u003csup\u003e3\u003c/sup\u003e In addition, the increased prevalence of obesity is a major public health concern, owing to its significant association with chronic diseases, such as hypertension (HTN), diabetes mellitus (DM), and dyslipidemia (DL).\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e Metabolic syndrome contributes to a decline in renal function and the development of proteinuria.\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e Measures of central obesity, such as waist circumference (WC) or fat mass, may predict CKD incidence more accurately than BMI.\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e This is particularly relevant because patients with CKD are prone to malnutrition and muscle wasting, which can lead to an underestimation of adiposity when BMI alone is used.\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e,\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e In addition to obesity-related factors, demographic and physiological characteristics including age, sex differences, and muscle strength also influence CKD risk.\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e,\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eCKD often progresses silently, with minimal or no symptoms in its early stages, resulting in delayed diagnosis until significant renal impairment has occurred.\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e Patients with advanced CKD frequently require high-cost treatments, such as dialysis or kidney transplantation, owning to progression to end-stage kidney disease. Therefore, early identification of individuals at high risk for CKD is crucial, as timely intervention may delay disease progression and improve long-term outcomes and quality of life.\u003c/p\u003e \u003cp\u003eSeveral risk prediction models have been developed to facilitate early detection of CKD.\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e,\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e However, previous studies have reported inconsistent findings regarding the demographic and clinical predictors of CKD, depending on the population studied. Moreover, most conventional approaches have relied on linear regression models, which are limited in their ability to capture complex nonlinear relationships and interactions among risk factors. Therefore, in this study, we aimed to identify key predictors of CKD among obese adults using a machine learning-based approach and to develop a robust and interpretable prediction model that accounts for nonlinear effects and feature interactions.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eDemographic characteristics\u003c/h2\u003e \u003cp\u003eAmong the 133,375 participants, 37,240 were included in the final eligible population (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Participants with CKD were significantly older than those without CKD (63.5\u0026thinsp;\u0026plusmn;\u0026thinsp;12.9 vs. 52.2\u0026thinsp;\u0026plusmn;\u0026thinsp;15.2 years, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Although BMI did not significantly differ between the two groups, measures of central obesity, including WC and conicity index, were significantly higher in the CKD group than in the non-CKD group. Comorbid conditions, including HTN, DM, DL, metabolic syndrome, and cardiovascular disease, were significantly more common in the CKD group than in the non-CKD group (all p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Regarding lifestyle factors, current smoking exposure was greater in the CKD group than in the non-CKD group, as reflected by longer smoking duration. Individuals with CKD reported lower daily water intake, while also exhibiting a higher frequency of eating out. In contrast, sodium intake and physical activity levels, including walking time and resistance exercise, were generally higher in the non-CKD group than in the CKD group, although some variables showed modest differences.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDemographic characteristics of study population.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVariables\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCKD\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNon-CKD\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eP-value\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eN, %\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2,966\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e34,274\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAge\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e63.5\u0026thinsp;\u0026plusmn;\u0026thinsp;12.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e52.2\u0026thinsp;\u0026plusmn;\u0026thinsp;15.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMale\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,507 (50.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e16,739 (48.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.039\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBMI (kg/m\u003csup\u003e2\u003c/sup\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e27.1\u0026thinsp;\u0026plusmn;\u0026thinsp;2.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e27.2\u0026thinsp;\u0026plusmn;\u0026thinsp;2.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.123\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWaist circumference (cm)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e93.5\u0026thinsp;\u0026plusmn;\u0026thinsp;7.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e91.5\u0026thinsp;\u0026plusmn;\u0026thinsp;7.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eConicity index\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1.3\u0026thinsp;\u0026plusmn;\u0026thinsp;0.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.26\u0026thinsp;\u0026plusmn;\u0026thinsp;0.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eComorbid diseases\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHypertension\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2,158 (72.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14,763 (43.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDiabetes mellitus\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,245 (42.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6,247 (18.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDyslipidemia\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2,124 (71.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e19,565 (57.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMetabolic syndrome\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2,247 (75.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e18,795 (54.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStroke\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e195 (6.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e673 (2.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMyocardial infarction\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e109 (3.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e302 (0.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAngina pectoris\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e176 (5.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e652 (1.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNAFLD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e892 (30.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e9,429 (27.5)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlcohol intake(g/day)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.8\u0026thinsp;\u0026plusmn;\u0026thinsp;1.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.0\u0026thinsp;\u0026plusmn;\u0026thinsp;1.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSmoking duration (pack per year)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e10.5\u0026thinsp;\u0026plusmn;\u0026thinsp;18.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e7.5\u0026thinsp;\u0026plusmn;\u0026thinsp;14.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSleep duration\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u0026le;\u0026thinsp;5 hours\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e535 (18.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4,899 (14.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u0026gt;\u0026thinsp;5 and \u0026lt;\u0026thinsp;9 hours\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,938 (65.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e25,071 (73.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u0026ge;\u0026thinsp;9 hours\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e380 (12.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3,150 (9.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWorking hours (h/week)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e37.8\u0026thinsp;\u0026plusmn;\u0026thinsp;20.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e42.1\u0026thinsp;\u0026plusmn;\u0026thinsp;18.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWalking time (min/week)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e341.9\u0026thinsp;\u0026plusmn;\u0026thinsp;1,105.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e367.3\u0026thinsp;\u0026plusmn;\u0026thinsp;1,044.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.302\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWeekly strength training days\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.015\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNever\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2,202 (74.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e24,775 (72.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1 day\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e89 (3.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,369 (4.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2 days\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e122 (4.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,617 (4.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3 days\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e119 (4.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,659 (4.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4 days\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e66 (2.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e815 (2.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5 or more days\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e212 (7.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2,382 (6.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAnnual household income (per 10,000/year)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e3,441.8\u0026thinsp;\u0026plusmn;\u0026thinsp;3,473.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4,582.3\u0026thinsp;\u0026plusmn;\u0026thinsp;3,694.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHousehold income quartile\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLow\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,070 (36.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6,638 (19.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLower-middle\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e773 (26.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e8,949 (26.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUpper-middle\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e559 (18.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e9,364 (27.3)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e521 (17.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e9,044 (26.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWater intake (g/day)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1,960.3\u0026thinsp;\u0026plusmn;\u0026thinsp;3,957.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2,272.1\u0026thinsp;\u0026plusmn;\u0026thinsp;5,824.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSodium intake (mg/day)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e3,512.1\u0026thinsp;\u0026plusmn;\u0026thinsp;2,476.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4,055.8\u0026thinsp;\u0026plusmn;\u0026thinsp;2,819.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFrequency of eating out\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTwice a day or more\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e79 (2.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1,912 (5.6)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOnce a day\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e229 (7.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4,348 (12.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u0026ndash;6 times a week\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e329 (11.1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5,211 (15.2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u0026ndash;4 times a week\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e387 (13.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e4,001 (11.7)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u0026ndash;2 times a week\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e620 (20.9)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e6,513 (19.0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u0026ndash;3 times a month\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e635 (21.4)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e5,419 (15.8)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"4\"\u003e\u003csup\u003ea\u003c/sup\u003e mean (IQR); \u003csup\u003eb\u003c/sup\u003e n (%)\u003c/td\u003e\u003c/tr\u003e \u003ctr\u003e\u003ctd colspan=\"4\"\u003eBMI, body mass index; CKD, chronic kidney disease; NAFLD, non-alcohol fatty liver disease\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eFeature selection\u003c/h3\u003e\n\u003cp\u003eRecursive feature elimination with cross-validation (RFECV) process demonstrated a gradual improvement in model performance as the number of selected features increased. The mean area under the precision-recall curve (AUPRC) steadily increased from approximately 0.221 with a smaller subset of features and reached its peak at 17 features: age; sex; BMI; WC; comorbid diseases, including HTN, DM, and DL; concomitant metabolic syndrome; conicity index; daily alcohol intake; smoking duration; daily water intake; daily sodium intake; frequency of eating out; weekly walking time; weekly working hours; and annual total household income. Beyond this point, the addition of further variables did not result in meaningful performance gains and instead showed a slight decline in AUPRC, indicating diminishing returns from additional features.\u003c/p\u003e \u003cp\u003eBased on these results, the optimal feature subset included 17 variables, which provided the best balance between predictive performance and model complexity. This pattern suggests that removing less informative variables improves model generalizability while reducing redundancy among predictors.\u003c/p\u003e\n\u003ch3\u003eModel development and evaluation\u003c/h3\u003e\n\u003cp\u003eBased on the confusion matrix of the model constructed using eXtreme Gradient Boosting (XGBoost), the model showed an accuracy of 84.51%, a precision of 0.244, a recall of 0.449, and an F1-score of 0.316. The final XGBoost model demonstrated acceptable discriminative performance, with an area under the receiver operating characteristic curve (AUROC) of 0.756 and an AUPRC of 0.240 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). The ROC curve indicated stable discrimination across thresholds, whereas the PR curve reflected the expected trade-off between precision and recall under class imbalance.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003ePermutation importance analysis of the final XGBoost model indicated that age, comorbidities, and conicity index were major contributors to CKD prediction (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). Feature importance ranking using Shapley Additive exPlanations (SHAP) analysis showed that age was the most strongly associated variable with CKD, followed by HTN, DM, DL, and conicity index (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Higher age and the presence of metabolic comorbidities were associated with increased CKD risk, whereas measures of central obesity showed stronger contributions than BMI. Lifestyle-related variables, including sodium intake, water intake, and frequency of eating out, demonstrated relatively low feature importance in the SHAP analysis. Although their individual contributions to model prediction were modest and showed heterogeneous patterns, these variables were consistently retained in the final model. Individual conditional expectation (ICE) plots further demonstrated that age had a nonlinear association with CKD risk, with a more pronounced increase at older ages (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). In contrast, HTN and DM showed consistent positive effects across individuals, indicating more uniform contributions to risk.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePermutation feature importance of predictors in the XGBoost model\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFeature\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eImportance\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAge\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0735\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0061\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDiabetes mellitus\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0315\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0054\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHypertension\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0309\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0055\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSex\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0117\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0024\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSmoking duration\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0091\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0024\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDyslipidemia\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0088\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0032\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eConicity index\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0067\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0036\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAverage weekly working hours\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0060\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0022\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlcohol intake\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0055\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0027\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSodium intake\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0040\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0013\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHousehold income quartile\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0036\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0018\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWater intake\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0033\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0019\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMetabolic syndrome\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0032\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0023\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWaist circumference\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0029\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0024\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eWeekly walking time\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0017\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0014\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFrequency of eating out\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.0011\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0013\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBody mass index\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e-0.0003\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0010\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eModel calibration and clinical utility\u003c/h3\u003e\n\u003cp\u003eProbability calibration using Platt scaling substantially improved the agreement between predicted probabilities and observed outcomes. The calibrated model showed a markedly lower Brier score compared with the uncalibrated model (0.067 vs. 0.193), and the calibration curve was more closely aligned with the ideal reference line, indicating improved reliability of probability estimates.\u003c/p\u003e \u003cp\u003eThe optimal classification threshold notably differed between models. Although the uncalibrated model achieved the maximum F1-score at a higher threshold (0.656), the calibrated model achieved a comparable F1-score (0.319) at a substantially lower threshold (0.169). This shift suggests that probability calibration redistributes predicted risks, enabling more sensitive identification of CKD without compromising overall performance.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e presents the decision curve analysis (DCA) for the raw and calibrated XGBoost models. Across a wide range of threshold probabilities (0.0\u0026ndash;1.0), both models showed similar net benefits (NBs) and largely overlapped with each other. Across the full range of threshold probabilities, the NB of both the raw and calibrated models remained close to zero and largely overlapped with the \u0026ldquo;treat none\u0026rdquo; strategy. However, the \u0026ldquo;treat all\u0026rdquo; strategy showed a marked decline in NB at higher threshold probabilities, whereas the model-based strategies remained relatively stable.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this study, we developed a machine learning model to predict CKD risk among obese adults using nationwide population-based data. The model demonstrated acceptable discriminative performance and identified key predictors, including age, metabolic comorbidities, and measures of central obesity. Through the integration of explainable artificial intelligence technology, we were able to gain insights into the relative importance and clinical effects of these predictor variables.\u003c/p\u003e \u003cp\u003eAge was the most influential factor, showing a nonlinear association with CKD risk. The ICE analysis indicated that the predicted probability of CKD increased more significantly at older ages, suggesting an accelerated risk pattern beyond 40 years. This finding is consistent with the results of previous epidemiological studies demonstrating age-related decline in renal function and increased susceptibility to kidney damage.\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e,\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e The nonlinear pattern further highlights the advantage of machine learning approaches in capturing complex risk relationships that may be inadequately modeled using conventional linear methods. The identified comorbidities of this study are not only primary causes of CKD but also can be exacerbated by declining renal function, reflecting a bidirectional and synergistic relationship.\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e,\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e As most obese individuals have these clustered comorbidities, the cumulative risk of CKD significantly increases.\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e Among the factors related to obesity, the conicity index demonstrated a greater influence on CKD prediction than BMI or WC, suggesting that renal impairment is more closely associated with central adiposity than with overall body mass.\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e The conicity index is an indicator that incorporates WC, body weight, and height to determine the extent to which the body shape approximates a double cone,\u003csup\u003e22\u003c/sup\u003e measuring the degree of abdominal obesity.\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e These findings emphasize that managing abdominal fat distribution, rather than focusing solely on weight loss, is a more important clinical strategy for delaying the onset and progression of CKD.\u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eUnlike non-modifiable factors, such as age, or partially modifiable conditions, such as established comorbidities, lifestyle-related variables, such as sodium intake, water intake, and frequency of eating out, showed relatively low but non-negligible contributions to CKD prediction. Variables related to physical activity and working conditions, including weekly working time and walking time, demonstrated minimal and non-directional contributions in the SHAP analysis. This may reflect the complex and multifactorial nature of lifestyle influences and potential measurement variability inherent in self-reported data. Nevertheless, the inclusion of these factors in the final model suggests that lifestyle factors represent modifiable risk components that can be targeted through behavioral interventions.\u003csup\u003e\u003cspan additionalcitationids=\"CR26\" citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eThe final model demonstrated moderate discriminative performance. Although Platt calibration improved probability reliability, it did not lead to a meaningful change in decision-level NB; however, discrimination metrics, such as AUROC, remained largely unchanged. This finding is expected, as probability calibration improves the accuracy of predicted probabilities rather than the model\u0026rsquo;s ability to distinguish between classes.\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e Notably, the Brier score substantially decreased from 0.193 in the uncalibrated model, indicating poor reliability of predicted probabilities, to 0.067 after calibration, reflecting a marked improvement in probabilistic accuracy. Given that CKD screening relies on accurate estimation of individual risk rather than binary classification alone, this improvement in calibration is clinically meaningful and may enhance the applicability of the model in risk-based decision-making.\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eDespite these improvements, DCA demonstrated that the overall NB of the model remained modest and close to that of the \u0026ldquo;treat-none\u0026rdquo; strategy across most threshold probabilities.\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e Several factors may explain this observation. First, CKD in population-based cohorts is characterized by relatively low short-term event rates, which limits the absolute number of true positives and constrains achievable NB.\u003csup\u003e31\u003c/sup\u003e Second, the threshold for intervention in CKD is typically low, favoring early detection strategies, thereby reducing the relative advantage of model-based decision-making compared with default approaches. Third, many of the key predictors identified in this study, such as age and metabolic comorbidities, are already incorporated into routine clinical judgment, limiting the incremental value of the model. Nevertheless, the model consistently outperformed the \u0026ldquo;treat-all\u0026rdquo; strategy at moderate to high threshold probabilities, suggesting a potential role in reducing unnecessary interventions.\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e In clinical practice, indiscriminate application of screening or treatment may lead to increased healthcare burden and potential harm. In this context, the model may be useful for selectively identifying individuals who are less likely to benefit from intervention, thereby improving resource allocation.\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eThis study has several strengths. It used a large, nationally representative dataset, enhancing generalizability within the target population. The use of a rigorous feature selection framework and nested cross-validation minimized overfitting and information leakage. Furthermore, the incorporation of SHAP and ICE analyses allowed for both global and individual-level interpretation, improving the transparency and clinical interpretability of the model. However, this study has some limitations. First, owing to the cross-sectional design, causal relationships cannot be established, and temporal prediction of CKD development is limited. In addition, the use of questionnaire-based variables may introduce measurement error and recall bias. Second, the proportion of CKD-positive cases was relatively low (approximately 8%), resulting in substantial class imbalance. In such settings, a model may achieve high overall accuracy by predominantly predicting the negative class, whereas its ability to correctly identify positive cases (e.g., precision and recall) remains limited. Because the AUPRC directly reflects class imbalance, the relatively low AUPRC observed in this study should be interpreted in this context. Third, the model was developed using data from a single country, and external validation in independent populations is required to confirm its generalizability. Finally, the modest NB observed in DCA suggests that further improvements in clinical utility may be achieved by incorporating longitudinal data or additional biomarkers in future studies.\u003c/p\u003e \u003cp\u003eIn conclusion, this study developed an interpretable machine learning model for CKD risk prediction among obese adults, highlighting the dominant roles of age, metabolic comorbidities, and central obesity. Although the model demonstrated acceptable performance and provided meaningful insights into risk factors, its clinical utility for guiding intervention remains modest. Future studie should focus on improving risk stratification through longitudinal modeling and integration of additional clinical and biological data.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eMethods\u003c/h2\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003eStudy design and data collection\u003c/h2\u003e \u003cp\u003eThis study was a cross-sectional analysis using publicly available data from the Korea National Health and Nutrition Examination Survey (KNHANES) conducted between 2007 and 2023 (n\u0026thinsp;=\u0026thinsp;133,375). As KNHANES data are publicly available and de-identified, institutional review board approval and informed consent were waived.\u003c/p\u003e \u003cp\u003ePrior to analysis, inclusion and exclusion criteria were predefined: Adults aged 19\u0026ndash;79 years were included, resulting in a final eligible population of 100,835 participants after applying the age criterion. Participants with missing values that precluded assessment of obesity status or CKD classification were excluded from the analysis. Obesity-related variables were defined and used as analytical features rather than inclusion criteria. Obesity was defined as a BMI\u0026thinsp;\u0026ge;\u0026thinsp;25 kg/m\u0026sup2; and/or abdominal obesity, defined as WC\u0026thinsp;\u0026ge;\u0026thinsp;90 cm for men and \u0026ge;\u0026thinsp;85 cm for women, according to Korean guidelines. CKD was defined based on the criteria of the Korean Society of Nephrology as any of the following: estimated glomerular filtration rate\u0026thinsp;\u0026lt;\u0026thinsp;60 mL/min/1.73 m\u0026sup2; calculated using the CKD-Epidemiology (CKD-EPI) or Modification of Diet in Renal Disease (MDRD) equation, presence of proteinuria on urinalysis, or self-reported physician-diagnosed CKD. The final analytic dataset was subsequently used for feature selection, model development, and evaluation. Complete-case analysis was applied for all subsequent modeling procedures.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e\n\u003ch3\u003eData preprocessing\u003c/h3\u003e\n\u003cp\u003eBased on literature review, we selected 24 factors that had been significantly associated with CKD incidence in existing literature and could be directly identified or calculated from the KNHANES raw data: age; sex; comorbidities, such as DM, HTN, DL, metabolic syndrome, non-alcoholic fatty liver disease, myocardial infarction, and angina; history of cerebrovascular disease; BMI; WC; conicity index; alcohol intake; smoking duration; sleep time; water intake; sodium intake; frequency of eating out per week; weekly walking time; average weekly working hours; frequency of resistance training per week; and household income quartile (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e).\u003csup\u003e\u003cspan additionalcitationids=\"CR35 CR36 CR37 CR38 CR39 CR40 CR41\" citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eTable 3. Definition of variables used in this study\u003c/p\u003e\n\u003ctable border=\"0\" cellspacing=\"0\" cellpadding=\"0\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eVariables\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003e\u003cstrong\u003eDefinition\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eObesity\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 468px;\"\u003e\n \u003cp\u003eBMI \u0026ge; 25(kg/m\u003csup\u003e2\u003c/sup\u003e) or WC \u0026ge; 90cm (if male), \u0026ge; 85 (if female)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eBody mass index\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 468px;\"\u003e\n \u003cp\u003eweight(kg) / height(m)\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eWaist circumference\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd valign=\"top\" style=\"width: 468px;\"\u003e\n \u003cp\u003e\u0026ge; 90 cm (if male), \u0026ge; 85 cm (if female)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eConicity index\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003e0.109\u003csup\u003e\u0026minus;1\u003c/sup\u003e x WC(m) x [weight(kg)/height(m)]\u003csup\u003e\u0026minus;1/2\u003c/sup\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eHypertension\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003ediagnosed with hypertension, taking antihypertensive medicines, SBP \u0026ge;140 mmHg or DBP \u0026ge;90 mmHg\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eDiabetes mellitus\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003ediagnosed with diabetes mellitus, taking antidiabetic medicines, random glucose \u0026ge;126 mg/dL, or HbA1c \u0026ge;6.5%\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eDyslipidemia\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003ediagnosed with dyslipidemia, taking lipid-lowering medications, total cholesterol \u0026ge;240 mg/dL, LDL cholesterol \u0026ge;160 mg/dL, HDL cholesterol \u0026lt;40 mg/dL, or triglycerides \u0026ge;200 mg/dL\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eNon-alcoholic fatty liver disease\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003eat least one of the three criteria applies:\u003c/p\u003e\n \u003cul start=\"50\"\u003e\n \u003cli\u003edaily alcohol intake criteria (WHO standard; 10 g per drink)\u003cul\u003e\n \u003cli\u003emen: \u0026lt; 30 g/day\u003c/li\u003e\n \u003cli\u003ewomen: \u0026lt; 20 g/day\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/li\u003e\n \u003cli\u003eHepatic Steatosis Index (HSI) cutoff: \u0026ge; 36\u003cul\u003e\n \u003cli\u003eHSI = 8 \u0026times; (ALT/AST ratio) + BMI + 2 (if diabetes) + 2 (if female)\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/li\u003e\n \u003cli\u003eFramingham Steatosis Index (FSI) cutoff: \u0026ge; 23\u003cul\u003e\n \u003cli\u003eFSI = -7.981 + 0.011 \u0026times; age (years) + 0.146 \u0026times; sex (female = 1, male = 0)\u003cbr\u003e\u0026nbsp;+ 0.173 \u0026times; BMI (kg/m\u0026sup2;) + 0.007 \u0026times; triglycerides (mg/dL)\u003cbr\u003e\u0026nbsp;+ 0.593 \u0026times; hypertension (yes = 1, no = 0) + 0.789 \u0026times; diabetes (yes = 1, no = 0)\u003cbr\u003e\u0026nbsp;+ 1.1 \u0026times; I(ALT:AST ratio \u0026ge; 1.33) (yes = 1, no = 0)\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eMetabolic syndrome\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003ehaving at least 3 out of 5 health conditions:\u003c/p\u003e\n \u003cul start=\"50\"\u003e\n \u003cli\u003eabdominal obesity: waist circumference for men \u0026ge;90 cm; for women \u0026ge;80 cm\u003c/li\u003e\n \u003cli\u003ehigh triglycerides: blood triglyceride levels \u0026ge;150 mg/dL\u003c/li\u003e\n \u003cli\u003elow HDL cholesterol: blood HDL levels for men \u0026lt;40 mg/dL; for women \u0026lt;50 mg/dL\u003c/li\u003e\n \u003cli\u003ehigh blood pressure: SBP/DBP \u0026ge;130/85 mmHg\u003c/li\u003e\n \u003cli\u003ehigh fasting blood sugar: blood glucose levels \u0026ge;100 mg/dL\u003c/li\u003e\n \u003c/ul\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eHistory of stroke\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003ediagnosed with stroke or currently under treatment\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eHistory of ischemic heart disease\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003ediagnosed with angina pectoris or myocardial infarction\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eAlcohol intake\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003eaverage daily alcohol intake (g)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eSmoking duration\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003e(number of packs smoked per day) \u0026times; (years of smoking)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eSleep duration group\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003eless or equal to 5 hours = 1; more than 5 and less than 9 hours = 2; more or equal to 9 hours = 3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eWater intake\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003ewater intake from foods (g) + drinking water intake (cups) \u0026times; 200 mL\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eSodium intake\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003eaverage daily sodium intake (mg)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eFrequency of eating out\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003etwice a day or more = 1; once a day = 2; 5-6 times a week = 3; 3-4 times a week = 4; 1-2 times a week = 5; 1-3 times a month = 6\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eWalking time\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003edaily total walking time \u0026times; number of walking days per week\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eResistance training time\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003enone = 0; 1 or 2 days per week = 1; 3 or 4 days per week = 2; 5 days or more per week = 3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd style=\"width: 132px;\"\u003e\n \u003cp\u003eTotal household income quartile\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd style=\"width: 468px;\"\u003e\n \u003cp\u003elow = 1; low-middle = 2; upper-middle = 3; high = 4\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eALT, alanine aminotransferase; AST, aspartate transaminase; BMI, body mass index; DBP, diastolic blood pressure; HDL, high density lipoprotein; LDL, low density lipoprotein; SBP, systolic blood pressure; WC, waist circumference; WHO, World Health Organization\u003c/p\u003e \u003cp\u003eTo address missing values in all variables included in the analysis, a data imputation pipeline was constructed using Python (version 3.11.7). In the raw KNHANES dataset, variables were then categorized into continuous and categorical types, and missing values were handled separately for each group. Continuous variables were normalized using robust scaling to reduce the influence of outliers. Missing values in continuous variables were subsequently imputed using k-nearest neighbors (KNN) imputation with the number of neighbors set to \u003cem\u003ek\u003c/em\u003e\u0026thinsp;=\u0026thinsp;5, based on proximity in the scaled feature space. For categorical variables, missing values were imputed using a KNN-based approach adapted for categorical data, following appropriate numerical encoding of categorical levels. After imputation, the separately processed continuous and categorical datasets were merged to restore the original variable order for subsequent analyses.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eFeature selection\u003c/h2\u003e \u003cp\u003eTo identify the most informative predictors for CKD classification while minimizing overfitting and information leakage, we performed RFECV within a nested cross-validation framework.\u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e A Decision Tree classifier was used as the internal estimator for recursive feature elimination (RFE), as it can capture nonlinear relationships and interactions between clinical variables. At each elimination step, feature importance was computed using the impurity-based importance scores derived from the Decision Tree model, and the feature with the lowest importance was iteratively removed. Hyperparameter optimization of the internal estimator was conducted using Bayesian optimization with Optuna. At each feature elimination step, 40 Optuna trials were performed under a stratified 5-fold cross-validation framework to identify the hyperparameter combination that maximized the mean AUPRC. Class distribution was preserved across all folds.\u003c/p\u003e \u003cp\u003eGiven the class imbalance inherent in CKD prediction tasks, AUPRC was selected as the primary optimization metric, as it provides a more informative assessment of model performance in imbalanced datasets. The AUROC was additionally reported as a complementary metric.\u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e Feature selection performance was evaluated using the cross-validated mean AUPRC derived from out-of-fold (OOF) predicted probabilities. The optimal number of features was defined as the subset yielding the highest mean AUPRC across folds.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eModel training and calibration using XGBoost\u003c/h2\u003e \u003cp\u003eModel training was performed using the XGBoost algorithm. To mitigate overfitting, both L1 (α) and L2 (λ) regularization terms on leaf weights were incorporated during model training.\u003csup\u003e\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e Hyperparameter optimization was conducted using a Bayesian optimization framework. The Expected Improvement (EI) acquisition function was used with an exploration\u0026ndash;exploitation trade-off parameter set to ξ\u0026thinsp;=\u0026thinsp;0.01. The optimization process consisted of three initial random evaluations followed by 50 iterative optimization steps.\u003c/p\u003e \u003cp\u003eModel performance at each evaluation point was assessed using stratified 5-fold cross-validation, with AUPRC selected as the primary optimization metric to appropriately account for class imbalance.\u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e A Gaussian Process Regressor was used as the surrogate model, combining a Mat\u0026eacute;rn kernel with a WhiteKernel to balance smoothness assumptions and noise tolerance within the hyperparameter search space.\u003csup\u003e\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u003c/sup\u003e The hyperparameter configuration yielding the highest cross-validated AUPRC was selected as the optimal set.\u003c/p\u003e \u003cp\u003eUsing the optimized hyperparameters, the final XGBoost model was retrained on the selected feature set. To improve the reliability of predicted probabilities, Platt scaling (sigmoid calibration) was applied. Calibration parameters were estimated via internal stratified 5-fold CV using the training data only. Model calibration performance before and after calibration was evaluated using the Brier score and calibration curves (reliability diagrams).\u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eModel evaluation: Threshold determination and performance evaluation\u003c/h2\u003e \u003cp\u003eThe optimal classification threshold was determined by maximizing the F1-score, following a commonly adopted approach for imbalanced binary classification tasks. Specifically, stratified 5-fold CV was performed on the training dataset, and OOF predicted probabilities were aggregated to construct a precision-recall curve. The probability threshold corresponding to the maximum F1-score was selected as the final cutoff and subsequently fixed for evaluation on the independent test set.\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:F1\\:score=\\frac{2\\times\\:Precision\\times\\:Recall}{Precision+Recall}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eFinal model performance was assessed on the test set by computing accuracy, precision, recall, F1-score, confusion matrix, AUROC, and AUPRC. AUROC was used to evaluate overall discriminative ability, whereas AUPRC was used to assess performance for the positive class.\u003c/p\u003e \u003cp\u003eClinical utility was further evaluated using DCA, which quantifies NB across a range of threshold probabilities.\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e Net benefit was calculated as follows:\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:NB\\left(pt\\right)=\\frac{TP}{N}-\\frac{FP}{N}\\times\\:\\frac{pt}{1-pt}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere pt denotes the threshold probability. Model performance was compared against treat-all and treat-none strategies to assess potential clinical value.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eModel explanation\u003c/h2\u003e \u003cp\u003eTo enhance model interpretability, two complementary approaches were used. First, permutation importance was calculated on the test set by randomly shuffling each feature and measuring the resulting decrease in AUPRC, thereby quantifying each feature\u0026rsquo;s contribution to predictive performance under unseen data conditions. Second, to enhance the interpretability of the final CKD prediction model, SHAP analysis was performed using TreeExplainer to leverage the tree-based structure of XGBoost. Feature importance was quantified using the mean absolute SHAP value across samples, and the most influential predictors were visualized using bar plots. In addition, SHAP summary (beeswarm) plots were generated to examine the directionality and distribution of each feature\u0026rsquo;s contribution to model predictions. Finally, ICE plots were generated to further investigate nonlinear feature effects.\u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e ICE curves were constructed for the top three features identified by permutation importance, using both training and test sets, to examine individual-level prediction patterns as feature values varied.\u003c/p\u003e \u003cp\u003eProbability calibration was performed using Platt scaling (sigmoid calibration). Specifically, the model\u0026rsquo;s raw output scores were mapped to calibrated probabilities via a logistic function \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:P(y=1\\mid\\:s)=1/(1+\\text{e}\\text{x}\\text{p}(As+B\\left)\\right)\\)\u003c/span\u003e\u003c/span\u003e, where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:A\\)\u003c/span\u003e\u003c/span\u003eand \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:B\\)\u003c/span\u003e\u003c/span\u003ewere estimated using cross-validated predictions within the training data to prevent information leakage. Calibration performance was assessed using the Brier score and calibration curves (reliability diagrams).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eStatistical analyses\u003c/h2\u003e \u003cp\u003eContinuous variables are summarized as means\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviations for normally distributed data or median with interquartile range (IQR) for non-normally distributed data. Normality was assessed using the Shapiro\u0026ndash;Wilk test. Between-group comparisons were performed using the independent \u003cem\u003et\u003c/em\u003e-test for normally distributed variables and the Mann-Whitney U test for non-normally distributed variables. Categorical variables are presented as frequencies (n) and percentages (%), and group differences were assessed using the chi-squared (χ\u0026sup2;) test.\u003c/p\u003e \u003cp\u003eAll statistical tests were two-tailed, and a p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05 was considered statistically significant. Conventional statistical analyses were performed to compare baseline characteristics, whereas machine learning-based modeling and evaluation were performed separately as described above.\u003c/p\u003e \u003cp\u003eAll model development and analyses were performed using Python (version 3.11, 64-bit). Key libraries included pandas, numpy, matplotlib, scikit-learn, xgboost, shap, and scipy. To ensure reproducibility, random seeds were fixed (random_state\u0026thinsp;=\u0026thinsp;42) throughout the analysis. Data preprocessing, model training, hyperparameter optimization, evaluation, and visualization were performed as an end-to-end automated pipeline, allowing the entire analysis to be reproducibly executed by specifying only the data paths and variable definitions.\u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eCompeting interests\u003c/h2\u003e \u003cp\u003eThe authors declare no competing interests.\u003c/p\u003e \u003c/p\u003e\u003cp\u003e \u003ch2\u003eEthics declarations\u003c/h2\u003e \u003cp\u003e This study was conducted in accordance with the Declaration of Helsinki and was approved by the Institutional Review Board of Jeju National University (IRB No. 2026-###-###). The data used in this study were obtained from the Korea National Health and Nutrition Examination Survey (KNHANES), which is publicly available. The requirement for informed consent was waived due to the use of anonymized secondary data.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eDeclaration\u003c/p\u003e \u003cp\u003eThis work was supported by the research grant of Jeju National University in 2023.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eN.H. conceived and designed the study. H.J.Y. and Y.G.S. performed data curation, formal analysis, and methodology development. H.J.Y. and developed the machine learning model and conducted statistical analyses. Y.G.S. drafted the manuscript and H.J.Y. contributed to data interpretation and critically revised the manuscript. All authors reviewed and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets used and/or analyzed during the current study are available from the corresponding author upon reasonable request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eChang, A. R. et al. Adiposity and risk of decline in glomerular filtration rate: meta-analysis of individual participant data in a global consortium. \u003cem\u003eBMJ\u003c/em\u003e \u003cb\u003e364\u003c/b\u003e, k5301. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/bmj.k5301\u003c/span\u003e\u003cspan address=\"10.1136/bmj.k5301\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRosenstock, J. L., Pommier, M., Stoffels, G., Patel, S. \u0026amp; Michelis, M. F. Prevalence of Proteinuria and Albuminuria in an Obese Population and Associated Risk Factors. \u003cem\u003eFront. Med. (Lausanne)\u003c/em\u003e. \u003cb\u003e5\u003c/b\u003e, 122. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fmed.2018.00122\u003c/span\u003e\u003cspan address=\"10.3389/fmed.2018.00122\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eElsayed, E. F. et al. Waist-to-hip ratio, body mass index, and subsequent kidney disease and death. \u003cem\u003eAm. J. Kidney Dis.\u003c/em\u003e \u003cb\u003e52\u003c/b\u003e, 29\u0026ndash;38. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1053/j.ajkd.2008.02.363\u003c/span\u003e\u003cspan address=\"10.1053/j.ajkd.2008.02.363\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJafari-Adli, S. et al. Prevalence of obesity and overweight in adults and children in Iran; a systematic review. \u003cem\u003eJ. Diabetes Metab. Disord\u003c/em\u003e. \u003cb\u003e13\u003c/b\u003e, 121. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s40200-014-0121-2\u003c/span\u003e\u003cspan address=\"10.1186/s40200-014-0121-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDespres, J. P. et al. Abdominal obesity and the metabolic syndrome: contribution to global cardiometabolic risk. \u003cem\u003eArterioscler. Thromb. Vasc Biol.\u003c/em\u003e \u003cb\u003e28\u003c/b\u003e, 1039\u0026ndash;1049. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1161/ATVBAHA.107.159228\u003c/span\u003e\u003cspan address=\"10.1161/ATVBAHA.107.159228\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThomas, G. et al. Metabolic syndrome and kidney disease: a systematic review and meta-analysis. \u003cem\u003eClin. J. Am. Soc. Nephrol.\u003c/em\u003e \u003cb\u003e6\u003c/b\u003e, 2364\u0026ndash;2373. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2215/CJN.02180311\u003c/span\u003e\u003cspan address=\"10.2215/CJN.02180311\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2011).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJaroszynski, A. et al. Association of anthropometric measures of obesity and chronic kidney disease in elderly women. \u003cem\u003eAnn. Agric. Environ. Med.\u003c/em\u003e \u003cb\u003e23\u003c/b\u003e, 636\u0026ndash;640. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.5604/12321966.1226859\u003c/span\u003e\u003cspan address=\"10.5604/12321966.1226859\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEvangelista, L. S., Cho, W. K. \u0026amp; Kim, Y. Obesity and chronic kidney disease: A population-based study among South Koreans. \u003cem\u003ePLoS One\u003c/em\u003e. \u003cb\u003e13\u003c/b\u003e, e0193559. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0193559\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0193559\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReis, J. P. et al. Comparison of overall obesity and body fat distribution in predicting risk of mortality. \u003cem\u003eObes. (Silver Spring)\u003c/em\u003e. \u003cb\u003e17\u003c/b\u003e, 1232\u0026ndash;1239. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/oby.2008.664\u003c/span\u003e\u003cspan address=\"10.1038/oby.2008.664\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2009).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim, D. W. et al. Reproducibility and validity of an FFQ developed for the Korea National Health and Nutrition Examination Survey (KNHANES). \u003cem\u003ePublic. Health Nutr.\u003c/em\u003e \u003cb\u003e18\u003c/b\u003e, 1369\u0026ndash;1377. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1017/S1368980014001712\u003c/span\u003e\u003cspan address=\"10.1017/S1368980014001712\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCarrero, J. J. Gender differences in chronic kidney disease: underpinnings and therapeutic implications. \u003cem\u003eKidney Blood Press. Res.\u003c/em\u003e \u003cb\u003e33\u003c/b\u003e, 383\u0026ndash;392. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1159/000320389\u003c/span\u003e\u003cspan address=\"10.1159/000320389\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen, C. C. et al. Impact of resistance exercise on patients with chronic kidney disease. \u003cem\u003eBMC Nephrol.\u003c/em\u003e \u003cb\u003e25\u003c/b\u003e, 115. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12882-024-03547-5\u003c/span\u003e\u003cspan address=\"10.1186/s12882-024-03547-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEvans, M. et al. A Narrative Review of Chronic Kidney Disease in Clinical Practice: Current Challenges and Future Perspectives. \u003cem\u003eAdv. Ther.\u003c/em\u003e \u003cb\u003e39\u003c/b\u003e, 33\u0026ndash;43. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s12325-021-01927-z\u003c/span\u003e\u003cspan address=\"10.1007/s12325-021-01927-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin, C. C. et al. Development and validation of a risk prediction model for chronic kidney disease among individuals with type 2 diabetes. \u003cem\u003eSci. Rep.\u003c/em\u003e \u003cb\u003e12\u003c/b\u003e, 4794. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41598-022-08284-z\u003c/span\u003e\u003cspan address=\"10.1038/s41598-022-08284-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eColli, V. A. et al. Chronic kidney disease risk prediction scores assessment and development in Mexican adult population. \u003cem\u003eFront. Med. (Lausanne)\u003c/em\u003e. \u003cb\u003e9\u003c/b\u003e, 903090. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fmed.2022.903090\u003c/span\u003e\u003cspan address=\"10.3389/fmed.2022.903090\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGlassock, R. J. \u0026amp; Rule, A. D. The implications of anatomical and functional changes of the aging kidney: with an emphasis on the glomeruli. \u003cem\u003eKidney Int.\u003c/em\u003e \u003cb\u003e82\u003c/b\u003e, 270\u0026ndash;277. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/ki.2012.65\u003c/span\u003e\u003cspan address=\"10.1038/ki.2012.65\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMallappallil, M., Friedman, E. A., Delano, B. G., McFarlane, S. I. \u0026amp; Salifu, M. O. Chronic kidney disease in the elderly: evaluation and management. \u003cem\u003eClin. Pract. (Lond)\u003c/em\u003e. \u003cb\u003e11\u003c/b\u003e, 525\u0026ndash;535. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2217/cpr.14.46\u003c/span\u003e\u003cspan address=\"10.2217/cpr.14.46\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKurella, M., Lo, J. C. \u0026amp; Chertow, G. M. Metabolic syndrome and the risk for chronic kidney disease among nondiabetic adults. \u003cem\u003eJ. Am. Soc. Nephrol.\u003c/em\u003e \u003cb\u003e16\u003c/b\u003e, 2134\u0026ndash;2140. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1681/ASN.2005010106\u003c/span\u003e\u003cspan address=\"10.1681/ASN.2005010106\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2005).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSowers, J. R., Epstein, M. \u0026amp; Frohlich, E. D. Diabetes, hypertension, and cardiovascular disease: an update. \u003cem\u003eHypertension\u003c/em\u003e \u003cb\u003e37\u003c/b\u003e, 1053\u0026ndash;1059. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1161/01.hyp.37.4.1053\u003c/span\u003e\u003cspan address=\"10.1161/01.hyp.37.4.1053\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2001).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKovesdy, C. P., Furth, S. L. \u0026amp; Zoccali, C. World Kidney Day Steering, C. Obesity and Kidney Disease: Hidden Consequences of the Epidemic. \u003cem\u003eCan. J. Kidney Health Dis.\u003c/em\u003e \u003cb\u003e4\u003c/b\u003e, 2054358117698669. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1177/2054358117698669\u003c/span\u003e\u003cspan address=\"10.1177/2054358117698669\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2017).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWahba, I. M. \u0026amp; Mak, R. H. Obesity and obesity-initiated metabolic syndrome: mechanistic links to chronic kidney disease. \u003cem\u003eClin. J. Am. Soc. Nephrol.\u003c/em\u003e \u003cb\u003e2\u003c/b\u003e, 550\u0026ndash;562. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.2215/CJN.04071206\u003c/span\u003e\u003cspan address=\"10.2215/CJN.04071206\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2007).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMartins, C. A. et al. Conicity index as an indicator of abdominal obesity in individuals with chronic kidney disease on hemodialysis. \u003cem\u003ePLoS One\u003c/em\u003e. \u003cb\u003e18\u003c/b\u003e, e0284059. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0284059\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0284059\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eValdez, R. A simple model-based index of abdominal adiposity. \u003cem\u003eJ. Clin. Epidemiol.\u003c/em\u003e \u003cb\u003e44\u003c/b\u003e, 955\u0026ndash;956. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/0895-4356(91)90059-i\u003c/span\u003e\u003cspan address=\"10.1016/0895-4356(91)90059-i\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (1991).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePostorino, M., Marino, C., Tripepi, G., Zoccali, C. \u0026amp; Group, C. W. Abdominal obesity and all-cause and cardiovascular mortality in end-stage renal disease. \u003cem\u003eJ. Am. Coll. Cardiol.\u003c/em\u003e \u003cb\u003e53\u003c/b\u003e, 1265\u0026ndash;1272. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jacc.2008.12.040\u003c/span\u003e\u003cspan address=\"10.1016/j.jacc.2008.12.040\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2009).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcMahon, E. J., Campbell, K. L., Bauer, J. D., Mudge, D. W. \u0026amp; Kelly, J. T. Altered dietary salt intake for people with chronic kidney disease. \u003cem\u003eCochrane Database Syst. Rev.\u003c/em\u003e \u003cb\u003e6\u003c/b\u003e, CD010070. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1002/14651858.CD010070.pub3\u003c/span\u003e\u003cspan address=\"10.1002/14651858.CD010070.pub3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eClark, W. F. et al. Effect of Coaching to Increase Water Intake on Kidney Function Decline in Adults With Chronic Kidney Disease: The CKD WIT Randomized Clinical Trial. \u003cem\u003eJAMA\u003c/em\u003e \u003cb\u003e319\u003c/b\u003e, 1870\u0026ndash;1879. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1001/jama.2018.4930\u003c/span\u003e\u003cspan address=\"10.1001/jama.2018.4930\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKelly, J. T., Su, G. \u0026amp; Carrero, J. J. Lifestyle interventions for preventing and ameliorating CKD in primary and secondary care. \u003cem\u003eCurr. Opin. Nephrol. Hypertens.\u003c/em\u003e \u003cb\u003e30\u003c/b\u003e, 538\u0026ndash;546. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/MNH.0000000000000745\u003c/span\u003e\u003cspan address=\"10.1097/MNH.0000000000000745\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSteyerberg, E. W. et al. Assessing the performance of prediction models: a framework for traditional and novel measures. \u003cem\u003eEpidemiology\u003c/em\u003e \u003cb\u003e21\u003c/b\u003e, 128\u0026ndash;138. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1097/EDE.0b013e3181c30fb2\u003c/span\u003e\u003cspan address=\"10.1097/EDE.0b013e3181c30fb2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVan Calster, B., McLernon, D. J., Van Smeden, M., Wynants, L. \u0026amp; Steyerberg, E. W. Calibration: the Achilles heel of predictive analytics. \u003cem\u003eBMC Med.\u003c/em\u003e \u003cb\u003e17\u003c/b\u003e, 230 (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVickers, A. J. \u0026amp; Elkin, E. B. Decision curve analysis: a novel method for evaluating prediction models. \u003cem\u003eMed. Decis. Mak.\u003c/em\u003e \u003cb\u003e26\u003c/b\u003e, 565\u0026ndash;574. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1177/0272989X06295361\u003c/span\u003e\u003cspan address=\"10.1177/0272989X06295361\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2006).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCook, N. R. Use and misuse of the receiver operating characteristic curve in risk prediction. \u003cem\u003eCirculation\u003c/em\u003e \u003cb\u003e115\u003c/b\u003e, 928\u0026ndash;935. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1161/CIRCULATIONAHA.106.672402\u003c/span\u003e\u003cspan address=\"10.1161/CIRCULATIONAHA.106.672402\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2007).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMoynihan, R., Doust, J. \u0026amp; Henry, D. Preventing overdiagnosis: how to stop harming the healthy. \u003cem\u003eBMJ\u003c/em\u003e \u003cb\u003e344\u003c/b\u003e, e3502. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1136/bmj.e3502\u003c/span\u003e\u003cspan address=\"10.1136/bmj.e3502\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKomenda, P. et al. Cost-effectiveness of primary screening for CKD: a systematic review. \u003cem\u003eAm. J. Kidney Dis.\u003c/em\u003e \u003cb\u003e63\u003c/b\u003e, 789\u0026ndash;797. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1053/j.ajkd.2013.12.012\u003c/span\u003e\u003cspan address=\"10.1053/j.ajkd.2013.12.012\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2014).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKang, S. H. et al. HbA1c Levels Are Associated with Chronic Kidney Disease in a Non-Diabetic Adult Population: A Nationwide Survey (KNHANES 2011\u0026ndash;2013). \u003cem\u003ePLoS One\u003c/em\u003e. \u003cb\u003e10\u003c/b\u003e, e0145827. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0145827\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0145827\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJang, S. Y. et al. Chronic kidney disease and metabolic syndrome in a general Korean population: the Third Korea National Health and Nutrition Examination Survey (KNHANES III) Study. \u003cem\u003eJ. Public. Health (Oxf)\u003c/em\u003e. \u003cb\u003e32\u003c/b\u003e, 538\u0026ndash;546. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/pubmed/fdp127\u003c/span\u003e\u003cspan address=\"10.1093/pubmed/fdp127\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSon, Y. B. et al. Smoking amplifies the risk of albuminuria in individuals with high sodium intake: the Korea National Health and Nutrition Examination Survey (KNHANES) 2008\u0026ndash;2011 and 2014\u0026ndash;2018. \u003cem\u003eKidney Res. Clin. Pract.\u003c/em\u003e \u003cb\u003e44\u003c/b\u003e, 452\u0026ndash;460. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.23876/j.krcp.22.133\u003c/span\u003e\u003cspan address=\"10.23876/j.krcp.22.133\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHan, E., Kim, M. K., Im, S. S., Jang, B. K. \u0026amp; Kim, H. S. Non-alcoholic fatty liver disease and sarcopenia is associated with the risk of albuminuria independent of insulin resistance, and obesity. \u003cem\u003eJ. Diabetes Complications\u003c/em\u003e. \u003cb\u003e36\u003c/b\u003e, 108253. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.jdiacomp.2022.108253\u003c/span\u003e\u003cspan address=\"10.1016/j.jdiacomp.2022.108253\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2022).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLo, J. A. et al. Impact of water consumption on renal function in the general population: a cross-sectional analysis of KNHANES data (2008\u0026ndash;2017). \u003cem\u003eClin. Exp. Nephrol.\u003c/em\u003e \u003cb\u003e25\u003c/b\u003e, 376\u0026ndash;384. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s10157-020-01997-3\u003c/span\u003e\u003cspan address=\"10.1007/s10157-020-01997-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu, J. H. et al. U-shaped association between sleep duration and urinary albumin excretion in Korean adults: 2011\u0026ndash;2014 Korea National Health and Nutrition Examination Survey. \u003cem\u003ePLoS One\u003c/em\u003e. \u003cb\u003e13\u003c/b\u003e, e0192980. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0192980\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0192980\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeon, B. et al. Association between dyslipidemia and the risk of incident chronic kidney disease affected by genetic susceptibility: Polygenic risk score analysis. \u003cem\u003ePLoS One\u003c/em\u003e. \u003cb\u003e19\u003c/b\u003e, e0299605. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1371/journal.pone.0299605\u003c/span\u003e\u003cspan address=\"10.1371/journal.pone.0299605\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee, Y., Seo, E., Mun, E. \u0026amp; Lee, W. A longitudinal study of working hours and chronic kidney disease in healthy workers: The Kangbuk Samsung Health Study. \u003cem\u003eJ. Occup. Health\u003c/em\u003e. \u003cb\u003e63\u003c/b\u003e, e12266. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1002/1348-9585.12266\u003c/span\u003e\u003cspan address=\"10.1002/1348-9585.12266\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee, D. Y. \u0026amp; Shin, S. Association between Chronic Kidney Disease and Dynapenia in Elderly Koreans. \u003cem\u003eHealthc. (Basel)\u003c/em\u003e. \u003cb\u003e11\u003c/b\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/healthcare11222976\u003c/span\u003e\u003cspan address=\"10.3390/healthcare11222976\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuyon, I., Weston, J., Barnhill, S. \u0026amp; Vapnik, V. Gene selection for cancer classification using support vector machines. \u003cem\u003eMach. Learn.\u003c/em\u003e \u003cb\u003e46\u003c/b\u003e, 389\u0026ndash;422 (2002).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaito, T. \u0026amp; Rehmsmeier, M. The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets. \u003cem\u003ePloS one\u003c/em\u003e. \u003cb\u003e10\u003c/b\u003e, e0118432 (2015).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen, T. \u0026amp; Guestrin, C. in \u003cem\u003eProceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining.\u003c/em\u003e 785\u0026ndash;794 (2016).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee, W. V. et al. Development of a machine learning model for precision prognosis of rapid kidney function decline in people with diabetes and chronic kidney disease. \u003cem\u003eDiabetes Res. Clin. Pract.\u003c/em\u003e \u003cb\u003e217\u003c/b\u003e, 111897. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.diabres.2024.111897\u003c/span\u003e\u003cspan address=\"10.1016/j.diabres.2024.111897\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSnoek, J., Larochelle, H. \u0026amp; Adams, R. P. Practical bayesian optimization of machine learning algorithms. \u003cem\u003eAdvances neural Inform. Process. systems\u003c/em\u003e 25 (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGoldstein, A., Kapelner, A., Bleich, J. \u0026amp; Pitkin, E. Peeking inside the black box: Visualizing statistical learning with plots of individual conditional expectation. \u003cem\u003eJ. Comput. Graphical Stat.\u003c/em\u003e \u003cb\u003e24\u003c/b\u003e, 44\u0026ndash;65 (2015).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"chronic kidney disease, obesity, machine learning, risk prevention","lastPublishedDoi":"10.21203/rs.3.rs-9308066/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9308066/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eChronic kidney disease (CKD) is a major public health concern, particularly among individuals with obesity; however, early identification of high-risk individuals remains challenging. This study aimed to develop an interpretable machine learning model to predict CKD risk and identify key contributing factors using nationwide data. We analyzed 37,240 obese adults aged 19\u0026ndash;79 years from the Korea National Health and Nutrition Examination Survey (2007\u0026ndash;2023). Feature selection was performed using recursive feature elimination with cross-validation (RFECV), and an eXtreme Gradient Boosting (XGBoost) model was developed. Model performance was evaluated using cross-validation, calibration, and decision curve analysis, and interpretability was assessed using Shapley Additive exPlanations (SHAP) and individual conditional expectation analyses. Seventeen predictors were selected, with age, metabolic comorbidities, and central obesity measures identified as the most important factors. The final model showed moderate discriminative performance (AUROC 0.756) and probability calibration improved reliability. Lifestyle-related variables, including water and sodium intake, and frequency of eating out, showed modest contributions but were retained as modifiable predictors. Decision curve analysis indicated modest clinical utility, with the model primarily helping reduce unnecessary interventions. This interpretable model highlights the importance of central obesity and metabolic comorbidities while supporting the role of modifiable factors in CKD risk stratification.\u003c/p\u003e","manuscriptTitle":"An interpretable machine learning model for predicting chronic kidney disease risk among obese adults: a nationwide population-based study","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-23 10:48:48","doi":"10.21203/rs.3.rs-9308066/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-05-08T04:40:07+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-07T07:48:10+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-16T03:43:38+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"543295148833614803264984497550995699","date":"2026-04-16T00:35:38+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"72189792388711714284286135806762057461","date":"2026-04-15T16:58:28+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-04-15T13:49:50+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-04-08T14:41:45+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-04-06T12:13:33+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-04-06T12:13:08+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2026-04-03T02:37:05+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2772d730-f572-4341-8293-954f69875995","owner":[],"postedDate":"April 23rd, 2026","published":true,"recentEditorialEvents":[{"type":"decision","content":"Revision requested","date":"2026-05-08T04:40:07+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-05-07T07:48:10+00:00","index":74,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"in-revision","subjectAreas":[{"id":66670724,"name":"Health sciences/Diseases"},{"id":66670725,"name":"Health sciences/Health care"},{"id":66670726,"name":"Health sciences/Medical research"},{"id":66670727,"name":"Health sciences/Nephrology"},{"id":66670728,"name":"Health sciences/Risk factors"}],"tags":[],"updatedAt":"2026-05-08T04:54:03+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-23 10:48:48","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9308066","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9308066","identity":"rs-9308066","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.