Multinational, Calibrated, Non-Laboratory Prevalent Disease Prediction and Survival Modeling for Diabetes, CKD, and CVD

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Reliable non-laboratory tools for assessment of probability of prevalent disease (PPD) are essential for scalable prevention, yet existing models are typically specific to single diseases, require laboratory tests, and show no or limited calibration across PPD strata, limiting scaling and public health utilization. We developed and validated a unified, non-invasive machine-learning model for simultaneous prediction of diabetes, chronic kidney disease, and cardiovascular disease PPD non-invasive predictors. The model was trained on 2011–2016 National Health and Nutrition Examination Survey data (n = 29,903) and evaluated on an independent 2017–2020 test set (n = 15,559). It demonstrated moderate-to-strong discrimination (C-statistic = 0.80–0.90), stable precision–recall performance, and moderate to strong calibration (slope > 0.94). Validation in an independent Korean population showed no or minimal degradation in discrimination and calibration performance, though more extensive validation is warranted. Predicted PPD was associated with cause-specific mortality over up to 7 years of follow-up, consistent with a predictor of latent disease burden. Each 10-percentage-point increase in predicted PPD was associated with roughly a two-fold higher hazard of disease-specific death (HR 2.00–2.20). We conclude that this model has potential as scalable, low-burden screening/surveillance aid, but note that it is not intended as a diagnostic or prognostic tool.
Full text 92,456 characters · extracted from preprint-html · click to expand
Multinational, Calibrated, Non-Laboratory Prevalent Disease Prediction and Survival Modeling for Diabetes, CKD, and CVD | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Multinational, Calibrated, Non-Laboratory Prevalent Disease Prediction and Survival Modeling for Diabetes, CKD, and CVD Arthur Moreira Costa, Iris Badezet-Delory This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8311243/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Reliable non-laboratory tools for assessment of probability of prevalent disease (PPD) are essential for scalable prevention, yet existing models are typically specific to single diseases, require laboratory tests, and show no or limited calibration across PPD strata, limiting scaling and public health utilization. We developed and validated a unified, non-invasive machine-learning model for simultaneous prediction of diabetes, chronic kidney disease, and cardiovascular disease PPD non-invasive predictors. The model was trained on 2011–2016 National Health and Nutrition Examination Survey data (n = 29,903) and evaluated on an independent 2017–2020 test set (n = 15,559). It demonstrated moderate-to-strong discrimination (C-statistic = 0.80–0.90), stable precision–recall performance, and moderate to strong calibration (slope > 0.94). Validation in an independent Korean population showed no or minimal degradation in discrimination and calibration performance, though more extensive validation is warranted. Predicted PPD was associated with cause-specific mortality over up to 7 years of follow-up, consistent with a predictor of latent disease burden. Each 10-percentage-point increase in predicted PPD was associated with roughly a two-fold higher hazard of disease-specific death (HR 2.00–2.20). We conclude that this model has potential as scalable, low-burden screening/surveillance aid, but note that it is not intended as a diagnostic or prognostic tool. Health sciences/Biomarkers Health sciences/Diseases Health sciences/Endocrinology Health sciences/Health care Health sciences/Medical research Health sciences/Nephrology Health sciences/Risk factors Diabetes Mellitus Cardiovascular Diseases Renal Insufficiency Chronic Machine Learning Diet Prevalent Disease Public Health Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Noncommunicable diseases (NCDs) remain the leading cause of global mortality, accounting for nearly three-quarters of non-pandemic-related deaths worldwide [ 1 ]. Their impact continues to intensify, with the Americas experiencing a 43% rise in NCD-related deaths since 2000 and exceeding six million deaths annually. In the United States, diabetes, chronic kidney disease (CKD), and cardiovascular disease (CVD) affect approximately 14.7%, 14%, and 9.9% of adults, respectively, with disproportionately higher burdens observed among older adults, Black individuals, and socioeconomically disadvantaged populations [ 1 , 2 , 3 ]. Because these chronic diseases share many modifiable factors and remain largely asymptomatic in early stages, scalable and low-burden approaches to predicting the probability of prevalent disease (PPD) are essential for effective prevention and public-health planning. Despite considerable progress in PPD modeling, existing non-laboratory tools suffer from several constraints that limit their practical utility. Many rely on detailed self-reported dietary intake, which is known to exhibit substantial day-to-day variability, systematic underreporting, and high respondent burden [ 4 , 5 ]. Other widely used tools lack calibration, restricting their use to coarse PPD categories rather than providing individualized, continuous PPD estimates necessary for informed decision-making and communication [ 6 ]. Moreover, most existing models are disease-specific, requiring individuals and public-health systems to administer separate tools for diabetes, CKD, and CVD despite substantial overlap in their non-laboratory predictors [ 7 , 8 ]. Furthermore, most existing models are not externally validated and are thus limited to application in the country they were trained in. Together, these limitations reduce scalability, impede interpretability, and increase user burden, especially in population-level screening contexts. To address these challenges, we evaluate whether a unified, low-burden, non-laboratory model can accurately predict PPD for diabetes, CKD, and CVD simultaneously. This model predicts PPD, not future incidence, and mortality associations reflect correlation with latent disease burden. Our models achieved discrimination comparable to established tools as well as calibration slopes close to 1 in both Korean and American populations. We also observe significant and consistent associations between model-predicted PPD and cause-specific mortality, supporting the role of our models as indicators of latent disease burden. The calibration of our models enables the use of continuous, individualized PPD estimates for three chronic diseases simultaneously, a significant improvement upon the typical coarse PPD estimates provided by the majority of the current approaches. The model estimates the probability that an individual already meets diagnostic criteria at the time of assessment using only non-laboratory predictors, but it does not estimate future risk or progression. It is intended as a scalable, low-burden screening and surveillance aid, not a diagnostic or prognostic tool. Results Population Description Table S1 shows the characteristics of the weighted 2011–2020 National Health and Nutrition Examination Survey (NHANES) sample based on sex. There was an even distribution of sex, with participants averaging 37.74 ± 22.51 years old. Additionally, over 60 percent of the participants were identified as non-Hispanic whites, which is consistent with the demographic composition of the United States [ 9 ]. Furthermore, most participants had completed at least a high school education. Most women were non-smokers, but nearly half of men were smokers. Model Design Fig. 1 summarizes the modeling workflow, including data sources, preprocessing, model development, internal and external performance evaluation, and mortality association analysis. More detailed flow diagrams of model design and exclusion criteria are available in supplementary figures S1 and S2. Model Discrimination Fig. 2 shows receiver operating characteristic (ROC) curves for diabetes, CKD, and CVD PPD prediction models and Area Under the ROC Curve (AUC), sensitivity, and specificity comparisons. All three models showed moderate to strong discrimination (AUC = 0.80–0.90) and sensitivity (0.79–0.89), with relatively weaker specificity (0.60–0.77). Precision-Recall and Calibration Figure 3 displays precision–recall curves and calibration analyses after isotonic recalibration with 5-fold cross-validation, using training data for calibration. All of our models showed significantly higher PR-AUC than that of prevalence-based null models. Calibration curves indicated good alignment between predicted and observed PPD across deciles, with minimal over- or underprediction. External Validation and Model Comparisons Further evaluation of calibration performance was undertaken by quantitative comparisons of model Brier scores to prevalence-based controls (Table 1 ). All models achieved calibration slopes near 1 (slopes = 0.94–0.99). Diabetes had a particularly high intercept (0.288). All models were significantly better calibrated than null prevalence-based PPD estimators. Calibration performance in the 2015 Korean NHANES (KNHANES) cohort (Fig. S4; Table 1 ) showed strong calibration for diabetes and CVD (slopes > 0.99, |intercepts| < 0.02). To contextualize performance, we compared our models’ AUCs on NHANES and KNHANES with those reported for established non-laboratory models for diabetes, CKD, and CVD (Fig. 4 ). Table 1 Cross-country comparison of discrimination, precision-recall, and calibration U.S. South Korea Metric Diabetes CKD CVD Diabetes CKD CVD N 9,908 9,475 9,231 5,824 7,380 6,459 Prevalence (%) 13.7 7.6 12.7 8.5 2.2 4.3 AUC 0.819 ± 0.010 0.895 ± 0.011 0.798 ± 0.012 0.768 ± 0.018 0.886 ± 0.019 0.868 ± 0.017 PR AUC 0.380 ± 0.025 0.405 ± 0.035 0.337 ± 0.026 0.183 ± 0.021 0.123 ± 0.029 0.188 ± 0.033 Null Brier 0.1179 0.07 0.1112 0.0779 0.0215 0.0408 Brier 0.1001 ± 0.0042 0.0544 ± 0.0032 0.0966 ± 0.0042 0.0723 ± 0.0052 0.0199 ± 0.0027 0.0365 ± 0.0033 ECE 0.0288 0.0061 0.0139 < 0.001 < 0.001 < 0.001 Calibration Slope 0.98 0.988 0.943 0.996 0.995 0.996 Calibration Intercept 0.288 –0.097 0.058 –0.008 –0.014 –0.008 Survival analysis Figure 5 shows validation of our predicted PPD against future cause-specific mortality. A 10-percentage-point increase in predicted diabetes, CKD, or CVD PPD was associated with approximately a two-fold higher hazard of disease-specific death (HR range: 2.00–2.20; p < 0.001 for all). When participants were stratified into quartiles of predicted prevalent CVD PPD, 7-year cumulative CVD mortality demonstrated clear monotonic separation (Fig. 5 ). Mortality increased from 0.0% (0.0–0.0) in Q1 to 1.2% (0.8–1.6) in Q4, representing more than a ten-fold gradient across quartiles (Table 2 ). Table 2 Hazard ratio and mortality by predicted-PPD quartiles Group N used Events HR † / % 7-year mortality* Diabetes mortality 22,734 155 2.00 † (1.82–2.20) CKD mortality 22,734 25 2.20 † (1.85–2.61) CVD mortality 22,734 359 2.13 † (2.03–2.24) Q1 CVD PPD 2,686 10 0.0* (0.0–0.0) Q2 CVD PPD 2,697 25 0.1* (0.0–0.2) Q3 CVD PPD 2,680 52 0.3* (0.1–0.6) Q4 CVD PPD 2,671 227 1.2* (0.8–1.6) †: Hazard-ratio (95% CI); *: % 7-year mortality (95% CI) Sensitivity Analyses Sensitivity analyses showed that age, BMI, and waist circumference were the strongest predictors across diseases, consistent with partial dependence patterns. Feature-ablation experiments, including the removal or addition of dietary predictors, use of simpler tree ensembles, and alternative imputation strategies, did not materially change discrimination or calibration. Full results are provided in Supplementary Tables S2–S3 and Supplementary Figure S4. Discussion We were able to create a model with comparable discrimination to current leading models, while also being well-calibrated, externally validated, and associated with future cause-specific mortality. It is also a single pipeline that enables simultaneous PPD prediction for three chronic diseases using a single set of general non-invasive predictors. When externally validating our model on 2015 KNHANES data without retraining, discrimination and calibration performance were comparable to the test set on 2017–2020 holdout test data from NHANES. Further comparison with established models indicates that our unified, three-disease model achieves discrimination on par with the leading noninvasive tools. Our CKD model showed particularly high discrimination, significantly higher than the leading models we analyzed. These results strengthen our claim that our tool can allow for decreased user burden by simultaneously screening for three diseases without losses in performance. Our survival analyses showed significant associations between predicted prevalent-disease probabilities and cause-specific mortality (Fig. 5 ; Table 2 ). These associations suggest that the model’s outputs do indeed correlate with underlying disease. We then investigated how our lack of dietary predictors and our choice of boosted trees over simpler models were affecting model performance. We found that diet-aware models did not show improved calibration or discrimination. This finding is likely attributed to the low consistency and reliability of short-term dietary recall data collected by NHANES, rather than a physiological independence between diet and chronic disease PPD. We also found that simpler bagged tree models were unable to match the discrimination of our boosted tree models. Finally, we found near-identical performance of KNN and mean/mode imputation, which suggests that missing-data handling was not a major driver of model behavior, reducing concern that imputation introduced bias. It is important to note, however, that all of our analyses are subject to significant limitations. Our population-level visualizations shown in figures S3 and S5 were generated using an unweighted population from our 2011–2016 NHANES training data. This population is not representative of the true US population and should not be extrapolated to make conclusions about disease-burdens on a full population. Additionally, although our model was ported to a Korean population (KNHANES), portability to lower-middle-income countries was not analyzed and should be addressed in future work. Another limitation is that our analyses are all association-based, and causal inferences cannot be drawn from our model. Similarly, our model does not indicate longitudinal or incident risk. Because the model is trained on cross-sectional data, predicted probabilities should not be interpreted as forecasts of future incidence or as estimates of the effects of behavioral or clinical interventions. Future work should assess this framework in prospective cohorts, incorporate repeated measurements, or evaluate whether longitudinal changes in predicted probabilities are associated with subsequent differences in long-term outcomes. Materials and Methods Study population and weighting We used NHANES 2011–2020 data to train and test our model [ 24 ]. NHANES 2011–2016 data was used for training and calibration (n = 29903), and 2017–2020 data was used for testing (n = 15,559). Sampling weights were applied for descriptive statistics but not for machine learning model development. We did not apply survey weights during model training because our goal was to learn the individual-level conditional probability P(Y = 1∣X), not to estimate population-level prevalence. Moreover, weighting increases population representativeness but doesn’t improve predictive accuracy can degrade individual-level risk estimation. Furthermore, NHANES weights can distort model fitting by over-emphasizing subgroups that make up larger proportions of the US population. Unweighted training therefore provides a cleaner estimate of the covariate-outcome relationship, while absolute risk alignment is handled separately through calibration and external validation. Weights for descriptive statistics were calculated according to NHANES guidelines using mobile examination center weights [ 24 ]. Weights were normalized to a mean value of 1.00, which is a standard procedure used in other NHANES-based analyses [e.g. 25, 26]. Predictors and outcome Standard predictors included demographic variables (age, sex, income-to-poverty ratio), anthropometric measures (body mass index (BMI), waist circumference), and physiologic measures (blood pressure, heart rate). Dietary predictors included dietary sugar, saturated fat, carbohydrate, fiber, polyunsaturated fat, mineral, vitamin, caffeine, theobromine, alcohol, and caloric intake. Diabetes status was defined as glycohemoglobin (HbA1c) ≥ 6.5%, consistent with the recommendation of the American Diabetes Association [ 27 ]. CKD status was determined according to KDIGO 2012 criteria as eGFR < 60 mL/min/1.73m², using the CKD-EPI equation [ 28 ]. CVD status was designated as a self-reported history of coronary heart disease, congestive heart failure, angina pectoris, heart attack, or stroke. For external validation on the KNHANES data, CVD was defined as a positive self-reported history of myocardial infarction or angina. Missing predictor values were imputed using training-set means for continuous variables, and modes for categorical variables. Participants missing HbA1c were excluded from diabetes analyses. Those missing any variables needed to compute eGFR (most commonly serum creatinine) were excluded from CKD analyses. Those missing any variables used to define CVD status were excluded from CVD analyses. Predictors in external validation were manually converted to NHANES-coded variables. A variable dictionary and missingness table are shown in tables S4 and S5. Model training and evaluation An initial model screening, trained on NHANES 2011–2016 data and tested on independent 2017–2020 data, was conducted. Models tested included logistic regression, regularized logistic regression, naïve Bayes, k-nearest neighbors, decision trees, random forests, bagged trees, boosted trees (AdaBoost and gradient boosting), and support vector machines with linear and nonlinear kernels. Across these models, tree-based classifiers showed the lowest classification error. The gradient boosting framework XGboost was selected for further development because it achieved the highest discrimination while offering stable performance across resamples, favorable calibration behavior with minimal hyperparameter tuning, robustness to moderate shifts in predictor distributions between NHANES and KNHANES, and clean integration with permutation-based interpretability and partial-dependence analyses. Model hyperparameters are described in the supplementary information, with key parameters listed in table S6. Model performance was evaluated using sensitivity, specificity, area under the Receiver Operating Characteristic (ROC) curve (AUC), and Brier score. To investigate whether we could exclude dietary variables without sacrificing predictive performance, four models were explored for each disease: a 48-hour-recall diet model (which averages day 1 and day 2 recalls into mean intake variables), a 24-hour-recall diet model, a standard predictor model, and a standard predictor + 48-hour dietary recall model. All model thresholds were chosen to maximize Youden’s J (i.e. sensitivity + specificity − 1). To account for total energy intake, we created energy-adjusted dietary variables (nutrient intake per 1,000 kcal) and repeated diet-only models with these adjusted features. AUC was calculated for each model and was used to compare discrimination performance between models. In addition to ROC analysis, we evaluated model performance using precision–recall (PR) curves and area under the PR curve (PR-AUC), which is more informative under class imbalance. Calibration (reliability) curves were generated using decile bins to assess agreement between predicted and observed PPD. We also computed bootstrap 95% confidence intervals (CIs) for AUC using bootstrap resampling. Null Brier scores for each chronic disease were calculated by creating hypothetical models that assign each individual in the test population a PPD estimate equal to the prevalence in the studied population/subpopulation. Our external validation was done by curating the 2015 KNHANES dataset such that variable labels and values were matched to NHANES data used to train the models, except for CVD, which could only be partially matched, as discussed above [ 29 ]. We then compared our model to other external models, selected via a literature search for non-invasive diabetes, CKD, and CVD PPD models, prioritizing varied recent and well-established models (e.g. FINDRISC, INTERHEART). Partial-dependence analyses To examine how predicted probabilities vary with individual predictors, we computed partial-dependence sensitivity analyses using the observed NHANES training sample. For each grid value, we replaced the predictor value for all individuals in the training set while leaving all other predictors unchanged, passed the modified dataset through the trained model, and calculated the mean predicted probability for each disease. This yields a one-dimensional partial-dependence curve that summarizes how the model’s predictions vary with that predictor, averaged over the joint distribution of all other predictors in the training sample. Cox Models, Discrimination, and Survival Across Predicted PPD Quartiles Cause-specific mortality was obtained from the 2019 NHANES-linked National Death Index files. Follow-up time was calculated from the MEC examination using PERMTH_INT (months) and converted to years. Participants outside the 2011–2016 survey cycles or missing survival time, cause-of-death indicators, or predicted-PPD estimates were excluded. To ensure meaningful exposure windows for chronic-disease mortality and to match established NHANES cardiovascular mortality analyses, we restricted the survival analyses to adults aged ≥ 40 years. Event indicators were defined for diabetes, CKD, and cardiovascular mortality using the underlying-cause classification provided in the linked files. For each mortality endpoint, we fit pooled Cox proportional hazards models across all eligible NHANES cycles, including cohort indicators and scaling predicted PPDs so hazard ratios represent the effect of a 10-percentage-point increase in model-predicted PPD. Cox coefficients were estimated using a numerically stable Breslow partial-likelihood implementation with ridge penalization, which improves convergence for sparse endpoints (e.g., CKD mortality). To evaluate PPD stratification, predicted CVD PPD was divided into quartiles, and 7-year cumulative CVD mortality was estimated within each quartile using Kaplan–Meier curves and bootstrap-derived confidence intervals. Similar analyses were not conducted for diabetes or CVD due to insufficient cause-specific mortality. Abbreviations PPD, probability of prevalent disease; NCD, noncommunicable disease; DM, diabetes mellitus, CVD, cardiovascular disease; CKD, chronic kidney disease; NHANES, National Health and Nutrition Examination Survey; CI, confidence interval; AUC, area under curve; ROC, receiver operating characteristic; BMI, body mass index; HbA1c, glycohemoglobin; KDIGO, Kidney Disease Improving Global Outcomes; PR, precision-recall; PR-AUC, area under the precision-recall curve; eGFR, estimated glomerular filtration rate; CDC, Centers for Disease Control; PDP, partial dependence plot Declarations Acknowledgments: The authors thank Gabriela Gomez, Dr. Ismael Haddadian, and Dr. Sandra C. Fuchs for helpful feedback. Author Contributions: A.M.C completed the conceptualization, methodology, and formal analysis. Investigation and data curation was completed by A.M.C and I.B.D. The first draft of the manuscript was written and edited by A.M.C and I.B.D. All authors have read and agreed to the published version of the manuscript. Competing interests: A.M.C. and I.B.D. are inventors on a provisional patent application related to the methods described in this manuscript. The authors declare no other competing interests. Ethics statement: NHANES and KNHANES are publicly available, deidentified national health surveys that obtain informed consent from all participants and are approved by their respective institutional review boards. This secondary analysis of fully anonymized public data was exempt from institutional review board review. Informed Consent Statement: All participants in NHANES and KNHANES provided written informed consent. Data Availability: All data used in this study are publicly accessible. NHANES 2011–2020 data can be obtained from the U.S. Centers for Disease Control and Prevention (https://www.cdc.gov/nchs/nhanes). KNHANES 2015 microdata are available from the Korea Disease Control and Prevention Agency via the KNHANES website (https://knhanes.kdca.go.kr) following free online registration; the authors do not redistribute these files. Mortality data were obtained from the NHANES-linked National Death Index. No proprietary or restricted-access datasets were used. Code availability: All analysis code used in this study are available in repositories linked in the supplementary information. Funding: This research received no external funding. References Centers for Disease Control and Prevention. Diabetes: Data & Research. Atlanta, GA: CDC. (accessed 4 Sept 2025). Centers for Disease Control and Prevention. Chronic Kidney Disease: Data & Research (Fast Stats). Atlanta, GA: CDC. (accessed 4 Sept 2025). Martin SS, et al. Heart Disease and Stroke Statistics—2024 Update: A Report From the American Heart Association. Circulation. 2024;149:e347–e913. Poslusna K, et al. Misreporting of Energy and Micronutrient Intake Estimated by Food Records and 24 h Recalls, Control and Adjustment Methods in Practice. Br J Nutr. 2009;101(S2):S73–S85. Ahluwalia N, et al. Update on NHANES Dietary Data: Focus on Collection, Release, Analytical Considerations, and Uses to Inform Public Policy. Adv Nutr. 2016;7:121–134. Van Calster B, et al. Calibration: The Achilles Heel of Predictive Analytics. BMC Med. 2019;17:230. Lindström J, Tuomilehto J. The Finnish Diabetes Risk Score (FINDRISC): A Valid Tool to Estimate the Risk of Type 2 Diabetes. Diabetes Care. 2003;26(3):725–731. D’Agostino RB Sr, et al. General Cardiovascular Risk Profile for Use in Primary Care: The Framingham Heart Study. Circulation. 2008;117:743–753. CDC/NCHS. Health, United States—Sources and Definitions: National Health and Nutrition Examination Survey (NHANES). (accessed 4 Sept 2025). Zhou X, Qiao Q, Ji L, et al. Non-laboratory-based risk assessment algorithm for prevalent type 2 diabetes developed on a nation-wide diabetes survey. Diabetes Care. 2013;36(12):3944-3952. Abdallah L, Naja F, Choufani J, et al. Validation of the Finnish Diabetes Risk Score (FINDRISC) in the Lebanese population. Prim Care Diabetes. 2020;14(3):271-277. Mugume I, Kibirige D, Ssebunya R, et al. Performance of the Finnish Diabetes Risk Score in detecting prevalent type 2 diabetes and dysglycaemia in a rural Ugandan population. PLOS One. 2023;18(11):e0276858. Ha KH, Kim DJ. Development and validation of the Korean Diabetes Risk Score: a prospective study. Diabetologia. 2018;61(4):949-958. Zhou X, Li W, Wong CKH, et al. Simple non-laboratory and laboratory-based risk assessment algorithms for prevalent diabetes in a Chinese population. J Diabetes Investig. 2016;7(5):657-666. Bang H, Vupputuri S, Shoham DA, et al. SCreening for Occult REnal Disease (SCORED): a simple prediction model for chronic kidney disease. Arch Intern Med. 2007;167(4):374-381. O’Seaghdha CM, Lyass A, Massaro JM, et al. A risk score for chronic kidney disease in the general population. Am J Med. 2012;125(3):270-277. Tangri N, Stevens LA, Griffith J, et al. A predictive model for progression of chronic kidney disease to kidney failure. JAMA. 2011;305(15):1553-1559. Gaziano TA, Young CR, Fitzmaurice G, Atwood S, Gaziano JM. Laboratory-based versus non-laboratory-based method for assessment of cardiovascular disease risk: the NHANES I Follow-up Study cohort. Lancet. 2008;371(9616):923-931. Pandya A, Weinstein MC, Gaziano TA. A comparative assessment of non-laboratory-based versus commonly used laboratory-based cardiovascular disease risk scores in the NHANES III population. PLOS One. 2011;6(5):e20416. McGorrian C, Yusuf S, Islam S, et al. Estimating modifiable coronary heart disease risk in multiple regions of the world: the INTERHEART Modifiable Risk Score. Eur Heart J. 2011;32(5):581-589. Kaptoge S, Pennells L, De Bacquer D, et al. World Health Organization cardiovascular disease risk charts: revised models to estimate risk in 21 global regions. Lancet Glob Health. 2019;7(10):e1332-e1345. Rezaei F, Yadegarfar G, Gohari K, et al. Agreement between laboratory-based and non-laboratory-based Framingham risk scores in a large Middle Eastern cohort. Sci Rep. 2021;11:10687. Jia, G.; Aroor, A.R.; Sowers, J.R. Arterial Stiffness: A Nexus between Cardiac and Renal Disease. Cardiorenal Med. 2014, 4, 60–71. https://doi.org/10.1159/000360867. University of Texas at Arlington. National Health and Nutrition Examination Survey (NHANES)—Big Data for Epidemiology. Available online: https://uta.pressbooks.pub/bigdataforepidemiology/chapter/chapter10-nhanes/ (accessed on 4 September 2025). Costa, A.M.; Sias, R.J.; Fuchs, S.C. Effect of Whole Blood Dietary Mineral Concentrations on Erythrocytes: Selenium, Manganese, and Chromium: NHANES Data. Nutrients 2024, 16, 3653. https://doi.org/10.3390/nu16213653. Chen, J.; Kan, M.; Ratnasekera, P.; Deol, L.K.; Thakkar, V.; Davison, K.M. Blood Chromium Levels and Their Association with Cardiovascular Diseases, Diabetes, and Depression: NHANES 2015–2016. Nutrients 2022, 14, 2687. https://doi.org/10.3390/nu14132687. American Diabetes Association. Diabetes Diagnosis. Available online: https://diabetes.org/about-diabetes/diagnosis (accessed on 4 September 2025). World Health Organization. Noncommunicable Diseases – Fact Sheet. Geneva: WHO. (accessed 4 Sept 2025). Kweon, Sanghui et al. "Data Resource Profile: The Korea National Health and Nutrition Examination Survey (KNHANES)" vol. 43, no. 1, 2014 Additional Declarations Competing interest reported. A.M.C. and I.B.D. are inventors on a U.S. provisional patent application (Application No. 63/933,737) related to the methods described in this manuscript. A.M.C. and I.B.D. are the filing applicants. Parts of this manuscript included in the patent application include: methods for PPD prediction and a public-facing research and indivudal-use PPD prediction tool. The authors declare no other competing interests. Supplementary Files NSR2025SI12082025.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8311243","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":557403904,"identity":"d0338043-7758-42d6-9231-e7c411d32e95","order_by":0,"name":"Arthur Moreira Costa","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA/UlEQVRIiWNgGAWjYHACNhiD8QEDwwEeKIcZjw5mNphOZgOStbBJALXAhXFq4J99/thj3j12cvLzm59V89TckeFvP37xAUOFdWIDDi0S55LZjXmeJRsbHGMzu81z7BmPxJmcYgOGM+k4tTCcYWaT5jlwIHEDG4PZbd6GwzwGDDlpEoxth3FqkYdpmd/G/q0YrIX/DVDLP9xaDGBaGo7xmDGDtUikH5NgbMCtxfAMs5nknAMgv+QUS845dphH4sYbZoOEY+nGuLTInWF8JvHmADDEmo9v/PCm5rA9f3/6wwcfaqxlcXofCwCGQAIJykGA/QGJGkbBKBgFo2CYAwCAxlVrpGO5NwAAAABJRU5ErkJggg==","orcid":"","institution":"University of Chicago","correspondingAuthor":true,"prefix":"","firstName":"Arthur","middleName":"Moreira","lastName":"Costa","suffix":""},{"id":557403905,"identity":"fb0f5bf3-0b57-4520-a7e9-5682c7bc88ad","order_by":1,"name":"Iris Badezet-Delory","email":"","orcid":"","institution":"University of Chicago","correspondingAuthor":false,"prefix":"","firstName":"Iris","middleName":"","lastName":"Badezet-Delory","suffix":""}],"badges":[],"createdAt":"2025-12-08 22:08:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8311243/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8311243/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":97898532,"identity":"17b4b2ea-c78d-4c73-bcf4-080ebf275572","added_by":"auto","created_at":"2025-12-10 15:39:16","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6536549,"visible":true,"origin":"","legend":"","description":"","filename":"NSR2025Manuscript12082025.docx","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/dd2a64d14ffb2d0ff2a21cbc.docx"},{"id":97899304,"identity":"5997898a-0a71-4b98-bfbb-2af515ab107b","added_by":"auto","created_at":"2025-12-10 15:42:53","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5437,"visible":true,"origin":"","legend":"","description":"","filename":"73e615b1d6b14a06956f6caec1ea06b9.json","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/caa211c0cf32bda6a7356669.json"},{"id":97851604,"identity":"d445ffab-192e-4bb1-b32f-821d6495cd68","added_by":"auto","created_at":"2025-12-10 07:00:27","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":876281,"visible":true,"origin":"","legend":"","description":"","filename":"NSR2025SI12082025.docx","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/623e9dcc17808292f5baea66.docx"},{"id":97898849,"identity":"242a6553-af61-4572-9a3d-4750665052df","added_by":"auto","created_at":"2025-12-10 15:39:51","extension":"xml","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":73822,"visible":true,"origin":"","legend":"","description":"","filename":"73e615b1d6b14a06956f6caec1ea06b91enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/0b4ebcf27e8509ed9b20ee95.xml"},{"id":97899558,"identity":"e61e0a41-fb78-4b01-bbe6-84ee7b6ec260","added_by":"auto","created_at":"2025-12-10 15:44:44","extension":"png","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":121774,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/695d01c64bf0ec32078e30f3.png"},{"id":97851617,"identity":"f164b8d7-14fc-40c8-bafe-f402e79da008","added_by":"auto","created_at":"2025-12-10 07:00:28","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":279084,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/2de7b74c4a3a3912e2802b3d.png"},{"id":97851620,"identity":"1664aaa8-75b6-4c42-a84a-89ef6c26ead3","added_by":"auto","created_at":"2025-12-10 07:00:28","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":203744,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/e10e061bc0cf03cf05096565.png"},{"id":97851615,"identity":"bfd93049-e0be-429f-bfe0-5ba83b8f2c45","added_by":"auto","created_at":"2025-12-10 07:00:28","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":216521,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/321f656c3a92cefe5ee6e403.png"},{"id":97897386,"identity":"83be8e2e-17fc-4c75-93cc-75a1dfcbedf3","added_by":"auto","created_at":"2025-12-10 15:37:47","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":204022,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/75ca984e13fb14ee5a2bdd8c.png"},{"id":97897419,"identity":"fc89f929-a1cd-4e03-adf5-6720115beb3a","added_by":"auto","created_at":"2025-12-10 15:37:48","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":32835,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/f5a4d70e61394c0eaf9949e3.png"},{"id":97851613,"identity":"e729409e-75f3-4f58-b814-0f5f1c2ed556","added_by":"auto","created_at":"2025-12-10 07:00:28","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":36183,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/86e2451dedb5c1880059a269.png"},{"id":97851606,"identity":"b85d1013-1a6b-48d0-a42f-a2a232f1478a","added_by":"auto","created_at":"2025-12-10 07:00:28","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":28329,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/4cde27fb69cf2fad4dc47f90.png"},{"id":97851611,"identity":"75a94348-4f87-4da1-bfb7-544f4da1951d","added_by":"auto","created_at":"2025-12-10 07:00:28","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":41877,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/44a2f8a3ceb02e7263bf6f5a.png"},{"id":97851608,"identity":"ec387f3b-bb31-45d5-b6a5-54da0579aa68","added_by":"auto","created_at":"2025-12-10 07:00:28","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":35779,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/f114e497ffc484c4f7381234.png"},{"id":97898774,"identity":"c2db8614-742d-46eb-8adb-1da7f2f7b887","added_by":"auto","created_at":"2025-12-10 15:39:37","extension":"xml","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":71253,"visible":true,"origin":"","legend":"","description":"","filename":"73e615b1d6b14a06956f6caec1ea06b91structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/cb6d8fb3fad9b589351b27a6.xml"},{"id":97900470,"identity":"38f5da59-f3cf-4a71-af5f-8954a4836966","added_by":"auto","created_at":"2025-12-10 15:45:33","extension":"html","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":80876,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/350909ff3201546add17321a.html"},{"id":97851598,"identity":"938840dc-4b0c-4367-a138-176bbff3ec55","added_by":"auto","created_at":"2025-12-10 07:00:27","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":79609,"visible":true,"origin":"","legend":"\u003cp\u003eModel design pipeline.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/6366261efa867d7cd2f3391f.png"},{"id":97851599,"identity":"cf75df48-81f0-42e5-8ca4-afbd15f9d3ff","added_by":"auto","created_at":"2025-12-10 07:00:27","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":105203,"visible":true,"origin":"","legend":"\u003cp\u003eModel performance metrics. A) ROC curves for the diabetes, CKD, and CVD PPD-prediction models. B) Bar charts comparing AUC, sensitivity, and specificity between the models for each disease. CI bars and confidence ribbons represent the bootstrapped 95% CI with 1000 resamples of the test population with replacement.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/61c8153c3a9f0516c8ab39b6.png"},{"id":97851600,"identity":"66d2bf6f-e9b6-4236-8f82-df5cfffffb3b","added_by":"auto","created_at":"2025-12-10 07:00:27","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":85400,"visible":true,"origin":"","legend":"\u003cp\u003ePrecision-recall and calibration curves. A) PR-AUC is reported and PR curves are compared to prevalence baselines, show in dotted lines. B) Model calibration is compared to perfect calibration across decile predicted PPD bins. Bins with less than 50 individuals are omitted. Error bars represent 95% CIs.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/aea4aa1afdbb5c3eeba176ba.png"},{"id":97851602,"identity":"9140d4fc-4977-4273-af1a-a348d862824e","added_by":"auto","created_at":"2025-12-10 07:00:27","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":95757,"visible":true,"origin":"","legend":"\u003cp\u003eAUC comparisons between our model’s when tested in various countries, and other leading PPD prediction models [10-22]. Our models showed comparable calibration to leading predictive models when tested on both NHANES and KNHANES data. 95% CIs from external models are shown in models where they were reported.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/f7637243aecf4a85b44c28e5.png"},{"id":97851603,"identity":"56c46fef-2b07-405f-a09b-1974c3e155c4","added_by":"auto","created_at":"2025-12-10 07:00:27","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":82105,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eValidation of predicted PPD against cause-specific mortality and longitudinal CVD outcomes.\u003c/strong\u003e\u003cbr\u003e\n \u003cstrong\u003eA)\u003c/strong\u003e Hazard ratios (HRs) for cause-specific mortality per 10-percentage-point increase in predicted \u003cem\u003ePPD\u003c/em\u003e for diabetes, CKD, and CVD. Error bars indicate 95% CIs. \u003cstrong\u003eB)\u003c/strong\u003eCumulative incidence of CVD mortality over 7 years, stratified by quartiles of predicted prevalent CVD PPD at baseline (Q1–Q4).\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/5e4f6420c760c19eaaea02c9.png"},{"id":106412402,"identity":"51f6eab7-f8d5-4d36-9d00-507b918d6399","added_by":"auto","created_at":"2026-04-08 09:59:43","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1083954,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/a5442842-7ba0-413b-a6b5-e534ab30695a.pdf"},{"id":97898865,"identity":"e13d3e9c-4905-4be1-9d4d-7a2bf5944584","added_by":"auto","created_at":"2025-12-10 15:39:54","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":876281,"visible":true,"origin":"","legend":"","description":"","filename":"NSR2025SI12082025.docx","url":"https://assets-eu.researchsquare.com/files/rs-8311243/v1/f252debc9c272c42b3ede29c.docx"}],"financialInterests":"Competing interest reported. A.M.C. and I.B.D. are inventors on a U.S. provisional patent application (Application No. 63/933,737) related to the methods described in this manuscript. A.M.C. and I.B.D. are the filing applicants. Parts of this manuscript included in the patent application include: methods for PPD prediction and a public-facing research and indivudal-use PPD prediction tool. The authors declare no other competing interests.","formattedTitle":"Multinational, Calibrated, Non-Laboratory Prevalent Disease Prediction and Survival Modeling for Diabetes, CKD, and CVD ","fulltext":[{"header":"Introduction","content":"\u003cp\u003eNoncommunicable diseases (NCDs) remain the leading cause of global mortality, accounting for nearly three-quarters of non-pandemic-related deaths worldwide [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Their impact continues to intensify, with the Americas experiencing a 43% rise in NCD-related deaths since 2000 and exceeding six million deaths annually. In the United States, diabetes, chronic kidney disease (CKD), and cardiovascular disease (CVD) affect approximately 14.7%, 14%, and 9.9% of adults, respectively, with disproportionately higher burdens observed among older adults, Black individuals, and socioeconomically disadvantaged populations [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Because these chronic diseases share many modifiable factors and remain largely asymptomatic in early stages, scalable and low-burden approaches to predicting the probability of prevalent disease (PPD) are essential for effective prevention and public-health planning.\u003c/p\u003e\u003cp\u003eDespite considerable progress in PPD modeling, existing non-laboratory tools suffer from several constraints that limit their practical utility. Many rely on detailed self-reported dietary intake, which is known to exhibit substantial day-to-day variability, systematic underreporting, and high respondent burden [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Other widely used tools lack calibration, restricting their use to coarse PPD categories rather than providing individualized, continuous PPD estimates necessary for informed decision-making and communication [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Moreover, most existing models are disease-specific, requiring individuals and public-health systems to administer separate tools for diabetes, CKD, and CVD despite substantial overlap in their non-laboratory predictors [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Furthermore, most existing models are not externally validated and are thus limited to application in the country they were trained in. Together, these limitations reduce scalability, impede interpretability, and increase user burden, especially in population-level screening contexts.\u003c/p\u003e\u003cp\u003eTo address these challenges, we evaluate whether a unified, low-burden, non-laboratory model can accurately predict PPD for diabetes, CKD, and CVD simultaneously. This model predicts PPD, not future incidence, and mortality associations reflect correlation with latent disease burden. Our models achieved discrimination comparable to established tools as well as calibration slopes close to 1 in both Korean and American populations. We also observe significant and consistent associations between model-predicted PPD and cause-specific mortality, supporting the role of our models as indicators of latent disease burden. The calibration of our models enables the use of continuous, individualized PPD estimates for three chronic diseases simultaneously, a significant improvement upon the typical coarse PPD estimates provided by the majority of the current approaches. The model estimates the probability that an individual already meets diagnostic criteria at the time of assessment using only non-laboratory predictors, but it does not estimate future risk or progression. It is intended as a scalable, low-burden screening and surveillance aid, not a diagnostic or prognostic tool.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n \u003ch2\u003ePopulation Description\u003c/h2\u003e\n \u003cp\u003eTable \u003cspan class=\"InternalRef\"\u003eS1\u003c/span\u003e shows the characteristics of the weighted 2011\u0026ndash;2020 National Health and Nutrition Examination Survey (NHANES) sample based on sex. There was an even distribution of sex, with participants averaging 37.74\u0026thinsp;\u0026plusmn;\u0026thinsp;22.51 years old. Additionally, over 60 percent of the participants were identified as non-Hispanic whites, which is consistent with the demographic composition of the United States [\u003cspan class=\"CitationRef\"\u003e9\u003c/span\u003e]. Furthermore, most participants had completed at least a high school education. Most women were non-smokers, but nearly half of men were smokers.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eModel Design\u003c/strong\u003e Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e summarizes the modeling workflow, including data sources, preprocessing, model development, internal and external performance evaluation, and mortality association analysis. More detailed flow diagrams of model design and exclusion criteria are available in supplementary figures \u003cspan class=\"InternalRef\"\u003eS1\u003c/span\u003e and S2.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eModel Discrimination\u003c/strong\u003e Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e shows receiver operating characteristic (ROC) curves for diabetes, CKD, and CVD PPD prediction models and Area Under the ROC Curve (AUC), sensitivity, and specificity comparisons. All three models showed moderate to strong discrimination (AUC\u0026thinsp;=\u0026thinsp;0.80\u0026ndash;0.90) and sensitivity (0.79\u0026ndash;0.89), with relatively weaker specificity (0.60\u0026ndash;0.77).\u003c/p\u003e\n\u003c/div\u003e\n\u003ch3\u003ePrecision-Recall and Calibration\u003c/h3\u003e\n\u003cp\u003eFigure \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e displays precision\u0026ndash;recall curves and calibration analyses after isotonic recalibration with 5-fold cross-validation, using training data for calibration. All of our models showed significantly higher PR-AUC than that of prevalence-based null models. Calibration curves indicated good alignment between predicted and observed PPD across deciles, with minimal over- or underprediction.\u003c/p\u003e\n\u003ch3\u003eExternal Validation and Model Comparisons\u003c/h3\u003e\n\u003cp\u003eFurther evaluation of calibration performance was undertaken by quantitative comparisons of model Brier scores to prevalence-based controls (Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e). All models achieved calibration slopes near 1 (slopes\u0026thinsp;=\u0026thinsp;0.94\u0026ndash;0.99). Diabetes had a particularly high intercept (0.288). All models were significantly better calibrated than null prevalence-based PPD estimators. Calibration performance in the 2015 Korean NHANES (KNHANES) cohort (Fig. S4; Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e) showed strong calibration for diabetes and CVD (slopes\u0026thinsp;\u0026gt;\u0026thinsp;0.99, |intercepts| \u0026lt; 0.02). To contextualize performance, we compared our models\u0026rsquo; AUCs on NHANES and KNHANES with those reported for established non-laboratory models for diabetes, CKD, and CVD (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e).\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n \u003ctable id=\"Tab1\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eCross-country comparison of discrimination, precision-recall, and calibration\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\" colspan=\"3\"\u003e\n \u003cp\u003eU.S.\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colspan=\"3\"\u003e\n \u003cp\u003eSouth Korea\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMetric\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eDiabetes\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eCKD\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eCVD\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eDiabetes\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eCKD\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eCVD\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eN\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e9,908\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e9,475\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e9,231\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e5,824\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7,380\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e6,459\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003ePrevalence (%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e13.7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e7.6\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e12.7\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e8.5\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e2.2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e4.3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eAUC\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.819\u0026thinsp;\u0026plusmn;\u0026thinsp;0.010\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.895\u0026thinsp;\u0026plusmn;\u0026thinsp;0.011\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.798\u0026thinsp;\u0026plusmn;\u0026thinsp;0.012\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.768\u0026thinsp;\u0026plusmn;\u0026thinsp;0.018\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.886\u0026thinsp;\u0026plusmn;\u0026thinsp;0.019\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.868\u0026thinsp;\u0026plusmn;\u0026thinsp;0.017\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003ePR AUC\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.380\u0026thinsp;\u0026plusmn;\u0026thinsp;0.025\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.405\u0026thinsp;\u0026plusmn;\u0026thinsp;0.035\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.337\u0026thinsp;\u0026plusmn;\u0026thinsp;0.026\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.183\u0026thinsp;\u0026plusmn;\u0026thinsp;0.021\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.123\u0026thinsp;\u0026plusmn;\u0026thinsp;0.029\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.188\u0026thinsp;\u0026plusmn;\u0026thinsp;0.033\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eNull Brier\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.1179\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.1112\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.0779\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.0215\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.0408\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eBrier\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.1001\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0042\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.0544\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0032\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.0966\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0042\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.0723\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0052\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.0199\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0027\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.0365\u0026thinsp;\u0026plusmn;\u0026thinsp;0.0033\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eECE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.0288\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.0061\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.0139\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eCalibration Slope\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.98\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.988\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.943\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.996\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.995\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.996\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eCalibration Intercept\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.288\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026ndash;0.097\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.058\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026ndash;0.008\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026ndash;0.014\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026ndash;0.008\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003ch3\u003eSurvival analysis\u003c/h3\u003e\n\u003cp\u003eFigure \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e shows validation of our predicted PPD against future cause-specific mortality. A 10-percentage-point increase in predicted diabetes, CKD, or CVD PPD was associated with approximately a two-fold higher hazard of disease-specific death (HR range: 2.00\u0026ndash;2.20; p\u0026thinsp;\u0026lt;\u0026thinsp;0.001 for all). When participants were stratified into quartiles of predicted prevalent CVD PPD, 7-year cumulative CVD mortality demonstrated clear monotonic separation (Fig. \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e). Mortality increased from 0.0% (0.0\u0026ndash;0.0) in Q1 to 1.2% (0.8\u0026ndash;1.6) in Q4, representing more than a ten-fold gradient across quartiles (Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n \u003ctable id=\"Tab2\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eHazard ratio and mortality by predicted-PPD quartiles\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eGroup\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eN used\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eEvents\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eHR \u0026dagger; / % 7-year mortality*\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eDiabetes mortality\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e22,734\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e155\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2.00 \u0026dagger; (1.82\u0026ndash;2.20)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCKD mortality\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e22,734\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2.20 \u0026dagger; (1.85\u0026ndash;2.61)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCVD mortality\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e22,734\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e359\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2.13 \u0026dagger; (2.03\u0026ndash;2.24)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eQ1 CVD PPD\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2,686\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0* (0.0\u0026ndash;0.0)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eQ2 CVD PPD\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2,697\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1* (0.0\u0026ndash;0.2)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eQ3 CVD PPD\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2,680\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e52\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.3* (0.1\u0026ndash;0.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eQ4 CVD PPD\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2,671\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e227\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1.2* (0.8\u0026ndash;1.6)\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003ctfoot\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"4\"\u003e\u0026dagger;: Hazard-ratio (95% CI); *: % 7-year mortality (95% CI)\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tfoot\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003cstrong\u003eSensitivity Analyses\u003c/strong\u003e Sensitivity analyses showed that age, BMI, and waist circumference were the strongest predictors across diseases, consistent with partial dependence patterns. Feature-ablation experiments, including the removal or addition of dietary predictors, use of simpler tree ensembles, and alternative imputation strategies, did not materially change discrimination or calibration. Full results are provided in Supplementary Tables S2\u0026ndash;S3 and Supplementary Figure S4.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eWe were able to create a model with comparable discrimination to current leading models, while also being well-calibrated, externally validated, and associated with future cause-specific mortality. It is also a single pipeline that enables simultaneous PPD prediction for three chronic diseases using a single set of general non-invasive predictors. When externally validating our model on 2015 KNHANES data without retraining, discrimination and calibration performance were comparable to the test set on 2017\u0026ndash;2020 holdout test data from NHANES. Further comparison with established models indicates that our unified, three-disease model achieves discrimination on par with the leading noninvasive tools. Our CKD model showed particularly high discrimination, significantly higher than the leading models we analyzed. These results strengthen our claim that our tool can allow for decreased user burden by simultaneously screening for three diseases without losses in performance. Our survival analyses showed significant associations between predicted prevalent-disease probabilities and cause-specific mortality (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e; Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). These associations suggest that the model\u0026rsquo;s outputs do indeed correlate with underlying disease.\u003c/p\u003e\u003cp\u003eWe then investigated how our lack of dietary predictors and our choice of boosted trees over simpler models were affecting model performance. We found that diet-aware models did not show improved calibration or discrimination. This finding is likely attributed to the low consistency and reliability of short-term dietary recall data collected by NHANES, rather than a physiological independence between diet and chronic disease PPD. We also found that simpler bagged tree models were unable to match the discrimination of our boosted tree models. Finally, we found near-identical performance of KNN and mean/mode imputation, which suggests that missing-data handling was not a major driver of model behavior, reducing concern that imputation introduced bias.\u003c/p\u003e\u003cp\u003eIt is important to note, however, that all of our analyses are subject to significant limitations. Our population-level visualizations shown in figures S3 and S5 were generated using an unweighted population from our 2011\u0026ndash;2016 NHANES training data. This population is not representative of the true US population and should not be extrapolated to make conclusions about disease-burdens on a full population. Additionally, although our model was ported to a Korean population (KNHANES), portability to lower-middle-income countries was not analyzed and should be addressed in future work. Another limitation is that our analyses are all association-based, and causal inferences cannot be drawn from our model. Similarly, our model does not indicate longitudinal or incident risk. Because the model is trained on cross-sectional data, predicted probabilities should not be interpreted as forecasts of future incidence or as estimates of the effects of behavioral or clinical interventions. Future work should assess this framework in prospective cohorts, incorporate repeated measurements, or evaluate whether longitudinal changes in predicted probabilities are associated with subsequent differences in long-term outcomes.\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003eStudy population and weighting\u003c/h2\u003e\u003cp\u003eWe used NHANES 2011\u0026ndash;2020 data to train and test our model [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. NHANES 2011\u0026ndash;2016 data was used for training and calibration (n\u0026thinsp;=\u0026thinsp;29903), and 2017\u0026ndash;2020 data was used for testing (n\u0026thinsp;=\u0026thinsp;15,559). Sampling weights were applied for descriptive statistics but not for machine learning model development. We did not apply survey weights during model training because our goal was to learn the individual-level conditional probability P(Y\u0026thinsp;=\u0026thinsp;1∣X), not to estimate population-level prevalence. Moreover, weighting increases population representativeness but doesn\u0026rsquo;t improve predictive accuracy can degrade individual-level risk estimation. Furthermore, NHANES weights can distort model fitting by over-emphasizing subgroups that make up larger proportions of the US population. Unweighted training therefore provides a cleaner estimate of the covariate-outcome relationship, while absolute risk alignment is handled separately through calibration and external validation. Weights for descriptive statistics were calculated according to NHANES guidelines using mobile examination center weights [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Weights were normalized to a mean value of 1.00, which is a standard procedure used in other NHANES-based analyses [e.g. 25, 26].\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003ePredictors and outcome\u003c/h3\u003e\n\u003cp\u003eStandard predictors included demographic variables (age, sex, income-to-poverty ratio), anthropometric measures (body mass index (BMI), waist circumference), and physiologic measures (blood pressure, heart rate). Dietary predictors included dietary sugar, saturated fat, carbohydrate, fiber, polyunsaturated fat, mineral, vitamin, caffeine, theobromine, alcohol, and caloric intake. Diabetes status was defined as glycohemoglobin (HbA1c)\u0026thinsp;\u0026ge;\u0026thinsp;6.5%, consistent with the recommendation of the American Diabetes Association [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. CKD status was determined according to KDIGO 2012 criteria as eGFR\u0026thinsp;\u0026lt;\u0026thinsp;60 mL/min/1.73m\u0026sup2;, using the CKD-EPI equation [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. CVD status was designated as a self-reported history of coronary heart disease, congestive heart failure, angina pectoris, heart attack, or stroke. For external validation on the KNHANES data, CVD was defined as a positive self-reported history of myocardial infarction or angina. Missing predictor values were imputed using training-set means for continuous variables, and modes for categorical variables. Participants missing HbA1c were excluded from diabetes analyses. Those missing any variables needed to compute eGFR (most commonly serum creatinine) were excluded from CKD analyses. Those missing any variables used to define CVD status were excluded from CVD analyses. Predictors in external validation were manually converted to NHANES-coded variables. A variable dictionary and missingness table are shown in tables S4 and S5.\u003c/p\u003e\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003eModel training and evaluation\u003c/h2\u003e\u003cp\u003eAn initial model screening, trained on NHANES 2011\u0026ndash;2016 data and tested on independent 2017\u0026ndash;2020 data, was conducted. Models tested included logistic regression, regularized logistic regression, na\u0026iuml;ve Bayes, k-nearest neighbors, decision trees, random forests, bagged trees, boosted trees (AdaBoost and gradient boosting), and support vector machines with linear and nonlinear kernels. Across these models, tree-based classifiers showed the lowest classification error. The gradient boosting framework XGboost was selected for further development because it achieved the highest discrimination while offering stable performance across resamples, favorable calibration behavior with minimal hyperparameter tuning, robustness to moderate shifts in predictor distributions between NHANES and KNHANES, and clean integration with permutation-based interpretability and partial-dependence analyses. Model hyperparameters are described in the supplementary information, with key parameters listed in table S6.\u003c/p\u003e\u003cp\u003eModel performance was evaluated using sensitivity, specificity, area under the Receiver Operating Characteristic (ROC) curve (AUC), and Brier score. To investigate whether we could exclude dietary variables without sacrificing predictive performance, four models were explored for each disease: a 48-hour-recall diet model (which averages day 1 and day 2 recalls into mean intake variables), a 24-hour-recall diet model, a standard predictor model, and a standard predictor\u0026thinsp;+\u0026thinsp;48-hour dietary recall model. All model thresholds were chosen to maximize Youden\u0026rsquo;s J (i.e. sensitivity\u0026thinsp;+\u0026thinsp;specificity \u0026minus;\u0026thinsp;1). To account for total energy intake, we created energy-adjusted dietary variables (nutrient intake per 1,000 kcal) and repeated diet-only models with these adjusted features. AUC was calculated for each model and was used to compare discrimination performance between models.\u003c/p\u003e\u003cp\u003eIn addition to ROC analysis, we evaluated model performance using precision\u0026ndash;recall (PR) curves and area under the PR curve (PR-AUC), which is more informative under class imbalance. Calibration (reliability) curves were generated using decile bins to assess agreement between predicted and observed PPD. We also computed bootstrap 95% confidence intervals (CIs) for AUC using bootstrap resampling. Null Brier scores for each chronic disease were calculated by creating hypothetical models that assign each individual in the test population a PPD estimate equal to the prevalence in the studied population/subpopulation.\u003c/p\u003e\u003cp\u003eOur external validation was done by curating the 2015 KNHANES dataset such that variable labels and values were matched to NHANES data used to train the models, except for CVD, which could only be partially matched, as discussed above [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. We then compared our model to other external models, selected via a literature search for non-invasive diabetes, CKD, and CVD PPD models, prioritizing varied recent and well-established models (e.g. FINDRISC, INTERHEART).\u003c/p\u003e\u003cp\u003e\u003cb\u003ePartial-dependence analyses\u003c/b\u003e To examine how predicted probabilities vary with individual predictors, we computed partial-dependence sensitivity analyses using the observed NHANES training sample. For each grid value, we replaced the predictor value for all individuals in the training set while leaving all other predictors unchanged, passed the modified dataset through the trained model, and calculated the mean predicted probability for each disease. This yields a one-dimensional partial-dependence curve that summarizes how the model\u0026rsquo;s predictions vary with that predictor, averaged over the joint distribution of all other predictors in the training sample.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\u003ch2\u003eCox Models, Discrimination, and Survival Across Predicted PPD Quartiles\u003c/h2\u003e\u003cp\u003eCause-specific mortality was obtained from the 2019 NHANES-linked National Death Index files. Follow-up time was calculated from the MEC examination using PERMTH_INT (months) and converted to years. Participants outside the 2011\u0026ndash;2016 survey cycles or missing survival time, cause-of-death indicators, or predicted-PPD estimates were excluded. To ensure meaningful exposure windows for chronic-disease mortality and to match established NHANES cardiovascular mortality analyses, we restricted the survival analyses to adults aged\u0026thinsp;\u0026ge;\u0026thinsp;40 years. Event indicators were defined for diabetes, CKD, and cardiovascular mortality using the underlying-cause classification provided in the linked files.\u003c/p\u003e\u003cp\u003eFor each mortality endpoint, we fit pooled Cox proportional hazards models across all eligible NHANES cycles, including cohort indicators and scaling predicted PPDs so hazard ratios represent the effect of a 10-percentage-point increase in model-predicted PPD. Cox coefficients were estimated using a numerically stable Breslow partial-likelihood implementation with ridge penalization, which improves convergence for sparse endpoints (e.g., CKD mortality). To evaluate PPD stratification, predicted CVD PPD was divided into quartiles, and 7-year cumulative CVD mortality was estimated within each quartile using Kaplan\u0026ndash;Meier curves and bootstrap-derived confidence intervals. Similar analyses were not conducted for diabetes or CVD due to insufficient cause-specific mortality.\u003c/p\u003e\u003c/div\u003e"},{"header":"Abbreviations ","content":"\u003cp\u003ePPD, probability of prevalent disease; NCD, noncommunicable disease; DM, diabetes mellitus, CVD, cardiovascular disease; CKD, chronic kidney disease; NHANES, National Health and Nutrition Examination Survey; CI, confidence interval; \u0026nbsp;AUC, area under curve; ROC, receiver operating characteristic; BMI, body mass index; HbA1c, glycohemoglobin; KDIGO, Kidney Disease Improving Global Outcomes; PR, precision-recall; PR-AUC, area under the precision-recall curve; eGFR, estimated glomerular filtration rate; CDC, Centers for Disease Control; PDP, partial dependence plot\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgments:\u003c/strong\u003e The authors thank Gabriela Gomez, Dr. Ismael Haddadian, and Dr. Sandra C. Fuchs for helpful feedback.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions:\u003c/strong\u003e A.M.C completed the conceptualization, methodology, and formal analysis. Investigation and data curation was completed by A.M.C and I.B.D. The first draft of the manuscript was written and edited by A.M.C and I.B.D. All authors have read and agreed to the published version of the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests:\u0026nbsp;\u003c/strong\u003eA.M.C. and I.B.D. are inventors on a provisional patent application related to the methods described in this manuscript. The authors declare no other competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics statement:\u0026nbsp;\u003c/strong\u003eNHANES and KNHANES are publicly available, deidentified national health surveys that obtain informed consent from all participants and are approved by their respective institutional review boards. This secondary analysis of fully anonymized public data was exempt from institutional review board review.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eInformed Consent Statement:\u0026nbsp;\u003c/strong\u003eAll participants in NHANES and KNHANES provided written informed consent.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability:\u0026nbsp;\u003c/strong\u003eAll data used in this study are publicly accessible. NHANES 2011\u0026ndash;2020 data can be obtained from the U.S. Centers for Disease Control and Prevention (https://www.cdc.gov/nchs/nhanes). KNHANES 2015 microdata are available from the Korea Disease Control and Prevention Agency via the KNHANES website (https://knhanes.kdca.go.kr) following free online registration; the authors do not redistribute these files. Mortality data were obtained from the NHANES-linked National Death Index. No proprietary or restricted-access datasets were used.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode availability:\u0026nbsp;\u003c/strong\u003eAll analysis code used in this study are available in repositories linked in the supplementary information.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u003c/strong\u003e This research received no external funding.\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eCenters for Disease Control and Prevention. Diabetes: Data \u0026amp; Research. Atlanta, GA: CDC. (accessed 4 Sept 2025).\u003c/li\u003e\n \u003cli\u003eCenters for Disease Control and Prevention. Chronic Kidney Disease: Data \u0026amp; Research (Fast Stats). Atlanta, GA: CDC. (accessed 4 Sept 2025).\u003c/li\u003e\n \u003cli\u003eMartin SS, et al. Heart Disease and Stroke Statistics\u0026mdash;2024 Update: A Report From the American Heart Association. Circulation. 2024;149:e347\u0026ndash;e913.\u003c/li\u003e\n \u003cli\u003ePoslusna K, et al. Misreporting of Energy and Micronutrient Intake Estimated by Food Records and 24 h Recalls, Control and Adjustment Methods in Practice. Br J Nutr. 2009;101(S2):S73\u0026ndash;S85.\u003c/li\u003e\n \u003cli\u003eAhluwalia N, et al. Update on NHANES Dietary Data: Focus on Collection, Release, Analytical Considerations, and Uses to Inform Public Policy. Adv Nutr. 2016;7:121\u0026ndash;134.\u003c/li\u003e\n \u003cli\u003eVan Calster B, et al. Calibration: The Achilles Heel of Predictive Analytics. BMC Med. 2019;17:230.\u003c/li\u003e\n \u003cli\u003eLindstr\u0026ouml;m J, Tuomilehto J. The Finnish Diabetes Risk Score (FINDRISC): A Valid Tool to Estimate the Risk of Type 2 Diabetes. Diabetes Care. 2003;26(3):725\u0026ndash;731.\u003c/li\u003e\n \u003cli\u003eD\u0026rsquo;Agostino RB Sr, et al. General Cardiovascular Risk Profile for Use in Primary Care: The Framingham Heart Study. Circulation. 2008;117:743\u0026ndash;753.\u003c/li\u003e\n \u003cli\u003eCDC/NCHS. Health, United States\u0026mdash;Sources and Definitions: National Health and Nutrition Examination Survey (NHANES). (accessed 4 Sept 2025).\u003c/li\u003e\n \u003cli\u003eZhou X, Qiao Q, Ji L, et al. Non-laboratory-based risk assessment algorithm for prevalent type 2 diabetes developed on a nation-wide diabetes survey. Diabetes Care. 2013;36(12):3944-3952.\u003c/li\u003e\n \u003cli\u003eAbdallah L, Naja F, Choufani J, et al. Validation of the Finnish Diabetes Risk Score (FINDRISC) in the Lebanese population. Prim Care Diabetes. 2020;14(3):271-277.\u003c/li\u003e\n \u003cli\u003eMugume I, Kibirige D, Ssebunya R, et al. Performance of the Finnish Diabetes Risk Score in detecting prevalent type 2 diabetes and dysglycaemia in a rural Ugandan population. PLOS One. 2023;18(11):e0276858.\u003c/li\u003e\n \u003cli\u003eHa KH, Kim DJ. Development and validation of the Korean Diabetes Risk Score: a prospective study. Diabetologia. 2018;61(4):949-958.\u003c/li\u003e\n \u003cli\u003eZhou X, Li W, Wong CKH, et al. Simple non-laboratory and laboratory-based risk assessment algorithms for prevalent diabetes in a Chinese population. J Diabetes Investig. 2016;7(5):657-666.\u003c/li\u003e\n \u003cli\u003eBang H, Vupputuri S, Shoham DA, et al. SCreening for Occult REnal Disease (SCORED): a simple prediction model for chronic kidney disease. Arch Intern Med. 2007;167(4):374-381.\u003c/li\u003e\n \u003cli\u003eO\u0026rsquo;Seaghdha CM, Lyass A, Massaro JM, et al. A risk score for chronic kidney disease in the general population. Am J Med. 2012;125(3):270-277.\u003c/li\u003e\n \u003cli\u003eTangri N, Stevens LA, Griffith J, et al. A predictive model for progression of chronic kidney disease to kidney failure. JAMA. 2011;305(15):1553-1559.\u003c/li\u003e\n \u003cli\u003eGaziano TA, Young CR, Fitzmaurice G, Atwood S, Gaziano JM. Laboratory-based versus non-laboratory-based method for assessment of cardiovascular disease risk: the NHANES I Follow-up Study cohort. Lancet. 2008;371(9616):923-931.\u003c/li\u003e\n \u003cli\u003ePandya A, Weinstein MC, Gaziano TA. A comparative assessment of non-laboratory-based versus commonly used laboratory-based cardiovascular disease risk scores in the NHANES III population. PLOS One. 2011;6(5):e20416.\u003c/li\u003e\n \u003cli\u003eMcGorrian C, Yusuf S, Islam S, et al. Estimating modifiable coronary heart disease risk in multiple regions of the world: the INTERHEART Modifiable Risk Score. Eur Heart J. 2011;32(5):581-589.\u003c/li\u003e\n \u003cli\u003eKaptoge S, Pennells L, De Bacquer D, et al. World Health Organization cardiovascular disease risk charts: revised models to estimate risk in 21 global regions. Lancet Glob Health. 2019;7(10):e1332-e1345.\u003c/li\u003e\n \u003cli\u003eRezaei F, Yadegarfar G, Gohari K, et al. Agreement between laboratory-based and non-laboratory-based Framingham risk scores in a large Middle Eastern cohort. Sci Rep. 2021;11:10687.\u003c/li\u003e\n \u003cli\u003eJia, G.; Aroor, A.R.; Sowers, J.R. Arterial Stiffness: A Nexus between Cardiac and Renal Disease. Cardiorenal Med. 2014, 4, 60\u0026ndash;71. https://doi.org/10.1159/000360867.\u003c/li\u003e\n \u003cli\u003eUniversity of Texas at Arlington. National Health and Nutrition Examination Survey (NHANES)\u0026mdash;Big Data for Epidemiology. Available online: https://uta.pressbooks.pub/bigdataforepidemiology/chapter/chapter10-nhanes/ (accessed on 4 September 2025).\u003c/li\u003e\n \u003cli\u003eCosta, A.M.; Sias, R.J.; Fuchs, S.C. Effect of Whole Blood Dietary Mineral Concentrations on Erythrocytes: Selenium, Manganese, and Chromium: NHANES Data. Nutrients 2024, 16, 3653. https://doi.org/10.3390/nu16213653.\u003c/li\u003e\n \u003cli\u003eChen, J.; Kan, M.; Ratnasekera, P.; Deol, L.K.; Thakkar, V.; Davison, K.M. Blood Chromium Levels and Their Association with Cardiovascular Diseases, Diabetes, and Depression: NHANES 2015\u0026ndash;2016. Nutrients 2022, 14, 2687. https://doi.org/10.3390/nu14132687.\u003c/li\u003e\n \u003cli\u003eAmerican Diabetes Association. Diabetes Diagnosis. Available online: https://diabetes.org/about-diabetes/diagnosis (accessed on 4 September 2025).\u003c/li\u003e\n \u003cli\u003eWorld Health Organization. Noncommunicable Diseases \u0026ndash; Fact Sheet. Geneva: WHO. (accessed 4 Sept 2025).\u003c/li\u003e\n \u003cli\u003eKweon, Sanghui et al. \u0026quot;Data Resource Profile: The Korea National Health and Nutrition Examination Survey (KNHANES)\u0026quot; vol. 43, no. 1, 2014\u0026nbsp;\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Diabetes Mellitus, Cardiovascular Diseases, Renal Insufficiency, Chronic, Machine Learning, Diet, Prevalent Disease, Public Health","lastPublishedDoi":"10.21203/rs.3.rs-8311243/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8311243/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eReliable non-laboratory tools for assessment of probability of prevalent disease (PPD) are essential for scalable prevention, yet existing models are typically specific to single diseases, require laboratory tests, and show no or limited calibration across PPD strata, limiting scaling and public health utilization. We developed and validated a unified, non-invasive machine-learning model for simultaneous prediction of diabetes, chronic kidney disease, and cardiovascular disease PPD non-invasive predictors. The model was trained on 2011\u0026ndash;2016 National Health and Nutrition Examination Survey data (n\u0026thinsp;=\u0026thinsp;29,903) and evaluated on an independent 2017\u0026ndash;2020 test set (n\u0026thinsp;=\u0026thinsp;15,559). It demonstrated moderate-to-strong discrimination (C-statistic\u0026thinsp;=\u0026thinsp;0.80\u0026ndash;0.90), stable precision\u0026ndash;recall performance, and moderate to strong calibration (slope\u0026thinsp;\u0026gt;\u0026thinsp;0.94). Validation in an independent Korean population showed no or minimal degradation in discrimination and calibration performance, though more extensive validation is warranted. Predicted PPD was associated with cause-specific mortality over up to 7 years of follow-up, consistent with a predictor of latent disease burden. Each 10-percentage-point increase in predicted PPD was associated with roughly a two-fold higher hazard of disease-specific death (HR 2.00\u0026ndash;2.20). We conclude that this model has potential as scalable, low-burden screening/surveillance aid, but note that it is not intended as a diagnostic or prognostic tool.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e","manuscriptTitle":"Multinational, Calibrated, Non-Laboratory Prevalent Disease Prediction and Survival Modeling for Diabetes, CKD, and CVD","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-12-10 07:00:19","doi":"10.21203/rs.3.rs-8311243/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"ee13577d-bf0f-49fe-b44d-5ba597fbe5af","owner":[],"postedDate":"December 10th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":59324158,"name":"Health sciences/Biomarkers"},{"id":59324159,"name":"Health sciences/Diseases"},{"id":59324160,"name":"Health sciences/Endocrinology"},{"id":59324161,"name":"Health sciences/Health care"},{"id":59324162,"name":"Health sciences/Medical research"},{"id":59324163,"name":"Health sciences/Nephrology"},{"id":59324164,"name":"Health sciences/Risk factors"}],"tags":[],"updatedAt":"2026-04-08T09:42:36+00:00","versionOfRecord":[],"versionCreatedAt":"2025-12-10 07:00:19","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8311243","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8311243","identity":"rs-8311243","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-06-05T02:00:03.366016+00:00
License: CC-BY-4.0