{"paper_id":"0d17209f-fb7a-4de5-87e5-d8900281fcfe","body_text":"Early prediction of severity progression in patients with chronic kidney disease: A Machine Learning Predictive Modelling analysis with retrospective data of a tertiary care hospital | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Early prediction of severity progression in patients with chronic kidney disease: A Machine Learning Predictive Modelling analysis with retrospective data of a tertiary care hospital Saurav Nayak, Gautom Kumar Saharia, Sandip Kumar Panda, Manaswini Mangaraj This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7198292/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background Chronic Kidney Disease (CKD) represents a growing health burden, particularly in low- and middle-income countries. Its progression to end-stage renal disease (ESRD) necessitates resource-intensive interventions. The inappropriate activation of the complement system, particularly Complement 3 (C3) and Complement 4 (C4), has been implicated in renal injury. While these markers are biologically relevant, their predictive value for CKD severity has not been adequately explored. This study aimed to assess the potential of serum C3 and C4 in predicting CKD severity using machine learning (ML) models. Methods A retrospective dataset comprising 2,279 adults (> 17 years) was extracted from the laboratory records of AIIMS Bhubaneswar. CKD severity was classified using the 2021 CKD-EPI Creatinine formula, with Stage G3b and above defined as severe CKD. Of these, 1,331 complete records (with C3, C4, and creatinine) formed the internal dataset, while two external datasets were used for validation—one with clinically confirmed CKD and the other age- and sex-matched controls. Predictive models were developed using five ML algorithms: Random Forest (RDF), XGBoost (XGB), Gradient Boosting (GB), Decision Tree (DT), and Artificial Neural Networks (ANN). Model performance was evaluated using accuracy, F1 score, R², diagnostic odds ratio (DOR), and other metrics. Results Recursive Feature Elimination identified age, C3, and the C3/C4 ratio as the most influential predictors. Among the models, RDF performed best (F1: 0.984, Accuracy: 0.991, DOR: 9348, R²: 0.954). External validation confirmed its high diagnostic power (DOR: 494 and 465 for CKD and control datasets, respectively). A web-based tool was developed to aid clinicians in estimating CKD severity using age, C3, and C4 values. Conclusion This study demonstrates that serum complements C3 and C4 can serve as early predictive biomarkers for CKD severity when interpreted via machine learning models. The RDF-based prediction pipeline offers a clinically relevant, non-invasive tool for stratifying CKD patients, potentially reducing the burden of late-stage interventions. Further prospective studies are warranted to validate these findings longitudinally. Figures Figure 1 Figure 2 Figure 3 Introduction Chronic kidney disease (CKD) which is defined as a glomerular filtration rate (GFR) of < 60mL/min/1.73 m 2 , is a significant public health issue worldwide. 1 Previously thought to be a health issue primarily in economically prosperous nations, 4 out of every 5 chronic illness deaths now occur in low- and middle-income countries. 2 In the Indian population, the actual illness burden of CKD/ESRD cannot be accurately assessed due to the lack of a renal registry; however, the number of fatalities attributable to chronic diseases has increased from 3.78 million in 1990 (40.4% of all deaths) to 7.63 million in 2020 (66.7% of all deaths). CKD has a strong connection to other major diseases, such as cardiovascular disease and diabetes, which continue to be the primary causes of mortality and morbidity in the Indian population. 3 Worldwide, a significant economic burden is associated with CKD, particularly once kidney failure occurs. This is because of the need for resource-intensive Kidney Replacement Therapy (KRT) resulting in a cost escalation as well as a significant societal burden due to the requirement of either dialysis or kidney transplantation. With the increasing prevalence of CKD along with the huge scope for improvement in the management for CKD, effective solutions are required to help healthcare decision-makers. 4 , 5 The inappropriate activation of the complement system has been implicated in kidney disease. 6 Complement C3 serves as a critical hub in activating the complement cascade. 7 Aside from its role in the complement cascade and inflammation, multiple observational studies have found that high C3 levels in plasma are related to the development of hypertension and diabetes, the two predominant causes of CKD. 4 , 8 Complement activation at the glomerular and systemic levels is critical in the development and clinical manifestations of renal diseases. 9 To assess the degree of complement activation, serum complement levels are used as surrogate markers, with Complement 3 (C3) and Complement 4 (C4) being the most commonly reviewed measures in clinical practice. 10 There is sufficient data in clinical settings to establish the role of complement activation in renal injury that exhibits altered circulating C3 levels, renal C3 deposits, and C3 genetic abnormalities. 11 Machine Learning and Artificial Intelligence models are efficient in finding patterns that are inconspicuous to normal statistical methods. The prediction accuracy can be improved in comparison to conventional biomarkers, by machine learning algorithms as it can learn complex and nonlinear interactions. 12 , 13 Early detection of CKD severity will, therefore, be a vital tool for the prevention and management of associated morbidity and mortality. Therefore, in this study, we will be utilising the serum complements C3 and C4 to predict severity in CKD patients. Methodology The ethical clearance for the study was obtained from the Institutional Ethical Committee of AIIMS Bhubaneswar vide approval letter no. T/IM-NF/Biochem/23/162 dated 29 Jan 2024. The study was carried out in the Immunology Laboratory under the Department of Biochemistry by revisiting laboratory e-records in conjunction with the Department of Nephrology. Serum Creatinine and Complements were measured in the Beckman AU Autoanalyzer. Estimation of eGFR and Classification Estimated GFR (eGFR) was calculated based on the 2021 CKD-EPI Creatinine formula. 14 The formula is eGFR = 142 × min(S cr /κ, 1) α × max(S cr /κ, 1) −1.200 × 0.9938 Age × 1.012 (if female) where κ = 0.7 (females) or 0.9 (males); an α = -0.241 (females) or -0.302 (males); S cr = serum creatinine in mg/dL; Age = in years The value of eGFR was rounded to the nearest whole number. The classification of kidney disease based on the Kidney Disease: Improving Global Outcomes (KDIGO) Classification based on eGFR has been tabulated in Table 1 . 15 Table 1 KDIGO Classification of Kidney Disease based on eGFR CKD Stage Description eGFR Range (mL/min/1.73 m²) G1 Normal or high ≥ 90 G2 Mildly decreased 60–89 G3a Mildly to moderately decreased 45–59 G3b Moderately to severely decreased 30–44 G4 Severely decreased 15–29 G5 Kidney failure < 15 Based on this classification, severity was determined to be Severe if the CKD Stage is G3b or above. Rest was considered to be non-severe. Patient and Public Engagement Summary The data involved in the study were acquired from the AIIMS Bhubaneswar Laboratory CDAC e-Records retrospectively over a period spanning 24 months. Only those unique data points for which age was greater than 17 years of age were selected for which all Serum Creatinine, C3, and C4 values existed. Any stray errors, blanks or other error-ridden data were duly excluded. Thus, each patient represented one datapoint, totalling to 2279, out of which 238 data points were clinically diagnosed with CKD by a Nephrologist. All data were stored electronically and anonymously. Dataset Separation : The dataset was divided into 3 datasets: 1 dataset to be used as an Internal Dataset and 2 External Datasets. The first external dataset comprised all CKD-diagnosed patients, and the other dataset comprised a 3:1 age- and sex-matched control group for whom the calculated CKD classification was G1. The internal dataset had overall data points that had all 3 parameters, i.e., C3, C4 and Serum Creatinine and was used for predictive model generation. Machine Learning Models : The internal dataset was deployed to generate predictive models based on 5 machine learning methods: Random Forest Classifier (RDF), XGBoost Classifier (XGB), Gradient-Boosting Classifier (GB), Decision Tree (DT) and Artificial Neural Network (ANN). Recursive Feature Elimination (RFE) was employed to show the feature importance of 5 parameters: Age, Sex, C3, C4 and C3/C4 ratio. The target is “Severity Classification”. The model is generated and evaluated by a 70 − 30 train-test split and validated by predictions made to the external datasets. Evaluation Metrics The models are evaluated based on Accuracy, Precision, Recall, F1 Score, MCC, AUC, Sensitivity, Specificity, PPV, NPV, LR+, LR-, DOR, R 2 . TP = True Positive. TN = True Negative. FP = False Positive. FN = False Negative. Accuracy = (TP + TN) / (TP + TN + FP + FN). Accuracy is the ratio of correctly predicted instances (both true positives and true negatives) to the total instances. It is a general measure of how well the model is performing. Precision or Positive Predictive Value (PPV) = (TP) / (TP + FP). Precision is the ratio of true positive instances to the sum of true positive and false positive instances. It indicates how many of the predicted positive instances are actually positive. Recall or Sensitivity or True Positive Rate (TPR) = (TP) / (TP + FN). Recall is the ratio of true positive instances to the sum of true positive and false negative instances. It measures the model's ability to identify positive instances correctly. F1 Score = 2 x [(Precision x Recall) / (Precision + Recall)] = The F1 Score is the harmonic mean of precision and recall. It provides a balance between the two metrics and is useful when you need to account for both false positives and false negatives. Matthews Correlation Coefficient (MCC) = [(TP × TN)-(FP × FN)] / {√[(TP + FP) × (TP + FN) × (TN + FP) × (TN + FN)]} MCC is a measure of the quality of binary classifications, considering all four confusion matrix categories. It is a balanced measure and is considered more informative than accurate in imbalanced datasets. Specificity = (TN) / (TN + FP). Specificity measures the proportion of actual negatives that are correctly identified. Negative Predictive Value (NPV) = (TN) / (TN + FN). NPV measures the proportion of negative results that are true negatives. Positive Likelihood Ratio (LR+) = (Sensitivity) / (1 – Specificity). LR + indicates how much the odds of the disease increase when a test is positive. A higher LR + value suggests a more useful test. Negative Likelihood Ratio (LR-) = (1 – Sensitivity) / (Specificity). LR- indicates how much the odds of the disease decrease when a test is negative. A lower LR value suggests a more useful test. Diagnostic Odds Ratio (DOR) = LR+ / LR-. DOR is the ratio of the odds of positivity in diseased to non-diseased. It provides a single metric to evaluate the performance of a diagnostic test. Coefficient of Determination (R 2 ) = 1 – [(Sum of Residuals) / (Total Sum of Squares)]. R² measures the proportion of the variance in the dependent variable that is predictable from the independent variable(s). It ranges from 0 to 1, where 1 indicates that the regression predictions perfectly fit the data. The decision on the model was made based on its efficiency. Pipeline for Prediction The most effective model, as per the evaluation metrics, will be utilised to create a web-based platform to predict CKD severity based on the provided Complement values. Results The Internal dataset had 1,331 defined data points of various stages of classification based on their eGFR values. The External – CKD dataset had 237 data points, and the External – Controls dataset had 711 data points defined in them. The classification parameters of these 3 datasets have been tabulated in Table 2 . Table 2 Class distribution in the analysed datasets. Dataset Total Datapoints Classification Severity G1 n (%) G2 n (%) G3a n (%) G3b n (%) G4 n (%) G5 n (%) Non-Severe n (%) Severe n (%) Internal 1,331 595 (44.7) 274 (20.6) 81 (6.1) 94 (7.1) 105 (7.9) 181 (13.6) 951 (71.4) 380 (28.6) External – CKD 237 23 (90.7) 20 (8.4) 20 (8.4) 25 (10.5) 44 (18.6) 105 (44.3) 63 (26.6) 174 (73.4) External - Controls 711 711 (100) - - - - - 711 (100) - RFE showed that the most important feature in determination was age (0.27), followed by C3 (0.24), C3/C4 (0.22) ratio and C4(0.21). Sex (0.06) was the least important parameter, with a normalised importance score of less than 25%. The Normalized Feature importances have been detailed in Fig. 1 . The difference in the parameter levels has been reported in Table 3 . Table 3 Comparison of Features between Severe and Non-Severe groups Parameter Severe [N = 380] Non-Severe [N = 951] p-value * Age (in years) 37 (28–47) 36 (25–46) 0.526 C3 (in mg/dL) 1.1 (0.92–1.29) 1.2 (0.76–1.48) 0.427 C4 (in mg/dL) 1.17 (0.8–1.37) 1.19 (0.84–1.43) 0.705 C3/C4 Ratio 0.92 (0.74–1.38) 1.01 (0.65–1.49) 0.397 *: Mann Whitney U test The models were generated using the four parameters: Age, C3, C4, and the calculated parameter C3/C4 ratio. The final decision of model selection was based on the evaluation metrics, primarily focused on the F1 Score, R 2 value and DOR. Based on these, the best models were the ones defined by RDF, XGB, and DT. The evaluation metrics on the test set of the Internal dataset have been tabulated in Table 4 . Table 4 Evaluation metrics of the Test Set (30%) of the Internal Dataset Metrics RDF XGB GB DT ANN Accuracy 0.991 0.961 0.792 0.990 0.725 Precision 0.983 0.953 0.753 0.981 0.535 Recall 0.984 0.907 0.389 0.984 0.195 F1 Score 0.984 0.930 0.513 0.983 0.286 MCC 0.977 0.903 0.432 0.976 0.191 AUC 0.999 0.989 0.843 0.988 0.742 Sensitivity 0.984 0.907 0.389 0.984 0.195 Specificity 0.993 0.983 0.950 0.992 0.934 PPV 0.983 0.953 0.753 0.981 0.535 NPV 0.994 0.963 0.798 0.994 0.747 LR+ 148 53 8 131 3 LR- 0.02 0.10 0.64 0.02 0.86 DOR 9348 553 12 8241 3 R 2 0.954 0.809 -0.029 0.951 -0.357 As the RDF model was determined to be the best model, it was applied to the two external datasets, and performance evaluation was measured in the form of evaluation metrics. The model predicted both the CKD and non-CKD datasets with very robust and high accuracy and diagnostic power. The DOR was 494 and 465 respectively. The metrics are shown in Fig. 2 . To aid the clinicians, we went a step further to create an ML model-based pipeline where the values for age, C3, and C4 can be entered, and the model provides the probability of severity as an outcome. The same can be accessed at https://bit.ly/c3c4ckdsevereity . A screenshot of the webpage is shown in Fig. 3 . Discussion Our study aimed to understand the early predictive power of complements in relation to the severity of CKD patients. This is one of the first studies that aims to understand the role of C3 and C4 through machine learning models. In this regard, it was elucidated that complements have a very accurate and emerging role in assessing the severity of CKD in patients. Showing very sensitivity, specificity and, high accuracy, these models provide a great deal of insight into predictive power. The ML RDF model aids in analysing the non-linear relationships between the various parameters and thus captures with high accuracy the prediction of CKD as well as its potential in differentiating severity in both CKD and Non-CKD Datasets. There is a scarcity of articles that have exclusively dealt with the participation of complements in CKD. Both the kidneys and abdominal fat tissue exhibit high levels of C3 expression. Notably, the kidney itself plays a major role in circulating C3 and is a large source of extrahepatic C3. In response to damage or cellular stress, the production of C3 is upregulated in podocytes, which are essential for maintaining the integrity of the glomerular filtration barrier. 16 , 17 The degree of kidney tissue damage is directly correlated with local C3 generation, most likely through increased T and B-cell function, the recruitment of pro-inflammatory and pro-fibrotic cytokines, the elimination of immune complexes and apoptotic cells, etc. Due to its in part renal origin, plasma C3 may be a sign of chronic kidney cell damage, which leads to chronic kidney disease (CKD) . 11 Higher levels of circulating C3 have been reported to be associated with cardiovascular diseases in patients with renal disorders, thus adding another requirement to the need for assessing severity as early as possible. 18 Similar associations, though scanty, also exist for C4 and C3/C4 ratios. 19 , 20 A study by Liu et al. primarily suggested that the progression of CKD and mortality is related to an increase in C4 levels in serum. 21 Also, an association of C3, C4, as well as the C3/C4 ratio, has been reported to aggravate renal dysfunction into End Stage Renal Diseases (ESRD), thus affecting morbidity and mortality in patients. 22 Previous machine learning models have been utilised to predict the progression of CKD in patients, and they have been done with high accuracy by random forest classification. 23 Early prediction of CKD has also been reported based on clinical and physiological data by Islam et al. in a similar study. 24 These factors underline the strength of our study, as it is one of the first of its kind, and it tries to assess severity predictions for patients, which will drastically reduce the burden both clinically and financially. This also comes against the backdrop of a rise in kidney diseases in India, as well as rapid progression of severity due to lax treatment and patient awareness. Notwithstanding, our study was limited by the scope of the quantity of data points available as patients of CKD were very few in our centre who had been advised for a Complement profile. As we advance, a longitudinal study with followed-up measurement of complements, along with treatment, aimed at monitoring eGFR levels and disease progression can aid in better predictive modelling. Conclusion Machine learning predictive modelling can assess with high accuracy the severity progression in early CKD, thus preventing morbidity and mortality in patients. The online predictive pipeline will aid clinicians in determining severity as well, providing a holistic approach to the management of CKD in such cases. Declarations Ethics approval and consent to participate The ethical clearance for the study was obtained from the Institutional Ethical Committee of AIIMS Bhubaneswar vide approval letter no. T/IM-NF/Biochem/23/162 dated 29 Jan 2024. This has been done in accordance with the Declaration of Helsinki. In accordance to the IEC waiver for retrospective study, there was no requirement for Informed Consent Form for the participants. Consent for publication Not Applicable as Retrospective Study. Waiver received from IEC. Availability of data and materials The datasets used and/or analysed during the current study are available from the corresponding author on reasonable request. Competing interests The authors declare that they have no competing interests Funding No funding was received. Authors’ Contributions SN and GK conceptualized the study and the methodology. SN developed the algorithm for model development and interpreted the data. SKP provided expert opinion on CKD cases and their diagnosis and determination. MM and GKS drafted the manuscript and proofread it. All authors read and approved the final manuscript. References Liyanage T, Toyama T, Hockham C, Ninomiya T, Perkovic V, Woodward M, et al. Prevalence of chronic kidney disease in Asia: a systematic review and analysis. BMJ Glob Health. 2022;7(1):e007525. Lozano R, Naghavi M, Foreman K, Lim S, Shibuya K, Aboyans V, et al. Global and regional mortality from 235 causes of death for 20 age groups in 1990 and 2010: a systematic analysis for the Global Burden of Disease Study 2010. Lancet. 2012;380(9859):2095–128. Agarwal SK, Srivastava RK. Chronic Kidney Disease in India: Challenges and Solutions. Nephron Clin Pract. 2009;111(3):c197–203. Jha V, Ur-Rashid H, Agarwal SK, Akhtar SF, Kafle RK, Sheriff R. The state of nephrology in South Asia. Kidney Int. 2019;95(1):31–7. Elshahat S, Cockwell P, Maxwell AP, Griffin M, O’Brien T, O’Neill C. The impact of chronic kidney disease on developed countries from a health economics perspective: A systematic scoping review. Barretti P, editor. PLOS ONE. 2020;15(3):e0230512. Kaartinen K, Safa A, Kotha S, Ratti G, Meri S. Complement dysregulation in glomerulonephritis. Semin Immunol. 2019;45:101331. Bao X, Borné Y, Muhammad IF, Schulz CA, Persson M, Orho-Melander M, et al. Complement C3 and incident hospitalization due to chronic kidney disease: a population-based cohort study. BMC Nephrol. 2019;20(1):61. Wlazlo N, Van Greevenbroek MMJ, Ferreira I, Feskens EJM, Van Der Kallen CJH, Schalkwijk CG, et al. Complement Factor 3 Is Associated With Insulin Resistance and With Incident Type 2 Diabetes Over a 7-Year Follow-up Period: The CODAM Study. Diabetes Care. 2014;37(7):1900–9. Lhotta K, Schlogl A, Kronenberg F, Joannidis M, Konig P. Glomerular deposition of the complement C4 isotypes C4A and C4B in glomeruonephritis. Nephrol Dial Transplant Off Publ Eur Dial Transpl Assoc -. Eur Ren Assoc. 1996;11(6):1024–8. Pan M, Zhou Q, Zheng S, You X, Li D, Zhang J, et al. Serum C3/C4 ratio is a novel predictor of renal prognosis in patients with IgA nephropathy: a retrospective study. Immunol Res. 2018;66(3):381–91. Thurman JM. Complement in Kidney Disease: Core Curriculum 2015. Am J Kidney Dis. 2015;65(1):156–68. Kumari S, Tripathy S, Nayak S, Rajasimman AS. Machine learning–aided algorithm design for prediction of severity from clinical, demographic, biochemical and immunological parameters: Our COVID-19 experience from the pandemic. J Fam Med Prim Care. 2024;13(5):1937–43. Nayak S, Singh A, Mangaraj M, Saharia GK. Predicting immune risk in treatment-naïve HIV patients using a machine learning algorithm: a decision tree algorithm based on micronutrients and inversion of the CD4/CD8 ratio. Front Nutr. 2024;11:1443076. Miller WG, Kaufman HW, Levey AS, Straseski JA, Wilhelms KW, Yu HY, Elsie. National Kidney Foundation Laboratory Engagement Working Group Recommendations for Implementing the CKD-EPI 2021 Race-Free Equations for Estimated Glomerular Filtration Rate: Practical Guidance for Clinical Laboratories. Clin Chem. 2022;68(4):511–20. Chen TK, Knicely DH, Grams ME. Chronic Kidney Disease Diagnosis and Management: A Review. JAMA. 2019;322(13):1294. Zhou W, Marsh JE, Sacks SH. Intrarenal synthesis of complement. Kidney Int. 2001;59(4):1227–35. Tang S, Zhou W, Sheerin NS, Vaughan RW, Sacks SH. Contribution of renal secreted complement C3 to the circulating pool in humans. J Immunol Baltim Md. 1950. 1999;162(7):4336–41. Lines SW, Richardson VR, Thomas B, Dunn EJ, Wright MJ, Carter AM. Complement and Cardiovascular Disease - The Missing Link in Haemodialysis Patients? Nephron. 2016;132(1):5–14. Xing Z, Wang Y, Gong K, Chen Y. Plasma C4 level was associated with mortality, cardiovascular and cerebrovascular complications in hemodialysis patients. BMC Nephrol. 2022;23(1):232. Zhang Y, Duan SW, Chen P, Yin Z, Wang Y, Cai GY et al. Relationship between serum C3/C4 ratio and prognosis of immunoglobulin A nephropathy based on propensity score matching. Chin Med J (Engl). 2020;(6):631–7. Liu J, Zha Y, Zhang P, He P, He L. The Association Between Serum Complement 4 and Kidney Disease Progression in Idiopathic Membranous Nephropathy: A Multicenter Retrospective Cohort Study. Front Immunol. 2022;13:896654. Matsuda S, Oe K, Kotani T, Okazaki A, Kiboshi T, Suzuka T, et al. Serum Complement C4 Levels Are a Useful Biomarker for Predicting End-Stage Renal Disease in Microscopic Polyangiitis. Int J Mol Sci. 2023;24(19):14436. Ferguson T, Ravani P, Sood MM, Clarke A, Komenda P, Rigatto C, et al. Development and External Validation of a Machine Learning Model for Progression of CKD. Kidney Int Rep. 2022;7(8):1772–81. Islam MA, Hussein MdA. Chronic kidney disease prediction based on machine learning algorithms. J Pathol Inf. 2023;14:100189. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {\"props\":{\"pageProps\":{\"initialData\":{\"identity\":\"rs-7198292\",\"acceptedTermsAndConditions\":true,\"allowDirectSubmit\":true,\"archivedVersions\":[],\"articleType\":\"Research Article\",\"associatedPublications\":[],\"authors\":[{\"id\":504444334,\"identity\":\"fb2206d8-b90e-4e97-8008-d01ad817fc02\",\"order_by\":0,\"name\":\"Saurav Nayak\",\"email\":\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5klEQVRIie2OoQrCUBSGz2VgumIVFOcj3DGwqa+iDGZZUCwTg6ZrUqu+xcAXGBzYygHrBQ2mWQyKRcGgINZtNsH7lf+E/+P8ABrNb1Ngh88Z5lUM8bVSKOfqmTNqnO9yb4o4ckc3vwWlWciwn6II8uz1XCZWQG604+RAmTqAqzQFPBs4IQvCntwxaQAoAORpw5Ynmz0I28H2KAd3OQEzSwHl2Qb3sRsoN4KiRBBZilDJ0Kj66KxV4lQ4xdyi7jRjmLNhJ4HNxda1Ljd/XKvFiNfUYR/q4TtfZTbNI7ze5expNBrNH/IEP1VQhn9X4SMAAAAASUVORK5CYII=\",\"orcid\":\"\",\"institution\":\"IMS \\u0026 SUM Hospital, SOA University\",\"correspondingAuthor\":true,\"prefix\":\"\",\"firstName\":\"Saurav\",\"middleName\":\"\",\"lastName\":\"Nayak\",\"suffix\":\"\"},{\"id\":504444335,\"identity\":\"8673a7c7-de41-4f86-8064-2d81472db800\",\"order_by\":1,\"name\":\"Gautom Kumar Saharia\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"AIIMS\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Gautom\",\"middleName\":\"Kumar\",\"lastName\":\"Saharia\",\"suffix\":\"\"},{\"id\":504444336,\"identity\":\"b32d0407-d0af-49d9-9995-b07af4be42d0\",\"order_by\":2,\"name\":\"Sandip Kumar Panda\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"AIIMS\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Sandip\",\"middleName\":\"Kumar\",\"lastName\":\"Panda\",\"suffix\":\"\"},{\"id\":504444338,\"identity\":\"1e88f9fb-0d34-4393-ad79-c259c1f2883a\",\"order_by\":3,\"name\":\"Manaswini Mangaraj\",\"email\":\"\",\"orcid\":\"\",\"institution\":\"AIIMS\",\"correspondingAuthor\":false,\"prefix\":\"\",\"firstName\":\"Manaswini\",\"middleName\":\"\",\"lastName\":\"Mangaraj\",\"suffix\":\"\"}],\"badges\":[],\"createdAt\":\"2025-07-23 16:08:13\",\"currentVersionCode\":1,\"declarations\":\"\",\"doi\":\"10.21203/rs.3.rs-7198292/v1\",\"doiUrl\":\"https://doi.org/10.21203/rs.3.rs-7198292/v1\",\"draftVersion\":[],\"editorialEvents\":[],\"editorialNote\":\"\",\"failedWorkflow\":false,\"files\":[{\"id\":90307566,\"identity\":\"ca21cf0f-dbda-478d-af7e-9b5643521caa\",\"added_by\":\"auto\",\"created_at\":\"2025-09-01 09:33:09\",\"extension\":\"jpeg\",\"order_by\":1,\"title\":\"Figure 1\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":23606,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eNormalized Feature Importance of analysed parameters to predict Severity\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"groupimage1.jpeg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7198292/v1/3451d54fe6d3d49e1079d14e.jpeg\"},{\"id\":90307568,\"identity\":\"9cf6bfeb-d0aa-4386-a3b4-8a747815e16f\",\"added_by\":\"auto\",\"created_at\":\"2025-09-01 09:33:09\",\"extension\":\"jpeg\",\"order_by\":2,\"title\":\"Figure 2\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":36411,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003ePerformance metrics of RDF model in both CKD and Non-CKD External Datasets\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"groupimage2.jpeg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7198292/v1/7bb3c32924923a55a5443909.jpeg\"},{\"id\":90309925,\"identity\":\"549afb3f-b452-4f1c-ba7d-828992d8706c\",\"added_by\":\"auto\",\"created_at\":\"2025-09-01 09:41:09\",\"extension\":\"jpeg\",\"order_by\":3,\"title\":\"Figure 3\",\"display\":\"\",\"copyAsset\":false,\"role\":\"figure\",\"size\":13127,\"visible\":true,\"origin\":\"\",\"legend\":\"\\u003cp\\u003eScreenshot of CKD Severity Predictor based on RDF Model\\u003c/p\\u003e\",\"description\":\"\",\"filename\":\"groupimage3.jpeg\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7198292/v1/5406aa7ec3b3d27d314eb745.jpeg\"},{\"id\":92918334,\"identity\":\"5f099712-f7b8-4794-a823-bb7721a656f5\",\"added_by\":\"auto\",\"created_at\":\"2025-10-07 06:10:42\",\"extension\":\"pdf\",\"order_by\":0,\"title\":\"\",\"display\":\"\",\"copyAsset\":false,\"role\":\"manuscript-pdf\",\"size\":736160,\"visible\":true,\"origin\":\"\",\"legend\":\"\",\"description\":\"\",\"filename\":\"manuscript.pdf\",\"url\":\"https://assets-eu.researchsquare.com/files/rs-7198292/v1/f8926781-cd90-4b97-87cf-13260c4b9d89.pdf\"}],\"financialInterests\":\"No competing interests reported.\",\"formattedTitle\":\"Early prediction of severity progression in patients with chronic kidney disease: A Machine Learning Predictive Modelling analysis with retrospective data of a tertiary care hospital\",\"fulltext\":[{\"header\":\"Introduction\",\"content\":\"\\u003cp\\u003eChronic kidney disease (CKD) which is defined as a glomerular filtration rate (GFR) of \\u0026lt;\\u0026thinsp;60mL/min/1.73 m\\u003csup\\u003e\\u003cspan citationid=\\\"CR2\\\" class=\\\"CitationRef\\\"\\u003e2\\u003c/span\\u003e\\u003c/sup\\u003e, is a significant public health issue worldwide.\\u003csup\\u003e\\u003cspan citationid=\\\"CR1\\\" class=\\\"CitationRef\\\"\\u003e1\\u003c/span\\u003e\\u003c/sup\\u003e Previously thought to be a health issue primarily in economically prosperous nations, 4 out of every 5 chronic illness deaths now occur in low- and middle-income countries.\\u003csup\\u003e\\u003cspan citationid=\\\"CR2\\\" class=\\\"CitationRef\\\"\\u003e2\\u003c/span\\u003e\\u003c/sup\\u003e In the Indian population, the actual illness burden of CKD/ESRD cannot be accurately assessed due to the lack of a renal registry; however, the number of fatalities attributable to chronic diseases has increased from 3.78\\u0026nbsp;million in 1990 (40.4% of all deaths) to 7.63\\u0026nbsp;million in 2020 (66.7% of all deaths). CKD has a strong connection to other major diseases, such as cardiovascular disease and diabetes, which continue to be the primary causes of mortality and morbidity in the Indian population.\\u003csup\\u003e\\u003cspan citationid=\\\"CR3\\\" class=\\\"CitationRef\\\"\\u003e3\\u003c/span\\u003e\\u003c/sup\\u003e Worldwide, a significant economic burden is associated with CKD, particularly once kidney failure occurs. This is because of the need for resource-intensive Kidney Replacement Therapy (KRT) resulting in a cost escalation as well as a significant societal burden due to the requirement of either dialysis or kidney transplantation. With the increasing prevalence of CKD along with the huge scope for improvement in the management for CKD, effective solutions are required to help healthcare decision-makers.\\u003csup\\u003e\\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e4\\u003c/span\\u003e, \\u003cspan citationid=\\\"CR5\\\" class=\\\"CitationRef\\\"\\u003e5\\u003c/span\\u003e\\u003c/sup\\u003e\\u003c/p\\u003e\\u003cp\\u003eThe inappropriate activation of the complement system has been implicated in kidney disease.\\u003csup\\u003e\\u003cspan citationid=\\\"CR6\\\" class=\\\"CitationRef\\\"\\u003e6\\u003c/span\\u003e\\u003c/sup\\u003e Complement C3 serves as a critical hub in activating the complement cascade.\\u003csup\\u003e\\u003cspan citationid=\\\"CR7\\\" class=\\\"CitationRef\\\"\\u003e7\\u003c/span\\u003e\\u003c/sup\\u003e Aside from its role in the complement cascade and inflammation, multiple observational studies have found that high C3 levels in plasma are related to the development of hypertension and diabetes, the two predominant causes of CKD.\\u003csup\\u003e\\u003cspan citationid=\\\"CR4\\\" class=\\\"CitationRef\\\"\\u003e4\\u003c/span\\u003e, \\u003cspan citationid=\\\"CR8\\\" class=\\\"CitationRef\\\"\\u003e8\\u003c/span\\u003e\\u003c/sup\\u003e Complement activation at the glomerular and systemic levels is critical in the development and clinical manifestations of renal diseases.\\u003csup\\u003e\\u003cspan citationid=\\\"CR9\\\" class=\\\"CitationRef\\\"\\u003e9\\u003c/span\\u003e\\u003c/sup\\u003e To assess the degree of complement activation, serum complement levels are used as surrogate markers, with Complement 3 (C3) and Complement 4 (C4) being the most commonly reviewed measures in clinical practice.\\u003csup\\u003e\\u003cspan citationid=\\\"CR10\\\" class=\\\"CitationRef\\\"\\u003e10\\u003c/span\\u003e\\u003c/sup\\u003e There is sufficient data in clinical settings to establish the role of complement activation in renal injury that exhibits altered circulating C3 levels, renal C3 deposits, and C3 genetic abnormalities.\\u003csup\\u003e\\u003cspan citationid=\\\"CR11\\\" class=\\\"CitationRef\\\"\\u003e11\\u003c/span\\u003e\\u003c/sup\\u003e\\u003c/p\\u003e\\u003cp\\u003eMachine Learning and Artificial Intelligence models are efficient in finding patterns that are inconspicuous to normal statistical methods. The prediction accuracy can be improved in comparison to conventional biomarkers, by machine learning algorithms as it can learn complex and nonlinear interactions.\\u003csup\\u003e\\u003cspan citationid=\\\"CR12\\\" class=\\\"CitationRef\\\"\\u003e12\\u003c/span\\u003e, \\u003cspan citationid=\\\"CR13\\\" class=\\\"CitationRef\\\"\\u003e13\\u003c/span\\u003e\\u003c/sup\\u003e Early detection of CKD severity will, therefore, be a vital tool for the prevention and management of associated morbidity and mortality. Therefore, in this study, we will be utilising the serum complements C3 and C4 to predict severity in CKD patients.\\u003c/p\\u003e\"},{\"header\":\"Methodology\",\"content\":\"\\u003cp\\u003eThe ethical clearance for the study was obtained from the Institutional Ethical Committee of AIIMS Bhubaneswar vide approval letter no. T/IM-NF/Biochem/23/162 dated 29 Jan 2024. The study was carried out in the Immunology Laboratory under the Department of Biochemistry by revisiting laboratory e-records in conjunction with the Department of Nephrology. Serum Creatinine and Complements were measured in the Beckman AU Autoanalyzer.\\u003c/p\\u003e\\u003cp\\u003e\\u003cstrong\\u003eEstimation of eGFR and Classification\\u003c/strong\\u003e\\u003cp\\u003eEstimated GFR (eGFR) was calculated based on the 2021 CKD-EPI Creatinine formula.\\u003csup\\u003e\\u003cspan citationid=\\\"CR14\\\" class=\\\"CitationRef\\\"\\u003e14\\u003c/span\\u003e\\u003c/sup\\u003e The formula is\\u003c/p\\u003e\\u003c/p\\u003e\\u003cp\\u003eeGFR\\u0026thinsp;=\\u0026thinsp;142 \\u0026times; min(S\\u003csub\\u003ecr\\u003c/sub\\u003e/κ, 1)\\u003csup\\u003eα\\u003c/sup\\u003e \\u0026times; max(S\\u003csub\\u003ecr\\u003c/sub\\u003e/κ, 1)\\u003csup\\u003e\\u0026minus;1.200\\u003c/sup\\u003e \\u0026times; 0.9938\\u003csup\\u003eAge\\u003c/sup\\u003e \\u0026times; 1.012 (if female)\\u003c/p\\u003e\\u003cp\\u003ewhere κ\\u0026thinsp;=\\u0026thinsp;0.7 (females) or 0.9 (males);\\u003c/p\\u003e\\u003cp\\u003ean α = -0.241 (females) or -0.302 (males);\\u003c/p\\u003e\\u003cp\\u003eS\\u003csub\\u003ecr\\u003c/sub\\u003e = serum creatinine in mg/dL;\\u003c/p\\u003e\\u003cp\\u003eAge\\u0026thinsp;=\\u0026thinsp;in years\\u003c/p\\u003e\\u003cp\\u003eThe value of eGFR was rounded to the nearest whole number. The classification of kidney disease based on the Kidney Disease: Improving Global Outcomes (KDIGO) Classification based on eGFR has been tabulated in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e.\\u003csup\\u003e\\u003cspan citationid=\\\"CR15\\\" class=\\\"CitationRef\\\"\\u003e15\\u003c/span\\u003e\\u003c/sup\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab1\\\" border=\\\"1\\\"\\u003e\\u003ccaption language=\\\"En\\\"\\u003e\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 1\\u003c/div\\u003e\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\u003cp\\u003eKDIGO Classification of Kidney Disease based on eGFR\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/caption\\u003e\\u003ccolgroup cols=\\\"3\\\"\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e\\u003cthead\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eCKD Stage\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eDescription\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eeGFR Range\\u003c/p\\u003e\\u003cp\\u003e(mL/min/1.73 m\\u0026sup2;)\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003c/thead\\u003e\\u003ctbody\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eG1\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eNormal or high\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e\\u0026ge;\\u0026thinsp;90\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eG2\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eMildly decreased\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e60\\u0026ndash;89\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eG3a\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eMildly to moderately decreased\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e45\\u0026ndash;59\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eG3b\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eModerately to severely decreased\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e30\\u0026ndash;44\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eG4\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eSeverely decreased\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e15\\u0026ndash;29\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eG5\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eKidney failure\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e\\u0026lt;\\u0026thinsp;15\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003c/tbody\\u003e\\u003c/colgroup\\u003e\\u003c/table\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003eBased on this classification, severity was determined to be Severe if the CKD Stage is G3b or above. Rest was considered to be non-severe.\\u003c/p\\u003e\\u003cp\\u003e\\u003cstrong\\u003ePatient and Public Engagement Summary\\u003c/strong\\u003e\\u003cp\\u003eThe data involved in the study were acquired from the AIIMS Bhubaneswar Laboratory CDAC e-Records retrospectively over a period spanning 24 months. Only those unique data points for which age was greater than 17 years of age were selected for which all Serum Creatinine, C3, and C4 values existed. Any stray errors, blanks or other error-ridden data were duly excluded. Thus, each patient represented one datapoint, totalling to 2279, out of which 238 data points were clinically diagnosed with CKD by a Nephrologist. All data were stored electronically and anonymously.\\u003c/p\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cem\\u003eDataset Separation\\u003c/em\\u003e: The dataset was divided into 3 datasets: 1 dataset to be used as an Internal Dataset and 2 External Datasets. The first external dataset comprised all CKD-diagnosed patients, and the other dataset comprised a 3:1 age- and sex-matched control group for whom the calculated CKD classification was G1. The internal dataset had overall data points that had all 3 parameters, i.e., C3, C4 and Serum Creatinine and was used for predictive model generation.\\u003c/p\\u003e\\u003cp\\u003e\\u003cem\\u003eMachine Learning Models\\u003c/em\\u003e: The internal dataset was deployed to generate predictive models based on 5 machine learning methods: Random Forest Classifier (RDF), XGBoost Classifier (XGB), Gradient-Boosting Classifier (GB), Decision Tree (DT) and Artificial Neural Network (ANN). Recursive Feature Elimination (RFE) was employed to show the feature importance of 5 parameters: Age, Sex, C3, C4 and C3/C4 ratio. The target is \\u0026ldquo;Severity Classification\\u0026rdquo;. The model is generated and evaluated by a 70\\u0026thinsp;\\u0026minus;\\u0026thinsp;30 train-test split and validated by predictions made to the external datasets.\\u003c/p\\u003e\\u003cp\\u003e\\u003cstrong\\u003eEvaluation Metrics\\u003c/strong\\u003e\\u003cp\\u003eThe models are evaluated based on Accuracy, Precision, Recall, F1 Score, MCC, AUC, Sensitivity, Specificity, PPV, NPV, LR+, LR-, DOR, R\\u003csup\\u003e\\u003cspan citationid=\\\"CR2\\\" class=\\\"CitationRef\\\"\\u003e2\\u003c/span\\u003e\\u003c/sup\\u003e.\\u003c/p\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cul\\u003e\\u003cli\\u003e\\u003cp\\u003eTP\\u0026thinsp;=\\u0026thinsp;True Positive. TN\\u0026thinsp;=\\u0026thinsp;True Negative. FP\\u0026thinsp;=\\u0026thinsp;False Positive. FN\\u0026thinsp;=\\u0026thinsp;False Negative.\\u003c/p\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cp\\u003eAccuracy = (TP\\u0026thinsp;+\\u0026thinsp;TN) / (TP\\u0026thinsp;+\\u0026thinsp;TN\\u0026thinsp;+\\u0026thinsp;FP\\u0026thinsp;+\\u0026thinsp;FN). Accuracy is the ratio of correctly predicted instances (both true positives and true negatives) to the total instances. It is a general measure of how well the model is performing.\\u003c/p\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cp\\u003ePrecision or Positive Predictive Value (PPV) = (TP) / (TP\\u0026thinsp;+\\u0026thinsp;FP). Precision is the ratio of true positive instances to the sum of true positive and false positive instances. It indicates how many of the predicted positive instances are actually positive.\\u003c/p\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cp\\u003eRecall or Sensitivity or True Positive Rate (TPR) = (TP) / (TP\\u0026thinsp;+\\u0026thinsp;FN). Recall is the ratio of true positive instances to the sum of true positive and false negative instances. It measures the model's ability to identify positive instances correctly.\\u003c/p\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cp\\u003eF1 Score\\u0026thinsp;=\\u0026thinsp;2 x [(Precision x Recall) / (Precision\\u0026thinsp;+\\u0026thinsp;Recall)]\\u0026thinsp;=\\u0026thinsp;The F1 Score is the harmonic mean of precision and recall. It provides a balance between the two metrics and is useful when you need to account for both false positives and false negatives.\\u003c/p\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cp\\u003eMatthews Correlation Coefficient (MCC) =\\u003c/p\\u003e\\u003c/li\\u003e\\u003c/ul\\u003e\\u003cdiv class=\\\"BlockQuote\\\"\\u003e\\u003cp\\u003e[(TP \\u0026times; TN)-(FP \\u0026times; FN)] / {\\u0026radic;[(TP\\u0026thinsp;+\\u0026thinsp;FP) \\u0026times; (TP\\u0026thinsp;+\\u0026thinsp;FN) \\u0026times; (TN\\u0026thinsp;+\\u0026thinsp;FP) \\u0026times; (TN\\u0026thinsp;+\\u0026thinsp;FN)]}\\u003c/p\\u003e\\u003cp\\u003eMCC is a measure of the quality of binary classifications, considering all four confusion matrix categories. It is a balanced measure and is considered more informative than accurate in imbalanced datasets.\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cul\\u003e\\u003cli\\u003e\\u003cp\\u003eSpecificity = (TN) / (TN\\u0026thinsp;+\\u0026thinsp;FP). Specificity measures the proportion of actual negatives that are correctly identified.\\u003c/p\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cp\\u003eNegative Predictive Value (NPV) = (TN) / (TN\\u0026thinsp;+\\u0026thinsp;FN). NPV measures the proportion of negative results that are true negatives.\\u003c/p\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cp\\u003ePositive Likelihood Ratio (LR+) = (Sensitivity) / (1 \\u0026ndash; Specificity). LR\\u0026thinsp;+\\u0026thinsp;indicates how much the odds of the disease increase when a test is positive. A higher LR\\u0026thinsp;+\\u0026thinsp;value suggests a more useful test.\\u003c/p\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cp\\u003eNegative Likelihood Ratio (LR-) = (1 \\u0026ndash; Sensitivity) / (Specificity). LR- indicates how much the odds of the disease decrease when a test is negative. A lower LR value suggests a more useful test.\\u003c/p\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cp\\u003eDiagnostic Odds Ratio (DOR)\\u0026thinsp;=\\u0026thinsp;LR+ / LR-. DOR is the ratio of the odds of positivity in diseased to non-diseased. It provides a single metric to evaluate the performance of a diagnostic test.\\u003c/p\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cp\\u003eCoefficient of Determination (R\\u003csup\\u003e2\\u003c/sup\\u003e)\\u0026thinsp;=\\u0026thinsp;1 \\u0026ndash; [(Sum of Residuals) / (Total Sum of Squares)]. R\\u0026sup2; measures the proportion of the variance in the dependent variable that is predictable from the independent variable(s). It ranges from 0 to 1, where 1 indicates that the regression predictions perfectly fit the data.\\u003c/p\\u003e\\u003c/li\\u003e\\u003c/ul\\u003e\\u003c/p\\u003e\\u003cp\\u003eThe decision on the model was made based on its efficiency.\\u003c/p\\u003e\\u003cp\\u003e\\u003cstrong\\u003ePipeline for Prediction\\u003c/strong\\u003e\\u003cp\\u003eThe most effective model, as per the evaluation metrics, will be utilised to create a web-based platform to predict CKD severity based on the provided Complement values.\\u003c/p\\u003e\\u003c/p\\u003e\"},{\"header\":\"Results\",\"content\":\"\\u003cp\\u003eThe Internal dataset had 1,331 defined data points of various stages of classification based on their eGFR values. The External \\u0026ndash; CKD dataset had 237 data points, and the External \\u0026ndash; Controls dataset had 711 data points defined in them. The classification parameters of these 3 datasets have been tabulated in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e.\\u003c/p\\u003e\\u003cp\\u003e\\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab2\\\" border=\\\"1\\\"\\u003e\\u003ccaption language=\\\"En\\\"\\u003e\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 2\\u003c/div\\u003e\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\u003cp\\u003eClass distribution in the analysed datasets.\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/caption\\u003e\\u003ccolgroup cols=\\\"10\\\"\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c4\\\" colnum=\\\"4\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c5\\\" colnum=\\\"5\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c6\\\" colnum=\\\"6\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c7\\\" colnum=\\\"7\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c8\\\" colnum=\\\"8\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c9\\\" colnum=\\\"9\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c10\\\" colnum=\\\"10\\\"\\u003e\\u003c/div\\u003e\\u003cthead\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c1\\\" morerows=\\\"1\\\" rowspan=\\\"2\\\"\\u003e\\u003cp\\u003eDataset\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c2\\\" morerows=\\\"1\\\" rowspan=\\\"2\\\"\\u003e\\u003cp\\u003eTotal Datapoints\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colspan=\\\"6\\\" nameend=\\\"c8\\\" namest=\\\"c3\\\"\\u003e\\u003cp\\u003eClassification\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colspan=\\\"2\\\" nameend=\\\"c10\\\" namest=\\\"c9\\\"\\u003e\\u003cp\\u003eSeverity\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eG1\\u003c/p\\u003e\\u003cp\\u003en (%)\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eG2\\u003c/p\\u003e\\u003cp\\u003en (%)\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eG3a\\u003c/p\\u003e\\u003cp\\u003en (%)\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eG3b\\u003c/p\\u003e\\u003cp\\u003en (%)\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003eG4\\u003c/p\\u003e\\u003cp\\u003en (%)\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003eG5\\u003c/p\\u003e\\u003cp\\u003en (%)\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c9\\\"\\u003e\\u003cp\\u003eNon-Severe\\u003c/p\\u003e\\u003cp\\u003en (%)\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c10\\\"\\u003e\\u003cp\\u003eSevere\\u003c/p\\u003e\\u003cp\\u003en (%)\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003c/thead\\u003e\\u003ctbody\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eInternal\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e1,331\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e595 (44.7)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e274 (20.6)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e81 (6.1)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e94 (7.1)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003e105 (7.9)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003e181 (13.6)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c9\\\"\\u003e\\u003cp\\u003e951 (71.4)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c10\\\"\\u003e\\u003cp\\u003e380 (28.6)\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eExternal \\u0026ndash; CKD\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e237\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e23 (90.7)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e20 (8.4)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e20 (8.4)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e25 (10.5)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003e44 (18.6)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003e105 (44.3)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c9\\\"\\u003e\\u003cp\\u003e63 (26.6)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c10\\\"\\u003e\\u003cp\\u003e174 (73.4)\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eExternal - Controls\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e711\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e711 (100)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e-\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e-\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e-\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c7\\\"\\u003e\\u003cp\\u003e-\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c8\\\"\\u003e\\u003cp\\u003e-\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c9\\\"\\u003e\\u003cp\\u003e711 (100)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c10\\\"\\u003e\\u003cp\\u003e-\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003c/tbody\\u003e\\u003c/colgroup\\u003e\\u003c/table\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003eRFE showed that the most important feature in determination was age (0.27), followed by C3 (0.24), C3/C4 (0.22) ratio and C4(0.21). Sex (0.06) was the least important parameter, with a normalised importance score of less than 25%. The Normalized Feature importances have been detailed in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig1\\\" class=\\\"InternalRef\\\"\\u003e1\\u003c/span\\u003e. The difference in the parameter levels has been reported in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab3\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e.\\u003c/p\\u003e\\u003cp\\u003e\\u003c/p\\u003e\\u003cp\\u003e\\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab3\\\" border=\\\"1\\\"\\u003e\\u003ccaption language=\\\"En\\\"\\u003e\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 3\\u003c/div\\u003e\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\u003cp\\u003eComparison of Features between Severe and Non-Severe groups\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/caption\\u003e\\u003ccolgroup cols=\\\"4\\\"\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"char\\\" char=\\\".\\\" class=\\\"colspec\\\" colname=\\\"c4\\\" colnum=\\\"4\\\"\\u003e\\u003c/div\\u003e\\u003cthead\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eParameter\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eSevere [N\\u0026thinsp;=\\u0026thinsp;380]\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eNon-Severe [N\\u0026thinsp;=\\u0026thinsp;951]\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003ep-value *\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003c/thead\\u003e\\u003ctbody\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eAge (in years)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e37 (28\\u0026ndash;47)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e36 (25\\u0026ndash;46)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.526\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eC3 (in mg/dL)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e1.1 (0.92\\u0026ndash;1.29)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e1.2 (0.76\\u0026ndash;1.48)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.427\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eC4 (in mg/dL)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e1.17 (0.8\\u0026ndash;1.37)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e1.19 (0.84\\u0026ndash;1.43)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.705\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eC3/C4 Ratio\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.92 (0.74\\u0026ndash;1.38)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e1.01 (0.65\\u0026ndash;1.49)\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"char\\\" char=\\\".\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.397\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003c/tbody\\u003e\\u003c/colgroup\\u003e\\u003c/table\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003e*: Mann Whitney U test\\u003c/p\\u003e\\u003cp\\u003eThe models were generated using the four parameters: Age, C3, C4, and the calculated parameter C3/C4 ratio. The final decision of model selection was based on the evaluation metrics, primarily focused on the F1 Score, R\\u003csup\\u003e\\u003cspan citationid=\\\"CR2\\\" class=\\\"CitationRef\\\"\\u003e2\\u003c/span\\u003e\\u003c/sup\\u003e value and DOR. Based on these, the best models were the ones defined by RDF, XGB, and DT. The evaluation metrics on the test set of the Internal dataset have been tabulated in Table\\u0026nbsp;\\u003cspan refid=\\\"Tab4\\\" class=\\\"InternalRef\\\"\\u003e4\\u003c/span\\u003e.\\u003c/p\\u003e\\u003cp\\u003e\\u003cdiv class=\\\"gridtable\\\"\\u003e\\u003ctable float=\\\"Yes\\\" id=\\\"Tab4\\\" border=\\\"1\\\"\\u003e\\u003ccaption language=\\\"En\\\"\\u003e\\u003cdiv class=\\\"CaptionNumber\\\"\\u003eTable 4\\u003c/div\\u003e\\u003cdiv class=\\\"CaptionContent\\\"\\u003e\\u003cp\\u003eEvaluation metrics of the Test Set (30%) of the Internal Dataset\\u003c/p\\u003e\\u003c/div\\u003e\\u003c/caption\\u003e\\u003ccolgroup cols=\\\"6\\\"\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c1\\\" colnum=\\\"1\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c2\\\" colnum=\\\"2\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c3\\\" colnum=\\\"3\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c4\\\" colnum=\\\"4\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c5\\\" colnum=\\\"5\\\"\\u003e\\u003c/div\\u003e\\u003cdiv align=\\\"left\\\" class=\\\"colspec\\\" colname=\\\"c6\\\" colnum=\\\"6\\\"\\u003e\\u003c/div\\u003e\\u003cthead\\u003e\\u003ctr\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003eMetrics\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003eRDF\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003eXGB\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003eGB\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003eDT\\u003c/p\\u003e\\u003c/th\\u003e\\u003cth align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003eANN\\u003c/p\\u003e\\u003c/th\\u003e\\u003c/tr\\u003e\\u003c/thead\\u003e\\u003ctbody\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eAccuracy\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.991\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.961\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.792\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.990\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.725\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003ePrecision\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.983\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.953\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.753\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.981\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.535\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eRecall\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.984\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.907\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.389\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.984\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.195\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eF1 Score\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.984\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.930\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.513\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.983\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.286\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eMCC\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.977\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.903\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.432\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.976\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.191\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eAUC\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.999\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.989\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.843\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.988\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.742\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eSensitivity\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.984\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.907\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.389\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.984\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.195\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eSpecificity\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.993\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.983\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.950\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.992\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.934\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003ePPV\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.983\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.953\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.753\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.981\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.535\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eNPV\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.994\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.963\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.798\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.994\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.747\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eLR+\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e148\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e53\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e8\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e131\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e3\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eLR-\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.02\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.10\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e0.64\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.02\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e0.86\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eDOR\\u003c/b\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e9348\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e553\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e12\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e8241\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e3\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003ctr\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c1\\\"\\u003e\\u003cp\\u003e\\u003cb\\u003eR\\u003c/b\\u003e\\u003csup\\u003e\\u003cb\\u003e\\u003cspan citationid=\\\"CR2\\\" class=\\\"CitationRef\\\"\\u003e2\\u003c/span\\u003e\\u003c/b\\u003e\\u003c/sup\\u003e\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c2\\\"\\u003e\\u003cp\\u003e0.954\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c3\\\"\\u003e\\u003cp\\u003e0.809\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c4\\\"\\u003e\\u003cp\\u003e-0.029\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c5\\\"\\u003e\\u003cp\\u003e0.951\\u003c/p\\u003e\\u003c/td\\u003e\\u003ctd align=\\\"left\\\" colname=\\\"c6\\\"\\u003e\\u003cp\\u003e-0.357\\u003c/p\\u003e\\u003c/td\\u003e\\u003c/tr\\u003e\\u003c/tbody\\u003e\\u003c/colgroup\\u003e\\u003c/table\\u003e\\u003c/div\\u003e\\u003c/p\\u003e\\u003cp\\u003eAs the RDF model was determined to be the best model, it was applied to the two external datasets, and performance evaluation was measured in the form of evaluation metrics. The model predicted both the CKD and non-CKD datasets with very robust and high accuracy and diagnostic power. The DOR was 494 and 465 respectively. The metrics are shown in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig2\\\" class=\\\"InternalRef\\\"\\u003e2\\u003c/span\\u003e.\\u003c/p\\u003e\\u003cp\\u003e\\u003c/p\\u003e\\u003cp\\u003eTo aid the clinicians, we went a step further to create an ML model-based pipeline where the values for age, C3, and C4 can be entered, and the model provides the probability of severity as an outcome. The same can be accessed at \\u003cspan class=\\\"ExternalRef\\\"\\u003e\\u003cspan class=\\\"RefSource\\\"\\u003ehttps://bit.ly/c3c4ckdsevereity\\u003c/span\\u003e\\u003cspan address=\\\"https://bit.ly/c3c4ckdsevereity\\\" targettype=\\\"URL\\\" class=\\\"RefTarget\\\"\\u003e\\u003c/span\\u003e\\u003c/span\\u003e. A screenshot of the webpage is shown in Fig.\\u0026nbsp;\\u003cspan refid=\\\"Fig3\\\" class=\\\"InternalRef\\\"\\u003e3\\u003c/span\\u003e.\\u003c/p\\u003e\\u003cp\\u003e\\u003c/p\\u003e\"},{\"header\":\"Discussion\",\"content\":\"\\u003cp\\u003eOur study aimed to understand the early predictive power of complements in relation to the severity of CKD patients. This is one of the first studies that aims to understand the role of C3 and C4 through machine learning models. In this regard, it was elucidated that complements have a very accurate and emerging role in assessing the severity of CKD in patients. Showing very sensitivity, specificity and, high accuracy, these models provide a great deal of insight into predictive power. The ML RDF model aids in analysing the non-linear relationships between the various parameters and thus captures with high accuracy the prediction of CKD as well as its potential in differentiating severity in both CKD and Non-CKD Datasets.\\u003c/p\\u003e\\u003cp\\u003eThere is a scarcity of articles that have exclusively dealt with the participation of complements in CKD. Both the kidneys and abdominal fat tissue exhibit high levels of C3 expression. Notably, the kidney itself plays a major role in circulating C3 and is a large source of extrahepatic C3. In response to damage or cellular stress, the production of C3 is upregulated in podocytes, which are essential for maintaining the integrity of the glomerular filtration barrier.\\u003csup\\u003e\\u003cspan citationid=\\\"CR16\\\" class=\\\"CitationRef\\\"\\u003e16\\u003c/span\\u003e, \\u003cspan citationid=\\\"CR17\\\" class=\\\"CitationRef\\\"\\u003e17\\u003c/span\\u003e\\u003c/sup\\u003e The degree of kidney tissue damage is directly correlated with local C3 generation, most likely through increased T and B-cell function, the recruitment of pro-inflammatory and pro-fibrotic cytokines, the elimination of immune complexes and apoptotic cells, etc. Due to its in part renal origin, plasma C3 may be a sign of chronic kidney cell damage, which leads to chronic kidney disease (CKD) .\\u003csup\\u003e\\u003cspan citationid=\\\"CR11\\\" class=\\\"CitationRef\\\"\\u003e11\\u003c/span\\u003e\\u003c/sup\\u003e Higher levels of circulating C3 have been reported to be associated with cardiovascular diseases in patients with renal disorders, thus adding another requirement to the need for assessing severity as early as possible.\\u003csup\\u003e\\u003cspan citationid=\\\"CR18\\\" class=\\\"CitationRef\\\"\\u003e18\\u003c/span\\u003e\\u003c/sup\\u003e Similar associations, though scanty, also exist for C4 and C3/C4 ratios.\\u003csup\\u003e\\u003cspan citationid=\\\"CR19\\\" class=\\\"CitationRef\\\"\\u003e19\\u003c/span\\u003e, \\u003cspan citationid=\\\"CR20\\\" class=\\\"CitationRef\\\"\\u003e20\\u003c/span\\u003e\\u003c/sup\\u003e A study by Liu et al. primarily suggested that the progression of CKD and mortality is related to an increase in C4 levels in serum.\\u003csup\\u003e\\u003cspan citationid=\\\"CR21\\\" class=\\\"CitationRef\\\"\\u003e21\\u003c/span\\u003e\\u003c/sup\\u003e Also, an association of C3, C4, as well as the C3/C4 ratio, has been reported to aggravate renal dysfunction into End Stage Renal Diseases (ESRD), thus affecting morbidity and mortality in patients.\\u003csup\\u003e\\u003cspan citationid=\\\"CR22\\\" class=\\\"CitationRef\\\"\\u003e22\\u003c/span\\u003e\\u003c/sup\\u003e\\u003c/p\\u003e\\u003cp\\u003ePrevious machine learning models have been utilised to predict the progression of CKD in patients, and they have been done with high accuracy by random forest classification.\\u003csup\\u003e\\u003cspan citationid=\\\"CR23\\\" class=\\\"CitationRef\\\"\\u003e23\\u003c/span\\u003e\\u003c/sup\\u003e Early prediction of CKD has also been reported based on clinical and physiological data by Islam et al. in a similar study.\\u003csup\\u003e\\u003cspan citationid=\\\"CR24\\\" class=\\\"CitationRef\\\"\\u003e24\\u003c/span\\u003e\\u003c/sup\\u003e These factors underline the strength of our study, as it is one of the first of its kind, and it tries to assess severity predictions for patients, which will drastically reduce the burden both clinically and financially. This also comes against the backdrop of a rise in kidney diseases in India, as well as rapid progression of severity due to lax treatment and patient awareness. Notwithstanding, our study was limited by the scope of the quantity of data points available as patients of CKD were very few in our centre who had been advised for a Complement profile. As we advance, a longitudinal study with followed-up measurement of complements, along with treatment, aimed at monitoring eGFR levels and disease progression can aid in better predictive modelling.\\u003c/p\\u003e\"},{\"header\":\"Conclusion\",\"content\":\"\\u003cp\\u003eMachine learning predictive modelling can assess with high accuracy the severity progression in early CKD, thus preventing morbidity and mortality in patients. The online predictive pipeline will aid clinicians in determining severity as well, providing a holistic approach to the management of CKD in such cases.\\u003c/p\\u003e\"},{\"header\":\"Declarations\",\"content\":\"\\u003cp\\u003e\\u003cem\\u003eEthics approval and consent to participate\\u003c/em\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eThe ethical clearance for the study was obtained from the Institutional Ethical Committee of AIIMS Bhubaneswar vide approval letter no. T/IM-NF/Biochem/23/162 dated 29 Jan 2024. This has been done in accordance with the Declaration of Helsinki. In accordance to the IEC waiver for retrospective study, there was no requirement for Informed Consent Form for the participants.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cem\\u003eConsent for publication\\u003c/em\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eNot Applicable as Retrospective Study. Waiver received from IEC.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cem\\u003eAvailability of data and materials\\u003c/em\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eThe datasets used and/or analysed during the current study are available from the corresponding author on reasonable request.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cem\\u003eCompeting interests\\u003c/em\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eThe authors declare that they have no competing interests\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cem\\u003eFunding\\u003c/em\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eNo funding was received.\\u003c/p\\u003e\\n\\u003cp\\u003e\\u003cem\\u003eAuthors’ Contributions\\u003c/em\\u003e\\u003c/p\\u003e\\n\\u003cp\\u003eSN and GK conceptualized the study and the methodology. SN developed the algorithm for model development and interpreted the data. SKP provided expert opinion on CKD cases and their diagnosis and determination. MM and GKS drafted the manuscript and proofread it. All authors read and approved the final manuscript.\\u003c/p\\u003e\"},{\"header\":\"References\",\"content\":\"\\u003col\\u003e\\u003cli\\u003e\\u003cspan\\u003eLiyanage T, Toyama T, Hockham C, Ninomiya T, Perkovic V, Woodward M, et al. Prevalence of chronic kidney disease in Asia: a systematic review and analysis. BMJ Glob Health. 2022;7(1):e007525.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eLozano R, Naghavi M, Foreman K, Lim S, Shibuya K, Aboyans V, et al. Global and regional mortality from 235 causes of death for 20 age groups in 1990 and 2010: a systematic analysis for the Global Burden of Disease Study 2010. Lancet. 2012;380(9859):2095\\u0026ndash;128.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eAgarwal SK, Srivastava RK. Chronic Kidney Disease in India: Challenges and Solutions. Nephron Clin Pract. 2009;111(3):c197\\u0026ndash;203.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eJha V, Ur-Rashid H, Agarwal SK, Akhtar SF, Kafle RK, Sheriff R. The state of nephrology in South Asia. Kidney Int. 2019;95(1):31\\u0026ndash;7.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eElshahat S, Cockwell P, Maxwell AP, Griffin M, O\\u0026rsquo;Brien T, O\\u0026rsquo;Neill C. The impact of chronic kidney disease on developed countries from a health economics perspective: A systematic scoping review. Barretti P, editor. PLOS ONE. 2020;15(3):e0230512.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eKaartinen K, Safa A, Kotha S, Ratti G, Meri S. Complement dysregulation in glomerulonephritis. Semin Immunol. 2019;45:101331.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eBao X, Born\\u0026eacute; Y, Muhammad IF, Schulz CA, Persson M, Orho-Melander M, et al. Complement C3 and incident hospitalization due to chronic kidney disease: a population-based cohort study. BMC Nephrol. 2019;20(1):61.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eWlazlo N, Van Greevenbroek MMJ, Ferreira I, Feskens EJM, Van Der Kallen CJH, Schalkwijk CG, et al. Complement Factor 3 Is Associated With Insulin Resistance and With Incident Type 2 Diabetes Over a 7-Year Follow-up Period: The CODAM Study. Diabetes Care. 2014;37(7):1900\\u0026ndash;9.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eLhotta K, Schlogl A, Kronenberg F, Joannidis M, Konig P. Glomerular deposition of the complement C4 isotypes C4A and C4B in glomeruonephritis. Nephrol Dial Transplant Off Publ Eur Dial Transpl Assoc -. Eur Ren Assoc. 1996;11(6):1024\\u0026ndash;8.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003ePan M, Zhou Q, Zheng S, You X, Li D, Zhang J, et al. Serum C3/C4 ratio is a novel predictor of renal prognosis in patients with IgA nephropathy: a retrospective study. Immunol Res. 2018;66(3):381\\u0026ndash;91.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eThurman JM. Complement in Kidney Disease: Core Curriculum 2015. Am J Kidney Dis. 2015;65(1):156\\u0026ndash;68.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eKumari S, Tripathy S, Nayak S, Rajasimman AS. Machine learning\\u0026ndash;aided algorithm design for prediction of severity from clinical, demographic, biochemical and immunological parameters: Our COVID-19 experience from the pandemic. J Fam Med Prim Care. 2024;13(5):1937\\u0026ndash;43.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eNayak S, Singh A, Mangaraj M, Saharia GK. Predicting immune risk in treatment-na\\u0026iuml;ve HIV patients using a machine learning algorithm: a decision tree algorithm based on micronutrients and inversion of the CD4/CD8 ratio. Front Nutr. 2024;11:1443076.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eMiller WG, Kaufman HW, Levey AS, Straseski JA, Wilhelms KW, Yu HY, Elsie. National Kidney Foundation Laboratory Engagement Working Group Recommendations for Implementing the CKD-EPI 2021 Race-Free Equations for Estimated Glomerular Filtration Rate: Practical Guidance for Clinical Laboratories. Clin Chem. 2022;68(4):511\\u0026ndash;20.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eChen TK, Knicely DH, Grams ME. Chronic Kidney Disease Diagnosis and Management: A Review. JAMA. 2019;322(13):1294.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eZhou W, Marsh JE, Sacks SH. Intrarenal synthesis of complement. Kidney Int. 2001;59(4):1227\\u0026ndash;35.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eTang S, Zhou W, Sheerin NS, Vaughan RW, Sacks SH. Contribution of renal secreted complement C3 to the circulating pool in humans. J Immunol Baltim Md. 1950. 1999;162(7):4336\\u0026ndash;41.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eLines SW, Richardson VR, Thomas B, Dunn EJ, Wright MJ, Carter AM. Complement and Cardiovascular Disease - The Missing Link in Haemodialysis Patients? Nephron. 2016;132(1):5\\u0026ndash;14.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eXing Z, Wang Y, Gong K, Chen Y. Plasma C4 level was associated with mortality, cardiovascular and cerebrovascular complications in hemodialysis patients. BMC Nephrol. 2022;23(1):232.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eZhang Y, Duan SW, Chen P, Yin Z, Wang Y, Cai GY et al. Relationship between serum C3/C4 ratio and prognosis of immunoglobulin A nephropathy based on propensity score matching. Chin Med J (Engl). 2020;(6):631\\u0026ndash;7.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eLiu J, Zha Y, Zhang P, He P, He L. The Association Between Serum Complement 4 and Kidney Disease Progression in Idiopathic Membranous Nephropathy: A Multicenter Retrospective Cohort Study. Front Immunol. 2022;13:896654.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eMatsuda S, Oe K, Kotani T, Okazaki A, Kiboshi T, Suzuka T, et al. Serum Complement C4 Levels Are a Useful Biomarker for Predicting End-Stage Renal Disease in Microscopic Polyangiitis. Int J Mol Sci. 2023;24(19):14436.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eFerguson T, Ravani P, Sood MM, Clarke A, Komenda P, Rigatto C, et al. Development and External Validation of a Machine Learning Model for Progression of CKD. Kidney Int Rep. 2022;7(8):1772\\u0026ndash;81.\\u003c/span\\u003e\\u003c/li\\u003e\\u003cli\\u003e\\u003cspan\\u003eIslam MA, Hussein MdA. Chronic kidney disease prediction based on machine learning algorithms. J Pathol Inf. 2023;14:100189.\\u003c/span\\u003e\\u003c/li\\u003e\\u003c/ol\\u003e\"}],\"fulltextSource\":\"\",\"fullText\":\"\",\"funders\":[],\"hasAdminPriorityOnWorkflow\":false,\"hasManuscriptDocX\":true,\"hasOptedInToPreprint\":true,\"hasPassedJournalQc\":\"\",\"hasAnyPriority\":false,\"hideJournal\":true,\"highlight\":\"\",\"institution\":\"\",\"isAcceptedByJournal\":false,\"isAuthorSuppliedPdf\":false,\"isDeskRejected\":\"\",\"isHiddenFromSearch\":false,\"isInQc\":false,\"isInWorkflow\":false,\"isPdf\":false,\"isPdfUpToDate\":true,\"isWithdrawnOrRetracted\":false,\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"researchsquare\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":true,\"externalIdentity\":\"\",\"sideBox\":\"\",\"snPcode\":\"\",\"submissionUrl\":\"/submission\",\"title\":\"Research Square\",\"twitterHandle\":\"researchsquare\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"\",\"reportingPortfolio\":\"\",\"inReviewEnabled\":false,\"inReviewRevisionsEnabled\":true},\"keywords\":\"\",\"lastPublishedDoi\":\"10.21203/rs.3.rs-7198292/v1\",\"lastPublishedDoiUrl\":\"https://doi.org/10.21203/rs.3.rs-7198292/v1\",\"license\":{\"name\":\"CC BY 4.0\",\"url\":\"https://creativecommons.org/licenses/by/4.0/\"},\"manuscriptAbstract\":\"\\u003ch2\\u003eBackground\\u003c/h2\\u003e\\u003cp\\u003eChronic Kidney Disease (CKD) represents a growing health burden, particularly in low- and middle-income countries. Its progression to end-stage renal disease (ESRD) necessitates resource-intensive interventions. The inappropriate activation of the complement system, particularly Complement 3 (C3) and Complement 4 (C4), has been implicated in renal injury. While these markers are biologically relevant, their predictive value for CKD severity has not been adequately explored. This study aimed to assess the potential of serum C3 and C4 in predicting CKD severity using machine learning (ML) models.\\u003c/p\\u003e\\u003ch2\\u003eMethods\\u003c/h2\\u003e\\u003cp\\u003eA retrospective dataset comprising 2,279 adults (\\u0026gt;\\u0026thinsp;17 years) was extracted from the laboratory records of AIIMS Bhubaneswar. CKD severity was classified using the 2021 CKD-EPI Creatinine formula, with Stage G3b and above defined as severe CKD. Of these, 1,331 complete records (with C3, C4, and creatinine) formed the internal dataset, while two external datasets were used for validation\\u0026mdash;one with clinically confirmed CKD and the other age- and sex-matched controls. Predictive models were developed using five ML algorithms: Random Forest (RDF), XGBoost (XGB), Gradient Boosting (GB), Decision Tree (DT), and Artificial Neural Networks (ANN). Model performance was evaluated using accuracy, F1 score, R\\u0026sup2;, diagnostic odds ratio (DOR), and other metrics.\\u003c/p\\u003e\\u003ch2\\u003eResults\\u003c/h2\\u003e\\u003cp\\u003eRecursive Feature Elimination identified age, C3, and the C3/C4 ratio as the most influential predictors. Among the models, RDF performed best (F1: 0.984, Accuracy: 0.991, DOR: 9348, R\\u0026sup2;: 0.954). External validation confirmed its high diagnostic power (DOR: 494 and 465 for CKD and control datasets, respectively). A web-based tool was developed to aid clinicians in estimating CKD severity using age, C3, and C4 values.\\u003c/p\\u003e\\u003ch2\\u003eConclusion\\u003c/h2\\u003e\\u003cp\\u003eThis study demonstrates that serum complements C3 and C4 can serve as early predictive biomarkers for CKD severity when interpreted via machine learning models. The RDF-based prediction pipeline offers a clinically relevant, non-invasive tool for stratifying CKD patients, potentially reducing the burden of late-stage interventions. Further prospective studies are warranted to validate these findings longitudinally.\\u003c/p\\u003e\",\"manuscriptTitle\":\"Early prediction of severity progression in patients with chronic kidney disease: A Machine Learning Predictive Modelling analysis with retrospective data of a tertiary care hospital\",\"msid\":\"\",\"msnumber\":\"\",\"nonDraftVersions\":[{\"code\":1,\"date\":\"2025-09-01 09:33:05\",\"doi\":\"10.21203/rs.3.rs-7198292/v1\",\"editorialEvents\":[{\"type\":\"communityComments\",\"content\":0}],\"status\":\"published\",\"journal\":{\"display\":true,\"email\":\"info@researchsquare.com\",\"identity\":\"researchsquare\",\"isNatureJournal\":false,\"hasQc\":true,\"allowDirectSubmit\":true,\"externalIdentity\":\"\",\"sideBox\":\"\",\"snPcode\":\"\",\"submissionUrl\":\"/submission\",\"title\":\"Research Square\",\"twitterHandle\":\"researchsquare\",\"acdcEnabled\":true,\"dfaEnabled\":false,\"editorialSystem\":\"\",\"reportingPortfolio\":\"\",\"inReviewEnabled\":false,\"inReviewRevisionsEnabled\":true}}],\"origin\":\"\",\"ownerIdentity\":\"93976d51-e891-4c21-88f7-5fdf89dc8828\",\"owner\":[],\"postedDate\":\"September 1st, 2025\",\"published\":true,\"recentEditorialEvents\":[],\"rejectedJournal\":[],\"revision\":\"\",\"amendment\":\"\",\"status\":\"posted\",\"subjectAreas\":[],\"tags\":[],\"updatedAt\":\"2025-10-07T05:54:13+00:00\",\"versionOfRecord\":[],\"versionCreatedAt\":\"2025-09-01 09:33:05\",\"video\":\"\",\"vorDoi\":\"\",\"vorDoiUrl\":\"\",\"workflowStages\":[]},\"version\":\"v1\",\"identity\":\"rs-7198292\",\"journalConfig\":\"researchsquare\"},\"__N_SSP\":true},\"page\":\"/article/[identity]/[[...version]]\",\"query\":{\"redirect\":\"/article/rs-7198292\",\"identity\":\"rs-7198292\",\"version\":[\"v1\"]},\"buildId\":\"8U1c8b4HqxoKbykW_rLl7\",\"isFallback\":false,\"isExperimentalCompile\":false,\"dynamicIds\":[84888],\"gssp\":true,\"scriptLoader\":[]}","source_license":"CC-BY-4.0","license_restricted":false}