Handling Missing Values and Outliers in Advanced Data Pre-processing: An Enhancement of Diabetes Classification Accuracy

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background The rising global threat of diabetes demands timely detection to prevent its complications. Data scientists and practitioners are seen to be used AI and some other classification models on different aspects. Nevertheless, addressing missing data and outlier’s accurate predictions may be questionable. As such incorporating ML and AI for early diagnosis has gained attention. This study integrates medical knowledge and what types of advanced technology to develop a comprehensive diabetes classification model, focusing on handling missing values and outliers to achieve improved accuracy in early disease identification.Methods The researcher’s methodology prioritized meticulous data pre-processing to enhance analysis quality. To address missing data, the researchers utilized the missForest method, employing a multistage imputation process that minimizes data loss and distortions. Outlier detection relied on Mahalanobis squared distances, identifying anomalous data points. Instead of outright removal, the researchers strategically leveraged the missForest method, known for its robust imputation capabilities. Temporarily replacing outliers with missing values, this approach seamlessly integrated imputation. The ensuing hybrid data, minus extreme outliers and enriched via missForest, formed the foundation for subsequent analysis and modelling. Model selection and evaluation were performed on pre-processed data. This analysis incorporated two-step CV: initial dataset partition (80% training, 20% testing) and ten iterations of ten-fold cross-validation for model stability and parameter optimization. A diverse array of ML models—LogitBoost, mlpWeightDecayML, avNNet, and others—were assessed. Metrics such as sensitivity, specificity, precision, recall, F1-score, AUC, accuracy, and Kappa score were scrutinized.Results Among the models examined, LogitBoost emerged as a strong contender with a sensitivity of 0.8095, specificity of 0.9464, precision of 0.85, recall of 0.8095, F1-score of 0.8293, AUC of 0.7888, accuracy of 0.9091, and Kappa score of 0.7674. However, the comparative results showcase varying performances across different metrics and models. Sensitivity ranged from 0.6792 to 0.9057, specificity from 0.6 to 0.9464, and precision from 0.5455 to 0.85.Conclusions In summation, the methodical approach has illuminated the path toward enhanced diabetes classification accuracy. By diligently addressing missing values through the robust missForest method and tactfully managing outliers using the hybrid approach, the researchers have elevated the integrity and quality of the PIMA dataset. This strategic handling of missing values and outliers has not only fortified the dataset against potential distortions but has also culminated in improved accuracy in diabetes classification. Through the synergy of meticulous pre-processing, strategic outlier management, and comprehensive model evaluation, the researchers have contributed valuable insights into the realm of early diabetes detection.
Full text 219,294 characters · extracted from preprint-html · click to expand
Handling Missing Values and Outliers in Advanced Data Pre-processing: An Enhancement of Diabetes Classification Accuracy | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Handling Missing Values and Outliers in Advanced Data Pre-processing: An Enhancement of Diabetes Classification Accuracy Md. Hossain, Astami Devnath, Provash Karmokar This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3364064/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background The rising global threat of diabetes demands timely detection to prevent its complications. Data scientists and practitioners are seen to be used AI and some other classification models on different aspects. Nevertheless, addressing missing data and outlier’s accurate predictions may be questionable. As such incorporating ML and AI for early diagnosis has gained attention. This study integrates medical knowledge and what types of advanced technology to develop a comprehensive diabetes classification model, focusing on handling missing values and outliers to achieve improved accuracy in early disease identification. Methods The researcher’s methodology prioritized meticulous data pre-processing to enhance analysis quality. To address missing data, the researchers utilized the missForest method, employing a multistage imputation process that minimizes data loss and distortions. Outlier detection relied on Mahalanobis squared distances, identifying anomalous data points. Instead of outright removal, the researchers strategically leveraged the missForest method, known for its robust imputation capabilities. Temporarily replacing outliers with missing values, this approach seamlessly integrated imputation. The ensuing hybrid data, minus extreme outliers and enriched via missForest, formed the foundation for subsequent analysis and modelling. Model selection and evaluation were performed on pre-processed data. This analysis incorporated two-step CV: initial dataset partition (80% training, 20% testing) and ten iterations of ten-fold cross-validation for model stability and parameter optimization. A diverse array of ML models—LogitBoost, mlpWeightDecayML, avNNet, and others—were assessed. Metrics such as sensitivity, specificity, precision, recall, F1-score, AUC, accuracy, and Kappa score were scrutinized. Results Among the models examined, LogitBoost emerged as a strong contender with a sensitivity of 0.8095, specificity of 0.9464, precision of 0.85, recall of 0.8095, F1-score of 0.8293, AUC of 0.7888, accuracy of 0.9091, and Kappa score of 0.7674. However, the comparative results showcase varying performances across different metrics and models. Sensitivity ranged from 0.6792 to 0.9057, specificity from 0.6 to 0.9464, and precision from 0.5455 to 0.85. Conclusions In summation, the methodical approach has illuminated the path toward enhanced diabetes classification accuracy. By diligently addressing missing values through the robust missForest method and tactfully managing outliers using the hybrid approach, the researchers have elevated the integrity and quality of the PIMA dataset. This strategic handling of missing values and outliers has not only fortified the dataset against potential distortions but has also culminated in improved accuracy in diabetes classification. Through the synergy of meticulous pre-processing, strategic outlier management, and comprehensive model evaluation, the researchers have contributed valuable insights into the realm of early diabetes detection. Health sciences/Risk factors Scientific community and society/Scientific community/Research management Diabetes Early Detection Machine Learning Artificial Intelligence Missing Data Outliers missForest and Classification Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Figure 13 Figure 14 Figure 15 Figure 16 Figure 17 Figure 18 Figure 19 Figure 20 Figure 21 Background The escalating global prevalence of diabetes underscores the urgency of timely detection and intervention to mitigate its far-reaching consequences [ 1 ]. The integration of ML has emerged as a promising avenue for early diagnosis, attracting significant attention from researchers and healthcare professionals alike. However, accurate disease prediction hinges not only on advanced technological tools but also on the meticulous handling of underlying data complexities, including missing values and outliers. In this context, the present study positions itself as a critical nexus between medical expertise and cutting-edge ML technology. The driving force behind this endeavour is to develop a comprehensive diabetes classification model that not only leverages the potential of ML but also addresses the pivotal challenge of managing missing data and outliers. These issues are particularly salient in medical datasets, where the quality of predictions can be undermined by data irregularities. As the study delves into these complexities, it sheds light on the paramount importance of data pre-processing. It seeks to establish a methodology that not only enhances the quality of analysis but also ensures the robustness of diabetes classification models. By focusing on the strategic treatment of missing values and outliers, the study aims to achieve improved accuracy in the early identification of diabetes, thereby contributing to effective preventive measures and optimized patient care. The significance of this research lies not only in its potential to refine disease classification techniques but also in its broader implications for healthcare decision-making. The amalgamation of medical insights with advanced ML technology is poised to revolutionize the early detection of diabetes, which, in turn, can lead to better patient outcomes and reduced healthcare burdens. As the study navigates the intricate terrain of data pre-processing, modelling, and analysis, it strives to illuminate a pathway toward more accurate and reliable diabetes prediction, amplifying the impact of both medical research and technological innovation. The urgent need for a predictive tool to aid early disease detection and recommend lifestyle changes is evident in the context of diabetes. With diabetes claiming approximately 1.6 million lives annually, its significance cannot be understated [ 2 ]. Diabetes occurs when blood glucose levels become excessively high, stemming from either insufficient insulin production or poor insulin utilization by the body's cells [ 3 ]. Glucose, produced during digestion, is regulated by insulin, which facilitates glucose absorption into cells for energy. Inadequate insulin results in glucose build-up in the blood, leading to elevated blood glucose levels and diabetes [ 4 ]. High blood glucose manifests in symptoms like excessive thirst and frequent urination. A typical adult's glucose range is 70 to 99 mg/dl, while values exceeding 126 mg/dl indicate diabetes [ 5 ]. Persistent high blood glucose can trigger severe complications such as heart disease, renal failure, stroke, and nerve damage [ 6 ], [ 7 ]. Despite advancements, diabetes remains incurable, with long-term cases leading to macrovascular and microvascular complications. Macrovascular issues involve large blood arteries, while microvascular complications impact smaller blood vessels, contributing to kidney, eye, foot, and nerve complications [ 8 ]. Early detection and management are vital to curbing diabetes progression [ 9 ]. Exercise and dietary habits play pivotal roles in preventing and managing diabetes [ 5 ]. Healthcare data, encompassing patient records and examination findings, holds the potential for predictive insights. Automation powered by ML is essential for detecting hidden patterns, enabling more accurate decision-making in diabetes diagnosis [ 10 ]–[ 12 ]. ML and data mining technologies have transformed the healthcare landscape, aiding in feature selection and automation of diabetes prediction [ 13 ]. These techniques uncover patterns in complex datasets, fostering reliable decision-making [ 14 ]. Current research, however, often lacks advanced missing value treatment methods and thorough evaluation of ML algorithms. The researchers propose using supervised and unsupervised ML techniques, as well as AI and DL, incorporating non-parametric missing value imputation using random forests. The researchers aim to enhance diabetes categorization accuracy across various metrics, ultimately establishing a robust predictive model for diabetes classification. To fulfil the SDG-03: “Ensure healthy lives and promote well-being for all at all ages” and to cope pace with the artificial intelligence-driven fourth industrial revolution which is “A technological shift affecting cultures and economies all over the globe”, this study must be a dire need worldwide, especially for Bangladesh [ 15 ], [ 16 ]. At present, diabetes has increasingly been listed in the top position as a major reason for death. International Diabetes Federation declared that 382 million people are breathing with diabetes globally like Bangladesh. Diabetes if untreated may turn deadly and directly or indirectly invites a lot of other diseases. However, diabetes is largely avoidable and can be avoided by lifestyle changes. These changes may also lower the probability of developing cancer and heart disease. Diabetes imposes an incredible socioeconomic burden on patients and the family unit. Diabetes disease-associated expenditures influence the family unit’s everyday work and further compliance, which is directly halting the fulfillment of SDG-03. All knows “prevention is better than cure”. So, if all is known the proper lifestyle and ready to accept the lifestyle, then it may reduce the rate of diabetes patients, which will be helpful for both socioeconomic and family welfare. Table 1 Some Related Studies, Their Methods and Contributions Ref Data Method/Model Metric Contribution [ 17 ] PIDD Removing outliers > Feature selection: Pearson’s correlation method > Normalization > k-fold CV > ML, AI, DL ACC 88.6 Introducing DL and Weka tools in health [ 18 ] PIDD Min-max normalization > Soft voting classifier > Ensemble ACC 79.08 Ensemble of ML algorithms [ 19 ] PIDD Removing outliers > Dealing with missing values > Data standardization > Web application development using flask ACC 80.26 Web application development using flask to fit model [ 20 ] PIDD Missing values as zero > Univariate feature selection > Creating new features > Ensemble technique: Max Voting ACC 78 Creating new features by categorizing some of the variables helps to increase the accuracy [ 21 ] T2DM from CBHS health funds company in Australia Outlier treatment > Data cleaning > Cohort selection > Network analysis > Network feature > ML, NN ACC 82.52 Bipartite graph and projection of the patient network [ 22 ] T2DM from CPCSSN An approximation of missing values > Hidden Markov models > Newton’s Divide Difference Method ACC 80.4 Proposed an algorithm of the polynomial function implemented to estimating the missed values [ 23 ] T2DM, PLA General Hospital Demographic and clinical variables as predictive risk factors > 3 years of follow-up > Cleaning big medical data derived from EMR ACC 76.8 ML of big medical data derived from electronic medical records in the real-world [ 24 ] Pregnancy Exercise and Nutrition, PEARS Data cleaning > Feature selection > Comparative study > Modelling BAL ACC 79.4 Ethnicity is introduced in diabetes classification, a comparative study using ML [ 25 ] Single-payer health system in Ontario Cohort characteristics > Feature extraction > Model performance > Feature contribution > Cost analysis AUC 77.7 Diabetes classification using big data introduces cost [ 26 ] PIDD ADAP ACC 65.9 Initially collected and analyzed dataset [ 27 ] PIDD DT, NB, KNN, SVM ACC 76.7 (SVM) Compared algorithms [ 28 ] PIDD DBN ensemble ACC 78 Proposed ensemble model [ 29 ] PIDD DL AUC: 80.6 Developed DL [ 30 ] PIDD DT, NB, NN ACC 75.47 (NN) Compared algorithms [ 31 ] PIDD DT ACC 76.47 Used DT algorithm [ 32 ] PIDD NN ACC 77 Used NN [ 33 ] PIDD NN ACC 74.9 (NN) Used NN and statistical methods [ 34 ] PIDD SVM ACC 76.6 Used SVM with combined kernels [ 35 ] PIDD SVM ACC 83.3 Used SVM with feature selection [ 36 ] PIDD Adversarial Learning AUC: 79 Reduced disparity using adversarial learning [ 37 ] PIDD ML algorithms ACC 76.1 (SVM) Compared ML algorithms [ 38 ] PIDD Ant Colony ACC 74.6 Used Ant Colony optimization [ 39 ] PIDD PNN ACC 78.3 Used PNN with feature extraction [ 40 ] PIDD GRNN ACC 76.8 (GRNN) Compared GRNN algorithms The Table 1 provides a summary of various research studies related to diabetes classification and prediction methods. Each entry includes the dataset used, data pre-processing steps, ML techniques applied, and the corresponding accuracy or performance metric. Additionally, some studies mention unique contributions or innovations in their approaches, such as introducing DL, web application development, or network analysis. The studies cover different aspects of diabetes classification and prediction, ranging from data cleaning and feature selection to the use of advanced techniques like DL and ML. Tejas and Pramila [ 41 ] picked the two methods LR and SVM to construct a diabetes prediction model. Pre-processing the data was done to get better outcomes. SVM outperformed other methods with an accuracy of 79%, they discovered. Using three separate ML algorithms—RF, DT, and NB—in Hadoop-based clusters, Yuvaraj and Sripreethaa [ 42 ] developed a diabetes prediction model. On the dataset, they used pre-processing methods. The findings demonstrated that the RF algorithm produced the greatest accuracy rate of 94%. DT, SVM, and Naive Bayes algorithms were employed by Deepti and Dilip [ 43 ]. Performance was enhanced with the use of 10-CV. The Naive Bayes model produced results with the best accuracy (76.30%). The PIDD was utilized in both of these articles. Both, Olaniyi and Adnan [ 44 ], and Swapna et al. [ 13 ] made use of DL techniques for diabetes prediction. In the former, a multilayer feed-forward NN was used. The model's training was carried out using the back-propagation approach. In order to achieve numerical stability, they also employed the PIMA Indian dataset and normalized it prior to pre-processing. They achieved an accuracy of 82%. The latter employed two models utilizing CNN and CNN-LSTM on a dataset named Electrocardiograms. There were 142,000 samples in the dataset, along with eight characteristics. With a five-fold cross validation for both models, they were able to achieve accuracy of 93.6% with the CNN model and 95.1% with the CNN-LSTM model. From the Table 1 , for PIDD, ADAP algorithm achieved an accuracy of 65.9% accuracy. Several studies have compared algorithms like DT, NB, KNN and SVM showing the best performance of 76.7% accuracy. A deep belief network-based ensemble classifier was proposed and obtained 78% accuracy. A DL model developed for this dataset achieved an AUC of 80.6%. Related study compared DT, NB and NN, finding NN gave maximum accuracy of 75.47%. DTs were applied with an accuracy result of 76.47%. NN alone obtained 77% accuracy. NN combined with statistical methods achieved 74.9% accuracy. SVM with combined kernels attained 76.6% accuracy. SVM with feature selection yielded higher accuracy of 83.3%. Adversarial learning reduced disparity and improved AUC to 79%. A comparison found SVM had the best accuracy of 76.1%. An ant colony optimization classifier developed achieved 74.6% accuracy. PNN with feature extraction gave 78.3% accuracy. A study analysed different generalized regression NN algorithms, reporting GRNN achieved 76.8% accuracy. All of the aforementioned research provided a performance comparison of several ML algorithms. While some of them utilized data pre-processing and cross-validation strategies to increase accuracy, all of them were more concerned with comparing the performance of several models than they were with perfecting a single one. In this study, the researchers have focused on a single model and investigated strategies that can boost performance by enhancing both execution speed and accuracy with a special attention to algorithm selection, pre- and post-processing of the data have a significant impact on the model's overall improvement of models. The main objective of this study is to compare the accuracies of diabetes classification among the AI algorithms using associated risk factors. The specific objectives are: 1. To identify the risk factors for diabetes and arising awareness among the people; 2. To show how parameter tuning can increase the accuracies of the models; 3. To show how missing values can be imputed and may have contributory role for the betterment of ML methods; 4. To show how data pre-processing can increase the accuracy of ML in public health; 5. To propose an appropriate model to classify diabetes accurately in this ground; Materials and Methods In this study, the experimentation was conducted using the Pima Indian Dataset, obtained from various sources including ML Databases, ML Repository, and the NIDDK [ 45 ]–[ 47 ]. The dataset serves as the foundation for the this analysis, containing crucial insights into diabetes diagnosis. Comprising a total of 768 rows, the dataset encompasses both non-diabetic and diabetic subjects, with 500 instances falling into the former category and 268 instances into the latter. This binary classification is determined by an output column, with a value of 0 indicating absence of diabetes and a value of 1 denoting its presence. The dataset features nine columns encapsulating distinct attributes: pregnant month, glucose level, plasma concentration, blood pressure, triceps skinfold thickness, insulin amount, BMI, pedigree function, and the patient’s age. These attributes collectively contribute to the predictive analysis, aiding in the differentiation between individuals with and without diabetes. Conceptual Framework A conceptual framework is a theoretical structure comprised of interrelated concepts, serving as a foundation for understanding complex subjects, visual diagram, illustrating the sequential steps, decisions, and actions within a process or system. The conceptual framework and flowchart of the whole study is given below in the Fig. 1 : Data Pre-processing In order to extract usable information from this raw data and feed it into the training model for effective medical choices, diagnoses, and treatments, data pre-processing is a necessary step. To enhance the quality of the data, the pre-processing conducts a number of tasks, including outlier treatment, filling in missing values, data normalization, and feature selection. 500 samples in the dataset were identified as not having diabetes, compared to 268 diabetic samples. Label encoding is the first method employed. This method is used to determine if a person has diabetes or not, which is the dependent variable. As a result, the output class is determined by replacing all of the string values in the output variable with 0 and 1. The subsequent data pre-processing method uses the ML algorithm missForest to address missing values [ 48 ]. The dataset had many missing values for many variables, including 227 missing values for the skin thickness parameter, 374 missing values for the insulin attribute, and 111 missing values for pregnancy. Missing Values Treatment The elimination of the rows or columns with null values is one method of addressing missing values. The data analyst can remove the whole column if any columns contain more than 50% null values. Similar to how columns can be discarded if one or more of their values are null, rows can likewise be deleted. Imputation approach is the alternative method of treating missing values. If a column has fewer than 50% of its values be null, the missing value can be added to the entire column using the imputation approach. In this study used missForest [ 48 ] technique to input missing value. Let \(X\) be \(n\times p\) matrix of predictors that requires imputation: $$\varvec{X} = ( {\varvec{X}}_{1}, {\varvec{X}}_{2}, \dots , {\varvec{X}}_{\varvec{p}})= \left[\begin{array}{ccc}{\varvec{x}}_{11}& {\varvec{x}}_{12}& \begin{array}{cc}\dots & {\varvec{x}}_{1\varvec{p}}\end{array}\\ {\varvec{x}}_{21}& {\varvec{x}}_{22}& \begin{array}{cc}\dots & {\varvec{x}}_{2\varvec{p}}\end{array}\\ \begin{array}{c}⋮\\ {\varvec{x}}_{\varvec{n}1}\end{array}& \begin{array}{c}⋮\\ {\varvec{x}}_{\varvec{n}2}\end{array} & \begin{array}{c}⋮\\ {\dots \varvec{x}}_{\varvec{n}\varvec{p}}\end{array}\end{array}\right]$$ An arbitrary variable \({\varvec{X}}_{\varvec{s}}\) contains missing values at entries \({\varvec{i}}_{\varvec{m}\varvec{i}\varvec{s}}^{\left(\varvec{s}\right)} \subseteq \{1, 2, \dots , n\}\) . For every variable \({\mathbf{X}}_{\varvec{s}}\) that contains missing values, the researchers can separate the dataset into 4 categories: 1. The non-missing values of variable \({\mathbf{X}}_{s}\) , denoted by \({\mathbf{y}}_{\mathbf{o}\mathbf{b}\mathbf{s}}^{\left(\mathbf{s}\right)}\) 2. The missing values of variable \({\mathbf{X}}_{\varvec{s}}\) , denoted by \({\mathbf{y}}_{\mathbf{m}\mathbf{i}\mathbf{s}}^{\left(\mathbf{s}\right)}\) . 3. The variables other than \({\mathbf{X}}_{\varvec{s}}\) , with observations \({\mathbf{i}}_{\mathbf{o}\mathbf{b}\mathbf{s}}^{\left(\mathbf{s}\right)}\) = \(\frac{\left\{1, 2, \dots , n\right\}}{{\varvec{i}}_{\varvec{m}\varvec{i}\varvec{s}}^{\left(\varvec{s}\right)}}\) , denoted by \({\mathbf{x}}_{\mathbf{o}\mathbf{b}\mathbf{s}}^{\left(\mathbf{s}\right)}\) 4. The variables other than \({\mathbf{X}}_{\varvec{s}}\) , with observations \({\mathbf{i}}_{\mathbf{m}\mathbf{i}\mathbf{s}}^{\left(\mathbf{s}\right)}\) , denoted by \({\mathbf{x}}_{\mathbf{m}\mathbf{i}\mathbf{s}}^{\left(\mathbf{s}\right)}\) Outlier Treatment An outlier is a data point in statistics that dramatically deviates from other observations. An outlier may be caused by measurement variability, a sign of unique data, or an experimental error; the latter is occasionally eliminated from the data set. While an outlier may signal an intriguing potential, it may also seriously impair statistical analysis. In any distribution, outliers can happen by accident, but they can also point to unexpected behaviour or structures in the data set, measurement error, or a heavy-tailed distribution in the population. Heavy-tailed distributions suggest that the distribution has substantial skewness, and one should be extremely cautious when using tools or intuitions that presume a normal distribution. In the case of measurement error, one desires to reject them or use statistics that are resilient to outliers. A mixture of two distributions, which may represent two different subpopulations or may represent "right trial" vs "measurement mistake," is a frequent source of outliers and is represented by a mixture model. Most often, in bigger data samples, certain data points will be further distant from the sample mean than what is regarded as fair. It can be because some observations are distant from the centre of the data, inadvertent systematic mistake, or problems with the theory that produced the presumed family of probability distributions. Therefore, outlier points may hint to flawed data, flawed processes, or locations where a certain hypothesis may not hold true. However, it is normal to expect a few outliers in big datasets (and not due to any anomalous condition). Analysis Tools The researchers used R and RStudio for data processing and data analysis, data visualization, Microsoft Office for documentation, and Mendeley for reference management, creating a unified framework for in-depth exploration of diabetes risk factors using the Pima diabetes dataset. Statistical Tests for Association, Mean Comparison, and Odds Ratios The analysis employs a spectrum of tests to explore associations, mean disparities, and ORs across diverse scenarios. The Chi-Square Test, utilized for categorical variables in contingency tables, detects associations, particularly with small sample sizes or when chi-square assumptions are unmet. The t-test, a parametric counterpart to compare means, evaluates differences among various independent samples for continuous data. OR quantifies associations in case-control studies, while likelihood ratios compare event odds between groups; an OR above 1 indicates higher odds in the first group. AOR, within LR, gauges predictor-outcome connections while considering confounders, offering deeper insights. These tests, spanning categorical, continuous, and ordinal data, equip researchers with nuanced insights into variable relationships, enriching conclusions with comprehensive considerations. Classification Algorithms 1. NNET: NNs are AI techniques inspired by the human brain. DL, a form of ML, uses interconnected nodes to simulate brain structures. 2. AVNNET: AVNNET utilizes NNs with various seeds for prediction. Model scores are averaged for regression, and helps building more reliable models by decreasing training variance. 3. PCANNET: PCANNET uses principal component analysis to reduce data dimensions. NNs then employ these components for training and prediction, ensuring predictors have enough variability. 4. MULTINOM: MULTINOM extends bias reduction techniques to multinomial LR models. It employs penalized maximum likelihood estimates and profile confidence intervals for hypothesis testing. 5. RBF: Radial basis function (RBF) networks are popular for function approximation. They use radial basis functions as activation functions and find use in approximation, classification, prediction, and control. 6. RBFDDA: RBF networks with dynamic decay adjustment technique are simpler and specialize in categorization. They start with minimal parameters and add units gradually, making them easier to use than standard RBF networks. 7. RPART: Recursive partitioning is a tool for creating decision rules in data mining. The RPART generates classification and regression trees, aiding data exploration and predictive modelling. The resulting CART model provides user-friendly predictions based on predictor variables. 8. XGBLINEAR: XGBoost is a potent ML library with unique algorithms, each controlled by hyperparameters. It models tasks using trees, linear functions, and regularization techniques. These hyperparameters influence model behaviour and outcomes. 9. RRF: Regularized Random Forest is used for feature selection with a single ensemble. It evaluates features in tree nodes using subsets of training data, eliminating the need for data normalization. 10. LOGITBOOST: LogitBoost is a boosting classification algorithm similar to AdaBoost. Both perform additive LR, but LogitBoost minimizes logistic loss while AdaBoost minimizes exponential loss. 11. RANGER: Ranger is a fast implementation of recursive partitioning and random forests, suitable for large datasets. It supports classification, regression, and survival forests, including examples of quantile regression forests and very random trees. 12. RF: Random Forest is a popular supervised ML approach for classification and regression. It creates an ensemble of DTs on random subsets of input data, enhancing prediction accuracy through voting. 13. MLP: A multilayer perceptron is a fully connected feedforward ANN. It comprises at least three layers: input, hidden, and output. Nonlinear activation functions are used in the hidden and output layers, and backpropagation is employed for supervised learning. MLP can handle non-linearly separable data due to its multiple layers and non-linear activations. 14. MLPWEIGHTDECAY: In DL, models can generalize better through data augmentation. But what about during training? MLPWEIGHTDECAY introduces weight and decay parameters to enhance model generalization. 15. MLPML: Similar to MLP, MLPML consists of three layers: input, hidden, and output. It uses non-linear activation functions and backpropagation for training. This architecture allows MLPML to distinguish non-linearly separable data. 16. MLPWEIGHTDECAYML: MLPWEIGHTDECAYML is a fully connected feedforward ANN that includes settings for decay and weight parameters. MLPs are sometimes called "vanilla" NNs, especially with one hidden layer. The term "MLP" can refer broadly to any feedforward ANN or strictly to networks with multiple layers of perceptrons. Performance Metrics in Classification Performance metrics serve as vital tools for the evaluation of classification models, enabling researchers to gauge how effectively the model predicts outcomes compared to real values. A range of commonly used performance metrics exist to aid in this assessment. Sensitivity, also known as the True Positive Rate or Recall, quantifies the proportion of actual positive cases correctly identified by the model. Specificity measures the model’s ability to accurately pinpoint actual negative cases. Precision, or Positive Predictive Value, gauges the accuracy of positive predictions by calculating the ratio of true positives to the total predicted positives. Recall, akin to Sensitivity, indicates the model’s capacity to correctly predict actual positive cases. The F1 Score, a harmonic mean of precision and recall, offers a balanced evaluation, particularly valuable when dealing with class imbalances. AUC-ROC graphically represents the model’s discriminatory ability between positive and negative classes, providing a threshold-independent evaluation. Accuracy captures the overall correctness of predictions by calculating the ratio of correctly predicted cases to the total. Cohen’s Kappa adjusts accuracy by considering chance agreement, thus quantifying agreement between predicted and actual outcomes while accounting for random agreement. These diverse performance metrics collectively illuminate the strengths and weaknesses of classification models, offering insights tailored to analysis goals, dataset characteristics, and the desired balance between evaluation criteria. Results In this section, the authors present the key findings and outcomes of this study, which contribute to a deeper understanding of the data and diabetes risk factors with the hybrid data preparation technique. This research endeavors, guided by rigorous methodology and comprehensive analysis, have yielded valuable insights into the intricacies. This findings are presented in a clear and organized manner to facilitate comprehension and offer a foundation for further discussion and interpretation. Data Treatment In this research work, the Pima Indian Dataset has been considered for the experimentation [ 45 ]–[ 47 ]. The dataset comprises nine columns, and one output column has a binary value to indicate whether the subject has diabetes or not. It comprises 768 rows, 500 of which contain non-diabetics and 268 of which have diabetic patients. Nine feature columns, including pregnant month, glucose, plasma, blood pressure, fold thickness of the triceps skin, amount of insulin, BMI, pedigree function, age of patients, and one goal column, are included in the dataset (0 or 1). Missing Value Treatment In this dataset there are many missing values which are presented in the Table 2 : Table 2 Missing Values Variable No. of Missing Values Pregnant 0 Glucose 5 Pressure 35 Triceps 227 Insulin 374 Mass 11 Pedigree 0 Age 0 Diabetes 0 The Table 2 presents a dataset with various variables related to diabetes prediction, indicating the number of missing values for each variable. Notably, the "Pregnant," "Pedigree," and "Age" variables have no missing values, while others like "Insulin" and "Triceps" have a substantial number of missing values, suggesting incomplete data records. Addressing these missing values is crucial for accurate analysis and modelling. Researchers typically employ imputation or data pre-processing techniques to handle missing data and ensure the dataset's quality for predictive modelling or statistical analysis. Imputation approach is the alternative method of treating missing values. If a column has fewer than 50% of its values be null, the missing value can be added to the entire column using the imputation approach. In this study, the missing value was entered using the missForest method [ 48 ]. Before applying missForest, the missing plot is given below. From the Fig. 2 , the researchers may notice that, there is 9% missing values in the dataset. Now the researchers handle the dataset by using missForest. Outlier Treatment An outlier is a data point in statistics that dramatically deviates from other observations. An outlier may be caused by measurement variability, a sign of unique data, or an experimental error; the latter is occasionally eliminated from the data set. While an outlier may signal an intriguing potential, it may also seriously impair statistical analysis. The visualization of the presence of outliers by Mahalanobis squared distances to detect outliers in this dataset is given below. From the Fig. 3 , the researchers can suspect that there are many outliers in the dataset. The researcher’s aim is to treat the outliers. The visualization of the presence of outliers by Mahalanobis squared distances to detect outliers in this dataset is given below. From the Fig. 4 , the researchers can suspect that there are many outliers in the dataset. The researcher’s focus is to treat the outliers. In this study, the researchers replaced the detected outliers as missing values and predicted them with multistage hybrid ML with optimum number of iterations, where the researchers got 13% outliers in this dataset. From the Fig. 5 , there is 13% outliers in this dataset. After outlier treatment by the above method the visualization of the absence of outliers by Mahalanobis squared distances in this dataset is given below. From the Fig. 6 , the researchers can see there is no harmful/potential suspected outliers in the dataset. After outlier treatment by the above method the visualization of the absence of outliers by Mahalanobis squared distances in this dataset is given below. From the Fig. 7 , the researchers may declare that, now the dataset is free from the presence of the suspected outliers. Box Plot Box plots without outlier treatment are given below. This Fig. 8 of the box plots are drawn on the original dataset without outlier treatment. Box plots with outlier treatment are given below. From the Fig. 9 of the box plots, the researchers can see that after the treatment, there is no outlier is suspected. Mean Test t-Test p-values for each independent numerical variable with the levels of the dependent variable are presented in the Table 3 which is given below. Table 3 t-Test p-values for each Independent Variable Diabetes neg pos p Pregnant Mean (SD) 3.3 (3.0) 4.9 (3.7) < 0.001 Glucose Mean (SD) 110.6 (24.7) 142.4 (29.5) < 0.001 Pressure Mean (SD) 70.8 (12.0) 75.3 (12.0) < 0.001 Triceps Mean (SD) 27.0 (9.2) 32.5 (9.0) < 0.001 Insulin Mean (SD) 129.6 (83.3) 207.2 (103.9) < 0.001 Mass Mean (SD) 30.8 (6.5) 35.4 (6.6) < 0.001 Pedigree Mean (SD) 0.4 (0.3) 0.6 (0.4) < 0.001 Age Mean (SD) 31.2 (11.7) 37.1 (11.0) < 0.001 The Table 3 presents a comprehensive comparison between individuals with negative and positive cases of diabetes across various health-related variables. Notably, individuals with positive diabetes cases exhibit significantly higher average values in several key parameters, including glucose levels, insulin levels, triceps skinfold thickness, BMI, pedigree function, and age, when compared to those with negative cases. These stark differences are underscored by the extremely low p-values (< 0.001) associated with each variable, indicating strong statistical significance. Specifically, higher glucose and insulin levels highlight the metabolic impact of diabetes, while the elevated age and BMI among positive cases may suggest long-term health implications and potential risk factors. These findings collectively emphasize the importance of monitoring and managing these variables as critical aspects of diabetes prevention and management. To visualize the mean differences the means plots with levels of the dependent variable are given below. The grouped histograms are given below in the Fig. 11 : The grouped box-plots are given below in the Fig. 12 : Grouped density plots are given below in the Fig. 13 : According the above Table 3 , Fig. 10 , Fig. 11 , Fig. 12 , and Fig. 13 the researchers can say that, the mean differences between negative and positive levels of the variable Diabetes for any other numeric variables are statistically significant. The means for positive level of Diabetes is significantly higher than negative level. The Table 3 summarizes the outcomes of t-tests conducted to compare various health indicators between individuals with negative (no diabetes) and positive (diabetes) outcomes. The means and standard deviations of several features were examined for both groups. The results indicate significant differences in almost all parameters. Individuals with diabetes exhibit notably higher mean values for key factors including glucose level, blood pressure, triceps skinfold thickness, insulin level, BMI, pedigree score, and age. These disparities are all statistically significant, as evidenced by the p-values (< 0.001) accompanying each comparison. This suggests that these health indicators could potentially serve as meaningful differentiators between those affected by diabetes and those who are not. Correlation Matrix The correlation matrix is given below: From the Fig. 14 , the researchers may get an overview of the data. Trivariate Analysis Scatterplot matrix of the features of original data is given below in the Fig. 15 : Scatterplot matrix of the features of the outlier treated data is given below in the Fig. 16 : From the Fig. 14 , Fig. 15 , and Fig. 16 , the researchers found that here multicollinearity may arise. So classical modelling is not appropriate for this dataset. ML may be a suitable replacement of the classical models. Classifiers In this study the researchers used 16 classifiers to classify diabetes using PIMA dataset. The performance metrics are presented in the Table 4 . Table 4 Model Comparison Model Sensitivity Specificity Precision Recall F1 AUC Accuracy Kappa LogitBoost 0.8095 0.9464 0.85 0.8095 0.8293 0.7888 0.9091 0.7674 mlpWeightDecayML 0.7925 0.77 0.6462 0.7925 0.7119 0.8196 0.7778 0.534 avNNet 0.8302 0.75 0.6377 0.8302 0.7213 0.8442 0.7778 0.5418 mlpML 0.7736 0.78 0.6508 0.7736 0.7069 0.8228 0.7778 0.5301 ranger 0.717 0.8 0.6552 0.717 0.6847 0.8476 0.7712 0.5058 nnet 0.7736 0.77 0.6406 0.7736 0.7009 0.8411 0.7712 0.5183 RRF 0.7547 0.77 0.6349 0.7547 0.6897 0.84 0.7647 0.5024 RF 0.717 0.79 0.6441 0.717 0.6786 0.8458 0.7647 0.4938 multinom 0.7547 0.77 0.6349 0.7547 0.6897 0.8438 0.7647 0.5024 rbf 0.8113 0.74 0.6232 0.8113 0.7049 0.8387 0.7647 0.5148 xgbLinear 0.6981 0.79 0.6379 0.6981 0.6667 0.8134 0.7582 0.4775 rbfDDA 0.6792 0.8 0.6429 0.6792 0.6606 0.7742 0.7582 0.473 mlp 0.7925 0.73 0.6087 0.7925 0.6885 0.8355 0.7516 0.4878 pcaNNet 0.7925 0.72 0.6 0.7925 0.6829 0.8409 0.7451 0.4765 mlpWeightDecay 0.6792 0.76 0.6 0.6792 0.6372 0.7945 0.732 0.426 rpart 0.9057 0.6 0.5455 0.9057 0.6809 0.7696 0.7059 0.4377 The accuracy and kappa metrics are visualised in the Fig. 17 : From the Table 4 and Fig. 17 : Sensitivity measures the proportion of actual positive cases that are correctly identified by the model. It's also called the true positive rate or recall. A higher sensitivity indicates that the model is better at identifying positive cases. For instance, in the first row ("LogitBoost"), the sensitivity is 0.8095, which means that the model correctly identifies around 81% of the actual positive cases. Specificity measures the proportion of actual negative cases that are correctly identified as negative by the model. A higher specificity indicates that the model is better at identifying negative cases. For example, in the "LogitBoost" model, the specificity is 0.9464, which means that around 95% of the actual negative cases are correctly identified as negative. Precision measures the proportion of positive predictions made by the model that are actually correct. It's a measure of the accuracy of positive predictions. A higher precision indicates that the positive predictions are more likely to be accurate. In the "LogitBoost" model, the precision is 0.85, which means that around 85% of the positive predictions are correct. Recall is another term for sensitivity, as mentioned above. It's the proportion of actual positive cases that the model correctly identifies. The F1 score is the harmonic mean of precision and recall. It combines both measures and provides a balanced assessment of a model's performance. A higher F1 score indicates a good balance between precision and recall. The AUC represents the area under the Receiver Operating Characteristic (ROC) curve. The ROC curve plots the true positive rate against the false positive rate for different thresholds. A higher AUC indicates better overall discrimination ability of the model. Accuracy measures the proportion of correctly classified instances (both true positives and true negatives) out of all instances. It's a general measure of the model's correctness. Cohen's Kappa is a measure of agreement between predicted and actual classifications, considering the possibility of agreement by chance. A higher Kappa indicates a higher agreement between predicted and actual classifications than what would be expected by chance. These metrics collectively provide insights into how well each model is performing on the PIDD. The "LogitBoost" model seems to have relatively high sensitivity, specificity, precision, and F1 score, suggesting a good overall performance for this dataset. However, the choice of the best model also depends on the specific goals of the analysis and the trade-offs the researchers are willing to make between different evaluation measures. Comparative Study with LR The objective is to gauge the model's performance in relation to various factors, both directly linked to diabetes and those encompassing broader health considerations. By conducting this evaluation, the researchers seek to discern the relative strengths and weaknesses of LR in comparison to alternative approaches, shedding light on its suitability for diabetes diagnosis and its applicability to broader health-related inquiries. Table 5 OR from LR Model Diabetes Predictors OR std. Error CI Statistic p (Intercept) 0.00 0.00 0.00–0.00 -10.83 < 0.001 Age 1.01 0.01 0.99–1.03 1.18 0.237 Glucose 1.04 0.00 1.03–1.05 8.34 < 0.001 Insulin 1.00 0.00 1.00–1.00 0.44 0.663 Mass 1.09 0.02 1.05–1.14 4.26 < 0.001 Pedigree 2.36 0.70 1.32–4.24 2.89 0.004 Pregnant 1.13 0.04 1.06–1.21 3.85 < 0.001 Pressure 0.99 0.01 0.98–1.01 -0.88 0.377 Triceps 1.01 0.01 0.98–1.03 0.44 0.659 Observations 768 \({R}^{2}\) Tjur 0.336 AIC 729.301 log-Likelihood -355.650 From the Table 5 , the researchers obtained: Age: The OR is 1.01, which suggests that for a one-unit increase in age, the odds of having diabetes increase by 1.01 times. However, the p-value (0.237) indicates that the effect is not statistically significant, as it's greater than 0.05. Glucose: The OR is 1.04, and the p-value is < 0.001, indicating that higher glucose levels are associated with increased odds of diabetes in a statistically significant way. Insulin: The OR is 1.00, with a p-value of 0.663. This suggests that insulin levels don't have a significant effect on diabetes presence. Mass: The OR is 1.09, and the p-value is < 0.001. This indicates that higher mass (body weight) is associated with increased odds of diabetes significantly. Pedigree: The OR is 2.36, and the p-value is 0.004. This suggests that a higher pedigree (a measure of diabetes heredity) is associated with significantly higher odds of diabetes. Pregnant: The OR is 1.13, and the p-value is < 0.001. This indicates that being pregnant is associated with higher odds of diabetes. Pressure: The OR is 0.99, with a p-value of 0.377. Blood pressure doesn't seem to have a significant effect on diabetes. Triceps: The OR is 1.01, and the p-value is 0.659. Triceps skinfold thickness doesn't have a significant effect on diabetes. The \({R}^{2}\) Tjur is a measure of goodness-of-fit in the context of LR. It's used to evaluate how well the model fits the data, just like the R-squared ( \({R}^{2}\) ) statistic in linear regression. However, the \({R}^{2}\) Tjur is specifically designed for LR, where the response variable is binary. In essence, \({R}^{2}\) Tjur quantifies the proportion of explained variance in the dependent variable by the predictor variables in the model. A value of 0 means that the predictors have no explanatory power, and a value of 1 indicates that they completely explain the variability in the response. In the Table 5 , the value is 0.336, suggesting that the predictors in the model collectively explain about 33.6% of the variance in the presence of diabetes. The log-likelihood is a measure used to assess how well the statistical model fits the observed data. In the context of LR, it's a measure of how likely the observed outcomes are, given the parameter estimates of the model. Higher log-likelihood values indicate that the model is better at explaining the observed data. In the Table 5 , the log-likelihood value is -355.650. While the absolute value itself doesn't hold a direct interpretation, comparing log-likelihood values between different models can help to determine which model fits the data better. The AIC is a metric used to compare the quality of different models, taking into account both goodness of fit and model complexity. It's a tool for model selection, helping to choose the model that best balances explanatory power and parsimony. The AIC is calculated using the log-likelihood value and the number of parameters in the model. Lower AIC values indicate a better trade-off between model fit and complexity. In the Table 5 , the AIC value is 729.301, which can be used to compare this model to other models. Lower AIC values, when comparing models, suggest that the model with the lower AIC is a better fit for the data. In summary, these three metrics— \({R}^{2}\) Tjur, log-likelihood, and AIC—provide insights into different aspects of the LR model's performance. \({R}^{2}\) Tjur assesses the overall fit and explained variance, log-likelihood gauges how well the model fits the data, and AIC helps to compare different models based on both goodness of fit and complexity. Based on this analysis, higher glucose levels, higher body mass, higher pedigree, and pregnancy are associated with increased odds of diabetes, while other factors like age, insulin levels, blood pressure, and triceps thickness do not show a statistically significant relationship. From the Fig. 18 , the researchers can say, glucose, mass, pedigree, pregnancy have a positive relation with the chance of having diabetes. From the Fig. 19 , the researchers can say the high-risk factors are glucose, mass, pregnant, pedigree. The Fig. 20 of ROC curve, with an Area Under the Curve (AUC) of 84.6493, indicates that binary classification model exhibits strong discriminatory power. With an AUC significantly above the random guessing threshold of 0.5, it demonstrates an ability to effectively distinguish between positive and negative instances across various decision thresholds. While not perfect, this AUC score suggests that the model performs well on average, making it a promising tool for making binary predictions. Here, the researchers want to show variable importance according to the LR model and best subset model by backward stepwise elimination method. And the researchers have shown Odds plot to do the expected comparative analysis. Discussion This study employs a sophisticated approach to unravel the intricate web of factors contributing to the susceptibility of diabetes. Through the utilization of advanced statistical methodologies, including LR, ORs, and AORs, the investigation reveals profound insights into the predictive landscape of this health condition. The ORs estimated from the LR model in the current study are largely consistent with findings from prior analyses of the Pima Indian dataset [ 26 ], [ 27 ], [ 39 ], [ 49 ]–[ 52 ], [ 28 ], [ 29 ], [ 31 ]–[ 34 ], [ 37 ], [ 38 ]. Specifically, glucose level was found to be a significant predictor of diabetes with an OR of 1.04 (95% CI: 1.03–1.05), which falls within the range of 1.03–1.05 reported in previous studies examining this relationship [ 27 ], [ 29 ], [ 38 ]. Similarly, the OR for mass of 1.09 (95% CI: 1.05–1.14) is comparable to past estimates of 1.02–1.09 [ 32 ], [ 33 ]. The significant OR of 2.36 (95% CI: 1.32–4.24) for pedigree function aligns with the significance and magnitude of heredity measures observed in other investigations [ 28 ], [ 38 ]. Some variables like age, insulin level, and blood pressure showed non-significant or weaker associations here as in portions of the earlier literature [ 29 ], [ 34 ], [ 37 ], [ 49 ], [ 50 ]. Overall, the pattern of predictor significance and OR estimates replicate established relationships demonstrated across multiple prior LR analyses of the PIDD [ 26 ], [ 51 ], [ 52 ]. Therefore, the current results corroborate past research findings using for this well-studied population. A systematic methodology centered on robust pre-processing of missing data and outliers to develop diabetes have been implemented classification models using the Pima Indian dataset. Prior investigations typically omitted missing values or removed outliers without supplementation, which may have compromised resultant model accuracy [ 26 ], [ 27 ], [ 36 ], [ 38 ]–[ 40 ], [ 53 ], [ 28 ]–[ 35 ] (Summary of these studies are presented in Table 1 ). In contrast, this study applied the missForest imputation method to minimize distortions from missing data while retaining potentially informative outlier observations via temporary replacement with missing indicators followed by missForest imputation. This novel hybrid pre-processing approach strengthened data integrity relative to complete outlier exclusion. Among the evaluated ML algorithms, the LogitBoost classifier demonstrated particularly strong performance, achieving a sensitivity of 0.8095, specificity of 0.9464, and accuracy of 0.9091 (From Table 4 ). These metrics represent improvements over past LogitBoost [ 27 ], [ 28 ], [ 38 ] and other single model implementations such as DT, NN [ 33 ], SVM [ 36 ], and DL [ 30 ] for this dataset, according to available reporting. While direct comparison must be made cautiously, consistently superior performance across multiple predictive metrics provides evidence the systematic pre-processing protocols implemented herein may enhance diabetes classification. A key gap addressed was the lack of standardized procedures for missing and anomalous data handling prior to modelling, which had been noted as potentially compromising accuracy in earlier studies. Through meticulous data pre-processing centered on judicious imputation and outlier management followed by rigorous model selection and validation, the current study substantively enhanced data integrity and strengthened resultant predictive performance relative to the past efforts studied herewith. Comprehensive model evaluation further allowed elucidation of appropriate algorithmic choices. A noteworthy contribution of the current study was the development and implementation of a systematic and rigorous framework for strategic pre-processing of diabetes-related clinical data prior to modelling. By judiciously applying the missForest imputation technique and thoughtfully incorporating outliers through a temporary replacement approach, critical missing and anomalous data points were retained without compromising dataset integrity. This enabled more accurate identification and extraction of informative patterns from the data compared to previous methods involving complete outlier removal or omission of missing values. Comprehensive evaluation of diverse ML algorithms further permitted elucidation of optimal techniques for classification. Critically, the combined impact of these strategic pre-processing protocols and hybrid data-driven techniques elevate this research above prior work, addressing longstanding limitations to substantially advance both understanding and predictive performance for early diabetes detection using ML. Therefore, the current study makes a significant methodological contribution toward maximizing the potential of supervised learning techniques applied to clinical datasets. The methodical pre-processing framework and hybrid data-driven techniques applied here notably elevated diabetes risk prediction for this seminal cohort above prior levels, addressing longstanding weaknesses to inform best practices maximizing ML accuracy from this dataset. The research fundamentally advances understanding of diabetes classification. Conclusion In summary, this study implemented a rigorous methodological framework centering on strategic pre-processing to enhance ML -based diabetes classification utilizing the Pima Indian dataset. By judiciously applying missForest imputation and thoughtfully incorporating outliers through a temporary replacement approach, critical missing and anomalous data were retained without compromising data integrity. Comprehensive evaluation of multiple algorithms then permitted elucidation of optimal techniques for predictive model development. Critically, these systematic pre-processing protocols and hybrid data-driven techniques substantively strengthened data quality and classification performance compared to previous investigations that typically excluded incomplete observations. Notably, the LogitBoost classifier emerged as highly accurate for prediction based on its consistent achievement of sensitivity of 0.8095, specificity of 0.9464, and accuracy of 0.9091 - improvements over past single model implementations for this cohort. Overall, the research addresses longstanding weaknesses in diabetes risk prediction through its novel, carefully optimized methodology. The workflow developed makes a significant advancement by maximizing extraction of meaningful patterns from clinical health records to inform early disease detection. Finally, this work evidences the merit of implementing standardized, rigorous data pre-processing centered on judicious imputation and outlier management prior to predictive modeling. The protocol introduced herein fundamentally advances both understanding and capacity for accurate ML -based identification of diabetes risk. Future research employing these systematic techniques on additional patient cohorts will continue to elucidate best practices and algorithmic choices optimizing predictive healthcare applications. List of abbreviations Below is an alphabetically ordered list of abbreviations utilized in this study: AI: Artificial Intelligence ACC: Accuracy ANN: Artificial Neural Network ANOVA: Analysis of Variance AOR: Adjusted Odds Ratio AUC: Area Under the Curve avNNet: Model Averaging Neural Network BAL: Balanced BAYESGLM: Bayesian Generalized Linear Model BMI: Body Mass Index CNN: Convolutional Neural Network CV: Cross Validation Direct-CNN: Direct Convolutional Neural Network DL: Deep Learning DT: Decision Tree FDA: Flexible Discriminant Analysis gcvEarth: Multivariate Adaptive Regression Splines with Generalized Cross-Validation HRF: High Risk Factor HTN: Hypertension IDE: Integrated Development Environment lda: Linear Discriminant Analysis LR: Logistic Regression ML: Machine Learning mlpML: Multilayer Perceptron with Multiple Layers NN/nnet: Neural Network NIDDK: National Institute of Diabetes and Digestive and Kidney Disease OR: Odds Ratio PIDD: Pima Indian Dataset PNN: Probabilistic Neural Network RF: Random Forest RRFglobal: Random Forest using the Global Feature Selection Algorithm ROC: Receiver Operating Characteristic SVM: Support Vector Machine svmPoly: Support Vector Machine with Polynomial Kernel SVM-RFE: Support Vector Machine-Recursive Feature Elimination Declarations Consent for Publication In obtaining consent for publication, the researchers ensured transparency and respect for the contributions of all involved parties. Authors collectively agreed to the submission of this work for potential publication, and any potential conflicts of interest were disclosed. This collaborative approach aligns with the author’s commitment to responsible and accountable research dissemination. Competing Interests The authors declare no competing interests in relation to the research, authorship, and publication of this paper. This ensures that the study’s outcomes and interpretations are solely driven by scientific rigor and unbiased exploration of the topic at hand. Funding This research was self-funded by the authors, reflecting their commitment to advancing the understanding of diabetes prediction through rigorous investigation and analysis. Author’s Contributions The collaborative research on diabetes was led by three key contributors: MSH, AD, and PKK. MSH conceived and designed the study, conducted all methodological work including data pre-processing, ML model development and evaluation, statistical analysis, visualization, writing of the initial manuscript draft, and revisions. PKK and AD contributed to interpretation of results and revisions to the final manuscript. All authors discussed the results, contributed to the intellectual content, and approved the final manuscript. The distribution of tasks among the authors effectively utilized their expertise, resulting in a comprehensive and well-rounded research paper. Furthermore, the meticulous report writing was primarily undertaken by MSH, ensuring the clarity, coherence, and accuracy of the manuscript’s content. Acknowledgements The authors would like to thank Late Professor Dr. Md. Jahanur Rahman for his support and initial ideas of this work. The authors also appreciate the efforts of the reviewers who provided constructive comments that helped improve the quality of this manuscript. Finally, authors are grateful to Agro Environmetrics Research Lab for providing Lab facilities and Computers during the research and writing process. Authors’ Information Md. Sifat Hossain holds an MSc degree and MPhil Research Fellow in the Department of Statistics at the University of Rajshahi, Bangladesh. With expertise in data analysis and a keen interest in healthcare research, he contributed significantly to the statistical analysis and interpretation of this study. Email: [email protected] Astami Devnath possesses an MSc degree and affiliated as an MPhil Research Fellow at the Institute of Environmental Science, University of Rajshahi, Bangladesh. Her expertise is particularly focused on the domains of econometrics and environmental health. Email: [email protected] Provash Kumar Karmokar, PhD is a Professor in the Department of Statistics at the University of Rajshahi, Bangladesh. With a wealth of knowledge in statistical methodologies and health economics, he played a pivotal role in guiding the statistical analysis and ensuring its rigor. Email: [email protected] References W. H. Herman, “The global burden of diabetes: an overview,” Diabetes Mellit. Dev. Ctries. underserved communities , pp. 1–5, 2017. “World Health Organization Diabetes.” https://www.who.int/health-topics/diabetes. “How is the pancreas involved in diabetes?” https://www.medicalnewstoday.com/articles/325018#how-is-the-pancreas-linked-with-diabetes. “Type 2 Diabetes Causes and Risk Factors.” https://www.webmd.com/diabetes/diabetes-causes. “Prediabetes – Your Chance to Prevent Type 2 Diabetes.” https://www.cdc.gov/diabetes/basics/prediabetes.html. “National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK).” https://www.niddk.nih.gov/health-information/diabetes. “Blood Sugar Level Ranges.” https://www.diabetes.co.uk/diabetes_care/blood-sugar-level-ranges.html. “Diabetes - Long-Term Effects.” https://www.betterhealth.vic.gov.au/health/conditionsandtreatments/diabetes-long-term-effects. S. A. Kaveeshwar and J. Cornwall, “The current state of diabetes mellitus in India,” Australas. Med. J. , vol. 7, no. 1, p. 45, 2014. J. Chaki, S. T. Ganesh, S. K. Cidham, and S. A. Theertan, “Machine learning and artificial intelligence based Diabetes Mellitus detection and self-management: A systematic review,” J. King Saud Univ. Inf. Sci. , vol. 34, no. 6, pp. 3204–3225, 2022. C.-L. Huang, M.-C. Chen, and C.-J. Wang, “Credit scoring with a data mining approach based on support vector machines,” Expert Syst. Appl. , vol. 33, no. 4, pp. 847–856, 2007. I. Contreras and J. Vehi, “Artificial intelligence for diabetes management and decision support: literature review,” J. Med. Internet Res. , vol. 20, no. 5, p. e10775, 2018. G. Swapna, S. Kp, and R. Vinayakumar, “Automated detection of diabetes using CNN and CNN-LSTM network and heart rate signals,” Procedia Comput. Sci. , vol. 132, pp. 1253–1262, 2018. M. W. Craven and J. W. Shavlik, “Using neural networks for data mining,” Futur. Gener. Comput. Syst. , vol. 13, no. 2–3, pp. 211–229, 1997. J. D. B. Gil, P. Reidsma, K. Giller, L. Todman, A. Whitmore, and M. van Ittersum, “Sustainable development goal 2: Improved targets and indicators for agriculture and food security,” Ambio , vol. 48, no. 7, pp. 685–698, 2019. M. Lee et al. , “How to respond to the fourth industrial revolution, or the second information technology revolution? Dynamic new combinations between technology, market, and society through open innovation,” J. Open Innov. Technol. Mark. Complex. , vol. 4, no. 3, p. 21, 2018. N. Ahmed et al. , “Machine learning based diabetes prediction and development of smart web application,” vol. 2, pp. 229–241, 2021. Z. Dong et al. , “Prediction of 3-year risk of diabetic kidney disease using machine learning based on electronic medical records,” vol. 20, no. 1, pp. 1–10, 2022. Y. Du, A. R. Rafferty, F. M. McAuliffe, L. Wei, and C. %J S. R. Mooney, “An explainable machine learning-based clinical decision support system for prediction of gestational diabetes mellitus,” vol. 12, no. 1, pp. 1–14, 2022. H. Gupta, H. Varshney, T. K. Sharma, N. Pachauri, O. P. %J C. Verma, and I. Systems, “Comparative performance analysis of quantum machine learning with deep learning for diabetes prediction,” vol. 8, no. 4, pp. 3073–3087, 2022. J. J. Khanam and S. Y. %J I. C. T. E. Foo, “A comparison of machine learning algorithms for diabetes prediction,” vol. 7, no. 4, pp. 432–439, 2021. S. Kumari, D. Kumar, and M. %J I. J. of C. C. in E. Mittal, “An ensemble approach for classification and prediction of diabetes mellitus using soft voting classifier,” vol. 2, pp. 40–46, 2021. H. Lu, S. Uddin, F. Hajati, M. A. Moni, and M. %J A. I. Khushi, “A patient network-based machine learning model for disease prediction: The case of type 2 diabetes mellitus,” vol. 52, no. 3, pp. 2411–2422, 2022. P. Rajendra, S. %J C. M. Latifi, and P. in B. Update, “Prediction of diabetes using logistic regression and ensemble techniques,” vol. 1, p. 100032, 2021. M. Ravaut et al. , “Predicting adverse outcomes due to diabetes complications with machine learning using administrative health data,” vol. 4, no. 1, pp. 1–12, 2021. J. W. Smith, J. E. Everhart, W. C. Dickson, W. C. Knowler, and R. S. Johannes, “Using the ADAP learning algorithm to forecast the onset of diabetes mellitus,” Proc. Symp. Comput. Appl. Med. Care , vol. 12, no. 1, pp. 261–265, 1988. H. M. El-Bakry, M. El-Dib, and A. Kamal, “A comparative study of machine learning techniques for diabetes disease prediction,” in 2019 8th International Conference on Computer and Knowledge Engineering (ICCKE) , 2019, pp. 179–184. P. V Lakshmi and N. Chilamkurti, “Deep belief network based ensemble classifier for diabetes disease prediction,” Comput. Electr. Eng. , vol. 72, pp. 418–430, 2019. Y. Shang, Z. Chen, and G. Jiang, “A deep learning model to predict diabetes through electronic health records,” IEEE Access , vol. 7, pp. 54445–54452, 2019. C. S. Preetha, “A comparative study of data mining techniques for prediction of diabetes,” arXiv Prepr. arXiv1211.5730 , 2012. K. S. Abdul Nazeer and M. P. Sebastian, “Detecting diabetes on set of biological data using decision tree algorithm,” Far East J. Theor. Stat. , vol. 27, no. 1, pp. 1–14, 2009. Ö. N. Geomat, “Classification of Pima Indian diabetes dataset using neural networks,” Procedia-Social Behav. Sci. , vol. 195, pp. 1408–1417, 2015. M. Güler and K. Polat, “Detecting Pima Indians diabetes using neural networks and estimated statistical classification functions,” Expert Syst. Appl. , vol. 28, no. 4, pp. 707–715, 2005. M. A. Tahir, A. Bouridane, and F. Kurugollu, “Classifying medical data using SVM with combined kernel functions,” J. Appl. Clin. Med. Phys. , vol. 12, no. 1, p. 3475, 2011. M. F. Akay, “Support vector machines combined with feature selection for breast cancer diagnosis,” Expert Syst. Appl. , vol. 36, no. 2, pp. 3240–3247, 2009. P. Vepakomma, O. Gupta, A. Dewan, and P. Roux, “Reducing disparity in diabetes prediction models using adversarial representation learning,” arXiv Prepr. arXiv1807.00540 , 2018. R. Agrawal and A. Choudhary, “Comparison of supervised machine learning algorithms for disease prediction,” IOSR J. Comput. Eng. , pp. 18–24, 2016. N. Zhang and J. Li, “Pima Indians diabetes prediction based on ant colony optimization classifier,” in Lecture Notes in Computer Science , vol. 3173, Springer, 2003, pp. 230–235. M. S. A. Kumar, V. Ravi, and K. B. Raja, “Prediction of diabetes using probabilistic neural network with feature extraction,” Measurement , vol. 71, pp. 53–60, 2015. İ. Karabulut, “Comparison of generalized regression neural network algorithms for diabetes disease diagnosis,” J. Intell. Syst. , vol. 22, no. 2, pp. 247–256, 2013. T. N. Joshi and P. M. Chawan, “Logistic regression and svm based diabetes prediction system,” Int. J. Technol. Res. Eng. , vol. 5, pp. 4347–4350, 2018. N. Yuvaraj and K. R. SriPreethaa, “Diabetes prediction in healthcare systems using machine learning algorithms on Hadoop cluster,” Cluster Comput. , vol. 22, no. Suppl 1, pp. 1–9, 2019. D. Sisodia and D. S. Sisodia, “Prediction of diabetes using classification algorithms,” Procedia Comput. Sci. , vol. 132, pp. 1578–1585, 2018. E. O. Olaniyi and K. Adnan, “Onset diabetes diagnosis using artificial neural network,” Int J Sci Eng Res , vol. 5, no. 10, pp. 754–759, 2014. “Machine Learning Databases.” ftp://ftp.ics.uci.edu/pub/machine-learning-databases. “Machine Learning Repository.” http://www.ics.uci.edu/~mlearn/MLRepository.html. “National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK).” https://www.niddk.nih.gov/health-information/diabetes. D. J. Stekhoven and P. %J B. Bühlmann, “MissForest—non-parametric missing value imputation for mixed-type data,” vol. 28, no. 1, pp. 112–118, 2012. A. G. Karegowda, A. S. Manjunath, and M. A. Jayaram, “Comparative study of attribute selection using GA,” Int. J. Adv. Soft Comput. its Appl. , vol. 2, no. 1, pp. 45–68, 2010. X. Wu et al. , “Top 10 algorithms in data mining,” Knowl. Inf. Syst. , vol. 14, no. 1, pp. 1–37, 2008. M. Ahmed, A. N. Mahmood, and M. R. Islam, “A machine learning approach for early diagnosis of diabetes disease,” in 2012 international conference on informatics, electronics & vision (ICIEV) , 2012, pp. 1–5. S. Raschka, Python Machine Learning . Packt Publishing Ltd, 2018. D. Koley and D. Saha, “Comparative study of supervised machine learning algorithms for Pima indians diabetes data set,” J. Theor. Appl. Inf. Technol. , vol. 95, no. 16, 2016. Additional Declarations There is NO Competing Interest. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3364064","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":234380189,"identity":"319b61f4-10b4-44cd-b275-b82588e9adef","order_by":0,"name":"Md. Hossain","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA4klEQVRIie2PsQrCMBCGTw7SJdL1BLGvEBCK4Ms4dWp3B8ki6CL6CL6CLs6Bw0w+QEdBcHASCuIkRkXc2roJ5ht+Evg/7g7A4/lFsLF/v4wLatdQUIHrEggxeCiyzpi3ItXjW61EU2wUZ9A6XM6KUz7qSQh4uypTFCOSASayzU0/tW4xmSR5qYKhcYohcEo3FU4hGZcq0RjxakBTZOWxm95qKMAo3BQkZSUeskkNxd0iejvFrbVNYszmJEXVLdGCMR8OddhhPhTpRXfCgG35Yq9ZzxT0zOr6Bzx/0/Z4PJ7/4Q4b/j+0pfCVsAAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0009-0000-7265-5143","institution":"Agro Environmetrics Research Lab, Department of Statistics, University of Rajshahi","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Md.","middleName":"","lastName":"Hossain","suffix":""},{"id":234380190,"identity":"89f0ae19-6cba-4e1d-9919-587922a234fe","order_by":1,"name":"Astami Devnath","email":"","orcid":"","institution":"Agro Environmetrics Research Lab, Department of Statistics, University of Rajshahi","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Astami","middleName":"","lastName":"Devnath","suffix":""},{"id":234380191,"identity":"565a3757-6370-4f8f-a253-5a6c108ab7d1","order_by":2,"name":"Provash Karmokar","email":"","orcid":"","institution":"Agro Environmetrics Research Lab, Department of Statistics, University of Rajshahi","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Provash","middleName":"","lastName":"Karmokar","suffix":""}],"badges":[],"createdAt":"2023-09-17 22:00:28","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3364064/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3364064/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":43648296,"identity":"a75d8850-44b1-402a-b442-eb149aad154b","added_by":"auto","created_at":"2023-09-25 17:07:59","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":43653,"visible":true,"origin":"","legend":"\u003cp\u003eConceptual Framework\u003c/p\u003e","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/8f4246bb225405c0cbda9dc3.jpg"},{"id":43647857,"identity":"9b85194f-4d87-4d64-b8d0-fd157c7e6cda","added_by":"auto","created_at":"2023-09-25 16:59:59","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":56013,"visible":true,"origin":"","legend":"\u003cp\u003eMissing Plot\u003c/p\u003e","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/c2e3bcc89a37d7c92d1e621d.jpg"},{"id":43648447,"identity":"f16e30cc-4530-4482-bbb1-e764fda601d0","added_by":"auto","created_at":"2023-09-25 17:15:59","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":39535,"visible":true,"origin":"","legend":"\u003cp\u003eMahalanobis Squared Distances to Detect Outliers\u003c/p\u003e","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/2b9d8cbd5411e1fc8520d788.jpg"},{"id":43647862,"identity":"2b3fdd5f-c676-44c6-b342-78729ebba488","added_by":"auto","created_at":"2023-09-25 17:00:00","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":122371,"visible":true,"origin":"","legend":"\u003cp\u003eShowing the Presence of Outliers\u003c/p\u003e","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/880fb79ca4e1c3eeabe4125c.jpg"},{"id":43648449,"identity":"65009d1f-cd52-42b0-bcff-8aa9a7809f3d","added_by":"auto","created_at":"2023-09-25 17:16:00","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":87608,"visible":true,"origin":"","legend":"\u003cp\u003eOutlier Map as Missingness\u003c/p\u003e","description":"","filename":"5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/feddb723354c3b7e2f981d33.jpg"},{"id":43647858,"identity":"4d5d81e9-9a73-47a4-9883-8654510a8a41","added_by":"auto","created_at":"2023-09-25 16:59:59","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":38036,"visible":true,"origin":"","legend":"\u003cp\u003eMahalanobis Squared Distances to Show the Absence of Outliers\u003c/p\u003e","description":"","filename":"6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/f40235f72a6901f001d5d81c.jpg"},{"id":43647867,"identity":"a3991bc5-b4d6-47d1-84f4-d8e07a0ad54a","added_by":"auto","created_at":"2023-09-25 17:00:00","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":193243,"visible":true,"origin":"","legend":"\u003cp\u003eShowing the Absence of Outliers\u003c/p\u003e","description":"","filename":"7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/2bc17c481d259919f1e29f3d.jpg"},{"id":43648303,"identity":"4c66be78-dd83-4336-8c43-b8867f50cd5f","added_by":"auto","created_at":"2023-09-25 17:08:00","extension":"jpg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":114440,"visible":true,"origin":"","legend":"\u003cp\u003eBox Plots without Outlier Treatment\u003c/p\u003e","description":"","filename":"8.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/95cf72d270549fef6d471277.jpg"},{"id":43648641,"identity":"f9ed4f0a-e541-4ca7-b8d7-d3e16d8646f3","added_by":"auto","created_at":"2023-09-25 17:23:59","extension":"jpg","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":119643,"visible":true,"origin":"","legend":"\u003cp\u003eBox Plots with Outlier Treatment\u003c/p\u003e","description":"","filename":"9.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/0de1064b1c2ac7be10626d00.jpg"},{"id":43647875,"identity":"4e8ec448-0fc0-4cf8-aa55-8865542e4d20","added_by":"auto","created_at":"2023-09-25 17:00:00","extension":"jpg","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":158009,"visible":true,"origin":"","legend":"\u003cp\u003eMeans Plot\u003c/p\u003e","description":"","filename":"10.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/0a5b717252ae2a7305735b60.jpg"},{"id":43648299,"identity":"cba31d25-b39a-4894-b227-25dcfa4d8f26","added_by":"auto","created_at":"2023-09-25 17:08:00","extension":"jpg","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":101682,"visible":true,"origin":"","legend":"\u003cp\u003eGrouped Histogram\u003c/p\u003e","description":"","filename":"11.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/5c900f8c3f908f3a3cfda564.jpg"},{"id":43648304,"identity":"2a553de8-5abf-42de-a451-4d2c13f68690","added_by":"auto","created_at":"2023-09-25 17:08:00","extension":"jpg","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":158442,"visible":true,"origin":"","legend":"\u003cp\u003eGrouped Box Plot\u003c/p\u003e","description":"","filename":"12.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/86f956af3131946a9497ccd8.jpg"},{"id":43647877,"identity":"cc08712c-1b5f-4a7d-b935-bf4959c37b8b","added_by":"auto","created_at":"2023-09-25 17:00:00","extension":"jpg","order_by":13,"title":"Figure 13","display":"","copyAsset":false,"role":"figure","size":222951,"visible":true,"origin":"","legend":"\u003cp\u003eGrouped Density Plot\u003c/p\u003e","description":"","filename":"13.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/ab831ab27ad7eaa951853ce1.jpg"},{"id":43648450,"identity":"dbb50d60-0da7-49d3-be1b-a242cb002de0","added_by":"auto","created_at":"2023-09-25 17:16:00","extension":"jpg","order_by":14,"title":"Figure 14","display":"","copyAsset":false,"role":"figure","size":56545,"visible":true,"origin":"","legend":"\u003cp\u003eCorrelation Matrix\u003c/p\u003e","description":"","filename":"14.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/98cdddf3ad277df13b436636.jpg"},{"id":43648452,"identity":"ed81f6d1-1a21-4b4a-b8ae-ed66cd08ddad","added_by":"auto","created_at":"2023-09-25 17:16:00","extension":"jpg","order_by":15,"title":"Figure 15","display":"","copyAsset":false,"role":"figure","size":407451,"visible":true,"origin":"","legend":"\u003cp\u003eVisualization of the Original Data\u003c/p\u003e","description":"","filename":"15.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/52c31c4b518f08e01f0001fe.jpg"},{"id":43648307,"identity":"12ac0765-2389-4687-a5db-67b058f19570","added_by":"auto","created_at":"2023-09-25 17:08:00","extension":"jpg","order_by":16,"title":"Figure 16","display":"","copyAsset":false,"role":"figure","size":479525,"visible":true,"origin":"","legend":"\u003cp\u003eVisualization of the Data After Outlier Treatment\u003c/p\u003e","description":"","filename":"16.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/145e8d27f264e72fcd020501.jpg"},{"id":43647876,"identity":"2fd8aa24-a1d9-41a1-aa02-02dd36b0db7c","added_by":"auto","created_at":"2023-09-25 17:00:00","extension":"png","order_by":17,"title":"Figure 17","display":"","copyAsset":false,"role":"figure","size":42460,"visible":true,"origin":"","legend":"\u003cp\u003eAccuracy and Kappa Visualization without Hybrid Transformation\u003c/p\u003e","description":"","filename":"17.png","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/d62ea3e5787fd6a4b880d44a.png"},{"id":43648453,"identity":"91e508f7-1048-4bde-912f-0bf892297107","added_by":"auto","created_at":"2023-09-25 17:16:00","extension":"jpg","order_by":18,"title":"Figure 18","display":"","copyAsset":false,"role":"figure","size":178747,"visible":true,"origin":"","legend":"\u003cp\u003eVisualization of LR Modell\u003c/p\u003e","description":"","filename":"18.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/e20fbeabd020643e07e19771.jpg"},{"id":43648642,"identity":"9474f6c0-e176-4859-b237-2cea5b898af0","added_by":"auto","created_at":"2023-09-25 17:24:00","extension":"jpg","order_by":19,"title":"Figure 19","display":"","copyAsset":false,"role":"figure","size":31089,"visible":true,"origin":"","legend":"\u003cp\u003eVariable Importance from LR\u003c/p\u003e","description":"","filename":"19.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/511a0051cb2fa3f464824228.jpg"},{"id":43647871,"identity":"79d5df57-52df-413f-9c3a-e8baf38e1ad6","added_by":"auto","created_at":"2023-09-25 17:00:00","extension":"jpg","order_by":20,"title":"Figure 20","display":"","copyAsset":false,"role":"figure","size":34395,"visible":true,"origin":"","legend":"\u003cp\u003eROC Curve of LR Model\u003c/p\u003e","description":"","filename":"20.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/e4a7740604046fddb4ebce3a.jpg"},{"id":43648305,"identity":"b6aee306-1fa6-476a-908d-82039966cf42","added_by":"auto","created_at":"2023-09-25 17:08:00","extension":"jpg","order_by":21,"title":"Figure 21","display":"","copyAsset":false,"role":"figure","size":36435,"visible":true,"origin":"","legend":"\u003cp\u003eOR Plot of Stepwise LR\u003c/p\u003e","description":"","filename":"21.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/c9206e2354c3a4c1c375d8b3.jpg"},{"id":43706303,"identity":"57edea24-4a26-4d4d-b7ce-80f12fed2707","added_by":"auto","created_at":"2023-09-26 14:56:00","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2291711,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3364064/v1/8214028e-936d-4f9b-8ae0-f2f476c51fd6.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Handling Missing Values and Outliers in Advanced Data Pre-processing: An Enhancement of Diabetes Classification Accuracy","fulltext":[{"header":"Background","content":"\u003cp\u003eThe escalating global prevalence of diabetes underscores the urgency of timely detection and intervention to mitigate its far-reaching consequences [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The integration of ML has emerged as a promising avenue for early diagnosis, attracting significant attention from researchers and healthcare professionals alike. However, accurate disease prediction hinges not only on advanced technological tools but also on the meticulous handling of underlying data complexities, including missing values and outliers.\u003c/p\u003e \u003cp\u003eIn this context, the present study positions itself as a critical nexus between medical expertise and cutting-edge ML technology. The driving force behind this endeavour is to develop a comprehensive diabetes classification model that not only leverages the potential of ML but also addresses the pivotal challenge of managing missing data and outliers. These issues are particularly salient in medical datasets, where the quality of predictions can be undermined by data irregularities.\u003c/p\u003e \u003cp\u003eAs the study delves into these complexities, it sheds light on the paramount importance of data pre-processing. It seeks to establish a methodology that not only enhances the quality of analysis but also ensures the robustness of diabetes classification models. By focusing on the strategic treatment of missing values and outliers, the study aims to achieve improved accuracy in the early identification of diabetes, thereby contributing to effective preventive measures and optimized patient care.\u003c/p\u003e \u003cp\u003eThe significance of this research lies not only in its potential to refine disease classification techniques but also in its broader implications for healthcare decision-making. The amalgamation of medical insights with advanced ML technology is poised to revolutionize the early detection of diabetes, which, in turn, can lead to better patient outcomes and reduced healthcare burdens. As the study navigates the intricate terrain of data pre-processing, modelling, and analysis, it strives to illuminate a pathway toward more accurate and reliable diabetes prediction, amplifying the impact of both medical research and technological innovation.\u003c/p\u003e \u003cp\u003eThe urgent need for a predictive tool to aid early disease detection and recommend lifestyle changes is evident in the context of diabetes. With diabetes claiming approximately 1.6\u0026nbsp;million lives annually, its significance cannot be understated [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. Diabetes occurs when blood glucose levels become excessively high, stemming from either insufficient insulin production or poor insulin utilization by the body's cells [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eGlucose, produced during digestion, is regulated by insulin, which facilitates glucose absorption into cells for energy. Inadequate insulin results in glucose build-up in the blood, leading to elevated blood glucose levels and diabetes [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. High blood glucose manifests in symptoms like excessive thirst and frequent urination. A typical adult's glucose range is 70 to 99 mg/dl, while values exceeding 126 mg/dl indicate diabetes [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Persistent high blood glucose can trigger severe complications such as heart disease, renal failure, stroke, and nerve damage [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eDespite advancements, diabetes remains incurable, with long-term cases leading to macrovascular and microvascular complications. Macrovascular issues involve large blood arteries, while microvascular complications impact smaller blood vessels, contributing to kidney, eye, foot, and nerve complications [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Early detection and management are vital to curbing diabetes progression [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Exercise and dietary habits play pivotal roles in preventing and managing diabetes [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eHealthcare data, encompassing patient records and examination findings, holds the potential for predictive insights. Automation powered by ML is essential for detecting hidden patterns, enabling more accurate decision-making in diabetes diagnosis [\u003cspan additionalcitationids=\"CR11\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]\u0026ndash;[\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. ML and data mining technologies have transformed the healthcare landscape, aiding in feature selection and automation of diabetes prediction [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. These techniques uncover patterns in complex datasets, fostering reliable decision-making [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eCurrent research, however, often lacks advanced missing value treatment methods and thorough evaluation of ML algorithms. The researchers propose using supervised and unsupervised ML techniques, as well as AI and DL, incorporating non-parametric missing value imputation using random forests. The researchers aim to enhance diabetes categorization accuracy across various metrics, ultimately establishing a robust predictive model for diabetes classification.\u003c/p\u003e \u003cp\u003eTo fulfil the SDG-03: \u0026ldquo;Ensure healthy lives and promote well-being for all at all ages\u0026rdquo; and to cope pace with the artificial intelligence-driven fourth industrial revolution which is \u0026ldquo;A technological shift affecting cultures and economies all over the globe\u0026rdquo;, this study must be a dire need worldwide, especially for Bangladesh [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e], [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. At present, diabetes has increasingly been listed in the top position as a major reason for death. International Diabetes Federation declared that 382\u0026nbsp;million people are breathing with diabetes globally like Bangladesh. Diabetes if untreated may turn deadly and directly or indirectly invites a lot of other diseases. However, diabetes is largely avoidable and can be avoided by lifestyle changes. These changes may also lower the probability of developing cancer and heart disease. Diabetes imposes an incredible socioeconomic burden on patients and the family unit. Diabetes disease-associated expenditures influence the family unit\u0026rsquo;s everyday work and further compliance, which is directly halting the fulfillment of SDG-03. All knows \u0026ldquo;prevention is better than cure\u0026rdquo;. So, if all is known the proper lifestyle and ready to accept the lifestyle, then it may reduce the rate of diabetes patients, which will be helpful for both socioeconomic and family welfare.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSome Related Studies, Their Methods and Contributions\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRef\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eData\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMethod/Model\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMetric\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eContribution\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRemoving outliers\u0026thinsp;\u0026gt;\u0026thinsp;Feature selection: Pearson\u0026rsquo;s correlation method\u0026thinsp;\u0026gt;\u0026thinsp;Normalization\u0026thinsp;\u0026gt;\u0026thinsp;k-fold CV\u0026thinsp;\u0026gt;\u0026thinsp;ML, AI, DL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 88.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eIntroducing DL and Weka tools in health\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMin-max normalization\u0026thinsp;\u0026gt;\u0026thinsp;Soft voting classifier\u0026thinsp;\u0026gt;\u0026thinsp;Ensemble\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 79.08\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eEnsemble of ML algorithms\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRemoving outliers\u0026thinsp;\u0026gt;\u0026thinsp;Dealing with missing values\u0026thinsp;\u0026gt;\u0026thinsp;Data standardization\u0026thinsp;\u0026gt;\u0026thinsp;Web application development using flask\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 80.26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eWeb application development using flask to fit model\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMissing values as zero\u0026thinsp;\u0026gt;\u0026thinsp;Univariate feature selection\u0026thinsp;\u0026gt;\u0026thinsp;Creating new features\u0026thinsp;\u0026gt;\u0026thinsp;Ensemble technique: Max Voting\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCreating new features by categorizing some of the variables helps to increase the accuracy\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eT2DM from CBHS health funds company in Australia\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eOutlier treatment\u0026thinsp;\u0026gt;\u0026thinsp;Data cleaning\u0026thinsp;\u0026gt;\u0026thinsp;Cohort selection\u0026thinsp;\u0026gt;\u0026thinsp;Network analysis\u0026thinsp;\u0026gt;\u0026thinsp;Network feature\u0026thinsp;\u0026gt;\u0026thinsp;ML, NN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 82.52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eBipartite graph and projection of the patient network\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eT2DM from CPCSSN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAn approximation of missing values\u0026thinsp;\u0026gt;\u0026thinsp;Hidden Markov models\u0026thinsp;\u0026gt;\u0026thinsp;Newton\u0026rsquo;s Divide Difference Method\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 80.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eProposed an algorithm of the polynomial function implemented to estimating the missed values\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eT2DM, PLA General Hospital\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDemographic and clinical variables as predictive risk factors\u0026thinsp;\u0026gt;\u0026thinsp;3 years of follow-up \u0026gt;\u0026thinsp;Cleaning big medical data derived from EMR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 76.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eML of big medical data derived from electronic medical records in the real-world\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePregnancy Exercise and Nutrition, PEARS\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eData cleaning\u0026thinsp;\u0026gt;\u0026thinsp;Feature selection\u0026thinsp;\u0026gt;\u0026thinsp;Comparative study\u0026thinsp;\u0026gt;\u0026thinsp;Modelling\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBAL ACC 79.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eEthnicity is introduced in diabetes classification, a comparative study using ML\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSingle-payer health system in Ontario\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCohort characteristics\u0026thinsp;\u0026gt;\u0026thinsp;Feature extraction\u0026thinsp;\u0026gt;\u0026thinsp;Model performance\u0026thinsp;\u0026gt;\u0026thinsp;Feature contribution\u0026thinsp;\u0026gt;\u0026thinsp;Cost analysis\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAUC 77.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDiabetes classification using big data introduces cost\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eADAP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 65.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eInitially collected and analyzed dataset\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDT, NB, KNN, SVM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 76.7 (SVM)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCompared algorithms\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDBN ensemble\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eProposed ensemble model\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAUC: 80.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eDeveloped DL\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDT, NB, NN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 75.47 (NN)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCompared algorithms\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 76.47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUsed DT algorithm\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 77\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUsed NN\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 74.9 (NN)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUsed NN and statistical methods\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSVM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 76.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUsed SVM with combined kernels\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSVM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 83.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUsed SVM with feature selection\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAdversarial Learning\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAUC: 79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eReduced disparity using adversarial learning\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eML algorithms\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 76.1 (SVM)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCompared ML algorithms\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eAnt Colony\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 74.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUsed Ant Colony optimization\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 78.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eUsed PNN with feature extraction\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e[\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePIDD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGRNN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eACC 76.8 (GRNN)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eCompared GRNN algorithms\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e provides a summary of various research studies related to diabetes classification and prediction methods. Each entry includes the dataset used, data pre-processing steps, ML techniques applied, and the corresponding accuracy or performance metric. Additionally, some studies mention unique contributions or innovations in their approaches, such as introducing DL, web application development, or network analysis. The studies cover different aspects of diabetes classification and prediction, ranging from data cleaning and feature selection to the use of advanced techniques like DL and ML.\u003c/p\u003e \u003cp\u003eTejas and Pramila [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e] picked the two methods LR and SVM to construct a diabetes prediction model. Pre-processing the data was done to get better outcomes. SVM outperformed other methods with an accuracy of 79%, they discovered.\u003c/p\u003e \u003cp\u003eUsing three separate ML algorithms\u0026mdash;RF, DT, and NB\u0026mdash;in Hadoop-based clusters, Yuvaraj and Sripreethaa [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e] developed a diabetes prediction model. On the dataset, they used pre-processing methods. The findings demonstrated that the RF algorithm produced the greatest accuracy rate of 94%. DT, SVM, and Naive Bayes algorithms were employed by Deepti and Dilip [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. Performance was enhanced with the use of 10-CV. The Naive Bayes model produced results with the best accuracy (76.30%). The PIDD was utilized in both of these articles.\u003c/p\u003e \u003cp\u003eBoth, Olaniyi and Adnan [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e], and Swapna et al. [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] made use of DL techniques for diabetes prediction. In the former, a multilayer feed-forward NN was used. The model's training was carried out using the back-propagation approach. In order to achieve numerical stability, they also employed the PIMA Indian dataset and normalized it prior to pre-processing. They achieved an accuracy of 82%. The latter employed two models utilizing CNN and CNN-LSTM on a dataset named Electrocardiograms. There were 142,000 samples in the dataset, along with eight characteristics. With a five-fold cross validation for both models, they were able to achieve accuracy of 93.6% with the CNN model and 95.1% with the CNN-LSTM model.\u003c/p\u003e \u003cp\u003eFrom the Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, for PIDD, ADAP algorithm achieved an accuracy of 65.9% accuracy. Several studies have compared algorithms like DT, NB, KNN and SVM showing the best performance of 76.7% accuracy. A deep belief network-based ensemble classifier was proposed and obtained 78% accuracy. A DL model developed for this dataset achieved an AUC of 80.6%. Related study compared DT, NB and NN, finding NN gave maximum accuracy of 75.47%. DTs were applied with an accuracy result of 76.47%. NN alone obtained 77% accuracy. NN combined with statistical methods achieved 74.9% accuracy. SVM with combined kernels attained 76.6% accuracy. SVM with feature selection yielded higher accuracy of 83.3%. Adversarial learning reduced disparity and improved AUC to 79%. A comparison found SVM had the best accuracy of 76.1%. An ant colony optimization classifier developed achieved 74.6% accuracy. PNN with feature extraction gave 78.3% accuracy. A study analysed different generalized regression NN algorithms, reporting GRNN achieved 76.8% accuracy.\u003c/p\u003e \u003cp\u003eAll of the aforementioned research provided a performance comparison of several ML algorithms. While some of them utilized data pre-processing and cross-validation strategies to increase accuracy, all of them were more concerned with comparing the performance of several models than they were with perfecting a single one. In this study, the researchers have focused on a single model and investigated strategies that can boost performance by enhancing both execution speed and accuracy with a special attention to algorithm selection, pre- and post-processing of the data have a significant impact on the model's overall improvement of models.\u003c/p\u003e \u003cp\u003eThe main objective of this study is to compare the accuracies of diabetes classification among the AI algorithms using associated risk factors. The specific objectives are:\u003c/p\u003e \u003cp\u003e1. To identify the risk factors for diabetes and arising awareness among the people;\u003c/p\u003e \u003cp\u003e2. To show how parameter tuning can increase the accuracies of the models;\u003c/p\u003e \u003cp\u003e3. To show how missing values can be imputed and may have contributory role for the betterment of ML methods;\u003c/p\u003e \u003cp\u003e4. To show how data pre-processing can increase the accuracy of ML in public health;\u003c/p\u003e \u003cp\u003e5. To propose an appropriate model to classify diabetes accurately in this ground;\u003c/p\u003e"},{"header":"Materials and Methods","content":"\u003cp\u003eIn this study, the experimentation was conducted using the Pima Indian Dataset, obtained from various sources including ML Databases, ML Repository, and the NIDDK [\u003cspan additionalcitationids=\"CR46\" citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]\u0026ndash;[\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e]. The dataset serves as the foundation for the this analysis, containing crucial insights into diabetes diagnosis. Comprising a total of 768 rows, the dataset encompasses both non-diabetic and diabetic subjects, with 500 instances falling into the former category and 268 instances into the latter. This binary classification is determined by an output column, with a value of 0 indicating absence of diabetes and a value of 1 denoting its presence. The dataset features nine columns encapsulating distinct attributes: pregnant month, glucose level, plasma concentration, blood pressure, triceps skinfold thickness, insulin amount, BMI, pedigree function, and the patient\u0026rsquo;s age. These attributes collectively contribute to the predictive analysis, aiding in the differentiation between individuals with and without diabetes.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eConceptual Framework\u003c/h2\u003e \u003cp\u003eA conceptual framework is a theoretical structure comprised of interrelated concepts, serving as a foundation for understanding complex subjects, visual diagram, illustrating the sequential steps, decisions, and actions within a process or system. The conceptual framework and flowchart of the whole study is given below in the Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e:\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eData Pre-processing\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eIn order to extract usable information from this raw data and feed it into the training model for effective medical choices, diagnoses, and treatments, data pre-processing is a necessary step. To enhance the quality of the data, the pre-processing conducts a number of tasks, including outlier treatment, filling in missing values, data normalization, and feature selection. 500 samples in the dataset were identified as not having diabetes, compared to 268 diabetic samples. Label encoding is the first method employed. This method is used to determine if a person has diabetes or not, which is the dependent variable. As a result, the output class is determined by replacing all of the string values in the output variable with 0 and 1. The subsequent data pre-processing method uses the ML algorithm missForest to address missing values [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e]. The dataset had many missing values for many variables, including 227 missing values for the skin thickness parameter, 374 missing values for the insulin attribute, and 111 missing values for pregnancy.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eMissing Values Treatment\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eThe elimination of the rows or columns with null values is one method of addressing missing values. The data analyst can remove the whole column if any columns contain more than 50% null values. Similar to how columns can be discarded if one or more of their values are null, rows can likewise be deleted. Imputation approach is the alternative method of treating missing values. If a column has fewer than 50% of its values be null, the missing value can be added to the entire column using the imputation approach. In this study used missForest [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e] technique to input missing value.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003eLet \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(X\\)\u003c/span\u003e\u003c/span\u003e be \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(n\\times p\\)\u003c/span\u003e\u003c/span\u003ematrix of predictors that requires imputation:\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\varvec{X} = ( {\\varvec{X}}_{1}, {\\varvec{X}}_{2}, \\dots , {\\varvec{X}}_{\\varvec{p}})= \\left[\\begin{array}{ccc}{\\varvec{x}}_{11}\u0026amp; {\\varvec{x}}_{12}\u0026amp; \\begin{array}{cc}\\dots \u0026amp; {\\varvec{x}}_{1\\varvec{p}}\\end{array}\\\\ {\\varvec{x}}_{21}\u0026amp; {\\varvec{x}}_{22}\u0026amp; \\begin{array}{cc}\\dots \u0026amp; {\\varvec{x}}_{2\\varvec{p}}\\end{array}\\\\ \\begin{array}{c}⋮\\\\ {\\varvec{x}}_{\\varvec{n}1}\\end{array}\u0026amp; \\begin{array}{c}⋮\\\\ {\\varvec{x}}_{\\varvec{n}2}\\end{array} \u0026amp; \\begin{array}{c}⋮\\\\ {\\dots \\varvec{x}}_{\\varvec{n}\\varvec{p}}\\end{array}\\end{array}\\right]$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eAn arbitrary variable \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\varvec{X}}_{\\varvec{s}}\\)\u003c/span\u003e\u003c/span\u003econtains missing values at entries \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\varvec{i}}_{\\varvec{m}\\varvec{i}\\varvec{s}}^{\\left(\\varvec{s}\\right)} \\subseteq \\{1, 2, \\dots , n\\}\\)\u003c/span\u003e\u003c/span\u003e. For every variable \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mathbf{X}}_{\\varvec{s}}\\)\u003c/span\u003e\u003c/span\u003ethat contains missing values, the researchers can separate the dataset into 4 categories:\u003c/p\u003e \u003cp\u003e1. The non-missing values of variable \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mathbf{X}}_{s}\\)\u003c/span\u003e\u003c/span\u003e, denoted by\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mathbf{y}}_{\\mathbf{o}\\mathbf{b}\\mathbf{s}}^{\\left(\\mathbf{s}\\right)}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003cp\u003e2. The missing values of variable \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mathbf{X}}_{\\varvec{s}}\\)\u003c/span\u003e\u003c/span\u003e, denoted by \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mathbf{y}}_{\\mathbf{m}\\mathbf{i}\\mathbf{s}}^{\\left(\\mathbf{s}\\right)}\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e3. The variables other than \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mathbf{X}}_{\\varvec{s}}\\)\u003c/span\u003e\u003c/span\u003e, with observations \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mathbf{i}}_{\\mathbf{o}\\mathbf{b}\\mathbf{s}}^{\\left(\\mathbf{s}\\right)}\\)\u003c/span\u003e\u003c/span\u003e= \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\frac{\\left\\{1, 2, \\dots , n\\right\\}}{{\\varvec{i}}_{\\varvec{m}\\varvec{i}\\varvec{s}}^{\\left(\\varvec{s}\\right)}}\\)\u003c/span\u003e\u003c/span\u003e, denoted by\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mathbf{x}}_{\\mathbf{o}\\mathbf{b}\\mathbf{s}}^{\\left(\\mathbf{s}\\right)}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003cp\u003e4. The variables other than \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mathbf{X}}_{\\varvec{s}}\\)\u003c/span\u003e\u003c/span\u003e, with observations \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mathbf{i}}_{\\mathbf{m}\\mathbf{i}\\mathbf{s}}^{\\left(\\mathbf{s}\\right)}\\)\u003c/span\u003e\u003c/span\u003e, denoted by\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\mathbf{x}}_{\\mathbf{m}\\mathbf{i}\\mathbf{s}}^{\\left(\\mathbf{s}\\right)}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eOutlier Treatment\u003c/h2\u003e \u003cp\u003eAn outlier is a data point in statistics that dramatically deviates from other observations. An outlier may be caused by measurement variability, a sign of unique data, or an experimental error; the latter is occasionally eliminated from the data set. While an outlier may signal an intriguing potential, it may also seriously impair statistical analysis. In any distribution, outliers can happen by accident, but they can also point to unexpected behaviour or structures in the data set, measurement error, or a heavy-tailed distribution in the population. Heavy-tailed distributions suggest that the distribution has substantial skewness, and one should be extremely cautious when using tools or intuitions that presume a normal distribution. In the case of measurement error, one desires to reject them or use statistics that are resilient to outliers. A mixture of two distributions, which may represent two different subpopulations or may represent \"right trial\" vs \"measurement mistake,\" is a frequent source of outliers and is represented by a mixture model. Most often, in bigger data samples, certain data points will be further distant from the sample mean than what is regarded as fair. It can be because some observations are distant from the centre of the data, inadvertent systematic mistake, or problems with the theory that produced the presumed family of probability distributions. Therefore, outlier points may hint to flawed data, flawed processes, or locations where a certain hypothesis may not hold true. However, it is normal to expect a few outliers in big datasets (and not due to any anomalous condition).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eAnalysis Tools\u003c/h2\u003e \u003cp\u003eThe researchers used R and RStudio for data processing and data analysis, data visualization, Microsoft Office for documentation, and Mendeley for reference management, creating a unified framework for in-depth exploration of diabetes risk factors using the Pima diabetes dataset.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Tests for Association, Mean Comparison, and Odds Ratios\u003c/h2\u003e \u003cp\u003eThe analysis employs a spectrum of tests to explore associations, mean disparities, and ORs across diverse scenarios. The Chi-Square Test, utilized for categorical variables in contingency tables, detects associations, particularly with small sample sizes or when chi-square assumptions are unmet. The t-test, a parametric counterpart to compare means, evaluates differences among various independent samples for continuous data. OR quantifies associations in case-control studies, while likelihood ratios compare event odds between groups; an OR above 1 indicates higher odds in the first group. AOR, within LR, gauges predictor-outcome connections while considering confounders, offering deeper insights. These tests, spanning categorical, continuous, and ordinal data, equip researchers with nuanced insights into variable relationships, enriching conclusions with comprehensive considerations.\u003c/p\u003e \u003cp\u003e \u003cb\u003eClassification Algorithms\u003c/b\u003e \u003c/p\u003e \u003cp\u003e1. NNET: NNs are AI techniques inspired by the human brain. DL, a form of ML, uses interconnected nodes to simulate brain structures.\u003c/p\u003e \u003cp\u003e2. AVNNET: AVNNET utilizes NNs with various seeds for prediction. Model scores are averaged for regression, and helps building more reliable models by decreasing training variance.\u003c/p\u003e \u003cp\u003e3. PCANNET: PCANNET uses principal component analysis to reduce data dimensions. NNs then employ these components for training and prediction, ensuring predictors have enough variability.\u003c/p\u003e \u003cp\u003e4. MULTINOM: MULTINOM extends bias reduction techniques to multinomial LR models. It employs penalized maximum likelihood estimates and profile confidence intervals for hypothesis testing.\u003c/p\u003e \u003cp\u003e5. RBF: Radial basis function (RBF) networks are popular for function approximation. They use radial basis functions as activation functions and find use in approximation, classification, prediction, and control.\u003c/p\u003e \u003cp\u003e6. RBFDDA: RBF networks with dynamic decay adjustment technique are simpler and specialize in categorization. They start with minimal parameters and add units gradually, making them easier to use than standard RBF networks.\u003c/p\u003e \u003cp\u003e7. RPART: Recursive partitioning is a tool for creating decision rules in data mining. The RPART generates classification and regression trees, aiding data exploration and predictive modelling. The resulting CART model provides user-friendly predictions based on predictor variables.\u003c/p\u003e \u003cp\u003e8. XGBLINEAR: XGBoost is a potent ML library with unique algorithms, each controlled by hyperparameters. It models tasks using trees, linear functions, and regularization techniques. These hyperparameters influence model behaviour and outcomes.\u003c/p\u003e \u003cp\u003e9. RRF: Regularized Random Forest is used for feature selection with a single ensemble. It evaluates features in tree nodes using subsets of training data, eliminating the need for data normalization.\u003c/p\u003e \u003cp\u003e10. LOGITBOOST: LogitBoost is a boosting classification algorithm similar to AdaBoost. Both perform additive LR, but LogitBoost minimizes logistic loss while AdaBoost minimizes exponential loss.\u003c/p\u003e \u003cp\u003e11. RANGER: Ranger is a fast implementation of recursive partitioning and random forests, suitable for large datasets. It supports classification, regression, and survival forests, including examples of quantile regression forests and very random trees.\u003c/p\u003e \u003cp\u003e12. RF: Random Forest is a popular supervised ML approach for classification and regression. It creates an ensemble of DTs on random subsets of input data, enhancing prediction accuracy through voting.\u003c/p\u003e \u003cp\u003e13. MLP: A multilayer perceptron is a fully connected feedforward ANN. It comprises at least three layers: input, hidden, and output. Nonlinear activation functions are used in the hidden and output layers, and backpropagation is employed for supervised learning. MLP can handle non-linearly separable data due to its multiple layers and non-linear activations.\u003c/p\u003e \u003cp\u003e14. MLPWEIGHTDECAY: In DL, models can generalize better through data augmentation. But what about during training? MLPWEIGHTDECAY introduces weight and decay parameters to enhance model generalization.\u003c/p\u003e \u003cp\u003e15. MLPML: Similar to MLP, MLPML consists of three layers: input, hidden, and output. It uses non-linear activation functions and backpropagation for training. This architecture allows MLPML to distinguish non-linearly separable data.\u003c/p\u003e \u003cp\u003e16. MLPWEIGHTDECAYML: MLPWEIGHTDECAYML is a fully connected feedforward ANN that includes settings for decay and weight parameters. MLPs are sometimes called \"vanilla\" NNs, especially with one hidden layer. The term \"MLP\" can refer broadly to any feedforward ANN or strictly to networks with multiple layers of perceptrons.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003ePerformance Metrics in Classification\u003c/h2\u003e \u003cp\u003ePerformance metrics serve as vital tools for the evaluation of classification models, enabling researchers to gauge how effectively the model predicts outcomes compared to real values. A range of commonly used performance metrics exist to aid in this assessment. Sensitivity, also known as the True Positive Rate or Recall, quantifies the proportion of actual positive cases correctly identified by the model. Specificity measures the model\u0026rsquo;s ability to accurately pinpoint actual negative cases. Precision, or Positive Predictive Value, gauges the accuracy of positive predictions by calculating the ratio of true positives to the total predicted positives. Recall, akin to Sensitivity, indicates the model\u0026rsquo;s capacity to correctly predict actual positive cases. The F1 Score, a harmonic mean of precision and recall, offers a balanced evaluation, particularly valuable when dealing with class imbalances. AUC-ROC graphically represents the model\u0026rsquo;s discriminatory ability between positive and negative classes, providing a threshold-independent evaluation. Accuracy captures the overall correctness of predictions by calculating the ratio of correctly predicted cases to the total. Cohen\u0026rsquo;s Kappa adjusts accuracy by considering chance agreement, thus quantifying agreement between predicted and actual outcomes while accounting for random agreement. These diverse performance metrics collectively illuminate the strengths and weaknesses of classification models, offering insights tailored to analysis goals, dataset characteristics, and the desired balance between evaluation criteria.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":"\u003cp\u003eIn this section, the authors present the key findings and outcomes of this study, which contribute to a deeper understanding of the data and diabetes risk factors with the hybrid data preparation technique. This research endeavors, guided by rigorous methodology and comprehensive analysis, have yielded valuable insights into the intricacies. This findings are presented in a clear and organized manner to facilitate comprehension and offer a foundation for further discussion and interpretation.\u003c/p\u003e\n\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n\u003ch2\u003eData Treatment\u003c/h2\u003e\n\u003cp\u003eIn this research work, the Pima Indian Dataset has been considered for the experimentation [\u003cspan class=\"CitationRef\"\u003e45\u003c/span\u003e]\u0026ndash;[\u003cspan class=\"CitationRef\"\u003e47\u003c/span\u003e]. The dataset comprises nine columns, and one output column has a binary value to indicate whether the subject has diabetes or not. It comprises 768 rows, 500 of which contain non-diabetics and 268 of which have diabetic patients. Nine feature columns, including pregnant month, glucose, plasma, blood pressure, fold thickness of the triceps skin, amount of insulin, BMI, pedigree function, age of patients, and one goal column, are included in the dataset (0 or 1).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\n\u003ch2\u003eMissing Value Treatment\u003c/h2\u003e\n\u003cp\u003eIn this dataset there are many missing values which are presented in the Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e:\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab2\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eMissing Values\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eVariable\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eNo. of Missing Values\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePregnant\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eGlucose\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e5\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePressure\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e35\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eTriceps\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e227\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eInsulin\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e374\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMass\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e11\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePedigree\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAge\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eDiabetes\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eThe Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e presents a dataset with various variables related to diabetes prediction, indicating the number of missing values for each variable. Notably, the \"Pregnant,\" \"Pedigree,\" and \"Age\" variables have no missing values, while others like \"Insulin\" and \"Triceps\" have a substantial number of missing values, suggesting incomplete data records. Addressing these missing values is crucial for accurate analysis and modelling. Researchers typically employ imputation or data pre-processing techniques to handle missing data and ensure the dataset's quality for predictive modelling or statistical analysis. Imputation approach is the alternative method of treating missing values. If a column has fewer than 50% of its values be null, the missing value can be added to the entire column using the imputation approach. In this study, the missing value was entered using the missForest method [\u003cspan class=\"CitationRef\"\u003e48\u003c/span\u003e].\u003c/p\u003e\n\u003cp\u003eBefore applying missForest, the missing plot is given below.\u003c/p\u003e\n\u003cp\u003eFrom the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e, the researchers may notice that, there is 9% missing values in the dataset. Now the researchers handle the dataset by using missForest.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\n\u003ch2\u003eOutlier Treatment\u003c/h2\u003e\n\u003cp\u003eAn outlier is a data point in statistics that dramatically deviates from other observations. An outlier may be caused by measurement variability, a sign of unique data, or an experimental error; the latter is occasionally eliminated from the data set. While an outlier may signal an intriguing potential, it may also seriously impair statistical analysis. The visualization of the presence of outliers by Mahalanobis squared distances to detect outliers in this dataset is given below.\u003c/p\u003e\n\u003cp\u003eFrom the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e, the researchers can suspect that there are many outliers in the dataset. The researcher\u0026rsquo;s aim is to treat the outliers.\u003c/p\u003e\n\u003cp\u003eThe visualization of the presence of outliers by Mahalanobis squared distances to detect outliers in this dataset is given below.\u003c/p\u003e\n\u003cp\u003eFrom the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e, the researchers can suspect that there are many outliers in the dataset. The researcher\u0026rsquo;s focus is to treat the outliers.\u003c/p\u003e\n\u003cp\u003eIn this study, the researchers replaced the detected outliers as missing values and predicted them with multistage hybrid ML with optimum number of iterations, where the researchers got 13% outliers in this dataset.\u003c/p\u003e\n\u003cp\u003eFrom the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e, there is 13% outliers in this dataset.\u003c/p\u003e\n\u003cp\u003eAfter outlier treatment by the above method the visualization of the absence of outliers by Mahalanobis squared distances in this dataset is given below.\u003c/p\u003e\n\u003cp\u003eFrom the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e, the researchers can see there is no harmful/potential suspected outliers in the dataset.\u003c/p\u003e\n\u003cp\u003eAfter outlier treatment by the above method the visualization of the absence of outliers by Mahalanobis squared distances in this dataset is given below.\u003c/p\u003e\n\u003cp\u003eFrom the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e7\u003c/span\u003e, the researchers may declare that, now the dataset is free from the presence of the suspected outliers.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\n\u003ch2\u003eBox Plot\u003c/h2\u003e\n\u003cp\u003eBox plots without outlier treatment are given below.\u003c/p\u003e\n\u003cp\u003eThis Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e8\u003c/span\u003e of the box plots are drawn on the original dataset without outlier treatment. Box plots with outlier treatment are given below.\u003c/p\u003e\n\u003cp\u003eFrom the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e9\u003c/span\u003e of the box plots, the researchers can see that after the treatment, there is no outlier is suspected.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\n\u003ch2\u003eMean Test\u003c/h2\u003e\n\u003cp\u003et-Test p-values for each independent numerical variable with the levels of the dependent variable are presented in the Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e which is given below.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab3\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003et-Test p-values for each Independent Variable\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth colspan=\"2\" align=\"left\"\u003e\n\u003cp\u003eDiabetes\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eneg\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003epos\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003ep\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003ePregnant\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMean (SD)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e3.3 (3.0)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e4.9 (3.7)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eGlucose\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMean (SD)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e110.6 (24.7)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e142.4 (29.5)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003ePressure\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMean (SD)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e70.8 (12.0)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e75.3 (12.0)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eTriceps\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMean (SD)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e27.0 (9.2)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e32.5 (9.0)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eInsulin\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMean (SD)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e129.6 (83.3)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e207.2 (103.9)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eMass\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMean (SD)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e30.8 (6.5)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e35.4 (6.6)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003ePedigree\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMean (SD)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.4 (0.3)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6 (0.4)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003eAge\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMean (SD)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e31.2 (11.7)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e37.1 (11.0)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e\u0026lt;\u0026thinsp;0.001\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eThe Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e presents a comprehensive comparison between individuals with negative and positive cases of diabetes across various health-related variables. Notably, individuals with positive diabetes cases exhibit significantly higher average values in several key parameters, including glucose levels, insulin levels, triceps skinfold thickness, BMI, pedigree function, and age, when compared to those with negative cases. These stark differences are underscored by the extremely low p-values (\u0026lt;\u0026thinsp;0.001) associated with each variable, indicating strong statistical significance. Specifically, higher glucose and insulin levels highlight the metabolic impact of diabetes, while the elevated age and BMI among positive cases may suggest long-term health implications and potential risk factors. These findings collectively emphasize the importance of monitoring and managing these variables as critical aspects of diabetes prevention and management.\u003c/p\u003e\n\u003cp\u003eTo visualize the mean differences the means plots with levels of the dependent variable are given below.\u003c/p\u003e\n\u003cp\u003eThe grouped histograms are given below in the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e11\u003c/span\u003e:\u003c/p\u003e\n\u003cp\u003eThe grouped box-plots are given below in the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e12\u003c/span\u003e:\u003c/p\u003e\n\u003cp\u003eGrouped density plots are given below in the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e13\u003c/span\u003e:\u003c/p\u003e\n\u003cp\u003eAccording the above Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e, Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e10\u003c/span\u003e, Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e11\u003c/span\u003e, Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e12\u003c/span\u003e, and Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e13\u003c/span\u003e the researchers can say that, the mean differences between negative and positive levels of the variable Diabetes for any other numeric variables are statistically significant. The means for positive level of Diabetes is significantly higher than negative level. The Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e summarizes the outcomes of t-tests conducted to compare various health indicators between individuals with negative (no diabetes) and positive (diabetes) outcomes. The means and standard deviations of several features were examined for both groups. The results indicate significant differences in almost all parameters. Individuals with diabetes exhibit notably higher mean values for key factors including glucose level, blood pressure, triceps skinfold thickness, insulin level, BMI, pedigree score, and age. These disparities are all statistically significant, as evidenced by the p-values (\u0026lt;\u0026thinsp;0.001) accompanying each comparison. This suggests that these health indicators could potentially serve as meaningful differentiators between those affected by diabetes and those who are not.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e\n\u003ch2\u003eCorrelation Matrix\u003c/h2\u003e\n\u003cp\u003eThe correlation matrix is given below:\u003c/p\u003e\n\u003cp\u003eFrom the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e14\u003c/span\u003e, the researchers may get an overview of the data.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e\n\u003ch2\u003eTrivariate Analysis\u003c/h2\u003e\n\u003cp\u003eScatterplot matrix of the features of original data is given below in the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e15\u003c/span\u003e:\u003c/p\u003e\n\u003cp\u003eScatterplot matrix of the features of the outlier treated data is given below in the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e16\u003c/span\u003e:\u003c/p\u003e\n\u003cp\u003eFrom the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e14\u003c/span\u003e, Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e15\u003c/span\u003e, and Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e16\u003c/span\u003e, the researchers found that here multicollinearity may arise. So classical modelling is not appropriate for this dataset. ML may be a suitable replacement of the classical models.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec18\" class=\"Section2\"\u003e\n\u003ch2\u003eClassifiers\u003c/h2\u003e\n\u003cp\u003eIn this study the researchers used 16 classifiers to classify diabetes using PIMA dataset. The performance metrics are presented in the Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab4\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eModel Comparison\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eModel\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eSensitivity\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eSpecificity\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003ePrecision\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eRecall\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eF1\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eAUC\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eAccuracy\u003c/p\u003e\n\u003c/th\u003e\n\u003cth align=\"left\"\u003e\n\u003cp\u003eKappa\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eLogitBoost\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8095\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.9464\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.85\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8095\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8293\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7888\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.9091\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7674\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003emlpWeightDecayML\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7925\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.77\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6462\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7925\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7119\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8196\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7778\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.534\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eavNNet\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8302\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.75\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6377\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8302\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7213\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8442\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7778\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.5418\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003emlpML\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7736\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.78\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6508\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7736\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7069\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8228\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7778\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.5301\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eranger\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.717\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6552\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.717\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6847\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8476\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7712\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.5058\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ennet\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7736\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.77\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6406\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7736\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7009\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8411\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7712\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.5183\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eRRF\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7547\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.77\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6349\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7547\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6897\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.84\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7647\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.5024\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eRF\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.717\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.79\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6441\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.717\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6786\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8458\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7647\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.4938\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003emultinom\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7547\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.77\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6349\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7547\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6897\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8438\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7647\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.5024\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003erbf\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8113\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.74\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6232\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8113\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7049\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8387\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7647\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.5148\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003exgbLinear\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6981\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.79\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6379\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6981\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6667\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8134\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7582\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.4775\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003erbfDDA\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6792\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6429\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6792\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6606\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7742\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7582\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.473\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003emlp\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7925\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.73\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6087\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7925\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6885\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8355\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7516\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.4878\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003epcaNNet\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7925\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.72\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7925\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6829\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.8409\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7451\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.4765\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003emlpWeightDecay\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6792\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.76\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6792\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6372\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7945\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.732\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.426\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003erpart\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.9057\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.5455\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.9057\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.6809\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7696\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.7059\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"char\" char=\".\"\u003e\n\u003cp\u003e0.4377\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eThe accuracy and kappa metrics are visualised in the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e17\u003c/span\u003e:\u003c/p\u003e\n\u003cp\u003eFrom the Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e and Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e17\u003c/span\u003e: Sensitivity measures the proportion of actual positive cases that are correctly identified by the model. It's also called the true positive rate or recall. A higher sensitivity indicates that the model is better at identifying positive cases. For instance, in the first row (\"LogitBoost\"), the sensitivity is 0.8095, which means that the model correctly identifies around 81% of the actual positive cases.\u003c/p\u003e\n\u003cp\u003eSpecificity measures the proportion of actual negative cases that are correctly identified as negative by the model. A higher specificity indicates that the model is better at identifying negative cases. For example, in the \"LogitBoost\" model, the specificity is 0.9464, which means that around 95% of the actual negative cases are correctly identified as negative.\u003c/p\u003e\n\u003cp\u003ePrecision measures the proportion of positive predictions made by the model that are actually correct. It's a measure of the accuracy of positive predictions. A higher precision indicates that the positive predictions are more likely to be accurate. In the \"LogitBoost\" model, the precision is 0.85, which means that around 85% of the positive predictions are correct.\u003c/p\u003e\n\u003cp\u003eRecall is another term for sensitivity, as mentioned above. It's the proportion of actual positive cases that the model correctly identifies.\u003c/p\u003e\n\u003cp\u003eThe F1 score is the harmonic mean of precision and recall. It combines both measures and provides a balanced assessment of a model's performance. A higher F1 score indicates a good balance between precision and recall.\u003c/p\u003e\n\u003cp\u003eThe AUC represents the area under the Receiver Operating Characteristic (ROC) curve. The ROC curve plots the true positive rate against the false positive rate for different thresholds. A higher AUC indicates better overall discrimination ability of the model.\u003c/p\u003e\n\u003cp\u003eAccuracy measures the proportion of correctly classified instances (both true positives and true negatives) out of all instances. It's a general measure of the model's correctness.\u003c/p\u003e\n\u003cp\u003eCohen's Kappa is a measure of agreement between predicted and actual classifications, considering the possibility of agreement by chance. A higher Kappa indicates a higher agreement between predicted and actual classifications than what would be expected by chance.\u003c/p\u003e\n\u003cp\u003eThese metrics collectively provide insights into how well each model is performing on the PIDD. The \"LogitBoost\" model seems to have relatively high sensitivity, specificity, precision, and F1 score, suggesting a good overall performance for this dataset. However, the choice of the best model also depends on the specific goals of the analysis and the trade-offs the researchers are willing to make between different evaluation measures.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec19\" class=\"Section2\"\u003e\n\u003ch2\u003eComparative Study with LR\u003c/h2\u003e\n\u003cp\u003eThe objective is to gauge the model's performance in relation to various factors, both directly linked to diabetes and those encompassing broader health considerations. By conducting this evaluation, the researchers seek to discern the relative strengths and weaknesses of LR in comparison to alternative approaches, shedding light on its suitability for diabetes diagnosis and its applicability to broader health-related inquiries.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003ctable id=\"Tab5\" border=\"1\"\u003e\u003ccaption\u003e\n\u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\n\u003cdiv class=\"CaptionContent\"\u003e\n\u003cp\u003eOR from LR Model\u003c/p\u003e\n\u003c/div\u003e\n\u003c/caption\u003e\n\u003cthead\u003e\n\u003ctr\u003e\n\u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n\u003cth colspan=\"5\" align=\"left\"\u003e\n\u003cp\u003eDiabetes\u003c/p\u003e\n\u003c/th\u003e\n\u003c/tr\u003e\n\u003c/thead\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003ePredictors\u003c/em\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eOR\u003c/em\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003estd. Error\u003c/em\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eCI\u003c/em\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003eStatistic\u003c/em\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cem\u003ep\u003c/em\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e(Intercept)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.00\u0026ndash;0.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-10.83\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026lt;\u0026thinsp;0.001\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAge\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.01\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.01\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.99\u0026ndash;1.03\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.18\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.237\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eGlucose\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.04\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.03\u0026ndash;1.05\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e8.34\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026lt;\u0026thinsp;0.001\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eInsulin\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.00\u0026ndash;1.00\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.44\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.663\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eMass\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.09\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.02\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.05\u0026ndash;1.14\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e4.26\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026lt;\u0026thinsp;0.001\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePedigree\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.36\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.70\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.32\u0026ndash;4.24\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e2.89\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e0.004\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePregnant\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.13\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.04\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.06\u0026ndash;1.21\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e3.85\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026lt;\u0026thinsp;0.001\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003ePressure\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.99\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.01\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.98\u0026ndash;1.01\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e-0.88\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.377\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eTriceps\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e1.01\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.01\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.98\u0026ndash;1.03\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.44\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e0.659\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eObservations\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"5\" align=\"left\"\u003e\n\u003cp\u003e768\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003e\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({R}^{2}\\)\u003c/span\u003e\u003c/span\u003e Tjur\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"5\" align=\"left\"\u003e\n\u003cp\u003e0.336\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003eAIC\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"5\" align=\"left\"\u003e\n\u003cp\u003e729.301\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd align=\"left\"\u003e\n\u003cp\u003elog-Likelihood\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"5\" align=\"left\"\u003e\n\u003cp\u003e-355.650\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eFrom the Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e, the researchers obtained:\u003c/p\u003e\n\u003c/div\u003e\n\u003cp\u003eAge: The OR is 1.01, which suggests that for a one-unit increase in age, the odds of having diabetes increase by 1.01 times. However, the p-value (0.237) indicates that the effect is not statistically significant, as it's greater than 0.05.\u003c/p\u003e\n\u003cp\u003eGlucose: The OR is 1.04, and the p-value is \u0026lt;\u0026thinsp;0.001, indicating that higher glucose levels are associated with increased odds of diabetes in a statistically significant way.\u003c/p\u003e\n\u003cp\u003eInsulin: The OR is 1.00, with a p-value of 0.663. This suggests that insulin levels don't have a significant effect on diabetes presence.\u003c/p\u003e\n\u003cp\u003eMass: The OR is 1.09, and the p-value is \u0026lt;\u0026thinsp;0.001. This indicates that higher mass (body weight) is associated with increased odds of diabetes significantly.\u003c/p\u003e\n\u003cp\u003ePedigree: The OR is 2.36, and the p-value is 0.004. This suggests that a higher pedigree (a measure of diabetes heredity) is associated with significantly higher odds of diabetes.\u003c/p\u003e\n\u003cp\u003ePregnant: The OR is 1.13, and the p-value is \u0026lt;\u0026thinsp;0.001. This indicates that being pregnant is associated with higher odds of diabetes.\u003c/p\u003e\n\u003cp\u003ePressure: The OR is 0.99, with a p-value of 0.377. Blood pressure doesn't seem to have a significant effect on diabetes.\u003c/p\u003e\n\u003cp\u003eTriceps: The OR is 1.01, and the p-value is 0.659. Triceps skinfold thickness doesn't have a significant effect on diabetes.\u003c/p\u003e\n\u003cp\u003eThe \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({R}^{2}\\)\u003c/span\u003e\u003c/span\u003e Tjur is a measure of goodness-of-fit in the context of LR. It's used to evaluate how well the model fits the data, just like the R-squared (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({R}^{2}\\)\u003c/span\u003e\u003c/span\u003e) statistic in linear regression. However, the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({R}^{2}\\)\u003c/span\u003e\u003c/span\u003e Tjur is specifically designed for LR, where the response variable is binary. In essence, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({R}^{2}\\)\u003c/span\u003e\u003c/span\u003e Tjur quantifies the proportion of explained variance in the dependent variable by the predictor variables in the model. A value of 0 means that the predictors have no explanatory power, and a value of 1 indicates that they completely explain the variability in the response. In the Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e, the value is 0.336, suggesting that the predictors in the model collectively explain about 33.6% of the variance in the presence of diabetes. The log-likelihood is a measure used to assess how well the statistical model fits the observed data. In the context of LR, it's a measure of how likely the observed outcomes are, given the parameter estimates of the model. Higher log-likelihood values indicate that the model is better at explaining the observed data. In the Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e, the log-likelihood value is -355.650. While the absolute value itself doesn't hold a direct interpretation, comparing log-likelihood values between different models can help to determine which model fits the data better. The AIC is a metric used to compare the quality of different models, taking into account both goodness of fit and model complexity. It's a tool for model selection, helping to choose the model that best balances explanatory power and parsimony. The AIC is calculated using the log-likelihood value and the number of parameters in the model. Lower AIC values indicate a better trade-off between model fit and complexity. In the Table\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e, the AIC value is 729.301, which can be used to compare this model to other models. Lower AIC values, when comparing models, suggest that the model with the lower AIC is a better fit for the data.\u003c/p\u003e\n\u003cp\u003eIn summary, these three metrics\u0026mdash;\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({R}^{2}\\)\u003c/span\u003e\u003c/span\u003e Tjur, log-likelihood, and AIC\u0026mdash;provide insights into different aspects of the LR model's performance. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({R}^{2}\\)\u003c/span\u003e\u003c/span\u003e Tjur assesses the overall fit and explained variance, log-likelihood gauges how well the model fits the data, and AIC helps to compare different models based on both goodness of fit and complexity. Based on this analysis, higher glucose levels, higher body mass, higher pedigree, and pregnancy are associated with increased odds of diabetes, while other factors like age, insulin levels, blood pressure, and triceps thickness do not show a statistically significant relationship.\u003c/p\u003e\n\u003cp\u003eFrom the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e18\u003c/span\u003e, the researchers can say, glucose, mass, pedigree, pregnancy have a positive relation with the chance of having diabetes.\u003c/p\u003e\n\u003cp\u003eFrom the Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e19\u003c/span\u003e, the researchers can say the high-risk factors are glucose, mass, pregnant, pedigree.\u003c/p\u003e\n\u003cp\u003eThe Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e20\u003c/span\u003e of ROC curve, with an Area Under the Curve (AUC) of 84.6493, indicates that binary classification model exhibits strong discriminatory power. With an AUC significantly above the random guessing threshold of 0.5, it demonstrates an ability to effectively distinguish between positive and negative instances across various decision thresholds. While not perfect, this AUC score suggests that the model performs well on average, making it a promising tool for making binary predictions.\u003c/p\u003e\n\u003cp\u003eHere, the researchers want to show variable importance according to the LR model and best subset model by backward stepwise elimination method. And the researchers have shown Odds plot to do the expected comparative analysis.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study employs a sophisticated approach to unravel the intricate web of factors contributing to the susceptibility of diabetes. Through the utilization of advanced statistical methodologies, including LR, ORs, and AORs, the investigation reveals profound insights into the predictive landscape of this health condition.\u003c/p\u003e \u003cp\u003eThe ORs estimated from the LR model in the current study are largely consistent with findings from prior analyses of the Pima Indian dataset [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e], [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e], [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e], [\u003cspan additionalcitationids=\"CR50 CR51\" citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e]\u0026ndash;[\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e], [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e], [\u003cspan additionalcitationids=\"CR32 CR33\" citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]\u0026ndash;[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e], [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e], [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. Specifically, glucose level was found to be a significant predictor of diabetes with an OR of 1.04 (95% CI: 1.03\u0026ndash;1.05), which falls within the range of 1.03\u0026ndash;1.05 reported in previous studies examining this relationship [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e], [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e], [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. Similarly, the OR for mass of 1.09 (95% CI: 1.05\u0026ndash;1.14) is comparable to past estimates of 1.02\u0026ndash;1.09 [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e], [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. The significant OR of 2.36 (95% CI: 1.32\u0026ndash;4.24) for pedigree function aligns with the significance and magnitude of heredity measures observed in other investigations [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. Some variables like age, insulin level, and blood pressure showed non-significant or weaker associations here as in portions of the earlier literature [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e], [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e], [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e], [\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e], [\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e]. Overall, the pattern of predictor significance and OR estimates replicate established relationships demonstrated across multiple prior LR analyses of the PIDD [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e], [\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e], [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e]. Therefore, the current results corroborate past research findings using for this well-studied population.\u003c/p\u003e \u003cp\u003eA systematic methodology centered on robust pre-processing of missing data and outliers to develop diabetes have been implemented classification models using the Pima Indian dataset. Prior investigations typically omitted missing values or removed outliers without supplementation, which may have compromised resultant model accuracy [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e], [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e], [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e], [\u003cspan additionalcitationids=\"CR39\" citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]\u0026ndash;[\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e], [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e], [\u003cspan additionalcitationids=\"CR29 CR30 CR31 CR32 CR33 CR34\" citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]\u0026ndash;[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e] (Summary of these studies are presented in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). In contrast, this study applied the missForest imputation method to minimize distortions from missing data while retaining potentially informative outlier observations via temporary replacement with missing indicators followed by missForest imputation. This novel hybrid pre-processing approach strengthened data integrity relative to complete outlier exclusion.\u003c/p\u003e \u003cp\u003eAmong the evaluated ML algorithms, the LogitBoost classifier demonstrated particularly strong performance, achieving a sensitivity of 0.8095, specificity of 0.9464, and accuracy of 0.9091 (From Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). These metrics represent improvements over past LogitBoost [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e], [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e] and other single model implementations such as DT, NN [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e], SVM [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e], and DL [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e] for this dataset, according to available reporting. While direct comparison must be made cautiously, consistently superior performance across multiple predictive metrics provides evidence the systematic pre-processing protocols implemented herein may enhance diabetes classification.\u003c/p\u003e \u003cp\u003eA key gap addressed was the lack of standardized procedures for missing and anomalous data handling prior to modelling, which had been noted as potentially compromising accuracy in earlier studies. Through meticulous data pre-processing centered on judicious imputation and outlier management followed by rigorous model selection and validation, the current study substantively enhanced data integrity and strengthened resultant predictive performance relative to the past efforts studied herewith. Comprehensive model evaluation further allowed elucidation of appropriate algorithmic choices.\u003c/p\u003e \u003cp\u003eA noteworthy contribution of the current study was the development and implementation of a systematic and rigorous framework for strategic pre-processing of diabetes-related clinical data prior to modelling. By judiciously applying the missForest imputation technique and thoughtfully incorporating outliers through a temporary replacement approach, critical missing and anomalous data points were retained without compromising dataset integrity. This enabled more accurate identification and extraction of informative patterns from the data compared to previous methods involving complete outlier removal or omission of missing values. Comprehensive evaluation of diverse ML algorithms further permitted elucidation of optimal techniques for classification. Critically, the combined impact of these strategic pre-processing protocols and hybrid data-driven techniques elevate this research above prior work, addressing longstanding limitations to substantially advance both understanding and predictive performance for early diabetes detection using ML. Therefore, the current study makes a significant methodological contribution toward maximizing the potential of supervised learning techniques applied to clinical datasets. The methodical pre-processing framework and hybrid data-driven techniques applied here notably elevated diabetes risk prediction for this seminal cohort above prior levels, addressing longstanding weaknesses to inform best practices maximizing ML accuracy from this dataset. The research fundamentally advances understanding of diabetes classification.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eIn summary, this study implemented a rigorous methodological framework centering on strategic pre-processing to enhance ML -based diabetes classification utilizing the Pima Indian dataset. By judiciously applying missForest imputation and thoughtfully incorporating outliers through a temporary replacement approach, critical missing and anomalous data were retained without compromising data integrity. Comprehensive evaluation of multiple algorithms then permitted elucidation of optimal techniques for predictive model development. Critically, these systematic pre-processing protocols and hybrid data-driven techniques substantively strengthened data quality and classification performance compared to previous investigations that typically excluded incomplete observations.\u003c/p\u003e \u003cp\u003eNotably, the LogitBoost classifier emerged as highly accurate for prediction based on its consistent achievement of sensitivity of 0.8095, specificity of 0.9464, and accuracy of 0.9091 - improvements over past single model implementations for this cohort. Overall, the research addresses longstanding weaknesses in diabetes risk prediction through its novel, carefully optimized methodology. The workflow developed makes a significant advancement by maximizing extraction of meaningful patterns from clinical health records to inform early disease detection.\u003c/p\u003e \u003cp\u003eFinally, this work evidences the merit of implementing standardized, rigorous data pre-processing centered on judicious imputation and outlier management prior to predictive modeling. The protocol introduced herein fundamentally advances both understanding and capacity for accurate ML -based identification of diabetes risk. Future research employing these systematic techniques on additional patient cohorts will continue to elucidate best practices and algorithmic choices optimizing predictive healthcare applications.\u003c/p\u003e"},{"header":"List of abbreviations","content":"\u003cp\u003eBelow is an alphabetically ordered list of abbreviations utilized in this study:\u003c/p\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003eAI: Artificial Intelligence\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eACC: Accuracy\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eANN: Artificial Neural Network\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eANOVA: Analysis of Variance\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAOR: Adjusted Odds Ratio\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eAUC: Area Under the Curve\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eavNNet: Model Averaging Neural Network\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eBAL: Balanced\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eBAYESGLM: Bayesian Generalized Linear Model\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eBMI: Body Mass Index\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eCNN: Convolutional Neural Network\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eCV: Cross Validation\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eDirect-CNN: Direct Convolutional Neural Network\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eDL: Deep Learning\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eDT: Decision Tree\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eFDA: Flexible Discriminant Analysis\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003egcvEarth: Multivariate Adaptive Regression Splines with Generalized Cross-Validation\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eHRF: High Risk Factor\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eHTN: Hypertension\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eIDE: Integrated Development Environment\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003elda: Linear Discriminant Analysis\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eLR: Logistic Regression\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eML: Machine Learning\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003emlpML: Multilayer Perceptron with Multiple Layers\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eNN/nnet: Neural Network\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eNIDDK: National Institute of Diabetes and Digestive and Kidney Disease\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eOR: Odds Ratio\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003ePIDD: Pima Indian Dataset\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003ePNN: Probabilistic Neural Network\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eRF: Random Forest\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eRRFglobal: Random Forest using the Global Feature Selection Algorithm\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eROC: Receiver Operating Characteristic\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eSVM: Support Vector Machine\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003esvmPoly: Support Vector Machine with Polynomial Kernel\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003eSVM-RFE: Support Vector Machine-Recursive Feature Elimination\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Declarations","content":"\u003ch2\u003eConsent for Publication\u003c/h2\u003e\n\u003cp\u003eIn obtaining consent for publication, the researchers ensured transparency and respect for the contributions of all involved parties. Authors collectively agreed to the submission of this work for potential publication, and any potential conflicts of interest were disclosed. This collaborative approach aligns with the author\u0026rsquo;s commitment to responsible and accountable research dissemination.\u003c/p\u003e\n\u003ch2\u003eCompeting Interests\u003c/h2\u003e\n\u003cp\u003eThe authors declare no competing interests in relation to the research, authorship, and publication of this paper. This ensures that the study\u0026rsquo;s outcomes and interpretations are solely driven by scientific rigor and unbiased exploration of the topic at hand.\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eThis research was self-funded by the authors, reflecting their commitment to advancing the understanding of diabetes prediction through rigorous investigation and analysis.\u003c/p\u003e\n\u003ch2\u003eAuthor\u0026rsquo;s Contributions\u003c/h2\u003e\n\u003cp\u003eThe collaborative research on diabetes was led by three key contributors: MSH, AD, and PKK. MSH conceived and designed the study, conducted all methodological work including data pre-processing, ML model development and evaluation, statistical analysis, visualization, writing of the initial manuscript draft, and revisions. PKK and AD contributed to interpretation of results and revisions to the final manuscript. All authors discussed the results, contributed to the intellectual content, and approved the final manuscript. The distribution of tasks among the authors effectively utilized their expertise, resulting in a comprehensive and well-rounded research paper. Furthermore, the meticulous report writing was primarily undertaken by MSH, ensuring the clarity, coherence, and accuracy of the manuscript\u0026rsquo;s content.\u003c/p\u003e\n\u003ch2\u003eAcknowledgements\u003c/h2\u003e\n\u003cp\u003eThe authors would like to thank Late Professor Dr. Md. Jahanur Rahman for his support and initial ideas of this work. The authors also appreciate the efforts of the reviewers who provided constructive comments that helped improve the quality of this manuscript. Finally, authors are grateful to Agro Environmetrics Research Lab for providing Lab facilities and Computers during the research and writing process.\u003c/p\u003e\n\u003ch2\u003eAuthors\u0026rsquo; Information\u003c/h2\u003e\n\u003cp\u003eMd. Sifat Hossain holds an MSc degree and MPhil Research Fellow in the Department of Statistics at the University of Rajshahi, Bangladesh. With expertise in data analysis and a keen interest in healthcare research, he contributed significantly to the statistical analysis and interpretation of this study. Email: [email protected]\u003c/p\u003e\n\u003cp\u003eAstami Devnath possesses an MSc degree and affiliated as an MPhil Research Fellow at the Institute of Environmental Science, University of Rajshahi, Bangladesh. Her expertise is particularly focused on the domains of econometrics and environmental health. Email: [email protected] \u003c/p\u003e\n\u003cp\u003eProvash Kumar Karmokar, PhD is a Professor in the Department of Statistics at the University of Rajshahi, Bangladesh. With a wealth of knowledge in statistical methodologies and health economics, he played a pivotal role in guiding the statistical analysis and ensuring its rigor. Email: [email protected] \u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eW. H. Herman, \u0026ldquo;The global burden of diabetes: an overview,\u0026rdquo; \u003cem\u003eDiabetes Mellit. Dev. Ctries. underserved communities\u003c/em\u003e, pp. 1\u0026ndash;5, 2017.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;World Health Organization Diabetes.\u0026rdquo; https://www.who.int/health-topics/diabetes.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;How is the pancreas involved in diabetes?\u0026rdquo; https://www.medicalnewstoday.com/articles/325018#how-is-the-pancreas-linked-with-diabetes.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;Type 2 Diabetes Causes and Risk Factors.\u0026rdquo; https://www.webmd.com/diabetes/diabetes-causes.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;Prediabetes \u0026ndash; Your Chance to Prevent Type 2 Diabetes.\u0026rdquo; https://www.cdc.gov/diabetes/basics/prediabetes.html.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK).\u0026rdquo; https://www.niddk.nih.gov/health-information/diabetes.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;Blood Sugar Level Ranges.\u0026rdquo; https://www.diabetes.co.uk/diabetes_care/blood-sugar-level-ranges.html.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;Diabetes - Long-Term Effects.\u0026rdquo; https://www.betterhealth.vic.gov.au/health/conditionsandtreatments/diabetes-long-term-effects.\u003c/li\u003e\n\u003cli\u003eS. A. Kaveeshwar and J. Cornwall, \u0026ldquo;The current state of diabetes mellitus in India,\u0026rdquo; \u003cem\u003eAustralas. Med. J.\u003c/em\u003e, vol. 7, no. 1, p. 45, 2014.\u003c/li\u003e\n\u003cli\u003eJ. Chaki, S. T. Ganesh, S. K. Cidham, and S. A. Theertan, \u0026ldquo;Machine learning and artificial intelligence based Diabetes Mellitus detection and self-management: A systematic review,\u0026rdquo; \u003cem\u003eJ. King Saud Univ. Inf. Sci.\u003c/em\u003e, vol. 34, no. 6, pp. 3204\u0026ndash;3225, 2022.\u003c/li\u003e\n\u003cli\u003eC.-L. Huang, M.-C. Chen, and C.-J. Wang, \u0026ldquo;Credit scoring with a data mining approach based on support vector machines,\u0026rdquo; \u003cem\u003eExpert Syst. Appl.\u003c/em\u003e, vol. 33, no. 4, pp. 847\u0026ndash;856, 2007.\u003c/li\u003e\n\u003cli\u003eI. Contreras and J. Vehi, \u0026ldquo;Artificial intelligence for diabetes management and decision support: literature review,\u0026rdquo; \u003cem\u003eJ. Med. Internet Res.\u003c/em\u003e, vol. 20, no. 5, p. e10775, 2018.\u003c/li\u003e\n\u003cli\u003eG. Swapna, S. Kp, and R. Vinayakumar, \u0026ldquo;Automated detection of diabetes using CNN and CNN-LSTM network and heart rate signals,\u0026rdquo; \u003cem\u003eProcedia Comput. Sci.\u003c/em\u003e, vol. 132, pp. 1253\u0026ndash;1262, 2018.\u003c/li\u003e\n\u003cli\u003eM. W. Craven and J. W. Shavlik, \u0026ldquo;Using neural networks for data mining,\u0026rdquo; \u003cem\u003eFutur. Gener. Comput. Syst.\u003c/em\u003e, vol. 13, no. 2\u0026ndash;3, pp. 211\u0026ndash;229, 1997.\u003c/li\u003e\n\u003cli\u003eJ. D. B. Gil, P. Reidsma, K. Giller, L. Todman, A. Whitmore, and M. van Ittersum, \u0026ldquo;Sustainable development goal 2: Improved targets and indicators for agriculture and food security,\u0026rdquo; \u003cem\u003eAmbio\u003c/em\u003e, vol. 48, no. 7, pp. 685\u0026ndash;698, 2019.\u003c/li\u003e\n\u003cli\u003eM. Lee \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;How to respond to the fourth industrial revolution, or the second information technology revolution? Dynamic new combinations between technology, market, and society through open innovation,\u0026rdquo; \u003cem\u003eJ. Open Innov. Technol. Mark. Complex.\u003c/em\u003e, vol. 4, no. 3, p. 21, 2018.\u003c/li\u003e\n\u003cli\u003eN. Ahmed \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Machine learning based diabetes prediction and development of smart web application,\u0026rdquo; vol. 2, pp. 229\u0026ndash;241, 2021.\u003c/li\u003e\n\u003cli\u003eZ. Dong \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Prediction of 3-year risk of diabetic kidney disease using machine learning based on electronic medical records,\u0026rdquo; vol. 20, no. 1, pp. 1\u0026ndash;10, 2022.\u003c/li\u003e\n\u003cli\u003eY. Du, A. R. Rafferty, F. M. McAuliffe, L. Wei, and C. %J S. R. Mooney, \u0026ldquo;An explainable machine learning-based clinical decision support system for prediction of gestational diabetes mellitus,\u0026rdquo; vol. 12, no. 1, pp. 1\u0026ndash;14, 2022.\u003c/li\u003e\n\u003cli\u003eH. Gupta, H. Varshney, T. K. Sharma, N. Pachauri, O. P. %J C. Verma, and I. Systems, \u0026ldquo;Comparative performance analysis of quantum machine learning with deep learning for diabetes prediction,\u0026rdquo; vol. 8, no. 4, pp. 3073\u0026ndash;3087, 2022.\u003c/li\u003e\n\u003cli\u003eJ. J. Khanam and S. Y. %J I. C. T. E. Foo, \u0026ldquo;A comparison of machine learning algorithms for diabetes prediction,\u0026rdquo; vol. 7, no. 4, pp. 432\u0026ndash;439, 2021.\u003c/li\u003e\n\u003cli\u003eS. Kumari, D. Kumar, and M. %J I. J. of C. C. in E. Mittal, \u0026ldquo;An ensemble approach for classification and prediction of diabetes mellitus using soft voting classifier,\u0026rdquo; vol. 2, pp. 40\u0026ndash;46, 2021.\u003c/li\u003e\n\u003cli\u003eH. Lu, S. Uddin, F. Hajati, M. A. Moni, and M. %J A. I. Khushi, \u0026ldquo;A patient network-based machine learning model for disease prediction: The case of type 2 diabetes mellitus,\u0026rdquo; vol. 52, no. 3, pp. 2411\u0026ndash;2422, 2022.\u003c/li\u003e\n\u003cli\u003eP. Rajendra, S. %J C. M. Latifi, and P. in B. Update, \u0026ldquo;Prediction of diabetes using logistic regression and ensemble techniques,\u0026rdquo; vol. 1, p. 100032, 2021.\u003c/li\u003e\n\u003cli\u003eM. Ravaut \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Predicting adverse outcomes due to diabetes complications with machine learning using administrative health data,\u0026rdquo; vol. 4, no. 1, pp. 1\u0026ndash;12, 2021.\u003c/li\u003e\n\u003cli\u003eJ. W. Smith, J. E. Everhart, W. C. Dickson, W. C. Knowler, and R. S. Johannes, \u0026ldquo;Using the ADAP learning algorithm to forecast the onset of diabetes mellitus,\u0026rdquo; \u003cem\u003eProc. Symp. Comput. Appl. Med. Care\u003c/em\u003e, vol. 12, no. 1, pp. 261\u0026ndash;265, 1988.\u003c/li\u003e\n\u003cli\u003eH. M. El-Bakry, M. El-Dib, and A. Kamal, \u0026ldquo;A comparative study of machine learning techniques for diabetes disease prediction,\u0026rdquo; in \u003cem\u003e2019 8th International Conference on Computer and Knowledge Engineering (ICCKE)\u003c/em\u003e, 2019, pp. 179\u0026ndash;184.\u003c/li\u003e\n\u003cli\u003eP. V Lakshmi and N. Chilamkurti, \u0026ldquo;Deep belief network based ensemble classifier for diabetes disease prediction,\u0026rdquo; \u003cem\u003eComput. Electr. Eng.\u003c/em\u003e, vol. 72, pp. 418\u0026ndash;430, 2019.\u003c/li\u003e\n\u003cli\u003eY. Shang, Z. Chen, and G. Jiang, \u0026ldquo;A deep learning model to predict diabetes through electronic health records,\u0026rdquo; \u003cem\u003eIEEE Access\u003c/em\u003e, vol. 7, pp. 54445\u0026ndash;54452, 2019.\u003c/li\u003e\n\u003cli\u003eC. S. Preetha, \u0026ldquo;A comparative study of data mining techniques for prediction of diabetes,\u0026rdquo; \u003cem\u003earXiv Prepr. arXiv1211.5730\u003c/em\u003e, 2012.\u003c/li\u003e\n\u003cli\u003eK. S. Abdul Nazeer and M. P. Sebastian, \u0026ldquo;Detecting diabetes on set of biological data using decision tree algorithm,\u0026rdquo; \u003cem\u003eFar East J. Theor. Stat.\u003c/em\u003e, vol. 27, no. 1, pp. 1\u0026ndash;14, 2009.\u003c/li\u003e\n\u003cli\u003e\u0026Ouml;. N. Geomat, \u0026ldquo;Classification of Pima Indian diabetes dataset using neural networks,\u0026rdquo; \u003cem\u003eProcedia-Social Behav. Sci.\u003c/em\u003e, vol. 195, pp. 1408\u0026ndash;1417, 2015.\u003c/li\u003e\n\u003cli\u003eM. G\u0026uuml;ler and K. Polat, \u0026ldquo;Detecting Pima Indians diabetes using neural networks and estimated statistical classification functions,\u0026rdquo; \u003cem\u003eExpert Syst. Appl.\u003c/em\u003e, vol. 28, no. 4, pp. 707\u0026ndash;715, 2005.\u003c/li\u003e\n\u003cli\u003eM. A. Tahir, A. Bouridane, and F. Kurugollu, \u0026ldquo;Classifying medical data using SVM with combined kernel functions,\u0026rdquo; \u003cem\u003eJ. Appl. Clin. Med. Phys.\u003c/em\u003e, vol. 12, no. 1, p. 3475, 2011.\u003c/li\u003e\n\u003cli\u003eM. F. Akay, \u0026ldquo;Support vector machines combined with feature selection for breast cancer diagnosis,\u0026rdquo; \u003cem\u003eExpert Syst. Appl.\u003c/em\u003e, vol. 36, no. 2, pp. 3240\u0026ndash;3247, 2009.\u003c/li\u003e\n\u003cli\u003eP. Vepakomma, O. Gupta, A. Dewan, and P. Roux, \u0026ldquo;Reducing disparity in diabetes prediction models using adversarial representation learning,\u0026rdquo; \u003cem\u003earXiv Prepr. arXiv1807.00540\u003c/em\u003e, 2018.\u003c/li\u003e\n\u003cli\u003eR. Agrawal and A. Choudhary, \u0026ldquo;Comparison of supervised machine learning algorithms for disease prediction,\u0026rdquo; \u003cem\u003eIOSR J. Comput. Eng.\u003c/em\u003e, pp. 18\u0026ndash;24, 2016.\u003c/li\u003e\n\u003cli\u003eN. Zhang and J. Li, \u0026ldquo;Pima Indians diabetes prediction based on ant colony optimization classifier,\u0026rdquo; in \u003cem\u003eLecture Notes in Computer Science\u003c/em\u003e, vol. 3173, Springer, 2003, pp. 230\u0026ndash;235.\u003c/li\u003e\n\u003cli\u003eM. S. A. Kumar, V. Ravi, and K. B. Raja, \u0026ldquo;Prediction of diabetes using probabilistic neural network with feature extraction,\u0026rdquo; \u003cem\u003eMeasurement\u003c/em\u003e, vol. 71, pp. 53\u0026ndash;60, 2015.\u003c/li\u003e\n\u003cli\u003eİ. Karabulut, \u0026ldquo;Comparison of generalized regression neural network algorithms for diabetes disease diagnosis,\u0026rdquo; \u003cem\u003eJ. Intell. Syst.\u003c/em\u003e, vol. 22, no. 2, pp. 247\u0026ndash;256, 2013.\u003c/li\u003e\n\u003cli\u003eT. N. Joshi and P. M. Chawan, \u0026ldquo;Logistic regression and svm based diabetes prediction system,\u0026rdquo; \u003cem\u003eInt. J. Technol. Res. Eng.\u003c/em\u003e, vol. 5, pp. 4347\u0026ndash;4350, 2018.\u003c/li\u003e\n\u003cli\u003eN. Yuvaraj and K. R. SriPreethaa, \u0026ldquo;Diabetes prediction in healthcare systems using machine learning algorithms on Hadoop cluster,\u0026rdquo; \u003cem\u003eCluster Comput.\u003c/em\u003e, vol. 22, no. Suppl 1, pp. 1\u0026ndash;9, 2019.\u003c/li\u003e\n\u003cli\u003eD. Sisodia and D. S. Sisodia, \u0026ldquo;Prediction of diabetes using classification algorithms,\u0026rdquo; \u003cem\u003eProcedia Comput. Sci.\u003c/em\u003e, vol. 132, pp. 1578\u0026ndash;1585, 2018.\u003c/li\u003e\n\u003cli\u003eE. O. Olaniyi and K. Adnan, \u0026ldquo;Onset diabetes diagnosis using artificial neural network,\u0026rdquo; \u003cem\u003eInt J Sci Eng Res\u003c/em\u003e, vol. 5, no. 10, pp. 754\u0026ndash;759, 2014.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;Machine Learning Databases.\u0026rdquo; ftp://ftp.ics.uci.edu/pub/machine-learning-databases.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;Machine Learning Repository.\u0026rdquo; http://www.ics.uci.edu/~mlearn/MLRepository.html.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK).\u0026rdquo; https://www.niddk.nih.gov/health-information/diabetes.\u003c/li\u003e\n\u003cli\u003eD. J. Stekhoven and P. %J B. B\u0026uuml;hlmann, \u0026ldquo;MissForest\u0026mdash;non-parametric missing value imputation for mixed-type data,\u0026rdquo; vol. 28, no. 1, pp. 112\u0026ndash;118, 2012.\u003c/li\u003e\n\u003cli\u003eA. G. Karegowda, A. S. Manjunath, and M. A. Jayaram, \u0026ldquo;Comparative study of attribute selection using GA,\u0026rdquo; \u003cem\u003eInt. J. Adv. Soft Comput. its Appl.\u003c/em\u003e, vol. 2, no. 1, pp. 45\u0026ndash;68, 2010.\u003c/li\u003e\n\u003cli\u003eX. Wu \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Top 10 algorithms in data mining,\u0026rdquo; \u003cem\u003eKnowl. Inf. Syst.\u003c/em\u003e, vol. 14, no. 1, pp. 1\u0026ndash;37, 2008.\u003c/li\u003e\n\u003cli\u003eM. Ahmed, A. N. Mahmood, and M. R. Islam, \u0026ldquo;A machine learning approach for early diagnosis of diabetes disease,\u0026rdquo; in \u003cem\u003e2012 international conference on informatics, electronics \u0026amp; vision (ICIEV)\u003c/em\u003e, 2012, pp. 1\u0026ndash;5.\u003c/li\u003e\n\u003cli\u003eS. Raschka, \u003cem\u003ePython Machine Learning\u003c/em\u003e. Packt Publishing Ltd, 2018.\u003c/li\u003e\n\u003cli\u003eD. Koley and D. Saha, \u0026ldquo;Comparative study of supervised machine learning algorithms for Pima indians diabetes data set,\u0026rdquo; \u003cem\u003eJ. Theor. Appl. Inf. Technol.\u003c/em\u003e, vol. 95, no. 16, 2016.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Diabetes, Early Detection, Machine Learning, Artificial Intelligence, Missing Data, Outliers, missForest, and Classification","lastPublishedDoi":"10.21203/rs.3.rs-3364064/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3364064/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cb\u003eBackground\u003c/b\u003e\u003c/p\u003e \u003cp\u003eThe rising global threat of diabetes demands timely detection to prevent its complications. Data scientists and practitioners are seen to be used AI and some other classification models on different aspects. Nevertheless, addressing missing data and outlier\u0026rsquo;s accurate predictions may be questionable. As such incorporating ML and AI for early diagnosis has gained attention. This study integrates medical knowledge and what types of advanced technology to develop a comprehensive diabetes classification model, focusing on handling missing values and outliers to achieve improved accuracy in early disease identification.\u003c/p\u003e\u003cp\u003e\u003cb\u003eMethods\u003c/b\u003e\u003c/p\u003e \u003cp\u003eThe researcher\u0026rsquo;s methodology prioritized meticulous data pre-processing to enhance analysis quality. To address missing data, the researchers utilized the missForest method, employing a multistage imputation process that minimizes data loss and distortions. Outlier detection relied on Mahalanobis squared distances, identifying anomalous data points. Instead of outright removal, the researchers strategically leveraged the missForest method, known for its robust imputation capabilities. Temporarily replacing outliers with missing values, this approach seamlessly integrated imputation. The ensuing hybrid data, minus extreme outliers and enriched via missForest, formed the foundation for subsequent analysis and modelling. Model selection and evaluation were performed on pre-processed data. This analysis incorporated two-step CV: initial dataset partition (80% training, 20% testing) and ten iterations of ten-fold cross-validation for model stability and parameter optimization. A diverse array of ML models\u0026mdash;LogitBoost, mlpWeightDecayML, avNNet, and others\u0026mdash;were assessed. Metrics such as sensitivity, specificity, precision, recall, F1-score, AUC, accuracy, and Kappa score were scrutinized.\u003c/p\u003e\u003cp\u003e\u003cb\u003eResults\u003c/b\u003e\u003c/p\u003e \u003cp\u003eAmong the models examined, LogitBoost emerged as a strong contender with a sensitivity of 0.8095, specificity of 0.9464, precision of 0.85, recall of 0.8095, F1-score of 0.8293, AUC of 0.7888, accuracy of 0.9091, and Kappa score of 0.7674. However, the comparative results showcase varying performances across different metrics and models. Sensitivity ranged from 0.6792 to 0.9057, specificity from 0.6 to 0.9464, and precision from 0.5455 to 0.85.\u003c/p\u003e\u003cp\u003e\u003cb\u003eConclusions\u003c/b\u003e\u003c/p\u003e \u003cp\u003eIn summation, the methodical approach has illuminated the path toward enhanced diabetes classification accuracy. By diligently addressing missing values through the robust missForest method and tactfully managing outliers using the hybrid approach, the researchers have elevated the integrity and quality of the PIMA dataset. This strategic handling of missing values and outliers has not only fortified the dataset against potential distortions but has also culminated in improved accuracy in diabetes classification. Through the synergy of meticulous pre-processing, strategic outlier management, and comprehensive model evaluation, the researchers have contributed valuable insights into the realm of early diabetes detection.\u003c/p\u003e","manuscriptTitle":"Handling Missing Values and Outliers in Advanced Data Pre-processing: An Enhancement of Diabetes Classification Accuracy","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-09-25 16:59:54","doi":"10.21203/rs.3.rs-3364064/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"506eef97-928c-4120-8aa8-393cca243536","owner":[],"postedDate":"September 25th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":24826221,"name":"Health sciences/Risk factors"},{"id":24826222,"name":"Scientific community and society/Scientific community/Research management"}],"tags":[],"updatedAt":"2023-09-25T16:59:57+00:00","versionOfRecord":[],"versionCreatedAt":"2023-09-25 16:59:54","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3364064","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3364064","identity":"rs-3364064","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00