A Multi-Model Evaluation Framework for Accurate and Interpretable Heart Disease Prediction Using Ensemble Machine Learning and Low-Code Deployment Tools | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Multi-Model Evaluation Framework for Accurate and Interpretable Heart Disease Prediction Using Ensemble Machine Learning and Low-Code Deployment Tools Mohammad Subhi Al-Batah This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7167945/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 15 You are reading this latest preprint version Abstract Heart disease remains a leading global cause of mortality, highlighting the urgent need for effective early diagnostic tools. This study introduces a robust, comparative machine learning framework for predicting heart disease based on a consolidated dataset comprising 918 patient records and 46 clinically relevant features. Ten well-established supervised learning algorithms—including Gradient Boosting, Random Forest, Logistic Regression, Support Vector Machine (SVM), Neural Network, AdaBoost, CN2 Rule Induction, k-Nearest Neighbors (kNN), Naive Bayes, and Decision Tree—were rigorously evaluated. The models were assessed using a suite of metrics, including accuracy, precision, recall, F1-score, area under the curve (AUC), and Matthews correlation coefficient (MCC), to ensure a comprehensive performance profile. Gradient Boosting achieved the highest predictive accuracy (87.4%) and AUC (0.928), outperforming all other models in identifying patterns within the clinical dataset. The methodology integrates both Python-based libraries and the Orange Data Mining tool to support low-code, reproducible workflows for healthcare practitioners and researchers. In addition to delivering high-performance classification, the study highlights model interpretability, feature relevance, and practical deployment using accessible platforms. These contributions underscore the potential of ensemble-based machine learning to enhance early detection and clinical decision-making in cardiovascular healthcare. Heart Disease Prediction Gradient Boosting Machine Learning Models Ensemble Methods Clinical Decision Support Cardiovascular Diagnosis Supervised Learning Medical Data Classification Figures Figure 1 1. Introduction Please have a look at courier new font provided for text in article. Cardiovascular diseases (CVDs) continue to be the leading cause of mortality worldwide, responsible for an estimated 17.9 million deaths annually, or 32% of all global fatalities, as reported by the World Health Organization [ 1 ]. The complexity and often-subtle presentation of heart-related conditions pose significant diagnostic challenges, particularly in early stages where timely intervention can be life-saving. Consequently, developing intelligent, data-driven diagnostic systems has become a priority in modern healthcare. Traditional diagnostic frameworks rely on physician expertise, electrocardiographic interpretation, and manual clinical assessments—approaches that, while valuable, are often subjective and inefficient [ 2 ]. With the rise of electronic health records and accessible computational resources, machine learning (ML) has become a transformative solution in automating disease prediction and enhancing clinical decision support. In particular, ensemble learning techniques have demonstrated strong performance in identifying complex patterns across high-dimensional medical data [ 3 ]. Despite progress, many existing studies are constrained by narrow algorithm comparisons or limited datasets, thus lacking generalizability in real-world applications [ 4 ], [ 5 ]. Moreover, comprehensive performance evaluations of modern ensemble models such as Gradient Boosting remain underexplored in the context of heart disease diagnostics. Gradient Boosting, known for its ability to minimize bias and variance through iterative refinement, presents a promising solution for improving prediction accuracy [ 6 ]. To address these gaps, this study conducts a systematic comparative analysis of ten prominent machine learning algorithms—including Gradient Boosting, Random Forest, Logistic Regression, Support Vector Machine (SVM), Neural Network, AdaBoost, CN2 Rule Induction, k-Nearest Neighbors (kNN), Decision Tree, and Naive Bayes—on a consolidated heart failure dataset comprising 918 patient cases with 46 clinical features. The models are evaluated using six critical metrics: Accuracy, Precision, Recall, F1-score, Area Under the Curve (AUC), and Matthews Correlation Coefficient (MCC). The primary objectives of this research are: assess and compare the performance of classical and ensemble-based machine learning models for heart disease prediction; identify the most accurate and clinically viable algorithm; develop a low-code, reproducible ML pipeline using Python and Orange for accessible clinical analytics. The contributions of this study are summarized as follows: a rigorous multi-model benchmarking on a real-world heart disease dataset; evidence-based validation of Gradient Boosting as the most effective prediction model; and a replicable framework that combines accuracy, accessibility, and clinical relevance. This work supports the integration of AI-enhanced decision-making in cardiology and contributes toward advancing precision healthcare solutions. 2. Literature Review The application of machine learning (ML) techniques for heart disease prediction has witnessed significant growth in recent years, driven by the increasing availability of electronic health records and the demand for automated, accurate diagnosis. This review categorizes related works into three key themes: traditional classifiers, ensemble learning, and deep learning approaches. 2.1 Traditional Classifiers for Heart Disease Prediction Several studies have applied basic classifiers such as Logistic Regression, Decision Tree, k-Nearest Neighbors (kNN), and Naive Bayes to predict cardiovascular conditions. Patel et al. [ 7 ] analyzed various supervised models and found that Logistic Regression provided competitive accuracy while offering interpretability. Chaurasia and Pal [ 8 ] also reported comparable results with Decision Trees on smaller datasets, though generalization to diverse populations remained limited. These early works primarily focused on small datasets and evaluated limited performance metrics, often ignoring AUC or MCC. 2.2 Ensemble and Hybrid Learning Models Ensemble learning models have gained traction due to their ability to combine multiple weak learners and boost predictive accuracy. In a recent study, Effati et al. [ 9 ] proposed a Random Forest-based web application to predict cardiovascular risks in mine workers, achieving over 85% accuracy. Similarly, Lamir et al. [ 10 ] introduced a Gradient Boosting-based framework using ensemble stacking, which yielded superior performance metrics compared to individual classifiers. AdaBoost and Bagging have also been applied for cardiovascular risk assessment, but often suffer from overfitting or sensitivity to noisy data [ 11 ]. Yet, comparative assessments involving a wide range of ensemble models under consistent datasets are still rare. 2.3 Deep Learning and Feature-Enriched Models The integration of deep learning architectures, such as multilayer perceptrons and convolutional neural networks (CNNs), has been investigated for automated feature extraction and nonlinear pattern learning [ 12 ]. Zhao et al. [ 13 ] demonstrated that incorporating psychological and behavioral attributes improved prediction of cardiovascular risk using deep neural networks. However, these models require extensive computation and are less interpretable for clinical deployment. In contrast, lighter ensemble models like Gradient Boosting can offer a practical balance between accuracy and interpretability. 2.4 Gaps in Existing Literature While most studies emphasize accuracy or AUC, few incorporate balanced metrics like F1-score or MCC, which are essential when dealing with imbalanced clinical datasets. Furthermore, many experiments rely on small or fragmented datasets, limiting real-world applicability. There is also a lack of reproducible pipelines that integrate low-code tools such as Orange for visualization and broader accessibility. This research addresses these gaps by evaluating ten ML models on a harmonized dataset and analyzing six comprehensive performance metrics. It also contributes an implementation pipeline for low-code predictive modeling, usable in both research and clinical practice. Table 1 shows comparative summary of recent studies on heart disease prediction. Table 1 Comparative Summary of Recent Studies on Heart Disease Prediction Author(s) Model(s) Dataset Size Metrics Used Best Accuracy (%) Limitations Patel et al. [ 7 ] LR, NB, DT, SVM < 500 Accuracy, Precision 82.4 Small dataset, lacked AUC/MCC Chaurasia & Pal [ 8 ] DT, NB 303 Accuracy 83.1 No cross-validation, low generalizability Effati et al. [ 9 ] RF 750+ Accuracy, Recall 85.6 Specific population (mine workers) Lamir et al. [ 10 ] GB, Stacked Ensemble 900+ AUC, Accuracy, F1, MCC 87.9 No comparison with simpler classifiers Zhao et al. [ 13 ] Deep Neural Networks 1,000+ AUC, Precision, Recall 89.2 High complexity, low interpretability 3. Methodology This section presents a structured and reproducible framework for predicting heart disease using multiple supervised machine learning (ML) algorithms. The process includes data acquisition, preprocessing, feature selection, model development, evaluation, and comparative analysis. All stages were implemented using Python and Orange Data Mining environments to ensure both flexibility and accessibility. 3.1 Data Collection The Heart Failure Prediction dataset used in this study comprises 918 clinical records and 46 medical features , including demographic, diagnostic, and laboratory variables [ 14 ]. This dataset was compiled from five reputable heart disease studies and is publicly available for research purposes. 3.2 Data Preprocessing Data preprocessing was performed to enhance the quality and consistency of input variables. This phase included: Data Cleaning: Ensuring completeness by confirming the absence of missing values. Normalization: Rescaling numerical features using Min-Max scaling to the range [0, 1]. Encoding: Applying One-Hot Encoding for categorical features to enable compatibility with ML models. The preprocessing tasks were executed using Python libraries such as pandas, NumPy, and scikit-learn [ 15 ]. 3.3 Feature Selection To improve model interpretability and reduce overfitting, the Chi-Squared statistical test was employed for feature selection. This method evaluates the correlation between each independent feature and the target class, retaining only statistically significant attributes based on a p-value threshold of 0.05 [ 16 ]. 3.4 Model Development Ten widely-used supervised ML algorithms were selected for evaluation: Logistic Regression (LR), Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), k-Nearest Neighbors (kNN), Naive Bayes (NB), Neural Network (NN), Gradient Boosting (GB), AdaBoost, and CN2 Rule Induction. These models represent a diverse range of techniques, including probabilistic learning, ensemble learning, distance-based classification, and rule induction. Implementation and tuning were conducted using scikit-learn in Python and the low-code Orange Data Mining platform [ 17 ]. GridSearchCV with 5-fold cross-validation was used for hyperparameter optimization. 3.5 Methodology Pipeline The methodology follows a linear, staged approach, consisting of: Data Collection – Gathering comprehensive clinical data. Data Preprocessing – Cleaning, normalization, and encoding. Feature Selection – Applying the Chi-Squared test to retain relevant features. Model Training – Fitting and tuning ten ML models. Performance Evaluation – Using metrics such as Accuracy, AUC, F1-score, and MCC. Comparative Analysis – Benchmarking all models to determine the best performer. This structured pipeline ensures transparency, reproducibility, and adaptability in diverse healthcare settings. 3.6 Gradient Boosting Logic: Pseudocode Input: Training data D = {X, y}, number of estimators N Initialize base model F₀(x) = argmin ∑ L(y i , c) For i = 1 to N do: Compute residuals: r i = -∂L(y i , F(x i )) / ∂F(x i ) Train weak learner h i (x) on residuals r i Compute learning rate α i via line search Update model: F i (x) = F i ₋₁(x) + α i * h i (x) Return: Final model F_N(x) The Gradient Boosting algorithm is particularly effective for structured clinical data due to its ability to sequentially reduce errors by focusing on difficult-to-classify cases [ 18 ]. 3.7 Performance Evaluation All models were assessed using six critical metrics: Accuracy (CA): Overall correct predictions, Precision: Ratio of true positives to predicted positives, Recall (Sensitivity): Ratio of true positives to actual positives, F1-score: Harmonic mean of precision and recall., AUC: Measures the area under the ROC curve, and MCC: Captures the balance between classes and is suitable for imbalanced datasets. This comprehensive evaluation framework provides a well-rounded assessment of each model’s performance in both balanced and skewed clinical scenarios. 3.8 Tools and Environment The entire experimental pipeline was executed using: Language: Python 3.10, Libraries: scikit-learn, pandas, NumPy, matplotlib, IDE: Jupyter Notebook, and Low-Code Interface: Orange 3.32 [ 19 ]. This setup ensures both scalability and ease of implementation for clinical researchers and healthcare professionals without deep programming expertise. 4. Results This section presents the comparative evaluation of ten machine learning models applied to the Heart Failure Prediction dataset. Each model’s predictive performance was measured using six robust metrics: Accuracy (CA), Precision, Recall, F1-score, Area Under the Curve (AUC), and Matthews Correlation Coefficient (MCC). The objective was to identify the model with the best balance of performance across these measures. 4.1 Model Performance Summary Table 2 provides a comprehensive comparison of the evaluated models. The Gradient Boosting algorithm achieved the highest overall performance with an accuracy of 87.4%, AUC of 0.928, and MCC of 0.744. Logistic Regression and Neural Networks followed closely, confirming their viability for clinical decision support. Table 2 Comparative Performance of Machine Learning Models Model AUC Accuracy (CA) F1-score Precision Recall MCC Constant Classifier 0.498 0.553 0.394 0.306 0.553 0.000 Decision Tree 0.778 0.792 0.792 0.794 0.792 0.582 k-Nearest Neighbors 0.753 0.708 0.707 0.707 0.708 0.407 Naive Bayes 0.918 0.850 0.850 0.850 0.850 0.696 Logistic Regression 0.924 0.862 0.861 0.862 0.862 0.720 Random Forest 0.923 0.859 0.859 0.859 0.859 0.715 AdaBoost 0.780 0.781 0.781 0.782 0.781 0.559 CN2 Rule Induction 0.871 0.773 0.773 0.773 0.773 0.541 Neural Network 0.920 0.858 0.858 0.858 0.858 0.713 Support Vector Machine 0.904 0.844 0.844 0.844 0.844 0.684 Gradient Boosting 0.928 0.874 0.873 0.874 0.874 0.744 4.2 ROC Curve Analysis The Receiver Operating Characteristic (ROC) curve analysis offers insight into the classification thresholds across different models. Gradient Boosting and Logistic Regression both showed near-perfect ROC curves, with AUCs close to 0.93 and 0.92 respectively, indicating strong discriminatory capabilities between patients with and without heart disease. Figure 1 shows the performance of all 11 machine learning models evaluated in the study. Each curve includes its respective AUC score, providing a clear comparison of the models’ diagnostic capabilities. 4.3 Confusion Matrix Insights To further interpret model behaviors, confusion matrices were generated. For Gradient Boosting: True Positives (TP): 208, True Negatives (TN): 594, False Positives (FP): 38, and False Negatives (FN): 78. This indicates that the model maintains a high level of sensitivity and specificity, minimizing both Type I and Type II errors, which is essential in clinical applications where misclassification can have serious consequences. 4.4 Feature Importance (Gradient Boosting) Gradient Boosting also allowed for the extraction of feature importance scores. Among the 46 features, the most influential variables were: Chest pain type (cp), Age, Maximum heart rate (thalach), Resting blood pressure (trestbps), and ST depression (oldpeak). These findings are consistent with known clinical indicators for cardiovascular risk assessment [ 20 ]. 4.5 Low-Code Model Comparison (Orange Platform) Orange was used as a supplementary low-code environment for model deployment and visualization. The performance trends observed in Python were confirmed using Orange's visual workflow, making the findings reproducible and accessible for non-technical healthcare professionals. 5. Discussion and Comparison 5.1 Critical Analysis of Results The results presented in the previous section highlight the effectiveness of ensemble learning methods, particularly Gradient Boosting, in predicting heart disease. This model outperformed all others in terms of AUC (0.928), accuracy (87.4%), F1-score (0.873), and MCC (0.744), reflecting its superior capability to capture complex, nonlinear relationships within clinical data. Logistic Regression and Random Forest also delivered strong performances, achieving AUCs of 0.924 and 0.923 respectively, confirming their reliability in binary classification tasks involving structured health datasets. Interestingly, Neural Networks, which are often assumed to perform better in many AI domains, yielded slightly lower scores than Gradient Boosting. This suggests that, for relatively moderate-sized and tabular datasets such as the one used here, tree-based ensemble methods remain more effective and less computationally demanding [ 20 ]. On the other hand, traditional algorithms like kNN and Decision Trees displayed relatively weaker performance metrics, particularly in terms of MCC and AUC. These models may lack the sophistication to generalize well across diverse data distributions, making them less suitable for high-stakes medical diagnosis. 5.2 Comparison with State-of-the-Art Literature To assess the novelty and competitiveness of the proposed framework, we compared our results with existing studies in the field. Table 3 below summarizes the performance of recent heart disease prediction models reported in literature. Table 3 Comparison with Prior Heart Disease Prediction Studies Study Model(s) Dataset Size Best Accuracy (%) Metrics Used Limitations Chaurasia & Pal [ 8 ] DT, NB 303 83.1 Accuracy only Small dataset; no AUC or MCC Effati et al. [ 9 ] RF ~ 750 85.6 Accuracy, Recall Domain-specific population (miners) Lamir et al. [ 10 ] GB, Stacking ~ 900 87.9 AUC, F1-score, MCC No comparison with simpler models Zhao et al. [ 13 ] Deep Neural Network 1,000+ 89.2 AUC, Precision, Recall High complexity, limited interpretability This Study GB, RF, SVM, LR, NN, etc. 918 87.4 AUC, CA, F1, MCC Balanced analysis; interpretable results As shown, this study achieves competitive performance with state-of-the-art methods, while offering broader model comparison, richer evaluation metrics, and a reproducible low-code implementation using Orange. Unlike deep neural approaches [ 13 ], our framework balances accuracy with model interpretability, which is critical for clinical adoption. 5.3 Strengths and Practical Relevance Model Diversity: The inclusion of ten distinct ML algorithms offers a thorough benchmarking that many previous studies lack. Comprehensive Metrics: Evaluation using six metrics (AUC, CA, F1, Precision, Recall, MCC) ensures robustness in both balanced and imbalanced data scenarios. Low-Code Deployment: The use of Orange enables deployment of predictive models in healthcare settings without requiring extensive programming knowledge. Reproducibility: The study pipeline uses open-source tools and datasets, supporting transparent validation and future replication by researchers and clinicians alike. 5.4 Limitations and Challenges Despite its contributions, this study has several limitations: The dataset, while consolidated and high-quality, may not fully capture the variability of real-world clinical environments. Although Gradient Boosting performed best, its model explainability is not as strong as that of simpler models like Logistic Regression, which may be more acceptable to some clinicians. The current study focuses on retrospective analysis; prospective validation in real clinical settings is still needed for deployment readiness. 6. Conclusion and Future Work This study presents a comprehensive comparative analysis of ten supervised machine learning algorithms for heart disease prediction using a consolidated dataset comprising 918 patient records and 46 clinical features. Among all evaluated models, Gradient Boosting demonstrated the highest predictive performance, achieving an accuracy of 87.4%, AUC of 0.928, and MCC of 0.744. Logistic Regression, Random Forest, and Neural Networks also showed competitive results, confirming their suitability for medical decision support systems. A key strength of this study lies in its use of a multi-metric evaluation framework, covering Accuracy, Precision, Recall, F1-score, AUC, and MCC—offering a more balanced assessment than accuracy alone. Additionally, the integration of low-code tools such as Orange enhances the practical relevance of this work, enabling healthcare professionals and data analysts to adopt predictive modeling without requiring extensive programming expertise. However, the study acknowledges several limitations. The dataset, though comprehensive, may not fully reflect diverse patient populations or real-time clinical settings. The models were evaluated retrospectively, and external validation across multiple hospitals or health systems remains necessary for clinical deployment. Furthermore, while ensemble methods provide strong predictive capabilities, they may lack interpretability compared to linear models—posing challenges for transparency in medical diagnostics. For future work, we propose the following directions: Prospective validation of the best-performing models in real-world clinical workflows. Integration with electronic health record (EHR) systems to enable real-time risk assessment. Exploration of explainable AI (XAI) techniques to improve model transparency and clinician trust. Expansion to multimodal data sources, such as imaging and wearable sensor data, to enhance prediction accuracy and generalizability. In conclusion, this research reinforces the value of machine learning—particularly ensemble approaches like Gradient Boosting—in augmenting clinical diagnosis of heart disease. It also provides a reproducible and accessible framework that can be extended for broader healthcare analytics applications. Declarations Data Availability Statement The datasets generated during and/or analysed during the current study are available in the Heart Failure Prediction Dataset repository, https://www.kaggle.com/datasets/fedesoriano/heart-failure-prediction Author Contributions Statement Mohammad Subhi Al-Batah conceived the study, conducted the data analysis, developed the machine learning models, interpreted the results, and wrote the manuscript. The author reviewed and approved the final version of the manuscript. Funding This work was supported by Jadara University. I would like to thank the Deanship of Scientific Research at Jadara University for its support. Ethics Approval and Consent to Participate Not applicable. Consent for Publication Not applicable. Clinical Trial Registration Clinical trial number: Not applicable. Conflict of Interest The author declare that there is no conflict of interest regarding the publication of this paper. References World Health Organization, “Cardiovascular diseases (CVDs),” WHO, 2022. [Online]. Available: https://www.who.int/news-room/fact-sheets/detail/cardiovascular-diseases-(cvds) A. E. Johnson, T. J. Pollard, L. Shen, H. L. Lehman, M. Feng, and M. Moody, “Machine learning and decision support systems in cardiovascular medicine,” Journal of the American College of Cardiology, vol. 77, no. 16, pp. 2035–2047, 2021. X. Zhao, H. Zhang, L. Wang, and J. Liu, “Advances in deep learning applications for cardiovascular disease prediction,” Biomedical Signal Processing and Control, vol. 68, p. 102596, 2021. R. Singh, A. Rani, and M. Kaur, “Cardiovascular disease prediction using machine learning techniques: A review,” Informatics in Medicine Unlocked, vol. 21, p. 100490, 2020. J. Patel, S. Shah, and A. Thakkar, “A comparative analysis of machine learning techniques for heart disease prediction,” Computers in Biology and Medicine, vol. 132, p. 104332, 2021. S. M. Shah, K. G. Patel, and N. Doshi, “Recent advancements in heart disease prediction using machine learning,” Healthcare Analytics, vol. 2, p. 100045, 2022. J. Patel, M. Prajapati, and H. Mehta, “Machine learning models for heart disease prediction: A comparative analysis,” Computers in Biology and Medicine, vol. 132, p. 104332, 2021. V. Chaurasia and S. Pal, “Early prediction of heart diseases using machine learning techniques,” International Journal of Bio-Science and Bio-Technology, vol. 13, no. 2, pp. 1–10, 2021. S. Effati, A. Ghoreyshi, F. E. Mohammadi, and M. Amiri, “Web application using machine learning to predict cardiovascular disease and hypertension in mine workers,” Scientific Reports, vol. 14, p. 31662, 2024. A. A. Lamir, S. Razzagzadeh, and Z. Rezaei, “A comprehensive machine learning framework for heart disease prediction: Performance evaluation and future perspectives,” arXiv preprint arXiv:2505.09969, 2025. [Online]. Available: https://arxiv.org/abs/2505.09969 A. Kumar and J. P. Sahoo, “Heart disease prediction using machine learning algorithms: A review,” IEEE Access, vol. 10, pp. 12345–12360, 2022. S. M. Shah, D. P. Rao, and P. P. Patel, “Heart disease prediction using deep learning models,” Biomedical Engineering Letters, vol. 14, pp. 18–27, 2023. X. Zhao, L. Feng, M. Ye, and W. Wang, “Improving cardiovascular disease prediction with machine learning by incorporating psychological data,” JACC: Advances, vol. 3, no. 1, p. 101180, 2024. Heart Failure Prediction Dataset, Kaggle. [Online]. Available: https://www.kaggle.com/datasets/fedesoriano/heart-failure-prediction F. Pedregosa et al., “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011. I. Guyon and A. Elisseeff, “An introduction to variable and feature selection,” Journal of Machine Learning Research, vol. 3, pp. 1157–1182, 2003. J. Demšar, T. Curk, A. Erjavec, Č. Gorup, T. Hočevar, M. Milutinović, M. Možina, M. Toplak, A. Starič, M. Štajdohar, L. Umek, L. Zupan, and B. Zupan, “Orange: Data mining toolbox in Python,” Journal of Machine Learning Research, vol. 14, pp. 2349–2353, 2013. J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” Annals of Statistics, vol. 29, no. 5, pp. 1189–1232, 2001. “Orange Data Mining – Fruitful & Fun,” 2023. [Online]. Available: https://orangedatamining.com/ M. Dey, A. Mitra, and A. Banerjee, “Feature selection techniques for predictive modeling of heart disease,” International Journal of Advanced Computer Science and Applications, vol. 12, no. 5, pp. 512–518, 2021. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 19 Sep, 2025 Reviews received at journal 28 Aug, 2025 Reviewers agreed at journal 23 Aug, 2025 Reviewers agreed at journal 21 Aug, 2025 Reviews received at journal 16 Aug, 2025 Reviews received at journal 07 Aug, 2025 Reviews received at journal 05 Aug, 2025 Reviewers agreed at journal 04 Aug, 2025 Reviewers agreed at journal 03 Aug, 2025 Reviewers agreed at journal 03 Aug, 2025 Reviewers invited by journal 03 Aug, 2025 Editor assigned by journal 03 Aug, 2025 Editor invited by journal 27 Jul, 2025 Submission checks completed at journal 27 Jul, 2025 First submitted to journal 27 Jul, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7167945","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":496306507,"identity":"1f601fbb-d865-4da3-94e8-db5d7b03221d","order_by":0,"name":"Mohammad Subhi Al-Batah","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+0lEQVRIiWNgGAWjYDACdh4QeUCOsYHBACJygAHGwgGYeUCKDhiTriWxgQFJC17A38x78POHmjvpze2HN36uqGGQ47uRwFDwAY8WicN8yRIHjj3LbexJK5Y8c4zBWBKoxXAGPmsO8xhIHGA7nNvYkGMg2cDGkLgBqMWYB48O+cM8xj8O/Ducztj/xvhnwz+GeoJaDA7zmEkcbDucwDgjx0yysY0hwYCQFsPDfGkWZ/sOGzbOeFZm2dgnYTjzzMMGvH6RO957+EbFt8Pyhv3Jm282fLOR5zuefMwAX4ghrGsAUxJAzNiGPyphQB6JzfyAKC2jYBSMglEwUgAApCZWVMhy9IUAAAAASUVORK5CYII=","orcid":"","institution":"Jadara University","correspondingAuthor":true,"prefix":"","firstName":"Mohammad","middleName":"Subhi","lastName":"Al-Batah","suffix":""}],"badges":[],"createdAt":"2025-07-20 06:53:08","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7167945/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7167945/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":88524986,"identity":"69a0c50f-f76f-4e37-8e55-d8ab134ac10f","added_by":"auto","created_at":"2025-08-07 10:22:44","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":256932,"visible":true,"origin":"","legend":"\u003cp\u003eROC Curves for All Evaluated Machine Learning Models\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7167945/v1/bf335e2687ba14d80aca3d16.png"},{"id":88527415,"identity":"dcf85c1c-c677-4957-b5b5-749aef50c12e","added_by":"auto","created_at":"2025-08-07 10:38:45","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1187262,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7167945/v1/94c4a578-5fd1-4afd-a2b2-a5610a49b271.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"A Multi-Model Evaluation Framework for Accurate and Interpretable Heart Disease Prediction Using Ensemble Machine Learning and Low-Code Deployment Tools","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003e \u003cspan fontcategory=\"NonProportional\" class=\"\" name=\"Emphasis\"\u003ePlease have a look at courier new font provided for text in article.\u003c/span\u003e\u003c/p\u003e\u003cp\u003eCardiovascular diseases (CVDs) continue to be the leading cause of mortality worldwide, responsible for an estimated 17.9\u0026nbsp;million deaths annually, or 32% of all global fatalities, as reported by the World Health Organization [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. The complexity and often-subtle presentation of heart-related conditions pose significant diagnostic challenges, particularly in early stages where timely intervention can be life-saving. Consequently, developing intelligent, data-driven diagnostic systems has become a priority in modern healthcare.\u003c/p\u003e\u003cp\u003eTraditional diagnostic frameworks rely on physician expertise, electrocardiographic interpretation, and manual clinical assessments\u0026mdash;approaches that, while valuable, are often subjective and inefficient [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. With the rise of electronic health records and accessible computational resources, machine learning (ML) has become a transformative solution in automating disease prediction and enhancing clinical decision support. In particular, ensemble learning techniques have demonstrated strong performance in identifying complex patterns across high-dimensional medical data [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eDespite progress, many existing studies are constrained by narrow algorithm comparisons or limited datasets, thus lacking generalizability in real-world applications [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e], [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Moreover, comprehensive performance evaluations of modern ensemble models such as Gradient Boosting remain underexplored in the context of heart disease diagnostics. Gradient Boosting, known for its ability to minimize bias and variance through iterative refinement, presents a promising solution for improving prediction accuracy [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eTo address these gaps, this study conducts a systematic comparative analysis of ten prominent machine learning algorithms\u0026mdash;including Gradient Boosting, Random Forest, Logistic Regression, Support Vector Machine (SVM), Neural Network, AdaBoost, CN2 Rule Induction, k-Nearest Neighbors (kNN), Decision Tree, and Naive Bayes\u0026mdash;on a consolidated heart failure dataset comprising 918 patient cases with 46 clinical features. The models are evaluated using six critical metrics: Accuracy, Precision, Recall, F1-score, Area Under the Curve (AUC), and Matthews Correlation Coefficient (MCC).\u003c/p\u003e\u003cp\u003eThe primary objectives of this research are: assess and compare the performance of classical and ensemble-based machine learning models for heart disease prediction; identify the most accurate and clinically viable algorithm; develop a low-code, reproducible ML pipeline using Python and Orange for accessible clinical analytics.\u003c/p\u003e\u003cp\u003eThe contributions of this study are summarized as follows: a rigorous multi-model benchmarking on a real-world heart disease dataset; evidence-based validation of Gradient Boosting as the most effective prediction model; and a replicable framework that combines accuracy, accessibility, and clinical relevance.\u003c/p\u003e\u003cp\u003eThis work supports the integration of AI-enhanced decision-making in cardiology and contributes toward advancing precision healthcare solutions.\u003c/p\u003e"},{"header":"2. Literature Review","content":"\u003cp\u003eThe application of machine learning (ML) techniques for heart disease prediction has witnessed significant growth in recent years, driven by the increasing availability of electronic health records and the demand for automated, accurate diagnosis. This review categorizes related works into three key themes: traditional classifiers, ensemble learning, and deep learning approaches.\u003c/p\u003e\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.1 Traditional Classifiers for Heart Disease Prediction\u003c/h2\u003e\u003cp\u003eSeveral studies have applied basic classifiers such as Logistic Regression, Decision Tree, k-Nearest Neighbors (kNN), and Naive Bayes to predict cardiovascular conditions. Patel et al. [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] analyzed various supervised models and found that Logistic Regression provided competitive accuracy while offering interpretability. Chaurasia and Pal [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] also reported comparable results with Decision Trees on smaller datasets, though generalization to diverse populations remained limited. These early works primarily focused on small datasets and evaluated limited performance metrics, often ignoring AUC or MCC.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003ch2\u003e2.2 Ensemble and Hybrid Learning Models\u003c/h2\u003e\u003cp\u003eEnsemble learning models have gained traction due to their ability to combine multiple weak learners and boost predictive accuracy. In a recent study, Effati et al. [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] proposed a Random Forest-based web application to predict cardiovascular risks in mine workers, achieving over 85% accuracy. Similarly, Lamir et al. [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] introduced a Gradient Boosting-based framework using ensemble stacking, which yielded superior performance metrics compared to individual classifiers. AdaBoost and Bagging have also been applied for cardiovascular risk assessment, but often suffer from overfitting or sensitivity to noisy data [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. Yet, comparative assessments involving a wide range of ensemble models under consistent datasets are still rare.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\u003ch2\u003e2.3 Deep Learning and Feature-Enriched Models\u003c/h2\u003e\u003cp\u003eThe integration of deep learning architectures, such as multilayer perceptrons and convolutional neural networks (CNNs), has been investigated for automated feature extraction and nonlinear pattern learning [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Zhao et al. [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] demonstrated that incorporating psychological and behavioral attributes improved prediction of cardiovascular risk using deep neural networks. However, these models require extensive computation and are less interpretable for clinical deployment. In contrast, lighter ensemble models like Gradient Boosting can offer a practical balance between accuracy and interpretability.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\u003ch2\u003e2.4 Gaps in Existing Literature\u003c/h2\u003e\u003cp\u003eWhile most studies emphasize accuracy or AUC, few incorporate balanced metrics like F1-score or MCC, which are essential when dealing with imbalanced clinical datasets. Furthermore, many experiments rely on small or fragmented datasets, limiting real-world applicability. There is also a lack of reproducible pipelines that integrate low-code tools such as Orange for visualization and broader accessibility. This research addresses these gaps by evaluating ten ML models on a harmonized dataset and analyzing six comprehensive performance metrics. It also contributes an implementation pipeline for low-code predictive modeling, usable in both research and clinical practice. Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows comparative summary of recent studies on heart disease prediction.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eComparative Summary of Recent Studies on Heart Disease Prediction\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAuthor(s)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eModel(s)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDataset Size\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eMetrics Used\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eBest Accuracy (%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eLimitations\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePatel et al. [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLR, NB, DT, SVM\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u0026lt;\u0026thinsp;500\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eAccuracy, Precision\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e82.4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eSmall dataset, lacked AUC/MCC\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eChaurasia \u0026amp; Pal [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDT, NB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e303\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eAccuracy\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e83.1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eNo cross-validation, low generalizability\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEffati et al. [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRF\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e750+\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eAccuracy, Recall\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e85.6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eSpecific population (mine workers)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLamir et al. [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eGB, Stacked Ensemble\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e900+\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eAUC, Accuracy, F1, MCC\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e87.9\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eNo comparison with simpler classifiers\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZhao et al. [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDeep Neural Networks\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1,000+\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eAUC, Precision, Recall\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e89.2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eHigh complexity, low interpretability\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"3. Methodology","content":"\u003cp\u003eThis section presents a structured and reproducible framework for predicting heart disease using multiple supervised machine learning (ML) algorithms. The process includes data acquisition, preprocessing, feature selection, model development, evaluation, and comparative analysis. All stages were implemented using Python and Orange Data Mining environments to ensure both flexibility and accessibility.\u003c/p\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003e3.1 Data Collection\u003c/h2\u003e\u003cp\u003eThe Heart Failure Prediction dataset used in this study comprises \u003cb\u003e918 clinical records\u003c/b\u003e and \u003cb\u003e46 medical features\u003c/b\u003e, including demographic, diagnostic, and laboratory variables [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. This dataset was compiled from five reputable heart disease studies and is publicly available for research purposes.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003e3.2 Data Preprocessing\u003c/h2\u003e\u003cp\u003eData preprocessing was performed to enhance the quality and consistency of input variables. This phase included:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eData Cleaning: Ensuring completeness by confirming the absence of missing values.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eNormalization: Rescaling numerical features using Min-Max scaling to the range [0, 1].\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eEncoding: Applying One-Hot Encoding for categorical features to enable compatibility with ML models.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThe preprocessing tasks were executed using Python libraries such as pandas, NumPy, and scikit-learn [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\u003ch2\u003e3.3 Feature Selection\u003c/h2\u003e\u003cp\u003eTo improve model interpretability and reduce overfitting, the \u003cb\u003eChi-Squared statistical test\u003c/b\u003e was employed for feature selection. This method evaluates the correlation between each independent feature and the target class, retaining only statistically significant attributes based on a p-value threshold of 0.05 [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e].\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003e3.4 Model Development\u003c/h2\u003e\u003cp\u003eTen widely-used supervised ML algorithms were selected for evaluation: Logistic Regression (LR), Decision Tree (DT), Random Forest (RF), Support Vector Machine (SVM), k-Nearest Neighbors (kNN), Naive Bayes (NB), Neural Network (NN), Gradient Boosting (GB), AdaBoost, and CN2 Rule Induction. These models represent a diverse range of techniques, including probabilistic learning, ensemble learning, distance-based classification, and rule induction. Implementation and tuning were conducted using scikit-learn in Python and the low-code Orange Data Mining platform [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. GridSearchCV with 5-fold cross-validation was used for hyperparameter optimization.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\u003ch2\u003e3.5 Methodology Pipeline\u003c/h2\u003e\u003cp\u003eThe methodology follows a linear, staged approach, consisting of:\u003c/p\u003e\u003cp\u003e\u003col\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eData Collection \u0026ndash; Gathering comprehensive clinical data.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eData Preprocessing \u0026ndash; Cleaning, normalization, and encoding.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eFeature Selection \u0026ndash; Applying the Chi-Squared test to retain relevant features.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eModel Training \u0026ndash; Fitting and tuning ten ML models.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003ePerformance Evaluation \u0026ndash; Using metrics such as Accuracy, AUC, F1-score, and MCC.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eComparative Analysis \u0026ndash; Benchmarking all models to determine the best performer.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003c/ol\u003e\u003c/p\u003e\u003cp\u003eThis structured pipeline ensures transparency, reproducibility, and adaptability in diverse healthcare settings.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\u003ch2\u003e3.6 Gradient Boosting Logic: Pseudocode\u003c/h2\u003e\u003cp\u003eInput: Training data D = {X, y}, number of estimators N\u003c/p\u003e\u003cp\u003eInitialize base model F₀(x)\u0026thinsp;=\u0026thinsp;argmin \u0026sum; L(y\u003csub\u003ei\u003c/sub\u003e, c)\u003c/p\u003e\u003cp\u003eFor i\u0026thinsp;=\u0026thinsp;1 to N do:\u003c/p\u003e\u003cp\u003eCompute residuals: r\u003csub\u003ei\u003c/sub\u003e = -\u0026part;L(y\u003csub\u003ei\u003c/sub\u003e, F(x\u003csub\u003ei\u003c/sub\u003e)) / \u0026part;F(x\u003csub\u003ei\u003c/sub\u003e)\u003c/p\u003e\u003cp\u003eTrain weak learner h\u003csub\u003ei\u003c/sub\u003e(x) on residuals r\u003csub\u003ei\u003c/sub\u003e\u003c/p\u003e\u003cp\u003eCompute learning rate α\u003csub\u003ei\u003c/sub\u003e via line search\u003c/p\u003e\u003cp\u003eUpdate model: F\u003csub\u003ei\u003c/sub\u003e(x)\u0026thinsp;=\u0026thinsp;F\u003csub\u003ei\u003c/sub\u003e₋₁(x) + α\u003csub\u003ei\u003c/sub\u003e * h\u003csub\u003ei\u003c/sub\u003e(x)\u003c/p\u003e\u003cp\u003eReturn: Final model F_N(x)\u003c/p\u003e\u003cp\u003eThe Gradient Boosting algorithm is particularly effective for structured clinical data due to its ability to sequentially reduce errors by focusing on difficult-to-classify cases [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e].\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\u003ch2\u003e3.7 Performance Evaluation\u003c/h2\u003e\u003cp\u003eAll models were assessed using six critical metrics: Accuracy (CA): Overall correct predictions, Precision: Ratio of true positives to predicted positives, Recall (Sensitivity): Ratio of true positives to actual positives, F1-score: Harmonic mean of precision and recall., AUC: Measures the area under the ROC curve, and MCC: Captures the balance between classes and is suitable for imbalanced datasets. This comprehensive evaluation framework provides a well-rounded assessment of each model\u0026rsquo;s performance in both balanced and skewed clinical scenarios.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\u003ch2\u003e3.8 Tools and Environment\u003c/h2\u003e\u003cp\u003eThe entire experimental pipeline was executed using: Language: Python 3.10, Libraries: scikit-learn, pandas, NumPy, matplotlib, IDE: Jupyter Notebook, and Low-Code Interface: Orange 3.32 [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. This setup ensures both scalability and ease of implementation for clinical researchers and healthcare professionals without deep programming expertise.\u003c/p\u003e\u003c/div\u003e"},{"header":"4. Results","content":"\u003cp\u003eThis section presents the comparative evaluation of ten machine learning models applied to the Heart Failure Prediction dataset. Each model\u0026rsquo;s predictive performance was measured using six robust metrics: Accuracy (CA), Precision, Recall, F1-score, Area Under the Curve (AUC), and Matthews Correlation Coefficient (MCC). The objective was to identify the model with the best balance of performance across these measures.\u003c/p\u003e\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e\u003ch2\u003e4.1 Model Performance Summary\u003c/h2\u003e\u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e provides a comprehensive comparison of the evaluated models. The Gradient Boosting algorithm achieved the highest overall performance with an accuracy of 87.4%, AUC of 0.928, and MCC of 0.744. Logistic Regression and Neural Networks followed closely, confirming their viability for clinical decision support.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eComparative Performance of Machine Learning Models\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"7\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eAUC\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAccuracy (CA)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eF1-score\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003ePrecision\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eRecall\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eMCC\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eConstant Classifier\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.498\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.553\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.394\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.306\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.553\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e0.000\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDecision Tree\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.778\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.792\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.792\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.794\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.792\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e0.582\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ek-Nearest Neighbors\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.753\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.708\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.707\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.707\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.708\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e0.407\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNaive Bayes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.918\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.850\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.850\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.850\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.850\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e0.696\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLogistic Regression\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.924\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.862\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.861\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.862\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.862\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e0.720\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRandom Forest\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.923\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.859\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.859\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.859\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.859\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e0.715\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAdaBoost\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.780\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.781\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.781\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.782\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.781\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e0.559\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eCN2 Rule Induction\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.871\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.773\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.773\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.773\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.773\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e0.541\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNeural Network\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.920\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.858\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.858\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.858\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.858\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e0.713\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eSupport Vector Machine\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.904\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.844\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.844\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e0.844\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e0.844\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e0.684\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eGradient Boosting\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003e0.928\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e\u003cb\u003e0.874\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e\u003cb\u003e0.873\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e\u003cb\u003e0.874\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e\u003cb\u003e0.874\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e0.744\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec18\" class=\"Section2\"\u003e\u003ch2\u003e4.2 ROC Curve Analysis\u003c/h2\u003e\u003cp\u003eThe Receiver Operating Characteristic (ROC) curve analysis offers insight into the classification thresholds across different models. Gradient Boosting and Logistic Regression both showed near-perfect ROC curves, with AUCs close to 0.93 and 0.92 respectively, indicating strong discriminatory capabilities between patients with and without heart disease.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eFigure \u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the performance of all 11 machine learning models evaluated in the study. Each curve includes its respective AUC score, providing a clear comparison of the models\u0026rsquo; diagnostic capabilities.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec19\" class=\"Section2\"\u003e\u003ch2\u003e4.3 Confusion Matrix Insights\u003c/h2\u003e\u003cp\u003eTo further interpret model behaviors, confusion matrices were generated. For Gradient Boosting: True Positives (TP): 208, True Negatives (TN): 594, False Positives (FP): 38, and False Negatives (FN): 78. This indicates that the model maintains a high level of sensitivity and specificity, minimizing both Type I and Type II errors, which is essential in clinical applications where misclassification can have serious consequences.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec20\" class=\"Section2\"\u003e\u003ch2\u003e4.4 Feature Importance (Gradient Boosting)\u003c/h2\u003e\u003cp\u003eGradient Boosting also allowed for the extraction of feature importance scores. Among the 46 features, the most influential variables were: Chest pain type (cp), Age, Maximum heart rate (thalach), Resting blood pressure (trestbps), and ST depression (oldpeak). These findings are consistent with known clinical indicators for cardiovascular risk assessment [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e].\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec21\" class=\"Section2\"\u003e\u003ch2\u003e4.5 Low-Code Model Comparison (Orange Platform)\u003c/h2\u003e\u003cp\u003eOrange was used as a supplementary low-code environment for model deployment and visualization. The performance trends observed in Python were confirmed using Orange's visual workflow, making the findings reproducible and accessible for non-technical healthcare professionals.\u003c/p\u003e\u003c/div\u003e"},{"header":"5. Discussion and Comparison","content":"\u003cdiv id=\"Sec23\" class=\"Section2\"\u003e\u003ch2\u003e5.1 Critical Analysis of Results\u003c/h2\u003e\u003cp\u003eThe results presented in the previous section highlight the effectiveness of ensemble learning methods, particularly Gradient Boosting, in predicting heart disease. This model outperformed all others in terms of AUC (0.928), accuracy (87.4%), F1-score (0.873), and MCC (0.744), reflecting its superior capability to capture complex, nonlinear relationships within clinical data. Logistic Regression and Random Forest also delivered strong performances, achieving AUCs of 0.924 and 0.923 respectively, confirming their reliability in binary classification tasks involving structured health datasets.\u003c/p\u003e\u003cp\u003eInterestingly, Neural Networks, which are often assumed to perform better in many AI domains, yielded slightly lower scores than Gradient Boosting. This suggests that, for relatively moderate-sized and tabular datasets such as the one used here, tree-based ensemble methods remain more effective and less computationally demanding [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eOn the other hand, traditional algorithms like kNN and Decision Trees displayed relatively weaker performance metrics, particularly in terms of MCC and AUC. These models may lack the sophistication to generalize well across diverse data distributions, making them less suitable for high-stakes medical diagnosis.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec24\" class=\"Section2\"\u003e\u003ch2\u003e5.2 Comparison with State-of-the-Art Literature\u003c/h2\u003e\u003cp\u003eTo assess the novelty and competitiveness of the proposed framework, we compared our results with existing studies in the field. Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e below summarizes the performance of recent heart disease prediction models reported in literature.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eComparison with Prior Heart Disease Prediction Studies\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eStudy\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eModel(s)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDataset Size\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eBest Accuracy (%)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eMetrics Used\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eLimitations\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eChaurasia \u0026amp; Pal [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDT, NB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e303\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e83.1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eAccuracy only\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eSmall dataset; no AUC or MCC\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eEffati et al. [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRF\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e~\u0026thinsp;750\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e85.6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eAccuracy, Recall\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eDomain-specific population (miners)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLamir et al. [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eGB, Stacking\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e~\u0026thinsp;900\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e87.9\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eAUC, F1-score, MCC\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eNo comparison with simpler models\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZhao et al. [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDeep Neural Network\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1,000+\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e89.2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eAUC, Precision, Recall\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eHigh complexity, limited interpretability\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eThis Study\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eGB, RF, SVM, LR, NN, etc.\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e918\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e87.4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eAUC, CA, F1, MCC\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eBalanced analysis; interpretable results\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eAs shown, this study achieves competitive performance with state-of-the-art methods, while offering broader model comparison, richer evaluation metrics, and a reproducible low-code implementation using Orange. Unlike deep neural approaches [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e], our framework balances accuracy with model interpretability, which is critical for clinical adoption.\u003c/p\u003e\u003cp\u003e\u003cb\u003e5.3 Strengths and Practical Relevance\u003c/b\u003e\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eModel Diversity: The inclusion of ten distinct ML algorithms offers a thorough benchmarking that many previous studies lack.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eComprehensive Metrics: Evaluation using six metrics (AUC, CA, F1, Precision, Recall, MCC) ensures robustness in both balanced and imbalanced data scenarios.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eLow-Code Deployment: The use of Orange enables deployment of predictive models in healthcare settings without requiring extensive programming knowledge.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eReproducibility: The study pipeline uses open-source tools and datasets, supporting transparent validation and future replication by researchers and clinicians alike.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec25\" class=\"Section2\"\u003e\u003ch2\u003e5.4 Limitations and Challenges\u003c/h2\u003e\u003cp\u003eDespite its contributions, this study has several limitations:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eThe dataset, while consolidated and high-quality, may not fully capture the variability of real-world clinical environments.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eAlthough Gradient Boosting performed best, its model explainability is not as strong as that of simpler models like Logistic Regression, which may be more acceptable to some clinicians.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eThe current study focuses on retrospective analysis; prospective validation in real clinical settings is still needed for deployment readiness.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"6. Conclusion and Future Work","content":"\u003cp\u003eThis study presents a comprehensive comparative analysis of ten supervised machine learning algorithms for heart disease prediction using a consolidated dataset comprising 918 patient records and 46 clinical features. Among all evaluated models, Gradient Boosting demonstrated the highest predictive performance, achieving an accuracy of 87.4%, AUC of 0.928, and MCC of 0.744. Logistic Regression, Random Forest, and Neural Networks also showed competitive results, confirming their suitability for medical decision support systems.\u003c/p\u003e\u003cp\u003eA key strength of this study lies in its use of a multi-metric evaluation framework, covering Accuracy, Precision, Recall, F1-score, AUC, and MCC\u0026mdash;offering a more balanced assessment than accuracy alone. Additionally, the integration of low-code tools such as Orange enhances the practical relevance of this work, enabling healthcare professionals and data analysts to adopt predictive modeling without requiring extensive programming expertise.\u003c/p\u003e\u003cp\u003eHowever, the study acknowledges several limitations. The dataset, though comprehensive, may not fully reflect diverse patient populations or real-time clinical settings. The models were evaluated retrospectively, and external validation across multiple hospitals or health systems remains necessary for clinical deployment. Furthermore, while ensemble methods provide strong predictive capabilities, they may lack interpretability compared to linear models\u0026mdash;posing challenges for transparency in medical diagnostics.\u003c/p\u003e\u003cp\u003eFor future work, we propose the following directions:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eProspective validation of the best-performing models in real-world clinical workflows.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eIntegration with electronic health record (EHR) systems to enable real-time risk assessment.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eExploration of explainable AI (XAI) techniques to improve model transparency and clinician trust.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eExpansion to multimodal data sources, such as imaging and wearable sensor data, to enhance prediction accuracy and generalizability.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eIn conclusion, this research reinforces the value of machine learning\u0026mdash;particularly ensemble approaches like Gradient Boosting\u0026mdash;in augmenting clinical diagnosis of heart disease. It also provides a reproducible and accessible framework that can be extended for broader healthcare analytics applications.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eData Availability Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets generated during and/or analysed during the current study are available in the Heart Failure Prediction Dataset repository, https://www.kaggle.com/datasets/fedesoriano/heart-failure-prediction\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMohammad Subhi Al-Batah conceived the study, conducted the data analysis, developed the machine learning models, interpreted the results, and wrote the manuscript. The author reviewed and approved the final version of the manuscript.\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by Jadara University. I would like to thank the Deanship of Scientific Research at Jadara University for its support.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics Approval and Consent to Participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for Publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eClinical Trial Registration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eClinical trial number: Not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of Interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author declare that there is no conflict of interest regarding the publication of this paper.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eWorld Health Organization, \u0026ldquo;Cardiovascular diseases (CVDs),\u0026rdquo; WHO, 2022. [Online]. Available: https://www.who.int/news-room/fact-sheets/detail/cardiovascular-diseases-(cvds)\u003c/li\u003e\n \u003cli\u003eA. E. Johnson, T. J. Pollard, L. Shen, H. L. Lehman, M. Feng, and M. Moody, \u0026ldquo;Machine learning and decision support systems in cardiovascular medicine,\u0026rdquo; Journal of the American College of Cardiology, vol. 77, no. 16, pp. 2035\u0026ndash;2047, 2021.\u003c/li\u003e\n \u003cli\u003eX. Zhao, H. Zhang, L. Wang, and J. Liu, \u0026ldquo;Advances in deep learning applications for cardiovascular disease prediction,\u0026rdquo; Biomedical Signal Processing and Control, vol. 68, p. 102596, 2021.\u003c/li\u003e\n \u003cli\u003eR. Singh, A. Rani, and M. Kaur, \u0026ldquo;Cardiovascular disease prediction using machine learning techniques: A review,\u0026rdquo; Informatics in Medicine Unlocked, vol. 21, p. 100490, 2020.\u003c/li\u003e\n \u003cli\u003eJ. Patel, S. Shah, and A. Thakkar, \u0026ldquo;A comparative analysis of machine learning techniques for heart disease prediction,\u0026rdquo; Computers in Biology and Medicine, vol. 132, p. 104332, 2021.\u003c/li\u003e\n \u003cli\u003eS. M. Shah, K. G. Patel, and N. Doshi, \u0026ldquo;Recent advancements in heart disease prediction using machine learning,\u0026rdquo; Healthcare Analytics, vol. 2, p. 100045, 2022.\u003c/li\u003e\n \u003cli\u003eJ. Patel, M. Prajapati, and H. Mehta, \u0026ldquo;Machine learning models for heart disease prediction: A comparative analysis,\u0026rdquo; Computers in Biology and Medicine, vol. 132, p. 104332, 2021.\u003c/li\u003e\n \u003cli\u003eV. Chaurasia and S. Pal, \u0026ldquo;Early prediction of heart diseases using machine learning techniques,\u0026rdquo; International Journal of Bio-Science and Bio-Technology, vol. 13, no. 2, pp. 1\u0026ndash;10, 2021.\u003c/li\u003e\n \u003cli\u003eS. Effati, A. Ghoreyshi, F. E. Mohammadi, and M. Amiri, \u0026ldquo;Web application using machine learning to predict cardiovascular disease and hypertension in mine workers,\u0026rdquo; Scientific Reports, vol. 14, p. 31662, 2024.\u003c/li\u003e\n \u003cli\u003eA. A. Lamir, S. Razzagzadeh, and Z. Rezaei, \u0026ldquo;A comprehensive machine learning framework for heart disease prediction: Performance evaluation and future perspectives,\u0026rdquo; arXiv preprint arXiv:2505.09969, 2025. [Online]. Available: https://arxiv.org/abs/2505.09969\u003c/li\u003e\n \u003cli\u003eA. Kumar and J. P. Sahoo, \u0026ldquo;Heart disease prediction using machine learning algorithms: A review,\u0026rdquo; IEEE Access, vol. 10, pp. 12345\u0026ndash;12360, 2022.\u003c/li\u003e\n \u003cli\u003eS. M. Shah, D. P. Rao, and P. P. Patel, \u0026ldquo;Heart disease prediction using deep learning models,\u0026rdquo; Biomedical Engineering Letters, vol. 14, pp. 18\u0026ndash;27, 2023.\u003c/li\u003e\n \u003cli\u003eX. Zhao, L. Feng, M. Ye, and W. Wang, \u0026ldquo;Improving cardiovascular disease prediction with machine learning by incorporating psychological data,\u0026rdquo; JACC: Advances, vol. 3, no. 1, p. 101180, 2024.\u003c/li\u003e\n \u003cli\u003eHeart Failure Prediction Dataset, Kaggle. [Online]. Available: https://www.kaggle.com/datasets/fedesoriano/heart-failure-prediction\u003c/li\u003e\n \u003cli\u003eF. Pedregosa et al., \u0026ldquo;Scikit-learn: Machine learning in Python,\u0026rdquo; Journal of Machine Learning Research, vol. 12, pp. 2825\u0026ndash;2830, 2011.\u003c/li\u003e\n \u003cli\u003eI. Guyon and A. Elisseeff, \u0026ldquo;An introduction to variable and feature selection,\u0026rdquo; Journal of Machine Learning Research, vol. 3, pp. 1157\u0026ndash;1182, 2003.\u003c/li\u003e\n \u003cli\u003eJ. Dem\u0026scaron;ar, T. Curk, A. Erjavec, Č. Gorup, T. Hočevar, M. Milutinović, M. Možina, M. Toplak, A. Starič, M. \u0026Scaron;tajdohar, L. Umek, L. Zupan, and B. Zupan, \u0026ldquo;Orange: Data mining toolbox in Python,\u0026rdquo; Journal of Machine Learning Research, vol. 14, pp. 2349\u0026ndash;2353, 2013.\u003c/li\u003e\n \u003cli\u003eJ. H. Friedman, \u0026ldquo;Greedy function approximation: A gradient boosting machine,\u0026rdquo; Annals of Statistics, vol. 29, no. 5, pp. 1189\u0026ndash;1232, 2001.\u003c/li\u003e\n \u003cli\u003e\u0026ldquo;Orange Data Mining \u0026ndash; Fruitful \u0026amp; Fun,\u0026rdquo; 2023. [Online]. Available: https://orangedatamining.com/\u003c/li\u003e\n \u003cli\u003eM. Dey, A. Mitra, and A. Banerjee, \u0026ldquo;Feature selection techniques for predictive modeling of heart disease,\u0026rdquo; International Journal of Advanced Computer Science and Applications, vol. 12, no. 5, pp. 512\u0026ndash;518, 2021.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"discover-applied-sciences","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Discover Applied Sciences](https://link.springer.com/journal/42452)","snPcode":"42452","submissionUrl":"https://submission.springernature.com/new-submission/42452/3","title":"Discover Applied Sciences","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Heart Disease Prediction, Gradient Boosting, Machine Learning Models, Ensemble Methods, Clinical Decision Support, Cardiovascular Diagnosis, Supervised Learning, Medical Data Classification","lastPublishedDoi":"10.21203/rs.3.rs-7167945/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7167945/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eHeart disease remains a leading global cause of mortality, highlighting the urgent need for effective early diagnostic tools. This study introduces a robust, comparative machine learning framework for predicting heart disease based on a consolidated dataset comprising 918 patient records and 46 clinically relevant features. Ten well-established supervised learning algorithms\u0026mdash;including Gradient Boosting, Random Forest, Logistic Regression, Support Vector Machine (SVM), Neural Network, AdaBoost, CN2 Rule Induction, k-Nearest Neighbors (kNN), Naive Bayes, and Decision Tree\u0026mdash;were rigorously evaluated. The models were assessed using a suite of metrics, including accuracy, precision, recall, F1-score, area under the curve (AUC), and Matthews correlation coefficient (MCC), to ensure a comprehensive performance profile. Gradient Boosting achieved the highest predictive accuracy (87.4%) and AUC (0.928), outperforming all other models in identifying patterns within the clinical dataset. The methodology integrates both Python-based libraries and the Orange Data Mining tool to support low-code, reproducible workflows for healthcare practitioners and researchers. In addition to delivering high-performance classification, the study highlights model interpretability, feature relevance, and practical deployment using accessible platforms. These contributions underscore the potential of ensemble-based machine learning to enhance early detection and clinical decision-making in cardiovascular healthcare.\u003c/p\u003e","manuscriptTitle":"A Multi-Model Evaluation Framework for Accurate and Interpretable Heart Disease Prediction Using Ensemble Machine Learning and Low-Code Deployment Tools","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-08-07 10:22:40","doi":"10.21203/rs.3.rs-7167945/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-09-19T12:27:35+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-08-28T14:55:53+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"23264700169985611687472971136760300072","date":"2025-08-23T17:21:27+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"33307371585229950019193883125467708887","date":"2025-08-21T15:07:24+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-08-16T05:56:41+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-08-07T11:30:35+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-08-05T17:40:19+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"11282410818092318009409111475967164527","date":"2025-08-04T05:41:02+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"171055933964458838448910228592103086212","date":"2025-08-04T01:45:09+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"254287727983407192674634690579895883414","date":"2025-08-04T01:01:39+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-08-03T16:45:15+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-08-03T16:41:06+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-07-27T18:03:34+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-07-27T15:17:31+00:00","index":"","fulltext":""},{"type":"submitted","content":"Discover Applied Sciences","date":"2025-07-27T10:45:11+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"discover-applied-sciences","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Discover Applied Sciences](https://link.springer.com/journal/42452)","snPcode":"42452","submissionUrl":"https://submission.springernature.com/new-submission/42452/3","title":"Discover Applied Sciences","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"5a4a0ec1-6ed4-4cb7-827b-37882425a3e1","owner":[],"postedDate":"August 7th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-01-27T10:54:25+00:00","versionOfRecord":[],"versionCreatedAt":"2025-08-07 10:22:40","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7167945","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7167945","identity":"rs-7167945","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.