Machine learning prediction accuracy score: The use of Feature selection techniques

preprint OA: closed
View at publisher

Abstract

Medical facilities around the world attract and collect personal and healthcare context-based data of millions and billions of people every year. These centers have become information-rich data-centers. Advances in technology coupled with its increased use within healthcare systems with documented results showing various performance evaluation metrics such as prediction accuracy scores has heightened research interests for additional knowledge discovery on insights into patterns of change and important contributory factors. A key challenge to achieving better performance outcomes is the identification of critical constituent factors and its impact. In the healthcare industry, collected data in many instances are of higher dimensionality and unstructured in some respects making their use for effective machine learning predictions difficult resulting in poor predictive outcomes. In this regard, various selection techniques have been employed to help address dataset dimensionality challenges. This article examines the impact of three most widely used feature selection techniques on prediction accuracy score with healthcare context-based dataset in the prediction of treatment default for hypertensive patients with comorbidities. Using tree based classification models; extreme gradient boosting classifier, gradient boosting classifier and random forest classifier together with learning curves we demonstrate how model performance behavior could be explained. Results obtained show that random forest classifier generalizes well on cross validation with minimum training examples. Prediction accuracy scores achieved for the various models with feature selection techniques such as Boruta, principal component analysis and mutual information gain are; for Boruta, extreme gradient boosting classifier 97.93% and 98.23%, Gradient Boosting Classifier 98.03% and 98.13% and Random Forest 97.94% and 97.94% respectively. For principal component analysis; extreme gradient boosting 97.94%, gradient boosting 98.13% and random forest 97.47% and for mutual information gain; extreme gradient boosting 98.56%, gradient boosting 98.63% and random forest 98.31%. Further evaluation metrics to determine performance using receiver operating characteristic curve show the following scores; extreme gradientboosting classifier achieves a score of 90.90%, gradient boosting classifier 83.70% and random forest 78.40%. Additional performance evaluation with learning curves to determine model performance behavior show that even though random forest achieves the least roc_auc score, it generalizes well on unseen data, has the least root mean square error score and shows no need for additional features to achieve optimum performance. These scores as recorded show the impact of applying feature selection techniques on prediction accuracy scores and explains model performance behavior in the prediction of treatment default for patients suffering from hypertension with comorbidities.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00