Efficacy of Small Molecules Blocking in Kv1.5 Potassium Channel From Machine Learning Models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Efficacy of Small Molecules Blocking in Kv1.5 Potassium Channel From Machine Learning Models Samiya Kabir Youme, Hossain Ahamed, Anika Mehjabin Oishi, Md.Tawfiq UZ-Zaman, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3263007/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Atrial fibrillation and associated cardiac problems may be treated with the development of potent potassium ion channel Kv 1.5 blockers. Since the use of these blockers provides therapeutic advantages and potential side effects, it is significant to identify Kv 1.5 channel blockers from compounds. In this work, we employed optimized machine learning models to predict the potential of small molecules in blocking the Kv 1.5 channel to address the limitations of traditional screening methods in the drug discovery process. Several machine learning classifiers and regression models were employed utilizing molecular descriptors and fingerprints incorporating with SMOTE oversampling technique to overcome the class imbalance in active and inactive molecules. The results show that distinct models excelled in predicting different molecular attributes. The regression models demonstrated superior performance with random forest regression (RFR) (root-mean-square error = 0.668) and Substructure-Count-HGBR (Histogram-based Gradient Boosting Regression) having adjusted R² of 39.50% for predicting binding affinity. The best-performing models among the fingerprint-based models were the k-Nearest Neighbors Classifier (KNNC) and Substructure-RFC (Random Forest Classifier), which both demonstrated well-balanced predictive models. The generalized machine learning models for Kv 1.5 can help researchers quickly narrow down drug candidates that are toxic or beneficial for treating atrial fibrillation in the early stages of drug discovery. drug discovery machine learning RDkit K+ channel blockers Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 1 Introduction Ion channels are ion-selective macromolecular pore-forming proteins that allow the passage of ions down their electrochemical gradient, thus establishing and controlling the slight voltage gradient across the plasma membrane of cells [ 1 ]. With the help of the inorganic ions, cells can transmit signals across the cell membrane or the cell’s surface. Traces of these ions can be found in cell membranes and intracellular component membranes, containing cell organelles like mitochondria, endoplasmic reticulum (ER), and nucleus membranes. The potassium (K+), sodium (Na+), and calcium (Ca2+) channels are the most common voltage-dependent ion channels. Voltage-gated channels generally play a crucial role in determining the shape and duration of action potentials, including the delayed rectifiers and transient outward potassium channels [ 2 ]. However, the ion channel’s function can be influenced directly by drugs binding to the channel protein and modifying its activity, or indirectly via G proteins and other intermediates [ 3 ]. Certain antiarrhythmic drugs, antihistamines, and antibiotics block cardiac Kv channels, prolong action potential duration, resulting in long QT syndrome, and induce ventricular arrhythmia such as torsade de Pointes (TdP) [ 4 ]. Throughout this research study, our work focuses on the potassium channel Kv1.5 , a commonly distributed channel in Atria. Thereby Kv channels allow a variety of physiological activities that take place in a cell. It depends on several factors, such as the type and location, the influence of ions, phospholipids, and binding proteins [ 5 ]. Kv channels are classified into seven subfamilies, Kv1.x - Kv7. x. Here x represents the number of members in each subfamily depending on the mediator and ion conduction characteristics. Kv beta subunits and all Kv channels are comprised of six helical transmembrane proteins [ 6 ]. It’s possible to find several voltage-gated potassium channels in cardiac myocytes. Kv 1.5 is one of the voltage-gated potassium channels responsible for the ultra-rapid delayed-rectifier current. The KCNA5 (location:12p13.32) gene encodes K v 1.5 protein in humans, with a molecular weight of 67kDA [ 7 ]. Kv 1.5 channels are also expressed in many other organs, including the pulmonary arteries, the brain, and skeletal muscle, and play an essential role in regulating the cell cycle [ 8 ]. In membrane repolarization, potassium channels play an essential role after sodium. In some instances, calcium channels depolarize the membrane during the action potential. In addition, the incredibly high ionic selectivity, fast rate of flux, and intricate gating mechanisms of potassium channels are used to achieve the necessary balance for this interplay [ 9 ]. Ultra-rapid delayed rectifier potassium current in cardiac cells, particularly in the atria oversees the repolarization phase of the action potential in atrial myocytes. Kv 1.5 has become a desirable therapeutic target for familial atrial fibrillation (AF) type 7 due to its selective expression in the atria and limited expression in the ventricles [ 10 ]. Frequent heart arrhythmia, AF causes inefficient atrial contractions as a result of fast and erratic atrial electrical activity. Atrial action potential duration can be influenced, and AF can be prevented or treated by modulating atrial electrophysiology without affecting ventricular function [ 11 ]. However, the inhibition of the Kv 1.5 channel can have negative effects on the prolonged action potential duration, in other situations. Research conducted in recent decades shows that mutations can induce a wider variety of human diseases, including long QT syndrome, short QT syndrome, TdP, Brugada syndrome, familial AF, and several similar genetic cardiac channelopathies. These diseases are all caused by inherited mutations in pore-forming subunits and accessory subunits of cardiac Kv channels, which can alter the atrial and ventricular action potential [ 12 ]. Many investigations have discovered this channel has been abnormally found in various human tumor cells over the last decade [ 13 ]. Several drugs can inhibit Kv channels, such as the most common 4-Aminopyridine (4AP) and Class III antiarrhythmic drugs, and these drugs can worsen cardiovascular toxicity, resulting in TdP [ 14 ] [ 15 ]. In addition, using the 4AP drug (a small molecule) pose a risk of severe Central Nervous System (CNS) side effects [ 16 ]. Due to significant adverse drug reactions, the early stages of drug discovery become vital for safety implications; otherwise, those drugs are withdrawn from the marketplace in the worst-case scenario [ 17 ]. Thus, all new medications should be evaluated preclinically for their Kv channel-blocking properties to regulate safety during nonclinical and clinical testing before being submitted to regulatory assessments in order to screen out toxic compounds, according to the guideline released by the International Conference of Harmonization (ICH) [ 18 ]. Nevertheless, this involves labor-intensive, expensive, and time-consuming experiments, such as electrophysiological assessments, fluorescence-based assays, patch-clamp, voltage clamp, and radioligand binding assays for screening channel blockers in existing in vivo and in vitro methods [ 19 ]. Therefore, in this work, in silico models are developed, as recent advancements in silico techniques have made it possible to screen out substantially toxic compounds in the early stages of drug discovery. Thus, computational techniques, such as computer-aided drug design, have evolved into common tools to increase the effectiveness of the drug development process and reduce unfavorable effects [ 20 ]. Cai et al. [ 21 ] developed a combined classifier framework to predict drug-induced cardiovascular complications utilizing a neural network (NN) algorithm through integration with single classifiers from machine learning (ML) models. The combined classifiers performed better when compared to single classifiers, with an AUC ranging from 0.784 to 0.842 in 5 cross-validations. In 2021, Meng et al. [ 22 ] constructed five machine-learning models using the combination of features for the prediction of binding affinities of hERG channel blockers with 9215 compounds. Among all models, Support Vector Regression (SVR) yielded better results when all features were combined, with a 0.585 Root-Mean- Square Error (RMSE) on ten-fold cross-validation. Arab et al. proposed 2D descriptor predictive Quantitative Structure-Activity Relationship (QSAR) models incorporating regression and multiclass classification ML models to screen out hERG and Nav 1.5 channels [ 23 ]. Most researchers have employed several ML techniques to develop strong models that predict mostly hERG and Na v 1.5 inhibition [ 24 – 26 ], however, there are very few works on the K + channel despite having a crucial role in AF treatment. In this study, we implemented various ML models for predicting inhibition of the Kv 1.5 channel using molecular descriptors and fingerprints in an imbalanced dataset since the blockade of Kv 1.5 significantly impacts several cardiovascular complications. We have performed a thorough investigation using various regression models for predicting the binding affinity of small molecules and classification models for identifying blockers and non-blockers by using molecular descriptors according to the Lipinski rules of five and 12 molecular fingerprints. In this work, we explored a resampling method and statistical approaches to improve the performance of ML models. 2. Materials and Methods 2.1. Data Collection and Preparation The data used in this research was collected from a publicly available ChEMBL [ 27 ] and PubChem bioactivity database [ 28 ]. Here the patch-clamp measurements were considered for assessing the K V 1.5 channel-blocking potency. Data from the ChEMBL database by the “chembl_webresource_client” package was selected randomly for substances having bioactivities for a specific target for blockers in Python. Initially, the dataset consisted of 889 entries, it was found that 89 entries had either duplicates or anomalies rendering them unusable. Particular attention was taken to removing redundant data because data curation is the most crucial step in the workflow for constructing models, resulting in a dataset containing 800 entries. Then, IC50 values were converted to negative logarithmic values, pIC50 to make the distribution of the data more uniform. Small molecules with IC50 values less than 1 M were classified as Kv 1.5 blockers, while molecules with IC50 values greater than 10 M were considered non-blockers. The rest of the compounds were taken as intermediate, which were eventually removed, resulting in 682 data. It can be seen from Fig. 1 that datasets are imbalanced which can lead to biased models and perform poorly in predicting the minority class. To overcome the problem of imbalanced classes, Synthetic Minority Oversampling Technique (SMOTE) was employed as SMOTE generates synthetic data of the minority class allowing the desired balance between the classes [ 29 ]. 2.2. Molecular Descriptors and Fingerprints Quantitative structure-activity relationship (QSAR) modeling, virtual screening, and property prediction all rely heavily on molecular descriptors. They make it possible to compare, evaluate, and choose molecules based on their structural and physicochemical characteristics, which helps predict their biological activities, ADME (absorption, distribution, metabolism, and excretion) qualities, toxicity, and other traits [ 30 ] [ 31 ]. We have computed molecular descriptors based on Lipinski’s rule using the RDKit library, comprising molecular weight (MW), an octanol-water partition coefficient (LogP), hydrogen bond donors (NumHDonors), and hydrogen bond acceptors (NumHAcceptors). The correlation between molecular descriptors according to the Lipinski rule is represented in Fig. 2 , where it can be seen that LogP has the strongest influence on pIC 50 among all molecular descriptors. Further, we have investigated the relationship between MW and LogP, shown in Fig. 3 . It is observed that as the molecular weight of a compound increases, its log P value tends to rise as well, in both cases. As a matter of fact, larger molecules often have more hydrophobic regions or groups, which tend to favor partitioning into the organic phase (e.g., octanol). Additionally, it can be depicted that active molecules have higher MW and LogP values than the non-blockers small molecules since larger molecules with higher molecular weight have a greater potential for forming numerous interactions with the target. Subsequently, twelve molecular fingerprints were calculated using PaDEL Descriptor, including Atom Pairs 2D (780 digits), Atom Pairs 2D count (780 digits), CDK fingerprints (1024 digits), CDK-extended fingerprints (1024 digits), CDK graph only fingerprint (1024 digits), Estate fingerprints (79 digits), PubChem fingerprints (881 digits), Klekota-Roth fingerprint (4860 digits), Klekota-Roth fingerprint count (4860 count), MACCS fingerprints (166 digits), Substructure fingerprints (307 digits) and Substructure Count fingerprints (307 digits). 2.3. Feature Extraction and Building of Machine Learning Models As feature selection is a crucial step in building sophisticated ML methods, we have conducted Mann-Whitney U Test [ 32 ] and calculated the p-value. All the descriptors based on the Lipinski rule have significant differences, as all the p-values were below the predetermined significance level (0.05). While selecting the features for the molecular fingerprints, the variance threshold technique was applied. We have set the threshold as p(1 − p), where p is 0.8. After employing statistical approaches, we split training and test sets with a ratio of 80:20 for the entire dataset. Several classification and regression models were employed using various ML methods using Lazy Classifier and Lazy Regressor, respectively from the Lazy Predict library. Lazy Regressor and Lazy Classifier produced 42 regression and 42 classification models, respectively, among which several models exhibited poor performance. Therefore, we selected the best ten regression models, including random forest regression (RFR), several variations of gradient boosting regression (GBR), different variations of support vector regression (SVR), k-nearest neighbor (KNNR), bagging regression (BR), and different variations of decision tree regression (DT). The classifiers that were selected for further analysis were random forest classifier (RFC), variations of decision tree classifier (DTC), bagging classifier (BC), label propagation classifier (LPC), label spreading classifier (LSC), k-nearest neighbor (KNNC), and several variations of gradient boosting classifier (GBC). Initially, we utilized default values for hyper-parameters. Then, we performed a Grid search cross-validation (GSCV) technique for the top common models, random forest, variations of gradient boosting, k-nearest neighbor, variations of decision tree, and label propagation for both classifier and regressors to tune with hyperparameters. 2.4. Performance Evaluation Metric We applied the most widely used metrics for traditional machine learning models, such as accuracy (AC), sensitivity (SN), specificity (SP), balanced accuracy (BA), F1 score (F1), Matthew’s correlation coefficient (MCC) to evaluate the classification models and for regression, these metrics were R-Squared (R²), adjusted R², and root-mean-square error (RMSE). Since accuracy serves as a benchmark for comparing the performance of different classification models or algorithms and reflects the proportion of correct predictions made by the model, we have selected the classification models for both fingerprint and descriptor based on the highest accuracy in the test set. However, accuracy can be misleading when dealing with considerably uneven datasets, as a model predicts the majority class most of the time can achieve high accuracy while underperforming in the minority class. For further analysis, we have considered the F1 score, balanced accuracy, and MCC, and computed the area under the curve (AUC) to gain a more comprehensive understanding of the top models’ performance. In addition, we also selected the top models with lower RMSE values in the test set for regression models. Accuracy = \(\frac{TP+TN}{TP+TN+FP+FN}\) Sensitivity = \(\frac{TP}{TP+FN}\) Specificity = \(\frac{TN}{TN+FP}\) Balanced Accuracy = \(\frac{\text{Sensitivity}+\text{Specificity}}{2}=\frac{1}{2}\left(\frac{TP}{TP+FN}+\frac{TN}{TN+FP}\right)\) $$F1\text{ score}=2\times \frac{\text{precision}\times \text{recall}}{\text{precision}+\text{recall}}=2\times \frac{TP}{2\times TP+FP+FN}$$ $$MCC=\frac{TP\times TN-FP\times FN}{\sqrt{\left(TP+FP\right)\left(TP+FN\right)\left(TN+FP\right)\left(TN+FN\right)}}$$ where TP, TN, FP, and FN stand for the true positives, true negatives, false positives, and false negatives, respectively. The following metrics are computed for regression models: $${R}^{2}=1-\frac{{\sum }_{i=1}^{n}{\left({y}_{i}-\stackrel{\prime }{{y}_{i}}\right)}^{2}}{{\sum }_{i=1}^{n}{\left({y}_{i}-\overline{y}\right)}^{2}}$$ $$\text{Adjusted }{\text{R}}^{2}=1-\left(1-{R}^{2}\right)\times \frac{n-1}{n-k-1}$$ $$\text{RMSE}=\sqrt{\frac{1}{n}{\sum }_{i=1}^{n}{\left({y}_{i}-\stackrel{\prime }{{y}_{i}}\right)}^{2}}$$ where y is the actual value, \(\stackrel{\prime }{{y}_{i}}\) is the corresponding prediction, \(\overline{y}\) is the mean of the actual values in the set and N is the number of test data points. 3. Results and Discussion The purpose of this study is to experiment with various chemical representations and algorithmic strategies that may impact how well the ML models predict blockers and non-blockers of Kv 1.5. After conducting all experiments, we initially selected the top ten models from 42 regression and classification models as good performance models based on the lowest RMSE and highest accuracy, respectively obtained in the test set. Out of ten models, ML models were further analyzed, as RMSE and accuracy measures were insufficient to adequately assess a model based on the dataset, where the number of active compounds was nearly twice as great as the number of inactive compounds. The detailed results are thoroughly analyzed in this section. In quantitative structure-activity relationship (QSAR) studies, using regression models based on descriptors and fingerprints is a common approach to predict the potency of a molecule (pIC50). These models aim to forecast the binding affinity using the chemical representation of molecular descriptors and fingerprints, for understanding potential activity (blockers and non-blockers) based on the predicted pIC50 values. 3.1. Regression Models 3.1.1. Lipinski Descriptor-based Regression Models Table 1 Metrics of the training and test set for the top three regression models using descriptor-based according to the Lipinski rule Dataset Model R-Squared Adjusted R-Squared RMSE RFR 0.882 0.881 0.285 Train GBR 0.650 0.647 0.490 HGBR 0.752 0.750 0.413 RFR 0.344 0.327 0.668 Test GBR 0.314 0.297 0.683 HGBR 0.309 0.290 0.685 When compared to the RMSE and R², all the generated models using molecular descriptors in this study perform similarly, as shown by the results in Fig. 3 . Among the models evaluated in this study, three ensemble-based methods namely RFR, GBR, and Hist Gradient Boosting Regressor (HGBR), demonstrated high efficacy in terms of predictive accuracy and goodness of fit. It is observed that GBR and HGBR yielded the same R² of 0.31, however, GBR achieved an RMSE of 0.68 in the test set, which is slightly better than HGBR, indicating its ability to iteratively improve, refine predictions and minimize the prediction errors. Overall, the RFR demonstrated the best performance, outperforming the other models in terms of RMSE and R² metrics. This suggests that the RFR model succeeded in capturing the underlying patterns and relationships within the data, among all other models. Since the predictive outcomes of these models are closer to a difference of 0.01 for RMSE, adjusted R² is utilized as a metric to provide insights into how well the models fit the data and penalizes the addition of redundant predictors that do not contribute significantly to the model’s predictive power. Table 1 clearly shows that the RFR regression model outperformed all other regression models in predicting binding affinity, as evidenced by its higher adjusted R² value of 0.327. 3.1.2. Fingerprint-based Regression Models Table 2 Metrics of the training and test set for regression models using fingerprint Fingerprint Models R-Squared Adjusted R-Squared RMSE HGBR 0.592 0.582 0.530 Estate-Train LGBR 0.579 0.570 0.538 RFR 0.699 0.692 0.455 HGBR 0.410 0.353 0.633 Estate-Test LGBR 0.394 0.336 0.641 RFR 0.344 0.281 0.668 SVR 0.548 0.537 0.557 Substructure-Train HGBR 0.577 0.567 0.539 LGBR 0.577 0.567 0.539 SVR 0.361 0.294 0.659 Substructure-Test HGBR 0.360 0.294 0.659 LGBR 0.360 0.294 0.659 GBR 0.743 0.731 0.420 SubstructCount-Train LGBR 0.851 0.844 0.319 HGBR 0.851 0.844 0.319 GBR 0.479 0.363 0.595 SubstructCount-Test LGBR 0.505 0.395 0.580 HGBR 0.505 0.395 0.580 Out of all the fingerprint-based regression methods, the models shown in Table 2 performed consistently overall metrics in the training and testing set. Notably, variants of the GBR and RFR performed better than every other model tested using fingerprint. Although RMSE values were comparatively closer to these nine models in the test set, they showed considerable differences in adjusted R². It can be seen that Estate-RFR performs poorly than all models with the highest RMSE of 0.668, and the lowest adjusted R² of 0.281. When incorporating the Substructure fingerprint, all of the models perform similarly in RMSE and adjusted R² with a value of 0.659, and 0.294, respectively, however, the Subsructure-SVR model’s R² value slightly differs. Based on the fact that SubstructureCount has the lowest RMSE of any model, it is a promising model in terms of RMSE. SubtructureCount-LGBR and SubtructureCount-HGBR perform equally well as all other models with similar results in RMSE, R², and adjusted R². Further, we have considered unrounded values of adjusted R² for models SubtructureCount-LGBR and SubtructureCount-HGBR, yielding values of 0.3946130914, and 0.3949747272 in the test set, respectively, indicating these values each account for around 39.46% and 39.50% of the variation in the binding affinity, respectively. Nevertheless, after examining the results for all the regressors with different fingerprints, it is obvious that the SubtructureCount-HGBR is the best regressor in terms of its RMSE, R², and adjusted R². 3.2. Classification Models SMOTE method is used throughout the section to balance active and inactive molecules in each experiment. As mentioned earlier, the performance of the classification models was evaluated using the accuracy for both descriptors and fingerprints, initially in correctly classifying instances to blockers and non-blockers of the Kv1.5 channel. Subsequently, we assessed and compared the best models using various metrics. 3.2.1. Lipinski Descriptor-based Classification Models Table 3 Metrics of the training and test set for classification models using descriptors to classify blockers and non-blockers. DatasetModels SN SP BA F1 MCC AUC XGBC 0.835 0.850 0.994 0.994 0.685 0.994 TrainDTC 0.873 0.880 0.994 0.994 0.753 0.994 KNNC 0.808 0.893 0.873 0.873 0.703 0.873 LPC 0.808 0.893 0.994 0.994 0.703 0.873 XGBC 0.871 0.639 0.677 0.765 0.510 0.812 TestDTC 0.851 0.500 0.657 0.741 0.362 0.723 KNNC 0.762 0.528 0.748 0.781 0.275 0.739 LPC 0.762 0.528 0.657 0.741 0.275 0.739 After conducting experiments with classification models using molecular descriptors, we obtained the top ten models with the highest AC values, as shown in Fig. 5 . It is observed that the Extreme Gradient Boosting classifier (XGBC), DTC, LPC, and RFC demonstrated high accuracy rates, where XGBC and DTC performed the same with greater AC in both train and test sets. While the KNNC model had relatively lower accuracy in classification, it depicted a significant characteristic regarding overfitting. In comparison to all the other models, it was discovered that the KNNC model’s difference between the train and test accuracies was substantially lower. This suggests that the KNNC model is less prone to overfitting since it generalizes better to unseen data. In addition, RFC and LPC also performed equally in test sets, albeit LPC generalizes better because the difference between the train and test sets is smaller than that of RFC. Further investigation of these models is depicted in Table 3 , where we have taken four models’ performances in train and test sets into consideration, namely, XGBC, DTC, LPC, and KNNC to validate the best descriptor classification model. It can be observed that KNNC and LPC both performed equally well in all metrics, with the exception that KNNC outperformed LPC in terms of BA values. Table 4 Metrics of the training and test set for classification models using fingerprints to classify blockers and non-blockers. The XGBC model demonstrated superior performance compared to the other models evaluated, having an MCC of 0.510 on the testing set, however, the KNNC model is a generalized model, indicating a strong ability to predict and classify the target variable. Additionally, the KNNC model appears to have a higher F1 score implies that the model has a good balance between precision and recall discriminating between active and inactive instances, although other measures values are significantly closer to the XGBC model, including the AUC value of 0.739, SN of 0.762, and SP of 0.528 in the test set. In summary, the KNNC model can generalize well than all the other models, including XGBC, DTC, and LPC, in terms of BA, F1, and accuracies, making it the most suitable model using molecular descriptors. 3.2.2. Fingerprint-based Classification Models The findings of fingerprint-based classification models with accuracy levels of at least 80% in the test set are shown in Table 4 . From the results presented in Table 4 , it can be noted that all the developed models in this study perform similarly regarding accuracy with minimal differences between these models, therefore, different metrics are considered for interpreting the predictive results. In particular, SN has remarkably greater values in the test set in maximum models, suggesting that the model has a higher ability to correctly recognize true positive cases. Conversely, as the number of positive instances is much higher than the number of negative instances, specificity has comparatively lower values in all models, indicating that the number of correctly identified negative classes is comparatively lower. AtomsPair2DCount-DTC obtained the highest accuracy in the test set, albeit other measures, including BA of 0.664, F1 of 0.723, and MCC of 0.376, failed to demonstrate adequate potential. While AtomsPair2D-RFC is the second-best model with just a 0.73% difference in accuracy, with the highest MCC of 0.515, BA of 0.747, F1 of 0.819, SN of 0.931, and SP of 0.528. However, when comparing train accuracies, the model has a higher training accuracy of 93.21%, it exhibits a larger drop in performance on the test set with a difference of 10.73%, leading to a potential issue with overfitting. Therefore, Substructure-RFC is a better suitable model with a training accuracy of 85.14% and test accuracy of 82.48%, and significant performance in other measures as well, since it presents a more balanced performance and reflects how well the model generalizes to unseen data. In predictive modeling, generalization is a crucial factor, especially in drug discovery. Generalized models are frequently utilized in the preliminary stages of drug development when researchers are searching a broad chemical space for new therapeutic compounds. Mainly, in this study, we have prioritized generalized models for Kv 1.5 that can help researchers quickly narrow down drug candidates that are toxic or beneficial for treating atrial fibrillation by saving time and resources. These findings contribute to the advancement of predictive models for binding affinity, facilitating more reliable predictions in drug discovery and development. Further studies are required to explore the robustness and scalability of these models, considering larger datasets and examining their effectiveness on diverse chemical compounds. 4. Conclusions In this work, we have experimented with and analyzed the performance of classification and regression models, trained on both descriptors and fingerprints along with oversampling technique to effectively predict inhibitors of Kv 1.5. We found that different models perform best in different molecular attributes, including RFR and KNNC for descriptor-based models, and SubtructureCount-HGBR, and Substructure-RFC for fingerprint-based models. The classification-based model seems capable of producing well-balanced predictive models. Despite having crucial effects on the prolonged action potential duration due to inhibition of Kv 1.5, studies specifically focused on this channel have been less common than those on the hERG channel and sodium channel. Thus, we have dealt with fewer datasets; nonetheless, we have achieved promising results in correctly identifying true positive values. Declarations Author Contributions: “Conceptualization, M.H.R, methodology, validation and formal analysis H.A, S.K.Y, A.M.O, M.T.Z; resources, M.H.R.; original draft preparation S.K.Y, H.A, A.M.O, R.A.R, M.H.R.; writing, review and editing S.K.Y, H.A, A.M.O, R.A.R, K.S.H, M.S.I and M.H.R. project administration M.H.R, supervision and funding acquisition M.S.I and M.H.R. Finally, the manuscript prepares from the input of all authors. All authors have read and agreed to the published version of the manuscript. Acknowledgment : M.H.R. acknowledges the NSU CTRG Research Grant CTRG-20/SEPS/06 and CTRG-21/SEPS/21. M. S. I acknowledge NSU CTRG Research Grant CTRG-21/SEPS/11. Competing Interests: The author declares no conflict of interest, be it financial or non-financial. References Tristani-Firouzi, M.; Chen, J.; Mitcheson, J.S.; Sanguinetti, M.C. Molecular biology of K+ channels and their role in cardiac arrhythmias. The American journal of medicine 2001, 110, 50–59. Litalien, C.; Beaulieu, P. Molecular Aspects of Drug Actions: From Receptors to Effectors. In Pediatric Critical Care; Elsevier, 2006; pp. 1659–1677. Roden, D. Mechanism and management of proarrhythmia. Am. J. Cardiol 1998, 82, 491–571. Kuang, Q.; Purhonen, P.; Hebert, H. Structure of potassium channels. Cellular and molecular life sciences 2015, 72, 3677–3693. Campomanes, C.R.; Carroll, K.I.; Manganas, L.N.; Hershberger, M.E.; Gong, B.; Antonucci, D.E.; Rhodes, K.J.; Trimmer, J.S. Kvβ subunit oxidoreductase activity and Kv1 potassium channel trafficking. Journal of Biological Chemistry 2002, 277, 8298–8305. Sansom, M.S. Ion channels: a first view of K+ channels in atomic glory. Current biology 1998, 8, R450–R452. Wettwer, E.; Terlau, H. Pharmacology of voltage-gated potassium channel Kv1. 5—impact on cardiac excitability. Current Opinion in Pharmacology 2014, 15, 115–121. Kim, D.M.; Nimigean, C.M. Voltage-gated potassium channels: a structural examination of selectivity and gating. Cold Spring Harbor perspectives in biology 2016, 8, a029231. Giudicessi, J.R.; Ackerman, M.J. Potassium-channel mutations and cardiac arrhythmias—diagnosis and therapy. Nature Reviews Cardiology 2012, 9, 319–332. Tamargo, J.; Caballero, R.; Gómez, R.; Delpón, E. IKur/Kv1. 5 channel blockers for the treatment of atrial fibrillation. Expert opinion on investigational drugs 2009, 18, 399–416. Pellman, J.; Sheikh, F. Atrial fibrillation: mechanisms, therapeutics, and future directions. Comprehensive Physiology 2015, 5, 649. Comes, N.; Bielanska, J.; Vallejo-Gracia, A.; Serrano-Albarrás, A.; Marruecos, L.; Gómez, D.; Soler, C.; Condom, E.; Ramón y Cajal, S.; Hernández-Losa, J.; et al. The voltage-dependent K+ channels Kv1. 3 and Kv1. 5 in human cancer. Frontiers in physiology 2013, 4, 283. Clarfield, A.M. Brocklehurst’s Textbook of Geriatric Medicine and Gerontology. JAMA 2010, 304, 1956–1957. Wang, Z.; Fermini, B.; Nattel, S. Effects of flecainide, quinidine, and 4-aminopyridine on transient outward and ultrarapid delayed rectifier currents in human atrial myocytes. Journal of Pharmacology and Experimental Therapeutics 1995, 272, 184–196. Brendorp, B.; Pedersen, O.D.; Torp-Pedersen, C.; Sahebzadah, N.; Køber, L. A benefit-risk assessment of class III antiarrhythmic agents. Drug safety 2002, 25, 847–865. King, A.M.; Menke, N.B.; Katz, K.D.; Pizon, A.F. 4-aminopyridine toxicity: a case report and review of the literature. Journal of Medical Toxicology 2012, 8, 314–321. Valentin, J.P.; Guth, B.; Hamlin, R.L.; Lainée, P.; Sarazan, D.; Skinner, M. Functional cardiac safety evaluation of novel therapeutics. Antitargets and Drug Safety 2015, pp. 199–234. Sager, P.T.; Nebout, T.; Darpo, B. ICH E14: a new regulatory guidance on the clinical evaluation of QT/QTc internal prolongation and proarrhythmic potential for non-antiarrhythmic drugs. Drug information journal: DIJ/Drug Information Association 2005, 39, 387–394. Priest, B.; Bell, I.M.; Garcia, M. Role of hERG potassium channel assays in drug development. Channels 2008, 2, 87–93. Kong, W.; Tu, X.; Huang, W.; Yang, Y.; Xie, Z.; Huang, Z. Prediction and optimization of NaV1. 7 sodium channel inhibitors based on machine learning and simulated annealing. Journal of Chemical Information and Modeling 2020, 60, 2739–2753. Cai, C.; Fang, J.; Guo, P.; Wang, Q.; Hong, H.; Moslehi, J.; Cheng, F. In silico pharmacoepidemiologic evaluation of drug-induced cardiovascular complications using combined classifiers. Journal of chemical information and modeling 2018, 58, 943–956. Meng, J.; Zhang, L.; Wang, L.; Li, S.; Xie, D.; Zhang, Y.; Liu, H. TSSF-hERG: A machine-learning-based hERG potassium channel-specific scoring function for chemical cardiotoxicity prediction. Toxicology 2021, 464, 153018. Arab, I.; Barakat, K. ToxTree: descriptor-based machine learning models for both hERG and Nav1. 5 cardiotoxicity liability predictions. arXiv preprint arXiv:2112.13467 2021. Khalifa, N.; Kumar Konda, L.S.; Kristam, R. Machine learning-based QSAR models to predict sodium ion channel (Nav 1.5) blockers. Future Medicinal Chemistry 2020, 12, 1829–1843. Cai, C.; Guo, P.; Zhou, Y.; Zhou, J.; Wang, Q.; Zhang, F.; Fang, J.; Cheng, F. Deep learning-based prediction of drug-i nduced cardiotoxicity. Journal of chemical information and modeling 2019, 59, 1073–1084. Ryu, J.Y.; Lee, M.Y.; Lee, J.H.; Lee, B.H.; Oh, K.S. DeepHIT: a deep learning framework for prediction of hERG-induced cardiotoxicity. Bioinformatics 2020, 36, 3049–3055. Gaulton, A.; Hersey, A.; Nowotka, M.; Bento, A.P.; Chambers, J.; Mendez, D.; Mutowo, P.; Atkinson, F.; Bellis, L.J.; Cibrián-Uhalte, E.; et al. The ChEMBL database in 2017. Nucleic acids research 2017, 45, D945–D954. Kim, S.; Thiessen, P.A.; Bolton, E.E.; Chen, J.; Fu, G.; Gindulyte, A.; Han, L.; He, J.; He, S.; Shoemaker, B.A.; et al. PubChem substance and compound databases. Nucleic acids research 2016, 44, D1202–D1213. Azlim Khan, A.K.; Ahamed Hassain Malim, N.H. Comparative Studies on Resampling Techniques in Machine Learning and Deep Learning Models for Drug-Target Interaction Prediction. Molecules 2023, 28, 1663. Bastikar, V.; Bastikar, A.; Gupta, P. Quantitative structure–activity relationship-based computational approaches. In Computational Approaches for Novel Therapeutic and Diagnostic Designing to Mitigate SARS-CoV2 Infection; Elsevier, 2022; pp. 191–205. Lagorce, D.; Douguet, D.; Miteva, M.A.; Villoutreix, B.O. Computational analysis of calculated physicochemical and ADMET properties of protein-protein interaction inhibitors. Scientific reports 2017, 7, 46277. McKnight, P.E.; Najab, J. Mann-Whitney U Test. The Corsini encyclopedia of psychology 2010, pp. 1–1. Table 4 Table 4 is available in Supplementary Files section. Additional Declarations No competing interests reported. Supplementary Files GraphicalAbstract.png Table4.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3263007","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":226582328,"identity":"98d20d4f-7623-41b6-af19-760c756f155a","order_by":0,"name":"Samiya Kabir Youme","email":"","orcid":"","institution":"North South University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Samiya","middleName":"Kabir","lastName":"Youme","suffix":""},{"id":226582329,"identity":"77acc0a9-16c6-470e-9a23-b42aa9356581","order_by":1,"name":"Hossain Ahamed","email":"","orcid":"","institution":"North South University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hossain","middleName":"","lastName":"Ahamed","suffix":""},{"id":226582330,"identity":"bc727f82-d1c0-45cb-80b9-1b363691bccb","order_by":2,"name":"Anika Mehjabin Oishi","email":"","orcid":"","institution":"North South University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Anika","middleName":"Mehjabin","lastName":"Oishi","suffix":""},{"id":226582331,"identity":"70bc7eff-7620-4407-a92c-f6e32ed584a3","order_by":3,"name":"Md.Tawfiq UZ-Zaman","email":"","orcid":"","institution":"","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Md.Tawfiq","middleName":"","lastName":"UZ-Zaman","suffix":""},{"id":226582332,"identity":"3b65dcb6-e176-4e07-88d1-35b5ef6653bb","order_by":4,"name":"Ramisha Anan Rahman","email":"","orcid":"","institution":"North South University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ramisha","middleName":"Anan","lastName":"Rahman","suffix":""},{"id":226582333,"identity":"8a1a5fc0-b75d-45d1-8ff3-1b08e1fcc664","order_by":5,"name":"Kazi Sumaiya Hoque","email":"","orcid":"","institution":"North South University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Kazi","middleName":"Sumaiya","lastName":"Hoque","suffix":""},{"id":226582334,"identity":"b3069a46-67f5-4760-a199-6df310342fe2","order_by":6,"name":"Md Shariful Islam Islam","email":"","orcid":"","institution":"North South University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Md","middleName":"Shariful Islam","lastName":"Islam","suffix":""},{"id":226582335,"identity":"70a08a1e-d3be-420c-a4f0-cbeed5b9a0ac","order_by":7,"name":"Md Harunur Rashid","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABJUlEQVRIie3QsUrEMBjA8a8U4pK7rik9+gyRQotcwVfJUWiXLiIUXIogxOXc61s4OrYE6tKra+FEO9XFwcPlFsFr7wSx1XN0yH8I4SM/EgIgk/3D6Jc9AoYI1tqtct6uuBvuI+5Eb09n7YL+QgD5Lk33EOdg0Rgn/BGcy7umriOBreVF/ra6jU36dJXCayS+k6N5YBvX/BQmRehQVgpsP+RekhXCovmYKUnZIzT1kTHiDAiEiMz4hlShBRlPZzc5puqI98l9syPac9MRK+lIvCXvA6T6vIUwe0N8TElH1C1RhkijTnHJMCEvNmGli0nle7DgwtLzkGbzMug/zFeWOGIm0YJGX0fkWEs8AWc8NseiOKzX0bT3y7vw8Dj96bxMJpPJfu0DKGJsYRomWh4AAAAASUVORK5CYII=","orcid":"","institution":"North South University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Md","middleName":"Harunur","lastName":"Rashid","suffix":""}],"badges":[],"createdAt":"2023-08-14 14:14:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3263007/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3263007/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":41870397,"identity":"cefe0ed2-0b0c-4741-abe8-57209f63a151","added_by":"auto","created_at":"2023-08-21 13:29:47","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":8712,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of the whole dataset. The length of the bars is the frequency of occurrence.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-3263007/v1/1df5d6fda70a0f29207c247f.png"},{"id":41872099,"identity":"6cb0a510-22a6-4776-bf15-89a9debb134f","added_by":"auto","created_at":"2023-08-21 13:37:47","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":28384,"visible":true,"origin":"","legend":"\u003cp\u003eHeat Maps of Pearson Correlation Coefficient representation of molecular descriptors.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-3263007/v1/770137e4dba5024e55578fbd.png"},{"id":41870402,"identity":"c7150543-95c3-4e5c-ac8d-f6fac6424a5f","added_by":"auto","created_at":"2023-08-21 13:29:47","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":164236,"visible":true,"origin":"","legend":"\u003cp\u003eRelationship of molecular weight (MW) and octanol-water partition coefficient (LogP) for blockers and non-blockers.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-3263007/v1/0a8fcbd66b02cffa46449f42.png"},{"id":41873856,"identity":"b39789d8-82c1-4c7e-89c4-4f76e315602a","added_by":"auto","created_at":"2023-08-21 13:45:47","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":47549,"visible":true,"origin":"","legend":"\u003cp\u003eRMSE and R² scores for the binding affinity of top ten regression models.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-3263007/v1/3806d6147d20125f8c4325ab.png"},{"id":41870403,"identity":"b37c0f76-0d91-41db-9c53-4db4c352b2b2","added_by":"auto","created_at":"2023-08-21 13:29:47","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":67758,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance of descriptor-based classification models.\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-3263007/v1/c726e2e62a6c13ac49d416ec.png"},{"id":43580767,"identity":"a3f3e4de-a505-4214-babc-40231313d723","added_by":"auto","created_at":"2023-09-24 03:07:21","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":706983,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3263007/v1/1bc5c319-b6e4-4213-83c9-4ac5b5910a80.pdf"},{"id":41872101,"identity":"d1f32c4b-62b8-4723-b425-204faef88a57","added_by":"auto","created_at":"2023-08-21 13:37:47","extension":"png","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":88095,"visible":true,"origin":"","legend":"","description":"","filename":"GraphicalAbstract.png","url":"https://assets-eu.researchsquare.com/files/rs-3263007/v1/2021e46fc00788963b536ec5.png"},{"id":41870400,"identity":"b392b3d4-1b62-4ccc-a7b4-e9a29b139b23","added_by":"auto","created_at":"2023-08-21 13:29:47","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":31612,"visible":true,"origin":"","legend":"","description":"","filename":"Table4.docx","url":"https://assets-eu.researchsquare.com/files/rs-3263007/v1/b689881e546617cb1f55c6de.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Efficacy of Small Molecules Blocking in Kv1.5 Potassium Channel From Machine Learning Models","fulltext":[{"header":"1 Introduction","content":"\u003cp\u003eIon channels are ion-selective macromolecular pore-forming proteins that allow the passage of ions down their electrochemical gradient, thus establishing and controlling the slight voltage gradient across the plasma membrane of cells [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. With the help of the inorganic ions, cells can transmit signals across the cell membrane or the cell\u0026rsquo;s surface. Traces of these ions can be found in cell membranes and intracellular component membranes, containing cell organelles like mitochondria, endoplasmic reticulum (ER), and nucleus membranes. The potassium (K+), sodium (Na+), and calcium (Ca2+) channels are the most common voltage-dependent ion channels.\u003c/p\u003e \u003cp\u003eVoltage-gated channels generally play a crucial role in determining the shape and duration of action potentials, including the delayed rectifiers and transient outward potassium channels [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. However, the ion channel\u0026rsquo;s function can be influenced directly by drugs binding to the channel protein and modifying its activity, or indirectly via G proteins and other intermediates [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Certain antiarrhythmic drugs, antihistamines, and antibiotics block cardiac \u003cem\u003eKv\u003c/em\u003echannels, prolong action potential duration, resulting in long QT syndrome, and induce ventricular arrhythmia such as torsade de Pointes (TdP) [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Throughout this research study, our work focuses on the potassium channel \u003cem\u003eKv1.5\u003c/em\u003e, a commonly distributed channel in Atria. Thereby \u003cem\u003eKv\u003c/em\u003echannels allow a variety of physiological activities that take place in a cell. It depends on several factors, such as the type and location, the influence of ions, phospholipids, and binding proteins [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. \u003cem\u003eKv\u003c/em\u003echannels are classified into seven subfamilies, Kv1.x - Kv7. x. Here x represents the number of members in each subfamily depending on the mediator and ion conduction characteristics. \u003cem\u003eKv\u003c/em\u003ebeta subunits and all \u003cem\u003eKv\u003c/em\u003echannels are comprised of six helical transmembrane proteins [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. It\u0026rsquo;s possible to find several voltage-gated potassium channels in cardiac myocytes. \u003cem\u003eKv\u003c/em\u003e1.5 is one of the voltage-gated potassium channels responsible for the ultra-rapid delayed-rectifier current. The KCNA5 (location:12p13.32) gene encodes \u003cem\u003eK\u003c/em\u003e\u003csub\u003e\u003cem\u003ev\u003c/em\u003e\u003c/sub\u003e1.5 protein in humans, with a molecular weight of 67kDA [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. \u003cem\u003eKv\u003c/em\u003e1.5 channels are also expressed in many other organs, including the pulmonary arteries, the brain, and skeletal muscle, and play an essential role in regulating the cell cycle [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn membrane repolarization, potassium channels play an essential role after sodium. In some instances, calcium channels depolarize the membrane during the action potential. In addition, the incredibly high ionic selectivity, fast rate of flux, and intricate gating mechanisms of potassium channels are used to achieve the necessary balance for this interplay [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Ultra-rapid delayed rectifier potassium current in cardiac cells, particularly in the atria oversees the repolarization phase of the action potential in atrial myocytes. \u003cem\u003eKv\u003c/em\u003e1.5 has become a desirable therapeutic target for familial atrial fibrillation (AF) type 7 due to its selective expression in the atria and limited expression in the ventricles [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Frequent heart arrhythmia, AF causes inefficient atrial contractions as a result of fast and erratic atrial electrical activity. Atrial action potential duration can be influenced, and AF can be prevented or treated by modulating atrial electrophysiology without affecting ventricular function [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eHowever, the inhibition of the \u003cem\u003eKv\u003c/em\u003e1.5 channel can have negative effects on the prolonged action potential duration, in other situations. Research conducted in recent decades shows that mutations can induce a wider variety of human diseases, including long QT syndrome, short QT syndrome, TdP, Brugada syndrome, familial AF, and several similar genetic cardiac channelopathies. These diseases are all caused by inherited mutations in pore-forming subunits and accessory subunits of cardiac \u003cem\u003eKv\u003c/em\u003echannels, which can alter the atrial and ventricular action potential [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Many investigations have discovered this channel has been abnormally found in various human tumor cells over the last decade [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Several drugs can inhibit \u003cem\u003eKv\u003c/em\u003echannels, such as the most common 4-Aminopyridine (4AP) and Class III antiarrhythmic drugs, and these drugs can worsen cardiovascular toxicity, resulting in TdP [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. In addition, using the 4AP drug (a small molecule) pose a risk of severe Central Nervous System (CNS) side effects [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eDue to significant adverse drug reactions, the early stages of drug discovery become vital for safety implications; otherwise, those drugs are withdrawn from the marketplace in the worst-case scenario [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. Thus, all new medications should be evaluated preclinically for their \u003cem\u003eKv\u003c/em\u003echannel-blocking properties to regulate safety during nonclinical and clinical testing before being submitted to regulatory assessments in order to screen out toxic compounds, according to the guideline released by the International Conference of Harmonization (ICH) [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Nevertheless, this involves labor-intensive, expensive, and time-consuming experiments, such as electrophysiological assessments, fluorescence-based assays, patch-clamp, voltage clamp, and radioligand binding assays for screening channel blockers in existing in vivo and in vitro methods [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Therefore, in this work, in silico models are developed, as recent advancements in silico techniques have made it possible to screen out substantially toxic compounds in the early stages of drug discovery. Thus, computational techniques, such as computer-aided drug design, have evolved into common tools to increase the effectiveness of the drug development process and reduce unfavorable effects [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eCai et al. [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] developed a combined classifier framework to predict drug-induced cardiovascular complications utilizing a neural network (NN) algorithm through integration with single classifiers from machine learning (ML) models. The combined classifiers performed better when compared to single classifiers, with an AUC ranging from 0.784 to 0.842 in 5 cross-validations. In 2021, Meng et al. [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e] constructed five machine-learning models using the combination of features for the prediction of binding affinities of hERG channel blockers with 9215 compounds. Among all models, Support Vector Regression (SVR) yielded better results when all features were combined, with a 0.585 Root-Mean- Square Error (RMSE) on ten-fold cross-validation. Arab et al. proposed 2D descriptor predictive Quantitative Structure-Activity Relationship (QSAR) models incorporating regression and multiclass classification ML models to screen out hERG and Nav 1.5 channels [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. Most researchers have employed several ML techniques to develop strong models that predict mostly hERG and \u003cem\u003eNa\u003c/em\u003e\u003csub\u003e\u003cem\u003ev\u003c/em\u003e\u003c/sub\u003e 1.5 inhibition [\u003cspan additionalcitationids=\"CR25\" citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e], however, there are very few works on the K\u0026thinsp;+\u0026thinsp;channel despite having a crucial role in AF treatment. In this study, we implemented various ML models for predicting inhibition of the \u003cem\u003eKv\u003c/em\u003e1.5 channel using molecular descriptors and fingerprints in an imbalanced dataset since the blockade of \u003cem\u003eKv\u003c/em\u003e1.5 significantly impacts several cardiovascular complications. We have performed a thorough investigation using various regression models for predicting the binding affinity of small molecules and classification models for identifying blockers and non-blockers by using molecular descriptors according to the Lipinski rules of five and 12 molecular fingerprints. In this work, we explored a resampling method and statistical approaches to improve the performance of ML models.\u003c/p\u003e"},{"header":"2. Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. Data Collection and Preparation\u003c/h2\u003e \u003cp\u003eThe data used in this research was collected from a publicly available ChEMBL [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e] and PubChem bioactivity database [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. Here the patch-clamp measurements were considered for assessing the \u003cem\u003eK\u003c/em\u003e\u003csub\u003e\u003cem\u003eV\u003c/em\u003e\u003c/sub\u003e1.5 channel-blocking potency. Data from the ChEMBL database by the \u0026ldquo;chembl_webresource_client\u0026rdquo; package was selected randomly for substances having bioactivities for a specific target for blockers in Python. Initially, the dataset consisted of 889 entries, it was found that 89 entries had either duplicates or anomalies rendering them unusable. Particular attention was taken to removing redundant data because data curation is the most crucial step in the workflow for constructing models, resulting in a dataset containing 800 entries. Then, IC50 values were converted to negative logarithmic values, pIC50 to make the distribution of the data more uniform. Small molecules with IC50 values less than 1 M were classified as \u003cem\u003eKv\u003c/em\u003e1.5 blockers, while molecules with IC50 values greater than 10 M were considered non-blockers. The rest of the compounds were taken as intermediate, which were eventually removed, resulting in 682 data.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIt can be seen from Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e that datasets are imbalanced which can lead to biased models and perform poorly in predicting the minority class. To overcome the problem of imbalanced classes, Synthetic Minority Oversampling Technique (SMOTE) was employed as SMOTE generates synthetic data of the minority class allowing the desired balance between the classes [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2. Molecular Descriptors and Fingerprints\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eQuantitative structure-activity relationship (QSAR) modeling, virtual screening, and property prediction all rely heavily on molecular descriptors. They make it possible to compare, evaluate, and choose molecules based on their structural and physicochemical characteristics, which helps predict their biological activities, ADME (absorption, distribution, metabolism, and excretion) qualities, toxicity, and other traits [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e] [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. We have computed molecular descriptors based on Lipinski\u0026rsquo;s rule using the RDKit library, comprising molecular weight (MW), an octanol-water partition coefficient (LogP), hydrogen bond donors (NumHDonors), and hydrogen bond acceptors (NumHAcceptors).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe correlation between molecular descriptors according to the Lipinski rule is represented in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, where it can be seen that LogP has the strongest influence on pIC\u003csub\u003e50\u003c/sub\u003e among all molecular descriptors. Further, we have investigated the relationship between MW and LogP, shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. It is observed that as the molecular weight of a compound increases, its log P value tends to rise as well, in both cases. As a matter of fact, larger molecules often have more hydrophobic regions or groups, which tend to favor partitioning into the organic phase (e.g., octanol). Additionally, it can be depicted that active molecules have higher MW and LogP values than the non-blockers small molecules since larger molecules with higher molecular weight have a greater potential for forming numerous interactions with the target.\u003c/p\u003e \u003cp\u003eSubsequently, twelve molecular fingerprints were calculated using PaDEL Descriptor, including Atom Pairs 2D (780 digits), Atom Pairs 2D count (780 digits), CDK fingerprints (1024 digits), CDK-extended fingerprints (1024 digits), CDK graph only fingerprint (1024 digits), Estate fingerprints (79 digits), PubChem fingerprints (881 digits), Klekota-Roth fingerprint (4860 digits), Klekota-Roth fingerprint count (4860 count), MACCS fingerprints (166 digits), Substructure fingerprints (307 digits) and Substructure Count fingerprints (307 digits).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3. Feature Extraction and Building of Machine Learning Models\u003c/h2\u003e \u003cp\u003eAs feature selection is a crucial step in building sophisticated ML methods, we have conducted Mann-Whitney U Test [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e] and calculated the p-value. All the descriptors based on the Lipinski rule have significant differences, as all the p-values were below the predetermined significance level (0.05). While selecting the features for the molecular fingerprints, the variance threshold technique was applied. We have set the threshold as p(1\u0026thinsp;\u0026minus;\u0026thinsp;p), where p is 0.8.\u003c/p\u003e \u003cp\u003eAfter employing statistical approaches, we split training and test sets with a ratio of 80:20 for the entire dataset. Several classification and regression models were employed using various ML methods using Lazy Classifier and Lazy Regressor, respectively from the Lazy Predict library. Lazy Regressor and Lazy Classifier produced 42 regression and 42 classification models, respectively, among which several models exhibited poor performance. Therefore, we selected the best ten regression models, including random forest regression (RFR), several variations of gradient boosting regression (GBR), different variations of support vector regression (SVR), k-nearest neighbor (KNNR), bagging regression (BR), and different variations of decision tree regression (DT). The classifiers that were selected for further analysis were random forest classifier (RFC), variations of decision tree classifier (DTC), bagging classifier (BC), label propagation classifier (LPC), label spreading classifier (LSC), k-nearest neighbor (KNNC), and several variations of gradient boosting classifier (GBC). Initially, we utilized default values for hyper-parameters. Then, we performed a Grid search cross-validation (GSCV) technique for the top common models, random forest, variations of gradient boosting, k-nearest neighbor, variations of decision tree, and label propagation for both classifier and regressors to tune with hyperparameters.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4. Performance Evaluation Metric\u003c/h2\u003e \u003cp\u003eWe applied the most widely used metrics for traditional machine learning models, such as accuracy (AC), sensitivity (SN), specificity (SP), balanced accuracy (BA), F1 score (F1), Matthew\u0026rsquo;s correlation coefficient (MCC) to evaluate the classification models and for regression, these metrics were R-Squared (R\u0026sup2;), adjusted R\u0026sup2;, and root-mean-square error (RMSE). Since accuracy serves as a benchmark for comparing the performance of different classification models or algorithms and reflects the proportion of correct predictions made by the model, we have selected the classification models for both fingerprint and descriptor based on the highest accuracy in the test set. However, accuracy can be misleading when dealing with considerably uneven datasets, as a model predicts the majority class most of the time can achieve high accuracy while underperforming in the minority class. For further analysis, we have considered the F1 score, balanced accuracy, and MCC, and computed the area under the curve (AUC) to gain a more comprehensive understanding of the top models\u0026rsquo; performance. In addition, we also selected the top models with lower RMSE values in the test set for regression models.\u003c/p\u003e \u003cp\u003eAccuracy =\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\frac{TP+TN}{TP+TN+FP+FN}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003cp\u003eSensitivity =\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\frac{TP}{TP+FN}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003cp\u003eSpecificity =\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\frac{TN}{TN+FP}\\)\u003c/span\u003e\u003c/span\u003e\u003c/p\u003e \u003cp\u003eBalanced Accuracy =\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\frac{\\text{Sensitivity}+\\text{Specificity}}{2}=\\frac{1}{2}\\left(\\frac{TP}{TP+FN}+\\frac{TN}{TN+FP}\\right)\\)\u003c/span\u003e\u003c/span\u003e\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$F1\\text{ score}=2\\times \\frac{\\text{precision}\\times \\text{recall}}{\\text{precision}+\\text{recall}}=2\\times \\frac{TP}{2\\times TP+FP+FN}$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$MCC=\\frac{TP\\times TN-FP\\times FN}{\\sqrt{\\left(TP+FP\\right)\\left(TP+FN\\right)\\left(TN+FP\\right)\\left(TN+FN\\right)}}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere TP, TN, FP, and FN stand for the true positives, true negatives, false positives, and false negatives, respectively. The following metrics are computed for regression models:\u003cdiv id=\"Equc\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$${R}^{2}=1-\\frac{{\\sum }_{i=1}^{n}{\\left({y}_{i}-\\stackrel{\\prime }{{y}_{i}}\\right)}^{2}}{{\\sum }_{i=1}^{n}{\\left({y}_{i}-\\overline{y}\\right)}^{2}}$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equd\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equd\" name=\"EquationSource\"\u003e\n$$\\text{Adjusted }{\\text{R}}^{2}=1-\\left(1-{R}^{2}\\right)\\times \\frac{n-1}{n-k-1}$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Eque\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Eque\" name=\"EquationSource\"\u003e\n$$\\text{RMSE}=\\sqrt{\\frac{1}{n}{\\sum }_{i=1}^{n}{\\left({y}_{i}-\\stackrel{\\prime }{{y}_{i}}\\right)}^{2}}$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere y is the actual value, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\stackrel{\\prime }{{y}_{i}}\\)\u003c/span\u003e\u003c/span\u003e is the corresponding prediction, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\overline{y}\\)\u003c/span\u003e\u003c/span\u003e is the mean of the actual values in the set and N is the number of test data points.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results and Discussion","content":"\u003cp\u003eThe purpose of this study is to experiment with various chemical representations and algorithmic strategies that may impact how well the ML models predict blockers and non-blockers of \u003cem\u003eKv\u003c/em\u003e1.5. After conducting all experiments, we initially selected the top ten models from 42 regression and classification models as good performance models based on the lowest RMSE and highest accuracy, respectively obtained in the test set. Out of ten models, ML models were further analyzed, as RMSE and accuracy measures were insufficient to adequately assess a model based on the dataset, where the number of active compounds was nearly twice as great as the number of inactive compounds. The detailed results are thoroughly analyzed in this section.\u003c/p\u003e \u003cp\u003eIn quantitative structure-activity relationship (QSAR) studies, using regression models based on descriptors and fingerprints is a common approach to predict the potency of a molecule (pIC50). These models aim to forecast the binding affinity using the chemical representation of molecular descriptors and fingerprints, for understanding potential activity (blockers and non-blockers) based on the predicted pIC50 values.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e3.1. Regression Models\u003c/h2\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003e3.1.1. Lipinski Descriptor-based Regression Models\u003c/h2\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMetrics of the training and test set for the top three regression models using descriptor-based according to the Lipinski rule\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDataset\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eR-Squared\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAdjusted R-Squared\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRFR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.882\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.881\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.285\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTrain\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.650\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.647\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.490\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.752\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.750\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.413\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRFR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.344\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.327\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.668\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.314\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.297\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.683\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.309\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.290\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.685\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eWhen compared to the RMSE and R\u0026sup2;, all the generated models using molecular descriptors in this study perform similarly, as shown by the results in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. Among the models evaluated in this study, three ensemble-based methods namely RFR, GBR, and Hist Gradient Boosting Regressor (HGBR), demonstrated high efficacy in terms of predictive accuracy and goodness of fit. It is observed that GBR and HGBR yielded the same R\u0026sup2; of 0.31, however, GBR achieved an RMSE of 0.68 in the test set, which is slightly better than HGBR, indicating its ability to iteratively improve, refine predictions and minimize the prediction errors. Overall, the RFR demonstrated the best performance, outperforming the other models in terms of RMSE and R\u0026sup2; metrics. This suggests that the RFR model succeeded in capturing the underlying patterns and relationships within the data, among all other models. Since the predictive outcomes of these models are closer to a difference of 0.01 for RMSE, adjusted R\u0026sup2; is utilized as a metric to provide insights into how well the models fit the data and penalizes the addition of redundant predictors that do not contribute significantly to the model\u0026rsquo;s predictive power. Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e clearly shows that the RFR regression model outperformed all other regression models in predicting binding affinity, as evidenced by its higher adjusted R\u0026sup2; value of 0.327.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section3\"\u003e \u003ch2\u003e3.1.2. Fingerprint-based Regression Models\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMetrics of the training and test set for regression models using fingerprint\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFingerprint\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eModels\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eR-Squared\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eAdjusted R-Squared\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.592\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.582\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.530\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEstate-Train\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.579\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.570\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.538\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRFR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.699\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.692\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.455\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.410\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.353\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.633\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEstate-Test\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.394\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.336\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.641\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRFR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.344\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.281\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.668\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSVR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.548\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.537\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.557\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSubstructure-Train\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.577\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.567\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.539\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.577\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.567\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.539\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSVR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.361\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.294\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.659\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSubstructure-Test\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.360\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.294\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.659\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.360\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.294\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.659\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.743\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.731\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.420\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSubstructCount-Train\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.851\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.844\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.319\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.851\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.844\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.319\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.479\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.363\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.595\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSubstructCount-Test\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eLGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.505\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.395\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.580\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHGBR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.505\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.395\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.580\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eOut of all the fingerprint-based regression methods, the models shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e performed consistently overall metrics in the training and testing set. Notably, variants of the GBR and RFR performed better than every other model tested using fingerprint. Although RMSE values were comparatively closer to these nine models in the test set, they showed considerable differences in adjusted R\u0026sup2;. It can be seen that Estate-RFR performs poorly than all models with the highest RMSE of 0.668, and the lowest adjusted R\u0026sup2; of 0.281. When incorporating the Substructure fingerprint, all of the models perform similarly in RMSE and adjusted R\u0026sup2; with a value of 0.659, and 0.294, respectively, however, the Subsructure-SVR model\u0026rsquo;s R\u0026sup2; value slightly differs. Based on the fact that SubstructureCount has the lowest RMSE of any model, it is a promising model in terms of RMSE. SubtructureCount-LGBR and SubtructureCount-HGBR perform equally well as all other models with similar results in RMSE, R\u0026sup2;, and adjusted R\u0026sup2;. Further, we have considered unrounded values of adjusted R\u0026sup2; for models SubtructureCount-LGBR and SubtructureCount-HGBR, yielding values of 0.3946130914, and 0.3949747272 in the test set, respectively, indicating these values each account for around 39.46% and 39.50% of the variation in the binding affinity, respectively. Nevertheless, after examining the results for all the regressors with different fingerprints, it is obvious that the SubtructureCount-HGBR is the best regressor in terms of its RMSE, R\u0026sup2;, and adjusted R\u0026sup2;.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e3.2. Classification Models\u003c/h2\u003e \u003cp\u003eSMOTE method is used throughout the section to balance active and inactive molecules in each experiment. As mentioned earlier, the performance of the classification models was evaluated using the accuracy for both descriptors and fingerprints, initially in correctly classifying instances to blockers and non-blockers of the Kv1.5 channel. Subsequently, we assessed and compared the best models using various metrics.\u003c/p\u003e \u003cdiv id=\"Sec12\" class=\"Section3\"\u003e \u003ch2\u003e3.2.1. Lipinski Descriptor-based Classification Models\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMetrics of the training and test set for classification models using descriptors to classify blockers and non-blockers.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDatasetModels\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSN\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eSP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eBA\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eF1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eMCC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXGBC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.835\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.850\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.994\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.994\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.685\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.994\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTrainDTC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.873\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.880\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.994\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.994\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.753\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.994\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKNNC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.808\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.893\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.873\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.873\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.703\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.873\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLPC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.808\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.893\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.994\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.994\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.703\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.873\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXGBC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.871\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.639\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.677\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.765\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.510\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.812\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTestDTC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.851\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.500\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.657\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.741\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.362\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.723\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKNNC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.762\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.528\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.748\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.781\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.275\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.739\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLPC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.762\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.528\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.657\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.741\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.275\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.739\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAfter conducting experiments with classification models using molecular descriptors, we obtained the top ten models with the highest AC values, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e. It is observed that the Extreme Gradient Boosting classifier (XGBC), DTC, LPC, and RFC demonstrated high accuracy rates, where XGBC and DTC performed the same with greater AC in both train and test sets. While the KNNC model had relatively lower accuracy in classification, it depicted a significant characteristic regarding overfitting. In comparison to all the other models, it was discovered that the KNNC model\u0026rsquo;s difference between the train and test accuracies was substantially lower. This suggests that the KNNC model is less prone to overfitting since it generalizes better to unseen data. In addition, RFC and LPC also performed equally in test sets, albeit LPC generalizes better because the difference between the train and test sets is smaller than that of RFC. Further investigation of these models is depicted in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, where we have taken four models\u0026rsquo; performances in train and test sets into consideration, namely, XGBC, DTC, LPC, and KNNC to validate the best descriptor classification model. It can be observed that KNNC and LPC both performed equally well in all metrics, with the exception that KNNC outperformed LPC in terms of BA values.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMetrics of the training and test set for classification models using fingerprints to classify blockers and non-blockers.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe XGBC model demonstrated superior performance compared to the other models evaluated, having an MCC of 0.510 on the testing set, however, the KNNC model is a generalized model, indicating a strong ability to predict and classify the target variable. Additionally, the KNNC model appears to have a higher F1 score implies that the model has a good balance between precision and recall discriminating between active and inactive instances, although other measures values are significantly closer to the XGBC model, including the AUC value of 0.739, SN of 0.762, and SP of 0.528 in the test set. In summary, the KNNC model can generalize well than all the other models, including XGBC, DTC, and LPC, in terms of BA, F1, and accuracies, making it the most suitable model using molecular descriptors.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section3\"\u003e \u003ch2\u003e3.2.2. Fingerprint-based Classification Models\u003c/h2\u003e \u003cp\u003eThe findings of fingerprint-based classification models with accuracy levels of at least 80% in the test set are shown in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. From the results presented in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e, it can be noted that all the developed models in this study perform similarly regarding accuracy with minimal differences between these models, therefore, different metrics are considered for interpreting the predictive results. In particular, SN has remarkably greater values in the test set in maximum models, suggesting that the model has a higher ability to correctly recognize true positive cases. Conversely, as the number of positive instances is much higher than the number of negative instances, specificity has comparatively lower values in all models, indicating that the number of correctly identified negative classes is comparatively lower. AtomsPair2DCount-DTC obtained the highest accuracy in the test set, albeit other measures, including BA of 0.664, F1 of 0.723, and MCC of 0.376, failed to demonstrate adequate potential.\u003c/p\u003e \u003cp\u003eWhile AtomsPair2D-RFC is the second-best model with just a 0.73% difference in accuracy, with the highest MCC of 0.515, BA of 0.747, F1 of 0.819, SN of 0.931, and SP of 0.528. However, when comparing train accuracies, the model has a higher training accuracy of 93.21%, it exhibits a larger drop in performance on the test set with a difference of 10.73%, leading to a potential issue with overfitting. Therefore, Substructure-RFC is a better suitable model with a training accuracy of 85.14% and test accuracy of 82.48%, and significant performance in other measures as well, since it presents a more balanced performance and reflects how well the model generalizes to unseen data.\u003c/p\u003e \u003cp\u003eIn predictive modeling, generalization is a crucial factor, especially in drug discovery. Generalized models are frequently utilized in the preliminary stages of drug development when researchers are searching a broad chemical space for new therapeutic compounds. Mainly, in this study, we have prioritized generalized models for \u003cem\u003eKv\u003c/em\u003e1.5 that can help researchers quickly narrow down drug candidates that are toxic or beneficial for treating atrial fibrillation by saving time and resources. These findings contribute to the advancement of predictive models for binding affinity, facilitating more reliable predictions in drug discovery and development. Further studies are required to explore the robustness and scalability of these models, considering larger datasets and examining their effectiveness on diverse chemical compounds.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"4. Conclusions","content":"\u003cp\u003eIn this work, we have experimented with and analyzed the performance of classification and regression models, trained on both descriptors and fingerprints along with oversampling technique to effectively predict inhibitors of \u003cem\u003eKv\u003c/em\u003e1.5. We found that different models perform best in different molecular attributes, including RFR and KNNC for descriptor-based models, and SubtructureCount-HGBR, and Substructure-RFC for fingerprint-based models. The classification-based model seems capable of producing well-balanced predictive models. Despite having crucial effects on the prolonged action potential duration due to inhibition of \u003cem\u003eKv\u003c/em\u003e1.5, studies specifically focused on this channel have been less common than those on the hERG channel and sodium channel. Thus, we have dealt with fewer datasets; nonetheless, we have achieved promising results in correctly identifying true positive values.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAuthor Contributions:\u0026nbsp;\u003c/strong\u003e\u0026ldquo;Conceptualization, M.H.R, methodology, validation and formal analysis H.A, S.K.Y, A.M.O, M.T.Z; resources, M.H.R.; original draft preparation S.K.Y, H.A, A.M.O, R.A.R, M.H.R.; writing, review and editing S.K.Y, H.A, A.M.O, R.A.R, K.S.H, M.S.I and M.H.R. project administration M.H.R, supervision and funding acquisition M.S.I and \u0026nbsp;M.H.R. Finally, the manuscript prepares from the input of all authors. \u0026nbsp;All authors have read and agreed to the published version of the manuscript.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003cstrong\u003eAcknowledgment\u003c/strong\u003e: M.H.R. acknowledges the NSU CTRG Research Grant CTRG-20/SEPS/06 and CTRG-21/SEPS/21. \u0026nbsp;M. S. I acknowledge NSU CTRG Research Grant CTRG-21/SEPS/11.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003cstrong\u003eCompeting Interests:\u003c/strong\u003e The author declares no conflict of interest, be it financial or non-financial.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eTristani-Firouzi, M.; Chen, J.; Mitcheson, J.S.; Sanguinetti, M.C. Molecular biology of K+ channels and their role in cardiac arrhythmias. The American journal of medicine 2001, 110, 50\u0026ndash;59.\u003c/li\u003e\n\u003cli\u003eLitalien, C.; Beaulieu, P. Molecular Aspects of Drug Actions: From Receptors to Effectors. In Pediatric Critical Care; Elsevier, 2006; pp. 1659\u0026ndash;1677.\u003c/li\u003e\n\u003cli\u003eRoden, D. Mechanism and management of proarrhythmia. Am. J. Cardiol 1998, 82, 491\u0026ndash;571.\u003c/li\u003e\n\u003cli\u003eKuang, Q.; Purhonen, P.; Hebert, H. Structure of potassium channels. Cellular and molecular life sciences 2015, 72, 3677\u0026ndash;3693.\u003c/li\u003e\n\u003cli\u003eCampomanes, C.R.; Carroll, K.I.; Manganas, L.N.; Hershberger, M.E.; Gong, B.; Antonucci, D.E.; Rhodes, K.J.; Trimmer, J.S. Kv\u0026beta; subunit oxidoreductase activity and Kv1 potassium channel trafficking. Journal of Biological Chemistry 2002, 277, 8298\u0026ndash;8305.\u003c/li\u003e\n\u003cli\u003eSansom, M.S. Ion channels: a first view of K+ channels in atomic glory. Current biology 1998, 8, R450\u0026ndash;R452.\u003c/li\u003e\n\u003cli\u003eWettwer, E.; Terlau, H. Pharmacology of voltage-gated potassium channel Kv1. 5\u0026mdash;impact on cardiac excitability. Current Opinion in Pharmacology 2014, 15, 115\u0026ndash;121.\u003c/li\u003e\n\u003cli\u003eKim, D.M.; Nimigean, C.M. Voltage-gated potassium channels: a structural examination of selectivity and gating. Cold Spring Harbor perspectives in biology 2016, 8, a029231.\u003c/li\u003e\n\u003cli\u003eGiudicessi, J.R.; Ackerman, M.J. Potassium-channel mutations and cardiac arrhythmias\u0026mdash;diagnosis and therapy. Nature Reviews Cardiology 2012, 9, 319\u0026ndash;332.\u003c/li\u003e\n\u003cli\u003eTamargo, J.; Caballero, R.; G\u0026oacute;mez, R.; Delp\u0026oacute;n, E. IKur/Kv1. 5 channel blockers for the treatment of atrial fibrillation. Expert opinion on investigational drugs 2009, 18, 399\u0026ndash;416.\u003c/li\u003e\n\u003cli\u003ePellman, J.; Sheikh, F. Atrial fibrillation: mechanisms, therapeutics, and future directions. Comprehensive Physiology 2015, 5, 649.\u003c/li\u003e\n\u003cli\u003eComes, N.; Bielanska, J.; Vallejo-Gracia, A.; Serrano-Albarr\u0026aacute;s, A.; Marruecos, L.; G\u0026oacute;mez, D.; Soler, C.; Condom, E.; Ram\u0026oacute;n y Cajal, S.; Hern\u0026aacute;ndez-Losa, J.; et al. The voltage-dependent K+ channels Kv1. 3 and Kv1. 5 in human cancer. Frontiers in physiology 2013, 4, 283.\u003c/li\u003e\n\u003cli\u003eClarfield, A.M. Brocklehurst\u0026rsquo;s Textbook of Geriatric Medicine and Gerontology. JAMA 2010, 304, 1956\u0026ndash;1957.\u003c/li\u003e\n\u003cli\u003eWang, Z.; Fermini, B.; Nattel, S. Effects of flecainide, quinidine, and 4-aminopyridine on transient outward and ultrarapid delayed rectifier currents in human atrial myocytes. Journal of Pharmacology and Experimental Therapeutics 1995, 272, 184\u0026ndash;196.\u003c/li\u003e\n\u003cli\u003eBrendorp, B.; Pedersen, O.D.; Torp-Pedersen, C.; Sahebzadah, N.; K\u0026oslash;ber, L. A benefit-risk assessment of class III antiarrhythmic agents. Drug safety 2002, 25, 847\u0026ndash;865.\u003c/li\u003e\n\u003cli\u003eKing, A.M.; Menke, N.B.; Katz, K.D.; Pizon, A.F. 4-aminopyridine toxicity: a case report and review of the literature. Journal of Medical Toxicology 2012, 8, 314\u0026ndash;321.\u003c/li\u003e\n\u003cli\u003eValentin, J.P.; Guth, B.; Hamlin, R.L.; Lain\u0026eacute;e, P.; Sarazan, D.; Skinner, M. Functional cardiac safety evaluation of novel therapeutics. Antitargets and Drug Safety 2015, pp. 199\u0026ndash;234.\u003c/li\u003e\n\u003cli\u003eSager, P.T.; Nebout, T.; Darpo, B. ICH E14: a new regulatory guidance on the clinical evaluation of QT/QTc internal prolongation and proarrhythmic potential for non-antiarrhythmic drugs. Drug information journal: DIJ/Drug Information Association 2005, 39, 387\u0026ndash;394.\u003c/li\u003e\n\u003cli\u003ePriest, B.; Bell, I.M.; Garcia, M. Role of hERG potassium channel assays in drug development. Channels 2008, 2, 87\u0026ndash;93.\u003c/li\u003e\n\u003cli\u003eKong, W.; Tu, X.; Huang, W.; Yang, Y.; Xie, Z.; Huang, Z. Prediction and optimization of NaV1. 7 sodium channel inhibitors based on machine learning and simulated annealing. Journal of Chemical Information and Modeling 2020, 60, 2739\u0026ndash;2753.\u003c/li\u003e\n\u003cli\u003eCai, C.; Fang, J.; Guo, P.; Wang, Q.; Hong, H.; Moslehi, J.; Cheng, F. In silico pharmacoepidemiologic evaluation of drug-induced cardiovascular complications using combined classifiers. Journal of chemical information and modeling 2018, 58, 943\u0026ndash;956.\u003c/li\u003e\n\u003cli\u003eMeng, J.; Zhang, L.; Wang, L.; Li, S.; Xie, D.; Zhang, Y.; Liu, H. TSSF-hERG: A machine-learning-based hERG potassium channel-specific scoring function for chemical cardiotoxicity prediction. Toxicology 2021, 464, 153018.\u003c/li\u003e\n\u003cli\u003eArab, I.; Barakat, K. ToxTree: descriptor-based machine learning models for both hERG and Nav1. 5 cardiotoxicity liability predictions. arXiv preprint arXiv:2112.13467 2021.\u003c/li\u003e\n\u003cli\u003eKhalifa, N.; Kumar Konda, L.S.; Kristam, R. Machine learning-based QSAR models to predict sodium ion channel (Nav 1.5) blockers. Future Medicinal Chemistry 2020, 12, 1829\u0026ndash;1843.\u003c/li\u003e\n\u003cli\u003eCai, C.; Guo, P.; Zhou, Y.; Zhou, J.; Wang, Q.; Zhang, F.; Fang, J.; Cheng, F. Deep learning-based prediction of drug-i nduced cardiotoxicity. Journal of chemical information and modeling 2019, 59, 1073\u0026ndash;1084.\u003c/li\u003e\n\u003cli\u003eRyu, J.Y.; Lee, M.Y.; Lee, J.H.; Lee, B.H.; Oh, K.S. DeepHIT: a deep learning framework for prediction of hERG-induced cardiotoxicity. Bioinformatics 2020, 36, 3049\u0026ndash;3055.\u003c/li\u003e\n\u003cli\u003eGaulton, A.; Hersey, A.; Nowotka, M.; Bento, A.P.; Chambers, J.; Mendez, D.; Mutowo, P.; Atkinson, F.; Bellis, L.J.; Cibri\u0026aacute;n-Uhalte, E.; et al. The ChEMBL database in 2017. Nucleic acids research 2017, 45, D945\u0026ndash;D954.\u003c/li\u003e\n\u003cli\u003eKim, S.; Thiessen, P.A.; Bolton, E.E.; Chen, J.; Fu, G.; Gindulyte, A.; Han, L.; He, J.; He, S.; Shoemaker, B.A.; et al. PubChem substance and compound databases. Nucleic acids research 2016, 44, D1202\u0026ndash;D1213.\u003c/li\u003e\n\u003cli\u003eAzlim Khan, A.K.; Ahamed Hassain Malim, N.H. Comparative Studies on Resampling Techniques in Machine Learning and Deep Learning Models for Drug-Target Interaction Prediction. Molecules 2023, 28, 1663.\u003c/li\u003e\n\u003cli\u003eBastikar, V.; Bastikar, A.; Gupta, P. Quantitative structure\u0026ndash;activity relationship-based computational approaches. In Computational Approaches for Novel Therapeutic and Diagnostic Designing to Mitigate SARS-CoV2 Infection; Elsevier, 2022; pp. 191\u0026ndash;205.\u003c/li\u003e\n\u003cli\u003eLagorce, D.; Douguet, D.; Miteva, M.A.; Villoutreix, B.O. Computational analysis of calculated physicochemical and ADMET properties of protein-protein interaction inhibitors. Scientific reports 2017, 7, 46277.\u003c/li\u003e\n\u003cli\u003eMcKnight, P.E.; Najab, J. Mann-Whitney U Test. The Corsini encyclopedia of psychology 2010, pp. 1\u0026ndash;1.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Table 4","content":"\u003cp\u003eTable 4 is available in Supplementary Files section.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"drug discovery, machine learning, RDkit, K+ channel, blockers","lastPublishedDoi":"10.21203/rs.3.rs-3263007/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3263007/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eAtrial fibrillation and associated cardiac problems may be treated with the development of potent potassium ion channel \u003cem\u003eKv\u003c/em\u003e1.5 blockers. Since the use of these blockers provides therapeutic advantages and potential side effects, it is significant to identify \u003cem\u003eKv\u003c/em\u003e1.5 channel blockers from compounds. In this work, we employed optimized machine learning models to predict the potential of small molecules in blocking the \u003cem\u003eKv\u003c/em\u003e1.5 channel to address the limitations of traditional screening methods in the drug discovery process. Several machine learning classifiers and regression models were employed utilizing molecular descriptors and fingerprints incorporating with SMOTE oversampling technique to overcome the class imbalance in active and inactive molecules. The results show that distinct models excelled in predicting different molecular attributes. The regression models demonstrated superior performance with random forest regression (RFR) (root-mean-square error = 0.668) and Substructure-Count-HGBR (Histogram-based Gradient Boosting Regression) having adjusted R² of 39.50% for predicting binding affinity. The best-performing models among the fingerprint-based models were the k-Nearest Neighbors Classifier (KNNC) and Substructure-RFC (Random Forest Classifier), which both demonstrated well-balanced predictive models. The generalized machine learning models for \u003cem\u003eKv\u003c/em\u003e1.5 can help researchers quickly narrow down drug candidates that are toxic or beneficial for treating atrial fibrillation in the early stages of drug discovery.\u003c/p\u003e","manuscriptTitle":"Efficacy of Small Molecules Blocking in Kv1.5 Potassium Channel From Machine Learning Models","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-08-21 13:29:42","doi":"10.21203/rs.3.rs-3263007/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"4264bba1-e857-4aa7-a564-d389bbdc954e","owner":[],"postedDate":"August 21st, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2023-11-15T10:01:08+00:00","versionOfRecord":[],"versionCreatedAt":"2023-08-21 13:29:42","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3263007","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3263007","identity":"rs-3263007","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.