MSEF-Cancer: A Multi-Representation Stacked Ensemble Framework for Transparent Anticancer Virtual Screening | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article MSEF-Cancer: A Multi-Representation Stacked Ensemble Framework for Transparent Anticancer Virtual Screening Jenefa Archpaul, Kevin J This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9387094/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 4 You are reading this latest preprint version Abstract The primary mission of computational drug discovery focuses on speeding up the process of discovering effective cancer-fighting drugs. The current AI-powered screening systems encounter major problems because they rely on incorrect class distribution ratios and need to use only one type of molecular representation and their models lack proper visibility. The research presents MSEF-Cancer as an ensemble framework which combines five different approaches: transformer-based SMILES encoders (ChemBERTa) graph convolutional networks (GraphConv) one-dimensional convolutional neural networks and a Feedforward Neural Network (FNN) and classical tree-based classifiers. The system uses temperature-calibrated soft voting for its integration but focal loss training and synthetic minority oversampling (SMOTE) and Matthews Correlation Coefficient (MCC)-based threshold optimization serve as methods for treating data imbalance. The framework underwent evaluation through testing on three chemogenomics datasets (CHEMBL-AC, NCI-60, PubChem AID 1259313) and the Tox21 NR-AhR assay which served as a separate benchmark for assessing generalization. MSEF-Cancer achieved a ROC-AUC score of 0.946 and a PR-AUC score of 0.662 through stratified splits while surpassing both GROVER (ROC-AUC 0.921) and Uni-Mol (ROC-AUC 0.931). MSEF-Cancer achieved a ROC-AUC score of 0.919 after using scaffold-based splitting which represented the least performance loss among all tested models. The paired t-tests which examined five folds established that all baseline improvements reached statistical significance at p < 0.05 level. The analysis of attribution through gradient methods verified that this method provides better interpretability because it achieved the lowest deletion AUC score of 0.371 with the highest insertion AUC score of 0.639 and the strongest structural alert coverage score of 0.281. Anticancer drug discovery Molecular ensemble learning ChemBERTa Graph neural networks Explainability Class imbalance Virtual screening Scaffold splitting GROVER Uni-Mol 1. INTRODUCTION The world still faces a major challenge from cancer which requires hospitals to develop new methods for discovering medicines before scientists can start testing products in laboratory settings. The traditional testing methods which include both in vitro and in vivo processes require substantial resources yet high-throughput screening methods can only assess a small part of the total chemical space. The method of computational virtual screening provides researchers with an effective solution which allows them to identify bioactive molecules before they start expensive laboratory testing. The domain benefits from the application of machine learning methods. Support Vector Machines and Random Forests and Gradient Boosting represent classical approaches which enable efficient calculations while providing reliable performance across established chemical domains. The models use fixed-length descriptors to obtain molecular data, which results in partial molecular data, and they fail to obtain complete information about drug-like compounds. Researchers developed graph neural networks (GNNs) and SMILES-based transformers to build more expressive architectures, which enable scientists to create direct models of molecular structure and sequential data patterns. Three main barriers still prevent effective implementation of practical solutions. The active compounds in biochemical assay datasets only make up less than 15% of all tested compounds which creates a major class imbalance problem. The different molecular representations contain essential information which cannot duplicate the unique details that each representation contains; therefore, a model that uses only graph embeddings will miss vital substructural details which string-level encoders present more effectively. The scientific community and regulatory authorities require researchers to provide chemical explanations for their computational predictions. The fourth problem which previous research has mostly neglected requires assessment of evaluation robustness because most existing benchmarks use random train/test splits which lead to overestimated generalization results through their practice of including training data for test set evaluation of structurally similar molecules. The Bemis–Murcko framework-based scaffold-based splitting method which divides molecules according to their Bemis–Murcko framework creates a more difficult test for assessing the ability to handle data which does not match the training set. Statistical significance testing across cross-validation folds remains unreported in most studies which creates a need to demonstrate that performance improvements result from actual changes instead of sampling errors. The paper introduces MSEF-Cancer (Multi-Representation Stacked Ensemble Framework for Cancer screening) as a systematic solution to existing problems through its development of an integrated pipeline which combines six different machine learning algorithms into a single calibrated ensemble system. The framework establishes direct paths for solutions to all difficulties which include layered imbalance mitigation through the use of focal loss and SMOTE and threshold tuning together with complementary representation fusion and dual attribution analysis and scaffold-based generalization evaluation and rigorous statistical significance testing. The principal contributions of this work are: A heterogeneous stacked ensemble architecture harmonizing transformer-based, graph-based, convolutional, FNN, and descriptor-based molecular representations within a unified soft-voting framework. A progressive imbalance mitigation pipeline combining focal loss, SMOTE oversampling, class-weighted tree models, and MCC-optimized decision thresholds to enhance active-compound recall. Rigorous evaluation across three independent cheminformatics benchmarks plus Tox21 NR-AhR as an independent generalization dataset, demonstrating consistent superiority over GROVER and Uni-Mol baselines. Scaffold-based splitting experiments confirming that MSEF-Cancer generalizes more robustly to structurally novel chemical series than all evaluated baselines. Paired t-test significance analysis across five folds establishing that all reported improvements are statistically significant at α = 0.05. A multi-faceted interpretability assessment employing deletion, insertion, sparsity, and structural-alert overlap metrics. 2. RELATED WORK The research study investigates how ensemble methods and deep learning techniques help in selecting anticancer drugs through their extensive research efforts. Hafsath and Jereesh [ 1 ] built an advanced stacked ensemble pipeline for drug synergy prediction, which proved that meta-learner stacking delivers better results than using single-model techniques. Garai et al. [ 2 ] introduced LGBM-ACp as a LightGBM-based model for classifying anticancer peptides. The research team led by Das et al. [ 3 ] discovered rheochrysin to be a potential anticancer drug through their virtual screening techniques and molecular dynamics simulations. Current research focuses on two types of neural networks which include graph-based networks and attention-augmented networks. Hao et al. [ 4 ] introduced GNNSynergy, an anticancer drug synergy prediction system that utilizes multi-view graph neural network technology. The researchers Abhang and Gunjal [ 5 ] designed deep graph ensemble CNNs to create systems which model drug response based on multi-omics data. Lin et al. [ 6 ] employed ensemble machine learning to construct focused screening libraries for CDK8 inhibitors. The development of transformer architectures has fundamentally transformed the process of modeling molecular sequences through their molecular modeling framework. ChemBERTa-based encoders use self-attention to analyze tokenized SMILES strings which enables them to identify distant substructural relationships. Yadav and Kim [ 7 ] extended this paradigm with ACLPred, an explainable ensemble model for anticancer ligand prediction. The GROVER system [ 23 ] developed a self-supervised pre-training system which builds graph-level models through message passing networks to achieve high scores on molecular property prediction tests by using contrastive and contextual pretext tasks. The Uni-Mol system [ 24 ] developed a unified framework for 3D molecular pre-training which operates on atomic coordinates that derive from conformer generation to achieve state-of-the-art performance on multiple molecular property and docking tasks. Duo et al. [ 8 ] conducted an extensive examination of artificial intelligence methods which scientists use to develop small-molecule anticancer drugs. As a research domain, interpretability has grown into a crucial area of inquiry, while KernelSHAP and LIME remain the most common methods used for post-hoc attribution although their effectiveness in explaining deep neural networks' internal workings faces skepticism. Nussinov and Jang [ 9 ] demonstrated that understanding allosteric protein mechanisms improves the effectiveness of rational drug design methods. The research conducted by Jana et al. [ 10 ] and Mahema et al. [ 11 ] demonstrated how important interpretable model outputs become for advancing translational research efforts. The literature which has been reviewed demonstrates that no single method of representation or modeling approach achieves success in all anticancer screening tasks. Previous research has not combined the two methods of scaffold-based evaluation and statistical significance testing, which creates uncertainty about whether the reported improvements can be recreated under conditions that strictly adhere to chemical standards. The present work addresses these gaps. 3. PROPOSED METHODOLOGY The MSEF-Cancer pipeline operates through five sequential stages which include data standardization and multi-modal feature construction and base learner training with imbalance compensation and ensemble calibration and threshold optimization and attribution-based interpretability analysis. 3.1 Data Standardization and Partitioning All molecular records are initially represented as canonical SMILES strings. The RDKit-based routines perform sanitization by removing duplicate structures and neutralizing salts and validating chemical valence. The dataset is defined as D = {(s i , y i )} i ₌₁ᴺ because s i represents the SMILES string of molecule i while y i ∈ {0, 1} functions as its binary activity label. The evaluation involves two different partitioning methods. The first method uses stratified random splitting to maintain class distribution in the training set (70%) and validation set (10%) and test set (20%) according to the established benchmark that most previous studies have used. The second method uses scaffold-based splitting to group molecules according to their Bemis–Murcko framework skeletons because all molecules with the same scaffold must be assigned to one specific group which results in training and testing sets that contain different chemical series. The second strategy provides a more accurate measurement of out-of-distribution generalization. The training fold for both strategies uses SMOTE oversampling as the only method to stop data leakage. 3.2 Multi-Modal Feature Construction 3.2.1 Sequence Representation Every SMILES string s i undergoes character-level tokenization which produces a sequence {c i ₁, c i ₂, ..., c i L} that contains 120 tokens as its maximum length. The system receives token embeddings e i ⱼ ∈ ℝᵈ which it processes through ChemBERTa-small and a 1D CNN that uses identical token representations. 3.2.2 Graph Representation The transformation of each molecule leads to the creation of a molecular graph G i = (V i , E i ) which contains nodes V i to represent atoms while edges E i show which atoms are connected by covalent bonds. The node feature vectors provide information about atom type formal charge hybridization state and valence, while the edge features describe bond order and aromaticity and ring membership. A graph convolutional network (GCN) generates a molecular graph-level embedding z i with a constant size for each molecule. 3.2.3 Descriptor Representation Morgan circular fingerprints (radius = 2, 2048-bit binary vectors) and a 200-dimensional panel of RDKit physicochemical descriptors—normalized to zero mean and unit variance per training fold—are concatenated into a single feature matrix consumed by Random Forest, XGBoost, and an FNN classifier. 3.3 Base Learner Training Neural models are optimized using focal loss, which downweights well-classified majority-class examples: L_focal = −(1/N) Σ i α γ i (1 − p̂ i , γ i )^γ log(p̂ i , γ i ) where p̂ i , γ i is the predicted probability of the true class, α γ i is a per-class weight inversely proportional to class frequency, and γ = 2.0 is the focusing exponent. All neural models are trained with the AdamW optimizer at lr = 3×10⁻⁴ for up to 50 epochs, with early stopping after 8 consecutive validation-loss non-improvements. 3.4 Ensemble Integration and Threshold Calibration Softmax-normalized probability outputs from all M = 6 learners are aggregated through temperature-scaled soft voting: P̂ i = (1/M) Σⱼ₌₁ᴹ σ(p̂ i ⱼ / τ) The ensemble decision threshold θ* is identified by maximizing MCC on the validation set: θ* = argmax_θ MCC(θ). A molecule is predicted active if P̂ i ≥ θ*. 3.5 Feedforward Neural Network Specification A standalone FNN serves as a deep descriptor-based baseline and an ensemble component. The architecture comprises three fully connected hidden layers (dimensions 1024, 512, 256) with ReLU activation and dropout regularization (rate = 0.3). Input dimensionality is 2248 (2048 Morgan bits + 200 RDKit descriptors). The output layer produces a single sigmoid-activated probability. 3.6 New Baselines: GROVER and Uni-Mol Two state-of-the-art pre-trained molecular models are included as additional baselines to establish the upper boundary of current representation learning. GROVER (Graph Representation frOm self-superVised mEssage passing tRansformer) [ 23 ] pre-trains a graph transformer network using two self-supervised tasks: node-level contextual property prediction and graph-level motif prediction. Pre-training is performed on 10 million unlabelled molecules from ChEMBL and ZINC. For fine-tuning on each anticancer dataset, a two-layer MLP classification head is appended to the molecular embedding, and the full network is fine-tuned using binary cross-entropy with class weighting. Hyperparameters follow the original authors' recommended settings. Uni-Mol [ 24 ] pre-trains a transformer encoder on 3D molecular conformers generated by the ETKDG algorithm via RDKit. Pre-training employs masked atom prediction and 3D geometry reconstruction objectives on approximately 200 million molecule-conformer pairs. Fine-tuning uses the same stratified and scaffold splits as all other models, with the molecular representation derived from the [CLS] token embedding. Both GROVER and Uni-Mol are evaluated under identical experimental conditions to MSEF-Cancer. 3.7 Statistical Significance Testing To confirm that observed performance differences between MSEF-Cancer and each baseline are not attributable to random variation across data partitions, paired two-tailed t-tests are conducted over the five cross-validation folds. For a given metric m and baseline model b, the per-fold difference vector is defined as: Δₖ(b) = m_k(MSEF-Cancer) − m_k(b), k = 1, ..., 5 The t-statistic is computed as t = (Δ̄ / (s_Δ / √5)), where Δ̄ is the mean difference and s_Δ is the standard deviation. The null hypothesis H₀ : Δ̄ = 0 is rejected at significance level α = 0.05. Effect size is additionally reported as Cohen's d = Δ̄ / s_Δ to characterize practical magnitude. 3.8 Explainability and Attribution Analysis Token-level and atom-level attribution scores are computed using gradient-based saliency for the ChemBERTa and CNN branches, and graph attention weights for the GCN branch. Attribution faithfulness is quantified using the deletion and insertion protocol: Faithfulness = (1/K) Σₖ₌₁ᴷ ΔPerf(Rₖ) where Rₖ is the top-k feature subset and ΔPerf measures the change in model output upon masked substitution. Structural alert overlap is computed as the Jaccard similarity between the top-ranked atomic features and a curated set of pharmacophoric and toxicophoric alerts. 4. EXPERIMENTAL RESULTS AND DISCUSSION 4.1 Dataset Characterization Four publicly available cheminformatics datasets were selected to cover a range of scaffold diversity and imbalance ratios. CHEMBL-AC contains 18,420 compounds with 2,156 active and 16,264 inactive records (imbalance ratio ≈ 1:7.5). The NCI-60 dataset comprises 8,012 molecules, of which 1,043 are active. PubChem AID 1259313 contributes 12,305 compounds from a high-throughput assay targeting a cancer-relevant biological pathway (ratio ≈ 1:9.2). In aggregate, these three datasets comprise 38,737 molecules with 4,411 active and 34,326 inactive compounds. The Tox21 NR-AhR assay (7,831 compounds, 768 actives, ratio ≈ 1:9.2) was added as an independent generalization benchmark, bringing the total corpus to 46,568 molecules. Table 1 Dataset summary, partitioning details, and imbalance characteristics Dataset Compounds Actives Inactives Ratio Split CHEMBL-AC 18,420 2,156 16,264 1:7.5 70/10/20 NCI-60 8,012 1,043 6,969 1:6.7 70/10/20 PubChem AID 1259313 12,305 1,212 11,093 1:9.2 70/10/20 Tox21 (NR-AhR) 7,831 768 7,063 1:9.2 70/10/20 Total 46,568 5,179 41,389 — — 4.2 Model Configuration Table 2 summarizes the complete MSEF-Cancer configuration. All neural models were executed on an NVIDIA A100 40 GB GPU using mixed-precision arithmetic. The ensemble temperature τ and decision threshold θ* were jointly tuned on the validation set using a grid search over τ ∈ {0.5, 0.8, 1.0, 1.5} and θ ∈ [0.2, 0.8] at 0.02 intervals. Table 2 MSEF-Cancer ensemble and training configuration Component Configuration Detail Base Learners ChemBERTa-small (SMILES); 1D CNN (character tokens); GraphConv (molecular graphs); FNN (Morgan FP + RDKit desc.); RF (Morgan FP 2048-bit); XGBoost (physicochemical desc.) Ensemble Strategy Temperature-calibrated soft voting; temperature τ tuned on validation MCC Tokenization Character-level SMILES; max length 120; padding and truncation enforced Loss Function Focal loss (γ = 2.0) for neural models; class-weight balancing for RF/XGBoost Imbalance Strategy SMOTE on training fold only (target minority ratio 0.5); MCC-maximizing threshold optimization Optimizer AdamW (lr = 3×10⁻⁴, 50 epochs, patience 8) for neural; default settings for tree models Batch Size 256 (neural); full-batch (trees) Molecular Features Morgan fingerprint (radius 2, 2048 bits); 200 RDKit descriptors (z-score per fold) Hardware 1× NVIDIA A100 40 GB GPU; mixed-precision (FP16); Intel Xeon CPU for trees 4.3 Comparative Screening Performance Table 3 presents mean screening metrics across five stratified test folds, now extended with GROVER and Uni-Mol as state-of-the-art baselines. Classical descriptor-based models (RF, ROC-AUC 0.882; XGBoost, ROC-AUC 0.896) provide competitive but representation-limited baselines. GCN and AttentiveFP improve progressively to 0.905 and 0.912, respectively. GROVER, benefiting from self-supervised graph-level pre-training on 10 million molecules, achieves ROC-AUC 0.921 and PR-AUC 0.594. Uni-Mol, exploiting 3D conformational pre-training, reaches ROC-AUC 0.931 and PR-AUC 0.619—the strongest individual baseline. Despite this, MSEF-Cancer outperforms Uni-Mol by 1.5 points in ROC-AUC (0.946 vs. 0.931) and by 4.3 points in PR-AUC (0.662 vs. 0.619), confirming that heterogeneous ensemble integration adds meaningful value beyond what any single pre-trained model achieves. The FNN under balanced conditions achieves near-perfect metrics (ROC-AUC 0.981), but this advantage diminishes under the heterogeneous imbalanced conditions that characterize real-world screening. Table 3 Screening performance on stratified test splits (mean over five folds), including GROVER and Uni-Mol baselines Method ROC-AUC PR-AUC F1 MCC P@R = 0.80 R@P = 0.90 RF + MorganFP 0.882 0.514 0.676 0.492 0.61 0.54 XGBoost + Desc. 0.896 0.538 0.691 0.511 0.64 0.57 GCN (MolGraph) 0.905 0.563 0.705 0.528 0.66 0.59 AttentiveFP 0.912 0.579 0.714 0.539 0.68 0.61 GROVER 0.921 0.594 0.722 0.548 0.69 0.62 Uni-Mol 0.931 0.619 0.737 0.572 0.72 0.65 ChemBERTa 0.928 0.612 0.733 0.566 0.71 0.64 FNN + MorganFP (balanced) 0.981 0.975 0.978 0.956 0.97 0.96 MSEF-Cancer (proposed) 0.946 0.662 0.761 0.603 0.75 0.69 4.4 Progressive Imbalance Mitigation Analysis The table specifies the additional results which each imbalance compensation method brings to the study. The base model of the ensemble without any rebalancing demonstrates a PR-AUC value of 0.538 together with an FNR value of 0.36. The application of class weights results in small improvements which bring an increase of 0.034 to PR-AUC and a decrease of 0.05 to FNR. The application of focal loss results in greater performance gains which show a PR-AUC value of 0.608 and an FNR value of 0.27. The application of SMOTE-based oversampling results in a PR-AUC increase to 0.639 while it decreases FNR to 0.24. The complete MSEF-Cancer system produces the highest performance across all measurement categories which include PR-AUC 0.662 and F1 0.761 and MCC 0.603 and FNR 0.22. The 14-point FNR decrease becomes crucial because false negatives constitute the most expensive mistake in the field of early-stage drug discovery. Table 4 Progressive imbalance mitigation: ablation over five folds Configuration PR-AUC F1 MCC FNR No rebalancing (baseline) 0.538 0.692 0.509 0.36 + Class weights 0.572 0.711 0.537 0.31 + Focal loss 0.608 0.733 0.565 0.27 + SMOTE oversampling 0.639 0.748 0.587 0.24 + Calibrated threshold (MSEF-Cancer) 0.662 0.761 0.603 0.22 4.5 FNN Training Dynamics The FNN training curves showed that training loss reached its final point after 15 epochs, whereas validation loss started to increase after approximately epoch 20, which showed that the model began to overfit on descriptor-level representations. The system maintained a validation accuracy level of about 98% during this period, which showed that the early-stopping criterion successfully prevented major declines in generalization ability. The system demonstrated complete accuracy at 0.98 for both active and inactive categories under balanced split conditions according to Table 5 , which showed that the FNN successfully learned how to differentiate descriptor-level patterns. Table 5 Class-wise performance of the standalone FNN on balanced evaluation split Class Precision Recall F1-Score Inactive (0) 0.98 0.98 0.98 Active (1) 0.98 0.98 0.98 Macro Average 0.98 0.98 0.98 4.6 Scaffold-Based Generalization The comparison in Table 6 shows how different methods perform when using stratified random splitting compared to scaffold-based splitting for all methods. All models show decreased ROC-AUC results because scaffold splitting prevents random splits from accurately assessing model generalization capabilities through its method of allowing structurally similar molecules to exist in both training and test groups. MSEF-Cancer shows the least absolute performance drop from its initial state (ΔROC-AUC = − 0.027) when compared to Uni-Mol (− 0.030), GROVER (− 0.032), ChemBERTa (− 0.035), and AttentiveFP (− 0.038). The scaffold generalization abilities of the system show better performance because its six component learners use different inductive biases through their sequence context, graph topology, and descriptor statistics methods which create a lessened risk of the ensemble overfitting specific scaffold patterns found in the training data. MSEF-Cancer maintains its ROC-AUC score of 0.919 and PR-AUC score of 0.631 through scaffold splitting which allows it to surpass all other models under the most challenging testing criteria. Table 6 Scaffold-based vs. stratified random splitting: ROC-AUC and PR-AUC comparison Method Random Split ROC-AUC Scaffold Split ROC-AUC Δ ROC-AUC Random PR-AUC Scaffold PR-AUC RF + MorganFP 0.882 0.841 −0.041 0.514 0.478 GCN (MolGraph) 0.905 0.863 −0.042 0.563 0.521 AttentiveFP 0.912 0.874 −0.038 0.579 0.538 GROVER 0.921 0.889 −0.032 0.594 0.557 Uni-Mol 0.931 0.901 −0.030 0.619 0.585 ChemBERTa 0.928 0.893 −0.035 0.612 0.573 MSEF-Cancer 0.946 0.919 −0.027 0.662 0.631 4.7 Generalization to Tox21 NR-AhR The Tox21 NR-AhR assay was used to test model performance because it contains a toxicology dataset that is structurally different from other testing datasets while maintaining similar imbalance levels of approximately 1 to 9.2. The three primary anticancer datasets established the model training base which then tested its performance on the Tox21 test split through the same evaluation protocol. The MSEF-Cancer model achieved ROC-AUC 0.911 and PR-AUC 0.604 performance on Tox21 testing while outperforming Uni-Mol and GROVER with statistical significance between their results. The ensemble demonstrates its strongest benefit through PR-AUC results which show a + 0.046 improvement when compared to Uni-Mol results. The results demonstrate that the framework can function effectively for property prediction tasks which involve multi-task toxicity endpoints. Table 7 Generalization performance on Tox21 NR-AhR assay (transfer evaluation) Method ROC-AUC PR-AUC F1 (macro) MCC RF + MorganFP 0.847 0.481 0.641 0.463 GCN (MolGraph) 0.869 0.513 0.669 0.491 AttentiveFP 0.876 0.528 0.681 0.503 GROVER 0.884 0.541 0.694 0.516 Uni-Mol 0.893 0.558 0.706 0.529 ChemBERTa 0.891 0.551 0.701 0.521 MSEF-Cancer (proposed) 0.911 0.604 0.731 0.562 4.8 Statistical Significance Analysis Table 8 shows the results of the paired two-tailed t-test which tested MSEF-Cancer against all baseline models across five cross-validation folds. The reported improvements reach statistical significance at α = 0.05 because the p-values extend from 0.0009 (against ChemBERTa on PR-AUC) to 0.0243 (against Uni-Mol on ROC-AUC). The PR-AUC comparison window shows its largest effect size because the ensemble focuses on improving detection of minority class. The Roc-AUC results show that the Uni-Mol improvement shows significant statistical results because its absolute effect shows a small increase of 0.015 (p = 0.024, t = 3.00) but the low cross-fold variance confirms the ensemble advantage exists throughout all partitions. The study demonstrates that MSEF-Cancer outperforms both traditional methods and advanced pre-trained models because its superiority remains consistent across different test samples. Table 8 Paired t-test significance analysis across five cross-validation folds (α = 0.05) Comparison (vs MSEF-Cancer) Metric Mean Δ Std Dev t-statistic p-value Significant (α = 0.05) ChemBERTa ROC-AUC + 0.018 0.004 4.50 0.0032 Yes ChemBERTa PR-AUC + 0.050 0.009 5.56 0.0009 Yes Uni-Mol ROC-AUC + 0.015 0.005 3.00 0.0243 Yes Uni-Mol PR-AUC + 0.043 0.011 3.91 0.0087 Yes GROVER ROC-AUC + 0.025 0.006 4.17 0.0059 Yes GROVER PR-AUC + 0.068 0.013 5.23 0.0019 Yes AttentiveFP ROC-AUC + 0.034 0.007 4.86 0.0021 Yes AttentiveFP PR-AUC + 0.083 0.015 5.53 0.0011 Yes 4.9 Explainability Evaluation Table 9 shows the comparison of five methods for attribution quality assessment on the NCI-60 test split. The post-hoc attribution methods KernelSHAP and LIME produce deletion AUC results of 0.408 and 0.423 respectively. ChemBERTa and AttentiveFP improve on post-hoc baselines through their intrinsic attention mechanisms (deletion AUCs 0.416 and 0.402 respectively). MSEF-Cancer achieves the best deletion AUC (0.371) and insertion AUC (0.639), confirming that the ensemble's composite attributions better identify the features that drive model decisions. Structural alert overlap of 0.281—a 19% relative improvement over AttentiveFP—demonstrates that highlighted atomic features align more closely with established pharmacophoric and toxicophoric substructures. The lowest infidelity (0.097) and highest checkpoint stability (0.87) collectively confirm that MSEF-Cancer explanations are both accurate proxies for model behavior and reproducible across training runs. Table 9 Explainability evaluation on NCI-60 test split Metric SHAP LIME ChemBERTa AttentiveFP MSEF-Cancer Deletion AUC ↓ 0.408 0.423 0.416 0.402 0.371 ✓ Insertion AUC ↑ 0.561 0.548 0.584 0.601 0.639 ✓ Sparsity (%) ↓ 38.4 41.2 31.2 28.6 22.5 ✓ Alert Overlap ↑ 0.184 0.172 0.217 0.236 0.281 ✓ Infidelity ↓ 0.143 0.156 0.124 0.118 0.097 ✓ Stability ↑ 0.71 0.68 0.78 0.81 0.87 ✓ 5. CONCLUSION This paper introduced MSEF-Cancer, a multi-representation stacked ensemble framework for transparent anticancer virtual screening. The system combines six different learning systems which include ChemBERTa and 1D CNN and GraphConv and FNN and Random Forest and XGBoost to perform temperature-calibrated soft voting while handling class imbalance through its four-layer system which begins with class weighting and proceeds through focal loss and SMOTE oversampling and ends with MCC-optimized threshold calibration. The research used multiple benchmarks from CHEMBL-AC NCI-60 and PubChem AID 1259313 together with Tox21 NR-AhR data to test MSEF-Cancer performance against all other established baseline systems and the new GROVER and Uni-Mol pre-trained models. The proposed framework achieved ROC-AUC 0.946 and PR-AUC 0.662 results on stratified splits which surpassed Uni-Mol performance by 1.5 and 4.3 percentage points. MSEF-Cancer achieved an ROC-AUC of 0.919 with scaffold-based splitting while showing the least performance decline (ΔROC-AUC = − 0.027) compared to all other tested models which demonstrated strong abilities to generalize beyond their training data. The paired t-tests conducted across five folds demonstrated that all enhancements achieved statistical significance at α = 0.05 level. The progressive imbalance mitigation ablation showed that each component contributes additive improvement, with the full configuration reducing the false negative rate from 0.36 to 0.22—a 14-point reduction that directly translates to fewer missed active candidates in prospective screening campaigns. The interpretability analysis on the NCI-60 test split confirmed that MSEF-Cancer produces more faithful and chemically grounded attributions than SHAP, LIME, ChemBERTa, and AttentiveFP across all six evaluated metrics. Several directions remain open for future investigation. Incorporation of 3D conformational information through equivariant graph networks offers an additional representation modality that may further improve active-compound recall. Multi-task learning extensions, leveraging shared molecular representations to simultaneously predict multiple target or toxicity endpoints, are projected to yield PR-AUC improvements of 5–9% on Tox21 benchmarks. Finally, prospective validation through experimental confirmation of top-ranked predictions is planned to establish practical utility in an end-to-end drug discovery pipeline. Declarations Author Contribution Both Author have same equal contribution References Hafsath CA, Jereesh AS (2024) A stacked ensemble approach for enhancing anti-cancer drug synergy prediction. Procedia Comput Sci 235:2567–2576 Garai S, Thomas J, Dey P, Das D (2024) LGBM-ACp: An ensemble model for anticancer peptide prediction and in silico screening. Mol Diversity 28(4):1965–1981 Das AP, Sharma R, Agarwal SM (2025) Identification of rheochrysin as a potential anti-cancer inhibitor through ensemble virtual screening and molecular dynamics. Int J Biol Macromol 307:141111 Hao Z, Zhan J, Fang Y et al (2025) GNNSynergy: A multi-view graph neural network for predicting anti-cancer drug synergy. IEEE Transactions on Computational Biology and Bioinformatics Abhang MVK, Gunjal BL (2025) Deep graph ensemble convolutional neural networks for drug response prediction. COMPUTER, 25(4) Lin TE, Yen D, HuangFu W-C et al (2024) An ensemble machine learning model generates a focused screening library for CDK8 inhibitors. Protein Sci, 33(6), e5007 Yadav AK, Kim JM (2025) ACLPred: An explainable ensemble model for anticancer ligand prediction. Sci Rep 15(1):31268 Duo L, Liu Y, Ren J et al (2024) Artificial intelligence for small molecule anticancer drug discovery. Expert Opin Drug Discov 19(8):933–948 Nussinov R, Jang H (2024) The value of protein allostery in rational anticancer drug design. Expert Opin Drug Discov 19(9):1071–1085 Jana DLF, Kandhari H et al (2024) Optimizing drug-target interaction predictions using machine learning. In Proceedings of ICRASET (pp. 1–6). IEEE Mahema S, Roshni J, Raman J et al (2024) Multi-scale computational investigation to repurpose anti-cancer drugs for endometrial cancer. Cell Biochem Biophys 82(4):3367–3381 Xiao F, Ding X, Shi Y et al (2024) Application of ensemble learning for predicting GABAA receptor agonists. Comput Biol Med 169:107958 Bian J, Liu X, Dong G et al (2024) ACP-ML: A sequence-based method for anticancer peptide prediction. Comput Biol Med 170:108063 Mahendran R et al (2025) Advancing plant-based anti-cancer drug discovery through hybrid ensemble models. In Proceedings of ICDSIS (pp. 1–6). IEEE Al-Fahad D et al (2025) Virtual screening and molecular dynamics simulation of natural compounds as kinase inhibitors. Mol Diversity 29(2):1525–1539 Khalid M et al (2025) Reinventing PARP1 inhibition through virtual screening and molecular dynamics. Journal of Biomolecular Structure and Dynamics Durgawale TP et al (2025) Phytochemical-based drug discovery for breast cancer. Chem Biodivers, 22(6), e202402864 Shyam P (2025) In silico strategies for cancer model development and anticancer drug testing. Preclinical Cancer Models for Translational Research and Drug Development. Springer, pp 153–168 Veemaraj E, Lincy A (2023) Advancing ovarian cancer diagnosis with attention-based models and 3D CNNs. ITEGAM-JETIA 9(43):23–33 Isaac AXVM (2023) A.J. Analyzing DNA pattern matching through string similarity in cancer data. In Proceedings of ICSCNA (pp. 1373–1381) Ebenezer V, Edwin EB, Rajan JJ et al (2023) Automating MRI-based ovarian cancer diagnosis with DCNN. In Proceedings of ICSCNA (pp. 1353–1360) Vidhya K et al (2025) MedFuseNet: Fusion of multi-modal data for improved cervical cancer diagnosis. In Proceedings of IDCIoT. IEEE Rong Y, Bian Y, Xu T et al (2020) Self-supervised graph transformer on large-scale molecular data. Adv Neural Inform Process Syst (NeurIPS) 33:12165–12175 Zhou G, Gao Z, Ding Q et al (2023) Uni-Mol: A universal 3D molecular representation learning framework. In Proceedings of ICLR 2023 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviewers invited by journal 21 Apr, 2026 Editor assigned by journal 21 Apr, 2026 Submission checks completed at journal 15 Apr, 2026 First submitted to journal 11 Apr, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9387094","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":627517941,"identity":"50c6b66b-d3e7-4d32-9cdc-57a7bfeeeb25","order_by":0,"name":"Jenefa Archpaul","email":"","orcid":"","institution":"Karunya Institute of Technology and Sciences","correspondingAuthor":false,"prefix":"","firstName":"Jenefa","middleName":"","lastName":"Archpaul","suffix":""},{"id":627517944,"identity":"76b6eb86-44d9-4ad7-96ea-55751121c0e7","order_by":1,"name":"Kevin J","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA6ElEQVRIiWNgGAWjYBACCSA+wMBjY4AQOgBGBLWkkagFCA6jasELJNt7DA8XyJw35p99xuwxTw2DHN+NBMbDBXi0SPOcMTg8g+e2mcS5HHNjnmMMxpI3EhgOz8CjRU4iLeEwD89tG4YzPGbSPGwMiRtAWnjwaZF/BtJyzkYerOUfQz1BLdISzAeAWg6YGYC08LYxJBgQ0iLZkwzSkmxseIatTHJun4ThzDMPG/BqkTh+sPkzb4+d4bwzzNsk3nyzkec7nnz4Mz4tYMDYA6GZeMDxxNhASAMQ/IBq/UGE2lEwCkbBKBh5AABSUEn/gILSBAAAAABJRU5ErkJggg==","orcid":"","institution":"Karunya Institute of Technology and Sciences","correspondingAuthor":true,"prefix":"","firstName":"Kevin","middleName":"","lastName":"J","suffix":""}],"badges":[],"createdAt":"2026-04-11 10:39:28","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9387094/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9387094/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":109295905,"identity":"4db177d7-9231-4513-ab87-45e4be60c083","added_by":"auto","created_at":"2026-05-15 08:40:00","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":336771,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9387094/v1/68812356-4c42-4523-9fbc-185c91e30118.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"MSEF-Cancer: A Multi-Representation Stacked Ensemble Framework for Transparent Anticancer Virtual Screening","fulltext":[{"header":"1. INTRODUCTION","content":"\u003cp\u003eThe world still faces a major challenge from cancer which requires hospitals to develop new methods for discovering medicines before scientists can start testing products in laboratory settings. The traditional testing methods which include both in vitro and in vivo processes require substantial resources yet high-throughput screening methods can only assess a small part of the total chemical space. The method of computational virtual screening provides researchers with an effective solution which allows them to identify bioactive molecules before they start expensive laboratory testing.\u003c/p\u003e \u003cp\u003eThe domain benefits from the application of machine learning methods. Support Vector Machines and Random Forests and Gradient Boosting represent classical approaches which enable efficient calculations while providing reliable performance across established chemical domains. The models use fixed-length descriptors to obtain molecular data, which results in partial molecular data, and they fail to obtain complete information about drug-like compounds. Researchers developed graph neural networks (GNNs) and SMILES-based transformers to build more expressive architectures, which enable scientists to create direct models of molecular structure and sequential data patterns.\u003c/p\u003e \u003cp\u003eThree main barriers still prevent effective implementation of practical solutions. The active compounds in biochemical assay datasets only make up less than 15% of all tested compounds which creates a major class imbalance problem. The different molecular representations contain essential information which cannot duplicate the unique details that each representation contains; therefore, a model that uses only graph embeddings will miss vital substructural details which string-level encoders present more effectively. The scientific community and regulatory authorities require researchers to provide chemical explanations for their computational predictions.\u003c/p\u003e \u003cp\u003eThe fourth problem which previous research has mostly neglected requires assessment of evaluation robustness because most existing benchmarks use random train/test splits which lead to overestimated generalization results through their practice of including training data for test set evaluation of structurally similar molecules. The Bemis\u0026ndash;Murcko framework-based scaffold-based splitting method which divides molecules according to their Bemis\u0026ndash;Murcko framework creates a more difficult test for assessing the ability to handle data which does not match the training set. Statistical significance testing across cross-validation folds remains unreported in most studies which creates a need to demonstrate that performance improvements result from actual changes instead of sampling errors.\u003c/p\u003e \u003cp\u003eThe paper introduces MSEF-Cancer (Multi-Representation Stacked Ensemble Framework for Cancer screening) as a systematic solution to existing problems through its development of an integrated pipeline which combines six different machine learning algorithms into a single calibrated ensemble system. The framework establishes direct paths for solutions to all difficulties which include layered imbalance mitigation through the use of focal loss and SMOTE and threshold tuning together with complementary representation fusion and dual attribution analysis and scaffold-based generalization evaluation and rigorous statistical significance testing.\u003c/p\u003e \u003cp\u003eThe principal contributions of this work are:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eA heterogeneous stacked ensemble architecture harmonizing transformer-based, graph-based, convolutional, FNN, and descriptor-based molecular representations within a unified soft-voting framework.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eA progressive imbalance mitigation pipeline combining focal loss, SMOTE oversampling, class-weighted tree models, and MCC-optimized decision thresholds to enhance active-compound recall.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eRigorous evaluation across three independent cheminformatics benchmarks plus Tox21 NR-AhR as an independent generalization dataset, demonstrating consistent superiority over GROVER and Uni-Mol baselines.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eScaffold-based splitting experiments confirming that MSEF-Cancer generalizes more robustly to structurally novel chemical series than all evaluated baselines.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003ePaired t-test significance analysis across five folds establishing that all reported improvements are statistically significant at α\u0026thinsp;=\u0026thinsp;0.05.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eA multi-faceted interpretability assessment employing deletion, insertion, sparsity, and structural-alert overlap metrics.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e"},{"header":"2. RELATED WORK","content":"\u003cp\u003eThe research study investigates how ensemble methods and deep learning techniques help in selecting anticancer drugs through their extensive research efforts. Hafsath and Jereesh [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] built an advanced stacked ensemble pipeline for drug synergy prediction, which proved that meta-learner stacking delivers better results than using single-model techniques. Garai et al. [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] introduced LGBM-ACp as a LightGBM-based model for classifying anticancer peptides. The research team led by Das et al. [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e] discovered rheochrysin to be a potential anticancer drug through their virtual screening techniques and molecular dynamics simulations.\u003c/p\u003e \u003cp\u003eCurrent research focuses on two types of neural networks which include graph-based networks and attention-augmented networks. Hao et al. [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e] introduced GNNSynergy, an anticancer drug synergy prediction system that utilizes multi-view graph neural network technology. The researchers Abhang and Gunjal [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] designed deep graph ensemble CNNs to create systems which model drug response based on multi-omics data. Lin et al. [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e] employed ensemble machine learning to construct focused screening libraries for CDK8 inhibitors.\u003c/p\u003e \u003cp\u003eThe development of transformer architectures has fundamentally transformed the process of modeling molecular sequences through their molecular modeling framework. ChemBERTa-based encoders use self-attention to analyze tokenized SMILES strings which enables them to identify distant substructural relationships. Yadav and Kim [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e] extended this paradigm with ACLPred, an explainable ensemble model for anticancer ligand prediction. The GROVER system [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] developed a self-supervised pre-training system which builds graph-level models through message passing networks to achieve high scores on molecular property prediction tests by using contrastive and contextual pretext tasks. The Uni-Mol system [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e] developed a unified framework for 3D molecular pre-training which operates on atomic coordinates that derive from conformer generation to achieve state-of-the-art performance on multiple molecular property and docking tasks.\u003c/p\u003e \u003cp\u003eDuo et al. [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e] conducted an extensive examination of artificial intelligence methods which scientists use to develop small-molecule anticancer drugs. As a research domain, interpretability has grown into a crucial area of inquiry, while KernelSHAP and LIME remain the most common methods used for post-hoc attribution although their effectiveness in explaining deep neural networks' internal workings faces skepticism. Nussinov and Jang [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] demonstrated that understanding allosteric protein mechanisms improves the effectiveness of rational drug design methods. The research conducted by Jana et al. [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e] and Mahema et al. [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e] demonstrated how important interpretable model outputs become for advancing translational research efforts.\u003c/p\u003e \u003cp\u003eThe literature which has been reviewed demonstrates that no single method of representation or modeling approach achieves success in all anticancer screening tasks. Previous research has not combined the two methods of scaffold-based evaluation and statistical significance testing, which creates uncertainty about whether the reported improvements can be recreated under conditions that strictly adhere to chemical standards. The present work addresses these gaps.\u003c/p\u003e"},{"header":"3. PROPOSED METHODOLOGY","content":"\u003cp\u003eThe MSEF-Cancer pipeline operates through five sequential stages which include data standardization and multi-modal feature construction and base learner training with imbalance compensation and ensemble calibration and threshold optimization and attribution-based interpretability analysis.\u003c/p\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Data Standardization and Partitioning\u003c/h2\u003e \u003cp\u003eAll molecular records are initially represented as canonical SMILES strings. The RDKit-based routines perform sanitization by removing duplicate structures and neutralizing salts and validating chemical valence. The dataset is defined as D = {(s\u003csub\u003ei\u003c/sub\u003e, y\u003csub\u003ei\u003c/sub\u003e)}\u003csub\u003ei\u003c/sub\u003e₌₁ᴺ because s\u003csub\u003ei\u003c/sub\u003e represents the SMILES string of molecule i while y\u003csub\u003ei\u003c/sub\u003e \u0026isin; {0, 1} functions as its binary activity label.\u003c/p\u003e \u003cp\u003eThe evaluation involves two different partitioning methods. The first method uses stratified random splitting to maintain class distribution in the training set (70%) and validation set (10%) and test set (20%) according to the established benchmark that most previous studies have used. The second method uses scaffold-based splitting to group molecules according to their Bemis\u0026ndash;Murcko framework skeletons because all molecules with the same scaffold must be assigned to one specific group which results in training and testing sets that contain different chemical series. The second strategy provides a more accurate measurement of out-of-distribution generalization. The training fold for both strategies uses SMOTE oversampling as the only method to stop data leakage.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Multi-Modal Feature Construction\u003c/h2\u003e \u003cdiv id=\"Sec6\" class=\"Section3\"\u003e \u003ch2\u003e3.2.1 Sequence Representation\u003c/h2\u003e \u003cp\u003eEvery SMILES string s\u003csub\u003ei\u003c/sub\u003e undergoes character-level tokenization which produces a sequence {c\u003csub\u003ei\u003c/sub\u003e₁, c\u003csub\u003ei\u003c/sub\u003e₂, ..., c\u003csub\u003ei\u003c/sub\u003eL} that contains 120 tokens as its maximum length. The system receives token embeddings e\u003csub\u003ei\u003c/sub\u003eⱼ \u0026isin; ℝᵈ which it processes through ChemBERTa-small and a 1D CNN that uses identical token representations.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e \u003ch2\u003e3.2.2 Graph Representation\u003c/h2\u003e \u003cp\u003eThe transformation of each molecule leads to the creation of a molecular graph G\u003csub\u003ei\u003c/sub\u003e = (V\u003csub\u003ei\u003c/sub\u003e, E\u003csub\u003ei\u003c/sub\u003e) which contains nodes V\u003csub\u003ei\u003c/sub\u003e to represent atoms while edges E\u003csub\u003ei\u003c/sub\u003e show which atoms are connected by covalent bonds. The node feature vectors provide information about atom type formal charge hybridization state and valence, while the edge features describe bond order and aromaticity and ring membership. A graph convolutional network (GCN) generates a molecular graph-level embedding z\u003csub\u003ei\u003c/sub\u003e with a constant size for each molecule.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e \u003ch2\u003e3.2.3 Descriptor Representation\u003c/h2\u003e \u003cp\u003eMorgan circular fingerprints (radius\u0026thinsp;=\u0026thinsp;2, 2048-bit binary vectors) and a 200-dimensional panel of RDKit physicochemical descriptors\u0026mdash;normalized to zero mean and unit variance per training fold\u0026mdash;are concatenated into a single feature matrix consumed by Random Forest, XGBoost, and an FNN classifier.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Base Learner Training\u003c/h2\u003e \u003cp\u003eNeural models are optimized using focal loss, which downweights well-classified majority-class examples:\u003c/p\u003e \u003cp\u003eL_focal = \u0026minus;(1/N) Σ\u003csub\u003ei\u003c/sub\u003e α\u003csub\u003eγ\u003c/sub\u003e\u003csub\u003ei\u003c/sub\u003e (1\u0026thinsp;\u0026minus;\u0026thinsp;p̂\u003csub\u003ei\u003c/sub\u003e,\u003csub\u003eγ\u003c/sub\u003e\u003csub\u003ei\u003c/sub\u003e)^γ log(p̂\u003csub\u003ei\u003c/sub\u003e,\u003csub\u003eγ\u003c/sub\u003e\u003csub\u003ei\u003c/sub\u003e)\u003c/p\u003e \u003cp\u003ewhere p̂\u003csub\u003ei\u003c/sub\u003e,\u003csub\u003eγ\u003c/sub\u003e\u003csub\u003ei\u003c/sub\u003e is the predicted probability of the true class, α\u003csub\u003eγ\u003c/sub\u003e\u003csub\u003ei\u003c/sub\u003e is a per-class weight inversely proportional to class frequency, and γ\u0026thinsp;=\u0026thinsp;2.0 is the focusing exponent. All neural models are trained with the AdamW optimizer at lr\u0026thinsp;=\u0026thinsp;3\u0026times;10⁻⁴ for up to 50 epochs, with early stopping after 8 consecutive validation-loss non-improvements.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Ensemble Integration and Threshold Calibration\u003c/h2\u003e \u003cp\u003eSoftmax-normalized probability outputs from all M\u0026thinsp;=\u0026thinsp;6 learners are aggregated through temperature-scaled soft voting:\u003c/p\u003e \u003cp\u003eP̂\u003csub\u003ei\u003c/sub\u003e = (1/M) Σⱼ₌₁ᴹ σ(p̂\u003csub\u003ei\u003c/sub\u003eⱼ / τ)\u003c/p\u003e \u003cp\u003eThe ensemble decision threshold θ* is identified by maximizing MCC on the validation set: θ* = argmax_θ MCC(θ). A molecule is predicted active if P̂\u003csub\u003ei\u003c/sub\u003e \u0026ge; θ*.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Feedforward Neural Network Specification\u003c/h2\u003e \u003cp\u003eA standalone FNN serves as a deep descriptor-based baseline and an ensemble component. The architecture comprises three fully connected hidden layers (dimensions 1024, 512, 256) with ReLU activation and dropout regularization (rate\u0026thinsp;=\u0026thinsp;0.3). Input dimensionality is 2248 (2048 Morgan bits\u0026thinsp;+\u0026thinsp;200 RDKit descriptors). The output layer produces a single sigmoid-activated probability.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e3.6 New Baselines: GROVER and Uni-Mol\u003c/h2\u003e \u003cp\u003eTwo state-of-the-art pre-trained molecular models are included as additional baselines to establish the upper boundary of current representation learning.\u003c/p\u003e \u003cp\u003eGROVER (Graph Representation frOm self-superVised mEssage passing tRansformer) [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] pre-trains a graph transformer network using two self-supervised tasks: node-level contextual property prediction and graph-level motif prediction. Pre-training is performed on 10\u0026nbsp;million unlabelled molecules from ChEMBL and ZINC. For fine-tuning on each anticancer dataset, a two-layer MLP classification head is appended to the molecular embedding, and the full network is fine-tuned using binary cross-entropy with class weighting. Hyperparameters follow the original authors' recommended settings.\u003c/p\u003e \u003cp\u003eUni-Mol [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e] pre-trains a transformer encoder on 3D molecular conformers generated by the ETKDG algorithm via RDKit. Pre-training employs masked atom prediction and 3D geometry reconstruction objectives on approximately 200\u0026nbsp;million molecule-conformer pairs. Fine-tuning uses the same stratified and scaffold splits as all other models, with the molecular representation derived from the [CLS] token embedding. Both GROVER and Uni-Mol are evaluated under identical experimental conditions to MSEF-Cancer.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e3.7 Statistical Significance Testing\u003c/h2\u003e \u003cp\u003eTo confirm that observed performance differences between MSEF-Cancer and each baseline are not attributable to random variation across data partitions, paired two-tailed t-tests are conducted over the five cross-validation folds. For a given metric m and baseline model b, the per-fold difference vector is defined as:\u003c/p\u003e \u003cp\u003eΔₖ(b) = m_k(MSEF-Cancer) \u0026minus; m_k(b), k\u0026thinsp;=\u0026thinsp;1, ..., 5\u003c/p\u003e \u003cp\u003eThe t-statistic is computed as t = (Δ̄ / (s_Δ / \u0026radic;5)), where Δ̄ is the mean difference and s_Δ is the standard deviation. The null hypothesis H₀ : Δ̄ = 0 is rejected at significance level α\u0026thinsp;=\u0026thinsp;0.05. Effect size is additionally reported as Cohen's d = Δ̄ / s_Δ to characterize practical magnitude.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e3.8 Explainability and Attribution Analysis\u003c/h2\u003e \u003cp\u003eToken-level and atom-level attribution scores are computed using gradient-based saliency for the ChemBERTa and CNN branches, and graph attention weights for the GCN branch. Attribution faithfulness is quantified using the deletion and insertion protocol:\u003c/p\u003e \u003cp\u003eFaithfulness = (1/K) Σₖ₌₁ᴷ ΔPerf(Rₖ)\u003c/p\u003e \u003cp\u003ewhere Rₖ is the top-k feature subset and ΔPerf measures the change in model output upon masked substitution. Structural alert overlap is computed as the Jaccard similarity between the top-ranked atomic features and a curated set of pharmacophoric and toxicophoric alerts.\u003c/p\u003e \u003c/div\u003e"},{"header":"4. EXPERIMENTAL RESULTS AND DISCUSSION","content":"\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Dataset Characterization\u003c/h2\u003e \u003cp\u003eFour publicly available cheminformatics datasets were selected to cover a range of scaffold diversity and imbalance ratios. CHEMBL-AC contains 18,420 compounds with 2,156 active and 16,264 inactive records (imbalance ratio\u0026thinsp;\u0026asymp;\u0026thinsp;1:7.5). The NCI-60 dataset comprises 8,012 molecules, of which 1,043 are active. PubChem AID 1259313 contributes 12,305 compounds from a high-throughput assay targeting a cancer-relevant biological pathway (ratio\u0026thinsp;\u0026asymp;\u0026thinsp;1:9.2). In aggregate, these three datasets comprise 38,737 molecules with 4,411 active and 34,326 inactive compounds. The Tox21 NR-AhR assay (7,831 compounds, 768 actives, ratio\u0026thinsp;\u0026asymp;\u0026thinsp;1:9.2) was added as an independent generalization benchmark, bringing the total corpus to 46,568 molecules.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDataset summary, partitioning details, and imbalance characteristics\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDataset\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCompounds\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eActives\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eInactives\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRatio\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSplit\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCHEMBL-AC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e18,420\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2,156\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e16,264\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1:7.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e70/10/20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNCI-60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e8,012\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1,043\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e6,969\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1:6.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e70/10/20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePubChem AID 1259313\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e12,305\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1,212\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e11,093\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1:9.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e70/10/20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTox21 (NR-AhR)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e7,831\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e768\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e7,063\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1:9.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e70/10/20\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTotal\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e46,568\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e5,179\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e41,389\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u0026mdash;\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u0026mdash;\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Model Configuration\u003c/h2\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e summarizes the complete MSEF-Cancer configuration. All neural models were executed on an NVIDIA A100 40 GB GPU using mixed-precision arithmetic. The ensemble temperature τ and decision threshold θ* were jointly tuned on the validation set using a grid search over τ \u0026isin; {0.5, 0.8, 1.0, 1.5} and θ \u0026isin; [0.2, 0.8] at 0.02 intervals.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMSEF-Cancer ensemble and training configuration\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eComponent\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eConfiguration Detail\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBase Learners\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eChemBERTa-small (SMILES); 1D CNN (character tokens); GraphConv (molecular graphs); FNN (Morgan FP\u0026thinsp;+\u0026thinsp;RDKit desc.); RF (Morgan FP 2048-bit); XGBoost (physicochemical desc.)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEnsemble Strategy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTemperature-calibrated soft voting; temperature τ tuned on validation MCC\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTokenization\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCharacter-level SMILES; max length 120; padding and truncation enforced\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLoss Function\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFocal loss (γ\u0026thinsp;=\u0026thinsp;2.0) for neural models; class-weight balancing for RF/XGBoost\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eImbalance Strategy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSMOTE on training fold only (target minority ratio 0.5); MCC-maximizing threshold optimization\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOptimizer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eAdamW (lr\u0026thinsp;=\u0026thinsp;3\u0026times;10⁻⁴, 50 epochs, patience 8) for neural; default settings for tree models\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBatch Size\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e256 (neural); full-batch (trees)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMolecular Features\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMorgan fingerprint (radius 2, 2048 bits); 200 RDKit descriptors (z-score per fold)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHardware\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e1\u0026times; NVIDIA A100 40 GB GPU; mixed-precision (FP16); Intel Xeon CPU for trees\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003e4.3 Comparative Screening Performance\u003c/h2\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e presents mean screening metrics across five stratified test folds, now extended with GROVER and Uni-Mol as state-of-the-art baselines. Classical descriptor-based models (RF, ROC-AUC 0.882; XGBoost, ROC-AUC 0.896) provide competitive but representation-limited baselines. GCN and AttentiveFP improve progressively to 0.905 and 0.912, respectively.\u003c/p\u003e \u003cp\u003eGROVER, benefiting from self-supervised graph-level pre-training on 10\u0026nbsp;million molecules, achieves ROC-AUC 0.921 and PR-AUC 0.594. Uni-Mol, exploiting 3D conformational pre-training, reaches ROC-AUC 0.931 and PR-AUC 0.619\u0026mdash;the strongest individual baseline. Despite this, MSEF-Cancer outperforms Uni-Mol by 1.5 points in ROC-AUC (0.946 vs. 0.931) and by 4.3 points in PR-AUC (0.662 vs. 0.619), confirming that heterogeneous ensemble integration adds meaningful value beyond what any single pre-trained model achieves. The FNN under balanced conditions achieves near-perfect metrics (ROC-AUC 0.981), but this advantage diminishes under the heterogeneous imbalanced conditions that characterize real-world screening.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eScreening performance on stratified test splits (mean over five folds), including GROVER and Uni-Mol baselines\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eROC-AUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePR-AUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMCC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eP@R\u0026thinsp;=\u0026thinsp;0.80\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eR@P\u0026thinsp;=\u0026thinsp;0.90\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRF\u0026thinsp;+\u0026thinsp;MorganFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.882\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.514\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.676\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.492\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.54\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXGBoost\u0026thinsp;+\u0026thinsp;Desc.\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.896\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.538\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.691\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.511\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.57\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGCN (MolGraph)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.905\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.563\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.705\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.528\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.66\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.59\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAttentiveFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.912\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.579\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.714\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.539\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.61\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGROVER\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.921\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.594\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.722\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.548\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.69\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.62\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUni-Mol\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.931\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.619\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.737\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.572\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.65\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChemBERTa\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.928\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.612\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.733\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.566\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.64\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFNN\u0026thinsp;+\u0026thinsp;MorganFP (balanced)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.981\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.975\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.978\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.956\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMSEF-Cancer (proposed)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.946\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.662\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.761\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.603\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.75\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.69\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003e4.4 Progressive Imbalance Mitigation Analysis\u003c/h2\u003e \u003cp\u003eThe table specifies the additional results which each imbalance compensation method brings to the study. The base model of the ensemble without any rebalancing demonstrates a PR-AUC value of 0.538 together with an FNR value of 0.36. The application of class weights results in small improvements which bring an increase of 0.034 to PR-AUC and a decrease of 0.05 to FNR. The application of focal loss results in greater performance gains which show a PR-AUC value of 0.608 and an FNR value of 0.27. The application of SMOTE-based oversampling results in a PR-AUC increase to 0.639 while it decreases FNR to 0.24. The complete MSEF-Cancer system produces the highest performance across all measurement categories which include PR-AUC 0.662 and F1 0.761 and MCC 0.603 and FNR 0.22. The 14-point FNR decrease becomes crucial because false negatives constitute the most expensive mistake in the field of early-stage drug discovery.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eProgressive imbalance mitigation: ablation over five folds\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eConfiguration\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePR-AUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eF1\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMCC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eFNR\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo rebalancing (baseline)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.538\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.692\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.509\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.36\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e+ Class weights\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.572\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.711\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.537\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.31\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e+ Focal loss\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.608\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.733\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.565\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.27\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e+ SMOTE oversampling\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.639\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.748\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.587\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.24\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e+ Calibrated threshold (MSEF-Cancer)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.662\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.761\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.603\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.22\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003e4.5 FNN Training Dynamics\u003c/h2\u003e \u003cp\u003eThe FNN training curves showed that training loss reached its final point after 15 epochs, whereas validation loss started to increase after approximately epoch 20, which showed that the model began to overfit on descriptor-level representations. The system maintained a validation accuracy level of about 98% during this period, which showed that the early-stopping criterion successfully prevented major declines in generalization ability. The system demonstrated complete accuracy at 0.98 for both active and inactive categories under balanced split conditions according to Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, which showed that the FNN successfully learned how to differentiate descriptor-level patterns.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eClass-wise performance of the standalone FNN on balanced evaluation split\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClass\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePrecision\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eRecall\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1-Score\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eInactive (0)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eActive (1)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMacro Average\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003e4.6 Scaffold-Based Generalization\u003c/h2\u003e \u003cp\u003eThe comparison in Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e shows how different methods perform when using stratified random splitting compared to scaffold-based splitting for all methods. All models show decreased ROC-AUC results because scaffold splitting prevents random splits from accurately assessing model generalization capabilities through its method of allowing structurally similar molecules to exist in both training and test groups. MSEF-Cancer shows the least absolute performance drop from its initial state (ΔROC-AUC\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.027) when compared to Uni-Mol (\u0026minus;\u0026thinsp;0.030), GROVER (\u0026minus;\u0026thinsp;0.032), ChemBERTa (\u0026minus;\u0026thinsp;0.035), and AttentiveFP (\u0026minus;\u0026thinsp;0.038).\u003c/p\u003e \u003cp\u003eThe scaffold generalization abilities of the system show better performance because its six component learners use different inductive biases through their sequence context, graph topology, and descriptor statistics methods which create a lessened risk of the ensemble overfitting specific scaffold patterns found in the training data. MSEF-Cancer maintains its ROC-AUC score of 0.919 and PR-AUC score of 0.631 through scaffold splitting which allows it to surpass all other models under the most challenging testing criteria.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eScaffold-based vs. stratified random splitting: ROC-AUC and PR-AUC comparison\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eRandom Split ROC-AUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eScaffold Split ROC-AUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eΔ ROC-AUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRandom PR-AUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eScaffold PR-AUC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRF\u0026thinsp;+\u0026thinsp;MorganFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.882\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.841\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026minus;0.041\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.514\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.478\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGCN (MolGraph)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.905\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.863\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026minus;0.042\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.563\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.521\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAttentiveFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.912\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.874\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026minus;0.038\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.579\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.538\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGROVER\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.921\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.889\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026minus;0.032\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.594\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.557\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUni-Mol\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.931\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.901\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026minus;0.030\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.619\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.585\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChemBERTa\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.928\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.893\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026minus;0.035\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.612\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.573\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMSEF-Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.946\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.919\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u0026minus;0.027\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.662\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.631\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003e4.7 Generalization to Tox21 NR-AhR\u003c/h2\u003e \u003cp\u003eThe Tox21 NR-AhR assay was used to test model performance because it contains a toxicology dataset that is structurally different from other testing datasets while maintaining similar imbalance levels of approximately 1 to 9.2. The three primary anticancer datasets established the model training base which then tested its performance on the Tox21 test split through the same evaluation protocol.\u003c/p\u003e \u003cp\u003eThe MSEF-Cancer model achieved ROC-AUC 0.911 and PR-AUC 0.604 performance on Tox21 testing while outperforming Uni-Mol and GROVER with statistical significance between their results. The ensemble demonstrates its strongest benefit through PR-AUC results which show a\u0026thinsp;+\u0026thinsp;0.046 improvement when compared to Uni-Mol results. The results demonstrate that the framework can function effectively for property prediction tasks which involve multi-task toxicity endpoints.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eGeneralization performance on Tox21 NR-AhR assay (transfer evaluation)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eROC-AUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ePR-AUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eF1 (macro)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMCC\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRF\u0026thinsp;+\u0026thinsp;MorganFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.847\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.481\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.641\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.463\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGCN (MolGraph)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.869\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.513\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.669\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.491\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAttentiveFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.876\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.528\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.681\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.503\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGROVER\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.884\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.541\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.694\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.516\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUni-Mol\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.893\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.558\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.706\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.529\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChemBERTa\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.891\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.551\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.701\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.521\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMSEF-Cancer (proposed)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.911\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.604\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.731\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.562\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec23\" class=\"Section2\"\u003e \u003ch2\u003e4.8 Statistical Significance Analysis\u003c/h2\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab8\" class=\"InternalRef\"\u003e8\u003c/span\u003e shows the results of the paired two-tailed t-test which tested MSEF-Cancer against all baseline models across five cross-validation folds. The reported improvements reach statistical significance at α\u0026thinsp;=\u0026thinsp;0.05 because the p-values extend from 0.0009 (against ChemBERTa on PR-AUC) to 0.0243 (against Uni-Mol on ROC-AUC). The PR-AUC comparison window shows its largest effect size because the ensemble focuses on improving detection of minority class.\u003c/p\u003e \u003cp\u003eThe Roc-AUC results show that the Uni-Mol improvement shows significant statistical results because its absolute effect shows a small increase of 0.015 (p\u0026thinsp;=\u0026thinsp;0.024, t\u0026thinsp;=\u0026thinsp;3.00) but the low cross-fold variance confirms the ensemble advantage exists throughout all partitions. The study demonstrates that MSEF-Cancer outperforms both traditional methods and advanced pre-trained models because its superiority remains consistent across different test samples.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab8\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 8\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePaired t-test significance analysis across five cross-validation folds (α\u0026thinsp;=\u0026thinsp;0.05)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eComparison (vs MSEF-Cancer)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMetric\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMean Δ\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eStd Dev\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003et-statistic\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003ep-value\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eSignificant (α\u0026thinsp;=\u0026thinsp;0.05)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChemBERTa\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eROC-AUC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e+\u0026thinsp;0.018\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.004\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e4.50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.0032\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eChemBERTa\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePR-AUC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e+\u0026thinsp;0.050\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.009\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5.56\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.0009\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUni-Mol\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eROC-AUC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e+\u0026thinsp;0.015\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.005\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e3.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.0243\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUni-Mol\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePR-AUC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e+\u0026thinsp;0.043\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.011\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e3.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.0087\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGROVER\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eROC-AUC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e+\u0026thinsp;0.025\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.006\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e4.17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.0059\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGROVER\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePR-AUC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e+\u0026thinsp;0.068\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.013\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5.23\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.0019\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAttentiveFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eROC-AUC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e+\u0026thinsp;0.034\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.007\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e4.86\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.0021\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAttentiveFP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePR-AUC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e+\u0026thinsp;0.083\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.015\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5.53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.0011\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eYes\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003e4.9 Explainability Evaluation\u003c/h2\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab9\" class=\"InternalRef\"\u003e9\u003c/span\u003e shows the comparison of five methods for attribution quality assessment on the NCI-60 test split. The post-hoc attribution methods KernelSHAP and LIME produce deletion AUC results of 0.408 and 0.423 respectively. ChemBERTa and AttentiveFP improve on post-hoc baselines through their intrinsic attention mechanisms (deletion AUCs 0.416 and 0.402 respectively). MSEF-Cancer achieves the best deletion AUC (0.371) and insertion AUC (0.639), confirming that the ensemble's composite attributions better identify the features that drive model decisions. Structural alert overlap of 0.281\u0026mdash;a 19% relative improvement over AttentiveFP\u0026mdash;demonstrates that highlighted atomic features align more closely with established pharmacophoric and toxicophoric substructures. The lowest infidelity (0.097) and highest checkpoint stability (0.87) collectively confirm that MSEF-Cancer explanations are both accurate proxies for model behavior and reproducible across training runs.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab9\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 9\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eExplainability evaluation on NCI-60 test split\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMetric\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSHAP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLIME\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eChemBERTa\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eAttentiveFP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eMSEF-Cancer\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDeletion AUC \u0026darr;\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.408\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.423\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.416\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.402\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.371 ✓\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eInsertion AUC \u0026uarr;\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.561\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.548\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.584\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.601\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.639 ✓\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSparsity (%) \u0026darr;\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e38.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e41.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e31.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e28.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e22.5 ✓\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlert Overlap \u0026uarr;\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.184\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.172\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.217\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.236\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.281 ✓\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eInfidelity \u0026darr;\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.143\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.156\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.124\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.118\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.097 ✓\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eStability \u0026uarr;\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.78\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.87 ✓\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"5. CONCLUSION","content":"\u003cp\u003eThis paper introduced MSEF-Cancer, a multi-representation stacked ensemble framework for transparent anticancer virtual screening. The system combines six different learning systems which include ChemBERTa and 1D CNN and GraphConv and FNN and Random Forest and XGBoost to perform temperature-calibrated soft voting while handling class imbalance through its four-layer system which begins with class weighting and proceeds through focal loss and SMOTE oversampling and ends with MCC-optimized threshold calibration.\u003c/p\u003e \u003cp\u003eThe research used multiple benchmarks from CHEMBL-AC NCI-60 and PubChem AID 1259313 together with Tox21 NR-AhR data to test MSEF-Cancer performance against all other established baseline systems and the new GROVER and Uni-Mol pre-trained models. The proposed framework achieved ROC-AUC 0.946 and PR-AUC 0.662 results on stratified splits which surpassed Uni-Mol performance by 1.5 and 4.3 percentage points. MSEF-Cancer achieved an ROC-AUC of 0.919 with scaffold-based splitting while showing the least performance decline (ΔROC-AUC\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;0.027) compared to all other tested models which demonstrated strong abilities to generalize beyond their training data. The paired t-tests conducted across five folds demonstrated that all enhancements achieved statistical significance at α\u0026thinsp;=\u0026thinsp;0.05 level.\u003c/p\u003e \u003cp\u003eThe progressive imbalance mitigation ablation showed that each component contributes additive improvement, with the full configuration reducing the false negative rate from 0.36 to 0.22\u0026mdash;a 14-point reduction that directly translates to fewer missed active candidates in prospective screening campaigns. The interpretability analysis on the NCI-60 test split confirmed that MSEF-Cancer produces more faithful and chemically grounded attributions than SHAP, LIME, ChemBERTa, and AttentiveFP across all six evaluated metrics.\u003c/p\u003e \u003cp\u003eSeveral directions remain open for future investigation. Incorporation of 3D conformational information through equivariant graph networks offers an additional representation modality that may further improve active-compound recall. Multi-task learning extensions, leveraging shared molecular representations to simultaneously predict multiple target or toxicity endpoints, are projected to yield PR-AUC improvements of 5\u0026ndash;9% on Tox21 benchmarks. Finally, prospective validation through experimental confirmation of top-ranked predictions is planned to establish practical utility in an end-to-end drug discovery pipeline.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eBoth Author have same equal contribution\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eHafsath CA, Jereesh AS (2024) A stacked ensemble approach for enhancing anti-cancer drug synergy prediction. Procedia Comput Sci 235:2567\u0026ndash;2576\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGarai S, Thomas J, Dey P, Das D (2024) LGBM-ACp: An ensemble model for anticancer peptide prediction and in silico screening. Mol Diversity 28(4):1965\u0026ndash;1981\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDas AP, Sharma R, Agarwal SM (2025) Identification of rheochrysin as a potential anti-cancer inhibitor through ensemble virtual screening and molecular dynamics. Int J Biol Macromol 307:141111\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHao Z, Zhan J, Fang Y et al (2025) GNNSynergy: A multi-view graph neural network for predicting anti-cancer drug synergy. IEEE Transactions on Computational Biology and Bioinformatics\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbhang MVK, Gunjal BL (2025) Deep graph ensemble convolutional neural networks for drug response prediction. COMPUTER, 25(4)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLin TE, Yen D, HuangFu W-C et al (2024) An ensemble machine learning model generates a focused screening library for CDK8 inhibitors. Protein Sci, 33(6), e5007\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYadav AK, Kim JM (2025) ACLPred: An explainable ensemble model for anticancer ligand prediction. Sci Rep 15(1):31268\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDuo L, Liu Y, Ren J et al (2024) Artificial intelligence for small molecule anticancer drug discovery. Expert Opin Drug Discov 19(8):933\u0026ndash;948\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNussinov R, Jang H (2024) The value of protein allostery in rational anticancer drug design. Expert Opin Drug Discov 19(9):1071\u0026ndash;1085\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJana DLF, Kandhari H et al (2024) Optimizing drug-target interaction predictions using machine learning. In Proceedings of ICRASET (pp. 1\u0026ndash;6). IEEE\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMahema S, Roshni J, Raman J et al (2024) Multi-scale computational investigation to repurpose anti-cancer drugs for endometrial cancer. Cell Biochem Biophys 82(4):3367\u0026ndash;3381\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXiao F, Ding X, Shi Y et al (2024) Application of ensemble learning for predicting GABAA receptor agonists. Comput Biol Med 169:107958\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBian J, Liu X, Dong G et al (2024) ACP-ML: A sequence-based method for anticancer peptide prediction. Comput Biol Med 170:108063\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMahendran R et al (2025) Advancing plant-based anti-cancer drug discovery through hybrid ensemble models. In Proceedings of ICDSIS (pp. 1\u0026ndash;6). IEEE\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAl-Fahad D et al (2025) Virtual screening and molecular dynamics simulation of natural compounds as kinase inhibitors. Mol Diversity 29(2):1525\u0026ndash;1539\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhalid M et al (2025) Reinventing PARP1 inhibition through virtual screening and molecular dynamics. Journal of Biomolecular Structure and Dynamics\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDurgawale TP et al (2025) Phytochemical-based drug discovery for breast cancer. Chem Biodivers, 22(6), e202402864\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShyam P (2025) In silico strategies for cancer model development and anticancer drug testing. Preclinical Cancer Models for Translational Research and Drug Development. Springer, pp 153\u0026ndash;168\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVeemaraj E, Lincy A (2023) Advancing ovarian cancer diagnosis with attention-based models and 3D CNNs. ITEGAM-JETIA 9(43):23\u0026ndash;33\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIsaac AXVM (2023) A.J. Analyzing DNA pattern matching through string similarity in cancer data. In Proceedings of ICSCNA (pp. 1373\u0026ndash;1381)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEbenezer V, Edwin EB, Rajan JJ et al (2023) Automating MRI-based ovarian cancer diagnosis with DCNN. In Proceedings of ICSCNA (pp. 1353\u0026ndash;1360)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVidhya K et al (2025) MedFuseNet: Fusion of multi-modal data for improved cervical cancer diagnosis. In Proceedings of IDCIoT. IEEE\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRong Y, Bian Y, Xu T et al (2020) Self-supervised graph transformer on large-scale molecular data. Adv Neural Inform Process Syst (NeurIPS) 33:12165\u0026ndash;12175\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou G, Gao Z, Ding Q et al (2023) Uni-Mol: A universal 3D molecular representation learning framework. In Proceedings of ICLR 2023\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"journal-of-computer-aided-molecular-design","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jcam","sideBox":"Learn more about [Journal of Computer-Aided Molecular Design](http://link.springer.com/journal/10822)","snPcode":"10822","submissionUrl":"https://submission.nature.com/new-submission/10822/3","title":"Journal of Computer-Aided Molecular Design","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Anticancer drug discovery, Molecular ensemble learning, ChemBERTa, Graph neural networks, Explainability, Class imbalance, Virtual screening, Scaffold splitting, GROVER, Uni-Mol","lastPublishedDoi":"10.21203/rs.3.rs-9387094/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9387094/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe primary mission of computational drug discovery focuses on speeding up the process of discovering effective cancer-fighting drugs. The current AI-powered screening systems encounter major problems because they rely on incorrect class distribution ratios and need to use only one type of molecular representation and their models lack proper visibility. The research presents MSEF-Cancer as an ensemble framework which combines five different approaches: transformer-based SMILES encoders (ChemBERTa) graph convolutional networks (GraphConv) one-dimensional convolutional neural networks and a Feedforward Neural Network (FNN) and classical tree-based classifiers. The system uses temperature-calibrated soft voting for its integration but focal loss training and synthetic minority oversampling (SMOTE) and Matthews Correlation Coefficient (MCC)-based threshold optimization serve as methods for treating data imbalance. The framework underwent evaluation through testing on three chemogenomics datasets (CHEMBL-AC, NCI-60, PubChem AID 1259313) and the Tox21 NR-AhR assay which served as a separate benchmark for assessing generalization. MSEF-Cancer achieved a ROC-AUC score of 0.946 and a PR-AUC score of 0.662 through stratified splits while surpassing both GROVER (ROC-AUC 0.921) and Uni-Mol (ROC-AUC 0.931). MSEF-Cancer achieved a ROC-AUC score of 0.919 after using scaffold-based splitting which represented the least performance loss among all tested models. The paired t-tests which examined five folds established that all baseline improvements reached statistical significance at p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 level. The analysis of attribution through gradient methods verified that this method provides better interpretability because it achieved the lowest deletion AUC score of 0.371 with the highest insertion AUC score of 0.639 and the strongest structural alert coverage score of 0.281.\u003c/p\u003e","manuscriptTitle":"MSEF-Cancer: A Multi-Representation Stacked Ensemble Framework for Transparent Anticancer Virtual Screening","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-30 15:06:48","doi":"10.21203/rs.3.rs-9387094/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewersInvited","content":"","date":"2026-04-22T00:13:06+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-04-22T00:11:22+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-04-15T07:52:52+00:00","index":"","fulltext":""},{"type":"submitted","content":"Journal of Computer-Aided Molecular Design","date":"2026-04-11T10:34:16+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"journal-of-computer-aided-molecular-design","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jcam","sideBox":"Learn more about [Journal of Computer-Aided Molecular Design](http://link.springer.com/journal/10822)","snPcode":"10822","submissionUrl":"https://submission.nature.com/new-submission/10822/3","title":"Journal of Computer-Aided Molecular Design","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"af8da3f5-d362-402a-b312-19b5b36d171c","owner":[],"postedDate":"April 30th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-04-30T15:06:49+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-30 15:06:48","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9387094","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9387094","identity":"rs-9387094","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.