Intelligent QSAR Approaches: Harnessing Machine Learning for Early Detection of Carcinogenic Agents

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Traditional carcinogenicity testing methods are costly and time-consuming, while existing Quantitative Structure-Activity Relationship (QSAR) models suffer from low accuracy and limited applicability domains. This study addresses these limitations by developing enhanced QSAR models using machine learning (ML) algorithms. A dataset of 805 compounds from the Carcinogenic Potency Database (CPD) was used to train classification and regression models, employing Bayesian classifiers, recursive partitioning, Kernel-based Partial Least Squares (KPLS), and deep learning techniques (Neural Networks, Random Forests). An independent validation dataset (105 compounds) was used to assess model performance. The DeepChem-based classification model achieved 81% test accuracy and 72% external validation accuracy, while the AutoQSAR regression model demonstrated an R² of 0.58 and Q² of 0.51, outperforming existing literature benchmarks. These models exhibit broad chemical space coverage, offering a robust, cost-effective alternative for carcinogenicity prediction.
Full text 96,950 characters · extracted from preprint-html · click to expand
Intelligent QSAR Approaches: Harnessing Machine Learning for Early Detection of Carcinogenic Agents | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Intelligent QSAR Approaches: Harnessing Machine Learning for Early Detection of Carcinogenic Agents Roy Tatenda Bisenti, Tanaka Denzel Chitsa, Amos Misi, Dr Paul Mushonga, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7438250/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Traditional carcinogenicity testing methods are costly and time-consuming, while existing Quantitative Structure-Activity Relationship (QSAR) models suffer from low accuracy and limited applicability domains. This study addresses these limitations by developing enhanced QSAR models using machine learning (ML) algorithms. A dataset of 805 compounds from the Carcinogenic Potency Database (CPD) was used to train classification and regression models, employing Bayesian classifiers, recursive partitioning, Kernel-based Partial Least Squares (KPLS), and deep learning techniques (Neural Networks, Random Forests). An independent validation dataset (105 compounds) was used to assess model performance. The DeepChem-based classification model achieved 81% test accuracy and 72% external validation accuracy, while the AutoQSAR regression model demonstrated an R² of 0.58 and Q² of 0.51, outperforming existing literature benchmarks. These models exhibit broad chemical space coverage, offering a robust, cost-effective alternative for carcinogenicity prediction. Computational Chemistry Medicinal Chemistry QSAR machine learning carcinogenicity deep learning predictive toxicology Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1. Introduction Carcinogenicity testing remains a critical challenge in toxicology and drug safety assessment, particularly given the widespread presence of potentially harmful substances in our environment, from food and air to water ( 1 ). Historically, the understanding of chemical risk, especially carcinogenicity, was often simplistic, leading to concepts like the "Zero Risk Concept" where it was believed that avoiding a few synthetic chemicals would suffice ( 2 ). The development of basic and applied science has focused more on the criticisms of such assumptions because of the need for comprehensive risk evaluation. Methods like rodent bioassays are costly and controversial from an ethical standpoint, while simultaneously and devoid of evidence, falling short of reliably predicting human responses. Because of these shortcomings, such methods are lagging behind the progress that has been made in non-animal testing methodologies ( 3 ). For this case, QSAR models are seen as new computing methodologies that link a biological activity of the compound to its structure ( 4 ). From a toxicological and pharmaceutical point of view, this discipline is of great importance because it aims to estimate possible toxicity for novel or unscreened compounds exclusively from their chemical architectures ( 4 , 5 ). Such models strongly improve capability for predicting toxic compounds and are instrumental in reducing the need for extensive animal testing (4). With regard to the drug development, QSAR models have become essential in providing accurate estimations of the biological activity of a compound in its early stages of development ( 4 , 5 ). This is accomplished through the use of mathematics and statistics to devise models that predict the bio-activities or other properties of the compounds with the prerequisite that any modification done will improve the activity of the compound ( 5 ). In predictive toxicology and in drug designing activities, the development of QSAR models is of great importance. Nonetheless, there are still data issues regarding the quality, interpretability, validation, precision, or even more concerning the QSAR models in ( 6 ). The first stage in the development of QSAR models was from 1990 to 2005, where it involved the application of topology and arithmetic calculations, said to be the first phase of QSAR model development ( 7 ). A number of models built during this period achieved approximately 75% precision ( 8 , 9 ). However, a lack of external validation datasets hindered their performance, which is well illustrated by the studies of previous researchers ( 10 – 13 ). The following period from 2006 to 2015 marked new advancements due to a shift towards reliance on databases and expansion of computational tool access, leading to the creation of sophisticated models comprised of complex descriptors and non-linear methods ( 14 – 16 ). From 2015 onwards, the precision of QSAR systems greatly advanced with deep learning and the addition of complex computational frameworks ( 17 ). However, even these advanced models still exhibit limitations in accuracy and applicability, with a scarcity of robust regression models for predicting carcinogenic potency ( 3 , 6 ).This study aims to address these persistent limitations by: Developing high-accuracy machine learning-based QSAR models for both carcinogenicity classification and potency prediction. Validating model performance rigorously using an independent external dataset. Comparing the results with existing models from the literature to demonstrate superiority in predictive accuracy and applicability domain. By integrating modern machine learning techniques, including deep learning and ensemble methods, into QSAR modeling, this research seeks to offer a scalable and robust solution for regulatory and industrial applications, bridging a critical knowledge gap in the early detection of carcinogenic agents. 2. Materials and Methods This study utilized two primary datasets for model development and validation. Dataset A comprised 805 compounds obtained from the Carcinogenic Potency Database (CPD), covering the period from 1990 to 2010, with associated TD50 rat endpoints. For external validation, an independent Dataset B, consisting of 105 compounds from the CPD (2011–2023), was employed. Prior to model development, molecular structures underwent preprocessing and optimization using LigPrep (Schrödinger Suite 2023-1, OPLS4 forcefield) at a pH of 7.0 ± 2.0. This ensured the selection of the ligand form most likely to exist under physiological conditions by choosing the compound with the lowest state penalty. Subsequently, 636 descriptors, encompassing topological, physicochemical, functional group counts, and semi-empirical properties, were computed for each compound using Schrödinger’s Informatics module (Schrödinger Suite 2023-1). For model development, both classification and regression approaches were pursued using three different software modules: Build QSAR Models, AutoQSAR, and DeepChem (all part of Schrödinger Suite 2023-1). Classification models included Bayesian Classifiers, which were evaluated with and without categorical descriptors, with smoothing coefficients varied from 0.0000001 to 0.001. Recursive Partitioning models were constructed as ensembles of 100 trees, optimized using Gini impurity or information gain, and X variables with correlation less than 0.1 were discarded. Deep learning-based classification was performed using DeepChem, with neural networks trained for 50 hours, utilizing the 636 descriptors and features identified by the neural networks. Regression models involved Kernel-based Partial Least Squares (KPLS), where the number of factors was varied from 1 to 16, and the kernel non-linearity was set to 0.05. Regression models using AutoQSAR and DeepChem employed the same settings as their classification counterparts, but with a numerical TD50 rat endpoint. Model validation was conducted through a combination of internal 5-fold cross-validation and external validation using the independent Dataset B, where each generated model was rigorously scrutinized. 3. Results 3.1 Data Preparation Table 1 Results of Dataset Preparation Dataset No. of Compounds Data Label Type State Penalty Endpoint Classification A 802 Integer Lowest TD50 rat Classification B 103 Integer Lowest TD50 rat Regression A 414 Real Lowest TD50 rat Regression B 63 Real Lowest TD50 rat Table 1 summarizes the characteristics of the datasets prepared for model training and validation. "Classification A" and "Regression A" were used for training classification and regression models, respectively, while "Classification B" and "Regression B" served as independent external validation sets. The use of integer data labels for classification simplifies the task to a binary problem (carcinogenic vs. non-carcinogenic), ensuring clear and unambiguous model outputs. For regression, real (numeric) data labels enable the prediction of the quantitative amount of compound required to induce tumors. The selection of the lowest state penalty for each compound aimed to ensure that the models reflect the most physiologically relevant ligand forms. The consistent use of the TD50 rat endpoint across all datasets provides a clear and well-defined measure of activity, analogous to IC50 in drug discovery. 3.2 Descriptor Generation A total of 636 molecular descriptors were generated for every compound in both datasets (Fig. 1 ), encompassing both categorical and numerical descriptors. This comprehensive set included semi-empirical, topological, functional group counts, and physiochemical descriptors. Semi-empirical descriptors, derived from quantum mechanical calculations, provide insights into the electronic structure of molecules. Topological descriptors describe molecular size and form based on atomic interconnectedness. Functional group counts reveal specific atomic groups that may influence biological activity, while physiochemical descriptors capture physical and chemical characteristics such as reactivity and solubility. This broad range of descriptors aligns with advanced intelligent QSAR development, leveraging extensive molecular characteristics to enable the model to learn complex relationships between structure and carcinogenic potency, thereby enhancing predictive power and versatility. 3.3 Model Training The model training phase involved the development of both classification and regression models using various machine learning algorithms. To begin, we started with Bayesian Classification models. With Model A, which lacked categorical descriptors, we achieved 60% accuracy on training data and 58% on testing data. Conversely, Model B which included categorical descriptors performed even worse as it achieved 56% accuracy on training data and 53% on testing data. These results as displayed in Fig. 2 suggest that the categorical descriptors either added too much noise or complexity to the model or that the dataset was poorly suited for the Bayesian method due to its heavy reliance on the independence assumption between descriptors (which is nearly always wrong in QSAR models, ( 18 )). Additionally, the high dimensionality of 636 descriptors would likely harm Bayesian models, resulting in critical information being lost due to excessive discretization of continuous variables and over fitting. Recursive Partitioning Models were created and seven models were analyzed in total. Table 2 illustrates that Models A-D were single models, whereas E-G were 100 tree ensemble models. Ensemble models E-G outperformed their single counterparts A-D in both training and test accuracy. This supports the observation that single models have limitations with respect to capturing intricate relationships within the data. Take, for example, Model E, an ensemble model that achieved test accuracy of 67% by applying Gini impurity while discarding weakly correlated variables. This model surpassed single Model B, which achieved only 53% accuracy with the use of information gain. As a rule, models built using Gini impurity performed better than those using information gain. Perhaps this is because the former is a weaker discriminator, and therefore less sensitive to fluctuations in class probabilities ( 19 ). Furthermore, the accuracy gain from skipping low correlation variables is likely a result from less dimensionality making the noise to signal ratio favor signal. Table 2 Recursive Partitioning Results Model Model Type Algorithm Discard Variables with low correlations Training Accuracy Test Accuracy A Single Gini Impurity no 88 56 B Single Information Gain no 89 53 C Single Information Gain yes 89 54 D Single Gini Impurity yes 87 62 E Ensemble Gini Impurity yes 98 67 F Ensemble Information Gain yes 98 63 G Ensemble Gini Impurity yes 95 64 Additional classification models were developed using AutoQSAR (Fig. 3 ), and all top 10 models employed the Bayes Naive algorithm. The training accuracies for these models ranged from 69–89%, while test accuracies were lower, ranging from 64–80%. One of the models achieving a training accuracy of 89% utilized dendritic fingerprint descriptors, but the 73% test accuracy suggested over fitting. Conversely, the model driven by informatics tools performed better in terms of generalization, yielding 80% test accuracy while only achieving 77% training accuracy. This underscores the importance of independent validation to assess an algorithm's predictive accuracy ( 20 ). The classification was done using DeepChem, which gave a high-performing model. The analysis from the Receiver Operating Characteristic (ROC) curve yielded an Area Under the Curve 0.81 (AUC-ROC) ( Fig. 4 ). An AUC-ROC of 0.81 indicates that there is an 81% chance the model accurately assesses whether a specific compound is carcinogenic or not, which confirms excellent predictive accuracy and robustness ( 21 ). From the high curvature of the ROC, not only is the AUC good, but the model also possesses a high true positive rate and a low false positive rate across many decision-making thresholds which is rare and usually very hard to find. This is due to DeepChem using neural networks on molecular information which uncovers intricate relationships and far outperforms traditional machine learning methods. Kernel-based Partial Least Squares (KPLS) was studied for its application in regression modeling. A total of four models were created. Despite achieving a respectable R² of 0.84 in Model A, it suffered from considerable over fitting as indicated by a severely negative Q² of -377. Predictive performance was significantly lower than that observed in the training data. Model D demonstrated the lowest R² of 0.08 among all KPLS models, but was notable for having the highest Q² of -0.36. These results (Fig. 5 ) illustrate that while PLS regression is a potent regression model, its performance can be greatly overshadowed by how the model is set up. In this case, all KPLS models failed to reach any acceptable benchmarks for regulatory predictive performance. In contrast, the performance of the AutoQSAR regression model in estimating the potency of carcinogenesis was considerably higher. It was built using 427 compounds and was tailored for 21 different types of tumors, yielding an R² of 0.58 and a Q² of 0.51 (Fig. 6 ). An R² of 0.58 means that the model predicts approximately 58% of the variability in carcinogenic potency, while Q² of 0.51 suggests that the model predicts about 51% of the variance in new data ( 20 ). Together, these values suggest a reasonable and balanced predictive capability, particularly given the additional difficulty posed by the numerous tumor types included in the dataset. The accuracy of AutoQSAR’s results demonstrated the effectiveness of its more sophisticated algorithms and its extensive computations of molecular descriptors that integrate fundamental synthetic molecular characteristics required for precise predictions. 3.4 Model Comparisons Table 3 Model Details Authors Year Model Type Endpoint Applicability Domain Model Codename Laboratory of Comparative Toxicology and Ecotoxicology 1991 Simple Equations TD/50 RAT Low, Limited to a small chemical space 1A Laboratory of Comparative Toxicology and Ecotoxicology 2002 Simple Equations TD/50 RAT Aromatic Amines 2B FDA's Center for Drug Evaluation and Research 2003 Statistical Model TD/50 RAT Wide Chemical Space 3C Contrera 2005 Statistical Model TD/50 RAT Pharmaceuticals 4D University of Porto 2008 Genetic Algorithm TD/50 RAT Nitroso Compounds 6F REACH 2010 Machine Learning TD/50 RAT Wide Chemical Space 7G Laboratory of Comparative Toxicology 2021 Deep Learning TD/50 RAT Wide Chemical Space 8H FDA 2021 Deep Learning TD/50 RAT Wide Chemical Space 9I Table 3 provides an overview of various QSAR models from the literature, detailing their authors, year of development, model type, endpoint, and applicability domain. This table serves as a reference for comparing the performance of the models developed in the current study against existing benchmarks. Table 4 Classification Models Comparison Model Training Accuracy Test Accuracy External Validation Accuracy 1991 72 65 2002 91 72 2003 5 73 2005 83 80 2010 91 73 66 2021 81 76 2021 91 77 BAYESA 58 58 58 RPB 95 64 69 AQA 83 72 62 DP 81 72 Table 4 presents a comparative analysis of various classification models, including those from the literature and the models developed in this study (BayesA, RPB, AQA, DP). The metrics include training accuracy, test accuracy, and external validation accuracy where available. This table allows for a direct comparison of how well each model performed in classifying compounds as carcinogenic or non-carcinogenic. Table 5 Regression Model Comparisons Model R^2 Q^2 2008 0.808 0.248 2021 0.756 PLSA 0.08 -0.36 PLSD 0.84 -377 AQ 0.58 0.51 DPC 0.02 Table 5 compares the performance of regression models for predicting carcinogenic potency, including those from the literature and the models developed in this study (PLSA, PLSD, AQ, and DPC). The key metrics for comparison are R² (coefficient of determination) and Q² (goodness of prediction), which indicate how well the model fits the training data and predicts new, unseen data, respectively. 4. Discussion The findings of this study demonstrate significant advancements in carcinogenicity prediction through the application of machine learning-enhanced QSAR models, particularly when compared to existing literature. The DeepChem classification model, achieving an 81% test accuracy and 72% external validation accuracy, notably surpasses the performance of earlier models. For instance, the FDA's 2005 model, while showing 80% training accuracy, yielded only 76% test accuracy ( 8 , 12 ). Similarly, the 2010 OECD model, a significant regulatory advancement, reported 73% accuracy on its test set with 66% external validation ( 14 ). The DeepChem model's superior performance can be attributed to its deep learning architecture, which is adept at capturing complex, non-linear relationships within molecular data that traditional QSAR methods might miss ( 17 , 21 ). This aligns with the broader trend in advanced QSAR models (2015–2023) that leverage deep learning to transition from traditional descriptor engineering to more sophisticated molecular embeddings, thereby enhancing prediction precision ( 17 ). Furthermore, the models developed in this study exhibit a broad applicability domain, covering 21 different tumor types. This is a crucial advancement over earlier niche approaches, such as the 2002 model specifically designed for aromatic amines ( 7 ) or the 2008 model focused solely on nitroso compounds ( 15 ). The comprehensive validation rigor, including the use of an independent external dataset, ensured the robustness of our models and addressed the common risk of over fitting observed in some prior studies ( 7 ). In terms of regression, the AutoQSAR model demonstrated a balanced and reliable performance with an R² of 0.58 and a Q² of 0.51. This represents a substantial improvement over earlier potency prediction models, such as the 2008 nitroso-compound model which reported a Q² of only 0.248 ( 15 ). The ability of AutoQSAR to achieve a Q² greater than 0.5 makes it suitable for regulatory applications, as it indicates a good predictive capability on unseen data ( 20 ). While other regression models like KPLS showed high R² values, their significantly negative Q² values indicated severe over fitting and poor generalization, consistent with challenges in developing robust regression models for carcinogenicity ( 6 ). The integration of diverse machine learning algorithms within AutoQSAR likely contributed to its ability to capture the underlying relationships across varied chemical structures and tumor types, a complexity that often limits the performance of simpler models. Despite these advancements, certain limitations warrant consideration. The primary dataset for this study was derived from rodent assays. While these are standard in carcinogenicity testing, the direct translation of findings to human toxicity can be limited due to species-specific differences in metabolism and response ( 3 ). Future research should explore expanding training data to include human-relevant endpoints to enhance direct applicability. Additionally, model performance is inherently dependent on the quality and selection of descriptors. While a comprehensive set of 636 descriptors was generated, future work could investigate the utility of more advanced, graph-based embeddings or other novel descriptor generation techniques to potentially capture more nuanced molecular features ( 22 ). 5. Conclusion This study successfully developed and validated machine learning-enhanced QSAR models for the early detection of carcinogenic agents, addressing key limitations of traditional methods and existing QSAR approaches. All defined objectives—data preparation, molecular descriptor generation, model building, and rigorous validation—were successfully met. The DeepChem classification model demonstrated superior performance, achieving an 81% test accuracy and a robust 72% external validation accuracy, showcasing its strong generalization capabilities. Concurrently, the AutoQSAR regression model provided reliable potency prediction with an R² of 0.58 and a Q² of 0.51, outperforming previous benchmarks and indicating its suitability for regulatory applications. Together, these models represent a pivotal step toward the application of advanced computational methods for the accurate and efficient testing of carcinogenicity. The growing use of advanced deep learning models demonstrates their potential in enhancing predictive toxicology tasks. These models comply with the OECD Regulations pertinent to QSAR Validation owing to their suitability for use in risk assessment. Further work in this area aims to improve model translatability by supplementing the training datasets with relevant human-exposure endpoints. Further, the application of XAI methods could reveal the mechanisms of predictions made which, in turn, would bolster the confidence and understanding in using these intelligent QSAR systems. Declarations CRediT author statement Tanaka Denzel Chitsa : Conceptualization, Methodology, Investigation, Writing- Original draft, Visualization, Formal analysis Amos Misi : Data curation, Supervision, Validation, Investigation, Project administration, Resources, Writing- Reviewing and Editing Dr Paul Mushonga, Dr Albert Wakandigara : Visualization, Investigation, Supervision , Writing- Reviewing and Editing Roy T Bisenti : Formal analysis, Writing- Reviewing and Editing, Visualization Declaration of interests The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. References Jacobs MN, Colacci A, Corvi R, Vaccari M, Aguila MC, Corvaro M et al (2020) Chemical carcinogen safety testing: OECD expert group international consensus on the development of an integrated approach for the testing and assessment of chemical non-genotoxic carcinogens. Arch Toxicol 94(8):2899–2923 Jain S A Comprehensive Study for The Regulations of Genotoxic Impurities in Pharmaceutical Drug Substances & Products with Special Reference to Their Detection, Identification and Evaluatione. Nirmauni.ac.in [Internet]. 2016 [cited 2025 Jan 1]; Available from: https://repository.nirmauni.ac.in/jspui/handle/123456789/6541 Chung E, Russo DP, Ciallella HL, Wang YT, Wu M, Aleksunes LM et al (2023) Data-driven quantitative structure–activity relationship modeling for human carcinogenicity by chronic oral exposure. Environ Sci Technol 57(16):6573–6588 Cherkasov A, Muratov EN, Fourches D, Varnek A, Baskin II, Cronin M et al (2014) QSAR modeling: where have you been? Where are you going to? J Med Chem 57(12):4977–5010 Roy K, Kar S, Das RN (2015) Understanding the basics of QSAR for applications in pharmaceutical sciences and risk assessment. Academic Toma C, Manganaro A, Raitano G, Marzo M, Gadaleta D, Baderna D et al (2020) QSAR models for human carcinogenicity: an assessment based on oral and inhalation slope factors. Molecules 26(1):127 Benigni R, Bossa C (2011) Flexible use of qsar models in predictive toxicology: A case study on aromatic amines. Environ Mol Mutagen 53(1):62–69 Halder AK, Moura AS, Cordeiro MNDS (2018) QSAR modelling: a therapeutic patent review 2010-present. Expert Opin Ther Pat 28(6):467–476 Mwanza C, Zhang WZ, Mulenga Kalulu, Ding SN (2024) Advancing green chemistry in environmental monitoring: the role of electropolymerized molecularly imprinted polymer-based electrochemical sensors. Green Chem Chen M, Bisgin H, Tong L, Hong H, Fang H, Borlak J et al (2014) Toward predictive models for drug-induced liver injury in humans: are we there yet? Biomark Med 8(2):201–213 Dargan S, Kumar M, Ayyagari MR, Kumar G (2019) A Survey of Deep Learning and Its Applications: A New Paradigm to Machine Learning. Arch Comput Methods Eng. ;27(4) Toropova AP, Toropov AA (2018) CORAL: QSAR models for carcinogenicity of organic compounds for male and female rats. Comput Biol Chem 72:26–32 Nag S, Baidya ATK, Mandal A, Mathew AT, Das B, Devi B et al (2022) Deep learning tools for advancing drug discovery and development. 3 Biotech. ;12(5) Fjodorova N, Vračko M, Novič M, Roncaglioni A, Benfenati E (2010) New public QSAR model for carcinogenicity. Chem Cent J 4(Suppl 1):S3 Kostal J, Voutchkova-Kostal A (2023) Quantum-Mechanical Approach to Predicting the Carcinogenic Potency of N-Nitroso Impurities in Pharmaceuticals. Chem Res Toxicol 36(2):291–304 Zuang V, Dura A, Bofill A, Leite B, Berggren E, Bernasconi C et al (2019) EURL ECVAM status report on the development, validation and regulatory acceptance of alternative methods and approaches Publications Office of the European Union Luxembourg; 2020 Tropsha A, Isayev O, Varnek A, Schneider G, Cherkasov A (2024) Integrating QSAR modelling and deep learning in drug discovery: the emergence of deep QSAR. Nat Rev Drug Discov 23(2):141–155 Du CJ, He HJ, Sun DW (2016) Object classification methods. Computer vision technology for food quality evaluation. Elsevier, pp 87–110 Dutta P, Paul S, Kumar A (2021) Comparative analysis of various supervised machine learning techniques for diagnosis of COVID-19. Electronic devices, circuits, and systems for biomedical applications. Elsevier, pp 521–540 Ciallella HL, Chung E, Russo DP, Zhu H (2022) Automatic quantitative structure–activity relationship modeling to fill data gaps in high-throughput screening. High-Throughput Screening Assays in Toxicology. Springer, pp 169–187 Dixon SL, Duan J, Smith E, Von Bargen CD, Sherman W, Repasky MP (2016) AutoQSAR: an automated machine learning tool for best-practice quantitative structure–activity relationship modeling. Future Med Chem 8(15):1825–1839 Bahia MS, Kaspi O, Touitou M, Binayev I, Dhail S, Spiegel J et al (2023) A comparison between 2D and 3D descriptors in QSAR modeling based on bio-active conformations. Mol Inf 42(4):2200186 Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7438250","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":504443525,"identity":"be9531c8-8b25-4e5a-8ba3-3dec9909a615","order_by":0,"name":"Roy Tatenda Bisenti","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABHElEQVRIie3SsUrEMBjA8S8Ecg4B18B57St8R8FD9GFSCrrcUBfBpSqB3CJ0reBD1CewEOgtxbngcnIv0MPlhBtMBqdroaND/kMDoT/6pQTA5/unUUC3kKeNez6OJ4woBOkIHUNcDJgYRRarZ6RpmgWLUOn7n72Z5dPcdERncLqq6Dc/JmdNg7RAE11ooj+5NNHLq2GCaAaikRD1ECGWSDlWcVlbAtLEZZswINq+2wIky0GSPThyu7fkvU2oHUxAaIkZJlSiJcDdV0QCdjAEbIGoPsLr1NizzMs6VlN+fRMVbXIu5Ifk8yZW9NBDJuptmx6yEM36a7e/upzlRbzturssCNbG7Iq+33yC1fGmBODDN2Gy6d/3+Xw+31+/0RVhWZqtB7EAAAAASUVORK5CYII=","orcid":"https://orcid.org/0009-0000-8970-4810","institution":"University of Zimbabwe","correspondingAuthor":true,"prefix":"","firstName":"Roy","middleName":"Tatenda","lastName":"Bisenti","suffix":""},{"id":504443526,"identity":"c1da4efc-49aa-404a-bf6a-43f94ce79535","order_by":1,"name":"Tanaka Denzel Chitsa","email":"","orcid":"","institution":"University of Zimbabwe","correspondingAuthor":false,"prefix":"","firstName":"Tanaka","middleName":"Denzel","lastName":"Chitsa","suffix":""},{"id":504443527,"identity":"aef7da03-3f2a-45fc-a044-cfe0367b76cf","order_by":2,"name":"Amos Misi","email":"","orcid":"","institution":"University of Zimbabwe","correspondingAuthor":false,"prefix":"","firstName":"Amos","middleName":"","lastName":"Misi","suffix":""},{"id":504443528,"identity":"5e27b9ba-3068-4698-bf32-4853dd02087f","order_by":3,"name":"Dr Paul Mushonga","email":"","orcid":"","institution":"University of Zimbabwe","correspondingAuthor":false,"prefix":"Dr","firstName":"Paul","middleName":"","lastName":"Mushonga","suffix":""},{"id":504443529,"identity":"be03ad3a-e7e0-440f-92b8-7922395d330a","order_by":4,"name":"Dr Albert Wakandigara","email":"","orcid":"","institution":"University of Zimbabwe","correspondingAuthor":false,"prefix":"Dr","firstName":"Albert","middleName":"","lastName":"Wakandigara","suffix":""}],"badges":[],"createdAt":"2025-08-23 03:30:22","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-7438250/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7438250/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":89902464,"identity":"b147619e-5b38-41d6-b256-6ae3f0309743","added_by":"auto","created_at":"2025-08-26 09:30:31","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":34844,"visible":true,"origin":"","legend":"\u003cp\u003eDescriptors Generation\u0026nbsp;\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-7438250/v1/fa7d44bdc9397c968c05e7a7.png"},{"id":89902462,"identity":"0bbbcc1f-2733-481f-b5cf-15c14995bfd9","added_by":"auto","created_at":"2025-08-26 09:30:31","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":21031,"visible":true,"origin":"","legend":"\u003cp\u003eBayes Model Results\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-7438250/v1/91fd8e3ee5f0c4d5b7f9a68b.png"},{"id":89904097,"identity":"47183ad1-1e41-4ed3-b397-8e0c13878bf7","added_by":"auto","created_at":"2025-08-26 09:46:31","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":179816,"visible":true,"origin":"","legend":"\u003cp\u003eAutoQsar Model Results\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-7438250/v1/7ffa1b2e21b791385161acb7.png"},{"id":89902467,"identity":"8328498c-1e02-460f-9911-38364b0f59b7","added_by":"auto","created_at":"2025-08-26 09:30:31","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":56219,"visible":true,"origin":"","legend":"\u003cp\u003eDeepChem Classification Model Results\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-7438250/v1/4fc5b5f8d87522e9687c964a.png"},{"id":89902468,"identity":"bc741d33-a4bc-4cdf-994a-7f5f8d825848","added_by":"auto","created_at":"2025-08-26 09:30:31","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":49762,"visible":true,"origin":"","legend":"\u003cp\u003eKPLS Model Results\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-7438250/v1/338bcb405114ef73ddfb034a.png"},{"id":89902471,"identity":"80defa7b-9eac-4fb2-b92e-66b86221de15","added_by":"auto","created_at":"2025-08-26 09:30:31","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":100164,"visible":true,"origin":"","legend":"\u003cp\u003eAutoQsar Regression Model Results\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-7438250/v1/a0b9844f9b7598d44bed1bba.png"},{"id":89905032,"identity":"7409b8ea-41f0-49f9-b926-5af2f692733b","added_by":"auto","created_at":"2025-08-26 09:54:31","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1053744,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7438250/v1/d052fa3f-f9ae-4b52-a4ea-80ed4b42c7cd.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e Intelligent QSAR Approaches: Harnessing Machine Learning for Early Detection of Carcinogenic Agents\u003c/p\u003e","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eCarcinogenicity testing remains a critical challenge in toxicology and drug safety assessment, particularly given the widespread presence of potentially harmful substances in our environment, from food and air to water (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e). Historically, the understanding of chemical risk, especially carcinogenicity, was often simplistic, leading to concepts like the \"Zero Risk Concept\" where it was believed that avoiding a few synthetic chemicals would suffice (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). The development of basic and applied science has focused more on the criticisms of such assumptions because of the need for comprehensive risk evaluation. Methods like rodent bioassays are costly and controversial from an ethical standpoint, while simultaneously and devoid of evidence, falling short of reliably predicting human responses. Because of these shortcomings, such methods are lagging behind the progress that has been made in non-animal testing methodologies (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eFor this case, QSAR models are seen as new computing methodologies that link a biological activity of the compound to its structure (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e). From a toxicological and pharmaceutical point of view, this discipline is of great importance because it aims to estimate possible toxicity for novel or unscreened compounds exclusively from their chemical architectures (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e). Such models strongly improve capability for predicting toxic compounds and are instrumental in reducing the need for extensive animal testing (4). With regard to the drug development, QSAR models have become essential in providing accurate estimations of the biological activity of a compound in its early stages of development (\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e). This is accomplished through the use of mathematics and statistics to devise models that predict the bio-activities or other properties of the compounds with the prerequisite that any modification done will improve the activity of the compound (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eIn predictive toxicology and in drug designing activities, the development of QSAR models is of great importance. Nonetheless, there are still data issues regarding the quality, interpretability, validation, precision, or even more concerning the QSAR models in (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e). The first stage in the development of QSAR models was from 1990 to 2005, where it involved the application of topology and arithmetic calculations, said to be the first phase of QSAR model development (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e). A number of models built during this period achieved approximately 75% precision (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e). However, a lack of external validation datasets hindered their performance, which is well illustrated by the studies of previous researchers (\u003cspan additionalcitationids=\"CR11 CR12\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). The following period from 2006 to 2015 marked new advancements due to a shift towards reliance on databases and expansion of computational tool access, leading to the creation of sophisticated models comprised of complex descriptors and non-linear methods (\u003cspan additionalcitationids=\"CR15\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e). From 2015 onwards, the precision of QSAR systems greatly advanced with deep learning and the addition of complex computational frameworks (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e). However, even these advanced models still exhibit limitations in accuracy and applicability, with a scarcity of robust regression models for predicting carcinogenic potency (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e).This study aims to address these persistent limitations by:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eDeveloping high-accuracy machine learning-based QSAR models for both carcinogenicity classification and potency prediction.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eValidating model performance rigorously using an independent external dataset.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eComparing the results with existing models from the literature to demonstrate superiority in predictive accuracy and applicability domain.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eBy integrating modern machine learning techniques, including deep learning and ensemble methods, into QSAR modeling, this research seeks to offer a scalable and robust solution for regulatory and industrial applications, bridging a critical knowledge gap in the early detection of carcinogenic agents.\u003c/p\u003e"},{"header":"2. Materials and Methods","content":"\u003cp\u003eThis study utilized two primary datasets for model development and validation. Dataset A comprised 805 compounds obtained from the Carcinogenic Potency Database (CPD), covering the period from 1990 to 2010, with associated TD50 rat endpoints. For external validation, an independent Dataset B, consisting of 105 compounds from the CPD (2011\u0026ndash;2023), was employed. Prior to model development, molecular structures underwent preprocessing and optimization using LigPrep (Schr\u0026ouml;dinger Suite 2023-1, OPLS4 forcefield) at a pH of 7.0\u0026thinsp;\u0026plusmn;\u0026thinsp;2.0. This ensured the selection of the ligand form most likely to exist under physiological conditions by choosing the compound with the lowest state penalty. Subsequently, 636 descriptors, encompassing topological, physicochemical, functional group counts, and semi-empirical properties, were computed for each compound using Schr\u0026ouml;dinger\u0026rsquo;s Informatics module (Schr\u0026ouml;dinger Suite 2023-1).\u003c/p\u003e\u003cp\u003eFor model development, both classification and regression approaches were pursued using three different software modules: Build QSAR Models, AutoQSAR, and DeepChem (all part of Schr\u0026ouml;dinger Suite 2023-1). Classification models included Bayesian Classifiers, which were evaluated with and without categorical descriptors, with smoothing coefficients varied from 0.0000001 to 0.001. Recursive Partitioning models were constructed as ensembles of 100 trees, optimized using Gini impurity or information gain, and X variables with correlation less than 0.1 were discarded. Deep learning-based classification was performed using DeepChem, with neural networks trained for 50 hours, utilizing the 636 descriptors and features identified by the neural networks. Regression models involved Kernel-based Partial Least Squares (KPLS), where the number of factors was varied from 1 to 16, and the kernel non-linearity was set to 0.05. Regression models using AutoQSAR and DeepChem employed the same settings as their classification counterparts, but with a numerical TD50 rat endpoint. Model validation was conducted through a combination of internal 5-fold cross-validation and external validation using the independent Dataset B, where each generated model was rigorously scrutinized.\u003c/p\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003ch2\u003e3.1 Data Preparation\u003c/h2\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eResults of Dataset Preparation\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDataset\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNo. of Compounds\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eData Label Type\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eState Penalty\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eEndpoint\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eClassification A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e802\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eInteger\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eLowest\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eTD50 rat\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eClassification B\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e103\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eInteger\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eLowest\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eTD50 rat\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRegression A\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e414\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eLowest\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eTD50 rat\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRegression B\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e63\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eReal\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eLowest\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eTD50 rat\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e summarizes the characteristics of the datasets prepared for model training and validation. \"Classification A\" and \"Regression A\" were used for training classification and regression models, respectively, while \"Classification B\" and \"Regression B\" served as independent external validation sets. The use of integer data labels for classification simplifies the task to a binary problem (carcinogenic vs. non-carcinogenic), ensuring clear and unambiguous model outputs. For regression, real (numeric) data labels enable the prediction of the quantitative amount of compound required to induce tumors. The selection of the lowest state penalty for each compound aimed to ensure that the models reflect the most physiologically relevant ligand forms. The consistent use of the TD50 rat endpoint across all datasets provides a clear and well-defined measure of activity, analogous to IC50 in drug discovery.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\u003ch2\u003e3.2 Descriptor Generation\u003c/h2\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eA total of 636 molecular descriptors were generated for every compound in both datasets (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), encompassing both categorical and numerical descriptors. This comprehensive set included semi-empirical, topological, functional group counts, and physiochemical descriptors. Semi-empirical descriptors, derived from quantum mechanical calculations, provide insights into the electronic structure of molecules. Topological descriptors describe molecular size and form based on atomic interconnectedness. Functional group counts reveal specific atomic groups that may influence biological activity, while physiochemical descriptors capture physical and chemical characteristics such as reactivity and solubility. This broad range of descriptors aligns with advanced intelligent QSAR development, leveraging extensive molecular characteristics to enable the model to learn complex relationships between structure and carcinogenic potency, thereby enhancing predictive power and versatility.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section2\"\u003e\u003ch2\u003e3.3 Model Training\u003c/h2\u003e\u003cp\u003eThe model training phase involved the development of both classification and regression models using various machine learning algorithms.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eTo begin, we started with Bayesian Classification models. With Model A, which lacked categorical descriptors, we achieved 60% accuracy on training data and 58% on testing data. Conversely, Model B which included categorical descriptors performed even worse as it achieved 56% accuracy on training data and 53% on testing data. These results as displayed in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e suggest that the categorical descriptors either added too much noise or complexity to the model or that the dataset was poorly suited for the Bayesian method due to its heavy reliance on the independence assumption between descriptors (which is nearly always wrong in QSAR models, (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e)). Additionally, the high dimensionality of 636 descriptors would likely harm Bayesian models, resulting in critical information being lost due to excessive discretization of continuous variables and over fitting.\u003c/p\u003e\u003cp\u003eRecursive Partitioning Models were created and seven models were analyzed in total. Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e illustrates that Models A-D were single models, whereas E-G were 100 tree ensemble models. Ensemble models E-G outperformed their single counterparts A-D in both training and test accuracy. This supports the observation that single models have limitations with respect to capturing intricate relationships within the data. Take, for example, Model E, an ensemble model that achieved test accuracy of 67% by applying Gini impurity while discarding weakly correlated variables. This model surpassed single Model B, which achieved only 53% accuracy with the use of information gain. As a rule, models built using Gini impurity performed better than those using information gain. Perhaps this is because the former is a weaker discriminator, and therefore less sensitive to fluctuations in class probabilities (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e). Furthermore, the accuracy gain from skipping low correlation variables is likely a result from less dimensionality making the noise to signal ratio favor signal.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eRecursive Partitioning Results\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eModel Type\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eAlgorithm\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDiscard Variables with low correlations\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eTraining Accuracy\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eTest Accuracy\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSingle\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGini Impurity\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eno\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e88\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e56\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSingle\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eInformation Gain\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eno\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e89\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e53\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eC\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSingle\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eInformation Gain\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eyes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e89\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e54\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eD\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSingle\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGini Impurity\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eyes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e87\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e62\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eE\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eEnsemble\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGini Impurity\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eyes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e98\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e67\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eF\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eEnsemble\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eInformation Gain\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eyes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e98\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e63\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eG\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eEnsemble\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGini Impurity\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eyes\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e95\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e64\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eAdditional classification models were developed using AutoQSAR (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e), and all top 10 models employed the Bayes Naive algorithm. The training accuracies for these models ranged from 69\u0026ndash;89%, while test accuracies were lower, ranging from 64\u0026ndash;80%. One of the models achieving a training accuracy of 89% utilized dendritic fingerprint descriptors, but the 73% test accuracy suggested over fitting. Conversely, the model driven by informatics tools performed better in terms of generalization, yielding 80% test accuracy while only achieving 77% training accuracy. This underscores the importance of independent validation to assess an algorithm's predictive accuracy (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eThe classification was done using DeepChem, which gave a high-performing model. The analysis from the Receiver Operating Characteristic (ROC) curve yielded an Area Under the Curve 0.81 (AUC-ROC) ( Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). An AUC-ROC of 0.81 indicates that there is an 81% chance the model accurately assesses whether a specific compound is carcinogenic or not, which confirms excellent predictive accuracy and robustness (\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). From the high curvature of the ROC, not only is the AUC good, but the model also possesses a high true positive rate and a low false positive rate across many decision-making thresholds which is rare and usually very hard to find. This is due to DeepChem using neural networks on molecular information which uncovers intricate relationships and far outperforms traditional machine learning methods.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eKernel-based Partial Least Squares (KPLS) was studied for its application in regression modeling. A total of four models were created. Despite achieving a respectable R\u0026sup2; of 0.84 in Model A, it suffered from considerable over fitting as indicated by a severely negative Q\u0026sup2; of -377. Predictive performance was significantly lower than that observed in the training data. Model D demonstrated the lowest R\u0026sup2; of 0.08 among all KPLS models, but was notable for having the highest Q\u0026sup2; of -0.36. These results (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e) illustrate that while PLS regression is a potent regression model, its performance can be greatly overshadowed by how the model is set up. In this case, all KPLS models failed to reach any acceptable benchmarks for regulatory predictive performance.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eIn contrast, the performance of the AutoQSAR regression model in estimating the potency of carcinogenesis was considerably higher. It was built using 427 compounds and was tailored for 21 different types of tumors, yielding an R\u0026sup2; of 0.58 and a Q\u0026sup2; of 0.51 (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e). An R\u0026sup2; of 0.58 means that the model predicts approximately 58% of the variability in carcinogenic potency, while Q\u0026sup2; of 0.51 suggests that the model predicts about 51% of the variance in new data (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eTogether, these values suggest a reasonable and balanced predictive capability, particularly given the additional difficulty posed by the numerous tumor types included in the dataset. The accuracy of AutoQSAR\u0026rsquo;s results demonstrated the effectiveness of its more sophisticated algorithms and its extensive computations of molecular descriptors that integrate fundamental synthetic molecular characteristics required for precise predictions.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\u003ch2\u003e3.4 Model Comparisons\u003c/h2\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eModel Details\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAuthors\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eYear\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eModel Type\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eEndpoint\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eApplicability Domain\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eModel Codename\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLaboratory of Comparative Toxicology and Ecotoxicology\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e1991\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSimple Equations\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eTD/50 RAT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eLow, Limited to a small chemical space\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e1A\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLaboratory of Comparative Toxicology and Ecotoxicology\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e2002\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eSimple Equations\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eTD/50 RAT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eAromatic Amines\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e2B\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFDA's Center for Drug Evaluation and Research\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e2003\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eStatistical Model\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eTD/50 RAT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eWide Chemical Space\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e3C\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eContrera\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e2005\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eStatistical Model\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eTD/50 RAT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003ePharmaceuticals\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e4D\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eUniversity of Porto\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e2008\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGenetic Algorithm\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eTD/50 RAT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eNitroso Compounds\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e6F\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eREACH\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e2010\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMachine Learning\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eTD/50 RAT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eWide Chemical Space\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e7G\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLaboratory of Comparative Toxicology\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e2021\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDeep Learning\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eTD/50 RAT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eWide Chemical Space\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e8H\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eFDA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e2021\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDeep Learning\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eTD/50 RAT\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eWide Chemical Space\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e9I\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e provides an overview of various QSAR models from the literature, detailing their authors, year of development, model type, endpoint, and applicability domain. This table serves as a reference for comparing the performance of the models developed in the current study against existing benchmarks.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eClassification Models Comparison\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTraining Accuracy\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eTest Accuracy\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eExternal Validation Accuracy\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e1991\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e72\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e65\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e2002\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e91\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e72\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e2003\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e73\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e2005\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e83\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e80\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e2010\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e91\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e73\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e66\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e2021\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e81\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e76\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e2021\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e91\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e77\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBAYESA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e58\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e58\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e58\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRPB\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e95\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e64\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e69\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAQA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e83\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e72\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e62\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDP\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e81\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e72\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e presents a comparative analysis of various classification models, including those from the literature and the models developed in this study (BayesA, RPB, AQA, DP). The metrics include training accuracy, test accuracy, and external validation accuracy where available. This table allows for a direct comparison of how well each model performed in classifying compounds as carcinogenic or non-carcinogenic.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eRegression Model Comparisons\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"3\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eR^2\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eQ^2\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e2008\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.808\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.248\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e2021\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.756\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePLSA\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.08\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e-0.36\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ePLSD\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.84\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e-377\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eAQ\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.58\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.51\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDPC\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e\u003cp\u003e0.02\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e compares the performance of regression models for predicting carcinogenic potency, including those from the literature and the models developed in this study (PLSA, PLSD, AQ, and DPC). The key metrics for comparison are R\u0026sup2; (coefficient of determination) and Q\u0026sup2; (goodness of prediction), which indicate how well the model fits the training data and predicts new, unseen data, respectively.\u003c/p\u003e\u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eThe findings of this study demonstrate significant advancements in carcinogenicity prediction through the application of machine learning-enhanced QSAR models, particularly when compared to existing literature. The DeepChem classification model, achieving an 81% test accuracy and 72% external validation accuracy, notably surpasses the performance of earlier models. For instance, the FDA's 2005 model, while showing 80% training accuracy, yielded only 76% test accuracy (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e). Similarly, the 2010 OECD model, a significant regulatory advancement, reported 73% accuracy on its test set with 66% external validation (\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e). The DeepChem model's superior performance can be attributed to its deep learning architecture, which is adept at capturing complex, non-linear relationships within molecular data that traditional QSAR methods might miss (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e). This aligns with the broader trend in advanced QSAR models (2015\u0026ndash;2023) that leverage deep learning to transition from traditional descriptor engineering to more sophisticated molecular embeddings, thereby enhancing prediction precision (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eFurthermore, the models developed in this study exhibit a broad applicability domain, covering 21 different tumor types. This is a crucial advancement over earlier niche approaches, such as the 2002 model specifically designed for aromatic amines (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e) or the 2008 model focused solely on nitroso compounds (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e). The comprehensive validation rigor, including the use of an independent external dataset, ensured the robustness of our models and addressed the common risk of over fitting observed in some prior studies (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eIn terms of regression, the AutoQSAR model demonstrated a balanced and reliable performance with an R\u0026sup2; of 0.58 and a Q\u0026sup2; of 0.51. This represents a substantial improvement over earlier potency prediction models, such as the 2008 nitroso-compound model which reported a Q\u0026sup2; of only 0.248 (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e). The ability of AutoQSAR to achieve a Q\u0026sup2; greater than 0.5 makes it suitable for regulatory applications, as it indicates a good predictive capability on unseen data (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e). While other regression models like KPLS showed high R\u0026sup2; values, their significantly negative Q\u0026sup2; values indicated severe over fitting and poor generalization, consistent with challenges in developing robust regression models for carcinogenicity (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e). The integration of diverse machine learning algorithms within AutoQSAR likely contributed to its ability to capture the underlying relationships across varied chemical structures and tumor types, a complexity that often limits the performance of simpler models.\u003c/p\u003e\u003cp\u003eDespite these advancements, certain limitations warrant consideration. The primary dataset for this study was derived from rodent assays. While these are standard in carcinogenicity testing, the direct translation of findings to human toxicity can be limited due to species-specific differences in metabolism and response (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). Future research should explore expanding training data to include human-relevant endpoints to enhance direct applicability. Additionally, model performance is inherently dependent on the quality and selection of descriptors. While a comprehensive set of 636 descriptors was generated, future work could investigate the utility of more advanced, graph-based embeddings or other novel descriptor generation techniques to potentially capture more nuanced molecular features (\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e).\u003c/p\u003e"},{"header":"5. Conclusion","content":"\u003cp\u003eThis study successfully developed and validated machine learning-enhanced QSAR models for the early detection of carcinogenic agents, addressing key limitations of traditional methods and existing QSAR approaches. All defined objectives\u0026mdash;data preparation, molecular descriptor generation, model building, and rigorous validation\u0026mdash;were successfully met. The DeepChem classification model demonstrated superior performance, achieving an 81% test accuracy and a robust 72% external validation accuracy, showcasing its strong generalization capabilities. Concurrently, the AutoQSAR regression model provided reliable potency prediction with an R\u0026sup2; of 0.58 and a Q\u0026sup2; of 0.51, outperforming previous benchmarks and indicating its suitability for regulatory applications. Together, these models represent a pivotal step toward the application of advanced computational methods for the accurate and efficient testing of carcinogenicity.\u003c/p\u003e\n\u003cp\u003eThe growing use of advanced deep learning models demonstrates their potential in enhancing predictive toxicology tasks. These models comply with the OECD Regulations pertinent to QSAR Validation owing to their suitability for use in risk assessment. Further work in this area aims to improve model translatability by supplementing the training datasets with relevant human-exposure endpoints. Further, the application of XAI methods could reveal the mechanisms of predictions made which, in turn, would bolster the confidence and understanding in using these intelligent QSAR systems.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eCRediT author statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTanaka Denzel Chitsa\u003c/strong\u003e: Conceptualization, Methodology, Investigation, Writing- Original draft, Visualization, Formal analysis\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAmos Misi\u003c/strong\u003e: Data curation, Supervision, Validation, Investigation, Project administration, Resources, Writing- Reviewing and Editing\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDr Paul Mushonga, Dr Albert Wakandigara\u003c/strong\u003e: Visualization, Investigation, Supervision , Writing- Reviewing and Editing\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRoy T Bisenti\u003c/strong\u003e: Formal analysis, Writing- Reviewing and Editing, Visualization\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDeclaration of interests\u003c/strong\u003eThe authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.\u003cbr\u003e\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eJacobs MN, Colacci A, Corvi R, Vaccari M, Aguila MC, Corvaro M et al (2020) Chemical carcinogen safety testing: OECD expert group international consensus on the development of an integrated approach for the testing and assessment of chemical non-genotoxic carcinogens. Arch Toxicol 94(8):2899\u0026ndash;2923\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJain S A Comprehensive Study for The Regulations of Genotoxic Impurities in Pharmaceutical Drug Substances \u0026amp; Products with Special Reference to Their Detection, Identification and Evaluatione. Nirmauni.ac.in [Internet]. 2016 [cited 2025 Jan 1]; Available from: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://repository.nirmauni.ac.in/jspui/handle/123456789/6541\u003c/span\u003e\u003cspan address=\"https://repository.nirmauni.ac.in/jspui/handle/123456789/6541\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChung E, Russo DP, Ciallella HL, Wang YT, Wu M, Aleksunes LM et al (2023) Data-driven quantitative structure\u0026ndash;activity relationship modeling for human carcinogenicity by chronic oral exposure. Environ Sci Technol 57(16):6573\u0026ndash;6588\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCherkasov A, Muratov EN, Fourches D, Varnek A, Baskin II, Cronin M et al (2014) QSAR modeling: where have you been? Where are you going to? J Med Chem 57(12):4977\u0026ndash;5010\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRoy K, Kar S, Das RN (2015) Understanding the basics of QSAR for applications in pharmaceutical sciences and risk assessment. Academic\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eToma C, Manganaro A, Raitano G, Marzo M, Gadaleta D, Baderna D et al (2020) QSAR models for human carcinogenicity: an assessment based on oral and inhalation slope factors. Molecules 26(1):127\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBenigni R, Bossa C (2011) Flexible use of qsar models in predictive toxicology: A case study on aromatic amines. Environ Mol Mutagen 53(1):62\u0026ndash;69\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHalder AK, Moura AS, Cordeiro MNDS (2018) QSAR modelling: a therapeutic patent review 2010-present. Expert Opin Ther Pat 28(6):467\u0026ndash;476\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMwanza C, Zhang WZ, Mulenga Kalulu, Ding SN (2024) Advancing green chemistry in environmental monitoring: the role of electropolymerized molecularly imprinted polymer-based electrochemical sensors. Green Chem\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChen M, Bisgin H, Tong L, Hong H, Fang H, Borlak J et al (2014) Toward predictive models for drug-induced liver injury in humans: are we there yet? Biomark Med 8(2):201\u0026ndash;213\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDargan S, Kumar M, Ayyagari MR, Kumar G (2019) A Survey of Deep Learning and Its Applications: A New Paradigm to Machine Learning. Arch Comput Methods Eng. ;27(4)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eToropova AP, Toropov AA (2018) CORAL: QSAR models for carcinogenicity of organic compounds for male and female rats. Comput Biol Chem 72:26\u0026ndash;32\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNag S, Baidya ATK, Mandal A, Mathew AT, Das B, Devi B et al (2022) Deep learning tools for advancing drug discovery and development. 3 Biotech. ;12(5)\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFjodorova N, Vračko M, Novič M, Roncaglioni A, Benfenati E (2010) New public QSAR model for carcinogenicity. Chem Cent J 4(Suppl 1):S3\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKostal J, Voutchkova-Kostal A (2023) Quantum-Mechanical Approach to Predicting the Carcinogenic Potency of N-Nitroso Impurities in Pharmaceuticals. Chem Res Toxicol 36(2):291\u0026ndash;304\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZuang V, Dura A, Bofill A, Leite B, Berggren E, Bernasconi C et al (2019) EURL ECVAM status report on the development, validation and regulatory acceptance of alternative methods and approaches Publications Office of the European Union Luxembourg; 2020\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTropsha A, Isayev O, Varnek A, Schneider G, Cherkasov A (2024) Integrating QSAR modelling and deep learning in drug discovery: the emergence of deep QSAR. Nat Rev Drug Discov 23(2):141\u0026ndash;155\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDu CJ, He HJ, Sun DW (2016) Object classification methods. Computer vision technology for food quality evaluation. Elsevier, pp 87\u0026ndash;110\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDutta P, Paul S, Kumar A (2021) Comparative analysis of various supervised machine learning techniques for diagnosis of COVID-19. Electronic devices, circuits, and systems for biomedical applications. Elsevier, pp 521\u0026ndash;540\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCiallella HL, Chung E, Russo DP, Zhu H (2022) Automatic quantitative structure\u0026ndash;activity relationship modeling to fill data gaps in high-throughput screening. High-Throughput Screening Assays in Toxicology. Springer, pp 169\u0026ndash;187\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDixon SL, Duan J, Smith E, Von Bargen CD, Sherman W, Repasky MP (2016) AutoQSAR: an automated machine learning tool for best-practice quantitative structure\u0026ndash;activity relationship modeling. Future Med Chem 8(15):1825\u0026ndash;1839\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBahia MS, Kaspi O, Touitou M, Binayev I, Dhail S, Spiegel J et al (2023) A comparison between 2D and 3D descriptors in QSAR modeling based on bio-active conformations. Mol Inf 42(4):2200186\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"University of Zimbabwe","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"QSAR, machine learning, carcinogenicity, deep learning, predictive toxicology","lastPublishedDoi":"10.21203/rs.3.rs-7438250/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7438250/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eTraditional carcinogenicity testing methods are costly and time-consuming, while existing Quantitative Structure-Activity Relationship (QSAR) models suffer from low accuracy and limited applicability domains. This study addresses these limitations by developing enhanced QSAR models using machine learning (ML) algorithms. A dataset of 805 compounds from the Carcinogenic Potency Database (CPD) was used to train classification and regression models, employing Bayesian classifiers, recursive partitioning, Kernel-based Partial Least Squares (KPLS), and deep learning techniques (Neural Networks, Random Forests). An independent validation dataset (105 compounds) was used to assess model performance. The DeepChem-based classification model achieved 81% test accuracy and 72% external validation accuracy, while the AutoQSAR regression model demonstrated an R\u0026sup2; of 0.58 and Q\u0026sup2; of 0.51, outperforming existing literature benchmarks. These models exhibit broad chemical space coverage, offering a robust, cost-effective alternative for carcinogenicity prediction.\u003c/p\u003e","manuscriptTitle":"Intelligent QSAR Approaches: Harnessing Machine Learning for Early Detection of Carcinogenic Agents","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-08-26 09:30:26","doi":"10.21203/rs.3.rs-7438250/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e1fd5eb6-d3d0-4d96-82fb-b335e0cfedf3","owner":[],"postedDate":"August 26th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":53595921,"name":"Computational Chemistry"},{"id":53595922,"name":"Medicinal Chemistry"}],"tags":[],"updatedAt":"2025-08-26T09:30:26+00:00","versionOfRecord":[],"versionCreatedAt":"2025-08-26 09:30:26","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7438250","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7438250","identity":"rs-7438250","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00