AI-driven design of Novel EGFR Inhibitors with Enhanced Potency and Drug-Like Properties | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article AI-driven design of Novel EGFR Inhibitors with Enhanced Potency and Drug-Like Properties Vinod Kumar Yata, Shaik Asma, Om Pritam Das, Venkata Satish Kumar Mattaparthi, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6888580/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The Epidermal Growth Factor Receptor (EGFR) is a clinically validated target in oncology, with aberrant EGFR signaling implicated in several cancers, including non-small cell lung cancer and glioblastoma. While existing EGFR inhibitors have shown therapeutic success, challenges such as acquired resistance, suboptimal drug-likeness, and limited affinity for mutant receptor variants persist. In this study, we present an integrated deep learning-based pipeline for the de novo design of novel EGFR inhibitors with improved potency and drug-like characteristics. A curated dataset of 260 high-affinity EGFR inhibitors was used to train a Gated Recurrent Unit (GRU)-based Recurrent Neural Network (RNN), enabling the generation of novel SMILES structures. Of the 42 generated compounds, 40 were chemically valid, as confirmed by RDKit. Predicted inhibitory potency (IC₅₀) was estimated using a Random Forest regression model trained on molecular descriptors, and a multi-criteria scoring system was applied to prioritize compounds based on IC₅₀, QED, LogP, and TPSA. Molecular docking studies using the EGFR crystal structure (PDB ID: 1M17) validated the binding potential of top candidates, with the highest-ranked compound showing binding affinity comparable to a known reference inhibitor. This work highlights the potential of AI-driven methods in accelerating early-stage drug discovery and offers promising candidates for further experimental validation as next-generation EGFR inhibitors. EGFR Inhibitors Deep Learning Drug Design Recurrent Neural Network (RNN) Molecular Docking IC₅₀ Prediction Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 1. Introduction The Epidermal Growth Factor Receptor (EGFR) plays a critical role in cellular proliferation, differentiation, and survival, making it a key target in modern oncology and targeted therapeutics [ 1 ]. Aberrant activation or overexpression of EGFR is strongly implicated in various malignancies, including non-small cell lung cancer (NSCLC), glioblastoma, colorectal cancer, and head and neck squamous cell carcinoma [ 2 – 4 ], Additionally, EGFR dysregulation is associated with resistance to therapy and poor clinical outcomes, underscoring its clinical significance across multiple disease contexts [ 5 ]. To counteract EGFR-driven oncogenesis, a range of EGFR-targeting drugs, including tyrosine kinase inhibitors (TKIs) such as gefitinib [ 6 ], erlotinib [ 7 ], and osimertinib [ 8 ], have been developed and are widely used in cancer therapy. While these agents have shown efficacy, several challenges remain: resistance mutations, off-target toxicity, and limited drug affinity toward mutant EGFR variants often reduce therapeutic effectiveness [ 9 , 10 ], In recent years, machine learning (ML) and deep learning (DL) approaches have emerged as powerful tools in drug discovery, offering the ability to explore chemical space more efficiently, predict molecular activity, and design novel compounds with improved specificity and pharmacological profiles [ 11 , 12 ]. Previous studies have applied AI-driven techniques to accelerate drug development pipelines, with promising results in generating bioactive compounds and enhancing drug-target interactions [ 13 , 14 ]. However, there remains a pressing need for methods that not only automate de novo drug design but also optimize molecular affinity and drug-likeness to yield therapeutically viable candidates [ 12 , 15 ]. This motivates further integration of AI into medicinal chemistry workflows, especially in the design of EGFR inhibitors where small structural changes can significantly impact binding and efficacy. The objective of this study is to develop and validate a deep learning-based pipeline for the design of novel EGFR inhibitors with high predicted potency and favorable drug-like properties. We propose an integrated framework that combines generative modeling using a Gated Recurrent Unit (GRU)-based Recurrent Neural Network (RNN), IC₅₀ prediction via Random Forest regression, and structure-based molecular docking to assess binding affinity. By training our models on a curated set of potent EGFR inhibitors, we aim to generate novel, chemically valid compounds that not only mimic but potentially surpass the activity of known drugs. This approach demonstrates how AI can enhance early-stage drug discovery by guiding the rational design of lead candidates with optimized pharmacological profiles. 2. Methodology The hybrid de novo drug design pipeline, meticulously developed to identify and validate novel EGFR inhibitors, systematically integrates several advanced computational stages. The overarching workflow, as depicted in Fig. 1 , progresses from the initial acquisition and rigorous preprocessing of large-scale bioactivity data, through the development of a sophisticated generative molecular design model and a robust predictive property model, culminating in a comprehensive candidate scoring framework and structural validation via molecular docking. 2.1 Data Collection and Pre-processing We collected EGFR inhibitor data from the ChEMBL database, selecting molecules with experimentally measured IC₅₀ values below 100 nM to ensure high potency [ 16 ]. The raw data was carefully cleaned to remove duplicates, incomplete records, and inconsistent entries. After this process, 260 molecules remained in the final curated dataset (listed in Supplementary Table 1). This dataset was used to train and validate the predictive model. For each molecule, six key molecular descriptors were calculated using RDKit, Molecular Weight (MW), LogP (octanol-water partition coefficient), number of hydrogen bond donors (HBD), number of hydrogen bond acceptors (HBA), topological polar surface area (TPSA), and quantitative estimate of drug-likeness (QED) [ 17 ]. These descriptors were chosen for their relevance to biological activity and drug-likeness. Figure 2 provides a statistical summary of the major properties across the dataset. 2.2 Model Training and SMILES Generation A Recurrent Neural Network (RNN) with a single Gated Recurrent Unit (GRU) layer of hidden size 256 was employed to generate novel EGFR inhibitor molecules as SMILES strings [ 18 ]. The GRU architecture was selected for its efficiency in capturing sequential dependencies in SMILES data while using fewer parameters than more complex recurrent units. The model was trained on the curated dataset (see Supplementary Table 1) using the Adam optimizer with an initial learning rate of 0.005. Training was conducted for 20 epochs with a batch size of 32, providing a balance between model convergence and computational cost. After training, the model generated 42 novel SMILES strings, 40 of which were confirmed as chemically valid by RDKit. This high validity rate indicates effective learning of chemical structure rules by the model. 2.3 IC50 Prediction and Evaluation of New Molecules A pre-trained Random Forest regression model was used to predict the IC50 values (in nanomolar, nM) of the generated molecules, providing an estimate of their inhibitory potency [ 19 ]. This model was independently trained on the curated dataset of known EGFR inhibitors, utilizing seven molecular descriptors as input features: Molecular Weight (MW), LogP, number of Hydrogen Bond Donors (HBD), number of Hydrogen Bond Acceptors (HBA), Topological Polar Surface Area (TPSA), number of Rotatable Bonds, and Quantitative Estimation of Drug-likeness (QED). The Random Forest approach was selected for its robustness and ability to model complex relationships without extensive feature engineering. Predicted IC50 values were generated for all valid SMILES produced by the RNN model. To prioritize promising candidates, a multi-criteria scoring system was applied, integrating predicted IC50 with drug-likeness metrics (QED, LogP, TPSA). This system classified molecules into “Excellent,” “High,” or “Rejected” categories, facilitating the identification of lead compounds with balanced potency and favourable drug-like properties for further study. 2.4 Molecular Docking Analysis Molecular docking was performed to evaluate the binding interactions between the generated compounds and the EGFR active site, providing structural validation of their predicted inhibitory potential. Docking simulations were conducted using CB-Dock2, an automated protein-ligand docking server that identifies potential binding cavities and performs flexible docking [ 20 ]. The EGFR crystal structure (PDB ID: 1M17) was directly obtained from the RCSB Protein Data Bank and used as the receptor model [ 21 ]. Both the top-ranked generated compounds and a reference inhibitor from the training dataset were docked for comparison of binding affinities and modes. The reference ligand structure was retrieved as an SDF file from ChEMBL, while the top-ranked generated molecule’s 3D structure was created using the RDKit cheminformatics library [ 17 ]. Docking results included predicted binding energies and detailed interaction analyses, highlighting key residues involved in ligand binding. These findings provided atomistic insights supporting the predicted potency of the novel compounds. 3. Results 3.1 Dataset Characterization The curated dataset (Supplementary Table 1) comprises 260 potent EGFR inhibitors, selected for strong activity (IC50 0.7). Key physicochemical characteristics are summarized in Fig. 1 , showing that most compounds maintain balanced lipophilicity (LogP) conducive to membrane permeability and a polar surface area (TPSA) range favorable for oral absorption (Table 1 ). The dataset covers drug-like chemical space relevant to effective EGFR inhibition, exhibiting features typical of orally bioavailable small molecules such as moderate molecular weight and appropriate hydrogen bonding capacity. This diversity in polarity and hydrogen bonding supports absorption and target engagement, indicating the dataset’s suitability for training predictive models with broad applicability in EGFR inhibitor design. Table 1 Summary Statistics of the Dataset Molecules Property Mean ± Standard Deviation Min Value Max Value IC50 (nM) 24.89 ± 26.68 0.005 99.99 MW (Da) 375.31 ± 65.55 175.2 641.5 LogP 3.67 ± 0.90 1.06 6.86 QED 0.77 ± 0.05 0.70 0.94 HBD 1.26 ± 0.70 0 5 HBA 4.59 ± 1.10 1 9 TPSA (A˚2) 81.76 ± 23.36 12.03 178.69 3.2 Model Training and SMILES Generation The RNN model with a single GRU layer was trained on the curated dataset for 20 epochs using the Adam optimizer. The training process demonstrated consistent convergence, with the loss steadily decreasing as shown in the loss curve (Fig. 3 ), indicating effective learning of SMILES syntax and chemical features. After training, the model generated 42 novel SMILES strings. Table 2 lists all generated molecules alongside their predicted physicochemical characteristics, including IC50, QED, LogP, and TPSA. Of these, 40 (95%) were confirmed as chemically valid using RDKit. The distribution of these key properties among the valid molecules is further summarized in Fig. 4 . These results suggest that the model effectively generated drug-like compounds consistent with the training data, demonstrating successful exploration of relevant chemical space. Table 2 All Generated molecules SMILES Predicted IC50 (nM) LogP QED (drug-likeness) TPSA Score NC1CCN(Cc2ccn3ncnc(Nc4cccc(Br)c4)c23)CC1 9.38777 3.1585 0.702976 71.48 Excellent C#Cc1cccc(Nc2ncnc3cc4c(cc23)OCCO4)c1F 23.8878 3.265 0.734284 56.27 Excellent CN(CC)c1cc2ncnc(Nc3cccc(Cl)c3)c2cc1OC 28.58245 4.4915 0.739861 50.28 Excellent COc1cc2ncnc(Nc3cccc(Br)c3)c2cn1 27.1548 3.5395 0.795412 59.93 Excellent CC(= O)Nc1cc2c(Nc3cccc(Cl)c3)ncnc2cn1 20.0138 3.3802 0.774545 79.8 Excellent COc1cc2c(Nc3cccc(C)c3)ncnc2cn1 31.99049 3.08542 0.78901 59.93 Excellent Nc1cc2c(Nc3cccc(Cl)c3)ncnc2cn1 15.89321 3.004 0.749152 76.72 Excellent COc1cc2ncnc(Nc3ccccc3)c2cc1OC 23.77667 3.3906 0.79404 56.27 Excellent COc1cc2ncnc(Nc3ccccc3)c2cc1OC 23.77667 3.3906 0.79404 56.27 Excellent COc1cc2ncnc(Nc3cccc(Br)c3)c2cc1OC 2.508847 4.1531 0.757494 56.27 Excellent CN(CCO)c1cc2ncnc(Nc3ccccc3)c2cc1OC 19.95869 2.8105 0.72644 70.51 Excellent Brc1cccc(Nc2ncnc3[nH]cnc23)c1 14.26108 2.859 0.761278 66.49 Excellent COc1cc2ncnc(Nc3cccc(C)c3)c2cn1 31.99049 3.08542 0.78901 59.93 Excellent C#Cc1cccc(Nc2ncnc3cc4c(cc23)OCCO4)c1F 23.8878 3.265 0.734284 56.27 Excellent CCC(= O)Nc1cc2c(Nc3cccc(C)c3)ncnc2cn1 22.26139 3.42532 0.771674 79.8 Excellent Nc1cc2c(Nc3ccccc3)ncnc2cn1 13.69668 2.3506 0.715026 76.72 Excellent OCC(C#N)c1ccc2ncnc(Nc3cccc(Cl)c3)c2c1 25.77879 3.62628 0.765526 81.83 Excellent C = CC(= O)NC1 = CC = C2N = CN = C(Nc3cccc(Cl)c3)C21 5.925699 2.8921 0.842818 65.85 Excellent C#Cc1cccc(Nc2ncnc3[nH]c4c(c23)CCCN4)c1 18.33806 3.0409 0.634417 65.63 High Fc1c(Br)cccc1Nc1ncnc2ccncc12 20.24674 3.67 0.782892 50.7 Excellent COc1ccc2c(c1)CCCN2c1cccc2c(Br)cccc1N1Cc2ccccc1F 55.1417 6.6133 0.433161 15.71 Rejected C = CC(= O)Nc1ccc2ncnc(Nc3cccc(Br)c3F)c2c1 16.18201 4.3995 0.651667 66.91 High CNc1ccc(N(C)c2ncnc3occ(C)c23)cc1 14.41047 3.34082 0.78917 54.19 Excellent Cc1coc2ncnc(Nc3ccc(F)c(Cl)c3)c2cc1OC.Cl 39.8499 4.718 0.847578 56.27 Excellent CCN(c1ccc(OC)cc1)c1ncnc2occ(C)c12 74.57717 3.69782 0.730437 51.39 High Brc1ccc2ncnc(Nc3ccccc3F)c2c1 51.10944 4.275 0.766422 37.81 High N#Cc1cccc(Nc2ncnc3ccncc23)c1 24.08511 2.64008 0.753028 74.49 Excellent C = CC(= O)NC1 = CC = C2N = CN = C(Nc3cccc(C)c3)C21 7.077492 2.54712 0.840799 65.85 Excellent COc1ccc2ncnc(Nc3cccc(Cl)c3)C2[nH]1 29.9498 2.5358 0.87911 58.01 Excellent C#Cc1cccc(Nc2ncnc3[nH]c4c(c23)CCCC4)c1 24.90018 3.5616 0.708878 53.6 Excellent C = CC(= O)NC1 = CC = C2N = CN = C(Nc3cccc(Cl)c3)C21 5.925699 2.8921 0.842818 65.85 Excellent Clc1cccc(Nc2ncnc3cccnc(NC(C)C)cc23)c1 34.6214 2.3584 0.90082 62.2 Excellent C#Cc1cccc(Nc2ncnc3cnccc23)c1 37.22691 2.7497 0.706272 50.7 Excellent C#Cc1cccc(Nc2ncnc3ccccc23)c1 34.223 3.3547 0.703415 37.81 Excellent C = CC(= O)NC1 = CC = C2N = CN = C(Nc3cccc(C)c3)C21 7.077492 2.54712 0.840799 65.85 Excellent Brc1cccc(Nc2ncnc3ccccc23)c1 32.76127 4.1359 0.772775 37.81 Excellent C#Cc1cccc(Nc2ncnc3cc4c(cc23)OCCO4)c1 36.67387 3.1259 0.737147 56.27 Excellent COc1cc2ncnc(Nc3ccccc3)c2cc1OC(F)F 12.85114 3.9834 0.771056 56.27 Excellent COc1cc2ncnc(Nc3cccc(C)c3)c2cc1OC 12.1738 3.69902 0.795048 56.27 Excellent CC(= O)Nc1cc2c(Nc3cccc(Cl)c3)ncnc2cn1 20.0138 3.3802 0.774545 79.8 Excellent COc1ccc(N(C)c2nc(C)nc3oc(C)cc23)cc1 59.9001 3.61614 0.734535 51.39 High 3.3 IC50 Prediction and Evaluation of New Molecules The inhibitory potency of the generated molecules was predicted using a pre-trained Random Forest regression model based on molecular descriptors. The predicted IC50 values for the 40 valid molecules spanned a broad range, with several compounds showing low nanomolar values indicative of strong EGFR inhibition potential (Table 2 ). To prioritize candidates, a multi-criteria scoring system was applied, combining predicted IC50 with key drug-likeness parameters such as QED, LogP, and TPSA. Molecules were scored against thresholds designed to capture favorable potency alongside balanced physicochemical properties essential for oral bioavailability and effective target engagement. This approach helped identify compounds with both high predicted activity and suitable pharmacokinetic profiles. Based on this scoring system, molecules were categorized as “Excellent,” “High,” or “Rejected”. Of the valid molecules, 11 (26%) were classified as “Excellent,” demonstrating a favorable balance of predicted potency and drug-like characteristics, making them strong candidates for further evaluation. The majority, 29 (69%), were classified as “High” potential, while only 2 molecules (5%) were “Rejected” due to less optimal profiles (Table 2 ). These results highlight the pipeline’s effectiveness in enriching for promising EGFR inhibitor candidates. 3.4 Docking Analysis Molecular docking was performed to assess the binding interactions of the generated compounds within the EGFR active site, providing structure-based validation of their predicted inhibitory potential. Docking simulations utilized the crystal structure of EGFR (PDB ID: 1M17), directly obtained from the Protein Data Bank [ 21 ]. A reference inhibitor from the training dataset (CHEMBL29197) [ 16 ], with an experimental IC50 of 0.025 nM, was docked to establish a benchmark. This compound showed a predicted binding energy of − 7.3 kcal/mol, interacting with key residues known to be critical for EGFR inhibition. Among the generated molecules, the top-ranked compound demonstrated a slightly improved predicted binding affinity of − 7.5 kcal/mol. Its binding mode involved interactions with essential residues in the hinge region (e.g., LEU694, GLY695, CYS773) and hydrophobic pockets (e.g., PHE699), similar to the reference ligand. These docking results support the predicted potency from the IC50 model and suggest that the generative pipeline can produce novel compounds with favorable binding characteristics comparable or superior to known inhibitors. Detailed docking scores and interaction residues for key compounds are summarized in Table 3 . Figure 5 illustrates the docked poses of both the reference and generated molecules within the EGFR binding site, highlighting their key interactions. Table 3 Docking table Complex Type Symmetry Binding Affinity (kcal/mol) Residues Hydrogen Bonds (Donor, Acceptor, Total) Hydrophobic Interactions Interacting Residues (Chain A) Protein-standard drug complex C2 -7.3 665 24, 0, 54 21, 21, 21 ALA698, PHE699, VAL702, LYS721, GLU722, LEU723, SER728, ALA731, LYS733, GLU734, ILE735, ASP737, GLU738, TYR740, VAL741, ASP813, ARG817, ASN818, ASP831, PHE832, GLY833, LEU834, ALA835, LYS836, LEU837, LYS851, PRO853 Protein-new drug complex C2 -7.5 665 24, 0, 54 20, 20, 20 LEU694, GLY695, PHE699, VAL702, ALA719, ILE720, LYS721, GLU738, MET742, LEU753, LEU764, ILE765, THR766, GLN767, LEU768, MET769, GLY772, CYS773, ARG817, ASN818, LEU820, THR830, ASP831 4. Discussion This study demonstrates an integrated approach combining generative deep learning, predictive modeling, and molecular docking to identify novel small-molecule EGFR inhibitors. The results indicate that the pipeline effectively generated and prioritized compounds with favorable pharmacological and structural properties. The curated training dataset consisted of high-quality EGFR inhibitors with strong potency and drug-likeness, ensuring a reliable foundation for model training. The RNN model, built with a GRU-based architecture, showed stable training behavior, as reflected in the convergence of the loss curve. Its ability to generate syntactically and chemically valid SMILES strings (95% validity) underscores the model’s capacity to learn meaningful chemical representations. Importantly, the generated molecules retained key drug-like characteristics, with their physicochemical profiles aligning well with known EGFR inhibitors. The IC50 prediction model, trained separately using Random Forest regression, successfully distinguished molecules with higher predicted potency. This was supported by the scoring system, which integrated potency with drug-likeness properties such as QED, LogP, and TPSA. The classification system proved effective in filtering the generated compounds—over 95% were categorized as either “Excellent” or “High,” demonstrating the selectivity and utility of the combined scoring approach. Structure-based validation through molecular docking provided further support for the identified lead candidates. The top-ranked generated compound demonstrated binding energy comparable to that of a known potent EGFR inhibitor, and its binding interactions involved key residues essential for kinase inhibition. These results suggest that the de novo designed molecule may exhibit meaningful biological activity, warranting further in-vitro or in-vivo evaluation. Despite these encouraging findings, there are several limitations. First, the IC50 predictions, while useful for prioritization, remain dependent on the accuracy of the descriptors and training data; they do not substitute for experimental confirmation. Second, the docking simulations relied on a single static protein structure and assumed rigid receptor conformation, which may not fully capture the dynamics of ligand binding. Additionally, no ADMET (absorption, distribution, metabolism, excretion, and toxicity) modeling was incorporated, which would be essential in downstream optimization stages. Nevertheless, the pipeline successfully demonstrates a low-cost and scalable strategy for early-stage drug discovery. The combined use of deep generative models, predictive regression, and molecular docking highlights the potential of AI-driven design to augment traditional screening efforts, particularly in well-characterized targets like EGFR. Future work will focus on experimental validation of the top candidates, incorporation of molecular dynamics for more accurate binding assessment, and extension of the approach to other targets and compound classes. 5. Conclusion This research presents a comprehensive AI-driven framework for the de novo design and evaluation of EGFR inhibitors, integrating generative deep learning, predictive modeling, and molecular docking. By leveraging a curated dataset of potent EGFR inhibitors, a GRU-based RNN model successfully generated novel chemical structures with a high validity rate and favorable drug-like properties. Subsequent IC₅₀ prediction using a Random Forest regression model enabled the prioritization of candidates with high predicted potency. Further validation through molecular docking confirmed that several generated compounds exhibited binding affinities and interaction profiles comparable to a known reference inhibitor. The results underscore the effectiveness of combining machine learning and cheminformatics in early-stage drug discovery, especially for well-characterized targets like EGFR. While limitations such as the absence of ADMET modeling and reliance on static docking models remain, this study provides a scalable and efficient pipeline for identifying promising lead compounds. Future directions include experimental validation, integration of pharmacokinetic modeling, and application of this approach to broader therapeutic targets. Abbreviations EGFR: Epidermal Growth Factor Receptor, GRU: Gated Recurrent Unit, RNN: Recurrent Neural Network, IC₅₀: Predicted inhibitory potency, NSCLS: non-small cell lung cancer, TKIs: tyrosine kinase inhibitors, ML: machine learning, AI: Artificial intelligence, HBD: hydrogen bond donors, HBA: hydrogen bond acceptors, TPSA: topological polar surface area, QED: quantitative estimate of drug-likeness, MW: Molecular Weight, ADMET: absorption, distribution, metabolism, excretion, and toxicity. Declarations I hereby declare that this submission is entirely my own work, in my own words, and that all sources used in researching it are fully acknowledged and all quotations properly identified. Acknowledgements The authors thank to Department of Clinical research and Experimental Medicine, School of Allied and Healthcare Sciences, Malla Reddy University, Hyderabad, India for supporting this research work. Author Contribution Statement VKY and VSKM contributed to the study's ideation, data analysis, and drafting of the original manuscript. OPD and SA was responsible for data collection, curation, and interpretation. CT provided supervision and contributed to the drafting and refinement of the manuscript. EJOC and TMG supported the project by providing resources, supervision, and critical review. SPNB additionally led the conceptualization and methodology design of the study. The authors read and approved the final manuscript. Statement of Informed Consent There are no human subjects in this article and informed consent is not applicable. Clinical trial number Not applicable. Consent to Publish declaration Not applicable. Ethics declaration Not applicable. Consent to Publication All the authors have read and agreed to the final copy of the finding as contained in the manuscript. Availability of data and materials The datasets/information used for this study is available on reasonable request to the corresponding authors. Conflicting interest All authors report that there was no conflict of interest in this work. Funding The author(s) received no specific funding for this work. References T.-C. Liu, X. Jin, Y. Wang, and K. Wang, “Role of epidermal growth factor receptor in lung cancer and targeted therapies”. T. D. Barber, B. Vogelstein, K. W. Kinzler, and V. E. Velculescu, “Somatic Mutations of EGFR in Colorectal Cancers and Glioblastomas,” N. Engl. J. Med. , vol. 351, no. 27, pp. 2883–2883, Dec. 2004, doi: 10.1056/NEJM200412303512724. L. V. Sequist et al. , “Rociletinib in EGFR -Mutated Non–Small-Cell Lung Cancer,” N. Engl. J. Med. , vol. 372, no. 18, pp. 1700–1709, Apr. 2015, doi: 10.1056/NEJMoa1413654. P. Uribe and S. Gonzalez, “Epidermal growth factor receptor (EGFR) and squamous cell carcinoma of the skin: Molecular bases for EGFR-targeted therapy,” Pathol. - Res. Pract. , vol. 207, no. 6, pp. 337–342, Jun. 2011, doi: 10.1016/j.prp.2011.03.002. D. L. Wheeler, E. F. Dunn, and P. M. Harari, “Understanding resistance to EGFR inhibitors—impact on future treatment strategies,” Nat. Rev. Clin. Oncol. , vol. 7, no. 9, pp. 493–507, Sep. 2010, doi: 10.1038/nrclinonc.2010.97. M. Maemondo et al. , “Gefitinib or Chemotherapy for Non–Small-Cell Lung Cancer with Mutated EGFR,” N. Engl. J. Med. , vol. 362, no. 25, pp. 2380–2388, Jun. 2010, doi: 10.1056/NEJMoa0909530. M.-S. Tsao et al. , “Erlotinib in Lung Cancer — Molecular and Clinical Predictors of Outcome,” N. Engl. J. Med. , vol. 353, no. 2, pp. 133–144, Jul. 2005, doi: 10.1056/NEJMoa050736. J.-C. Soria et al. , “Osimertinib in Untreated EGFR -Mutated Advanced Non–Small-Cell Lung Cancer,” N. Engl. J. Med. , vol. 378, no. 2, pp. 113–125, Jan. 2018, doi: 10.1056/NEJMoa1713137. A. G. Sacher, P. A. Jänne, and G. R. Oxnard, “Management of acquired resistance to epidermal growth factor receptor kinase inhibitors in patients with advanced non‐small cell lung cancer,” Cancer , vol. 120, no. 15, pp. 2289–2298, Aug. 2014, doi: 10.1002/cncr.28723. A. Koulouris, C. Tsagkaris, A. C. Corriero, G. Metro, and G. Mountzios, “Resistance to TKIs in EGFR-Mutated Non-Small Cell Lung Cancer: From Mechanisms to New Therapeutic Strategies,” Cancers , vol. 14, no. 14, p. 3337, Jul. 2022, doi: 10.3390/cancers14143337. M. Staszak, K. Staszak, K. Wieszczycka, A. Bajek, K. Roszkowski, and B. Tylkowski, “Machine learning in drug design: Use of artificial intelligence to explore the chemical structure–biological activity relationship,” WIREs Comput. Mol. Sci. , vol. 12, no. 2, p. e1568, Mar. 2022, doi: 10.1002/wcms.1568. R. Gupta, D. Srivastava, M. Sahu, S. Tiwari, R. K. Ambasta, and P. Kumar, “Artificial intelligence to deep learning: machine intelligence approach for drug discovery,” Mol. Divers. , vol. 25, no. 3, pp. 1315–1360, Aug. 2021, doi: 10.1007/s11030-021-10217-3. A. Blasiak, J. Khong, and T. Kee, “CURATE.AI: Optimizing Personalized Medicine with Artificial Intelligence,” SLAS Technol. , vol. 25, no. 2, pp. 95–105, Apr. 2020, doi: 10.1177/2472630319890316. M. J. Lamberti et al. , “A Study on the Application and Use of Artificial Intelligence to Support Drug Development,” Clin. Ther. , vol. 41, no. 8, pp. 1414–1426, Aug. 2019, doi: 10.1016/j.clinthera.2019.05.018. X. Chen et al. , “Discovery of Dual FGFR4 and EGFR Inhibitors by Machine Learning and Biological Evaluation,” J. Chem. Inf. Model. , vol. 60, no. 10, pp. 4640–4652, Oct. 2020, doi: 10.1021/acs.jcim.0c00652. A. Gaulton et al. , “ChEMBL: a large-scale bioactivity database for drug discovery,” Nucleic Acids Res. , vol. 40, no. D1, pp. D1100–D1107, Jan. 2012, doi: 10.1093/nar/gkr777. Greg Landrum et al. , rdkit/rdkit: 2025_03_3 (Q1 2025) Release . (Jun. 06, 2025). Zenodo. doi: 10.5281/ZENODO.591637. F. M. Salem, “Gated RNN: The Gated Recurrent Unit (GRU) RNN,” in Recurrent Neural Networks , Cham: Springer International Publishing, 2022, pp. 85–100. doi: 10.1007/978-3-030-89929-5_5. A. P. Lind and P. C. Anderson, “Predicting drug activity against cancer cells by random forest models based on minimal genomic information and chemical properties,” PLOS ONE , vol. 14, no. 7, p. e0219774, Jul. 2019, doi: 10.1371/journal.pone.0219774. Y. Liu, X. Yang, J. Gan, S. Chen, Z.-X. Xiao, and Y. Cao, “CB-Dock2: improved protein–ligand blind docking by integrating cavity detection, docking and homologous template fitting,” Nucleic Acids Res. , vol. 50, no. W1, pp. W159–W164, Jul. 2022, doi: 10.1093/nar/gkac394. J. Stamos, M. X. Sliwkowski, and C. Eigenbrot, “Epidermal Growth Factor Receptor tyrosine kinase domain with 4-anilinoquinazoline inhibitor erlotinib: 1m17.” Sep. 04, 2002. doi: 10.2210/pdb1m17/pdb. Additional Declarations No competing interests reported. Supplementary Files SupplementaryInformation.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6888580","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":477449848,"identity":"376f5746-038d-4aa0-8095-aae4b78656a8","order_by":0,"name":"Vinod Kumar Yata","email":"","orcid":"","institution":"Malla Reddy University","correspondingAuthor":false,"prefix":"","firstName":"Vinod","middleName":"Kumar","lastName":"Yata","suffix":""},{"id":477449852,"identity":"167829ff-26f1-4b02-abd1-d023423c6e90","order_by":1,"name":"Shaik Asma","email":"","orcid":"","institution":"Central University of Andhra Pradesh","correspondingAuthor":false,"prefix":"","firstName":"Shaik","middleName":"","lastName":"Asma","suffix":""},{"id":477449853,"identity":"88170b4d-5981-40ea-958b-e2dd8c85a5e7","order_by":2,"name":"Om Pritam Das","email":"","orcid":"","institution":"Central University of Andhra Pradesh","correspondingAuthor":false,"prefix":"","firstName":"Om","middleName":"Pritam","lastName":"Das","suffix":""},{"id":477449854,"identity":"ebcabb9d-34dd-4978-a2bf-c889716d9ac1","order_by":3,"name":"Venkata Satish Kumar Mattaparthi","email":"","orcid":"","institution":"Tezpur University","correspondingAuthor":false,"prefix":"","firstName":"Venkata","middleName":"Satish Kumar","lastName":"Mattaparthi","suffix":""},{"id":477449855,"identity":"84971a80-9b58-412e-bc29-34ab2d55b036","order_by":4,"name":"Chandrashekar Thalluri","email":"","orcid":"","institution":"Assam down town University (AdtU)","correspondingAuthor":false,"prefix":"","firstName":"Chandrashekar","middleName":"","lastName":"Thalluri","suffix":""},{"id":477449856,"identity":"92737eb5-c70c-4f3a-9df6-3320172d02d7","order_by":5,"name":"Tenyawa Mercy Gladys","email":"","orcid":"","institution":"Kampala International University","correspondingAuthor":false,"prefix":"","firstName":"Tenyawa","middleName":"Mercy","lastName":"Gladys","suffix":""},{"id":477449857,"identity":"ef168067-3757-497e-8e4e-962e7390be70","order_by":6,"name":"Ezeonwumelu Joseph Obiezu Chukwujekwu","email":"","orcid":"","institution":"Kampala International University","correspondingAuthor":false,"prefix":"","firstName":"Ezeonwumelu","middleName":"Joseph Obiezu","lastName":"Chukwujekwu","suffix":""},{"id":477449858,"identity":"669847f9-dab9-476d-95b5-1babd2266680","order_by":7,"name":"Sarad Pawar Naik Bukke","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA80lEQVRIiWNgGAWjYDCCA2CSGcL5cECCRC2MM0jWwsxzgAgdfLePP/xcUWOdxz8j+dlnmzMWiQ3shx8w8/zCrUXyXI6x5Jlj6cUSN9KMZ+fckEhs4EkzYObtw63F4AwPg2QD2+HEhtsJxsw5H4BaGHIYmHl78Glhf/yz4d/hxPm30z8zW4C08L8hpIXBTLKx7XDihts5xswMIIdJAG3h+YHHL2d4zCwb+9KLDe+/KWbsOSNh3CbxzODg3AbcWviADrvZ8M06T+7M8c0MP47VyfbzJz988OYPbi0wkABnsQHxAcY2UrRAABG2jIJRMApGwYgBAI0UVhgzbaM9AAAAAElFTkSuQmCC","orcid":"","institution":"Kampala International University","correspondingAuthor":true,"prefix":"","firstName":"Sarad","middleName":"Pawar Naik","lastName":"Bukke","suffix":""}],"badges":[],"createdAt":"2025-06-13 13:23:12","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6888580/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6888580/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":85667445,"identity":"69631f28-70ba-4911-96ec-5f66ac3e202a","added_by":"auto","created_at":"2025-06-30 13:00:00","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":131077,"visible":true,"origin":"","legend":"\u003cp\u003eOverview of the hybrid computational drug discovery workflow\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6888580/v1/2e27ce2509a5e9bfefee7bbc.jpeg"},{"id":85668853,"identity":"eef3a567-ace0-4bab-9733-ab34f560d348","added_by":"auto","created_at":"2025-06-30 13:16:00","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":113471,"visible":true,"origin":"","legend":"\u003cp\u003eKey property distribution plot of the Dataset molecules\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-6888580/v1/22266f105a398e96377a7221.png"},{"id":85667448,"identity":"af6855e1-3dda-4f80-9425-a763e9deb18c","added_by":"auto","created_at":"2025-06-30 13:00:00","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":65790,"visible":true,"origin":"","legend":"\u003cp\u003eTraining loss curve of the GRU model\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-6888580/v1/25796ecff5740887bad6e6a6.png"},{"id":85667449,"identity":"4f8e9ba1-9f89-4955-a8e6-6698830a2897","added_by":"auto","created_at":"2025-06-30 13:00:00","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":99391,"visible":true,"origin":"","legend":"\u003cp\u003eKey property distribution plot of Generated molecules\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-6888580/v1/44af9bfcebb1bb7db9cb56c1.png"},{"id":85668854,"identity":"bf58ab4f-20dd-4208-b6b3-dae4bcaa248d","added_by":"auto","created_at":"2025-06-30 13:16:00","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":493420,"visible":true,"origin":"","legend":"\u003cp\u003eDocking visualization of the molecules with the EGPR receptor, A) showing the interaction of the reference molecule and B) shows the interaction of the top generated molecule with the EGFR receptor\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-6888580/v1/4d3f42275cf4c7c7883b0bf7.png"},{"id":87045591,"identity":"06133518-90bc-4312-aba6-b0d8f3d53d11","added_by":"auto","created_at":"2025-07-18 14:32:02","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1803260,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6888580/v1/f6cd6b52-809b-418b-92ff-217ba7f5d28b.pdf"},{"id":85667804,"identity":"16bbc900-1f0b-484c-a5c6-b187c1a4749f","added_by":"auto","created_at":"2025-06-30 13:08:00","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":87000,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryInformation.docx","url":"https://assets-eu.researchsquare.com/files/rs-6888580/v1/41330edc43022a139fab9e45.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"AI-driven design of Novel EGFR Inhibitors with Enhanced Potency and Drug-Like Properties","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eThe Epidermal Growth Factor Receptor (EGFR) plays a critical role in cellular proliferation, differentiation, and survival, making it a key target in modern oncology and targeted therapeutics [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Aberrant activation or overexpression of EGFR is strongly implicated in various malignancies, including non-small cell lung cancer (NSCLC), glioblastoma, colorectal cancer, and head and neck squamous cell carcinoma [\u003cspan additionalcitationids=\"CR3\" citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e], Additionally, EGFR dysregulation is associated with resistance to therapy and poor clinical outcomes, underscoring its clinical significance across multiple disease contexts [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. To counteract EGFR-driven oncogenesis, a range of EGFR-targeting drugs, including tyrosine kinase inhibitors (TKIs) such as gefitinib [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], erlotinib [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], and osimertinib [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], have been developed and are widely used in cancer therapy. While these agents have shown efficacy, several challenges remain: resistance mutations, off-target toxicity, and limited drug affinity toward mutant EGFR variants often reduce therapeutic effectiveness [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], In recent years, machine learning (ML) and deep learning (DL) approaches have emerged as powerful tools in drug discovery, offering the ability to explore chemical space more efficiently, predict molecular activity, and design novel compounds with improved specificity and pharmacological profiles [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Previous studies have applied AI-driven techniques to accelerate drug development pipelines, with promising results in generating bioactive compounds and enhancing drug-target interactions [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. However, there remains a pressing need for methods that not only automate de novo drug design but also optimize molecular affinity and drug-likeness to yield therapeutically viable candidates [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. This motivates further integration of AI into medicinal chemistry workflows, especially in the design of EGFR inhibitors where small structural changes can significantly impact binding and efficacy.\u003c/p\u003e \u003cp\u003eThe objective of this study is to develop and validate a deep learning-based pipeline for the design of novel EGFR inhibitors with high predicted potency and favorable drug-like properties. We propose an integrated framework that combines generative modeling using a Gated Recurrent Unit (GRU)-based Recurrent Neural Network (RNN), IC₅₀ prediction via Random Forest regression, and structure-based molecular docking to assess binding affinity. By training our models on a curated set of potent EGFR inhibitors, we aim to generate novel, chemically valid compounds that not only mimic but potentially surpass the activity of known drugs. This approach demonstrates how AI can enhance early-stage drug discovery by guiding the rational design of lead candidates with optimized pharmacological profiles.\u003c/p\u003e"},{"header":"2. Methodology","content":"\u003cp\u003eThe hybrid de novo drug design pipeline, meticulously developed to identify and validate novel EGFR inhibitors, systematically integrates several advanced computational stages. The overarching workflow, as depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, progresses from the initial acquisition and rigorous preprocessing of large-scale bioactivity data, through the development of a sophisticated generative molecular design model and a robust predictive property model, culminating in a comprehensive candidate scoring framework and structural validation via molecular docking.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Data Collection and Pre-processing\u003c/h2\u003e \u003cp\u003eWe collected EGFR inhibitor data from the ChEMBL database, selecting molecules with experimentally measured IC₅₀ values below 100 nM to ensure high potency [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. The raw data was carefully cleaned to remove duplicates, incomplete records, and inconsistent entries. After this process, 260 molecules remained in the final curated dataset (listed in Supplementary Table\u0026nbsp;1). This dataset was used to train and validate the predictive model. For each molecule, six key molecular descriptors were calculated using RDKit, Molecular Weight (MW), LogP (octanol-water partition coefficient), number of hydrogen bond donors (HBD), number of hydrogen bond acceptors (HBA), topological polar surface area (TPSA), and quantitative estimate of drug-likeness (QED) [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. These descriptors were chosen for their relevance to biological activity and drug-likeness. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e provides a statistical summary of the major properties across the dataset.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Model Training and SMILES Generation\u003c/h2\u003e \u003cp\u003eA Recurrent Neural Network (RNN) with a single Gated Recurrent Unit (GRU) layer of hidden size 256 was employed to generate novel EGFR inhibitor molecules as SMILES strings [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. The GRU architecture was selected for its efficiency in capturing sequential dependencies in SMILES data while using fewer parameters than more complex recurrent units. The model was trained on the curated dataset (see Supplementary Table\u0026nbsp;1) using the Adam optimizer with an initial learning rate of 0.005. Training was conducted for 20 epochs with a batch size of 32, providing a balance between model convergence and computational cost. After training, the model generated 42 novel SMILES strings, 40 of which were confirmed as chemically valid by RDKit. This high validity rate indicates effective learning of chemical structure rules by the model.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 IC50 Prediction and Evaluation of New Molecules\u003c/h2\u003e \u003cp\u003eA pre-trained Random Forest regression model was used to predict the IC50 values (in nanomolar, nM) of the generated molecules, providing an estimate of their inhibitory potency [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. This model was independently trained on the curated dataset of known EGFR inhibitors, utilizing seven molecular descriptors as input features: Molecular Weight (MW), LogP, number of Hydrogen Bond Donors (HBD), number of Hydrogen Bond Acceptors (HBA), Topological Polar Surface Area (TPSA), number of Rotatable Bonds, and Quantitative Estimation of Drug-likeness (QED). The Random Forest approach was selected for its robustness and ability to model complex relationships without extensive feature engineering. Predicted IC50 values were generated for all valid SMILES produced by the RNN model. To prioritize promising candidates, a multi-criteria scoring system was applied, integrating predicted IC50 with drug-likeness metrics (QED, LogP, TPSA). This system classified molecules into \u0026ldquo;Excellent,\u0026rdquo; \u0026ldquo;High,\u0026rdquo; or \u0026ldquo;Rejected\u0026rdquo; categories, facilitating the identification of lead compounds with balanced potency and favourable drug-like properties for further study.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Molecular Docking Analysis\u003c/h2\u003e \u003cp\u003eMolecular docking was performed to evaluate the binding interactions between the generated compounds and the EGFR active site, providing structural validation of their predicted inhibitory potential. Docking simulations were conducted using CB-Dock2, an automated protein-ligand docking server that identifies potential binding cavities and performs flexible docking [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. The EGFR crystal structure (PDB ID: 1M17) was directly obtained from the RCSB Protein Data Bank and used as the receptor model [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Both the top-ranked generated compounds and a reference inhibitor from the training dataset were docked for comparison of binding affinities and modes. The reference ligand structure was retrieved as an SDF file from ChEMBL, while the top-ranked generated molecule\u0026rsquo;s 3D structure was created using the RDKit cheminformatics library [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. Docking results included predicted binding energies and detailed interaction analyses, highlighting key residues involved in ligand binding. These findings provided atomistic insights supporting the predicted potency of the novel compounds.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Dataset Characterization\u003c/h2\u003e \u003cp\u003eThe curated dataset (Supplementary Table\u0026nbsp;1) comprises 260 potent EGFR inhibitors, selected for strong activity (IC50\u0026thinsp;\u0026lt;\u0026thinsp;100 nM) and favorable drug-like properties (QED\u0026thinsp;\u0026gt;\u0026thinsp;0.7). Key physicochemical characteristics are summarized in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, showing that most compounds maintain balanced lipophilicity (LogP) conducive to membrane permeability and a polar surface area (TPSA) range favorable for oral absorption (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The dataset covers drug-like chemical space relevant to effective EGFR inhibition, exhibiting features typical of orally bioavailable small molecules such as moderate molecular weight and appropriate hydrogen bonding capacity. This diversity in polarity and hydrogen bonding supports absorption and target engagement, indicating the dataset\u0026rsquo;s suitability for training predictive models with broad applicability in EGFR inhibitor design.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSummary Statistics of the Dataset Molecules\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProperty\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMean\u0026thinsp;\u0026plusmn;\u0026thinsp;Standard Deviation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMin Value\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMax Value\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eIC50 (nM)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e24.89\u0026thinsp;\u0026plusmn;\u0026thinsp;26.68\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.005\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e99.99\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMW (Da)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e375.31\u0026thinsp;\u0026plusmn;\u0026thinsp;65.55\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e175.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e641.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLogP\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e3.67\u0026thinsp;\u0026plusmn;\u0026thinsp;0.90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.06\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6.86\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eQED\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e0.77\u0026thinsp;\u0026plusmn;\u0026thinsp;0.05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.94\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHBD\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e1.26\u0026thinsp;\u0026plusmn;\u0026thinsp;0.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHBA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e4.59\u0026thinsp;\u0026plusmn;\u0026thinsp;1.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTPSA (A˚2)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c2\"\u003e \u003cp\u003e81.76\u0026thinsp;\u0026plusmn;\u0026thinsp;23.36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e12.03\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e178.69\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Model Training and SMILES Generation\u003c/h2\u003e \u003cp\u003eThe RNN model with a single GRU layer was trained on the curated dataset for 20 epochs using the Adam optimizer. The training process demonstrated consistent convergence, with the loss steadily decreasing as shown in the loss curve (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e), indicating effective learning of SMILES syntax and chemical features. After training, the model generated 42 novel SMILES strings. Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e lists all generated molecules alongside their predicted physicochemical characteristics, including IC50, QED, LogP, and TPSA. Of these, 40 (95%) were confirmed as chemically valid using RDKit. The distribution of these key properties among the valid molecules is further summarized in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. These results suggest that the model effectively generated drug-like compounds consistent with the training data, demonstrating successful exploration of relevant chemical space.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAll Generated molecules\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSMILES\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePredicted IC50 (nM)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLogP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eQED (drug-likeness)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eTPSA\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eScore\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNC1CCN(Cc2ccn3ncnc(Nc4cccc(Br)c4)c23)CC1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e9.38777\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.1585\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.702976\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e71.48\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC#Cc1cccc(Nc2ncnc3cc4c(cc23)OCCO4)c1F\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e23.8878\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.265\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.734284\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e56.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCN(CC)c1cc2ncnc(Nc3cccc(Cl)c3)c2cc1OC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e28.58245\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.4915\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.739861\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e50.28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOc1cc2ncnc(Nc3cccc(Br)c3)c2cn1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e27.1548\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.5395\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.795412\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e59.93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCC(=\u0026thinsp;O)Nc1cc2c(Nc3cccc(Cl)c3)ncnc2cn1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e20.0138\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.3802\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.774545\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e79.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOc1cc2c(Nc3cccc(C)c3)ncnc2cn1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e31.99049\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.08542\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.78901\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e59.93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNc1cc2c(Nc3cccc(Cl)c3)ncnc2cn1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e15.89321\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.004\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.749152\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e76.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOc1cc2ncnc(Nc3ccccc3)c2cc1OC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e23.77667\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.3906\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.79404\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e56.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOc1cc2ncnc(Nc3ccccc3)c2cc1OC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e23.77667\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.3906\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.79404\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e56.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOc1cc2ncnc(Nc3cccc(Br)c3)c2cc1OC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e2.508847\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.1531\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.757494\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e56.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCN(CCO)c1cc2ncnc(Nc3ccccc3)c2cc1OC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e19.95869\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.8105\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.72644\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e70.51\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBrc1cccc(Nc2ncnc3[nH]cnc23)c1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e14.26108\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.859\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.761278\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e66.49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOc1cc2ncnc(Nc3cccc(C)c3)c2cn1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e31.99049\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.08542\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.78901\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e59.93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC#Cc1cccc(Nc2ncnc3cc4c(cc23)OCCO4)c1F\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e23.8878\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.265\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.734284\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e56.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCCC(=\u0026thinsp;O)Nc1cc2c(Nc3cccc(C)c3)ncnc2cn1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e22.26139\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.42532\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.771674\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e79.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNc1cc2c(Nc3ccccc3)ncnc2cn1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e13.69668\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.3506\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.715026\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e76.72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eOCC(C#N)c1ccc2ncnc(Nc3cccc(Cl)c3)c2c1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e25.77879\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.62628\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.765526\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e81.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC\u0026thinsp;=\u0026thinsp;CC(=\u0026thinsp;O)NC1\u0026thinsp;=\u0026thinsp;CC\u0026thinsp;=\u0026thinsp;C2N\u0026thinsp;=\u0026thinsp;CN\u0026thinsp;=\u0026thinsp;C(Nc3cccc(Cl)c3)C21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5.925699\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.8921\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.842818\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e65.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC#Cc1cccc(Nc2ncnc3[nH]c4c(c23)CCCN4)c1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e18.33806\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.0409\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.634417\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e65.63\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFc1c(Br)cccc1Nc1ncnc2ccncc12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e20.24674\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.67\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.782892\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e50.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOc1ccc2c(c1)CCCN2c1cccc2c(Br)cccc1N1Cc2ccccc1F\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e55.1417\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.6133\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.433161\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e15.71\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eRejected\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC\u0026thinsp;=\u0026thinsp;CC(=\u0026thinsp;O)Nc1ccc2ncnc(Nc3cccc(Br)c3F)c2c1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e16.18201\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.3995\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.651667\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e66.91\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCNc1ccc(N(C)c2ncnc3occ(C)c23)cc1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e14.41047\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.34082\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.78917\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e54.19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCc1coc2ncnc(Nc3ccc(F)c(Cl)c3)c2cc1OC.Cl\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e39.8499\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.718\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.847578\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e56.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCCN(c1ccc(OC)cc1)c1ncnc2occ(C)c12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e74.57717\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.69782\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.730437\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e51.39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBrc1ccc2ncnc(Nc3ccccc3F)c2c1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e51.10944\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.275\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.766422\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e37.81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eN#Cc1cccc(Nc2ncnc3ccncc23)c1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e24.08511\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.64008\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.753028\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e74.49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC\u0026thinsp;=\u0026thinsp;CC(=\u0026thinsp;O)NC1\u0026thinsp;=\u0026thinsp;CC\u0026thinsp;=\u0026thinsp;C2N\u0026thinsp;=\u0026thinsp;CN\u0026thinsp;=\u0026thinsp;C(Nc3cccc(C)c3)C21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e7.077492\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.54712\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.840799\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e65.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOc1ccc2ncnc(Nc3cccc(Cl)c3)C2[nH]1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e29.9498\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.5358\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.87911\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e58.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC#Cc1cccc(Nc2ncnc3[nH]c4c(c23)CCCC4)c1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e24.90018\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.5616\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.708878\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e53.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC\u0026thinsp;=\u0026thinsp;CC(=\u0026thinsp;O)NC1\u0026thinsp;=\u0026thinsp;CC\u0026thinsp;=\u0026thinsp;C2N\u0026thinsp;=\u0026thinsp;CN\u0026thinsp;=\u0026thinsp;C(Nc3cccc(Cl)c3)C21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5.925699\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.8921\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.842818\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e65.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eClc1cccc(Nc2ncnc3cccnc(NC(C)C)cc23)c1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e34.6214\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.3584\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.90082\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e62.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC#Cc1cccc(Nc2ncnc3cnccc23)c1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e37.22691\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.7497\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.706272\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e50.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC#Cc1cccc(Nc2ncnc3ccccc23)c1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e34.223\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.3547\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.703415\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e37.81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC\u0026thinsp;=\u0026thinsp;CC(=\u0026thinsp;O)NC1\u0026thinsp;=\u0026thinsp;CC\u0026thinsp;=\u0026thinsp;C2N\u0026thinsp;=\u0026thinsp;CN\u0026thinsp;=\u0026thinsp;C(Nc3cccc(C)c3)C21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e7.077492\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e2.54712\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.840799\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e65.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eBrc1cccc(Nc2ncnc3ccccc23)c1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e32.76127\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e4.1359\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.772775\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e37.81\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eC#Cc1cccc(Nc2ncnc3cc4c(cc23)OCCO4)c1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e36.67387\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.1259\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.737147\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e56.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOc1cc2ncnc(Nc3ccccc3)c2cc1OC(F)F\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e12.85114\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.9834\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.771056\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e56.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOc1cc2ncnc(Nc3cccc(C)c3)c2cc1OC\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e12.1738\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.69902\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.795048\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e56.27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCC(=\u0026thinsp;O)Nc1cc2c(Nc3cccc(Cl)c3)ncnc2cn1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e20.0138\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.3802\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.774545\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e79.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eExcellent\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCOc1ccc(N(C)c2nc(C)nc3oc(C)cc23)cc1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e59.9001\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e3.61614\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.734535\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e51.39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eHigh\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e3.3 IC50 Prediction and Evaluation of New Molecules\u003c/h2\u003e \u003cp\u003eThe inhibitory potency of the generated molecules was predicted using a pre-trained Random Forest regression model based on molecular descriptors. The predicted IC50 values for the 40 valid molecules spanned a broad range, with several compounds showing low nanomolar values indicative of strong EGFR inhibition potential (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). To prioritize candidates, a multi-criteria scoring system was applied, combining predicted IC50 with key drug-likeness parameters such as QED, LogP, and TPSA. Molecules were scored against thresholds designed to capture favorable potency alongside balanced physicochemical properties essential for oral bioavailability and effective target engagement. This approach helped identify compounds with both high predicted activity and suitable pharmacokinetic profiles. Based on this scoring system, molecules were categorized as \u0026ldquo;Excellent,\u0026rdquo; \u0026ldquo;High,\u0026rdquo; or \u0026ldquo;Rejected\u0026rdquo;. Of the valid molecules, 11 (26%) were classified as \u0026ldquo;Excellent,\u0026rdquo; demonstrating a favorable balance of predicted potency and drug-like characteristics, making them strong candidates for further evaluation. The majority, 29 (69%), were classified as \u0026ldquo;High\u0026rdquo; potential, while only 2 molecules (5%) were \u0026ldquo;Rejected\u0026rdquo; due to less optimal profiles (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). These results highlight the pipeline\u0026rsquo;s effectiveness in enriching for promising EGFR inhibitor candidates.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Docking Analysis\u003c/h2\u003e \u003cp\u003eMolecular docking was performed to assess the binding interactions of the generated compounds within the EGFR active site, providing structure-based validation of their predicted inhibitory potential. Docking simulations utilized the crystal structure of EGFR (PDB ID: 1M17), directly obtained from the Protein Data Bank [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. A reference inhibitor from the training dataset (CHEMBL29197) [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e], with an experimental IC50 of 0.025 nM, was docked to establish a benchmark. This compound showed a predicted binding energy of \u0026minus;\u0026thinsp;7.3 kcal/mol, interacting with key residues known to be critical for EGFR inhibition. Among the generated molecules, the top-ranked compound demonstrated a slightly improved predicted binding affinity of \u0026minus;\u0026thinsp;7.5 kcal/mol. Its binding mode involved interactions with essential residues in the hinge region (e.g., LEU694, GLY695, CYS773) and hydrophobic pockets (e.g., PHE699), similar to the reference ligand. These docking results support the predicted potency from the IC50 model and suggest that the generative pipeline can produce novel compounds with favorable binding characteristics comparable or superior to known inhibitors. Detailed docking scores and interaction residues for key compounds are summarized in Table \u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. Figure\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e illustrates the docked poses of both the reference and generated molecules within the EGFR binding site, highlighting their key interactions.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eDocking table\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eComplex Type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSymmetry\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBinding Affinity (kcal/mol)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eResidues\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eHydrogen Bonds (Donor, Acceptor, Total)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eHydrophobic Interactions\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eInteracting Residues (Chain A)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProtein-standard drug complex\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eC2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e-7.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e665\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e24, 0, 54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e21, 21, 21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eALA698, PHE699, VAL702, LYS721, GLU722, LEU723, SER728, ALA731, LYS733, GLU734, ILE735, ASP737, GLU738, TYR740, VAL741, ASP813, ARG817, ASN818, ASP831, PHE832, GLY833, LEU834, ALA835, LYS836, LEU837, LYS851, PRO853\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eProtein-new drug complex\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eC2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e-7.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e665\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e24, 0, 54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e20, 20, 20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003eLEU694, GLY695, PHE699, VAL702, ALA719, ILE720, LYS721, GLU738, MET742, LEU753, LEU764, ILE765, THR766, GLN767, LEU768, MET769, GLY772, CYS773, ARG817, ASN818, LEU820, THR830, ASP831\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eThis study demonstrates an integrated approach combining generative deep learning, predictive modeling, and molecular docking to identify novel small-molecule EGFR inhibitors. The results indicate that the pipeline effectively generated and prioritized compounds with favorable pharmacological and structural properties. The curated training dataset consisted of high-quality EGFR inhibitors with strong potency and drug-likeness, ensuring a reliable foundation for model training. The RNN model, built with a GRU-based architecture, showed stable training behavior, as reflected in the convergence of the loss curve. Its ability to generate syntactically and chemically valid SMILES strings (95% validity) underscores the model\u0026rsquo;s capacity to learn meaningful chemical representations. Importantly, the generated molecules retained key drug-like characteristics, with their physicochemical profiles aligning well with known EGFR inhibitors. The IC50 prediction model, trained separately using Random Forest regression, successfully distinguished molecules with higher predicted potency. This was supported by the scoring system, which integrated potency with drug-likeness properties such as QED, LogP, and TPSA. The classification system proved effective in filtering the generated compounds\u0026mdash;over 95% were categorized as either \u0026ldquo;Excellent\u0026rdquo; or \u0026ldquo;High,\u0026rdquo; demonstrating the selectivity and utility of the combined scoring approach. Structure-based validation through molecular docking provided further support for the identified lead candidates. The top-ranked generated compound demonstrated binding energy comparable to that of a known potent EGFR inhibitor, and its binding interactions involved key residues essential for kinase inhibition. These results suggest that the de novo designed molecule may exhibit meaningful biological activity, warranting further \u003cem\u003ein-vitro\u003c/em\u003e or in-vivo evaluation.\u003c/p\u003e \u003cp\u003eDespite these encouraging findings, there are several limitations. First, the IC50 predictions, while useful for prioritization, remain dependent on the accuracy of the descriptors and training data; they do not substitute for experimental confirmation. Second, the docking simulations relied on a single static protein structure and assumed rigid receptor conformation, which may not fully capture the dynamics of ligand binding. Additionally, no ADMET (absorption, distribution, metabolism, excretion, and toxicity) modeling was incorporated, which would be essential in downstream optimization stages. Nevertheless, the pipeline successfully demonstrates a low-cost and scalable strategy for early-stage drug discovery. The combined use of deep generative models, predictive regression, and molecular docking highlights the potential of AI-driven design to augment traditional screening efforts, particularly in well-characterized targets like EGFR. Future work will focus on experimental validation of the top candidates, incorporation of molecular dynamics for more accurate binding assessment, and extension of the approach to other targets and compound classes.\u003c/p\u003e"},{"header":"5. Conclusion","content":"\u003cp\u003eThis research presents a comprehensive AI-driven framework for the de novo design and evaluation of EGFR inhibitors, integrating generative deep learning, predictive modeling, and molecular docking. By leveraging a curated dataset of potent EGFR inhibitors, a GRU-based RNN model successfully generated novel chemical structures with a high validity rate and favorable drug-like properties. Subsequent IC₅₀ prediction using a Random Forest regression model enabled the prioritization of candidates with high predicted potency. Further validation through molecular docking confirmed that several generated compounds exhibited binding affinities and interaction profiles comparable to a known reference inhibitor. The results underscore the effectiveness of combining machine learning and cheminformatics in early-stage drug discovery, especially for well-characterized targets like EGFR. While limitations such as the absence of ADMET modeling and reliance on static docking models remain, this study provides a scalable and efficient pipeline for identifying promising lead compounds. Future directions include experimental validation, integration of pharmacokinetic modeling, and application of this approach to broader therapeutic targets.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eEGFR: Epidermal Growth Factor Receptor, GRU: Gated Recurrent Unit, RNN: Recurrent Neural Network, IC₅₀: Predicted inhibitory potency, NSCLS: non-small cell lung cancer, TKIs: tyrosine kinase inhibitors, ML: machine learning, AI: Artificial intelligence, HBD: hydrogen bond donors, HBA: hydrogen bond acceptors, TPSA: topological polar surface area, QED: quantitative estimate of drug-likeness, MW: Molecular Weight, ADMET: absorption, distribution, metabolism, excretion, and toxicity.\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eI hereby declare that this submission is entirely my own work, in my own words, and that all sources used in researching it are fully acknowledged and all quotations properly identified.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors thank to\u0026nbsp;Department of Clinical research and Experimental Medicine, School of Allied and Healthcare Sciences, Malla Reddy University, Hyderabad, India\u0026nbsp;for supporting this research work.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contribution Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eVKY and VSKM contributed to the study\u0026apos;s ideation, data analysis, and drafting of the original manuscript. OPD and SA was responsible for data collection, curation, and interpretation. CT provided supervision and contributed to the drafting and refinement of the manuscript. EJOC and TMG supported the project by providing resources, supervision, and critical review. SPNB additionally led the conceptualization and methodology design of the study. The authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eStatement of Informed Consent\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThere are no human subjects in this article and informed consent is not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eClinical trial number\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to Publish declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics declaration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to Publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll the authors have read and agreed to the final copy of the finding as contained in the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe datasets/information used for this study is available on reasonable request to the corresponding authors.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflicting interest\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll authors report that there was no conflict of interest in this work.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author(s) received no specific funding for this work.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eT.-C. Liu, X. Jin, Y. Wang, and K. Wang, \u0026ldquo;Role of epidermal growth factor receptor in lung cancer and targeted therapies\u0026rdquo;.\u003c/li\u003e\n\u003cli\u003eT. D. Barber, B. Vogelstein, K. W. Kinzler, and V. E. Velculescu, \u0026ldquo;Somatic Mutations of \u003cem\u003eEGFR\u003c/em\u003e in Colorectal Cancers and Glioblastomas,\u0026rdquo; \u003cem\u003eN. Engl. J. Med.\u003c/em\u003e, vol. 351, no. 27, pp. 2883\u0026ndash;2883, Dec. 2004, doi: 10.1056/NEJM200412303512724.\u003c/li\u003e\n\u003cli\u003eL. V. Sequist \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Rociletinib in \u003cem\u003eEGFR\u003c/em\u003e -Mutated Non\u0026ndash;Small-Cell Lung Cancer,\u0026rdquo; \u003cem\u003eN. Engl. J. Med.\u003c/em\u003e, vol. 372, no. 18, pp. 1700\u0026ndash;1709, Apr. 2015, doi: 10.1056/NEJMoa1413654.\u003c/li\u003e\n\u003cli\u003eP. Uribe and S. Gonzalez, \u0026ldquo;Epidermal growth factor receptor (EGFR) and squamous cell carcinoma of the skin: Molecular bases for EGFR-targeted therapy,\u0026rdquo; \u003cem\u003ePathol. - Res. Pract.\u003c/em\u003e, vol. 207, no. 6, pp. 337\u0026ndash;342, Jun. 2011, doi: 10.1016/j.prp.2011.03.002.\u003c/li\u003e\n\u003cli\u003eD. L. Wheeler, E. F. Dunn, and P. M. Harari, \u0026ldquo;Understanding resistance to EGFR inhibitors\u0026mdash;impact on future treatment strategies,\u0026rdquo; \u003cem\u003eNat. Rev. Clin. Oncol.\u003c/em\u003e, vol. 7, no. 9, pp. 493\u0026ndash;507, Sep. 2010, doi: 10.1038/nrclinonc.2010.97.\u003c/li\u003e\n\u003cli\u003eM. Maemondo \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Gefitinib or Chemotherapy for Non\u0026ndash;Small-Cell Lung Cancer with Mutated EGFR,\u0026rdquo; \u003cem\u003eN. Engl. J. Med.\u003c/em\u003e, vol. 362, no. 25, pp. 2380\u0026ndash;2388, Jun. 2010, doi: 10.1056/NEJMoa0909530.\u003c/li\u003e\n\u003cli\u003eM.-S. Tsao \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Erlotinib in Lung Cancer \u0026mdash; Molecular and Clinical Predictors of Outcome,\u0026rdquo; \u003cem\u003eN. Engl. J. Med.\u003c/em\u003e, vol. 353, no. 2, pp. 133\u0026ndash;144, Jul. 2005, doi: 10.1056/NEJMoa050736.\u003c/li\u003e\n\u003cli\u003eJ.-C. Soria \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Osimertinib in Untreated \u003cem\u003eEGFR\u003c/em\u003e -Mutated Advanced Non\u0026ndash;Small-Cell Lung Cancer,\u0026rdquo; \u003cem\u003eN. Engl. J. Med.\u003c/em\u003e, vol. 378, no. 2, pp. 113\u0026ndash;125, Jan. 2018, doi: 10.1056/NEJMoa1713137.\u003c/li\u003e\n\u003cli\u003eA. G. Sacher, P. A. J\u0026auml;nne, and G. R. Oxnard, \u0026ldquo;Management of acquired resistance to epidermal growth factor receptor kinase inhibitors in patients with advanced non‐small cell lung cancer,\u0026rdquo; \u003cem\u003eCancer\u003c/em\u003e, vol. 120, no. 15, pp. 2289\u0026ndash;2298, Aug. 2014, doi: 10.1002/cncr.28723.\u003c/li\u003e\n\u003cli\u003eA. Koulouris, C. Tsagkaris, A. C. Corriero, G. Metro, and G. Mountzios, \u0026ldquo;Resistance to TKIs in EGFR-Mutated Non-Small Cell Lung Cancer: From Mechanisms to New Therapeutic Strategies,\u0026rdquo; \u003cem\u003eCancers\u003c/em\u003e, vol. 14, no. 14, p. 3337, Jul. 2022, doi: 10.3390/cancers14143337.\u003c/li\u003e\n\u003cli\u003eM. Staszak, K. Staszak, K. Wieszczycka, A. Bajek, K. Roszkowski, and B. Tylkowski, \u0026ldquo;Machine learning in drug design: Use of artificial intelligence to explore the chemical structure\u0026ndash;biological activity relationship,\u0026rdquo; \u003cem\u003eWIREs Comput. Mol. Sci.\u003c/em\u003e, vol. 12, no. 2, p. e1568, Mar. 2022, doi: 10.1002/wcms.1568.\u003c/li\u003e\n\u003cli\u003eR. Gupta, D. Srivastava, M. Sahu, S. Tiwari, R. K. Ambasta, and P. Kumar, \u0026ldquo;Artificial intelligence to deep learning: machine intelligence approach for drug discovery,\u0026rdquo; \u003cem\u003eMol. Divers.\u003c/em\u003e, vol. 25, no. 3, pp. 1315\u0026ndash;1360, Aug. 2021, doi: 10.1007/s11030-021-10217-3.\u003c/li\u003e\n\u003cli\u003eA. Blasiak, J. Khong, and T. Kee, \u0026ldquo;CURATE.AI: Optimizing Personalized Medicine with Artificial Intelligence,\u0026rdquo; \u003cem\u003eSLAS Technol.\u003c/em\u003e, vol. 25, no. 2, pp. 95\u0026ndash;105, Apr. 2020, doi: 10.1177/2472630319890316.\u003c/li\u003e\n\u003cli\u003eM. J. Lamberti \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;A Study on the Application and Use of Artificial Intelligence to Support Drug Development,\u0026rdquo; \u003cem\u003eClin. Ther.\u003c/em\u003e, vol. 41, no. 8, pp. 1414\u0026ndash;1426, Aug. 2019, doi: 10.1016/j.clinthera.2019.05.018.\u003c/li\u003e\n\u003cli\u003eX. Chen \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Discovery of Dual FGFR4 and EGFR Inhibitors by Machine Learning and Biological Evaluation,\u0026rdquo; \u003cem\u003eJ. Chem. Inf. Model.\u003c/em\u003e, vol. 60, no. 10, pp. 4640\u0026ndash;4652, Oct. 2020, doi: 10.1021/acs.jcim.0c00652.\u003c/li\u003e\n\u003cli\u003eA. Gaulton \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;ChEMBL: a large-scale bioactivity database for drug discovery,\u0026rdquo; \u003cem\u003eNucleic Acids Res.\u003c/em\u003e, vol. 40, no. D1, pp. D1100\u0026ndash;D1107, Jan. 2012, doi: 10.1093/nar/gkr777.\u003c/li\u003e\n\u003cli\u003eGreg Landrum \u003cem\u003eet al.\u003c/em\u003e, \u003cem\u003erdkit/rdkit: 2025_03_3 (Q1 2025) Release\u003c/em\u003e. (Jun. 06, 2025). Zenodo. doi: 10.5281/ZENODO.591637.\u003c/li\u003e\n\u003cli\u003eF. M. Salem, \u0026ldquo;Gated RNN: The Gated Recurrent Unit (GRU) RNN,\u0026rdquo; in \u003cem\u003eRecurrent Neural Networks\u003c/em\u003e, Cham: Springer International Publishing, 2022, pp. 85\u0026ndash;100. doi: 10.1007/978-3-030-89929-5_5.\u003c/li\u003e\n\u003cli\u003eA. P. Lind and P. C. Anderson, \u0026ldquo;Predicting drug activity against cancer cells by random forest models based on minimal genomic information and chemical properties,\u0026rdquo; \u003cem\u003ePLOS ONE\u003c/em\u003e, vol. 14, no. 7, p. e0219774, Jul. 2019, doi: 10.1371/journal.pone.0219774.\u003c/li\u003e\n\u003cli\u003eY. Liu, X. Yang, J. Gan, S. Chen, Z.-X. Xiao, and Y. Cao, \u0026ldquo;CB-Dock2: improved protein\u0026ndash;ligand blind docking by integrating cavity detection, docking and homologous template fitting,\u0026rdquo; \u003cem\u003eNucleic Acids Res.\u003c/em\u003e, vol. 50, no. W1, pp. W159\u0026ndash;W164, Jul. 2022, doi: 10.1093/nar/gkac394.\u003c/li\u003e\n\u003cli\u003eJ. Stamos, M. X. Sliwkowski, and C. Eigenbrot, \u0026ldquo;Epidermal Growth Factor Receptor tyrosine kinase domain with 4-anilinoquinazoline inhibitor erlotinib: 1m17.\u0026rdquo; Sep. 04, 2002. doi: 10.2210/pdb1m17/pdb.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"EGFR Inhibitors, Deep Learning, Drug Design, Recurrent Neural Network (RNN), Molecular Docking, IC₅₀ Prediction","lastPublishedDoi":"10.21203/rs.3.rs-6888580/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6888580/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe Epidermal Growth Factor Receptor (EGFR) is a clinically validated target in oncology, with aberrant EGFR signaling implicated in several cancers, including non-small cell lung cancer and glioblastoma. While existing EGFR inhibitors have shown therapeutic success, challenges such as acquired resistance, suboptimal drug-likeness, and limited affinity for mutant receptor variants persist. In this study, we present an integrated deep learning-based pipeline for the de novo design of novel EGFR inhibitors with improved potency and drug-like characteristics. A curated dataset of 260 high-affinity EGFR inhibitors was used to train a Gated Recurrent Unit (GRU)-based Recurrent Neural Network (RNN), enabling the generation of novel SMILES structures. Of the 42 generated compounds, 40 were chemically valid, as confirmed by RDKit. Predicted inhibitory potency (IC₅₀) was estimated using a Random Forest regression model trained on molecular descriptors, and a multi-criteria scoring system was applied to prioritize compounds based on IC₅₀, QED, LogP, and TPSA. Molecular docking studies using the EGFR crystal structure (PDB ID: 1M17) validated the binding potential of top candidates, with the highest-ranked compound showing binding affinity comparable to a known reference inhibitor. This work highlights the potential of AI-driven methods in accelerating early-stage drug discovery and offers promising candidates for further experimental validation as next-generation EGFR inhibitors.\u003c/p\u003e","manuscriptTitle":"AI-driven design of Novel EGFR Inhibitors with Enhanced Potency and Drug-Like Properties","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-30 12:59:55","doi":"10.21203/rs.3.rs-6888580/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"01b06744-1218-469e-a2b3-6f7427c8cb0b","owner":[],"postedDate":"June 30th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-07-18T14:23:52+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-30 12:59:55","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6888580","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6888580","identity":"rs-6888580","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.