Computational Prediction of Drug-Induced Hematotoxicity: Mechanisms and Model Development | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Computational Prediction of Drug-Induced Hematotoxicity: Mechanisms and Model Development Jian Xu, Mengfei Zhang, Leyu Tao, Shuying Zhang This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6295317/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Hematotoxicity, encompassing adverse effects such as anemia, leukopenia, thrombocytopenia, and coagulation disorders, is a critical yet underexplored area of toxicological research. These toxic effects can lead to severe clinical outcomes, including heightened risks of infection, bleeding, and mortality. Despite its significance, hematotoxicity research lags behind general toxicities like hepatotoxicity and nephrotoxicity, and traditional evaluation methods such as animal models and in vitro assays often fail to accurately predict human responses. To address this gap, we curated a dataset of thousands of compounds with and without hematotoxic effects and performed in-depth analyses of their molecular properties and mechanisms using clustering and target prediction. These analyses revealed key pathways and targets underlying hematotoxicity, demonstrating its complex and multifactorial nature. We developed predictive models using fingerprints, combined with machine learning algorithms such as Random Forest (RF), Support Vector Machine (SVM), and XGBoost. The best-performing model achieved an AUC of 0.78, highlighting its potential for accurately identifying hematotoxic compounds. This study provides a computational framework for understanding hematotoxicity mechanisms and offers a practical tool for early screening of blood-toxic compounds during drug development. These advancements pave the way for safer therapeutic strategies, improved patient safety, and reduced risks of drug-induced hematological disorders. Hematotoxicity Machine learning Network Pharmacology Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Figure 13 Figure 14 1. Introduction Blood plays a crucial role in the human body by transporting oxygen, nutrients, and waste products, maintaining normal physiological activities[ 1 ]. Hematotoxicity refers to the potential adverse effects of certain drugs on the hematological system, including abnormalities in red blood cells, white blood cells, and platelets, as well as disruptions in coagulation functions[ 2 ]. The main clinical manifestations of this toxicity include anemia, leukopenia, thrombocytopenia, and a tendency to bleed. Since the impact of drugs on the blood system is often overlooked, hematotoxicity can lead to serious consequences such as increased risk of infection, hemorrhage, and even death[ 3 ]. The mechanisms may involve direct damage to hematopoietic stem cells, immune-mediated cell destruction, or interference with signaling pathways that regulate cell proliferation and differentiation[ 4 ]. Several drugs have been reported to exhibit hematotoxicity. For example, chemotherapy agents like cyclophosphamide and doxorubicin can cause bone marrow suppression, leading to pancytopenia[ 5 ]. Antibiotics such as chloramphenicol may induce aplastic anemia[ 6 ]. Additionally, some antiepileptic drugs and non-steroidal anti-inflammatory drugs (NSAIDs) are associated with hematological abnormalities[ 7 ]. Compared to general toxicities like hepatotoxicity, nephrotoxicity, and hematotoxicity, research on hematotoxicity is relatively underdeveloped. Early identification of potential blood-toxic drugs during drug development can not only minimize adverse effects on patients but also reduce economic losses due to market withdrawal or usage restrictions, ensuring patient safety[ 8 ]. Current evaluations of hematotoxicity primarily rely on animal models and in vitro cell assays; however, these methods often fail to accurately predict human responses to drugs[ 9 ]. In recent years, computational methods have been increasingly applied to the study of biological problems[ 10 – 12 ], especially toxicological assessments[ 13 – 15 ], which provides an efficient and economical alternative[ 16 ]. Nevertheless, specialized tools for predicting drug-induced hematotoxicity remain scarce. Even comprehensive toxicity prediction platforms like ADMETlab[ 17 ] and SwissADME[ 18 ] do not include hematotoxicity prediction functions. Therefore, this study aims to establish predictive models for drug-induced hematotoxicity to facilitate the identification and screening of blood-toxic compounds. In this research, we collected thousands of compounds with and without hematotoxicity and conducted in-depth analyses using clustering, target prediction, and machine learning techniques. Our study revealed that the mechanisms underlying the action of hematotoxic molecules are intricate and engage multiple pathways. The principal targets identified include HSP90AA1 and SRC. The predominant signaling pathways implicated are the PI3K-Akt, Lipid and atherosclerosis, and MAPK cascades, which are integral to the regulation of fundamental cellular processes such as proliferation, survival, angiogenesis, and signal transduction. By utilizing molecular descriptors and integrating models such as Random Forest (RF), Support Vector Machine (SVM), and XGBoost, we constructed a predictive model for hematotoxicity. The best model achieved an AUC of 0.78, providing a convenient tool for hematotoxicity prediction in drug development. These advancements offer strong support for developing safer therapeutic strategies, reducing drug-related hematotoxicity, and ultimately improving patients' quality of life. Our research workflow is illustrated in Fig. 1. 2. Materials and Methods 2.1 Data Collection and Standardization To construct a comprehensive dataset for predicting drug-induced hematotoxicity, we collected and curated data from multiple reputable sources. Initially, we gathered known hematotoxic compounds from scientific literature, specifically the publication by Hua et al. [ 19 ], and from established databases including SIDER [ 20 ], OCHEM[ 21 ], and PubChem BioAssay[ 22 ]. Our focus was on drugs associated with hematological adverse effects such as anemia, leukopenia, thrombocytopenia, osteomyelitis, and other blood-related disorders. hematotoxicity Data Collection: We performed keyword searches in the SIDER and FAERS databases using terms related to hematotoxicity, including "anemia," "leukopenia," "thrombocytopenia," and "coagulopathy." The SIDER database catalogs drugs with documented side effects, providing valuable information on drug-induced adverse events. Data from the FAERS database were accessed through OpenVigil, a tool that cleans and organizes FAERS data for analysis [ 23 ]. OpenVigil enables Proportional Reporting Ratio (PRR) analysis, allowing us to statistically assess the association between specific drugs and hematological adverse events. Drugs with significant PRR values indicating a strong correlation with hematological adverse drug reactions (ADRs) were selected as hematotoxic compounds. For each identified compound, we retrieved canonical SMILES[ 24 ] and InChI strings[ 25 ] using PubChemPy[ 26 ]. Compounds lacking structural information, such as mixtures, inorganic substances, and macromolecules that could not be converted to SMILES format, were excluded from the dataset. InChI strings served as unique identifiers for deduplication to eliminate redundant entries. Further data cleaning involved removing counterions, solvent fragments, and salts. Hydrogen atoms were added, and molecular structures were standardized using the "wash" function in the Molecular Operating Environment software (MOE, version 2022.02, Chemical Computing Group, Montreal, QC, Canada)[ 27 ]. This process included deprotonating strong acids, protonating strong bases, and adding explicit hydrogens to ensure uniformity and consistency across all compounds. Compounds with molecular weights less than 30 Daltons or greater than 1,000 Daltons were removed to focus on small-molecule drugs relevant to clinical use. Additionally, elemental substances and inorganic materials were excluded to retain only organic compounds pertinent to hematotoxicity. After these rigorous preprocessing steps, we obtained a refined set of unique hematotoxic compounds suitable for model development. Non-Toxic Data Collection: To create a balanced dataset, we collected small-molecule drugs approved for clinical use from the DrugBank database[ 28 ]. The same cleaning and standardization procedures were applied to these compounds. Drugs known to have reported hematological adverse effects in the SIDER or FAERS databases were excluded to ensure that the non-toxic dataset consisted of compounds without documented hematotoxicity. To maintain balance between the toxic and non-toxic classes, we randomly selected an equal number of non-toxic compounds to match the number of hematotoxic compounds. The final dataset comprised 589 hematotoxic compounds (positive set) and 589 non-hematotoxic compounds (negative set), totaling 1,178 unique chemicals. This balanced dataset enhances the reliability of the predictive models by preventing bias toward either class. The dataset was then randomly split into training and test sets in an 80:20 ratio, ensuring that both sets were representative of the overall data distribution and that the models could be effectively trained and validated. By meticulously collecting and cleaning the data, we established a robust foundation for developing accurate predictive models for drug-induced hematotoxicity. This comprehensive dataset enables the application of advanced machine learning techniques to identify potential hematotoxic compounds early in the drug development process. 2.2 Feature Extraction and Data Visualization We utilized the RDKit cheminformatics library[ 29 ] to interpret the SMILES strings of each molecule in our dataset, transforming them into molecular objects. A detailed set of molecular descriptors was subsequently computed using the Descriptors[ 30 ] and MoleculeDescriptors modules of RDKit. Among these descriptors were molecular weight (MolWt), topological polar surface area (TPSA), the count of hydrogen bond donors (NumHDonors), the count of hydrogen bond acceptors (NumHAcceptors), LogP (MolLogP), and molecular refractivity (MolMR). Once these properties were calculated for all molecules, we standardized the descriptors using StandardScaler to maintain consistency across the features. Following this, we visualized the distribution of each property by creating histograms with Seaborn’s histplot function, organizing the data according to toxicity status (toxic versus non-toxic). 2.3 Clustering and Target Analysis of Toxic Molecules Clustering analysis was performed to analyze the similarity of hematotoxic small molecules. Firstly, we extracted the SMILES strings from the original dataset and converted them into molecular objects using the rcdk package[ 31 ]. We then encoded each molecule into binary fingerprint vectors using the extended fingerprint type, with a length of 2,048 bits, where each bit marks specific structural features of the molecule. To visualize the high-dimensional fingerprint data, we applied the t-SNE algorithm[ 32 ] to reduce the 2,048-dimensional vectors to three dimensions. The t-SNE parameters were set as follows: dimensions (dims) = 3, perplexity = 40, and θ (theta) = 0.0. The t-SNE-reduced data were then subjected to hierarchical clustering using the ward.D2 method. We calculated the Euclidean distances between samples and constructed a hierarchical cluster tree, dividing the molecules into four clusters. Each molecule was assigned to its respective cluster and labeled accordingly. For each cluster, we identified representative molecules by calculating the Tanimoto[ 33 ] similarity matrix within the cluster. Specifically, we computed pairwise similarities for all molecules within a cluster and determined the average similarity for each molecule. The molecule with the highest average similarity was selected as the representative of that cluster. To ensure uniqueness, only one representative molecule was identified for each cluster. Finally, we visualized the t-SNE-reduced three-dimensional data using the plotly package[ 34 ], with clusters distinguished by different colors. Representative molecules were labeled as "REP" in the plot for easy identification. 2.4 Construction and Analysis of the Protein-Protein Interaction Network In this investigation, potential target genes associated with hematotoxicity were first identified from four authoritative sources: OMIM[ 35 ], GeneCards[ 36 ], DrugBank[ 37 ], and UniProt[ 38 ]. After performing a comprehensive comparison of these databases and eliminating duplicates, we selected a total of 1,864 potential target genes for subsequent analysis. To predict possible interactions between small molecules causing hematotoxicity and the identified target genes, we employed target prediction methods for the representative molecules through SEA, SuperPred[ 39 ], and SwissTargetPrediction[ 40 ]. These platforms facilitated an in-depth prediction regarding intersecting target genes for the hematotoxic small molecules. To examine the potential interactions among significant target proteins, a comprehensive Protein-Protein Interaction (PPI) network[ 41 ] was constructed using the STRING database[ 42 ]. Visualization of this network was carried out with the latest version of Cytoscape software[ 43 ]. A thorough analysis of the topological characteristics of the network was performed with advanced analytical tools integrated within Cytoscape, allowing us to identify ten hub genes that are vital to the disease mechanism. 2.5 Comprehensive Enrichment Analysis of GO and KEGG Pathways We conducted enrichment analyses for the Gene Ontology (GO)[ 44 ] and Kyoto Encyclopedia of Genes and Genomes (KEGG)[ 45 ] pathways with R packages such as clusterProfiler[ 46 ], org.Hs.eg.db[ 47 ], and pathview[ 48 ]. Gene lists were transformed into Entrez IDs and subsequently analyzed for enrichment through enrichGO and enrichKEGG. The results were illustrated using dot and bar plots, emphasizing significant pathways and terms. Moreover, pathview was employed to map genes onto KEGG pathways, allowing for more detailed visualization. 2.6 Batch Molecular Docking of Active Components The initial receptor protein model derived from HSP90AA1 (PDB ID: 7KRJ) was created, comprising 732 residues. Homology modeling was executed with MODELER 10.1 to add any missing residues and domains. Following this, molecular docking was performed using AutoDock Vina 1.2.0[ 49 ]. The docking box was configured with dimensions of 46 Å x 68 Å x 70 Å and centered at coordinates x = 98.73, y = 125.889, z = 124.831. A second receptor protein model based on SRC (PDB ID: 8JN8) was also developed, consisting of 536 residues. Similar to the first model, homology modeling was completed using MODELER 10.1 to include any missing residues and domains. Molecular docking for this model was likewise performed using AutoDock Vina 1.2.0, with the docking box dimensions set at 54.0 Å x 48.0 Å x 36.0 Å and centered at coordinates x = -25.336, y = 83.729, z = 110.88. The small molecules, initially represented in SMILES format, were converted to pdbqt format with Open Babel[ 50 ]. Batch docking of four representative small molecules against the two target proteins was carried out. This approach allowed for a comprehensive assessment of the small molecules by examining their predicted interaction strengths with the respective targets[ 51 ]. 2.7 Machine Learning In this research, data was imported from CSV files and preprocessed to remove any missing values. The dataset comprised SMILES strings that represent molecular structures alongside their corresponding toxicity classification labels. To describe these molecules, we utilized five different types of molecular fingerprints and an array of molecular descriptors, all calculated using the open-source RDKit software. The molecular fingerprints included Morgan fingerprints (1024 bits), which are akin to Extended Connectivity Fingerprints (ECFP)[ 52 ] and capture circular substructures within the molecules; RDKit fingerprints, which are based on pathway data and represent substructure information; MACCS keys (166 bits)[ 53 ], a predefined collection of structural keys commonly employed in cheminformatics for quick screening; Atom Pairs fingerprints (1024 bits)[ 54 ], which encode pairs of atoms along with their topological distances to detail relationships in the molecular structure; and Topological Pharmacophore Atom Triplets (TPBFP, 1024 bits)[ 55 ], designed to encode triplet interactions based on pharmacophore features present in the molecule. To improve the predictive capabilities of our models, we selected four distinct machine learning algorithms and refined their hyperparameters. The Random Forest (RF) model[ 56 ], an ensemble method for recursive partitioning, was fine-tuned for five hyperparameters: the number of trees (n_estimators: 50, 100, 200), the criterion for splitting (criterion: "gini", "entropy"), the maximum depth of the tree (max_depth: None, 5, 10, 15), the minimum number of samples for leaves (min_samples_leaf: 1–10), and the number of features considered for splits (max_features: "log2", "auto", "sqrt"). This Random Forest model enhances both accuracy and robustness in predictions by generating multiple decision trees, each one constructed from a random selection of training data. The Support Vector Machine (SVM) model[ 57 ] was also applied, with the objective of finding the optimal hyperplane in the feature space by maximizing the margin between different classes, effectively distinguishing between objects labeled differently. We optimized two hyperparameters for the SVM: the coefficient for the kernel function (gamma: 'scale', 'auto') and the penalty parameter for error terms (C: 0.01, 0.1, 1, 10). The SVM model performs especially well in high-dimensional environments and can manage nonlinear classification tasks by transforming the data into higher-dimensional space to locate the best separating hyperplane. In addition, we utilized the Multilayer Perceptron (MLP)[ 58 ], a kind of artificial neural network designed for the recognition of complex patterns. We modified hyperparameters such as hidden layer sizes (hidden_layer_sizes: (50,), (100,), (50, 50)), activation functions (activation: 'tanh', 'relu'), solvers (solver: 'sgd', 'adam'), initial learning rates (learning_rate_init: 0.001, 0.01), and maximum iterations (max_iter: 10000). The MLP learns complex patterns and relationships through multiple hidden layers and nonlinear activation functions. Lastly, we optimized the Extreme Gradient Boosting (XGBoost) model[ 59 ], a sophisticated ensemble learning technique that operates on a gradient boosting framework. Hyperparameters, including the maximum tree depth (max_depth: 3, 5, 7, 10, 15), were tuned for improved performance by successively adding new trees to address the errors made by earlier trees. Throughout the training and evaluation processes, various key performance metrics were employed to evaluate the effectiveness of the models. The Area Under the ROC Curve (AUC) offers a comprehensive assessment of the model's classification abilities across all potential thresholds. Sensitivity (SE) gauges the model's proficiency in accurately identifying true positive cases, while Specificity (SP) evaluates its success in correctly detecting true negative instances. The Matthews Correlation Coefficient (MCC) considers all four sections of the confusion matrix, measuring the quality of binary classifications on a scale from − 1 (complete misclassification) to 1 (perfect classification). Accuracy (ACC) indicates the ratio of correctly classified samples, and Precision (P) refers to the share of true positives among those predicted as positive. The F1 Score (F1), which is the harmonic mean of precision and recall, is particularly beneficial for imbalanced datasets. Balanced Accuracy (BA), calculated as the average of SE and SP, provides a just evaluation for imbalanced classes. These metrics allow for a thorough assessment of each model's performance in predicting hematotoxicity, ensuring precise identification of both positive and negative samples. 3. Results and Discussion 3.1 Dataset Molecular Feature Evaluation The histograms in Fig. 2 illustrate the distribution of six molecular descriptors—molecular weight (MolWt), topological polar surface area (TPSA), number of hydrogen bond donors (NumHDonors), number of hydrogen bond acceptors (NumHAcceptors), LogP, and molecular volume (MolVolume)—for toxic and non-toxic compounds. For molecular weight (Fig. 2A), toxic compounds show a slightly higher tendency toward larger molecular weights compared to non-toxic compounds, though the overall distribution is highly overlapping. In terms of topological polar surface area (Fig. 2B), toxic compounds exhibit a broader range with a shift toward higher values, indicating that polar surface area may play a role in toxicity. The distribution of the number of hydrogen bond donors (Fig. 2C) suggests that toxic compounds tend to have slightly more donors than non-toxic compounds, although the majority of both groups have relatively low values. A similar trend is observed for the number of hydrogen bond acceptors (Fig. 2D), where toxic compounds display a wider range and higher values on average, suggesting a potential correlation between acceptor count and toxicity. For LogP (Fig. 2E), the distribution is nearly identical between toxic and non-toxic compounds, indicating that lipophilicity alone may not be a significant determinant of toxicity. Finally, the molecular volume distribution (Fig. 2F) reveals a slight shift toward larger volumes in toxic compounds compared to non-toxic ones, although the overlap remains substantial. These findings highlight the potential influence of certain molecular descriptors, particularly TPSA, NumHDonors, and NumHAcceptors, in distinguishing toxic compounds from non-toxic ones. 3.2 Targets of hematotoxic small molecules Clustering analysis grouped the molecules into four distinct categories, as shown in Fig. 3 and Table S1. The molecules are grouped into four distinct clusters, each represented in a different color. The representative molecules (labeled "REP") for each cluster are highlighted, with their chemical structures displayed alongside. The representative molecule for the first cluster is L-Penicillamine, which contains 221 molecules. The second cluster is represented by Acetanilide, consisting of 460 molecules. The third cluster is represented by 1-[4-(4-Amino-7-propan-2-ylpyrrolo[2,3-d]pyrimidin-5-yl)phenyl]-3-(5-tert-butyl-1,2-oxazol-3-yl)urea, containing 40 molecules. The fourth cluster is represented by (7bR,8aS)-6-methyl-4-oxo-2-(5,6,7-trimethoxy-1H-indole-2-carbonyl)-1,2,4,5,8,8a-hexahydrocyclopropa[c]pyrrolo[3,2-e]indole-7-carbaldehyde, with 38 molecules. Cluster analysis showed that the molecular data set has intrinsic chemical heterogeneity, which may lead to multiple mechanisms of toxicity. Identifying representative compounds helps to understand structural diversity and select candidate compounds for further analysis. Table 1. Molecular information of the four clusters. Table 1 Molecular information of the four clusters. SMILES cluster SC([C@H](N)C(= O)O)(C)C 1 O = C(Nc1ccccc1)C 2 O = C(Nc1noc(C(C)(C)C)c1)Nc1ccc(-c2c3c(N)ncnc3n(C(C)C)c2)cc1 3 O = C(N1C = 2[C@]3([C@@H](C1)C3)c1c(C = O)c(C)[nH]c1C(= O)C = 2)c1[nH]c2c(OC)c(OC)c(OC)cc2c1 4 Figure 4 presented similarity heatmaps for the molecules within each cluster. The heatmaps visualize pairwise similarity scores among molecules, with darker red regions indicating higher similarity. Clusters with more intense red regions have molecules that share more common structural features, suggesting a tighter chemical space. Clusters 1, 3 and 4 display a relatively higher degree of internal similarity, implying that their molecules have closely related structures. Heatmap of Cluster 2 shows more diversity, with less intense red regions, indicating a broader chemical space within cluster. Using multiple databases, we predicted targets and intersected the results with collected disease targets related to hematotoxicity. As shown in Fig. 5, cluster_1 intersected with 47 genes, cluster_2 with 148 genes, cluster_3 with 123 genes, and cluster_4 with 110 genes. Using the maximum clique centrality (MCC) algorithm in the cytoHubba toolkit, we identified key nodes in the hematotoxicity interactome. The MCC score represents the strength of connectivity and is visually presented in different color intensities, where darker shades indicate higher correlation with hematologic toxicity. We then cataloged the targets for each small molecule, revealing key players such as HSP90AA1 and SRC, as shown in Fig. 6. 3.3 KEGG and GO analysis The Fig. 7, Fig. 8, Fig. 9, and Fig. 10 collectively present the results of GO and KEGG pathway enrichment analyses for genes related to small molecules and hematologic toxicity. The GO enrichment analyses (A of each figure) identify key biological processes, cellular components, and molecular functions. In the first figure, functions and signaling pathways related to cancer, immune regulation, cell migration, and wound healing were significantly enriched. The results of KEGG further emphasized the importance of cancer-related signaling pathways and immune signals, especially infection-related pathways. This provides a possible biological background for further study of these genes. In the second figure, enriched biological processes include response to xenobiotic stimulus, lipopolysaccharide, and bacterial molecules, emphasizing cellular responses to abiotic and biotic stimuli. Processes related to oxidative stress, radiation, and UV response are also significant, alongside regulation of apoptotic signaling pathways and hormone metabolic processes. The key KEGG pathways include fluid shear stress and atherosclerosis, AGE-RAGE signaling pathway in diabetic complications, and PD-L1 expression and PD-1 checkpoint pathway in cancer. Additionally, pathways like serotonergic synapse, chemical carcinogenesis via DNA adducts, and epithelial cell signaling in Helicobacter pylori infection are enriched. Specific disease-related pathways such as Chagas disease, tuberculosis, and toxoplasmosis further emphasize the clinical relevance of these findings. The third figure, enriched biological processes related to blood toxicity include chemotaxis and cell signaling, highlighting pathways such as protein autophosphorylation, MAPK cascade regulation, and cellular responses to peptide hormones, all of which are crucial in hematopoietic and vascular systems. Key KEGG pathways emphasize platelet activation, VEGF signaling, and chemokine signaling, which are directly associated with blood and vascular responses. Cancer-related pathways like acute myeloid leukemia and non-small-cell lung cancer, as well as the PD-L1 checkpoint pathway, are relevant to the hematopoietic system’s regulation under toxic stress. In the last figure, enriched biological processes associated with blood toxicity include leukocyte migration, chemotaxis, and blood coagulation, highlighting critical responses to hypoxia and oxygen level changes. Processes like peptide hormone response and regulation of body fluid levels further connect to vascular and hematopoietic system function. Key KEGG pathways emphasize platelet activation, VEGF signaling, and phosphatidylinositol 3-kinase/protein kinase B (PI3K-Akt) signaling, which are crucial for blood flow regulation and toxicity responses. Cancer-related pathways, such as PD-L1 checkpoint, melanoma, and chronic myeloid leukemia, suggest links between hematological abnormalities and toxic stress. Collectively, these analyses suggest that the interaction between hematotoxic factors and blood toxicity involves intricate networks of inflammation, cellular signaling, and metabolic regulation. Enriched pathways highlight chemotaxis, coagulation, and oxygen response, alongside signaling cascades such as PI3K-Akt, Lipid and atherosclerosis, and MAPK, which are critical for vascular integrity and hematopoietic balance. The findings also reveal links to cancer, immune dysfunction, and systemic stress responses, emphasizing broad implications for cancer progression, immune regulation, and metabolic disorders in the context of hematotoxicity. 3.4 Molecular docking Due to the inclusion of HSP90AA1 and SRC as the key targets for the four categories of molecules and considering the importance of HSP90AA1 and SRC in hematotoxicity, we performed molecular docking to screen the affinity of representative hematotoxic small molecules for HSP90AA1 and SRC. The docking energy results are shown in Table 2. Table 2 Molecules and docking energy results. Molecule Target Affinity (kcal/mol) cluster_1 HSP90AA1 -4.9 cluster_2 HSP90AA1 -5.6 cluster_3 HSP90AA1 -8.9 cluster_4 HSP90AA1 -9.6 cluster_1 SRC -4.6 cluster_2 SRC -5.8 cluster_3 SRC -9.3 cluster_4 SRC -9.6 Table 2. Molecules and docking energy results. The docking results of four representative molecules and two proteins are shown in Fig. 11. The ligand forms more diverse interactions with HSP90AA1, including halogen bonding and pi-sulfur interactions, which suggest a potentially higher binding specificity or affinity for this target. Pi-Pi Stacked and Pi-Alkyl interactions involve aromatic and hydrophobic residues such as PHE138, TYR200, and LEU107.Conventional Hydrogen Bonds are formed with residues like ASP93 and GLN123, stabilizing the ligand. Pi-Sulfur interactions likely involve CYS110, contributing to unique stabilization. Halogen Bonding is observed with residues like SER52 or nearby polar residues capable of accommodating halogen atoms. For SRC, the interactions rely heavily on amide stacking, hydrogen bonding, and hydrophobic contacts, indicating a strong but potentially more hydrophobic-driven binding. Amide-Pi Stacking interactions involve backbone or side-chain amides of residues like GLN276 and ASN350. Conventional Hydrogen Bonds are formed with residues such as HIS327 and LYS295, enhancing specificity. Hydrophobic Interactions involve residues like VAL268 and ILE314, which stabilize the ligand through Pi-Alkyl interactions. Pi-Cation Interactions likely involve positively charged residues like ARG386 or LYS295, further enhancing binding affinity. These residue-specific interactions underline the molecular basis for the ligand’s differential binding affinity and specificity for HSP90AA1 and SRCd, highlighting their potential as dual-target ligands with complementary binding mechanisms. 3.5 Machine Learning The heatmap in Fig. 12 presents the performance metrics for various machine learning models trained on different molecular fingerprints. The metrics evaluated include sensitivity (SE), specificity (SP), accuracy (ACC), F1 score (F1), precision (P), balanced accuracy (BA), and area under the receiver operating characteristic curve (AUC). Among the models, the Random Forest algorithm consistently demonstrates strong performance across multiple fingerprint types, with the MACCS and RDKit fingerprints yielding the highest AUC values (0.76 and 0.77, respectively). Similarly, the AtomPairsFP-RandomForest combination achieves competitive results, with an AUC of 0.77. The Support Vector Machine (SVM) models perform slightly lower overall but still achieve reasonable specificity, particularly with the AtomPairsFP-SVM and Morgan-SVM models, both reaching an AUC of 0.75. The Multi-Layer Perceptron (MLP) models show moderate performance, with the MACCS-MLP model achieving the highest AUC (0.75) among MLP-based models. XGBoost models generally provide consistent results, with AtomPairsFP-XGBoost standing out with the highest overall AUC (0.78) among all model and fingerprint combinations. Sensitivity (SE) is lower for Morgan-SVM (0.50), indicating a potential limitation in identifying true positive cases for this combination. In contrast, specificity (SP) scores are notably high for Random Forest models with Morgan and AtomPairsFP fingerprints, suggesting robust performance in distinguishing true negative cases. Overall, the combination of Random Forest and AtomPairsFP yields the most balanced performance across metrics, making it a suitable choice for predicting toxicity with high reliability. Figure 13 displays the Receiver Operating Characteristic (ROC) curves and Area Under the Curve (AUC) values for four machine learning models—Multi-Layer Perceptron (MLP), Random Forest, Support Vector Machine (SVM), and XGBoost—trained using five different molecular fingerprints. In Panel A, the ROC curves for MLP models show that the MACCS fingerprint achieves the highest AUC (0.75), followed by Morgan and RDKit fingerprints, both with AUCs of 0.74. Topological Torsion demonstrates the lowest performance with an AUC of 0.68. Panel B illustrates the ROC curves for Random Forest models, where RDKit and AtomPairsFP fingerprints yield the highest AUCs (0.77), closely followed by MACCS and Morgan (both at 0.76). Panel C presents the SVM model results, where MACCS achieves the highest AUC (0.77), followed by RDKit (0.76) and AtomPairsFP (0.75), while Topological Torsion again exhibits the lowest AUC (0.68). Panel D highlights the XGBoost models, where AtomPairsFP achieves the best overall performance with an AUC of 0.78, while MACCS and RDKit follow closely at 0.77 and 0.76, respectively. Across all models, the Topological Torsion fingerprint consistently underperforms, while AtomPairsFP and MACCS fingerprints are among the top-performing representations. The XGBoost model paired with AtomPairsFP demonstrates the highest predictive capability overall, emphasizing its suitability for this classification task. Figure 14. Confusion matrices illustrating the classification performance of machine learning models paired with various molecular fingerprints. Panels represent results for (A) AtomPairsFP, (B) MACCS, (C) Morgan, (D) RDKit, and (E) Topological Torsion fingerprints across Multi-Layer Perceptron (MLP), Random Forest, Support Vector Machine (SVM), and XGBoost algorithms. Figure 14 illustrates the confusion matrices for various model and fingerprint combinations, emphasizing their ability to differentiate toxic (hematotoxic) and non-toxic compounds based on true positive rates (recall for the toxic class) and true negative rates (specificity for the non-toxic class). In Panel A, models using the AtomPairsFP fingerprint exhibit noticeable variability across algorithms. The AtomPairsFP-XGBoost model achieves the highest true positive rate (67.11%) while maintaining a reasonable true negative rate (75.17%), highlighting its strength in identifying hematotoxic compounds. In contrast, the AtomPairsFP-SVM model shows a significantly lower true positive rate (57.89%), potentially limiting its effectiveness in detecting toxic samples. The AtomPairsFP-RandomForest model excels in identifying non-toxic compounds, with a high true negative rate of 76.51%, underscoring its reliability in excluding non-hematotoxic molecules. Panel B evaluates models trained on MACCS fingerprints. The MACCS-XGBoost model delivers balanced performance, with a true positive rate of 65.13% and a true negative rate of 75.17%, making it suitable for identifying both hematotoxic and non-toxic compounds. Similarly, the MACCS-RandomForest model shows strong specificity (71.14%) and moderate sensitivity (68.42%), indicating it is slightly biased toward non-toxic samples but remains effective in predicting hematotoxicity. Panel C focuses on Morgan fingerprints, where the Morgan-RandomForest model demonstrates the highest true negative rate (77.85%), emphasizing its robustness in identifying non-toxic compounds. However, the Morgan-SVM model underperforms in terms of recall for hematotoxic compounds, achieving a lower true positive rate (59.87%), which may reduce its overall effectiveness in hematotoxicity classification. Panel D analyzes RDKit fingerprints, with both RDKit-XGBoost and RDKit-RandomForest models showing balanced performance. The RDKit-XGBoost model achieves a true positive rate of 65.13% and a true negative rate of 71.14%, indicating it is well-suited for predicting hematotoxicity. The RDKit-RandomForest model slightly improves on specificity, achieving 73.15%, but has a comparable recall for toxic compounds. Panel E evaluates Topological Torsion fingerprints, which generally underperform compared to other fingerprints across all models. The TopologicalTorsion-XGBoost model achieves a relatively high true negative rate (67.11%) but a modest true positive rate (61.18%), suggesting its limited utility for accurately identifying hematotoxic compounds. Overall, XGBoost models paired with AtomPairsFP consistently demonstrate the best balance between true positive and true negative rates, making them highly effective for predicting both hematotoxic and non-toxic compounds. Random Forest models also show strong specificity, particularly when paired with Morgan fingerprints, making them reliable for identifying non-toxic samples. However, models using Topological Torsion fingerprints consistently underperform across metrics, highlighting their limited applicability in hematotoxicity classification. These findings reinforce the potential of fingerprint-driven machine learning models for early prediction of drug-induced hematotoxicity, providing valuable insights for safer drug development. 4. Conclusion In conclusion, this study presents an integrative framework to explore the mechanisms underlying hematotoxicity in small molecules by employing clustering, reverse target prediction, molecular docking, and machine learning models. This comprehensive approach reveals the complexity of hematotoxic mechanisms and provides an efficient predictive method for assessing molecular hematotoxicity. Clustering analysis grouped the molecules into distinct clusters, each characterized by unique structural features and representative compounds. This chemical heterogeneity suggests multiple potential mechanisms of toxicity, emphasizing the importance of considering structural diversity in toxicity assessment. The clustering results served as the foundation for reverse target prediction, enabling the identification of key proteins such as HSP90AA1 and SRC, which were identified as central nodes within hematotoxicity-related networks. Target prediction and network analysis highlighted the role of these key proteins in mediating toxic effects, providing potential targets for therapeutic intervention or risk assessment. Pathway enrichment analyses further revealed that the identified targets are involved in critical biological processes and pathways, such as inflammation, immune regulation, cell migration, and cancer signaling. These findings suggest that hematotoxic compounds may disrupt these processes, leading to adverse hematological effects. Molecular docking studies provided mechanistic insights by demonstrating significant binding affinities of representative compounds to key proteins such as HSP90AA1 and SRC. The diverse interactions observed, including hydrogen bonds, π–π stacking, and hydrophobic contacts, indicate that these compounds can effectively bind and potentially inhibit these proteins, leading to disruptions in critical signaling pathways. This highlights the complexity of the mechanisms involved in hematotoxicity. Finally, a series of machine learning models were constructed, particularly using the XGBoost algorithm combined with AtomPairsFP fingerprints, to provide an efficient predictive tool for hematotoxicity assessment. These models demonstrated high accuracy in classification, effectively balancing sensitivity and specificity. This enables the early identification of potentially toxic compounds and reduces the need for animal testing. In summary, this study advances our understanding of hematotoxicity by elucidating the complex relationships between molecular structures, biological targets, and toxicological outcomes. The integrative methodology presented here underscores the importance of a multidisciplinary approach, connecting structural characteristics to biological effects through interconnected analyses. This framework provides valuable insights for the development of safer drugs, enabling the early detection of toxic compounds, informing strategies to mitigate adverse effects, and guiding the design of new compounds with improved safety profiles. Declarations Author Contributions: Funding: None. Institutional Review Board Statement: Not applicable. Informed Consent Statement: Not applicable. Data Availability Statement: All data and material used to support the findings of this article are included within this article. Conflicts of Interest: The authors declare no conflicts of interest. References Mathew, J., P. Sankar, and M. Varacallo, Physiology, blood plasma. 2018. Mintzer, D.M., S.N. Billet, and L. Chmielewski, Drug‐induced hematologic syndromes. Advances in hematology, 2009. 2009 (1): p. 495863. Budinsky Jr, R.A., Hematotoxicity: chemically induced toxicity of the blood. Principles of Toxicology: Environmental and Industrial Applications, 2000: p. 87-109. Krishnan, K. and M. Pelekis, Hematotoxic interactions: occurrence, mechanisms and predictability. Toxicology, 1995. 105 (2-3): p. 355-364. Maxwell, M.B. and K.E. Maher. Chemotherapy-induced myelosuppression . in Seminars in oncology nursing . 1992. Bilinski, J., et al., Fecal microbiota transplantation in patients with blood disorders inhibits gut colonization with antibiotic-resistant bacteria: results of a prospective, single-center study. Clinical Infectious Diseases, 2017. 65 (3): p. 364-370. Robak, P., P. Smolewski, and T. Robak, The role of non-steroidal anti-inflammatory drugs in the risk of development and treatment of hematologic malignancies. Leukemia & lymphoma, 2008. 49 (8): p. 1452-1462. Greene, N. and R. Naven, Early toxicity screening strategies. Curr Opin Drug Discov Devel, 2009. 12 (1): p. 90-97. Van Norman, G.A., Limitations of animal studies for predicting toxicity in clinical trials: is it time to rethink our current approach? JACC: Basic to Translational Science, 2019. 4 (7): p. 845-854. Zhao, W., et al., Unveiling Anti-Diabetic Potential of Baicalin and Baicalein from Baikal Skullcap: LC-MS, In Silico, and In Vitro Studies. Int J Mol Sci, 2024. 25 (7). Wang, K., et al., Exploring the anti-gout potential of sunflower receptacles alkaloids: A computational and pharmacological analysis. Comput Biol Med, 2024. 172 : p. 108252. Su, J., et al., Integrating Computational and Experimental Methods to Identify Novel Sweet Peptides from Egg and Soy Proteins. Int J Mol Sci, 2024. 25 (10). Ekins, S., Progress in computational toxicology. Journal of pharmacological and toxicological methods, 2014. 69 (2): p. 115-140. Raies, A.B. and V.B. Bajic, In silico toxicology: computational methods for the prediction of chemical toxicity. Wiley Interdisciplinary Reviews: Computational Molecular Science, 2016. 6 (2): p. 147-172. Cui, H., et al., Computational Insights into Reproductive Toxicity: Clustering, Mechanism Analysis, and Predictive Models. Int J Mol Sci, 2024. 25 (14). Liu, K., et al., GPT4Kinase: High-accuracy prediction of inhibitor-kinase binding affinity utilizing large language model. Int J Biol Macromol, 2024. 282 (Pt 5): p. 137069. Dong, J., et al., ADMETlab: a platform for systematic ADMET evaluation based on a comprehensively collected ADMET database. Journal of cheminformatics, 2018. 10 : p. 1-11. Daina, A., O. Michielin, and V. Zoete, SwissADME: a free web tool to evaluate pharmacokinetics, drug-likeness and medicinal chemistry friendliness of small molecules. Scientific reports, 2017. 7 (1): p. 42717. Hua, Y., et al., In silico prediction of chemical-induced hematotoxicity with machine learning and deep learning methods. Molecular Diversity, 2021. 25 (3): p. 1585-1596. Kuhn, M., et al., The SIDER database of drugs and side effects. Nucleic acids research, 2016. 44 (D1): p. D1075-D1079. Sushko, I., et al., Online chemical modeling environment (OCHEM): web platform for data storage, model development and publishing of chemical information. Journal of computer-aided molecular design, 2011. 25 : p. 533-554. Wang, Y., et al., PubChem's BioAssay database. Nucleic acids research, 2012. 40 (D1): p. D400-D412. Böhm, R., et al., OpenVigil FDA–inspection of US American adverse drug events pharmacovigilance data and novel clinical applications. PloS one, 2016. 11 (6): p. e0157753. Weininger, D., SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. Journal of chemical information and computer sciences, 1988. 28 (1): p. 31-36. Heller, S.R., et al., InChI, the IUPAC international chemical identifier. Journal of cheminformatics, 2015. 7 : p. 1-34. Swain, M., PubChemPy: A way to interact with PubChem in Python . 2014. Vilar, S., G. Cozza, and S. Moro, Medicinal chemistry and the molecular operating environment (MOE): application of QSAR and molecular docking to drug discovery. Current topics in medicinal chemistry, 2008. 8 (18): p. 1555-1572. Wishart, D.S., et al., DrugBank 5.0: a major update to the DrugBank database for 2018. Nucleic acids research, 2018. 46 (D1): p. D1074-D1082. Bento, A.P., et al., An open source chemical structure curation pipeline using RDKit. Journal of Cheminformatics, 2020. 12 : p. 1-16. Landrum, G., Rdkit documentation. Release, 2013. 1 (1-79): p. 4. Voicu, A., et al., The rcdk and cluster R packages applied to drug candidate selection. Journal of Cheminformatics, 2020. 12 : p. 1-8. Arora, S., W. Hu, and P.K. Kothari. An analysis of the t-sne algorithm for data visualization . in Conference on learning theory . 2018. PMLR. Bajusz, D., A. Rácz, and K. Héberger, Why is Tanimoto index an appropriate choice for fingerprint-based similarity calculations? Journal of cheminformatics, 2015. 7 : p. 1-13. Sievert, C., Interactive web-based data visualization with R, plotly, and shiny . 2020: Chapman and Hall/CRC. Amberger, J.S., et al., OMIM. org: Online Mendelian Inheritance in Man (OMIM®), an online catalog of human genes and genetic disorders. Nucleic acids research, 2015. 43 (D1): p. D789-D798. Safran, M., et al., GeneCards Version 3: the human gene integrator. Database, 2010. 2010 : p. baq020. Knox, C., et al., DrugBank 6.0: the DrugBank knowledgebase for 2024. Nucleic acids research, 2024. 52 (D1): p. D1265-D1275. Consortium, U., UniProt: a worldwide hub of protein knowledge. Nucleic acids research, 2019. 47 (D1): p. D506-D515. Nickel, J., et al., SuperPred: update on drug classification and target prediction. Nucleic acids research, 2014. 42 (W1): p. W26-W31. Daina, A., O. Michielin, and V. Zoete, SwissTargetPrediction: updated data and new features for efficient prediction of protein targets of small molecules. Nucleic acids research, 2019. 47 (W1): p. W357-W364. Rao, V.S., et al., Protein‐protein interaction detection: methods and analysis. International journal of proteomics, 2014. 2014 (1): p. 147648. Szklarczyk, D., et al., The STRING database in 2021: customizable protein–protein networks, and functional characterization of user-uploaded gene/measurement sets. Nucleic acids research, 2021. 49 (D1): p. D605-D612. Shannon, P., et al., Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome research, 2003. 13 (11): p. 2498-2504. Ashburner, M., et al., Gene ontology: tool for the unification of biology. Nature genetics, 2000. 25 (1): p. 25-29. Kanehisa, M. and S. Goto, KEGG: kyoto encyclopedia of genes and genomes. Nucleic acids research, 2000. 28 (1): p. 27-30. Yu, G., et al., clusterProfiler: an R package for comparing biological themes among gene clusters. Omics: a journal of integrative biology, 2012. 16 (5): p. 284-287. Carlson, M., et al., org. Hs. eg. db: Genome wide annotation for Human. R package version, 2019. 3 (2): p. 3. Luo, W. and C. Brouwer, Pathview: an R/Bioconductor package for pathway-based data integration and visualization. Bioinformatics, 2013. 29 (14): p. 1830-1831. Eberhardt, J., et al., AutoDock Vina 1.2. 0: New docking methods, expanded force field, and python bindings. Journal of chemical information and modeling, 2021. 61 (8): p. 3891-3898. O'Boyle, N.M., et al., Open Babel: An open chemical toolbox. Journal of cheminformatics, 2011. 3 : p. 1-14. Copeland, R.A., Evaluation of enzyme inhibitors in drug discovery: a guide for medicinal chemists and pharmacologists . 2013: John Wiley & Sons. Rogers, D. and M. Hahn, Extended-connectivity fingerprints. Journal of chemical information and modeling, 2010. 50 (5): p. 742-754. Jow, H., et al., MELCOR accident consequence code system (MACCS) . 1990, Nuclear Regulatory Commission, Washington, DC (USA). Div. of Systems …. Pérez-Nueno, V.I., et al., APIF: a new interaction fingerprint based on atom pairs and its application to virtual screening. Journal of chemical information and modeling, 2009. 49 (5): p. 1245-1260. Bonachéra, F., et al., Fuzzy tricentric pharmacophore fingerprints. 1. Topological fuzzy pharmacophore triplets and adapted molecular similarity scoring schemes. Journal of chemical information and modeling, 2006. 46 (6): p. 2457-2477. Belgiu, M. and L. Drăguţ, Random forest in remote sensing: A review of applications and future directions. ISPRS journal of photogrammetry and remote sensing, 2016. 114 : p. 24-31. Wang, L., Support vector machines: theory and applications . Vol. 177. 2005: Springer Science & Business Media. Taud, H. and J.-F. Mas, Multilayer perceptron (MLP). Geomatic approaches for modeling land change scenarios, 2018: p. 451-455. Chen, T., Xgboost: extreme gradient boosting. R package version 0.4-2, 2015. 1 (4). Table S1 Table S1 is not available with this version. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6295317","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":447551196,"identity":"18c908f2-7196-4f28-801c-0566b721717e","order_by":0,"name":"Jian Xu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABH0lEQVRIie2QMUvDQBTHLxxkunDrhZP4FZ4EIi5+lh6BTIqdJGDBlMplUHFt8Us4nR0tB+1y0jWDQ4qfoJvFDCahk8TQ0eF+cBw83o//ew8hi+U/wurnZIhQ7IzLQTpqi/oQJfBzrKE0y6bmfCIE9Y97lRDWbuJv7touHPYp9Pkh2uzm52I2IVEqMhfR/D7hw6q6ovn6rTPkw5yGnonFEyZRIeZHiJl3xWcSzqYm7kwBdhFxR+I2pRDGrSuXinsZACq6B2sUfydvxYsm0VBI3Civ36QCOO5RmCd1CNpN0F5RnLgA8IfCiuSae3IV+BOs2cAsyX6XEE5MDJ0Xm8aqHuyGULoYb7/SUUDzR1VfLIBgtSg7Y35DDuqyWCwWSz8/1/JfNhavU6AAAAAASUVORK5CYII=","orcid":"","institution":"Shaoxing People's Hospital","correspondingAuthor":true,"prefix":"","firstName":"Jian","middleName":"","lastName":"Xu","suffix":""},{"id":447551197,"identity":"ed50d7fb-5e08-42c6-941b-48468dac1985","order_by":1,"name":"Mengfei Zhang","email":"","orcid":"","institution":"Shaoxing People's Hospital","correspondingAuthor":false,"prefix":"","firstName":"Mengfei","middleName":"","lastName":"Zhang","suffix":""},{"id":447551199,"identity":"6fb4fded-398d-4d41-ab82-751a8ce6c024","order_by":2,"name":"Leyu Tao","email":"","orcid":"","institution":"Shaoxing People's Hospital","correspondingAuthor":false,"prefix":"","firstName":"Leyu","middleName":"","lastName":"Tao","suffix":""},{"id":447551201,"identity":"12c8c4a2-22ac-4a4a-96ea-eb84847b9b6d","order_by":3,"name":"Shuying Zhang","email":"","orcid":"","institution":"Shaoxing People's Hospital","correspondingAuthor":false,"prefix":"","firstName":"Shuying","middleName":"","lastName":"Zhang","suffix":""}],"badges":[],"createdAt":"2025-03-24 12:23:29","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6295317/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6295317/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":82150868,"identity":"cd38c32d-3a56-4569-9047-6e610d44be2f","added_by":"auto","created_at":"2025-05-07 07:18:38","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":868648,"visible":true,"origin":"","legend":"\u003cp\u003eResearch flow chart.\u003c/p\u003e","description":"","filename":"F1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/48f021b2d89961671723c8f1.jpg"},{"id":82149643,"identity":"2ee8bdc2-9c22-476a-83df-89b9361fac0e","added_by":"auto","created_at":"2025-05-07 07:10:38","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1273815,"visible":true,"origin":"","legend":"\u003cp\u003eDistribution of molecular descriptors for toxic (green) and non-toxic (red) compounds. Panels represent the distributions of (A) molecular weight (MolWt), (B) topological polar surface area (TPSA), (C) number of hydrogen bond donors (NumHDonors), (D) number of hydrogen bond acceptors (NumHAcceptors), (E) LogP, and (F) molecular volume (MolVolume). Density plots are overlaid to illustrate the differences in descriptor distributions between toxic and non-toxic compounds\u003c/p\u003e","description":"","filename":"F2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/c651fea7f02aa212f627e3f9.jpg"},{"id":82147726,"identity":"d4a15563-eddf-4219-992b-454756a60234","added_by":"auto","created_at":"2025-05-07 07:02:38","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":305785,"visible":true,"origin":"","legend":"\u003cp\u003eMolecular Clustering Diagram. The diagram shows four clusters of small molecules.\u003c/p\u003e","description":"","filename":"F3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/642ce9c5b40709bccbb1a0d0.jpg"},{"id":82149646,"identity":"f5794325-0f80-4a72-aa90-15a5abc7da6f","added_by":"auto","created_at":"2025-05-07 07:10:38","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1339654,"visible":true,"origin":"","legend":"\u003cp\u003eSimilarity heatmaps for each cluster, visualizing pairwise similarity scores among molecules. Darker red regions indicate higher similarity.\u003c/p\u003e","description":"","filename":"F4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/1934c6fefb2d5fd0e01aae75.jpg"},{"id":82149644,"identity":"718dc9ab-1ee4-4f65-9be5-cde712bf4f3d","added_by":"auto","created_at":"2025-05-07 07:10:38","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":128776,"visible":true,"origin":"","legend":"\u003cp\u003eThe Venn diagram illustrates the shared intersection genes between small molecules and Hematotoxicity disease. A: The number of shared intersection genes between cluster_1 and disease is 47. B: The number of shared intersection genes between cluster_2 and disease is 148. C: The number of shared intersection genes between cluster_3 and disease is 123. D: The number of shared intersection genes between cluster_4 and disease is 110.\u003c/p\u003e","description":"","filename":"F5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/93dc7e70faee37cca8457bee.jpg"},{"id":82146130,"identity":"c2ab43c0-f0c6-4cea-a2d6-4d26403d05c2","added_by":"auto","created_at":"2025-05-07 06:54:38","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":2946118,"visible":true,"origin":"","legend":"\u003cp\u003eThis figure illustrates the protein interactions involved in hematotoxicity-related diseases. A: Gene interactions of cluster_1. B: Gene interactions of cluster_2. C: Gene interactions of cluster_3. D: Gene interactions of cluster_4.\u003c/p\u003e","description":"","filename":"F6.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/a9a7cb9a8e4a72ff4c66b4c8.jpg"},{"id":82146116,"identity":"81f01379-54cc-4c8f-bb08-ff8f1b698147","added_by":"auto","created_at":"2025-05-07 06:54:38","extension":"jpg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":271584,"visible":true,"origin":"","legend":"\u003cp\u003eThe results presented are based on the GO and KEGG pathway enrichment analyses of intersection genes between cluster_1 and hematotoxicity.\u003c/p\u003e","description":"","filename":"F7.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/e5318b20b6df11d491e1910b.jpg"},{"id":82146125,"identity":"71d3c2b1-d99c-40da-9e61-b6ecfbf459e3","added_by":"auto","created_at":"2025-05-07 06:54:38","extension":"jpg","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":272983,"visible":true,"origin":"","legend":"\u003cp\u003eThe results presented are based on the GO and KEGG pathway enrichment analyses of intersection genes between cluster_2 and hematotoxicity.\u003c/p\u003e","description":"","filename":"F8.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/564daeb3aeb7275d8f7284c2.jpg"},{"id":82153680,"identity":"deb2b1fa-9e5f-42d2-902e-e5be955f5081","added_by":"auto","created_at":"2025-05-07 07:26:38","extension":"jpg","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":280470,"visible":true,"origin":"","legend":"\u003cp\u003eThe results presented are based on the GO and KEGG pathway enrichment analyses of intersection genes between cluster_3 and hematotoxicity.\u003c/p\u003e","description":"","filename":"F9.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/4b5adc0fc6805cdf1d979b7f.jpg"},{"id":82147731,"identity":"1adbdab7-1fc9-4578-aeaa-3b7ab0a30f4b","added_by":"auto","created_at":"2025-05-07 07:02:38","extension":"jpg","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":267281,"visible":true,"origin":"","legend":"\u003cp\u003eThe results presented are based on the GO and KEGG pathway enrichment analyses of intersection genes between cluster_4 and hematotoxicity.\u003c/p\u003e","description":"","filename":"F10.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/f89d1d1aa25b7d0c73659586.jpg"},{"id":82147736,"identity":"d3d21dc4-3422-4e90-89ed-8fb2f016357e","added_by":"auto","created_at":"2025-05-07 07:02:38","extension":"jpg","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":1065633,"visible":true,"origin":"","legend":"\u003cp\u003eDocking Results of HSP90AA1 and SRC with Four Representative Molecules. A: Interaction between Molecules and HSP90AA1. B: Interaction between Molecules and SRC.\u003c/p\u003e","description":"","filename":"F11.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/669585a572ab0228e8eb053a.jpg"},{"id":82149648,"identity":"f3fe2bd3-3b5c-436f-b6f4-fe1bbd36af2a","added_by":"auto","created_at":"2025-05-07 07:10:38","extension":"jpg","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":1206650,"visible":true,"origin":"","legend":"\u003cp\u003eHeatmap illustrating the performance metrics of various machine learning models combined with different molecular fingerprints. The metrics evaluated include sensitivity (SE), specificity (SP), accuracy (ACC), F1 score (F1), precision (P), balanced accuracy (BA), and area under the receiver operating characteristic curve (AUC). The AtomPairsFP-XGBoost model achieves the highest AUC (0.78), while Random Forest models consistently perform well across multiple fingerprints. Color intensity indicates the magnitude of the metric, with darker red signifying higher values.\u003c/p\u003e","description":"","filename":"F12.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/ebceaecae65d63cea03b4e86.jpg"},{"id":82147743,"identity":"3476e0a1-d16d-4700-9fbf-0dc0cca32712","added_by":"auto","created_at":"2025-05-07 07:02:38","extension":"jpg","order_by":13,"title":"Figure 13","display":"","copyAsset":false,"role":"figure","size":843694,"visible":true,"origin":"","legend":"\u003cp\u003eROC curves and AUC values for four machine learning models—(A) Multi-Layer Perceptron (MLP), (B) Random Forest, (C) Support Vector Machine (SVM), and (D) XGBoost—trained on five molecular fingerprints (MACCS, Morgan, RDKit, Topological Torsion, and AtomPairsFP). Each curve represents the trade-off between true positive rate (sensitivity) and false positive rate (1-specificity) for a given model-fingerprint combination.\u003c/p\u003e","description":"","filename":"F13.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/cbf0737ba642b0e04d227622.jpg"},{"id":82149653,"identity":"d4f2bf47-d58d-4680-90f3-9d2b59f16b72","added_by":"auto","created_at":"2025-05-07 07:10:38","extension":"jpg","order_by":14,"title":"Figure 14","display":"","copyAsset":false,"role":"figure","size":2766741,"visible":true,"origin":"","legend":"\u003cp\u003eConfusion matrices illustrating the classification performance of machine learning models paired with various molecular fingerprints. Panels represent results for (A) AtomPairsFP, (B) MACCS, (C) Morgan, (D) RDKit, and (E) Topological Torsion fingerprints across Multi-Layer Perceptron (MLP), Random Forest, Support Vector Machine (SVM), and XGBoost algorithms.\u003c/p\u003e","description":"","filename":"F14.jpg","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/6ef51826f86e4cffb0c76a4c.jpg"},{"id":86083786,"identity":"65597de4-6beb-4bd1-9eee-c71ee3c80441","added_by":"auto","created_at":"2025-07-06 02:01:24","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":14712754,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6295317/v1/43b0142d-1458-4b49-a23e-221c5ccfe46c.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Computational Prediction of Drug-Induced Hematotoxicity: Mechanisms and Model Development","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eBlood plays a crucial role in the human body by transporting oxygen, nutrients, and waste products, maintaining normal physiological activities[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Hematotoxicity refers to the potential adverse effects of certain drugs on the hematological system, including abnormalities in red blood cells, white blood cells, and platelets, as well as disruptions in coagulation functions[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. The main clinical manifestations of this toxicity include anemia, leukopenia, thrombocytopenia, and a tendency to bleed. Since the impact of drugs on the blood system is often overlooked, hematotoxicity can lead to serious consequences such as increased risk of infection, hemorrhage, and even death[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. The mechanisms may involve direct damage to hematopoietic stem cells, immune-mediated cell destruction, or interference with signaling pathways that regulate cell proliferation and differentiation[\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSeveral drugs have been reported to exhibit hematotoxicity. For example, chemotherapy agents like cyclophosphamide and doxorubicin can cause bone marrow suppression, leading to pancytopenia[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Antibiotics such as chloramphenicol may induce aplastic anemia[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Additionally, some antiepileptic drugs and non-steroidal anti-inflammatory drugs (NSAIDs) are associated with hematological abnormalities[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Compared to general toxicities like hepatotoxicity, nephrotoxicity, and hematotoxicity, research on hematotoxicity is relatively underdeveloped. Early identification of potential blood-toxic drugs during drug development can not only minimize adverse effects on patients but also reduce economic losses due to market withdrawal or usage restrictions, ensuring patient safety[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Current evaluations of hematotoxicity primarily rely on animal models and in vitro cell assays; however, these methods often fail to accurately predict human responses to drugs[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. In recent years, computational methods have been increasingly applied to the study of biological problems[\u003cspan additionalcitationids=\"CR11\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e], especially toxicological assessments[\u003cspan additionalcitationids=\"CR14\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e], which provides an efficient and economical alternative[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. Nevertheless, specialized tools for predicting drug-induced hematotoxicity remain scarce. Even comprehensive toxicity prediction platforms like ADMETlab[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] and SwissADME[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] do not include hematotoxicity prediction functions. Therefore, this study aims to establish predictive models for drug-induced hematotoxicity to facilitate the identification and screening of blood-toxic compounds.\u003c/p\u003e \u003cp\u003eIn this research, we collected thousands of compounds with and without hematotoxicity and conducted in-depth analyses using clustering, target prediction, and machine learning techniques. Our study revealed that the mechanisms underlying the action of hematotoxic molecules are intricate and engage multiple pathways. The principal targets identified include HSP90AA1 and SRC. The predominant signaling pathways implicated are the PI3K-Akt, Lipid and atherosclerosis, and MAPK cascades, which are integral to the regulation of fundamental cellular processes such as proliferation, survival, angiogenesis, and signal transduction. By utilizing molecular descriptors and integrating models such as Random Forest (RF), Support Vector Machine (SVM), and XGBoost, we constructed a predictive model for hematotoxicity. The best model achieved an AUC of 0.78, providing a convenient tool for hematotoxicity prediction in drug development. These advancements offer strong support for developing safer therapeutic strategies, reducing drug-related hematotoxicity, and ultimately improving patients' quality of life. Our research workflow is illustrated in Fig.\u0026nbsp;1.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e"},{"header":"2. Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Data Collection and Standardization\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eTo construct a comprehensive dataset for predicting drug-induced hematotoxicity, we collected and curated data from multiple reputable sources. Initially, we gathered known hematotoxic compounds from scientific literature, specifically the publication by Hua et al. [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], and from established databases including SIDER [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], OCHEM[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e], and PubChem BioAssay[\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Our focus was on drugs associated with hematological adverse effects such as anemia, leukopenia, thrombocytopenia, osteomyelitis, and other blood-related disorders. hematotoxicity Data Collection: We performed keyword searches in the SIDER and FAERS databases using terms related to hematotoxicity, including \"anemia,\" \"leukopenia,\" \"thrombocytopenia,\" and \"coagulopathy.\" The SIDER database catalogs drugs with documented side effects, providing valuable information on drug-induced adverse events. Data from the FAERS database were accessed through OpenVigil, a tool that cleans and organizes FAERS data for analysis [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. OpenVigil enables Proportional Reporting Ratio (PRR) analysis, allowing us to statistically assess the association between specific drugs and hematological adverse events. Drugs with significant PRR values indicating a strong correlation with hematological adverse drug reactions (ADRs) were selected as hematotoxic compounds.\u003c/p\u003e \u003cp\u003eFor each identified compound, we retrieved canonical SMILES[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e] and InChI strings[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] using PubChemPy[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. Compounds lacking structural information, such as mixtures, inorganic substances, and macromolecules that could not be converted to SMILES format, were excluded from the dataset. InChI strings served as unique identifiers for deduplication to eliminate redundant entries. Further data cleaning involved removing counterions, solvent fragments, and salts. Hydrogen atoms were added, and molecular structures were standardized using the \"wash\" function in the Molecular Operating Environment software (MOE, version 2022.02, Chemical Computing Group, Montreal, QC, Canada)[\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. This process included deprotonating strong acids, protonating strong bases, and adding explicit hydrogens to ensure uniformity and consistency across all compounds. Compounds with molecular weights less than 30 Daltons or greater than 1,000 Daltons were removed to focus on small-molecule drugs relevant to clinical use. Additionally, elemental substances and inorganic materials were excluded to retain only organic compounds pertinent to hematotoxicity. After these rigorous preprocessing steps, we obtained a refined set of unique hematotoxic compounds suitable for model development. Non-Toxic Data Collection: To create a balanced dataset, we collected small-molecule drugs approved for clinical use from the DrugBank database[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. The same cleaning and standardization procedures were applied to these compounds. Drugs known to have reported hematological adverse effects in the SIDER or FAERS databases were excluded to ensure that the non-toxic dataset consisted of compounds without documented hematotoxicity. To maintain balance between the toxic and non-toxic classes, we randomly selected an equal number of non-toxic compounds to match the number of hematotoxic compounds.\u003c/p\u003e \u003cp\u003eThe final dataset comprised 589 hematotoxic compounds (positive set) and 589 non-hematotoxic compounds (negative set), totaling 1,178 unique chemicals. This balanced dataset enhances the reliability of the predictive models by preventing bias toward either class. The dataset was then randomly split into training and test sets in an 80:20 ratio, ensuring that both sets were representative of the overall data distribution and that the models could be effectively trained and validated. By meticulously collecting and cleaning the data, we established a robust foundation for developing accurate predictive models for drug-induced hematotoxicity. This comprehensive dataset enables the application of advanced machine learning techniques to identify potential hematotoxic compounds early in the drug development process.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Feature Extraction and Data Visualization\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eWe utilized the RDKit cheminformatics library[\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] to interpret the SMILES strings of each molecule in our dataset, transforming them into molecular objects. A detailed set of molecular descriptors was subsequently computed using the Descriptors[\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e] and MoleculeDescriptors modules of RDKit. Among these descriptors were molecular weight (MolWt), topological polar surface area (TPSA), the count of hydrogen bond donors (NumHDonors), the count of hydrogen bond acceptors (NumHAcceptors), LogP (MolLogP), and molecular refractivity (MolMR). Once these properties were calculated for all molecules, we standardized the descriptors using StandardScaler to maintain consistency across the features. Following this, we visualized the distribution of each property by creating histograms with Seaborn\u0026rsquo;s histplot function, organizing the data according to toxicity status (toxic versus non-toxic).\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Clustering and Target Analysis of Toxic Molecules\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eClustering analysis was performed to analyze the similarity of hematotoxic small molecules. Firstly, we extracted the SMILES strings from the original dataset and converted them into molecular objects using the rcdk package[\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. We then encoded each molecule into binary fingerprint vectors using the extended fingerprint type, with a length of 2,048 bits, where each bit marks specific structural features of the molecule. To visualize the high-dimensional fingerprint data, we applied the t-SNE algorithm[\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e] to reduce the 2,048-dimensional vectors to three dimensions. The t-SNE parameters were set as follows: dimensions (dims)\u0026thinsp;=\u0026thinsp;3, perplexity\u0026thinsp;=\u0026thinsp;40, and θ (theta)\u0026thinsp;=\u0026thinsp;0.0. The t-SNE-reduced data were then subjected to hierarchical clustering using the ward.D2 method. We calculated the Euclidean distances between samples and constructed a hierarchical cluster tree, dividing the molecules into four clusters. Each molecule was assigned to its respective cluster and labeled accordingly.\u003c/p\u003e \u003cp\u003eFor each cluster, we identified representative molecules by calculating the Tanimoto[\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e] similarity matrix within the cluster. Specifically, we computed pairwise similarities for all molecules within a cluster and determined the average similarity for each molecule. The molecule with the highest average similarity was selected as the representative of that cluster. To ensure uniqueness, only one representative molecule was identified for each cluster. Finally, we visualized the t-SNE-reduced three-dimensional data using the plotly package[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e], with clusters distinguished by different colors. Representative molecules were labeled as \"REP\" in the plot for easy identification.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Construction and Analysis of the Protein-Protein Interaction Network\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eIn this investigation, potential target genes associated with hematotoxicity were first identified from four authoritative sources: OMIM[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e], GeneCards[\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e], DrugBank[\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e], and UniProt[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. After performing a comprehensive comparison of these databases and eliminating duplicates, we selected a total of 1,864 potential target genes for subsequent analysis. To predict possible interactions between small molecules causing hematotoxicity and the identified target genes, we employed target prediction methods for the representative molecules through SEA, SuperPred[\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e], and SwissTargetPrediction[\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. These platforms facilitated an in-depth prediction regarding intersecting target genes for the hematotoxic small molecules.\u003c/p\u003e \u003cp\u003eTo examine the potential interactions among significant target proteins, a comprehensive Protein-Protein Interaction (PPI) network[\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e] was constructed using the STRING database[\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. Visualization of this network was carried out with the latest version of Cytoscape software[\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. A thorough analysis of the topological characteristics of the network was performed with advanced analytical tools integrated within Cytoscape, allowing us to identify ten hub genes that are vital to the disease mechanism.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5 Comprehensive Enrichment Analysis of GO and KEGG Pathways\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eWe conducted enrichment analyses for the Gene Ontology (GO)[\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e] and Kyoto Encyclopedia of Genes and Genomes (KEGG)[\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e] pathways with R packages such as clusterProfiler[\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e], org.Hs.eg.db[\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e], and pathview[\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e]. Gene lists were transformed into Entrez IDs and subsequently analyzed for enrichment through enrichGO and enrichKEGG. The results were illustrated using dot and bar plots, emphasizing significant pathways and terms. Moreover, pathview was employed to map genes onto KEGG pathways, allowing for more detailed visualization.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.6 Batch Molecular Docking of Active Components\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eThe initial receptor protein model derived from HSP90AA1 (PDB ID: 7KRJ) was created, comprising 732 residues. Homology modeling was executed with MODELER 10.1 to add any missing residues and domains. Following this, molecular docking was performed using AutoDock Vina 1.2.0[\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e]. The docking box was configured with dimensions of 46 \u0026Aring; x 68 \u0026Aring; x 70 \u0026Aring; and centered at coordinates x\u0026thinsp;=\u0026thinsp;98.73, y\u0026thinsp;=\u0026thinsp;125.889, z\u0026thinsp;=\u0026thinsp;124.831.\u003c/p\u003e \u003cp\u003eA second receptor protein model based on SRC (PDB ID: 8JN8) was also developed, consisting of 536 residues. Similar to the first model, homology modeling was completed using MODELER 10.1 to include any missing residues and domains. Molecular docking for this model was likewise performed using AutoDock Vina 1.2.0, with the docking box dimensions set at 54.0 \u0026Aring; x 48.0 \u0026Aring; x 36.0 \u0026Aring; and centered at coordinates x = -25.336, y\u0026thinsp;=\u0026thinsp;83.729, z\u0026thinsp;=\u0026thinsp;110.88. The small molecules, initially represented in SMILES format, were converted to pdbqt format with Open Babel[\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eBatch docking of four representative small molecules against the two target proteins was carried out. This approach allowed for a comprehensive assessment of the small molecules by examining their predicted interaction strengths with the respective targets[\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.7 Machine Learning\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eIn this research, data was imported from CSV files and preprocessed to remove any missing values. The dataset comprised SMILES strings that represent molecular structures alongside their corresponding toxicity classification labels. To describe these molecules, we utilized five different types of molecular fingerprints and an array of molecular descriptors, all calculated using the open-source RDKit software. The molecular fingerprints included Morgan fingerprints (1024 bits), which are akin to Extended Connectivity Fingerprints (ECFP)[\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e] and capture circular substructures within the molecules; RDKit fingerprints, which are based on pathway data and represent substructure information; MACCS keys (166 bits)[\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e], a predefined collection of structural keys commonly employed in cheminformatics for quick screening; Atom Pairs fingerprints (1024 bits)[\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e], which encode pairs of atoms along with their topological distances to detail relationships in the molecular structure; and Topological Pharmacophore Atom Triplets (TPBFP, 1024 bits)[\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e], designed to encode triplet interactions based on pharmacophore features present in the molecule.\u003c/p\u003e \u003cp\u003eTo improve the predictive capabilities of our models, we selected four distinct machine learning algorithms and refined their hyperparameters. The Random Forest (RF) model[\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e], an ensemble method for recursive partitioning, was fine-tuned for five hyperparameters: the number of trees (n_estimators: 50, 100, 200), the criterion for splitting (criterion: \"gini\", \"entropy\"), the maximum depth of the tree (max_depth: None, 5, 10, 15), the minimum number of samples for leaves (min_samples_leaf: 1\u0026ndash;10), and the number of features considered for splits (max_features: \"log2\", \"auto\", \"sqrt\"). This Random Forest model enhances both accuracy and robustness in predictions by generating multiple decision trees, each one constructed from a random selection of training data.\u003c/p\u003e \u003cp\u003eThe Support Vector Machine (SVM) model[\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e] was also applied, with the objective of finding the optimal hyperplane in the feature space by maximizing the margin between different classes, effectively distinguishing between objects labeled differently. We optimized two hyperparameters for the SVM: the coefficient for the kernel function (gamma: 'scale', 'auto') and the penalty parameter for error terms (C: 0.01, 0.1, 1, 10). The SVM model performs especially well in high-dimensional environments and can manage nonlinear classification tasks by transforming the data into higher-dimensional space to locate the best separating hyperplane.\u003c/p\u003e \u003cp\u003eIn addition, we utilized the Multilayer Perceptron (MLP)[\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e], a kind of artificial neural network designed for the recognition of complex patterns. We modified hyperparameters such as hidden layer sizes (hidden_layer_sizes: (50,), (100,), (50, 50)), activation functions (activation: 'tanh', 'relu'), solvers (solver: 'sgd', 'adam'), initial learning rates (learning_rate_init: 0.001, 0.01), and maximum iterations (max_iter: 10000). The MLP learns complex patterns and relationships through multiple hidden layers and nonlinear activation functions.\u003c/p\u003e \u003cp\u003eLastly, we optimized the Extreme Gradient Boosting (XGBoost) model[\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e], a sophisticated ensemble learning technique that operates on a gradient boosting framework. Hyperparameters, including the maximum tree depth (max_depth: 3, 5, 7, 10, 15), were tuned for improved performance by successively adding new trees to address the errors made by earlier trees.\u003c/p\u003e \u003cp\u003eThroughout the training and evaluation processes, various key performance metrics were employed to evaluate the effectiveness of the models. The Area Under the ROC Curve (AUC) offers a comprehensive assessment of the model's classification abilities across all potential thresholds. Sensitivity (SE) gauges the model's proficiency in accurately identifying true positive cases, while Specificity (SP) evaluates its success in correctly detecting true negative instances. The Matthews Correlation Coefficient (MCC) considers all four sections of the confusion matrix, measuring the quality of binary classifications on a scale from \u0026minus;\u0026thinsp;1 (complete misclassification) to 1 (perfect classification). Accuracy (ACC) indicates the ratio of correctly classified samples, and Precision (P) refers to the share of true positives among those predicted as positive. The F1 Score (F1), which is the harmonic mean of precision and recall, is particularly beneficial for imbalanced datasets. Balanced Accuracy (BA), calculated as the average of SE and SP, provides a just evaluation for imbalanced classes. These metrics allow for a thorough assessment of each model's performance in predicting hematotoxicity, ensuring precise identification of both positive and negative samples.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results and Discussion","content":"\u003cdiv id=\"Sec11\"\u003e\n \u003ch2\u003e3.1 Dataset Molecular Feature Evaluation\u003c/h2\u003e\n \u003cdiv\u003e\n \u003cp\u003eThe histograms in Fig. 2 illustrate the distribution of six molecular descriptors—molecular weight (MolWt), topological polar surface area (TPSA), number of hydrogen bond donors (NumHDonors), number of hydrogen bond acceptors (NumHAcceptors), LogP, and molecular volume (MolVolume)—for toxic and non-toxic compounds. For molecular weight (Fig. 2A), toxic compounds show a slightly higher tendency toward larger molecular weights compared to non-toxic compounds, though the overall distribution is highly overlapping. In terms of topological polar surface area (Fig. 2B), toxic compounds exhibit a broader range with a shift toward higher values, indicating that polar surface area may play a role in toxicity. The distribution of the number of hydrogen bond donors (Fig. 2C) suggests that toxic compounds tend to have slightly more donors than non-toxic compounds, although the majority of both groups have relatively low values. A similar trend is observed for the number of hydrogen bond acceptors (Fig. 2D), where toxic compounds display a wider range and higher values on average, suggesting a potential correlation between acceptor count and toxicity. For LogP (Fig. 2E), the distribution is nearly identical between toxic and non-toxic compounds, indicating that lipophilicity alone may not be a significant determinant of toxicity. Finally, the molecular volume distribution (Fig. 2F) reveals a slight shift toward larger volumes in toxic compounds compared to non-toxic ones, although the overlap remains substantial. These findings highlight the potential influence of certain molecular descriptors, particularly TPSA, NumHDonors, and NumHAcceptors, in distinguishing toxic compounds from non-toxic ones.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec12\"\u003e\n \u003ch2\u003e3.2 Targets of hematotoxic small molecules\u003c/h2\u003e\n \u003cdiv\u003e\n \u003cp\u003eClustering analysis grouped the molecules into four distinct categories, as shown in Fig.\u0026nbsp;3 and Table S1. The molecules are grouped into four distinct clusters, each represented in a different color. The representative molecules (labeled \"REP\") for each cluster are highlighted, with their chemical structures displayed alongside. The representative molecule for the first cluster is L-Penicillamine, which contains 221 molecules. The second cluster is represented by Acetanilide, consisting of 460 molecules. The third cluster is represented by 1-[4-(4-Amino-7-propan-2-ylpyrrolo[2,3-d]pyrimidin-5-yl)phenyl]-3-(5-tert-butyl-1,2-oxazol-3-yl)urea, containing 40 molecules. The fourth cluster is represented by (7bR,8aS)-6-methyl-4-oxo-2-(5,6,7-trimethoxy-1H-indole-2-carbonyl)-1,2,4,5,8,8a-hexahydrocyclopropa[c]pyrrolo[3,2-e]indole-7-carbaldehyde, with 38 molecules. Cluster analysis showed that the molecular data set has intrinsic chemical heterogeneity, which may lead to multiple mechanisms of toxicity. Identifying representative compounds helps to understand structural diversity and select candidate compounds for further analysis.\u003c/p\u003e\n \u003c/div\u003e\n \u003cp\u003eTable 1. Molecular information of the four clusters.\u003c/p\u003e\n \u003cdiv\u003e\n \u003ctable id=\"Tab1\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 1\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003eMolecular information of the four clusters.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"2\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eSMILES\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ecluster\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSC([C@H](N)C(= O)O)(C)C\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eO = C(Nc1ccccc1)C\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eO = C(Nc1noc(C(C)(C)C)c1)Nc1ccc(-c2c3c(N)ncnc3n(C(C)C)c2)cc1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eO = C(N1C = 2[C@]3([C@@H](C1)C3)c1c(C = O)c(C)[nH]c1C(= O)C = 2)c1[nH]c2c(OC)c(OC)c(OC)cc2c1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003eFigure 4 presented similarity heatmaps for the molecules within each cluster. The heatmaps visualize pairwise similarity scores among molecules, with darker red regions indicating higher similarity. Clusters with more intense red regions have molecules that share more common structural features, suggesting a tighter chemical space. Clusters 1, 3 and 4 display a relatively higher degree of internal similarity, implying that their molecules have closely related structures. Heatmap of Cluster 2 shows more diversity, with less intense red regions, indicating a broader chemical space within cluster.\u003c/p\u003e\n \u003c/div\u003e\n \u003cp\u003eUsing multiple databases, we predicted targets and intersected the results with collected disease targets related to hematotoxicity. As shown in Fig.\u0026nbsp;5, cluster_1 intersected with 47 genes, cluster_2 with 148 genes, cluster_3 with 123 genes, and cluster_4 with 110 genes. Using the maximum clique centrality (MCC) algorithm in the cytoHubba toolkit, we identified key nodes in the hematotoxicity interactome. The MCC score represents the strength of connectivity and is visually presented in different color intensities, where darker shades indicate higher correlation with hematologic toxicity. We then cataloged the targets for each small molecule, revealing key players such as HSP90AA1 and SRC, as shown in Fig.\u0026nbsp;6.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec13\"\u003e\n \u003ch2\u003e3.3 KEGG and GO analysis\u003c/h2\u003e\n \u003cp\u003eThe Fig.\u0026nbsp;7, Fig.\u0026nbsp;8, Fig.\u0026nbsp;9, and Fig.\u0026nbsp;10 collectively present the results of GO and KEGG pathway enrichment analyses for genes related to small molecules and hematologic toxicity. The GO enrichment analyses (A of each figure) identify key biological processes, cellular components, and molecular functions. In the first figure, functions and signaling pathways related to cancer, immune regulation, cell migration, and wound healing were significantly enriched. The results of KEGG further emphasized the importance of cancer-related signaling pathways and immune signals, especially infection-related pathways. This provides a possible biological background for further study of these genes. In the second figure, enriched biological processes include response to xenobiotic stimulus, lipopolysaccharide, and bacterial molecules, emphasizing cellular responses to abiotic and biotic stimuli. Processes related to oxidative stress, radiation, and UV response are also significant, alongside regulation of apoptotic signaling pathways and hormone metabolic processes. The key KEGG pathways include fluid shear stress and atherosclerosis, AGE-RAGE signaling pathway in diabetic complications, and PD-L1 expression and PD-1 checkpoint pathway in cancer. Additionally, pathways like serotonergic synapse, chemical carcinogenesis via DNA adducts, and epithelial cell signaling in Helicobacter pylori infection are enriched. Specific disease-related pathways such as Chagas disease, tuberculosis, and toxoplasmosis further emphasize the clinical relevance of these findings.\u003c/p\u003e\n \u003cp\u003eThe third figure, enriched biological processes related to blood toxicity include chemotaxis and cell signaling, highlighting pathways such as protein autophosphorylation, MAPK cascade regulation, and cellular responses to peptide hormones, all of which are crucial in hematopoietic and vascular systems. Key KEGG pathways emphasize platelet activation, VEGF signaling, and chemokine signaling, which are directly associated with blood and vascular responses. Cancer-related pathways like acute myeloid leukemia and non-small-cell lung cancer, as well as the PD-L1 checkpoint pathway, are relevant to the hematopoietic system’s regulation under toxic stress. In the last figure, enriched biological processes associated with blood toxicity include leukocyte migration, chemotaxis, and blood coagulation, highlighting critical responses to hypoxia and oxygen level changes. Processes like peptide hormone response and regulation of body fluid levels further connect to vascular and hematopoietic system function. Key KEGG pathways emphasize platelet activation, VEGF signaling, and phosphatidylinositol 3-kinase/protein kinase B (PI3K-Akt) signaling, which are crucial for blood flow regulation and toxicity responses. Cancer-related pathways, such as PD-L1 checkpoint, melanoma, and chronic myeloid leukemia, suggest links between hematological abnormalities and toxic stress.\u003c/p\u003e\n \u003cp\u003eCollectively, these analyses suggest that the interaction between hematotoxic factors and blood toxicity involves intricate networks of inflammation, cellular signaling, and metabolic regulation. Enriched pathways highlight chemotaxis, coagulation, and oxygen response, alongside signaling cascades such as PI3K-Akt, Lipid and atherosclerosis, and MAPK, which are critical for vascular integrity and hematopoietic balance. The findings also reveal links to cancer, immune dysfunction, and systemic stress responses, emphasizing broad implications for cancer progression, immune regulation, and metabolic disorders in the context of hematotoxicity.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec14\"\u003e\n \u003ch2\u003e3.4 Molecular docking\u003c/h2\u003e\n \u003cp\u003eDue to the inclusion of HSP90AA1 and SRC as the key targets for the four categories of molecules and considering the importance of HSP90AA1 and SRC in hematotoxicity, we performed molecular docking to screen the affinity of representative hematotoxic small molecules for HSP90AA1 and SRC. The docking energy results are shown in Table 2.\u003c/p\u003e\n \u003cdiv\u003e\n \u003ctable id=\"Tab2\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv\u003eTable 2\u003c/div\u003e\n \u003cdiv\u003e\n \u003cp\u003eMolecules and docking energy results.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"3\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMolecule\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eTarget\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAffinity (kcal/mol)\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ecluster_1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eHSP90AA1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e-4.9\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ecluster_2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eHSP90AA1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e-5.6\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ecluster_3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eHSP90AA1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e-8.9\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ecluster_4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eHSP90AA1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e-9.6\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ecluster_1\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSRC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e-4.6\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ecluster_2\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSRC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e-5.8\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ecluster_3\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSRC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e-9.3\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ecluster_4\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eSRC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e-9.6\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cp\u003eTable 2. Molecules and docking energy results.\u003c/p\u003e\n \u003cp\u003eThe docking results of four representative molecules and two proteins are shown in Fig.\u0026nbsp;11. The ligand forms more diverse interactions with HSP90AA1, including halogen bonding and pi-sulfur interactions, which suggest a potentially higher binding specificity or affinity for this target. Pi-Pi Stacked and Pi-Alkyl interactions involve aromatic and hydrophobic residues such as PHE138, TYR200, and LEU107.Conventional Hydrogen Bonds are formed with residues like ASP93 and GLN123, stabilizing the ligand. Pi-Sulfur interactions likely involve CYS110, contributing to unique stabilization. Halogen Bonding is observed with residues like SER52 or nearby polar residues capable of accommodating halogen atoms. For SRC, the interactions rely heavily on amide stacking, hydrogen bonding, and hydrophobic contacts, indicating a strong but potentially more hydrophobic-driven binding. Amide-Pi Stacking interactions involve backbone or side-chain amides of residues like GLN276 and ASN350. Conventional Hydrogen Bonds are formed with residues such as HIS327 and LYS295, enhancing specificity. Hydrophobic Interactions involve residues like VAL268 and ILE314, which stabilize the ligand through Pi-Alkyl interactions. Pi-Cation Interactions likely involve positively charged residues like ARG386 or LYS295, further enhancing binding affinity.\u003c/p\u003e\n \u003cp\u003eThese residue-specific interactions underline the molecular basis for the ligand’s differential binding affinity and specificity for HSP90AA1 and SRCd, highlighting their potential as dual-target ligands with complementary binding mechanisms.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec15\"\u003e\n \u003ch2\u003e3.5 Machine Learning\u003c/h2\u003e\n \u003cp\u003eThe heatmap in Fig.\u0026nbsp;12 presents the performance metrics for various machine learning models trained on different molecular fingerprints. The metrics evaluated include sensitivity (SE), specificity (SP), accuracy (ACC), F1 score (F1), precision (P), balanced accuracy (BA), and area under the receiver operating characteristic curve (AUC). Among the models, the Random Forest algorithm consistently demonstrates strong performance across multiple fingerprint types, with the MACCS and RDKit fingerprints yielding the highest AUC values (0.76 and 0.77, respectively). Similarly, the AtomPairsFP-RandomForest combination achieves competitive results, with an AUC of 0.77. The Support Vector Machine (SVM) models perform slightly lower overall but still achieve reasonable specificity, particularly with the AtomPairsFP-SVM and Morgan-SVM models, both reaching an AUC of 0.75. The Multi-Layer Perceptron (MLP) models show moderate performance, with the MACCS-MLP model achieving the highest AUC (0.75) among MLP-based models. XGBoost models generally provide consistent results, with AtomPairsFP-XGBoost standing out with the highest overall AUC (0.78) among all model and fingerprint combinations.\u003c/p\u003e\n \u003cp\u003eSensitivity (SE) is lower for Morgan-SVM (0.50), indicating a potential limitation in identifying true positive cases for this combination. In contrast, specificity (SP) scores are notably high for Random Forest models with Morgan and AtomPairsFP fingerprints, suggesting robust performance in distinguishing true negative cases. Overall, the combination of Random Forest and AtomPairsFP yields the most balanced performance across metrics, making it a suitable choice for predicting toxicity with high reliability.\u003c/p\u003e\n \u003cdiv\u003e\n \u003cp\u003eFigure 13 displays the Receiver Operating Characteristic (ROC) curves and Area Under the Curve (AUC) values for four machine learning models—Multi-Layer Perceptron (MLP), Random Forest, Support Vector Machine (SVM), and XGBoost—trained using five different molecular fingerprints. In Panel A, the ROC curves for MLP models show that the MACCS fingerprint achieves the highest AUC (0.75), followed by Morgan and RDKit fingerprints, both with AUCs of 0.74. Topological Torsion demonstrates the lowest performance with an AUC of 0.68. Panel B illustrates the ROC curves for Random Forest models, where RDKit and AtomPairsFP fingerprints yield the highest AUCs (0.77), closely followed by MACCS and Morgan (both at 0.76). Panel C presents the SVM model results, where MACCS achieves the highest AUC (0.77), followed by RDKit (0.76) and AtomPairsFP (0.75), while Topological Torsion again exhibits the lowest AUC (0.68). Panel D highlights the XGBoost models, where AtomPairsFP achieves the best overall performance with an AUC of 0.78, while MACCS and RDKit follow closely at 0.77 and 0.76, respectively. Across all models, the Topological Torsion fingerprint consistently underperforms, while AtomPairsFP and MACCS fingerprints are among the top-performing representations. The XGBoost model paired with AtomPairsFP demonstrates the highest predictive capability overall, emphasizing its suitability for this classification task.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eFigure 14.\u003c/strong\u003e Confusion matrices illustrating the classification performance of machine learning models paired with various molecular fingerprints. Panels represent results for (A) AtomPairsFP, (B) MACCS, (C) Morgan, (D) RDKit, and (E) Topological Torsion fingerprints across Multi-Layer Perceptron (MLP), Random Forest, Support Vector Machine (SVM), and XGBoost algorithms.\u003c/p\u003e\n \u003cp\u003eFigure 14 illustrates the confusion matrices for various model and fingerprint combinations, emphasizing their ability to differentiate toxic (hematotoxic) and non-toxic compounds based on true positive rates (recall for the toxic class) and true negative rates (specificity for the non-toxic class). In Panel A, models using the AtomPairsFP fingerprint exhibit noticeable variability across algorithms. The AtomPairsFP-XGBoost model achieves the highest true positive rate (67.11%) while maintaining a reasonable true negative rate (75.17%), highlighting its strength in identifying hematotoxic compounds. In contrast, the AtomPairsFP-SVM model shows a significantly lower true positive rate (57.89%), potentially limiting its effectiveness in detecting toxic samples. The AtomPairsFP-RandomForest model excels in identifying non-toxic compounds, with a high true negative rate of 76.51%, underscoring its reliability in excluding non-hematotoxic molecules. Panel B evaluates models trained on MACCS fingerprints. The MACCS-XGBoost model delivers balanced performance, with a true positive rate of 65.13% and a true negative rate of 75.17%, making it suitable for identifying both hematotoxic and non-toxic compounds. Similarly, the MACCS-RandomForest model shows strong specificity (71.14%) and moderate sensitivity (68.42%), indicating it is slightly biased toward non-toxic samples but remains effective in predicting hematotoxicity. Panel C focuses on Morgan fingerprints, where the Morgan-RandomForest model demonstrates the highest true negative rate (77.85%), emphasizing its robustness in identifying non-toxic compounds. However, the Morgan-SVM model underperforms in terms of recall for hematotoxic compounds, achieving a lower true positive rate (59.87%), which may reduce its overall effectiveness in hematotoxicity classification. Panel D analyzes RDKit fingerprints, with both RDKit-XGBoost and RDKit-RandomForest models showing balanced performance. The RDKit-XGBoost model achieves a true positive rate of 65.13% and a true negative rate of 71.14%, indicating it is well-suited for predicting hematotoxicity. The RDKit-RandomForest model slightly improves on specificity, achieving 73.15%, but has a comparable recall for toxic compounds. Panel E evaluates Topological Torsion fingerprints, which generally underperform compared to other fingerprints across all models. The TopologicalTorsion-XGBoost model achieves a relatively high true negative rate (67.11%) but a modest true positive rate (61.18%), suggesting its limited utility for accurately identifying hematotoxic compounds.\u003c/p\u003e\n \u003cp\u003eOverall, XGBoost models paired with AtomPairsFP consistently demonstrate the best balance between true positive and true negative rates, making them highly effective for predicting both hematotoxic and non-toxic compounds. Random Forest models also show strong specificity, particularly when paired with Morgan fingerprints, making them reliable for identifying non-toxic samples. However, models using Topological Torsion fingerprints consistently underperform across metrics, highlighting their limited applicability in hematotoxicity classification. These findings reinforce the potential of fingerprint-driven machine learning models for early prediction of drug-induced hematotoxicity, providing valuable insights for safer drug development.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e"},{"header":"4. Conclusion","content":"\u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eIn conclusion, this study presents an integrative framework to explore the mechanisms underlying hematotoxicity in small molecules by employing clustering, reverse target prediction, molecular docking, and machine learning models. This comprehensive approach reveals the complexity of hematotoxic mechanisms and provides an efficient predictive method for assessing molecular hematotoxicity.\u003c/p\u003e \u003cp\u003eClustering analysis grouped the molecules into distinct clusters, each characterized by unique structural features and representative compounds. This chemical heterogeneity suggests multiple potential mechanisms of toxicity, emphasizing the importance of considering structural diversity in toxicity assessment. The clustering results served as the foundation for reverse target prediction, enabling the identification of key proteins such as HSP90AA1 and SRC, which were identified as central nodes within hematotoxicity-related networks. Target prediction and network analysis highlighted the role of these key proteins in mediating toxic effects, providing potential targets for therapeutic intervention or risk assessment. Pathway enrichment analyses further revealed that the identified targets are involved in critical biological processes and pathways, such as inflammation, immune regulation, cell migration, and cancer signaling. These findings suggest that hematotoxic compounds may disrupt these processes, leading to adverse hematological effects. Molecular docking studies provided mechanistic insights by demonstrating significant binding affinities of representative compounds to key proteins such as HSP90AA1 and SRC. The diverse interactions observed, including hydrogen bonds, π\u0026ndash;π stacking, and hydrophobic contacts, indicate that these compounds can effectively bind and potentially inhibit these proteins, leading to disruptions in critical signaling pathways. This highlights the complexity of the mechanisms involved in hematotoxicity. Finally, a series of machine learning models were constructed, particularly using the XGBoost algorithm combined with AtomPairsFP fingerprints, to provide an efficient predictive tool for hematotoxicity assessment. These models demonstrated high accuracy in classification, effectively balancing sensitivity and specificity. This enables the early identification of potentially toxic compounds and reduces the need for animal testing.\u003c/p\u003e \u003cp\u003eIn summary, this study advances our understanding of hematotoxicity by elucidating the complex relationships between molecular structures, biological targets, and toxicological outcomes. The integrative methodology presented here underscores the importance of a multidisciplinary approach, connecting structural characteristics to biological effects through interconnected analyses. This framework provides valuable insights for the development of safer drugs, enabling the early detection of toxic compounds, informing strategies to mitigate adverse effects, and guiding the design of new compounds with improved safety profiles.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAuthor Contributions:\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u003c/strong\u003e None.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eInstitutional Review Board Statement:\u0026nbsp;\u003c/strong\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eInformed Consent Statement:\u0026nbsp;\u003c/strong\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability Statement:\u003c/strong\u003e All data and material used to support the findings of this article are included within this article.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflicts of Interest:\u003c/strong\u003e The authors declare no conflicts of interest.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eMathew, J., P. Sankar, and M. Varacallo, \u003cem\u003ePhysiology, blood plasma.\u003c/em\u003e 2018.\u003c/li\u003e\n\u003cli\u003eMintzer, D.M., S.N. Billet, and L. Chmielewski, \u003cem\u003eDrug‐induced hematologic syndromes.\u003c/em\u003e Advances in hematology, 2009. \u003cstrong\u003e2009\u003c/strong\u003e(1): p. 495863.\u003c/li\u003e\n\u003cli\u003eBudinsky Jr, R.A., \u003cem\u003eHematotoxicity: chemically induced toxicity of the blood.\u003c/em\u003e Principles of Toxicology: Environmental and Industrial Applications, 2000: p. 87-109.\u003c/li\u003e\n\u003cli\u003eKrishnan, K. and M. Pelekis, \u003cem\u003eHematotoxic interactions: occurrence, mechanisms and predictability.\u003c/em\u003e Toxicology, 1995. \u003cstrong\u003e105\u003c/strong\u003e(2-3): p. 355-364.\u003c/li\u003e\n\u003cli\u003eMaxwell, M.B. and K.E. Maher. \u003cem\u003eChemotherapy-induced myelosuppression\u003c/em\u003e. in \u003cem\u003eSeminars in oncology nursing\u003c/em\u003e. 1992.\u003c/li\u003e\n\u003cli\u003eBilinski, J., et al., \u003cem\u003eFecal microbiota transplantation in patients with blood disorders inhibits gut colonization with antibiotic-resistant bacteria: results of a prospective, single-center study.\u003c/em\u003e Clinical Infectious Diseases, 2017. \u003cstrong\u003e65\u003c/strong\u003e(3): p. 364-370.\u003c/li\u003e\n\u003cli\u003eRobak, P., P. Smolewski, and T. Robak, \u003cem\u003eThe role of non-steroidal anti-inflammatory drugs in the risk of development and treatment of hematologic malignancies.\u003c/em\u003e Leukemia \u0026amp; lymphoma, 2008. \u003cstrong\u003e49\u003c/strong\u003e(8): p. 1452-1462.\u003c/li\u003e\n\u003cli\u003eGreene, N. and R. Naven, \u003cem\u003eEarly toxicity screening strategies.\u003c/em\u003e Curr Opin Drug Discov Devel, 2009. \u003cstrong\u003e12\u003c/strong\u003e(1): p. 90-97.\u003c/li\u003e\n\u003cli\u003eVan Norman, G.A., \u003cem\u003eLimitations of animal studies for predicting toxicity in clinical trials: is it time to rethink our current approach?\u003c/em\u003e JACC: Basic to Translational Science, 2019. \u003cstrong\u003e4\u003c/strong\u003e(7): p. 845-854.\u003c/li\u003e\n\u003cli\u003eZhao, W., et al., \u003cem\u003eUnveiling Anti-Diabetic Potential of Baicalin and Baicalein from Baikal Skullcap: LC-MS, In Silico, and In Vitro Studies.\u003c/em\u003e Int J Mol Sci, 2024. \u003cstrong\u003e25\u003c/strong\u003e(7).\u003c/li\u003e\n\u003cli\u003eWang, K., et al., \u003cem\u003eExploring the anti-gout potential of sunflower receptacles alkaloids: A computational and pharmacological analysis.\u003c/em\u003e Comput Biol Med, 2024. \u003cstrong\u003e172\u003c/strong\u003e: p. 108252.\u003c/li\u003e\n\u003cli\u003eSu, J., et al., \u003cem\u003eIntegrating Computational and Experimental Methods to Identify Novel Sweet Peptides from Egg and Soy Proteins.\u003c/em\u003e Int J Mol Sci, 2024. \u003cstrong\u003e25\u003c/strong\u003e(10).\u003c/li\u003e\n\u003cli\u003eEkins, S., \u003cem\u003eProgress in computational toxicology.\u003c/em\u003e Journal of pharmacological and toxicological methods, 2014. \u003cstrong\u003e69\u003c/strong\u003e(2): p. 115-140.\u003c/li\u003e\n\u003cli\u003eRaies, A.B. and V.B. Bajic, \u003cem\u003eIn silico toxicology: computational methods for the prediction of chemical toxicity.\u003c/em\u003e Wiley Interdisciplinary Reviews: Computational Molecular Science, 2016. \u003cstrong\u003e6\u003c/strong\u003e(2): p. 147-172.\u003c/li\u003e\n\u003cli\u003eCui, H., et al., \u003cem\u003eComputational Insights into Reproductive Toxicity: Clustering, Mechanism Analysis, and Predictive Models.\u003c/em\u003e Int J Mol Sci, 2024. \u003cstrong\u003e25\u003c/strong\u003e(14).\u003c/li\u003e\n\u003cli\u003eLiu, K., et al., \u003cem\u003eGPT4Kinase: High-accuracy prediction of inhibitor-kinase binding affinity utilizing large language model.\u003c/em\u003e Int J Biol Macromol, 2024. \u003cstrong\u003e282\u003c/strong\u003e(Pt 5): p. 137069.\u003c/li\u003e\n\u003cli\u003eDong, J., et al., \u003cem\u003eADMETlab: a platform for systematic ADMET evaluation based on a comprehensively collected ADMET database.\u003c/em\u003e Journal of cheminformatics, 2018. \u003cstrong\u003e10\u003c/strong\u003e: p. 1-11.\u003c/li\u003e\n\u003cli\u003eDaina, A., O. Michielin, and V. Zoete, \u003cem\u003eSwissADME: a free web tool to evaluate pharmacokinetics, drug-likeness and medicinal chemistry friendliness of small molecules.\u003c/em\u003e Scientific reports, 2017. \u003cstrong\u003e7\u003c/strong\u003e(1): p. 42717.\u003c/li\u003e\n\u003cli\u003eHua, Y., et al., \u003cem\u003eIn silico prediction of chemical-induced hematotoxicity with machine learning and deep learning methods.\u003c/em\u003e Molecular Diversity, 2021. \u003cstrong\u003e25\u003c/strong\u003e(3): p. 1585-1596.\u003c/li\u003e\n\u003cli\u003eKuhn, M., et al., \u003cem\u003eThe SIDER database of drugs and side effects.\u003c/em\u003e Nucleic acids research, 2016. \u003cstrong\u003e44\u003c/strong\u003e(D1): p. D1075-D1079.\u003c/li\u003e\n\u003cli\u003eSushko, I., et al., \u003cem\u003eOnline chemical modeling environment (OCHEM): web platform for data storage, model development and publishing of chemical information.\u003c/em\u003e Journal of computer-aided molecular design, 2011. \u003cstrong\u003e25\u003c/strong\u003e: p. 533-554.\u003c/li\u003e\n\u003cli\u003eWang, Y., et al., \u003cem\u003ePubChem\u0026apos;s BioAssay database.\u003c/em\u003e Nucleic acids research, 2012. \u003cstrong\u003e40\u003c/strong\u003e(D1): p. D400-D412.\u003c/li\u003e\n\u003cli\u003eB\u0026ouml;hm, R., et al., \u003cem\u003eOpenVigil FDA\u0026ndash;inspection of US American adverse drug events pharmacovigilance data and novel clinical applications.\u003c/em\u003e PloS one, 2016. \u003cstrong\u003e11\u003c/strong\u003e(6): p. e0157753.\u003c/li\u003e\n\u003cli\u003eWeininger, D., \u003cem\u003eSMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules.\u003c/em\u003e Journal of chemical information and computer sciences, 1988. \u003cstrong\u003e28\u003c/strong\u003e(1): p. 31-36.\u003c/li\u003e\n\u003cli\u003eHeller, S.R., et al., \u003cem\u003eInChI, the IUPAC international chemical identifier.\u003c/em\u003e Journal of cheminformatics, 2015. \u003cstrong\u003e7\u003c/strong\u003e: p. 1-34.\u003c/li\u003e\n\u003cli\u003eSwain, M., \u003cem\u003ePubChemPy: A way to interact with PubChem in Python\u003c/em\u003e. 2014.\u003c/li\u003e\n\u003cli\u003eVilar, S., G. Cozza, and S. Moro, \u003cem\u003eMedicinal chemistry and the molecular operating environment (MOE): application of QSAR and molecular docking to drug discovery.\u003c/em\u003e Current topics in medicinal chemistry, 2008. \u003cstrong\u003e8\u003c/strong\u003e(18): p. 1555-1572.\u003c/li\u003e\n\u003cli\u003eWishart, D.S., et al., \u003cem\u003eDrugBank 5.0: a major update to the DrugBank database for 2018.\u003c/em\u003e Nucleic acids research, 2018. \u003cstrong\u003e46\u003c/strong\u003e(D1): p. D1074-D1082.\u003c/li\u003e\n\u003cli\u003eBento, A.P., et al., \u003cem\u003eAn open source chemical structure curation pipeline using RDKit.\u003c/em\u003e Journal of Cheminformatics, 2020. \u003cstrong\u003e12\u003c/strong\u003e: p. 1-16.\u003c/li\u003e\n\u003cli\u003eLandrum, G., \u003cem\u003eRdkit documentation.\u003c/em\u003e Release, 2013. \u003cstrong\u003e1\u003c/strong\u003e(1-79): p. 4.\u003c/li\u003e\n\u003cli\u003eVoicu, A., et al., \u003cem\u003eThe rcdk and cluster R packages applied to drug candidate selection.\u003c/em\u003e Journal of Cheminformatics, 2020. \u003cstrong\u003e12\u003c/strong\u003e: p. 1-8.\u003c/li\u003e\n\u003cli\u003eArora, S., W. Hu, and P.K. Kothari. \u003cem\u003eAn analysis of the t-sne algorithm for data visualization\u003c/em\u003e. in \u003cem\u003eConference on learning theory\u003c/em\u003e. 2018. PMLR.\u003c/li\u003e\n\u003cli\u003eBajusz, D., A. R\u0026aacute;cz, and K. H\u0026eacute;berger, \u003cem\u003eWhy is Tanimoto index an appropriate choice for fingerprint-based similarity calculations?\u003c/em\u003e Journal of cheminformatics, 2015. \u003cstrong\u003e7\u003c/strong\u003e: p. 1-13.\u003c/li\u003e\n\u003cli\u003eSievert, C., \u003cem\u003eInteractive web-based data visualization with R, plotly, and shiny\u003c/em\u003e. 2020: Chapman and Hall/CRC.\u003c/li\u003e\n\u003cli\u003eAmberger, J.S., et al., \u003cem\u003eOMIM. org: Online Mendelian Inheritance in Man (OMIM\u0026reg;), an online catalog of human genes and genetic disorders.\u003c/em\u003e Nucleic acids research, 2015. \u003cstrong\u003e43\u003c/strong\u003e(D1): p. D789-D798.\u003c/li\u003e\n\u003cli\u003eSafran, M., et al., \u003cem\u003eGeneCards Version 3: the human gene integrator.\u003c/em\u003e Database, 2010. \u003cstrong\u003e2010\u003c/strong\u003e: p. baq020.\u003c/li\u003e\n\u003cli\u003eKnox, C., et al., \u003cem\u003eDrugBank 6.0: the DrugBank knowledgebase for 2024.\u003c/em\u003e Nucleic acids research, 2024. \u003cstrong\u003e52\u003c/strong\u003e(D1): p. D1265-D1275.\u003c/li\u003e\n\u003cli\u003eConsortium, U., \u003cem\u003eUniProt: a worldwide hub of protein knowledge.\u003c/em\u003e Nucleic acids research, 2019. \u003cstrong\u003e47\u003c/strong\u003e(D1): p. D506-D515.\u003c/li\u003e\n\u003cli\u003eNickel, J., et al., \u003cem\u003eSuperPred: update on drug classification and target prediction.\u003c/em\u003e Nucleic acids research, 2014. \u003cstrong\u003e42\u003c/strong\u003e(W1): p. W26-W31.\u003c/li\u003e\n\u003cli\u003eDaina, A., O. Michielin, and V. Zoete, \u003cem\u003eSwissTargetPrediction: updated data and new features for efficient prediction of protein targets of small molecules.\u003c/em\u003e Nucleic acids research, 2019. \u003cstrong\u003e47\u003c/strong\u003e(W1): p. W357-W364.\u003c/li\u003e\n\u003cli\u003eRao, V.S., et al., \u003cem\u003eProtein‐protein interaction detection: methods and analysis.\u003c/em\u003e International journal of proteomics, 2014. \u003cstrong\u003e2014\u003c/strong\u003e(1): p. 147648.\u003c/li\u003e\n\u003cli\u003eSzklarczyk, D., et al., \u003cem\u003eThe STRING database in 2021: customizable protein\u0026ndash;protein networks, and functional characterization of user-uploaded gene/measurement sets.\u003c/em\u003e Nucleic acids research, 2021. \u003cstrong\u003e49\u003c/strong\u003e(D1): p. D605-D612.\u003c/li\u003e\n\u003cli\u003eShannon, P., et al., \u003cem\u003eCytoscape: a software environment for integrated models of biomolecular interaction networks.\u003c/em\u003e Genome research, 2003. \u003cstrong\u003e13\u003c/strong\u003e(11): p. 2498-2504.\u003c/li\u003e\n\u003cli\u003eAshburner, M., et al., \u003cem\u003eGene ontology: tool for the unification of biology.\u003c/em\u003e Nature genetics, 2000. \u003cstrong\u003e25\u003c/strong\u003e(1): p. 25-29.\u003c/li\u003e\n\u003cli\u003eKanehisa, M. and S. Goto, \u003cem\u003eKEGG: kyoto encyclopedia of genes and genomes.\u003c/em\u003e Nucleic acids research, 2000. \u003cstrong\u003e28\u003c/strong\u003e(1): p. 27-30.\u003c/li\u003e\n\u003cli\u003eYu, G., et al., \u003cem\u003eclusterProfiler: an R package for comparing biological themes among gene clusters.\u003c/em\u003e Omics: a journal of integrative biology, 2012. \u003cstrong\u003e16\u003c/strong\u003e(5): p. 284-287.\u003c/li\u003e\n\u003cli\u003eCarlson, M., et al., \u003cem\u003eorg. Hs. eg. db: Genome wide annotation for Human.\u003c/em\u003e R package version, 2019. \u003cstrong\u003e3\u003c/strong\u003e(2): p. 3.\u003c/li\u003e\n\u003cli\u003eLuo, W. and C. Brouwer, \u003cem\u003ePathview: an R/Bioconductor package for pathway-based data integration and visualization.\u003c/em\u003e Bioinformatics, 2013. \u003cstrong\u003e29\u003c/strong\u003e(14): p. 1830-1831.\u003c/li\u003e\n\u003cli\u003eEberhardt, J., et al., \u003cem\u003eAutoDock Vina 1.2. 0: New docking methods, expanded force field, and python bindings.\u003c/em\u003e Journal of chemical information and modeling, 2021. \u003cstrong\u003e61\u003c/strong\u003e(8): p. 3891-3898.\u003c/li\u003e\n\u003cli\u003eO\u0026apos;Boyle, N.M., et al., \u003cem\u003eOpen Babel: An open chemical toolbox.\u003c/em\u003e Journal of cheminformatics, 2011. \u003cstrong\u003e3\u003c/strong\u003e: p. 1-14.\u003c/li\u003e\n\u003cli\u003eCopeland, R.A., \u003cem\u003eEvaluation of enzyme inhibitors in drug discovery: a guide for medicinal chemists and pharmacologists\u003c/em\u003e. 2013: John Wiley \u0026amp; Sons.\u003c/li\u003e\n\u003cli\u003eRogers, D. and M. Hahn, \u003cem\u003eExtended-connectivity fingerprints.\u003c/em\u003e Journal of chemical information and modeling, 2010. \u003cstrong\u003e50\u003c/strong\u003e(5): p. 742-754.\u003c/li\u003e\n\u003cli\u003eJow, H., et al., \u003cem\u003eMELCOR accident consequence code system (MACCS)\u003c/em\u003e. 1990, Nuclear Regulatory Commission, Washington, DC (USA). Div. of Systems \u0026hellip;.\u003c/li\u003e\n\u003cli\u003eP\u0026eacute;rez-Nueno, V.I., et al., \u003cem\u003eAPIF: a new interaction fingerprint based on atom pairs and its application to virtual screening.\u003c/em\u003e Journal of chemical information and modeling, 2009. \u003cstrong\u003e49\u003c/strong\u003e(5): p. 1245-1260.\u003c/li\u003e\n\u003cli\u003eBonach\u0026eacute;ra, F., et al., \u003cem\u003eFuzzy tricentric pharmacophore fingerprints. 1. Topological fuzzy pharmacophore triplets and adapted molecular similarity scoring schemes.\u003c/em\u003e Journal of chemical information and modeling, 2006. \u003cstrong\u003e46\u003c/strong\u003e(6): p. 2457-2477.\u003c/li\u003e\n\u003cli\u003eBelgiu, M. and L. Drăguţ, \u003cem\u003eRandom forest in remote sensing: A review of applications and future directions.\u003c/em\u003e ISPRS journal of photogrammetry and remote sensing, 2016. \u003cstrong\u003e114\u003c/strong\u003e: p. 24-31.\u003c/li\u003e\n\u003cli\u003eWang, L., \u003cem\u003eSupport vector machines: theory and applications\u003c/em\u003e. Vol. 177. 2005: Springer Science \u0026amp; Business Media.\u003c/li\u003e\n\u003cli\u003eTaud, H. and J.-F. Mas, \u003cem\u003eMultilayer perceptron (MLP).\u003c/em\u003e Geomatic approaches for modeling land change scenarios, 2018: p. 451-455.\u003c/li\u003e\n\u003cli\u003eChen, T., \u003cem\u003eXgboost: extreme gradient boosting.\u003c/em\u003e R package version 0.4-2, 2015. \u003cstrong\u003e1\u003c/strong\u003e(4).\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Table S1","content":"\u003cp\u003eTable S1 is not available with this version.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Hematotoxicity, Machine learning, Network Pharmacology","lastPublishedDoi":"10.21203/rs.3.rs-6295317/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6295317/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eHematotoxicity, encompassing adverse effects such as anemia, leukopenia, thrombocytopenia, and coagulation disorders, is a critical yet underexplored area of toxicological research. These toxic effects can lead to severe clinical outcomes, including heightened risks of infection, bleeding, and mortality. Despite its significance, hematotoxicity research lags behind general toxicities like hepatotoxicity and nephrotoxicity, and traditional evaluation methods such as animal models and in vitro assays often fail to accurately predict human responses. To address this gap, we curated a dataset of thousands of compounds with and without hematotoxic effects and performed in-depth analyses of their molecular properties and mechanisms using clustering and target prediction. These analyses revealed key pathways and targets underlying hematotoxicity, demonstrating its complex and multifactorial nature. We developed predictive models using fingerprints, combined with machine learning algorithms such as Random Forest (RF), Support Vector Machine (SVM), and XGBoost. The best-performing model achieved an AUC of 0.78, highlighting its potential for accurately identifying hematotoxic compounds. This study provides a computational framework for understanding hematotoxicity mechanisms and offers a practical tool for early screening of blood-toxic compounds during drug development. These advancements pave the way for safer therapeutic strategies, improved patient safety, and reduced risks of drug-induced hematological disorders.\u003c/p\u003e","manuscriptTitle":"Computational Prediction of Drug-Induced Hematotoxicity: Mechanisms and Model Development","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-05-07 06:54:33","doi":"10.21203/rs.3.rs-6295317/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"81341b10-9a6d-463a-9ac5-7104fd395688","owner":[],"postedDate":"May 7th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-07-06T01:53:12+00:00","versionOfRecord":[],"versionCreatedAt":"2025-05-07 06:54:33","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6295317","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6295317","identity":"rs-6295317","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.