Pharmacogenomics-Driven Multimodal Data Integration Improves Predictions of Adverse Drug Reactions in Cancer Patients using Machine Learning

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-16

This study developed a pharmacogenomics-driven machine learning framework integrating genomic, environmental, and comorbidity data to improve prediction of adverse drug reactions in cancer patients.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-16 · read from full text

This preprint used UK Biobank data to study 26,235 cancer patients treated with antineoplastic drugs, aiming to predict adverse drug reactions (ADRs) defined by ICD-10 codes via a pharmacogenomics-driven machine learning framework that integrated GWAS-derived SNPs from 169 pharmacogenes, PharmGKB variants, demographics/lifestyle, laboratory biomarkers, and comorbidities. Across five supervised models (logistic regression, random forest, SVM, XGBoost, and multilayer perceptron), logistic regression and multilayer perceptrons performed best, with genetic data alone reaching AUC-ROC values up to 0.82–0.80 in drug-specific cohorts and improving when all features were added (e.g., up to 0.85–0.86); for secondary thrombocytopenia, performance was AUC-ROC 0.94 with genetics only and 0.97 with all features. The authors report that univariate and SHAP analyses highlighted female gender, elevated cystatin C, and alkaline phosphatase as key contributors, and that certain cancer types and comorbidities were enriched among ADR cases, while noting the study is a preprint and not peer reviewed. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Accurately predicting adverse drug reactions (ADRs) in cancer remains challenging. We applied a pharmacogenomics-driven machine learning framework that integrates genomic, environmental, and comorbidity data to enhance ADR prediction. Using UK Biobank, we analysed 26,235 antineoplastic-treated patients, identifying ADRs via ICD-10 codes. Features included GWAS-derived SNPs from 169 pharmacogenes, curated PharmGKB variants, demographics, lifestyle, laboratory biomarkers, and comorbidities. Five supervised models were trained; subgroup analyses assessed drug-specific and ADR-specific cohort performance. Logistic regression and multilayer perceptron models performed best. In drug-specific cohort, genetic data alone achieved AUC-ROC 0.82 (LR) and 0.80 (MLP), improving to 0.85 and 0.86 when all features were included. For secondary thrombocytopenia, LR and MLP achieved AUC-ROC 0.94 using genetic data only and 0.97 with all features. SHAP and univariate analyses highlighted female gender, elevated cystatin C, and alkaline phosphatase (all p < 0.001); haematologic and digestive cancers showed higher risk compared to other cancer types. This integrative approach supports data-driven clinical decision-making to reduce ADRs.
Full text 173,520 characters · extracted from preprint-html · click to expand
Pharmacogenomics-Driven Multimodal Data Integration Improves Predictions of Adverse Drug Reactions in Cancer Patients using Machine Learning | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Pharmacogenomics-Driven Multimodal Data Integration Improves Predictions of Adverse Drug Reactions in Cancer Patients using Machine Learning Anoop Joseph, Vignesh Arunachalam, Arabella Hart, Rodney Lea, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7431071/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Accurately predicting adverse drug reactions (ADRs) in cancer remains challenging. We applied a pharmacogenomics-driven machine learning framework that integrates genomic, environmental, and comorbidity data to enhance ADR prediction. Using UK Biobank, we analysed 26,235 antineoplastic-treated patients, identifying ADRs via ICD-10 codes. Features included GWAS-derived SNPs from 169 pharmacogenes, curated PharmGKB variants, demographics, lifestyle, laboratory biomarkers, and comorbidities. Five supervised models were trained; subgroup analyses assessed drug-specific and ADR-specific cohort performance. Logistic regression and multilayer perceptron models performed best. In drug-specific cohort, genetic data alone achieved AUC-ROC 0.82 (LR) and 0.80 (MLP), improving to 0.85 and 0.86 when all features were included. For secondary thrombocytopenia, LR and MLP achieved AUC-ROC 0.94 using genetic data only and 0.97 with all features. SHAP and univariate analyses highlighted female gender, elevated cystatin C, and alkaline phosphatase (all p < 0.001); haematologic and digestive cancers showed higher risk compared to other cancer types. This integrative approach supports data-driven clinical decision-making to reduce ADRs. Health sciences/Biomarkers Biological sciences/Cancer Biological sciences/Computational biology and bioinformatics Health sciences/Oncology Machine Learning Pharmacogenomics Adverse drug reactions cancer comorbidities personalized medicine prediction model Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Adverse drug reactions (ADRs), defined by the World Health Organization (WHO) as “any response to a drug which is noxious and unintended, and which occurs at doses normally used in humans for prophylaxis, diagnosis, or therapy of disease, or for the modification of physiological function” ( 1 ), account for nearly 13% of all hospital admissions worldwide, of which over 1% are fatal ( 2 – 9 ). ADRs rank as the seventh leading cause of death in Sweden ( 10 ) and fourth in the United States ( 11 ), and result in increased costs for healthcare systems ( 12 , 13 ). Notably, ADRs are extremely prevalent in individuals being treated for cancer ( 14 – 16 ). During cancer treatment, complex medication regimens and the concurrent use of multiple drugs heighten the risk of adverse effects, with over 40% of individuals experiencing three or more ADRs, 88% of which are preventable ( 17 – 19 ). While adverse event reporting systems and predictive modelling have proven to be helpful in predicting relative drug risk ( 16 , 20 , 21 ), a large body of evidence has demonstrated that individuals are genetically predisposed to ADRs, highlighting the need to better understand the role of genetics in ADRs at both the population and individual levels ( 12 , 22 – 24 ). Pharmacogenomics (PGx), the study of the genetic basis of drug responses, focuses on the link between genetic variations and adverse drug reactions to inform individualized drug selection, maximize efficacy, and increase safety ( 25 – 27 ). PGx accounts for approximately 80% of all variability in drug efficacy and safety, where approximately half of all clinically relevant genes are associated with ADRs ( 25 ). Genomic variants have been shown to significantly affect patient responses to cancer treatments ( 28 , 29 ), and most individuals carry actionable PGx alleles pivotal in drug-gene interactions ( 30 , 31 ). However, uncovering the genetic basis of drug response remains challenging owing to its complex, multimodal nature, which is shaped not only by genetic variation but also by environmental influences, gene-drug and gene-gene interactions, comorbidities, age, and ethnic and population-level diversity ( 24 , 32 – 35 ). This is especially crucial in cancer, where medication-induced adverse reactions can affect various organ systems, with common effects including fatigue, anorexia, alopecia, constipation, nausea, vomiting, and neuropathy ( 36 – 39 ). Yet, current predictive approaches are limited by their reliance on single- or few-gene analyses and failure to integrate broader clinical context. Furthermore, models often overlook the complex, multimodal nature of ADRs, which are influenced by a combination of genetic, environmental (demographic, lifestyle, and clinical), and comorbid factors. Advanced machine learning (ML) techniques are increasingly being applied to large-scale genomic and environmental datasets to address the complexity of predicting ADRs ( 40 , 41 ). ML algorithms excel in managing large datasets and identifying complex, non-linear relationships that traditional statistical models may not adequately explain ( 42 ). Despite the rise of complex models, logistic regression (LR), a widely used statistical and ML algorithm for binary classification tasks, offers interpretable insights into the interplay of genetic and clinical factors ( 43 – 45 ). LR frameworks that integrate single-nucleotide polymorphisms (SNPs) from genome-wide association studies (GWAS) have enabled the stratification of lung cancer risk in Chinese populations ( 46 ), while another study using LR achieved an area under the receiver operating characteristic curve (AUC-ROC) of 0.87 for ADR prediction ( 47 ). Sophisticated deep learning neural networks such as the multilayer perceptron (MLP) model have improved predictions of disease susceptibility using SNP data with advanced feature selection techniques, achieving an AUC-ROC of 0.94 ( 48 ), while integrating genetic data with electronic health records is highly effective at predicting preventable adverse drug events using the Extreme Gradient Boosting (XGBoost) model achieving an AUC of 0.972 ( 49 ). Artificial neural networks (ANN), support vector machines (SVM), and random forests (RF) have been compared in their ability to predict hepatotoxicity induced by anti-tuberculosis drugs, with the ANN model integrating both clinical and genomic information providing the best outcomes ( 50 ). To address prevailing shortcomings in predicting ADRs among cancer patients, the present study systematically investigated the potential of multimodal data integration - including PGx variants, environmental exposures, and comorbid characteristics - to enhance individual-level ADR prediction using machine learning models. We leveraged 169 curated pharmacogenes to extract genetic features relevant to drug response. Our approach combined advanced informatics with predictive modeling to develop and validate algorithms designed to facilitate informed pharmacogenomic decision-making in near real-time at the point of care. Employing a comprehensive ML framework, we compared five models: logistic regression, random forest, support vector machine, XGBoost, and multilayer perceptron. To our knowledge, this is the first study to systematically integrate whole-genome pharmacogenomic variants with environmental (demographic, lifestyle, and clinical) and comorbidity data (derived from ICD-10 codes) for ADR prediction, offering a more holistic and personalized approach to cancer treatment. Results Univariate analyses reveal clinical, demographic, and biological contributors to ADRs In the complete cohort, univariate analyses identified several variables significantly associated with ADRs. Gender differences were observed, with females more likely to experience ADRs than males (p < 0.001). Lifestyle-related factors were also more prevalent among cases, including current smoking (p = 0.013), major dietary changes (p < 0.001), and frequent dietary variation (p = 0.046). Additionally, individuals in the case group reported spending significantly less time outdoors (p = 0.017) (Supplementary Table 4). Notably, there was no significant variation in the number of prescribed medications between cases and controls (p = 0.371), indicating that medication load alone does not account for ADR occurrence. However, several haematological and biochemical parameters differed significantly between groups. Elevated white blood cell counts (p = 0.009) and lymphocyte levels (p = 0.001), in addition to lower haemoglobin (p < 0.001) and direct bilirubin concentrations (p = 0.001), were significantly associated with ADRs, highlighting potential biological correlates of ADR vulnerability. Further analyses revealed that ADR incidence varied significantly by cancer type. Notably, 63% of patients in the case group were diagnosed with ill-defined, secondary, or unspecified cancer sites, compared to only 23% in the control group - a 2.7 fold difference between cases and controls (Fig. 2 ). In addition, lymphoid and hematopoietic malignancies, as well as cancers of the digestive system, were more frequently observed in the case group, underscoring the potential role of cancer pathology in modulating drug tolerance. Finally, comorbidities were significantly more prevalent among ADR cases. Chronic kidney disease (ICD-10: N18.9, p < 0.001) and liver disease (ICD-10: K76.9, p < 0.001) were significantly more prevalent, suggesting that impaired drug excretion and metabolism may exacerbate ADR risk. Other conditions (p < 0.001) more prevalent in case group included hypothyroidism (ICD-10: E03.9), pneumonia (ICD-10: J18.9), and hydronephrosis (ICD-10: N13.3). A complete list of comorbid conditions and their ICD-10 codes is provided in Supplementary Table 4. Prioritized pharmacogenomic variants identified via feature selection Feature selection approaches identified a substantial number of SNPs associated with ADR risk from pharmacogene regions. In the complete cohort, 3,601 SNPs surpassed the GWAS significance threshold; 3,271 SNPs were identified in drug-specific cohort analyses. For ADR-specific subgroups, the number of significant SNPs was 2,315 for secondary thrombocytopenia, 2,655 for drug-induced polyneuropathy, 3,181 for drug-induced gastroenteritis and colitis, and 3,521 for skin eruption. All these SNPs were selected from PGx regions and converted to allele dosage format for subsequent ML model development. Key pharmacogenomic variants prioritized through feature selection across drug-specific, ADR-specific, and complete cohort analyses are summarized below, with detailed cohort information available in Supplementary Table 5. These include SNPs with well-established relevance to antineoplastic drug toxicity, such as rs67376798 (OR: 1.34 [95% CI: 1.06, 1.69]) in the DPYD gene, which has been significantly associated with ADR risk in our analysis. This variant is known to reduce the activity of the dihydropyridine dehydrogenase (DPD) enzyme, a key player in the metabolism of 5-fluorouracil, and has strong clinical relevance in predicting toxicity to cancer drugs ( 51 ). Additional significant variants included rs4673993 (OR: 0.79 [95% CI: 0.63, 0.98]) in ATIC , rs2298383 (OR: 0.72 [95% CI: 0.59, 0.87]) in ADORA2A , and rs1801394 (OR: 0.64 [95% CI: 0.45, 0.91]) in MTRR , all previously associated with methotrexate toxicity, supported by PharmGKB Level 2B & Level 3 evidence( 52 , 53 ). In addition, variants rs1056892 (OR: 1.26 [95% CI: 1.03, 1.54]) in CBR3 , rs11615 (OR: 1.37 [95% CI: 1.12, 1.67]) in ERCC1 , and rs4646 (OR: 1.26 [95% CI: 1.02, 1.55]) in CYP19A1 have been linked to toxicity from other antineoplastic agents, including doxorubicin, cisplatin, capecitabine, and tamoxifen. Furthermore, rs16969968 (OR: 1.50 [95% CI: 1.06, 2.14]) in CHRNA5 , rs1051730 (OR: 1.49 [95% CI: 1.05, 2.11]) in CHRNA3 , rs3745274 (OR: 1.33 [95% CI: 1.09, 1.64]) in CYP2B6 , rs2298383 in ADORA2A , and rs11615 in ERCC1 were identified in multiple cohorts, strengthening their potential as robust pharmacogenomic markers of ADR risk. A complete list of significant SNPs with known antineoplastic drug associations is provided in Supplementary Table 5. Superior performance of MLP and LR using multimodal features in the drug-specific cohort A GWAS conducted in the drug-specific cohort identified 3,271 SNPs from pharmacogenes. In addition, clinically relevant SNPs from PharmKGB were incorporated into model development, including SNPs associated with fluorouracil (171), capecitabine ( 68 ), methotrexate (119), and tamoxifen ( 15 ) (Supplementary Table 2). Model performance metrics are summarized in Table 1 , with corresponding AUC-ROC curves shown in Fig. 3 . When using genetic features alone, both MLP and LR models demonstrated strong performance, achieving AUC-ROC values of 0.80 and 0.82, respectively. The MLP model achieved higher overall accuracy (0.83) and specificity (0.93), whereas LR exhibited superior sensitivity (0.66), highlighting its effectiveness in correctly identifying ADR cases. In contrast, the random forest (RF) model, despite high specificity (0.96), demonstrated low sensitivity (0.18), reducing its utility for ADR risk prediction. Incorporating environmental features led to only marginal improvements across all models. However, the inclusion of comorbidity-related variables significantly enhanced predictive performance. With the combined multimodal feature set (genetic, environmental, and comorbidity), the AUC-ROC increased to 0.86 for MLP and 0.85 for LR. The MLP model maintained superior performance, achieving an accuracy of 0.86, sensitivity of 0.56, and specificity of 0.94. These results underscore the effectiveness in capturing the complex interplay of genetic, environmental, and comorbid factors influencing ADR risk. Table 1 Performance metrics of ML models across different feature sets in drug-specific cohort (N = 1098; 246 Cases, 852 Controls). Model evaluations using genetic-only data ML Models Accuracy Precision Sensitivity Specificity F1 Score AUC-ROC LR 0.80 0.55 0.66 0.84 0.60 0.82 RF 0.79 0.64 0.19 0.97 0.29 0.74 SVM 0.82 0.64 0.49 0.92 0.55 0.79 XGB 0.81 0.58 0.49 0.90 0.53 0.79 MLP 0.84 0.69 0.51 0.94 0.57 0.80 Model evaluations using genetic and environmental data ML Models Accuracy Precision Sensitivity Specificity F1 Score AUC-ROC LR 0.81 0.57 0.66 0.86 0.61 0.83 RF 0.78 0.62 0.11 0.98 0.18 0.74 SVM 0.81 0.59 0.49 0.90 0.53 0.81 XGB 0.78 0.52 0.46 0.88 0.49 0.78 MLP 0.83 0.67 0.51 0.93 0.58 0.82 Model evaluations using genetic, environmental, and comorbidity data ML Models Accuracy Precision Sensitivity Specificity F1 Score AUC-ROC LR 0.82 0.60 0.58 0.89 0.59 0.85 RF 0.80 0.75 0.16 0.98 0.27 0.75 SVM 0.83 0.64 0.51 0.92 0.57 0.83 XGB 0.79 0.54 0.46 0.89 0.50 0.78 MLP 0.86 0.75 0.57 0.95 0.65 0.86 LR: Logistic Regression, RF: Random Forest, SVM: Support Vector Machine, GXB: Gradient Boosting (XGBoost), MLP: Multi-Layer Perceptron Multimodal feature integration improves the predictive performance of ADR phenotypes ML models were trained to classify individuals as either at high risk or low risk of developing specific treatment-associated ADR phenotypes (i.e., drug-induced gastroenteritis and colitis (ICD-10: K52.1, N = 3129), drug-induced polyneuropathy (G62.0, N = 1436), secondary thrombocytopenia (D69.5, N = 419), and generalized or localized skin eruptions (L27.0 & L27.1, N = 1277)). Across all ADR subtypes, model performance improved consistently with the integration of environmental and comorbidity features alongside genetic data. For secondary thrombocytopenia (D69.5), both LR and MLP models demonstrated exceptional predictive performance, with AUC-ROC values increasing from .94 (genetic-only model) to 0.97 following the inclusion of environmental and comorbidity features (Fig. 4 ). This trend of enhancement with multimodal feature integration was observed across all ADR phenotypes. For drug-induced polyneuropathy (G62.0), the best models achieved AUC-ROC values above .92 in models with the full-feature set (i.e., genetic, environmental, and comorbidity data), exhibiting a clear benefit from multimodal data integration. Although baseline genetic-only models performed moderately well (AUC-ROC of .89), the addition of environmental and comorbid features during model training substantially improved overall performance. Similarly, AUC-ROC values in the full-feature model reached .87 – .90 for gastroenteritis & colitis (K52.1) and generalized skin eruptions (L27.0 & L27.1), with noticeable improvements in recall and F1 scores. Across all ADR phenotypes, LR and MLP consistently outperformed ensemble methods such as RF and XGB. While RF and XGB offered high specificity, their relatively low sensitivity limited their effectiveness in identifying high-risk individuals (Supplementary Table 6). Reduced predictive performance in the complete cohort is attributed to phenotypic heterogeneity The performance of ML models in the complete case-control cohort (N = 26,235; 5194 cases and 21,041 controls) was lower compared to ADR-specific models, reflecting the increased heterogeneity and complexity of the combined dataset. The best-performing model was LR, achieving an AUC-ROC value of .78 with the full feature set. MLP achieved a maximum AUC-ROC value of .70 with the full feature set, while SVM (.75) performed better than both XGBoost (.72) and RF (.67) models (Supplementary Table 7). Across models, LR and MLP demonstrated higher sensitivity and balanced performance, except in complete cohort, whereas RF consistently exhibited poor sensitivity despite high specificity, likely owing to class imbalance. While XGBoost achieved reasonable discrimination, this model exhibited comparatively lower recall, highlighting trade-offs in detecting true ADR cases. To assess the standalone contribution of non-genetic features, an LR model trained solely on environmental variables achieved limited discrimination (AUC-ROC: 0.58) but revealed clinically meaningful predictors via SHAP analysis (Fig. 5 ). Female gender emerged as the most influential demographic factor, along with clinical biomarkers such as low haemoglobin, elevated cystatin C, and high alkaline phosphatase levels, indicating the potential role of kidney and liver function markers in ADR susceptibility. Discussion To the best of our knowledge, this is the first study to systematically integrate whole-genome pharmacogenomic variants with environmental, and ICD-10-derived comorbidities in ADR prediction, offering a more holistic and personalised approach to cancer treatment. Furthermore, the present study leveraged ML models to create a multimodal framework that captures the complex, multifactorial drivers of ADR risk, addressing critical gaps in current PGx-focused approaches and advancing individualized pharmacotherapy. Notably, we used a biologically informed feature selection strategy focused on 169 curated pharmacogenes, harnessing domain-specific knowledge to refine genomic input and minimize the noise typically inherent in large-scale WGS datasets, thereby improving the interpretability and clinical relevance of ML models. Our study shows that integrating pharmacogenomic, environmental and comorbidities data significantly improves ADR prediction performance compared to single-modality models, underscoring the value of multimodal integration for more accurate and clinically actionable risk stratification in oncology. As ADRs associated with oncological treatments pose a substantial burden on patient quality of life, improved predictions of ADR risk is of considerable clinical relevance ( 54 ). Previous studies have demonstrated the utility of WGS in identifying clinically relevant mutations and pharmacogenetic variants associated with drug-induced toxicity in cancer patients, highlighting its potential for personalized medicine ( 55 , 56 ). For example, PGx analyses within the 100,000 Genomes Project have shown that germline WGS can identify actionable variants linked to ADRs in cancer treatments, informing prescribing decisions and reducing toxicity risks ( 56 ). Data in the literature support the use of PGx to improve treatment by tailoring therapeutic strategies based on individual genetic profiles, thereby minimizing toxicity, optimizing drug efficacy ( 57 ), and reducing healthcare costs. This targeted approach helped reduce the likelihood of Type I errors often encountered in genome-wide testing ( 58 , 59 ). Moreover, by applying a more lenient significance threshold within this biologically relevant region, we aimed to limit Type II error and retain informative variants for ML modelling. This strategy aligns with previous efforts that utilized SNP-based platforms to predict chemotherapy-induced ADRs, demonstrating efficacy in reducing treatment-related toxicity ( 29 , 60 ). Our study shows that females have a higher propensity for ADRs. This is in agreement with previous studies, which report a 1.5 to 1.7-fold increased risk of ADRs in females, likely owing to pharmacokinetic and pharmacodynamic differences and hormonal and immunological influences ( 61 , 62 ). Additionally, comorbid conditions, including chronic kidney and liver diseases were significantly more prevalent among ADR cases in the present study, highlighting the potential role of pre-existing organ dysfunction in altering drug metabolism and elimination ( 63 , 64 ). This was further supported by elevated levels of cystatin C and alkaline phosphatase in the case group, both indicative of impaired renal and hepatic function ( 65 – 67 ). Haematological and metabolic differences also emerged as biological correlates of ADR susceptibility. Cases exhibited higher white blood cell and lymphocyte counts, and lower haemoglobin and direct bilirubin levels, suggesting systemic inflammation and possible bone marrow suppression, consistent with hematologic toxicity commonly associated with antineoplastic agents ( 68 ). Additionally, patients with malignancies affecting the hematopoietic and digestive systems exhibited a higher risk of ADRs. This is in line with prior studies reporting frequent ADRs in hematologic cancers, with blood disorders and infections being the most common complications ( 68 , 69 ). Moreover, ADR incidence was disproportionately high among patients with ill-defined, secondary, or unspecified cancer types, possibly reflecting diagnostic uncertainty or advanced disease stages, both of which could contribute to higher treatment burden and ADR risk ( 19 , 70 ). Lifestyle-related factors also appeared to influence ADR susceptibility, albeit to a lesser extent. Less time spent outdoors was more common among cases, potentially reflecting poor overall health or lower activity levels. Major dietary changes and smoking status were also observed as contributing factors to ADR risk. These findings are supported by another study demonstrating that modifiable lifestyle factors, including smoking, physical activity, and diet, are significant risk contributors to adverse health outcomes and this may indirectly increase vulnerability to ADRs ( 71 ). Furthermore, the identification of a SNP in the DPYD gene, which influences fluoropyrimidine toxicity ( 51 ), reinforces the genetic basis of ADR susceptibility and highlights the importance of genetic polymorphisms in drug metabolism. Additionally, variants in genes ATIC, ADORA2A , and MTRR were associated with methotrexate toxicity, aligning with previous evidence from PharmGKB and supporting their potential relevance in clinical decision-making ( 52 , 53 ). Other notable genes identified included CBR3 and ERCC1 , linked to doxorubicin and cisplatin-induced toxicity, and CYP19A1 , associated with adverse responses to tamoxifen and capecitabine. Moreover, polymorphisms in CHRNA3, CHRNA5 , and CYP2B6 , which are involved in drug metabolism and nicotine dependence pathways ( 72 , 73 ), were consistently identified across more than one ADR cohort studied, suggesting a broader relevance in modulating treatment outcomes with antineoplastic agents. These findings not only validate known pharmacogenomic associations but also reinforce the utility of incorporating germline genomic profiles into predictive models for ADRs. Collectively, these insights underscore the importance of considering a holistic view encompassing genetic, demographic, clinical, lifestyle and comorbid factors to enhance the efficacy of personalized medicine and reduce the burden of treatment-related toxicity ( 74 , 75 ). One of the key strengths of this study is the use of a large, well-characterized cohort from the UK Biobank, enabling robust model training, validation, and subgroup analyses. Consistent with prior literature on ML with genetic data ( 48 , 50 , 76 ), deep learning model (MLP) emerged as the top performer in our analysis, particularly when leveraging multimodal features. In our drug-specific prediction tasks, MLP model achieved AUC-ROC values up to 0.86, and in ADR-specific context - predicting secondary thrombocytopenia, performance peaked at 0.97, comparable to or exceeding results reported in earlier studies (e.g., MLP achieving 0.94 for complex disease prediction using SNPs ( 48 ); XGBoost achieving 0.972 in predicting adverse drug events using EHRs ( 49 )). Interestingly, LR - a traditional model, outperformed more complex tree-based methods like RF and XGBoost in our analyses, particularly in the complete cohort where heterogeneity is higher. This aligns with observations from Ho et al. ( 77 ), who noted that simpler models can remain competitive when paired with appropriate feature selection and regularization strategies, especially under high-dimensional but structured data. Performance was comparatively lower in the complete cohort analysis, where greater phenotypic and treatment heterogeneity likely contributed to reduced classification performance. Nevertheless, LR remained the best-performing model with AUC-ROC of 0.78. Importantly, both LR and MLP maintained a more favourable balance between sensitivity and specificity, resulting in improved F1 scores and overall classification reliability - an essential consideration in clinical settings. In contrast, SVM and RF consistently exhibited suboptimal performance across datasets. The geometric constraints of SVM may hinder its effectiveness in high-dimensional genomic data where complex, non-linear patterns are present ( 78 ), while RF’s reliance on hierarchical splits may fail to capture subtle feature interactions relevant to ADR risk ( 79 ). The superior performance of MLP is consistent with its capacity to model non-linear relationships and manage high-dimensional data, aligning with previous findings supporting the utility of neural networks in clinical genomics ( 80 ). Further improvements in predictive performance could be achieved by increasing sample sizes, particularly for rare ADR events in specific cancer cohorts, and by integrating more detailed longitudinal treatment histories, complementing the robust performance already demonstrated by LR and MLP across analyses. Given their ability to handle multimodal data and generate reliable predictions, ML models for ADR prediction can be meaningfully integrated into clinical workflows by embedding them within EHR systems. This integration can be enabled through a secure application programming interface (API) that serves as a bridge between the ADR prediction model and the EHR. On one side, the API would pull relevant patient data - such as demographics, lifestyle, clinical characteristics, comorbid conditions, medication history, and genetic information - from the EHR. On the other, it would feed this data into the trained ADR model to compute risk scores in real time. These risk estimates could then be delivered back into the EHR interface, flagging high-risk patients directly within the clinician’s workflow. While this study offers valuable insights, there were several limitations. The analysis was limited to individuals of white European ancestry to minimize population stratification, which may restrict the generalizability of findings to other ethnic groups. Expanding future research to more diverse populations will be essential for equitable clinical application. Additionally, ADR identification was based on ICD-10 coding, which may miss subtle or misclassified cases owing to inconsistencies in clinical reporting ( 81 ). Emerging methods such as natural language processing (NLP), could help capture richer ADR-related information from electronic health records, though challenges such as lexical variability and data imbalance remain ( 82 ). Furthermore, polypharmacy - a common scenario in oncology - introduces significant complexity in isolating the causal factors of ADRs, as interactions between multiple drugs can confound the attribution of specific effects to individual agents. Addressing this challenge requires more granular data and advanced modelling approaches. Notably, the current models are designed as a broad, multiple-drug risk assessment, offering a valuable starting point for initial screening in oncology settings. While not individual drug-specific, our approach offers a wider understanding of factors contributing to ADR susceptibility across antineoplastic treatments to identify high-risk individuals who may benefit from closer monitoring ( 83 ). The superior predictive performance of our models not only advances the understanding of ADR mechanisms but also highlights their potential for real-world clinical application. Building upon previous efforts (e.g. Naranjo Adverse Drug Reaction Probability Scale and the ADR Risk Estimator tool ( 84 , 85 )), our study advances the field by integrating a broader and richer set of features encompassing pharmacogenomic variants, environmental characteristics, and comorbid conditions to train ML models specifically tailored to oncology population. This holistic approach has the potential to enhance clinical workflows by enabling personalized ADR risk stratification prior to initiating antineoplastic treatments, ultimately supporting safer prescribing, reducing treatment-related complications, and improving patient outcomes. Translating these predictive tools into routine clinical care will require prospective validation in independent cohorts. Embedding such models within clinical decision-support systems will be critical for realizing their full potential in guiding personalized cancer therapy. This study addresses the urgent need to tailor oncological treatments to minimize ADRs, leveraging advances in next-generation sequencing technology, curated datasets, PGx, and ML. By integrating genetic, environmental, and comorbidities data with treatment efficacy and ADR profiles, we present a novel framework for early risk stratification and personalized intervention. This approach has the potential to reduce ADR-related complications, improve treatment adherence, and optimize clinical outcomes while lowering healthcare costs. As oncology moves toward more holistic, data-driven care models, our work - one of the first applications of multimodal ML-driven ADR prediction in the field, offers a roadmap for advancing personalized oncology care to improve outcomes for people with cancer. Methods Study population and dataset This retrospective case-control study utilized de-identified data from the biomedical database UK Biobank (Application ID: 86460) to investigate the genomic, environmental (demographic, lifestyle, and clinical) features, and comorbid conditions associated with ADRs in cancer patients being treated with antineoplastic agents. The case group (N = 5577) comprised individuals diagnosed with malignant neoplasms who reported ADRs following the correct administration of antineoplastic agents at therapeutic or prophylactic doses. ADRs were identified using the International Classification of Diseases, Tenth Revision (ICD-10) ( 86 ) codes Y43.0 (antiallergic and antiemetic drugs), Y43.1 (antineoplastic antimetabolites), and Y43.3 (antineoplastic drugs). Cases of drug poisoning or overdose were excluded. The control group (N = 22,308) consisted of individuals diagnosed with malignant neoplasms (ICD-10 codes C00-C97) who did not report any ADRs. Cases and controls were matched at a 1:4 ratio, resulting in a total cohort of 27,885 participants. To reduce population stratification and enhance comparability, analyses were restricted to individuals who self-identified as being of white ethnic background. Ethnicity was determined using UK Biobank self-reported data (field 21000; codes 1 [White], 1001 [British], 1002 [Irish], or 1003 [any other White background]). Participants were further restricted to those aged 40–69 years. Data Preprocessing and Feature Selection Genetic Data A total of 755 participants were excluded from the initial total cohort owing to the unavailability of whole-genome sequencing (WGS) data. Additionally, directly applying ML techniques is currently computationally challenging owing to the large number of features (i.e., SNPs) relative to the number of samples. This imbalance has been shown to lead to cause overfitting, resulting in high training accuracy but poor generalization to new datasets ( 87 , 88 ). To address this, we first reduced genomic dimensionality by focusing on pharmacogenetic regions. A list of 169 pharmacogenes (Supplementary Table 1) was compiled from the Pharmacogenomics Knowledgebase (PharmGKB) ( 52 , 53 ) and the Clinical Pharmacogenetics Implementation Consortium (CPIC®) ( 89 ), and each gene was extended by ± 20 kilobases (kb) to capture potential regulatory variants. Only SNPs within these pharmacogenetic regions were retained for further analysis. Variant-level quality control (QC) was performed using PLINK v2.0 ( 90 ) and excluded SNPs with a minor allele frequency (MAF) ≤ 0.001, Hardy-Weinberg equilibrium p ≤ 1×10⁻⁸, and missing genotype rate > 2%. Additionally, individuals with > 5% missing genotypes were excluded. After QC, the final dataset consisted of 26,235 individuals (5194 cases and 21,041 controls). To further refine the genomic feature set for ML, we applied two complementary strategies for feature selection. First, we conducted GWAS restricted to pharmacogenetic regions (169 pharmacogenes extended by ± 20 kb), using mixed linear models adjusting for age, sex, and population stratification via the top 10 genotype principal components. By narrowing the analysis to pharmacogenes, we reduced the likelihood of Type I errors (i.e., false positives) commonly associated with conventional genome-wide testing( 58 , 59 ). To further reduce the risk of Type II error (i.e., false negatives), we applied a relaxed significance threshold (p < 5×10⁻²), retaining variants with potential biological relevance for downstream ML modelling. In the second strategy, we incorporated 373 SNPs from filtered pharmacogenetic regions previously associated with specific antineoplastic agents, curated from the PharmGKB database (Supplementary Table 2), to ensure the inclusion of variants supported by prior pharmacogenetic evidence. The resulting SNP set comprising variants identified through GWAS and curated knowledge, was converted to allele dosage format and used as the genomic feature set for model development. Clinical data The selection of environmental (demographic, lifestyle, clinical features) and comorbidity-related features was guided by their relevance to ADRs, existing literature, and potential influence in the univariate analyses using Welch’s t-test for continuous variables and Chi-square or Fisher’s exact test for categorical variables, as appropriate. Demographic factors included age and gender; clinical measures included blood biomarkers, and lifestyle factors (e.g., exercise, sleep, and dietary habits); and specific habits known to affect health and drug metabolism, including alcohol consumption and smoking, were also included (Supplementary table 4). A wide range of comorbid conditions including chronic kidney disease, liver disease, pneumonia, and others were identified using ICD-10 codes, allowing for a detailed representation of participants’ health conditions. Missing data were imputed using Multiple Imputation by Chained Equations (MICE) ( 91 ). ML model development We developed ADR risk prediction models using ML algorithms, based on a tiered analysis comprising three levels of participant cohorts consisting of cases and controls. Drug-Specific Cohort : Focused on participants with a history of taking one or more of four commonly prescribed cancer medications: fluorouracil, capecitabine, methotrexate, and tamoxifen - selected owing to their frequent clinical use and well established pharmacogenomic relevance ( 92 – 94 ). To improve predictive accuracy, this analysis incorporated genetic variants known to interact with these drugs (as catalogued in PharmGKB), along with variants from key pharmacogenes identified through univariate GWAS filtering (N = 1098; 246 cases, 852 controls). ADR-Specific Cohort : Four separate ADR-specific cohorts were examined, each defined by a distinct adverse drug reaction (ADR) with a minimum case count of N > 80. This threshold was chosen to ensure adequate statistical power, while excluding ADR categories with smaller case numbers that may yield unstable or unreliable estimates. The included ADRs were: i) drug-induced gastroenteritis and colitis (ICD-10: K52.1, N = 3129), ii) drug-induced polyneuropathy (G62.0, N = 1436), iii) secondary thrombocytopenia (D69.5, N = 419), iv) generalized or localized skin eruptions (L27.0 & L27.1, N = 1277). Complete Cohort : Included all participants who experienced any ADRs associated with antineoplastic drugs (N = 26,235; 5194 cases and 21,041 controls). Our approach to model development was incremental, assessing the impact of various feature sets on ADR prediction. We began with genetic features, specifically SNPs identified from 169 curated pharmacogenes. Next, we incorporated environmental variables, and finally, we added comorbid conditions to construct a comprehensive predictive model. We trained predictive models using five supervised machine learning algorithms: LR ( 95 ), RF ( 96 ), SVM ( 97 ), XGBoost ( 98 ), and MLP ( 99 ). Each model’s predictive performance was systematically assessed in three stages: first using only genetic data, then adding environmental features, and finally incorporating comorbid conditions to evaluate the incremental value of each feature set for ADR prediction. To further assess the independent contribution of environmental features, we also trained a logistic regression model using only environmental variables, excluding genetic and comorbidity data. SHapley Additive exPlanations (SHAP) values were computed to interpret the relative importance of individual features in ADR prediction. Models were trained using a split dataset, with 70% allocated to training and 30% to testing. Additionally, the distribution of patients who experienced ADRs (case group) in our cohort was markedly imbalanced relative to controls. To mitigate potential bias and improve classifier performance, we employed the Synthetic Minority Oversampling Technique (SMOTE), a widely validated oversampling approach ( 100 ). SMOTE was applied to the training set to generate synthetic instances of the minority class (cases) to achieve a more balanced distribution, thereby enhancing model sensitivity, specificity and predictive accuracy ( 101 ). The training process involved 5-fold cross-validation using GridSearchCV to optimize hyperparameters and ensure robust model evaluation. Within each fold, the model was trained on four subsets of the training data and validated on the fifth, iterating through all partitions. The final model, selected based on cross-validation performance, was then evaluated on the hold-out test set to assess generalizability. Model effectiveness and optimal hyperparameters in the testing set were assessed using the AUC-ROC. Additional metrics (sensitivity, specificity, accuracy, precision, and F1 score) were used to provide a comprehensive evaluation of classification performance in the testing set. Based on the evaluation metrics, the optimal model was identified primarily by its AUC-ROC performance, while also prioritizing a high F1 score, sensitivity and other metrics. The model development workflow is depicted in Fig. 1 , and detailed information of hyperparameters used for the models are presented in Supplementary Table 3. Declarations Ethics approval and consent to participate This study was conducted using data from the UK Biobank under approved application number 86460. All participants provided informed consent at recruitment. No additional ethical approval was required for this secondary data analysis. DATA AVAILABILITY The data used in this study were obtained from the UK Biobank. Access to the UK Biobank dataset is not publicly available but can be obtained to researchers through an application process. Funding This work was supported by the National Health and Medical Research Council (NHMRC) Ideas Grant [GNT2029756]. CODE AVAILABILITY Quantitative analysis and modelling were carried out using PLINK v2.0 and Python v3.11.3, with Bash used to automate data preprocessing and pipeline execution. Analyses in python made use of the following packages: NumPy, Pandas, Sklearn, Xgboost, Imblearn, Matplotlib, SciPy, Seaborn, Shap. Code will be made available upon reasonable request. ACKNOWLEDGEMENTS This research has been conducted using the UK Biobank resource under Application Number 86460. We thank the participants of the UK Biobank study, without whom this research would not have been possible. The authors would like to acknowledge Dr. Prathosh AP (IISc, Bengaluru) for his valuable guidance on the machine learning components of this work. We also thank Dr. Senthil Lingarathnam, Director of Pharmacy at Peter MacCallum Cancer Centre, for his insightful feedback and contributions to the clinical framing of cancer-related ADRs and antineoplastic drug selection. We are grateful to Kim N. Tran, Rohit Mishra, and Abhiram D. B. for their thoughtful feedback and manuscript revisions. AUTHOR CONTRIBUTIONS SN led the project. SN and RL oversaw the analysis. AJ and VA designed the research. AJ performed the analysis. VA and BH assisted in data curation. SN and AJ prepared the first draft of the manuscript. All authors contributed to manuscript revision, read and approved the submitted version. COMPETING INTERESTS The authors declare that they have no competing interests References Edwards IR, Aronson JK. Adverse drug reactions: definitions, diagnosis, and management. Lancet. 2000;356(9237):1255-9. Zhou ZW, Chen XW, Sneed KB, Yang YX, Zhang X, He ZX, et al. Clinical association between pharmacogenomics and adverse drug reactions. Drugs. 2015;75(6):589-631. Pirmohamed M, James S, Meakin S, Green C, Scott AK, Walley TJ, et al. Adverse drug reactions as cause of admission to hospital: prospective analysis of 18,820 patients. Bmj-Brit Med J. 2004;329(7456):15-9. Thuermann PA, Windecker R, Steffen J, Schaefer M, Tenter U, Reese E, et al. Detection of adverse drug reactions in a neurological department: comparison between intensified surveillance and a computer-assisted approach. Drug Saf. 2002;25(10):713-24. Bouvy JC, De Bruin ML, Koopmanschap MA. Epidemiology of adverse drug reactions in Europe: a review of recent observational studies. Drug Saf. 2015;38(5):437-53. Alexopoulou A, Dourakis SP, Mantzoukis D, Pitsariotis T, Kandyli A, Deutsch M, et al. Adverse drug reactions as a cause of hospital admissions: A 6-month experience in a single center in Greece. Eur J Intern Med. 2008;19(7):505-10. Angamo MT, Chalmers L, Curtain CM, Bereznicki LRE. Adverse-Drug-Reaction-Related Hospitalisations in Developed and Developing Countries: A Review of Prevalence and Contributing Factors. Drug Safety. 2016;39(9):847-57. Silva LT, Modesto ACF, Amaral RG, Lopes FM. Hospitalizations and deaths related to adverse drug events worldwide: Systematic review of studies with national coverage. European Journal of Clinical Pharmacology. 2022;78(3):435-66. de Bienassis K, Esmail L, Lopert R, Klazinga N. The economics of medication safety: Improving medication safety through collective, real-time learning. OECD Health Working Papers. 2022(147):0_1-85. Wester K, Jonsson AK, Spigset O, Druid H, Hagg S. Incidence of fatal adverse drug reactions: a population based study. British journal of clinical pharmacology. 2008;65(4):573-9. U.S. Food and Drug Administration. U.S. Food and Drug Administration 2022 [Available from: https://www.fda.gov. Empey PE. Genetic predisposition to adverse drug reactions in the intensive care unit. Crit Care Med. 2010;38(6 Suppl):S106-16. Sendekie AK, Netere AK, Tesfaye S, Dagnew EM, Belachew EA. Incidence and patterns of adverse drug reactions among adult patients hospitalized in the University of Gondar comprehensive specialized hospital: A prospective observational follow-up study. Plos One. 2023;18(2). Wahlang JB, Laishram PD, Brahma DK, Sarkar C, Lahon J, Nongkynrih BS. Adverse drug reactions due to cancer chemotherapy in a tertiary care teaching hospital. Therapeutic Advances in Drug Safety. 2017;8(2):61-6. Kuderer NM, Desai A, Lustberg MB, Lyman GH. Mitigating acute chemotherapy-associated adverse events in patients with cancer. Nature Reviews Clinical Oncology. 2022;19(11):681-97. Tajani BB, Maheswari E, Maka VV, Nair AS. Adverse drug reactions and drug-related problems with supportive care medications among the oncological population. Discov Oncol. 2024;15(1). Chopra D, Rehan HS, Sharma V, Mishra R. Chemotherapy-induced adverse drug reactions in oncology patients: A prospective observational survey. Indian J Med Paediatr Oncol. 2016;37(1):42-6. Du RF, Wang X, Ma LX, Larcher LM, Tang H, Zhou HY, et al. Adverse reactions of targeted therapy in cancer patients: a retrospective study of hospital medical data in China. Bmc Cancer. 2021;21(1). Lavan AH, O'Mahony D, Buckley M, O'Mahony D, Gallagher P. Adverse Drug Reactions in an Oncological Population: Prevalence, Predictability, and Preventability. Oncologist. 2019;24(9):e968-e77. Davies MR, Martinec M, Walls R, Schwarz R, Mirams GR, Wang K, et al. Use of Patient Health Records to Quantify Drug-Related Pro-arrhythmic Risk. Cell Reports Medicine. 2020;1(5):100076. Naranjo CA, Busto U, Sellers EM, Sandor P, Ruiz I, Roberts EA, et al. A method for estimating the probability of adverse drug reactions. Clinical Pharmacology & Therapeutics. 1981;30(2):239-45. Pirmohamed M, Park BK. Genetic susceptibility to adverse drug reactions. Trends in Pharmacological Sciences. 2001;22(6):298-305. Meyer UA. Pharmacogenetics and adverse drug reactions. The Lancet. 2000;356(9242):1667-71. Swen JJ, Van Der Wouden CH, Manson LE, Abdullah-Koolmees H, Blagec K, Blagus T, et al. A 12-gene pharmacogenetic panel to prevent adverse drug reactions: an open-label, multicentre, controlled, cluster-randomised crossover implementation study. The Lancet. 2023;401(10374):347-56. Cacabelos R, Cacabelos N, Carril JC. The role of pharmacogenomics in adverse drug reactions. Expert Rev Clin Pharmacol. 2019;12(5):407-42. Weinshilboum RM, Wang LW. Pharmacogenomics: Precision Medicine and Drug Response. Mayo Clin Proc. 2017;92(11):1711-22. Sadee W, Wang D, Hartmann K, Toland AE. Pharmacogenomics: Driving Personalized Medicine. Pharmacological Reviews. 2023;75(4):789-814. Palmirotta R, Carella C, Silvestris E, Cives M, Stucci SL, Tucci M, et al. SNPs in predicting clinical efficacy and toxicity of chemotherapy: walking through the quicksand. Oncotarget. 2018;9(38):25355-82. Cafiero C, Palmirotta R, Martinelli C, Micera A, Giaco L, Persiani F, et al. Oncological Treatment Adverse Reaction Prediction: Development and Initial Validation of a Pharmacogenetic Model in Non-Small-Cell Lung Cancer Patients. Genes (Basel). 2025;16(3). Gong L, Whirl-Carrillo M, Klein TE. PharmGKB, an Integrated Resource of Pharmacogenomic Knowledge. Curr Protoc. 2021;1(8):e226. Krzyszczyk P, Acevedo A, Davidoff EJ, Timmins LM, Marrero-Berrios I, Patel M, et al. The growing role of precision and personalized medicine for cancer treatment. Technology (Singap World Sci). 2018;6(3-4):79-100. Pirmohamed M. Personalized Pharmacogenomics: Predicting Efficacy and Adverse Drug Reactions. Annual Review of Genomics and Human Genetics. 2014;15(Volume 15, 2014):349-70. Johnson JA. Ethnic differences in cardiovascular drug response: potential contribution of pharmacogenetics. Circulation. 2008;118(13):1383-93. Dang MT, Hambleton J, Kayser SR. The influence of ethnicity on warfarin dosage requirement. Ann Pharmacother. 2005;39(6):1008-12. Magavern EF, Megase M, Thompson J, Marengo G, Jacobsen J, Smedley D, et al. Pharmacogenetics and adverse drug reports: Insights from a United Kingdom national pharmacovigilance database. PLOS Medicine. 2025;22(3):e1004565. Tang C, Livingston MJ, Safirstein R, Dong Z. Cisplatin nephrotoxicity: new insights and therapeutic implications. Nat Rev Nephrol. 2023;19(1):53-72. Hanoodi M, Mittal M. Methotrexate. StatPearls. Treasure Island (FL)2024. Gaytan SL, Lawan A, Chang J, Nurunnabi M, Bajpeyi S, Boyle JB, et al. The beneficial role of exercise in preventing doxorubicin-induced cardiotoxicity. Front Physiol. 2023;14:1133423. Tamang R, Bharati L, Khatiwada AP, Ozaki A, Shrestha S. Pattern of Adverse Drug Reactions Associated with the Use of Anticancer Drugs in an Oncology-Based Hospital of Nepal. JMA J. 2022;5(4):416-26. Monaco A, Pantaleo E, Amoroso N, Lacalamita A, Lo Giudice C, Fonzino A, et al. A primer on machine learning techniques for genomic applications. Comput Struct Biotechnol J. 2021;19:4345-59. Syrowatka A, Song W, Amato MG, Foer D, Edrees H, Co Z, et al. Key use cases for artificial intelligence to reduce the frequency of adverse drug events: a scoping review. The Lancet Digital Health. 2022;4(2):e137-e48. Bowe AK, Lightbody G, Staines A, Murray DM. Big data, machine learning, and population health: predicting cognitive outcomes in childhood. Pediatr Res. 2023;93(2):300-7. Kirasich K, Smith T, Sadler B. Random forest vs logistic regression: binary classification for heterogeneous datasets. SMU Data Science Review. 2018;1(3):9. Lai C, Zimmer AD, O'Connor R, Kim S, Chan R, van den Akker J, et al. LEAP: Using machine learning to support variant classification in a clinical setting. Human Mutation. 2020;41(6):1079-90. Huang RJ, Kwon NS-E, Tomizawa Y, Choi AY, Hernandez-Boussard T, Hwang JH. A Comparison of Logistic Regression Against Machine Learning Algorithms for Gastric Cancer Risk Prediction Within Real-World Clinical Data Streams. JCO Clinical Cancer Informatics. 2022(6):e2200039. Long C, Lv GT, Fu XM. Development of a general logistic model for disease risk prediction using multiple SNPs. Febs Open Bio. 2019;9(11):2006-12. Zhang F, Sun B, Diao XL, Zhao W, Shu T. Prediction of adverse drug reactions based on knowledge graph embedding. Bmc Med Inform Decis. 2021;21(1). Alzoubi H, Alzubi R, Ramzan N. Deep Learning Framework for Complex Disease Risk Prediction Using Genomic Variations. Sensors (Basel). 2023;23(9). Kidwai-Khan F, Rentsch CT, Pulk R, Alcorn C, Brandt CA, Justice AC. Pharmacogenomics driven decision support prototype with machine learning: A framework for improving patient care. Front Big Data. 2022;5:1059088. Lai NH, Shen WC, Lee CN, Chang JC, Hsu MC, Kuo LN, et al. Comparison of the predictive outcomes for anti-tuberculosis drug-induced hepatotoxicity by different machine learning techniques. Comput Methods Programs Biomed. 2020;188:105307. Ruzzo A, Graziano F, Galli F, Galli F, Rulli E, Lonardi S, et al. Dihydropyrimidine dehydrogenase pharmacogenetics for predicting fluoropyrimidine-related toxicity in the randomised, phase III adjuvant TOSCA trial in high-risk colon cancer patients. Brit J Cancer. 2017;117(9):1269-77. Whirl-Carrillo M, McDonagh EM, Hebert JM, Gong L, Sangkuhl K, Thorn CF, et al. Pharmacogenomics Knowledge for Personalized Medicine. Clinical Pharmacology & Therapeutics. 2012;92(4):414-7. Whirl-Carrillo M, Huddart R, Gong L, Sangkuhl K, Thorn CF, Whaley R, et al. An Evidence-Based Framework for Evaluating Pharmacogenomics Knowledge for Personalized Medicine. Clinical Pharmacology & Therapeutics. 2021;110(3):563-72. Yazbeck V, Alesi E, Myers J, Hackney MH, Cuttino L, Gewirtz DA. An overview of chemotoxicity and radiation toxicity in cancer therapy. Adv Cancer Res. 2022;155:1-27. Kinnersley B, Sud A, Everall A, Cornish AJ, Chubb D, Culliford R, et al. Analysis of 10,478 cancer genomes identifies candidate driver genes and opportunities for precision oncology. Nat Genet. 2024;56(9):1868-77. Leong IUS, Cabrera CP, Cipriani V, Ross PJ, Turner RM, Stuckey A, et al. Large-Scale Pharmacogenomics Analysis of Patients With Cancer Within the 100,000 Genomes Project Combining Whole-Genome Sequencing and Medical Records to Inform Clinical Practice. J Clin Oncol. 2024:JCO2302761. Pharoah PDP. Genetic susceptibility, predicting risk and preventing cancer. Recent Results Canc. 2003;163:7-18. Uffelmann E, Posthuma D, Peyrot WJ. Genome-wide association studies of polygenic risk score-derived phenotypes may lead to inflated false positive rates. Sci Rep. 2023;13(1):4219. Marees AT, de Kluiver H, Stringer S, Vorspan F, Curis E, Marie-Claire C, et al. A tutorial on conducting genome-wide association studies: Quality control and statistical analysis. Int J Methods Psychiatr Res. 2018;27(2):e1608. Bjorn N, Badam TVS, Spalinskas R, Branden E, Koyi H, Lewensohn R, et al. Whole-genome sequencing and gene network modules predict gemcitabine/carboplatin-induced myelosuppression in non-small cell lung cancer patients. NPJ Syst Biol Appl. 2020;6(1):25. Rademaker M. Do women have more adverse drug reactions? Am J Clin Dermatol. 2001;2(6):349-51. Watson S, Caster O, Rochon PA, den Ruijter H. Reported adverse drug reactions in women and men: Aggregated evidence from globally collected individual case reports during half a century. EClinicalMedicine. 2019;17:100188. Wang X, Chen X. Clinical Characteristics of 162 Patients with Drug-Induced Liver and/or Kidney Injury. Biomed Res Int. 2020;2020:3930921. Spanakis M, Roubedaki M, Tzanakis I, Zografakis-Sfakianakis M, Patelarou E, Patelarou A. Impact of Adverse Drug Reactions in Patients with End Stage Renal Disease in Greece. Int J Environ Res Public Health. 2020;17(23). Alosco ML, Spitznagel MB, Strain G, Devlin M, Cohen R, Crosby RD, et al. The effects of cystatin C and alkaline phosphatase changes on cognitive function 12-months after bariatric surgery. J Neurol Sci. 2014;345(1-2):176-80. Chai X, Huang HB, Feng G, Cao YH, Cheng QS, Li SH, et al. Baseline Serum Cystatin C Is a Potential Predictor for Acute Kidney Injury in Patients with Acute Pancreatitis. Dis Markers. 2018;2018. Buyukberber M, Koruk I, Cykman O, Koruk M, Küçükoglu ME, Sakman A, et al. Serum cystatin C measurement in differential diagnosis of intra and extrahepatic cholestatic diseases. Ann Hepatol. 2010;9(1):58-62. Yu B, Yan XD, Zhu YY, Luo T, Sohail M, Ning H, et al. Analysis of adverse drug reactions/events of cancer chemotherapy and the potential mechanism of Danggui Buxue decoction against bone marrow suppression induced by chemotherapy. Frontiers in Pharmacology. 2023;14. Amaro-Hosey K, Danés I, Vendrell L, Alonso L, Renedo B, Gros L, et al. Adverse Reactions to Drugs of Special Interest in a Pediatric Oncohematology Service. Frontiers in Pharmacology. 2021;12. Belachew SA, Erku DA, Mekuria AB, Gebresillassie BM. Pattern of chemotherapy-related adverse effects among adult cancer patients treated at Gondar university referral hospital, Ethiopia: a cross-sectional study. Drug Healthc Patient. 2016;8:83-90. Jackowska B, Wisniewski P, Noinski T, Bandosz P. Effects of lifestyle-related risk factors on life expectancy: A comprehensive model for use in early prevention of premature mortality from noncommunicable diseases. Plos One. 2024;19(3). Chmielowiec K, Chmielowiec J, Stronska-Pluta A, Trybek G, Smiarowska M, Suchanecka A, et al. Association of Polymorphism CHRNA5 and CHRNA3 Gene in People Addicted to Nicotine. Int J Environ Res Public Health. 2022;19(17). Muderrisoglu A, Babaoglu E, Korkmaz ET, Kalkisim S, Karabulut E, Emri S, et al. Comparative Assessment of Outcomes in Drug Treatment for Smoking Cessation and Role of Genetic Polymorphisms of Human Nicotinic Acetylcholine Receptor Subunits. Frontiers in Genetics. 2022;13. Farnoush A, Sedighi-Maman Z, Rasoolian B, Heath JJ, Fallah B. Prediction of adverse drug reactions using demographic and non-clinical drug characteristics in FAERS data. Sci Rep. 2024;14(1):23636. Alomar MJ. Factors affecting the development of adverse drug reactions (Review article). Saudi Pharm J. 2014;22(2):83-94. Mohsen A. Deep Learning Prediction of Adverse Drug Reactions in Drug Discovery Using Open TG–GATEs and FAERS Databases. Ho DSW, Schierding W, Wake M, Saffery R, O'Sullivan J. Machine Learning SNP Based Prediction for Precision Medicine. Front Genet. 2019;10:267. Clarke R, Ressom HW, Wang A, Xuan J, Liu MC, Gehan EA, et al. The properties of high-dimensional data spaces: implications for exploring gene and protein expression data. Nat Rev Cancer. 2008;8(1):37-49. Ghosh D, Cabrera J. Enriched Random Forest for High Dimensional Genomic Data. IEEE/ACM Trans Comput Biol Bioinform. 2022;19(5):2817-28. Anastopoulos IN, Herczeg CK, Davis KN, Dixit AC. Multi-Drug Featurization and Deep Learning Improve Patient-Specific Predictions of Adverse Events. Int J Env Res Pub He. 2021;18(5). Hohl CM, Karpov A, Reddekopp L, Stausberg J. ICD-10 codes used to identify adverse drug events in administrative data: a systematic review. J Am Med Inform Assn. 2014;21(3):547-57. Wieland-Jorna Y, van Kooten D, Verheij RA, de Man Y, Francke AL, Oosterveld-Vlug MG. Natural language processing systems for extracting information from electronic health records about activities of daily living. A systematic review. Jamia Open. 2024;7(2). Dsouza VS, Leyens L, Kurian JR, Brand A, Brand H. Artificial intelligence (AI) in pharmacovigilance: A systematic review on predicting adverse drug reactions (ADR) in hospitalized patients. Res Soc Admin Pharm. 2025;21(6):453-62. García-Cortés M, Lucena MI, Pachkoria K, Borraz Y, Hidalgo R, Andrade RJ, et al. Evaluation of Naranjo Adverse Drug Reactions Probability Scale in causality assessment of drug-induced liver injury. Aliment Pharm Therap. 2008;27(9):780-9. Valeanu A, Damian C, Marineci CD, Negres S. The development of a scoring and ranking strategy for a patient-tailored adverse drug reaction prediction in polypharmacy. Sci Rep-Uk. 2020;10(1). John HH. International Statistical Classification of Diseases and Related Health-Problems - World-Hlth-Org. J Roy Soc Health. 1994;114(6):339-. Pudjihartono N, Fadason T, Kempa-Liehr AW, O'Sullivan JM. A Review of Feature Selection Methods for Machine Learning-Based Disease Risk Prediction. Front Bioinform. 2022;2:927312. Silva PP, Gaudillo JD, Vilela JA, Roxas-Villanueva RML, Tiangco BJ, Domingo MR, et al. A machine learning-based SNP-set analysis approach for identifying disease-associated susceptibility loci. Sci Rep. 2022;12(1):15817. Relling MV, Klein TE. CPIC: Clinical Pharmacogenetics Implementation Consortium of the Pharmacogenomics Research Network. Clinical Pharmacology & Therapeutics. 2011;89(3):464-7. Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MA, Bender D, et al. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet. 2007;81(3):559-75. van Buuren S, Groothuis-Oudshoorn K. mice: Multivariate Imputation by Chained Equations in R. J Stat Softw. 2011;45(3):1-67. Rofaiel S, Muo EN, Mousa SA. Pharmacogenetics in breast cancer: steps toward personalized medicine in breast cancer management. Pharmacogn Pers Med. 2010;3:129-43. Sánchez-Bayona R, Catalán C, Cobos MA, Bergamino M. Pharmacogenomics in Solid Tumors: A Comprehensive Review of Genetic Variability and Its Clinical Implications. Cancers. 2025;17(6). Franczyk B, Rysz J, Gluba-Brzózka A. Pharmacogenetics of Drugs Used in the Treatment of Cancers. Genes-Basel. 2022;13(2). Mccullagh P. Generalized Linear-Models. Eur J Oper Res. 1984;16(3):285-92. Breiman L. Random forests. Mach Learn. 2001;45(1):5-32. Cortes C, Vapnik V. Support-Vector Networks. Mach Learn. 1995;20(3):273-97. Chen TQ, Guestrin C. XGBoost: A Scalable Tree Boosting System. Kdd'16: Proceedings of the 22nd Acm Sigkdd International Conference on Knowledge Discovery and Data Mining. 2016:785-94. Attali JG, Pages G. Approximations of functions by a multilayer perceptron: a new approach. Neural Networks. 1997;10(6):1069-81. Barua S, Islam MM, Yao X, Murase K. MWMOTE-Majority Weighted Minority Oversampling Technique for Imbalanced Data Set Learning. Ieee T Knowl Data En. 2014;26(2):405-25. Dablain D, Krawczyk B, Chawla N. DeepSMOTE: Fusing Deep Learning and SMOTE for Imbalanced Data. Ieee T Neur Net Lear. 2023;34(9):6390-404. Additional Declarations No competing interests reported. Supplementary Files Anoopet.al.Supplementarytables.docx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7431071","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":514847665,"identity":"1cdab323-48db-4724-94ee-2f1171868eab","order_by":0,"name":"Anoop Joseph","email":"","orcid":"","institution":"Queensland University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Anoop","middleName":"","lastName":"Joseph","suffix":""},{"id":514847666,"identity":"09bdc403-38c9-4329-9382-47d2b96221f7","order_by":1,"name":"Vignesh Arunachalam","email":"","orcid":"","institution":"Queensland University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Vignesh","middleName":"","lastName":"Arunachalam","suffix":""},{"id":514847667,"identity":"0c482b26-f85f-4f85-b360-38fec57cf42f","order_by":2,"name":"Arabella Hart","email":"","orcid":"","institution":"Walter and Eliza Hall Institute of Medical Research (WEHI)","correspondingAuthor":false,"prefix":"","firstName":"Arabella","middleName":"","lastName":"Hart","suffix":""},{"id":514847668,"identity":"4ad6bf9b-0436-47ac-a936-33e5aca60d5b","order_by":3,"name":"Rodney Lea","email":"","orcid":"","institution":"Queensland University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Rodney","middleName":"","lastName":"Lea","suffix":""},{"id":514847669,"identity":"c062308b-675c-48cb-b463-400bc1744644","order_by":4,"name":"Shivashankar H Nagaraj","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABAElEQVRIiWNgGAWjYDACCcYGONvgA5QmXovhDCCfCC1IbGYeYrTIz25uk2DcYWfXP/vwg2KbssN1DOzN2yQYag7j1GJw5yBQy5nk5Bnn0gyMc84dlmDgOVYmwXAMjxaJxGYDxjbmZIYzDAbGuW1ALRI5ZhIMbLi1yM8Aa6lPlj/D/sHYEqRF/g1Qyz/cWhhuJDY+YGw7bGdwhsfAmBFsC4+ZBJCB22EgLYltxxMMz/AUGPacS5ds40krtkjsS8fjsPQHBz62VdvLnWHfZvCjzJqfn/3wxhsfvlnjdhgIJDAwJDYwMLAZMLAxgBBYhCCwB2LmBxD1o2AUjIJRMApQAQA6DU7yRTudqwAAAABJRU5ErkJggg==","orcid":"","institution":"Queensland University of Technology","correspondingAuthor":true,"prefix":"","firstName":"Shivashankar","middleName":"H","lastName":"Nagaraj","suffix":""}],"badges":[],"createdAt":"2025-08-22 05:53:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7431071/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7431071/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":91958136,"identity":"94634512-9ba8-46e8-a85f-b7dbf9b36c74","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":935463,"visible":true,"origin":"","legend":"","description":"","filename":"Anoopet.al.Manuscript.docx","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/0d42325923be3df17ccf4b60.docx"},{"id":91958135,"identity":"aa4d2498-a8a5-4a32-970d-d9c3269d4379","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":7155,"visible":true,"origin":"","legend":"","description":"","filename":"bf75c67b1c624887bca8dbe5299cee3c.json","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/5da3d115e8af97704cf7cdd7.json"},{"id":91959906,"identity":"9de5ec0b-a38e-40b6-b417-b9965f5cff6d","added_by":"auto","created_at":"2025-09-23 07:40:38","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":83433,"visible":true,"origin":"","legend":"","description":"","filename":"Anoopet.al.Supplementarytables.docx","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/d28b6f0593d19d715e007348.docx"},{"id":91958153,"identity":"d8a5025d-a614-4173-a0ce-9a26b1b01b5e","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"xml","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":204048,"visible":true,"origin":"","legend":"","description":"","filename":"bf75c67b1c624887bca8dbe5299cee3c1enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/5554d3ae50dbc46e8fcf971f.xml"},{"id":91958144,"identity":"1cb2c943-fe2f-4390-9184-3256dae3807e","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"png","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":78958,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/e2c8d6d266b073ccd071405f.png"},{"id":91958138,"identity":"a7279e49-2ff6-4059-87fb-f0fe8ef1faba","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"png","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":44277,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/541e803f022e3db40cfec8b7.png"},{"id":91959915,"identity":"abeb3901-1402-40c1-883c-d86b9658d60e","added_by":"auto","created_at":"2025-09-23 07:40:38","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":227054,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/f92a6c8943958fce90546fcf.png"},{"id":91958150,"identity":"2f5a10e4-864b-4b4e-8fc8-7d3ef33c677c","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":92192,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/8f797bb7424d7184230ab286.png"},{"id":91958152,"identity":"87e60f3a-a6e2-4f96-9875-db09f30d6ed2","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":257931,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/ed2286d12ead279bd81456e7.png"},{"id":91960791,"identity":"b4673fa9-48c5-4153-ba95-f524348caf54","added_by":"auto","created_at":"2025-09-23 07:48:38","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":26996,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/c9470d85308ce50d6a7b75b9.png"},{"id":91959935,"identity":"ab1a3c1f-3758-4d7d-9e1b-098f28c732ea","added_by":"auto","created_at":"2025-09-23 07:40:38","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":19208,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/94e00ebeb69db015055c2f39.png"},{"id":91958156,"identity":"4a36b401-1cd1-46d6-a65a-7a8ff9c21717","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":75507,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/b893ed8b20dbe558bd676a7a.png"},{"id":91958147,"identity":"6ebc05d5-555f-4797-be73-9b6433072ae0","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":27490,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/c7c557f8343edafd03c6e97d.png"},{"id":91958154,"identity":"9a2d00fd-3dd4-4b5d-a495-97e02c6419d5","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":77059,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/91f25b25062ca73005ae451e.png"},{"id":91959938,"identity":"1eb1591e-a2de-40de-93af-ea41832a15a7","added_by":"auto","created_at":"2025-09-23 07:40:38","extension":"xml","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":202079,"visible":true,"origin":"","legend":"","description":"","filename":"bf75c67b1c624887bca8dbe5299cee3c1structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/83611c8c434cfdc853e907ef.xml"},{"id":91958157,"identity":"b47fdb77-bb95-41da-ba95-041ebc7794cd","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"html","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":216080,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/339d7d9e8a7079e66d642f99.html"},{"id":91960790,"identity":"c3616a84-3c0d-4fea-b678-0d4e893fb728","added_by":"auto","created_at":"2025-09-23 07:48:38","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":393914,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eWorkflow for developing and applying machine learning (ML) models to predict phenotype risk. \u003c/strong\u003e(A) Overview of the model development pipeline using a training dataset with labelled genotype data. The process begins with data pre-processing, including quality control and feature selection, followed by hyperparameter tuning and model training. Model performance is evaluated via cross-validation and subsequently validated on an independent test set to ensure generalizability and robustness. The optimal model is selected based on key evaluation metrics, primarily AUC-ROC. (B) Application of the final trained and validated ML model to new patient genotype data for phenotype risk prediction. The model classifies individuals as low- or high-risk based on their genetic and other feature profiles, enabling phenotype prediction for clinical or research use.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/5938d105281c27578219f099.png"},{"id":91960792,"identity":"41b6043d-b10d-446f-a674-665f0f81a17e","added_by":"auto","created_at":"2025-09-23 07:48:38","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":178184,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePercentage of patients within each cancer type by adverse drug reaction (ADR) status. \u003c/strong\u003ePercentages are based on the number of patients in each group (N= 26,235; 5194 cases and 21,041 controls)\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/d4eb33a7d41299329ce1a2d3.png"},{"id":91958140,"identity":"519d8ee5-fa97-4c75-974a-a4fa8ecb8bb5","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":162018,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eROC curves and AUC values for ML models in the drug-specific cohort.\u003c/strong\u003e ROC curves are shown for models trained on A) genetic-only data; B) genetic and environmental data; and C) genetic, environmental, and comorbidity data. The corresponding AUC-ROC values are displayed within each panel, quantifying model discrimination performance. The plots illustrate the true positive rate against the false positive rate, with the dashed line representing the line of no discrimination. Model colour coding remains consistent across panels. ROC: Receiver operating characteristic, AUC: Area Under the Curve, ML: Machine Learning, LR: logistic regression, RF: Random Forest, SVM: Support Vector Machine, XGB: eXtreme Gradient Boosting (i.e., XGBoost), and MLP: Multilayer Perceptron.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/bfbbdd3efa00a46760a47533.png"},{"id":91958139,"identity":"f0f6f178-fd00-4d77-ac2a-a03fb9ec73ca","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":626268,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eROC curves and AUC values for ML models across the four ADR specific cohorts.\u003c/strong\u003eColumns represent different feature sets: 1) genetic-only data; 2) genetic and environmental data; and 3) genetic, environmental, and comorbidity data. The corresponding AUC-ROC values are displayed within each panel, quantifying model discrimination performance.\u003cem\u003e \u003c/em\u003eRows correspond to ADR type. Each predictive model was trained separately for each feature set, with ROC curve colors corresponding to the legend for each ML model. The plots illustrate the true positive rate against the false positive rate, with the dashed line representing the line of no discrimination. Model colour coding remains consistent across panels. ROC: Receiver operating characteristic, AUC: Area Under the Curve, ML: Machine Learning, ADR: Adverse drug reaction, LR: logistic regression, RF: Random Forest, SVM: Support Vector Machine, XGB: eXtreme Gradient Boosting (i.e., XGBoost), and MLP: Multilayer Perceptron.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/2dbff29893b37b137137da8e.png"},{"id":91959899,"identity":"e10d3db2-f372-47a4-a94c-8032fb761158","added_by":"auto","created_at":"2025-09-23 07:40:38","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":133128,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSHAP summary plot for the LR model using environmental features in the complete cohort.\u003c/strong\u003e The plot illustrates the relative impact of each feature on the model’s prediction. Each point represents an individual subject, with the x-axis showing the SHAP value, indicating the contribution of the feature to the model output. Positive SHAP values indicate an increase in the predicted risk, while negative values indicate a decrease. The colour gradient reflects the original feature value, with red indicating high values and blue indicating low values. LR: Logistic Regression.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/d743c301e1c02f8521e8b1bf.png"},{"id":94474065,"identity":"ee2a2abb-909a-4a10-bac2-565d120e1f71","added_by":"auto","created_at":"2025-10-27 15:47:09","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2688818,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/a567c0fa-cd82-422e-b4ab-871aed70d9ff.pdf"},{"id":91958142,"identity":"046c07f4-1d35-478c-8f06-e083b80c3c3d","added_by":"auto","created_at":"2025-09-23 07:32:38","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":83433,"visible":true,"origin":"","legend":"","description":"","filename":"Anoopet.al.Supplementarytables.docx","url":"https://assets-eu.researchsquare.com/files/rs-7431071/v1/b262b249a53c7df7343c57b6.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Pharmacogenomics-Driven Multimodal Data Integration Improves Predictions of Adverse Drug Reactions in Cancer Patients using Machine Learning","fulltext":[{"header":"Introduction","content":"\u003cp\u003eAdverse drug reactions (ADRs), defined by the World Health Organization (WHO) as \u0026ldquo;any response to a drug which is noxious and unintended, and which occurs at doses normally used in humans for prophylaxis, diagnosis, or therapy of disease, or for the modification of physiological function\u0026rdquo; (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e), account for nearly 13% of all hospital admissions worldwide, of which over 1% are fatal (\u003cspan additionalcitationids=\"CR3 CR4 CR5 CR6 CR7 CR8\" citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e). ADRs rank as the seventh leading cause of death in Sweden (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e) and fourth in the United States (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e), and result in increased costs for healthcare systems (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e). Notably, ADRs are extremely prevalent in individuals being treated for cancer (\u003cspan additionalcitationids=\"CR15\" citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e). During cancer treatment, complex medication regimens and the concurrent use of multiple drugs heighten the risk of adverse effects, with over 40% of individuals experiencing three or more ADRs, 88% of which are preventable (\u003cspan additionalcitationids=\"CR18\" citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e). While adverse event reporting systems and predictive modelling have proven to be helpful in predicting relative drug risk (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e), a large body of evidence has demonstrated that individuals are genetically predisposed to ADRs, highlighting the need to better understand the role of genetics in ADRs at both the population and individual levels (\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan additionalcitationids=\"CR23\" citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e).\u003c/p\u003e\u003cp\u003ePharmacogenomics (PGx), the study of the genetic basis of drug responses, focuses on the link between genetic variations and adverse drug reactions to inform individualized drug selection, maximize efficacy, and increase safety (\u003cspan additionalcitationids=\"CR26\" citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e). PGx accounts for approximately 80% of all variability in drug efficacy and safety, where approximately half of all clinically relevant genes are associated with ADRs (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e). Genomic variants have been shown to significantly affect patient responses to cancer treatments (\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e), and most individuals carry actionable PGx alleles pivotal in drug-gene interactions (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e). However, uncovering the genetic basis of drug response remains challenging owing to its complex, multimodal nature, which is shaped not only by genetic variation but also by environmental influences, gene-drug and gene-gene interactions, comorbidities, age, and ethnic and population-level diversity (\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan additionalcitationids=\"CR33 CR34\" citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e). This is especially crucial in cancer, where medication-induced adverse reactions can affect various organ systems, with common effects including fatigue, anorexia, alopecia, constipation, nausea, vomiting, and neuropathy (\u003cspan additionalcitationids=\"CR37 CR38\" citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e). Yet, current predictive approaches are limited by their reliance on single- or few-gene analyses and failure to integrate broader clinical context. Furthermore, models often overlook the complex, multimodal nature of ADRs, which are influenced by a combination of genetic, environmental (demographic, lifestyle, and clinical), and comorbid factors.\u003c/p\u003e\u003cp\u003eAdvanced machine learning (ML) techniques are increasingly being applied to large-scale genomic and environmental datasets to address the complexity of predicting ADRs (\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e, \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e). ML algorithms excel in managing large datasets and identifying complex, non-linear relationships that traditional statistical models may not adequately explain (\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e). Despite the rise of complex models, logistic regression (LR), a widely used statistical and ML algorithm for binary classification tasks, offers interpretable insights into the interplay of genetic and clinical factors (\u003cspan additionalcitationids=\"CR44\" citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e). LR frameworks that integrate single-nucleotide polymorphisms (SNPs) from genome-wide association studies (GWAS) have enabled the stratification of lung cancer risk in Chinese populations (\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e), while another study using LR achieved an area under the receiver operating characteristic curve (AUC-ROC) of 0.87 for ADR prediction (\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e). Sophisticated deep learning neural networks such as the multilayer perceptron (MLP) model have improved predictions of disease susceptibility using SNP data with advanced feature selection techniques, achieving an AUC-ROC of 0.94 (\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e), while integrating genetic data with electronic health records is highly effective at predicting preventable adverse drug events using the Extreme Gradient Boosting (XGBoost) model achieving an AUC of 0.972 (\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e). Artificial neural networks (ANN), support vector machines (SVM), and random forests (RF) have been compared in their ability to predict hepatotoxicity induced by anti-tuberculosis drugs, with the ANN model integrating both clinical and genomic information providing the best outcomes (\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eTo address prevailing shortcomings in predicting ADRs among cancer patients, the present study systematically investigated the potential of multimodal data integration - including PGx variants, environmental exposures, and comorbid characteristics - to enhance individual-level ADR prediction using machine learning models. We leveraged 169 curated pharmacogenes to extract genetic features relevant to drug response. Our approach combined advanced informatics with predictive modeling to develop and validate algorithms designed to facilitate informed pharmacogenomic decision-making in near real-time at the point of care. Employing a comprehensive ML framework, we compared five models: logistic regression, random forest, support vector machine, XGBoost, and multilayer perceptron. To our knowledge, this is the first study to systematically integrate whole-genome pharmacogenomic variants with environmental (demographic, lifestyle, and clinical) and comorbidity data (derived from ICD-10 codes) for ADR prediction, offering a more holistic and personalized approach to cancer treatment.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003eUnivariate analyses reveal clinical, demographic, and biological contributors to ADRs\u003c/h2\u003e\u003cp\u003eIn the complete cohort, univariate analyses identified several variables significantly associated with ADRs. Gender differences were observed, with females more likely to experience ADRs than males (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001). Lifestyle-related factors were also more prevalent among cases, including current smoking (p\u0026thinsp;=\u0026thinsp;0.013), major dietary changes (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001), and frequent dietary variation (p\u0026thinsp;=\u0026thinsp;0.046). Additionally, individuals in the case group reported spending significantly less time outdoors (p\u0026thinsp;=\u0026thinsp;0.017) (Supplementary Table\u0026nbsp;4).\u003c/p\u003e\u003cp\u003eNotably, there was no significant variation in the number of prescribed medications between cases and controls (p\u0026thinsp;=\u0026thinsp;0.371), indicating that medication load alone does not account for ADR occurrence. However, several haematological and biochemical parameters differed significantly between groups. Elevated white blood cell counts (p\u0026thinsp;=\u0026thinsp;0.009) and lymphocyte levels (p\u0026thinsp;=\u0026thinsp;0.001), in addition to lower haemoglobin (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) and direct bilirubin concentrations (p\u0026thinsp;=\u0026thinsp;0.001), were significantly associated with ADRs, highlighting potential biological correlates of ADR vulnerability.\u003c/p\u003e\u003cp\u003eFurther analyses revealed that ADR incidence varied significantly by cancer type. Notably, 63% of patients in the case group were diagnosed with ill-defined, secondary, or unspecified cancer sites, compared to only 23% in the control group - a 2.7 fold difference between cases and controls (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e2\u003c/span\u003e). In addition, lymphoid and hematopoietic malignancies, as well as cancers of the digestive system, were more frequently observed in the case group, underscoring the potential role of cancer pathology in modulating drug tolerance.\u003c/p\u003e\u003cp\u003eFinally, comorbidities were significantly more prevalent among ADR cases. Chronic kidney disease (ICD-10: N18.9, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) and liver disease (ICD-10: K76.9, p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) were significantly more prevalent, suggesting that impaired drug excretion and metabolism may exacerbate ADR risk. Other conditions (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001) more prevalent in case group included hypothyroidism (ICD-10: E03.9), pneumonia (ICD-10: J18.9), and hydronephrosis (ICD-10: N13.3). A complete list of comorbid conditions and their ICD-10 codes is provided in Supplementary Table\u0026nbsp;4.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\n\u003ch3\u003ePrioritized pharmacogenomic variants identified via feature selection\u003c/h3\u003e\n\u003cp\u003eFeature selection approaches identified a substantial number of SNPs associated with ADR risk from pharmacogene regions. In the complete cohort, 3,601 SNPs surpassed the GWAS significance threshold; 3,271 SNPs were identified in drug-specific cohort analyses. For ADR-specific subgroups, the number of significant SNPs was 2,315 for secondary thrombocytopenia, 2,655 for drug-induced polyneuropathy, 3,181 for drug-induced gastroenteritis and colitis, and 3,521 for skin eruption. All these SNPs were selected from PGx regions and converted to allele dosage format for subsequent ML model development.\u003c/p\u003e\u003cp\u003eKey pharmacogenomic variants prioritized through feature selection across drug-specific, ADR-specific, and complete cohort analyses are summarized below, with detailed cohort information available in Supplementary Table\u0026nbsp;5. These include SNPs with well-established relevance to antineoplastic drug toxicity, such as rs67376798 (OR: 1.34 [95% CI: 1.06, 1.69]) in the DPYD gene, which has been significantly associated with ADR risk in our analysis. This variant is known to reduce the activity of the dihydropyridine dehydrogenase (DPD) enzyme, a key player in the metabolism of 5-fluorouracil, and has strong clinical relevance in predicting toxicity to cancer drugs (\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e). Additional significant variants included rs4673993 (OR: 0.79 [95% CI: 0.63, 0.98]) in \u003cem\u003eATIC\u003c/em\u003e, rs2298383 (OR: 0.72 [95% CI: 0.59, 0.87]) in \u003cem\u003eADORA2A\u003c/em\u003e, and rs1801394 (OR: 0.64 [95% CI: 0.45, 0.91]) in \u003cem\u003eMTRR\u003c/em\u003e, all previously associated with methotrexate toxicity, supported by PharmGKB Level 2B \u0026amp; Level 3 evidence(\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e, \u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e). In addition, variants rs1056892 (OR: 1.26 [95% CI: 1.03, 1.54]) in \u003cem\u003eCBR3\u003c/em\u003e, rs11615 (OR: 1.37 [95% CI: 1.12, 1.67]) in \u003cem\u003eERCC1\u003c/em\u003e, and rs4646 (OR: 1.26 [95% CI: 1.02, 1.55]) in \u003cem\u003eCYP19A1\u003c/em\u003e have been linked to toxicity from other antineoplastic agents, including doxorubicin, cisplatin, capecitabine, and tamoxifen. Furthermore, rs16969968 (OR: 1.50 [95% CI: 1.06, 2.14]) in \u003cem\u003eCHRNA5\u003c/em\u003e, rs1051730 (OR: 1.49 [95% CI: 1.05, 2.11]) in \u003cem\u003eCHRNA3\u003c/em\u003e, rs3745274 (OR: 1.33 [95% CI: 1.09, 1.64]) in \u003cem\u003eCYP2B6\u003c/em\u003e, rs2298383 in \u003cem\u003eADORA2A\u003c/em\u003e, and rs11615 in \u003cem\u003eERCC1\u003c/em\u003e were identified in multiple cohorts, strengthening their potential as robust pharmacogenomic markers of ADR risk. A complete list of significant SNPs with known antineoplastic drug associations is provided in Supplementary Table\u0026nbsp;5.\u003c/p\u003e\n\u003ch3\u003eSuperior performance of MLP and LR using multimodal features in the drug-specific cohort\u003c/h3\u003e\n\u003cp\u003eA GWAS conducted in the drug-specific cohort identified 3,271 SNPs from pharmacogenes. In addition, clinically relevant SNPs from PharmKGB were incorporated into model development, including SNPs associated with fluorouracil (171), capecitabine (\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e), methotrexate (119), and tamoxifen (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e) (Supplementary Table\u0026nbsp;2).\u003c/p\u003e\u003cp\u003eModel performance metrics are summarized in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, with corresponding AUC-ROC curves shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003e. When using genetic features alone, both MLP and LR models demonstrated strong performance, achieving AUC-ROC values of 0.80 and 0.82, respectively. The MLP model achieved higher overall accuracy (0.83) and specificity (0.93), whereas LR exhibited superior sensitivity (0.66), highlighting its effectiveness in correctly identifying ADR cases. In contrast, the random forest (RF) model, despite high specificity (0.96), demonstrated low sensitivity (0.18), reducing its utility for ADR risk prediction.\u003c/p\u003e\u003cp\u003eIncorporating environmental features led to only marginal improvements across all models. However, the inclusion of comorbidity-related variables significantly enhanced predictive performance. With the combined multimodal feature set (genetic, environmental, and comorbidity), the AUC-ROC increased to 0.86 for MLP and 0.85 for LR. The MLP model maintained superior performance, achieving an accuracy of 0.86, sensitivity of 0.56, and specificity of 0.94. These results underscore the effectiveness in capturing the complex interplay of genetic, environmental, and comorbid factors influencing ADR risk.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003ePerformance metrics of ML models across different feature sets in drug-specific cohort (N\u0026thinsp;=\u0026thinsp;1098; 246 Cases, 852 Controls).\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"7\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colspan=\"7\" nameend=\"c7\" namest=\"c1\"\u003e\u003cp\u003e\u003cem\u003eModel evaluations using genetic-only data\u003c/em\u003e\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eML Models\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eAccuracy\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cb\u003ePrecision\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cb\u003eSensitivity\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cb\u003eSpecificity\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e\u003cb\u003eF1 Score\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003eAUC-ROC\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eLR\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.80\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.55\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.66\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.84\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.60\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.82\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eRF\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.79\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.64\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.19\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.97\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.29\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.74\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eSVM\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.82\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.64\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.49\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.92\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.55\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.79\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eXGB\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.81\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.58\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.49\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.90\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.53\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.79\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eMLP\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.84\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.69\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.51\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.94\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.57\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.80\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"7\" nameend=\"c7\" namest=\"c1\"\u003e\u003cp\u003e\u003cem\u003eModel evaluations using genetic and environmental data\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eML Models\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eAccuracy\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cb\u003ePrecision\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cb\u003eSensitivity\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cb\u003eSpecificity\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e\u003cb\u003eF1 Score\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003eAUC-ROC\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eLR\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.81\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.57\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.66\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.86\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.61\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.83\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eRF\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.78\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.62\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.11\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.98\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.18\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.74\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eSVM\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.81\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.59\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.49\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.90\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.53\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.81\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eXGB\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.78\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.52\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.46\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.88\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.49\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.78\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eMLP\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.83\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.67\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.51\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.93\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.58\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.82\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"7\" nameend=\"c7\" namest=\"c1\"\u003e\u003cp\u003e\u003cem\u003eModel evaluations using genetic, environmental, and comorbidity data\u003c/em\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eML Models\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eAccuracy\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cb\u003ePrecision\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cb\u003eSensitivity\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cb\u003eSpecificity\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e\u003cb\u003eF1 Score\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003eAUC-ROC\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eLR\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003e0.82\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cb\u003e0.60\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cb\u003e0.58\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cb\u003e0.89\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e\u003cb\u003e0.59\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e0.85\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eRF\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.80\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.75\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.16\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.98\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.27\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.75\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eSVM\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.83\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.64\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.51\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.92\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.57\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.83\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eXGB\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.79\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.54\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.46\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.89\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.50\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.78\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003e\u003cb\u003eMLP\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003e0.86\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cb\u003e0.75\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cb\u003e0.57\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cb\u003e0.95\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e\u003cb\u003e0.65\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e0.86\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"7\"\u003e\u003cem\u003eLR: Logistic Regression, RF: Random Forest, SVM: Support Vector Machine, GXB: Gradient Boosting (XGBoost), MLP: Multi-Layer Perceptron\u003c/em\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\n\u003ch3\u003eMultimodal feature integration improves the predictive performance of ADR phenotypes\u003c/h3\u003e\n\u003cp\u003eML models were trained to classify individuals as either at high risk or low risk of developing specific treatment-associated ADR phenotypes (i.e., drug-induced gastroenteritis and colitis (ICD-10: K52.1, N\u0026thinsp;=\u0026thinsp;3129), drug-induced polyneuropathy (G62.0, N\u0026thinsp;=\u0026thinsp;1436), secondary thrombocytopenia (D69.5, N\u0026thinsp;=\u0026thinsp;419), and generalized or localized skin eruptions (L27.0 \u0026amp; L27.1, N\u0026thinsp;=\u0026thinsp;1277)). Across all ADR subtypes, model performance improved consistently with the integration of environmental and comorbidity features alongside genetic data. For secondary thrombocytopenia (D69.5), both LR and MLP models demonstrated exceptional predictive performance, with AUC-ROC values increasing from .94 (genetic-only model) to 0.97 following the inclusion of environmental and comorbidity features (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e4\u003c/span\u003e). This trend of enhancement with multimodal feature integration was observed across all ADR phenotypes.\u003c/p\u003e\u003cp\u003eFor drug-induced polyneuropathy (G62.0), the best models achieved AUC-ROC values above .92 in models with the full-feature set (i.e., genetic, environmental, and comorbidity data), exhibiting a clear benefit from multimodal data integration. Although baseline genetic-only models performed moderately well (AUC-ROC of .89), the addition of environmental and comorbid features during model training substantially improved overall performance. Similarly, AUC-ROC values in the full-feature model reached .87 \u0026ndash; .90 for gastroenteritis \u0026amp; colitis (K52.1) and generalized skin eruptions (L27.0 \u0026amp; L27.1), with noticeable improvements in recall and F1 scores. Across all ADR phenotypes, LR and MLP consistently outperformed ensemble methods such as RF and XGB. While RF and XGB offered high specificity, their relatively low sensitivity limited their effectiveness in identifying high-risk individuals (Supplementary Table\u0026nbsp;6).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\n\u003ch3\u003eReduced predictive performance in the complete cohort is attributed to phenotypic heterogeneity\u003c/h3\u003e\n\u003cp\u003eThe performance of ML models in the complete case-control cohort (N\u0026thinsp;=\u0026thinsp;26,235; 5194 cases and 21,041 controls) was lower compared to ADR-specific models, reflecting the increased heterogeneity and complexity of the combined dataset. The best-performing model was LR, achieving an AUC-ROC value of .78 with the full feature set. MLP achieved a maximum AUC-ROC value of .70 with the full feature set, while SVM (.75) performed better than both XGBoost (.72) and RF (.67) models (Supplementary Table\u0026nbsp;7). Across models, LR and MLP demonstrated higher sensitivity and balanced performance, except in complete cohort, whereas RF consistently exhibited poor sensitivity despite high specificity, likely owing to class imbalance. While XGBoost achieved reasonable discrimination, this model exhibited comparatively lower recall, highlighting trade-offs in detecting true ADR cases.\u003c/p\u003e\u003cp\u003eTo assess the standalone contribution of non-genetic features, an LR model trained solely on environmental variables achieved limited discrimination (AUC-ROC: 0.58) but revealed clinically meaningful predictors via SHAP analysis (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). Female gender emerged as the most influential demographic factor, along with clinical biomarkers such as low haemoglobin, elevated cystatin C, and high alkaline phosphatase levels, indicating the potential role of kidney and liver function markers in ADR susceptibility.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eTo the best of our knowledge, this is the first study to systematically integrate whole-genome pharmacogenomic variants with environmental, and ICD-10-derived comorbidities in ADR prediction, offering a more holistic and personalised approach to cancer treatment. Furthermore, the present study leveraged ML models to create a multimodal framework that captures the complex, multifactorial drivers of ADR risk, addressing critical gaps in current PGx-focused approaches and advancing individualized pharmacotherapy. Notably, we used a biologically informed feature selection strategy focused on 169 curated pharmacogenes, harnessing domain-specific knowledge to refine genomic input and minimize the noise typically inherent in large-scale WGS datasets, thereby improving the interpretability and clinical relevance of ML models. Our study shows that integrating pharmacogenomic, environmental and comorbidities data significantly improves ADR prediction performance compared to single-modality models, underscoring the value of multimodal integration for more accurate and clinically actionable risk stratification in oncology.\u003c/p\u003e\u003cp\u003eAs ADRs associated with oncological treatments pose a substantial burden on patient quality of life, improved predictions of ADR risk is of considerable clinical relevance (\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e). Previous studies have demonstrated the utility of WGS in identifying clinically relevant mutations and pharmacogenetic variants associated with drug-induced toxicity in cancer patients, highlighting its potential for personalized medicine (\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e, \u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e). For example, PGx analyses within the 100,000 Genomes Project have shown that germline WGS can identify actionable variants linked to ADRs in cancer treatments, informing prescribing decisions and reducing toxicity risks (\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e). Data in the literature support the use of PGx to improve treatment by tailoring therapeutic strategies based on individual genetic profiles, thereby minimizing toxicity, optimizing drug efficacy (\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e), and reducing healthcare costs. This targeted approach helped reduce the likelihood of Type I errors often encountered in genome-wide testing (\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e, \u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e). Moreover, by applying a more lenient significance threshold within this biologically relevant region, we aimed to limit Type II error and retain informative variants for ML modelling. This strategy aligns with previous efforts that utilized SNP-based platforms to predict chemotherapy-induced ADRs, demonstrating efficacy in reducing treatment-related toxicity (\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e, \u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eOur study shows that females have a higher propensity for ADRs. This is in agreement with previous studies, which report a 1.5 to 1.7-fold increased risk of ADRs in females, likely owing to pharmacokinetic and pharmacodynamic differences and hormonal and immunological influences (\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e, \u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e). Additionally, comorbid conditions, including chronic kidney and liver diseases were significantly more prevalent among ADR cases in the present study, highlighting the potential role of pre-existing organ dysfunction in altering drug metabolism and elimination (\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e, \u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e). This was further supported by elevated levels of cystatin C and alkaline phosphatase in the case group, both indicative of impaired renal and hepatic function (\u003cspan additionalcitationids=\"CR66\" citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e). Haematological and metabolic differences also emerged as biological correlates of ADR susceptibility. Cases exhibited higher white blood cell and lymphocyte counts, and lower haemoglobin and direct bilirubin levels, suggesting systemic inflammation and possible bone marrow suppression, consistent with hematologic toxicity commonly associated with antineoplastic agents (\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e). Additionally, patients with malignancies affecting the hematopoietic and digestive systems exhibited a higher risk of ADRs. This is in line with prior studies reporting frequent ADRs in hematologic cancers, with blood disorders and infections being the most common complications (\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e, \u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e69\u003c/span\u003e). Moreover, ADR incidence was disproportionately high among patients with ill-defined, secondary, or unspecified cancer types, possibly reflecting diagnostic uncertainty or advanced disease stages, both of which could contribute to higher treatment burden and ADR risk (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e, \u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e). Lifestyle-related factors also appeared to influence ADR susceptibility, albeit to a lesser extent. Less time spent outdoors was more common among cases, potentially reflecting poor overall health or lower activity levels. Major dietary changes and smoking status were also observed as contributing factors to ADR risk. These findings are supported by another study demonstrating that modifiable lifestyle factors, including smoking, physical activity, and diet, are significant risk contributors to adverse health outcomes and this may indirectly increase vulnerability to ADRs (\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eFurthermore, the identification of a SNP in the \u003cem\u003eDPYD\u003c/em\u003e gene, which influences fluoropyrimidine toxicity (\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e), reinforces the genetic basis of ADR susceptibility and highlights the importance of genetic polymorphisms in drug metabolism. Additionally, variants in genes \u003cem\u003eATIC, ADORA2A\u003c/em\u003e, and \u003cem\u003eMTRR\u003c/em\u003e were associated with methotrexate toxicity, aligning with previous evidence from PharmGKB and supporting their potential relevance in clinical decision-making (\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e, \u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e). Other notable genes identified included \u003cem\u003eCBR3\u003c/em\u003e and \u003cem\u003eERCC1\u003c/em\u003e, linked to doxorubicin and cisplatin-induced toxicity, and \u003cem\u003eCYP19A1\u003c/em\u003e, associated with adverse responses to tamoxifen and capecitabine. Moreover, polymorphisms in \u003cem\u003eCHRNA3, CHRNA5\u003c/em\u003e, and \u003cem\u003eCYP2B6\u003c/em\u003e, which are involved in drug metabolism and nicotine dependence pathways (\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e, \u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e), were consistently identified across more than one ADR cohort studied, suggesting a broader relevance in modulating treatment outcomes with antineoplastic agents. These findings not only validate known pharmacogenomic associations but also reinforce the utility of incorporating germline genomic profiles into predictive models for ADRs. Collectively, these insights underscore the importance of considering a holistic view encompassing genetic, demographic, clinical, lifestyle and comorbid factors to enhance the efficacy of personalized medicine and reduce the burden of treatment-related toxicity (\u003cspan citationid=\"CR74\" class=\"CitationRef\"\u003e74\u003c/span\u003e, \u003cspan citationid=\"CR75\" class=\"CitationRef\"\u003e75\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eOne of the key strengths of this study is the use of a large, well-characterized cohort from the UK Biobank, enabling robust model training, validation, and subgroup analyses. Consistent with prior literature on ML with genetic data (\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e, \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e, \u003cspan citationid=\"CR76\" class=\"CitationRef\"\u003e76\u003c/span\u003e), deep learning model (MLP) emerged as the top performer in our analysis, particularly when leveraging multimodal features. In our drug-specific prediction tasks, MLP model achieved AUC-ROC values up to 0.86, and in ADR-specific context - predicting secondary thrombocytopenia, performance peaked at 0.97, comparable to or exceeding results reported in earlier studies (e.g., MLP achieving 0.94 for complex disease prediction using SNPs (\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e); XGBoost achieving 0.972 in predicting adverse drug events using EHRs (\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e)). Interestingly, LR - a traditional model, outperformed more complex tree-based methods like RF and XGBoost in our analyses, particularly in the complete cohort where heterogeneity is higher. This aligns with observations from Ho et al. (\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e77\u003c/span\u003e), who noted that simpler models can remain competitive when paired with appropriate feature selection and regularization strategies, especially under high-dimensional but structured data. Performance was comparatively lower in the complete cohort analysis, where greater phenotypic and treatment heterogeneity likely contributed to reduced classification performance. Nevertheless, LR remained the best-performing model with AUC-ROC of 0.78. Importantly, both LR and MLP maintained a more favourable balance between sensitivity and specificity, resulting in improved F1 scores and overall classification reliability - an essential consideration in clinical settings. In contrast, SVM and RF consistently exhibited suboptimal performance across datasets. The geometric constraints of SVM may hinder its effectiveness in high-dimensional genomic data where complex, non-linear patterns are present (\u003cspan citationid=\"CR78\" class=\"CitationRef\"\u003e78\u003c/span\u003e), while RF\u0026rsquo;s reliance on hierarchical splits may fail to capture subtle feature interactions relevant to ADR risk (\u003cspan citationid=\"CR79\" class=\"CitationRef\"\u003e79\u003c/span\u003e). The superior performance of MLP is consistent with its capacity to model non-linear relationships and manage high-dimensional data, aligning with previous findings supporting the utility of neural networks in clinical genomics (\u003cspan citationid=\"CR80\" class=\"CitationRef\"\u003e80\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eFurther improvements in predictive performance could be achieved by increasing sample sizes, particularly for rare ADR events in specific cancer cohorts, and by integrating more detailed longitudinal treatment histories, complementing the robust performance already demonstrated by LR and MLP across analyses. Given their ability to handle multimodal data and generate reliable predictions, ML models for ADR prediction can be meaningfully integrated into clinical workflows by embedding them within EHR systems. This integration can be enabled through a secure application programming interface (API) that serves as a bridge between the ADR prediction model and the EHR. On one side, the API would pull relevant patient data - such as demographics, lifestyle, clinical characteristics, comorbid conditions, medication history, and genetic information - from the EHR. On the other, it would feed this data into the trained ADR model to compute risk scores in real time. These risk estimates could then be delivered back into the EHR interface, flagging high-risk patients directly within the clinician\u0026rsquo;s workflow.\u003c/p\u003e\u003cp\u003eWhile this study offers valuable insights, there were several limitations. The analysis was limited to individuals of white European ancestry to minimize population stratification, which may restrict the generalizability of findings to other ethnic groups. Expanding future research to more diverse populations will be essential for equitable clinical application. Additionally, ADR identification was based on ICD-10 coding, which may miss subtle or misclassified cases owing to inconsistencies in clinical reporting (\u003cspan citationid=\"CR81\" class=\"CitationRef\"\u003e81\u003c/span\u003e). Emerging methods such as natural language processing (NLP), could help capture richer ADR-related information from electronic health records, though challenges such as lexical variability and data imbalance remain (\u003cspan citationid=\"CR82\" class=\"CitationRef\"\u003e82\u003c/span\u003e). Furthermore, polypharmacy - a common scenario in oncology - introduces significant complexity in isolating the causal factors of ADRs, as interactions between multiple drugs can confound the attribution of specific effects to individual agents. Addressing this challenge requires more granular data and advanced modelling approaches. Notably, the current models are designed as a broad, multiple-drug risk assessment, offering a valuable starting point for initial screening in oncology settings. While not individual drug-specific, our approach offers a wider understanding of factors contributing to ADR susceptibility across antineoplastic treatments to identify high-risk individuals who may benefit from closer monitoring (\u003cspan citationid=\"CR83\" class=\"CitationRef\"\u003e83\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eThe superior predictive performance of our models not only advances the understanding of ADR mechanisms but also highlights their potential for real-world clinical application. Building upon previous efforts (e.g. Naranjo Adverse Drug Reaction Probability Scale and the ADR Risk Estimator tool (\u003cspan citationid=\"CR84\" class=\"CitationRef\"\u003e84\u003c/span\u003e, \u003cspan citationid=\"CR85\" class=\"CitationRef\"\u003e85\u003c/span\u003e)), our study advances the field by integrating a broader and richer set of features encompassing pharmacogenomic variants, environmental characteristics, and comorbid conditions to train ML models specifically tailored to oncology population. This holistic approach has the potential to enhance clinical workflows by enabling personalized ADR risk stratification prior to initiating antineoplastic treatments, ultimately supporting safer prescribing, reducing treatment-related complications, and improving patient outcomes. Translating these predictive tools into routine clinical care will require prospective validation in independent cohorts. Embedding such models within clinical decision-support systems will be critical for realizing their full potential in guiding personalized cancer therapy.\u003c/p\u003e\u003cp\u003eThis study addresses the urgent need to tailor oncological treatments to minimize ADRs, leveraging advances in next-generation sequencing technology, curated datasets, PGx, and ML. By integrating genetic, environmental, and comorbidities data with treatment efficacy and ADR profiles, we present a novel framework for early risk stratification and personalized intervention. This approach has the potential to reduce ADR-related complications, improve treatment adherence, and optimize clinical outcomes while lowering healthcare costs. As oncology moves toward more holistic, data-driven care models, our work - one of the first applications of multimodal ML-driven ADR prediction in the field, offers a roadmap for advancing personalized oncology care to improve outcomes for people with cancer.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\u003ch2\u003eStudy population and dataset\u003c/h2\u003e\u003cp\u003eThis retrospective case-control study utilized de-identified data from the biomedical database UK Biobank (Application ID: 86460) to investigate the genomic, environmental (demographic, lifestyle, and clinical) features, and comorbid conditions associated with ADRs in cancer patients being treated with antineoplastic agents. The case group (N\u0026thinsp;=\u0026thinsp;5577) comprised individuals diagnosed with malignant neoplasms who reported ADRs following the correct administration of antineoplastic agents at therapeutic or prophylactic doses. ADRs were identified using the International Classification of Diseases, Tenth Revision (ICD-10) (\u003cspan citationid=\"CR86\" class=\"CitationRef\"\u003e86\u003c/span\u003e) codes Y43.0 (antiallergic and antiemetic drugs), Y43.1 (antineoplastic antimetabolites), and Y43.3 (antineoplastic drugs). Cases of drug poisoning or overdose were excluded.\u003c/p\u003e\u003cp\u003eThe control group (N\u0026thinsp;=\u0026thinsp;22,308) consisted of individuals diagnosed with malignant neoplasms (ICD-10 codes C00-C97) who did not report any ADRs. Cases and controls were matched at a 1:4 ratio, resulting in a total cohort of 27,885 participants. To reduce population stratification and enhance comparability, analyses were restricted to individuals who self-identified as being of white ethnic background. Ethnicity was determined using UK Biobank self-reported data (field 21000; codes 1 [White], 1001 [British], 1002 [Irish], or 1003 [any other White background]). Participants were further restricted to those aged 40\u0026ndash;69 years.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\u003ch2\u003eData Preprocessing and Feature Selection\u003c/h2\u003e\u003cdiv id=\"Sec12\" class=\"Section3\"\u003e\u003ch2\u003eGenetic Data\u003c/h2\u003e\u003cp\u003eA total of 755 participants were excluded from the initial total cohort owing to the unavailability of whole-genome sequencing (WGS) data. Additionally, directly applying ML techniques is currently computationally challenging owing to the large number of features (i.e., SNPs) relative to the number of samples. This imbalance has been shown to lead to cause overfitting, resulting in high training accuracy but poor generalization to new datasets (\u003cspan citationid=\"CR87\" class=\"CitationRef\"\u003e87\u003c/span\u003e, \u003cspan citationid=\"CR88\" class=\"CitationRef\"\u003e88\u003c/span\u003e). To address this, we first reduced genomic dimensionality by focusing on pharmacogenetic regions. A list of 169 pharmacogenes (Supplementary Table\u0026nbsp;1) was compiled from the Pharmacogenomics Knowledgebase (PharmGKB) (\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e, \u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e) and the Clinical Pharmacogenetics Implementation Consortium (CPIC\u0026reg;) (\u003cspan citationid=\"CR89\" class=\"CitationRef\"\u003e89\u003c/span\u003e), and each gene was extended by \u0026plusmn;\u0026thinsp;20 kilobases (kb) to capture potential regulatory variants. Only SNPs within these pharmacogenetic regions were retained for further analysis. Variant-level quality control (QC) was performed using PLINK v2.0 (\u003cspan citationid=\"CR90\" class=\"CitationRef\"\u003e90\u003c/span\u003e) and excluded SNPs with a minor allele frequency (MAF)\u0026thinsp;\u0026le;\u0026thinsp;0.001, Hardy-Weinberg equilibrium p\u0026thinsp;\u0026le;\u0026thinsp;1\u0026times;10⁻⁸, and missing genotype rate\u0026thinsp;\u0026gt;\u0026thinsp;2%. Additionally, individuals with \u0026gt;\u0026thinsp;5% missing genotypes were excluded. After QC, the final dataset consisted of 26,235 individuals (5194 cases and 21,041 controls).\u003c/p\u003e\u003cp\u003eTo further refine the genomic feature set for ML, we applied two complementary strategies for feature selection. First, we conducted GWAS restricted to pharmacogenetic regions (169 pharmacogenes extended by \u0026plusmn;\u0026thinsp;20 kb), using mixed linear models adjusting for age, sex, and population stratification via the top 10 genotype principal components. By narrowing the analysis to pharmacogenes, we reduced the likelihood of Type I errors (i.e., false positives) commonly associated with conventional genome-wide testing(\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e, \u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e). To further reduce the risk of Type II error (i.e., false negatives), we applied a relaxed significance threshold (p\u0026thinsp;\u0026lt;\u0026thinsp;5\u0026times;10⁻\u0026sup2;), retaining variants with potential biological relevance for downstream ML modelling. In the second strategy, we incorporated 373 SNPs from filtered pharmacogenetic regions previously associated with specific antineoplastic agents, curated from the PharmGKB database (Supplementary Table\u0026nbsp;2), to ensure the inclusion of variants supported by prior pharmacogenetic evidence. The resulting SNP set comprising variants identified through GWAS and curated knowledge, was converted to allele dosage format and used as the genomic feature set for model development.\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e\u003ch2\u003eClinical data\u003c/h2\u003e\u003cp\u003eThe selection of environmental (demographic, lifestyle, clinical features) and comorbidity-related features was guided by their relevance to ADRs, existing literature, and potential influence in the univariate analyses using Welch\u0026rsquo;s t-test for continuous variables and Chi-square or Fisher\u0026rsquo;s exact test for categorical variables, as appropriate. Demographic factors included age and gender; clinical measures included blood biomarkers, and lifestyle factors (e.g., exercise, sleep, and dietary habits); and specific habits known to affect health and drug metabolism, including alcohol consumption and smoking, were also included (Supplementary table 4). A wide range of comorbid conditions including chronic kidney disease, liver disease, pneumonia, and others were identified using ICD-10 codes, allowing for a detailed representation of participants\u0026rsquo; health conditions. Missing data were imputed using Multiple Imputation by Chained Equations (MICE) (\u003cspan citationid=\"CR91\" class=\"CitationRef\"\u003e91\u003c/span\u003e).\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec14\" class=\"Section2\"\u003e\u003ch2\u003eML model development\u003c/h2\u003e\u003cp\u003e We developed ADR risk prediction models using ML algorithms, based on a tiered analysis comprising three levels of participant cohorts consisting of cases and controls.\u003c/p\u003e\u003cp\u003e\u003col\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eDrug-Specific Cohort\u003c/b\u003e: Focused on participants with a history of taking one or more of four commonly prescribed cancer medications: fluorouracil, capecitabine, methotrexate, and tamoxifen - selected owing to their frequent clinical use and well established pharmacogenomic relevance (\u003cspan additionalcitationids=\"CR93\" citationid=\"CR92\" class=\"CitationRef\"\u003e92\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR94\" class=\"CitationRef\"\u003e94\u003c/span\u003e). To improve predictive accuracy, this analysis incorporated genetic variants known to interact with these drugs (as catalogued in PharmGKB), along with variants from key pharmacogenes identified through univariate GWAS filtering (N\u0026thinsp;=\u0026thinsp;1098; 246 cases, 852 controls).\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eADR-Specific Cohort\u003c/b\u003e: Four separate ADR-specific cohorts were examined, each defined by a distinct adverse drug reaction (ADR) with a minimum case count of N\u0026thinsp;\u0026gt;\u0026thinsp;80. This threshold was chosen to ensure adequate statistical power, while excluding ADR categories with smaller case numbers that may yield unstable or unreliable estimates. The included ADRs were: i) drug-induced gastroenteritis and colitis (ICD-10: K52.1, N\u0026thinsp;=\u0026thinsp;3129), ii) drug-induced polyneuropathy (G62.0, N\u0026thinsp;=\u0026thinsp;1436), iii) secondary thrombocytopenia (D69.5, N\u0026thinsp;=\u0026thinsp;419), iv) generalized or localized skin eruptions (L27.0 \u0026amp; L27.1, N\u0026thinsp;=\u0026thinsp;1277).\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003e\u003cb\u003eComplete Cohort\u003c/b\u003e: Included all participants who experienced any ADRs associated with antineoplastic drugs (N\u0026thinsp;=\u0026thinsp;26,235; 5194 cases and 21,041 controls).\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003c/ol\u003e\u003c/p\u003e\u003cp\u003eOur approach to model development was incremental, assessing the impact of various feature sets on ADR prediction. We began with genetic features, specifically SNPs identified from 169 curated pharmacogenes. Next, we incorporated environmental variables, and finally, we added comorbid conditions to construct a comprehensive predictive model. We trained predictive models using five supervised machine learning algorithms: LR (\u003cspan citationid=\"CR95\" class=\"CitationRef\"\u003e95\u003c/span\u003e), RF (\u003cspan citationid=\"CR96\" class=\"CitationRef\"\u003e96\u003c/span\u003e), SVM (\u003cspan citationid=\"CR97\" class=\"CitationRef\"\u003e97\u003c/span\u003e), XGBoost (\u003cspan citationid=\"CR98\" class=\"CitationRef\"\u003e98\u003c/span\u003e), and MLP (\u003cspan citationid=\"CR99\" class=\"CitationRef\"\u003e99\u003c/span\u003e). Each model\u0026rsquo;s predictive performance was systematically assessed in three stages: first using only genetic data, then adding environmental features, and finally incorporating comorbid conditions to evaluate the incremental value of each feature set for ADR prediction. To further assess the independent contribution of environmental features, we also trained a logistic regression model using only environmental variables, excluding genetic and comorbidity data. SHapley Additive exPlanations (SHAP) values were computed to interpret the relative importance of individual features in ADR prediction.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eModels were trained using a split dataset, with 70% allocated to training and 30% to testing. Additionally, the distribution of patients who experienced ADRs (case group) in our cohort was markedly imbalanced relative to controls. To mitigate potential bias and improve classifier performance, we employed the Synthetic Minority Oversampling Technique (SMOTE), a widely validated oversampling approach (\u003cspan citationid=\"CR100\" class=\"CitationRef\"\u003e100\u003c/span\u003e). SMOTE was applied to the training set to generate synthetic instances of the minority class (cases) to achieve a more balanced distribution, thereby enhancing model sensitivity, specificity and predictive accuracy (\u003cspan citationid=\"CR101\" class=\"CitationRef\"\u003e101\u003c/span\u003e). The training process involved 5-fold cross-validation using GridSearchCV to optimize hyperparameters and ensure robust model evaluation. Within each fold, the model was trained on four subsets of the training data and validated on the fifth, iterating through all partitions. The final model, selected based on cross-validation performance, was then evaluated on the hold-out test set to assess generalizability. Model effectiveness and optimal hyperparameters in the testing set were assessed using the AUC-ROC. Additional metrics (sensitivity, specificity, accuracy, precision, and F1 score) were used to provide a comprehensive evaluation of classification performance in the testing set. Based on the evaluation metrics, the optimal model was identified primarily by its AUC-ROC performance, while also prioritizing a high F1 score, sensitivity and other metrics. The model development workflow is depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e1\u003c/span\u003e, and detailed information of hyperparameters used for the models are presented in Supplementary Table\u0026nbsp;3.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"Declarations","content":"\u003ch5\u003eEthics approval and consent to participate\u003c/h5\u003e\n\u003cp\u003eThis study was conducted using data from the UK Biobank under approved application number 86460. All participants provided informed consent at recruitment. No additional ethical approval was required for this secondary data analysis.\u003c/p\u003e\n\u003ch5\u003eDATA AVAILABILITY\u003c/h5\u003e\n\u003cp\u003eThe data used in this study were obtained from the UK Biobank. Access to the UK Biobank dataset is not publicly available but can be obtained to researchers through an application process.\u003c/p\u003e\n\u003ch5\u003eFunding\u003c/h5\u003e\n\u003cp\u003eThis work was supported by the National Health and Medical Research Council (NHMRC) Ideas Grant [GNT2029756].\u003c/p\u003e\n\u003ch5\u003eCODE AVAILABILITY\u003c/h5\u003e\n\u003cp\u003eQuantitative analysis and modelling were carried out using PLINK v2.0 and Python v3.11.3, with Bash used to automate data preprocessing and pipeline execution. Analyses in python made use of the following packages: NumPy, Pandas, Sklearn, Xgboost, Imblearn, Matplotlib, SciPy, Seaborn, Shap. Code will be made available upon reasonable request.\u0026nbsp;\u003c/p\u003e\n\u003ch5\u003eACKNOWLEDGEMENTS\u003c/h5\u003e\n\u003cp\u003eThis research has been conducted using the UK Biobank resource under Application Number 86460. We thank the participants of the UK Biobank study, without whom this research would not have been possible. The authors would like to acknowledge Dr. Prathosh AP (IISc, Bengaluru) for his valuable guidance on the machine learning components of this work. We also thank Dr. Senthil Lingarathnam, Director of Pharmacy at Peter MacCallum Cancer Centre, for his insightful feedback and contributions to the clinical framing of cancer-related ADRs and antineoplastic drug selection. We are grateful to Kim N. Tran, Rohit Mishra, and Abhiram D. B. for their thoughtful feedback and manuscript revisions.\u003c/p\u003e\n\u003ch5\u003eAUTHOR CONTRIBUTIONS\u003c/h5\u003e\n\u003cp\u003eSN led the project. SN and RL oversaw the analysis. AJ and VA designed the research. AJ performed the analysis. VA and BH assisted in data curation. SN and AJ prepared the first draft of the manuscript. All authors contributed to manuscript revision, read and approved the submitted version.\u003c/p\u003e\n\u003ch5\u003eCOMPETING INTERESTS\u003c/h5\u003e\n\u003cp\u003eThe authors declare that they have no competing interests\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eEdwards IR, Aronson JK. Adverse drug reactions: definitions, diagnosis, and management. Lancet. 2000;356(9237):1255-9.\u003c/li\u003e\n\u003cli\u003eZhou ZW, Chen XW, Sneed KB, Yang YX, Zhang X, He ZX, et al. Clinical association between pharmacogenomics and adverse drug reactions. Drugs. 2015;75(6):589-631.\u003c/li\u003e\n\u003cli\u003ePirmohamed M, James S, Meakin S, Green C, Scott AK, Walley TJ, et al. Adverse drug reactions as cause of admission to hospital: prospective analysis of 18,820 patients. Bmj-Brit Med J. 2004;329(7456):15-9.\u003c/li\u003e\n\u003cli\u003eThuermann PA, Windecker R, Steffen J, Schaefer M, Tenter U, Reese E, et al. Detection of adverse drug reactions in a neurological department: comparison between intensified surveillance and a computer-assisted approach. Drug Saf. 2002;25(10):713-24.\u003c/li\u003e\n\u003cli\u003eBouvy JC, De Bruin ML, Koopmanschap MA. Epidemiology of adverse drug reactions in Europe: a review of recent observational studies. Drug Saf. 2015;38(5):437-53.\u003c/li\u003e\n\u003cli\u003eAlexopoulou A, Dourakis SP, Mantzoukis D, Pitsariotis T, Kandyli A, Deutsch M, et al. Adverse drug reactions as a cause of hospital admissions: A 6-month experience in a single center in Greece. Eur J Intern Med. 2008;19(7):505-10.\u003c/li\u003e\n\u003cli\u003eAngamo MT, Chalmers L, Curtain CM, Bereznicki LRE. Adverse-Drug-Reaction-Related Hospitalisations in Developed and Developing Countries: A Review of Prevalence and Contributing Factors. Drug Safety. 2016;39(9):847-57.\u003c/li\u003e\n\u003cli\u003eSilva LT, Modesto ACF, Amaral RG, Lopes FM. Hospitalizations and deaths related to adverse drug events worldwide: Systematic review of studies with national coverage. European Journal of Clinical Pharmacology. 2022;78(3):435-66.\u003c/li\u003e\n\u003cli\u003ede Bienassis K, Esmail L, Lopert R, Klazinga N. The economics of medication safety: Improving medication safety through collective, real-time learning. OECD Health Working Papers. 2022(147):0_1-85.\u003c/li\u003e\n\u003cli\u003eWester K, Jonsson AK, Spigset O, Druid H, Hagg S. Incidence of fatal adverse drug reactions: a population based study. British journal of clinical pharmacology. 2008;65(4):573-9.\u003c/li\u003e\n\u003cli\u003eU.S. Food and Drug Administration. U.S. Food and Drug Administration 2022 [Available from: https://www.fda.gov.\u003c/li\u003e\n\u003cli\u003eEmpey PE. Genetic predisposition to adverse drug reactions in the intensive care unit. Crit Care Med. 2010;38(6 Suppl):S106-16.\u003c/li\u003e\n\u003cli\u003eSendekie AK, Netere AK, Tesfaye S, Dagnew EM, Belachew EA. Incidence and patterns of adverse drug reactions among adult patients hospitalized in the University of Gondar comprehensive specialized hospital: A prospective observational follow-up study. Plos One. 2023;18(2).\u003c/li\u003e\n\u003cli\u003eWahlang JB, Laishram PD, Brahma DK, Sarkar C, Lahon J, Nongkynrih BS. Adverse drug reactions due to cancer chemotherapy in a tertiary care teaching hospital. Therapeutic Advances in Drug Safety. 2017;8(2):61-6.\u003c/li\u003e\n\u003cli\u003eKuderer NM, Desai A, Lustberg MB, Lyman GH. Mitigating acute chemotherapy-associated adverse events in patients with cancer. Nature Reviews Clinical Oncology. 2022;19(11):681-97.\u003c/li\u003e\n\u003cli\u003eTajani BB, Maheswari E, Maka VV, Nair AS. Adverse drug reactions and drug-related problems with supportive care medications among the oncological population. Discov Oncol. 2024;15(1).\u003c/li\u003e\n\u003cli\u003eChopra D, Rehan HS, Sharma V, Mishra R. Chemotherapy-induced adverse drug reactions in oncology patients: A prospective observational survey. Indian J Med Paediatr Oncol. 2016;37(1):42-6.\u003c/li\u003e\n\u003cli\u003eDu RF, Wang X, Ma LX, Larcher LM, Tang H, Zhou HY, et al. Adverse reactions of targeted therapy in cancer patients: a retrospective study of hospital medical data in China. Bmc Cancer. 2021;21(1).\u003c/li\u003e\n\u003cli\u003eLavan AH, O\u0026apos;Mahony D, Buckley M, O\u0026apos;Mahony D, Gallagher P. Adverse Drug Reactions in an Oncological Population: Prevalence, Predictability, and Preventability. Oncologist. 2019;24(9):e968-e77.\u003c/li\u003e\n\u003cli\u003eDavies MR, Martinec M, Walls R, Schwarz R, Mirams GR, Wang K, et al. Use of Patient Health Records to Quantify Drug-Related Pro-arrhythmic Risk. Cell Reports Medicine. 2020;1(5):100076.\u003c/li\u003e\n\u003cli\u003eNaranjo CA, Busto U, Sellers EM, Sandor P, Ruiz I, Roberts EA, et al. A method for estimating the probability of adverse drug reactions. Clinical Pharmacology \u0026amp; Therapeutics. 1981;30(2):239-45.\u003c/li\u003e\n\u003cli\u003ePirmohamed M, Park BK. Genetic susceptibility to adverse drug reactions. Trends in Pharmacological Sciences. 2001;22(6):298-305.\u003c/li\u003e\n\u003cli\u003eMeyer UA. Pharmacogenetics and adverse drug reactions. The Lancet. 2000;356(9242):1667-71.\u003c/li\u003e\n\u003cli\u003eSwen JJ, Van Der Wouden CH, Manson LE, Abdullah-Koolmees H, Blagec K, Blagus T, et al. A 12-gene pharmacogenetic panel to prevent adverse drug reactions: an open-label, multicentre, controlled, cluster-randomised crossover implementation study. The Lancet. 2023;401(10374):347-56.\u003c/li\u003e\n\u003cli\u003eCacabelos R, Cacabelos N, Carril JC. The role of pharmacogenomics in adverse drug reactions. Expert Rev Clin Pharmacol. 2019;12(5):407-42.\u003c/li\u003e\n\u003cli\u003eWeinshilboum RM, Wang LW. Pharmacogenomics: Precision Medicine and Drug Response. Mayo Clin Proc. 2017;92(11):1711-22.\u003c/li\u003e\n\u003cli\u003eSadee W, Wang D, Hartmann K, Toland AE. Pharmacogenomics: Driving Personalized Medicine. Pharmacological Reviews. 2023;75(4):789-814.\u003c/li\u003e\n\u003cli\u003ePalmirotta R, Carella C, Silvestris E, Cives M, Stucci SL, Tucci M, et al. SNPs in predicting clinical efficacy and toxicity of chemotherapy: walking through the quicksand. Oncotarget. 2018;9(38):25355-82.\u003c/li\u003e\n\u003cli\u003eCafiero C, Palmirotta R, Martinelli C, Micera A, Giaco L, Persiani F, et al. Oncological Treatment Adverse Reaction Prediction: Development and Initial Validation of a Pharmacogenetic Model in Non-Small-Cell Lung Cancer Patients. Genes (Basel). 2025;16(3).\u003c/li\u003e\n\u003cli\u003eGong L, Whirl-Carrillo M, Klein TE. PharmGKB, an Integrated Resource of Pharmacogenomic Knowledge. Curr Protoc. 2021;1(8):e226.\u003c/li\u003e\n\u003cli\u003eKrzyszczyk P, Acevedo A, Davidoff EJ, Timmins LM, Marrero-Berrios I, Patel M, et al. The growing role of precision and personalized medicine for cancer treatment. Technology (Singap World Sci). 2018;6(3-4):79-100.\u003c/li\u003e\n\u003cli\u003ePirmohamed M. Personalized Pharmacogenomics: Predicting Efficacy and Adverse Drug Reactions. Annual Review of Genomics and Human Genetics. 2014;15(Volume 15, 2014):349-70.\u003c/li\u003e\n\u003cli\u003eJohnson JA. Ethnic differences in cardiovascular drug response: potential contribution of pharmacogenetics. Circulation. 2008;118(13):1383-93.\u003c/li\u003e\n\u003cli\u003eDang MT, Hambleton J, Kayser SR. The influence of ethnicity on warfarin dosage requirement. Ann Pharmacother. 2005;39(6):1008-12.\u003c/li\u003e\n\u003cli\u003eMagavern EF, Megase M, Thompson J, Marengo G, Jacobsen J, Smedley D, et al. Pharmacogenetics and adverse drug reports: Insights from a United Kingdom national pharmacovigilance database. PLOS Medicine. 2025;22(3):e1004565.\u003c/li\u003e\n\u003cli\u003eTang C, Livingston MJ, Safirstein R, Dong Z. Cisplatin nephrotoxicity: new insights and therapeutic implications. Nat Rev Nephrol. 2023;19(1):53-72.\u003c/li\u003e\n\u003cli\u003eHanoodi M, Mittal M. Methotrexate. StatPearls. Treasure Island (FL)2024.\u003c/li\u003e\n\u003cli\u003eGaytan SL, Lawan A, Chang J, Nurunnabi M, Bajpeyi S, Boyle JB, et al. The beneficial role of exercise in preventing doxorubicin-induced cardiotoxicity. Front Physiol. 2023;14:1133423.\u003c/li\u003e\n\u003cli\u003eTamang R, Bharati L, Khatiwada AP, Ozaki A, Shrestha S. Pattern of Adverse Drug Reactions Associated with the Use of Anticancer Drugs in an Oncology-Based Hospital of Nepal. JMA J. 2022;5(4):416-26.\u003c/li\u003e\n\u003cli\u003eMonaco A, Pantaleo E, Amoroso N, Lacalamita A, Lo Giudice C, Fonzino A, et al. A primer on machine learning techniques for genomic applications. Comput Struct Biotechnol J. 2021;19:4345-59.\u003c/li\u003e\n\u003cli\u003eSyrowatka A, Song W, Amato MG, Foer D, Edrees H, Co Z, et al. Key use cases for artificial intelligence to reduce the frequency of adverse drug events: a scoping review. The Lancet Digital Health. 2022;4(2):e137-e48.\u003c/li\u003e\n\u003cli\u003eBowe AK, Lightbody G, Staines A, Murray DM. Big data, machine learning, and population health: predicting cognitive outcomes in childhood. Pediatr Res. 2023;93(2):300-7.\u003c/li\u003e\n\u003cli\u003eKirasich K, Smith T, Sadler B. Random forest vs logistic regression: binary classification for heterogeneous datasets. SMU Data Science Review. 2018;1(3):9.\u003c/li\u003e\n\u003cli\u003eLai C, Zimmer AD, O\u0026apos;Connor R, Kim S, Chan R, van den Akker J, et al. LEAP: Using machine learning to support variant classification in a clinical setting. Human Mutation. 2020;41(6):1079-90.\u003c/li\u003e\n\u003cli\u003eHuang RJ, Kwon NS-E, Tomizawa Y, Choi AY, Hernandez-Boussard T, Hwang JH. A Comparison of Logistic Regression Against Machine Learning Algorithms for Gastric Cancer Risk Prediction Within Real-World Clinical Data Streams. JCO Clinical Cancer Informatics. 2022(6):e2200039.\u003c/li\u003e\n\u003cli\u003eLong C, Lv GT, Fu XM. Development of a general logistic model for disease risk prediction using multiple SNPs. Febs Open Bio. 2019;9(11):2006-12.\u003c/li\u003e\n\u003cli\u003eZhang F, Sun B, Diao XL, Zhao W, Shu T. Prediction of adverse drug reactions based on knowledge graph embedding. Bmc Med Inform Decis. 2021;21(1).\u003c/li\u003e\n\u003cli\u003eAlzoubi H, Alzubi R, Ramzan N. Deep Learning Framework for Complex Disease Risk Prediction Using Genomic Variations. Sensors (Basel). 2023;23(9).\u003c/li\u003e\n\u003cli\u003eKidwai-Khan F, Rentsch CT, Pulk R, Alcorn C, Brandt CA, Justice AC. Pharmacogenomics driven decision support prototype with machine learning: A framework for improving patient care. Front Big Data. 2022;5:1059088.\u003c/li\u003e\n\u003cli\u003eLai NH, Shen WC, Lee CN, Chang JC, Hsu MC, Kuo LN, et al. Comparison of the predictive outcomes for anti-tuberculosis drug-induced hepatotoxicity by different machine learning techniques. Comput Methods Programs Biomed. 2020;188:105307.\u003c/li\u003e\n\u003cli\u003eRuzzo A, Graziano F, Galli F, Galli F, Rulli E, Lonardi S, et al. Dihydropyrimidine dehydrogenase pharmacogenetics for predicting fluoropyrimidine-related toxicity in the randomised, phase III adjuvant TOSCA trial in high-risk colon cancer patients. Brit J Cancer. 2017;117(9):1269-77.\u003c/li\u003e\n\u003cli\u003eWhirl-Carrillo M, McDonagh EM, Hebert JM, Gong L, Sangkuhl K, Thorn CF, et al. Pharmacogenomics Knowledge for Personalized Medicine. Clinical Pharmacology \u0026amp; Therapeutics. 2012;92(4):414-7.\u003c/li\u003e\n\u003cli\u003eWhirl-Carrillo M, Huddart R, Gong L, Sangkuhl K, Thorn CF, Whaley R, et al. An Evidence-Based Framework for Evaluating Pharmacogenomics Knowledge for Personalized Medicine. Clinical Pharmacology \u0026amp; Therapeutics. 2021;110(3):563-72.\u003c/li\u003e\n\u003cli\u003eYazbeck V, Alesi E, Myers J, Hackney MH, Cuttino L, Gewirtz DA. An overview of chemotoxicity and radiation toxicity in cancer therapy. Adv Cancer Res. 2022;155:1-27.\u003c/li\u003e\n\u003cli\u003eKinnersley B, Sud A, Everall A, Cornish AJ, Chubb D, Culliford R, et al. Analysis of 10,478 cancer genomes identifies candidate driver genes and opportunities for precision oncology. Nat Genet. 2024;56(9):1868-77.\u003c/li\u003e\n\u003cli\u003eLeong IUS, Cabrera CP, Cipriani V, Ross PJ, Turner RM, Stuckey A, et al. Large-Scale Pharmacogenomics Analysis of Patients With Cancer Within the 100,000 Genomes Project Combining Whole-Genome Sequencing and Medical Records to Inform Clinical Practice. J Clin Oncol. 2024:JCO2302761.\u003c/li\u003e\n\u003cli\u003ePharoah PDP. Genetic susceptibility, predicting risk and preventing cancer. Recent Results Canc. 2003;163:7-18.\u003c/li\u003e\n\u003cli\u003eUffelmann E, Posthuma D, Peyrot WJ. Genome-wide association studies of polygenic risk score-derived phenotypes may lead to inflated false positive rates. Sci Rep. 2023;13(1):4219.\u003c/li\u003e\n\u003cli\u003eMarees AT, de Kluiver H, Stringer S, Vorspan F, Curis E, Marie-Claire C, et al. A tutorial on conducting genome-wide association studies: Quality control and statistical analysis. Int J Methods Psychiatr Res. 2018;27(2):e1608.\u003c/li\u003e\n\u003cli\u003eBjorn N, Badam TVS, Spalinskas R, Branden E, Koyi H, Lewensohn R, et al. Whole-genome sequencing and gene network modules predict gemcitabine/carboplatin-induced myelosuppression in non-small cell lung cancer patients. NPJ Syst Biol Appl. 2020;6(1):25.\u003c/li\u003e\n\u003cli\u003eRademaker M. Do women have more adverse drug reactions? Am J Clin Dermatol. 2001;2(6):349-51.\u003c/li\u003e\n\u003cli\u003eWatson S, Caster O, Rochon PA, den Ruijter H. Reported adverse drug reactions in women and men: Aggregated evidence from globally collected individual case reports during half a century. EClinicalMedicine. 2019;17:100188.\u003c/li\u003e\n\u003cli\u003eWang X, Chen X. Clinical Characteristics of 162 Patients with Drug-Induced Liver and/or Kidney Injury. Biomed Res Int. 2020;2020:3930921.\u003c/li\u003e\n\u003cli\u003eSpanakis M, Roubedaki M, Tzanakis I, Zografakis-Sfakianakis M, Patelarou E, Patelarou A. Impact of Adverse Drug Reactions in Patients with End Stage Renal Disease in Greece. Int J Environ Res Public Health. 2020;17(23).\u003c/li\u003e\n\u003cli\u003eAlosco ML, Spitznagel MB, Strain G, Devlin M, Cohen R, Crosby RD, et al. The effects of cystatin C and alkaline phosphatase changes on cognitive function 12-months after bariatric surgery. J Neurol Sci. 2014;345(1-2):176-80.\u003c/li\u003e\n\u003cli\u003eChai X, Huang HB, Feng G, Cao YH, Cheng QS, Li SH, et al. Baseline Serum Cystatin C Is a Potential Predictor for Acute Kidney Injury in Patients with Acute Pancreatitis. Dis Markers. 2018;2018.\u003c/li\u003e\n\u003cli\u003eBuyukberber M, Koruk I, Cykman O, Koruk M, K\u0026uuml;\u0026ccedil;\u0026uuml;koglu ME, Sakman A, et al. Serum cystatin C measurement in differential diagnosis of intra and extrahepatic cholestatic diseases. Ann Hepatol. 2010;9(1):58-62.\u003c/li\u003e\n\u003cli\u003eYu B, Yan XD, Zhu YY, Luo T, Sohail M, Ning H, et al. Analysis of adverse drug reactions/events of cancer chemotherapy and the potential mechanism of Danggui Buxue decoction against bone marrow suppression induced by chemotherapy. Frontiers in Pharmacology. 2023;14.\u003c/li\u003e\n\u003cli\u003eAmaro-Hosey K, Dan\u0026eacute;s I, Vendrell L, Alonso L, Renedo B, Gros L, et al. Adverse Reactions to Drugs of Special Interest in a Pediatric Oncohematology Service. Frontiers in Pharmacology. 2021;12.\u003c/li\u003e\n\u003cli\u003eBelachew SA, Erku DA, Mekuria AB, Gebresillassie BM. Pattern of chemotherapy-related adverse effects among adult cancer patients treated at Gondar university referral hospital, Ethiopia: a cross-sectional study. Drug Healthc Patient. 2016;8:83-90.\u003c/li\u003e\n\u003cli\u003eJackowska B, Wisniewski P, Noinski T, Bandosz P. Effects of lifestyle-related risk factors on life expectancy: A comprehensive model for use in early prevention of premature mortality from noncommunicable diseases. Plos One. 2024;19(3).\u003c/li\u003e\n\u003cli\u003eChmielowiec K, Chmielowiec J, Stronska-Pluta A, Trybek G, Smiarowska M, Suchanecka A, et al. Association of Polymorphism CHRNA5 and CHRNA3 Gene in People Addicted to Nicotine. Int J Environ Res Public Health. 2022;19(17).\u003c/li\u003e\n\u003cli\u003eMuderrisoglu A, Babaoglu E, Korkmaz ET, Kalkisim S, Karabulut E, Emri S, et al. Comparative Assessment of Outcomes in Drug Treatment for Smoking Cessation and Role of Genetic Polymorphisms of Human Nicotinic Acetylcholine Receptor Subunits. Frontiers in Genetics. 2022;13.\u003c/li\u003e\n\u003cli\u003eFarnoush A, Sedighi-Maman Z, Rasoolian B, Heath JJ, Fallah B. Prediction of adverse drug reactions using demographic and non-clinical drug characteristics in FAERS data. Sci Rep. 2024;14(1):23636.\u003c/li\u003e\n\u003cli\u003eAlomar MJ. Factors affecting the development of adverse drug reactions (Review article). Saudi Pharm J. 2014;22(2):83-94.\u003c/li\u003e\n\u003cli\u003eMohsen A. Deep Learning Prediction of Adverse Drug Reactions in Drug Discovery Using Open TG\u0026ndash;GATEs and FAERS Databases.\u003c/li\u003e\n\u003cli\u003eHo DSW, Schierding W, Wake M, Saffery R, O\u0026apos;Sullivan J. Machine Learning SNP Based Prediction for Precision Medicine. Front Genet. 2019;10:267.\u003c/li\u003e\n\u003cli\u003eClarke R, Ressom HW, Wang A, Xuan J, Liu MC, Gehan EA, et al. The properties of high-dimensional data spaces: implications for exploring gene and protein expression data. Nat Rev Cancer. 2008;8(1):37-49.\u003c/li\u003e\n\u003cli\u003eGhosh D, Cabrera J. Enriched Random Forest for High Dimensional Genomic Data. IEEE/ACM Trans Comput Biol Bioinform. 2022;19(5):2817-28.\u003c/li\u003e\n\u003cli\u003eAnastopoulos IN, Herczeg CK, Davis KN, Dixit AC. Multi-Drug Featurization and Deep Learning Improve Patient-Specific Predictions of Adverse Events. Int J Env Res Pub He. 2021;18(5).\u003c/li\u003e\n\u003cli\u003eHohl CM, Karpov A, Reddekopp L, Stausberg J. ICD-10 codes used to identify adverse drug events in administrative data: a systematic review. J Am Med Inform Assn. 2014;21(3):547-57.\u003c/li\u003e\n\u003cli\u003eWieland-Jorna Y, van Kooten D, Verheij RA, de Man Y, Francke AL, Oosterveld-Vlug MG. Natural language processing systems for extracting information from electronic health records about activities of daily living. A systematic review. Jamia Open. 2024;7(2).\u003c/li\u003e\n\u003cli\u003eDsouza VS, Leyens L, Kurian JR, Brand A, Brand H. Artificial intelligence (AI) in pharmacovigilance: A systematic review on predicting adverse drug reactions (ADR) in hospitalized patients. Res Soc Admin Pharm. 2025;21(6):453-62.\u003c/li\u003e\n\u003cli\u003eGarc\u0026iacute;a-Cort\u0026eacute;s M, Lucena MI, Pachkoria K, Borraz Y, Hidalgo R, Andrade RJ, et al. Evaluation of Naranjo Adverse Drug Reactions Probability Scale in causality assessment of drug-induced liver injury. Aliment Pharm Therap. 2008;27(9):780-9.\u003c/li\u003e\n\u003cli\u003eValeanu A, Damian C, Marineci CD, Negres S. The development of a scoring and ranking strategy for a patient-tailored adverse drug reaction prediction in polypharmacy. Sci Rep-Uk. 2020;10(1).\u003c/li\u003e\n\u003cli\u003eJohn HH. International Statistical Classification of Diseases and Related Health-Problems - World-Hlth-Org. J Roy Soc Health. 1994;114(6):339-.\u003c/li\u003e\n\u003cli\u003ePudjihartono N, Fadason T, Kempa-Liehr AW, O\u0026apos;Sullivan JM. A Review of Feature Selection Methods for Machine Learning-Based Disease Risk Prediction. Front Bioinform. 2022;2:927312.\u003c/li\u003e\n\u003cli\u003eSilva PP, Gaudillo JD, Vilela JA, Roxas-Villanueva RML, Tiangco BJ, Domingo MR, et al. A machine learning-based SNP-set analysis approach for identifying disease-associated susceptibility loci. Sci Rep. 2022;12(1):15817.\u003c/li\u003e\n\u003cli\u003eRelling MV, Klein TE. CPIC: Clinical Pharmacogenetics Implementation Consortium of the Pharmacogenomics Research Network. Clinical Pharmacology \u0026amp; Therapeutics. 2011;89(3):464-7.\u003c/li\u003e\n\u003cli\u003ePurcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MA, Bender D, et al. PLINK: a tool set for whole-genome association and population-based linkage analyses. Am J Hum Genet. 2007;81(3):559-75.\u003c/li\u003e\n\u003cli\u003evan Buuren S, Groothuis-Oudshoorn K. mice: Multivariate Imputation by Chained Equations in R. J Stat Softw. 2011;45(3):1-67.\u003c/li\u003e\n\u003cli\u003eRofaiel S, Muo EN, Mousa SA. Pharmacogenetics in breast cancer: steps toward personalized medicine in breast cancer management. Pharmacogn Pers Med. 2010;3:129-43.\u003c/li\u003e\n\u003cli\u003eS\u0026aacute;nchez-Bayona R, Catal\u0026aacute;n C, Cobos MA, Bergamino M. Pharmacogenomics in Solid Tumors: A Comprehensive Review of Genetic Variability and Its Clinical Implications. Cancers. 2025;17(6).\u003c/li\u003e\n\u003cli\u003eFranczyk B, Rysz J, Gluba-Brz\u0026oacute;zka A. Pharmacogenetics of Drugs Used in the Treatment of Cancers. Genes-Basel. 2022;13(2).\u003c/li\u003e\n\u003cli\u003eMccullagh P. Generalized Linear-Models. Eur J Oper Res. 1984;16(3):285-92.\u003c/li\u003e\n\u003cli\u003eBreiman L. Random forests. Mach Learn. 2001;45(1):5-32.\u003c/li\u003e\n\u003cli\u003eCortes C, Vapnik V. Support-Vector Networks. Mach Learn. 1995;20(3):273-97.\u003c/li\u003e\n\u003cli\u003eChen TQ, Guestrin C. XGBoost: A Scalable Tree Boosting System. Kdd\u0026apos;16: Proceedings of the 22nd Acm Sigkdd International Conference on Knowledge Discovery and Data Mining. 2016:785-94.\u003c/li\u003e\n\u003cli\u003eAttali JG, Pages G. Approximations of functions by a multilayer perceptron: a new approach. Neural Networks. 1997;10(6):1069-81.\u003c/li\u003e\n\u003cli\u003eBarua S, Islam MM, Yao X, Murase K. MWMOTE-Majority Weighted Minority Oversampling Technique for Imbalanced Data Set Learning. Ieee T Knowl Data En. 2014;26(2):405-25.\u003c/li\u003e\n\u003cli\u003eDablain D, Krawczyk B, Chawla N. DeepSMOTE: Fusing Deep Learning and SMOTE for Imbalanced Data. Ieee T Neur Net Lear. 2023;34(9):6390-404.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Machine Learning, Pharmacogenomics, Adverse drug reactions, cancer, comorbidities, personalized medicine, prediction model","lastPublishedDoi":"10.21203/rs.3.rs-7431071/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7431071/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eAccurately predicting adverse drug reactions (ADRs) in cancer remains challenging. We applied a pharmacogenomics-driven machine learning framework that integrates genomic, environmental, and comorbidity data to enhance ADR prediction. Using UK Biobank, we analysed 26,235 antineoplastic-treated patients, identifying ADRs via ICD-10 codes. Features included GWAS-derived SNPs from 169 pharmacogenes, curated PharmGKB variants, demographics, lifestyle, laboratory biomarkers, and comorbidities. Five supervised models were trained; subgroup analyses assessed drug-specific and ADR-specific cohort performance. Logistic regression and multilayer perceptron models performed best. In drug-specific cohort, genetic data alone achieved AUC-ROC 0.82 (LR) and 0.80 (MLP), improving to 0.85 and 0.86 when all features were included. For secondary thrombocytopenia, LR and MLP achieved AUC-ROC 0.94 using genetic data only and 0.97 with all features. SHAP and univariate analyses highlighted female gender, elevated cystatin C, and alkaline phosphatase (all p\u0026thinsp;\u0026lt;\u0026thinsp;0.001); haematologic and digestive cancers showed higher risk compared to other cancer types. This integrative approach supports data-driven clinical decision-making to reduce ADRs.\u003c/p\u003e","manuscriptTitle":"Pharmacogenomics-Driven Multimodal Data Integration Improves Predictions of Adverse Drug Reactions in Cancer Patients using Machine Learning","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-09-23 07:32:33","doi":"10.21203/rs.3.rs-7431071/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"6cf1f0cf-bb2d-4943-bb04-89ed3fc86df5","owner":[],"postedDate":"September 23rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":54687629,"name":"Health sciences/Biomarkers"},{"id":54687630,"name":"Biological sciences/Cancer"},{"id":54687631,"name":"Biological sciences/Computational biology and bioinformatics"},{"id":54687632,"name":"Health sciences/Oncology"}],"tags":[],"updatedAt":"2025-10-27T14:34:30+00:00","versionOfRecord":[],"versionCreatedAt":"2025-09-23 07:32:33","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-7431071","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7431071","identity":"rs-7431071","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00