Accurate Clinical Toxicity Prediction using Multi-task Deep Neural Nets and Contrastive Molecular Explanations

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Explainable machine learning for molecular toxicity prediction is a promising approach for efficient drug development and chemical safety. A predictive ML model of toxicity can reduce experimental cost and time while mitigating ethical concerns by significantly reducing animal and clinical testing. Herein, we use a deep learning framework for simultaneously modeling in vitro, in vivo, and clinical toxicity data. Two different molecular input representations are used; Morgan fingerprints and pre-trained SMILES embeddings. A multi-task deep learning model accurately predicts toxicity for all endpoints, including clinical, as indicated by the area under the Receiver Operator Characteristic curve and balanced accuracy. In particular, pre-trained molecular SMILES embeddings as input to the multi-task model improved clinical toxicity predictions compared to existing models in MoleculeNet benchmark. Additionally, our multitask approach is comprehensive in the sense that it is comparable to state-of-the-art approaches for specific endpoints in in vitro, in vivo and clinical platforms. Through both the multi-task model and transfer learning, we were able to indicate the minimal need of in vivo data for clinical toxicity predictions. To provide confidence and explain the model’s predictions, we adapt a post-hoc contrastive explanation method that returns pertinent positive and negative features, which correspond well to known mutagenic and reactive toxicophores, such as unsubstituted bonded heteroatoms, aromatic amines, and Michael receptors. Furthermore, toxicophore recovery by pertinent feature analysis captures more of the in vitro (53%) and in vivo (56%), rather than of the clinical (8%), endpoints, and indeed uncovers a preference in known toxicophore data towards in vitro and in vivo experimental data. To our knowledge, this is the first contrastive explanation, using both present and absent substructures, for predictions of clinical and in vivo molecular toxicity.
Full text 15,366 characters · extracted from preprint-html · click to expand
Accurate Clinical Toxicity Prediction using Multi-task Deep Neural Nets and Contrastive Molecular Explanations | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Accurate Clinical Toxicity Prediction using Multi-task Deep Neural Nets and Contrastive Molecular Explanations Bhanushee Sharma, Vijil Chenthamarakshan, Amit Dhurandhar, Shiranee Pereira, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-1605700/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 25 Mar, 2023 Read the published version in Scientific Reports → Version 1 posted 8 You are reading this latest preprint version Abstract Explainable machine learning for molecular toxicity prediction is a promising approach for efficient drug development and chemical safety. A predictive ML model of toxicity can reduce experimental cost and time while mitigating ethical concerns by significantly reducing animal and clinical testing. Herein, we use a deep learning framework for simultaneously modeling in vitro, in vivo, and clinical toxicity data. Two different molecular input representations are used; Morgan fingerprints and pre-trained SMILES embeddings. A multi-task deep learning model accurately predicts toxicity for all endpoints, including clinical, as indicated by the area under the Receiver Operator Characteristic curve and balanced accuracy. In particular, pre-trained molecular SMILES embeddings as input to the multi-task model improved clinical toxicity predictions compared to existing models in MoleculeNet benchmark. Additionally, our multitask approach is comprehensive in the sense that it is comparable to state-of-the-art approaches for specific endpoints in in vitro, in vivo and clinical platforms. Through both the multi-task model and transfer learning, we were able to indicate the minimal need of in vivo data for clinical toxicity predictions. To provide confidence and explain the model’s predictions, we adapt a post-hoc contrastive explanation method that returns pertinent positive and negative features, which correspond well to known mutagenic and reactive toxicophores, such as unsubstituted bonded heteroatoms, aromatic amines, and Michael receptors. Furthermore, toxicophore recovery by pertinent feature analysis captures more of the in vitro (53%) and in vivo (56%), rather than of the clinical (8%), endpoints, and indeed uncovers a preference in known toxicophore data towards in vitro and in vivo experimental data. To our knowledge, this is the first contrastive explanation, using both present and absent substructures, for predictions of clinical and in vivo molecular toxicity. Full Text Additional Declarations No competing interests reported. Supplementary Files supplementary.pdf Cite Share Download PDF Status: Published Journal Publication published 25 Mar, 2023 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Major revision 27 Jun, 2022 Reviews received at journal 17 Jun, 2022 Reviewers agreed at journal 03 Jun, 2022 Reviewers invited by journal 02 Jun, 2022 Editor assigned by journal 27 May, 2022 Editor invited by journal 11 May, 2022 Submission checks completed at journal 11 May, 2022 First submitted to journal 28 Apr, 2022 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-1605700","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":105208839,"identity":"50e43d3a-0f3a-4694-96e1-2536d13d2618","order_by":0,"name":"Bhanushee Sharma","email":"","orcid":"","institution":"Chemical and Biological Engineering, RPI","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Bhanushee","middleName":"","lastName":"Sharma","suffix":""},{"id":105208840,"identity":"6c691d3a-6863-4cf9-b69c-5ca8a82116b1","order_by":1,"name":"Vijil Chenthamarakshan","email":"","orcid":"","institution":"IBM Research, Yorktown Heights","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Vijil","middleName":"","lastName":"Chenthamarakshan","suffix":""},{"id":105208841,"identity":"a327df7b-d0ab-4458-a692-cccc808aa68e","order_by":2,"name":"Amit Dhurandhar","email":"","orcid":"","institution":"IBM Research, Yorktown Heights","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Amit","middleName":"","lastName":"Dhurandhar","suffix":""},{"id":105208842,"identity":"8e124c1f-fcdb-468a-8357-474f3ea8671f","order_by":3,"name":"Shiranee Pereira","email":"","orcid":"","institution":"People for Animals, Chennai","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Shiranee","middleName":"","lastName":"Pereira","suffix":""},{"id":105208843,"identity":"7fe6ff65-2ff8-4b9a-b0be-a4a4350ec048","order_by":4,"name":"James A. Hendler","email":"","orcid":"","institution":"Computer Science, RPI","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"James","middleName":"A.","lastName":"Hendler","suffix":""},{"id":105208844,"identity":"9ff71804-a4e9-4567-8fa9-bb6545feaaed","order_by":5,"name":"Jonathan S. Dordick","email":"","orcid":"","institution":"Computer Science, RPI","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jonathan","middleName":"S.","lastName":"Dordick","suffix":""},{"id":105208845,"identity":"eb3f2715-c769-4a30-a30b-79dcc276f698","order_by":6,"name":"Payel Das","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAsElEQVRIiWNgGAWjYDCCAyCiAoiZeUjScoZkLYxtIBaxWviOH3/24Oe8w/YM7LxHNzDU2EQT1CJ5JsfcsHfb4cQGZr60GwzH0nIbCGkxOJDDJsG77XAC0C9mNxgbDhOh5fzzZ5J/5wAdRryWGwlm0rwNhxkbiNYieeONubHMsfTENpBfEojxC9/59GcP39RY2/Pznz1240ONDWEtQMCGIBOIUI7QMgpGwSgYBaMAJwAAqQ895ZqDQ1oAAAAASUVORK5CYII=","orcid":"","institution":"IBM Research, Yorktown Heights","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Payel","middleName":"","lastName":"Das","suffix":""}],"badges":[],"createdAt":"2022-04-28 16:44:19","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-1605700/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-1605700/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-023-31169-8","type":"published","date":"2023-03-25T20:08:13+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":21571073,"identity":"df11d09b-8911-4392-b520-2e85923ea264","added_by":"auto","created_at":"2022-05-17 16:45:44","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1518320,"visible":true,"origin":"","legend":"","description":"","filename":"toxicitypaper.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1605700/v1_covered.pdf"},{"id":21571064,"identity":"9706f0c6-7172-4177-97c0-544a3de9920b","added_by":"auto","created_at":"2022-05-17 16:45:31","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":3020236,"visible":true,"origin":"","legend":"","description":"","filename":"supplementary.pdf","url":"https://assets-eu.researchsquare.com/files/rs-1605700/v1/a915cddcea002e0c021481a3.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Accurate Clinical Toxicity Prediction using Multi-task Deep Neural Nets and Contrastive Molecular Explanations","fulltext":[{"header":"Full Text","content":"This preprint is available for \u003ca href='/article/rs-1605700/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e."}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-1605700/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-1605700/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"Explainable machine learning for molecular toxicity prediction is a promising approach for efficient drug development and chemical safety. A predictive ML model of toxicity can reduce experimental cost and time while mitigating ethical concerns by significantly reducing animal and clinical testing. Herein, we use a deep learning framework for simultaneously modeling in vitro, in vivo, and clinical toxicity data. Two different molecular input representations are used; Morgan fingerprints and pre-trained SMILES embeddings. A multi-task deep learning model accurately predicts toxicity for all endpoints, including clinical, as indicated by the area under the Receiver Operator Characteristic curve and balanced accuracy. In particular, pre-trained molecular SMILES embeddings as input to the multi-task model improved clinical toxicity predictions compared to existing models in MoleculeNet benchmark. Additionally, our multitask approach is comprehensive in the sense that it is comparable to state-of-the-art approaches for specific endpoints in in vitro, in vivo and clinical platforms. Through both the multi-task model and transfer learning, we were able to indicate the minimal need of in vivo data for clinical toxicity predictions. To provide confidence and explain the model’s predictions, we adapt a post-hoc contrastive explanation method that returns pertinent positive and negative features, which correspond well to known mutagenic and reactive toxicophores, such as unsubstituted bonded heteroatoms, aromatic amines, and Michael receptors. Furthermore, toxicophore recovery by pertinent feature analysis captures more of the in vitro (53%) and in vivo (56%), rather than of the clinical (8%), endpoints, and indeed uncovers a preference in known toxicophore data towards in vitro and in vivo experimental data. To our knowledge, this is the first contrastive explanation, using both present and absent substructures, for predictions of clinical and in vivo molecular toxicity.","manuscriptTitle":"Accurate Clinical Toxicity Prediction using Multi-task Deep Neural Nets and Contrastive Molecular Explanations","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2022-05-17 16:45:29","doi":"10.21203/rs.3.rs-1605700/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2022-06-27T07:41:37+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2022-06-17T07:54:33+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"7b2a13cd-d2ad-4d2e-8f61-287d551ec4a0","date":"2022-06-03T08:37:55+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2022-06-02T12:32:31+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2022-05-27T14:16:27+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2022-05-11T11:59:36+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2022-05-11T11:56:51+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2022-04-28T16:38:01+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"9dbe3248-a9fc-4c1c-8127-66fef2cadd2d","owner":[],"postedDate":"May 17th, 2022","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2023-10-16T20:14:04+00:00","versionOfRecord":{"articleIdentity":"rs-1605700","link":"https://doi.org/10.1038/s41598-023-31169-8","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2023-03-25 20:08:13","publishedOnDateReadable":"March 25th, 2023"},"versionCreatedAt":"2022-05-17 16:45:29","video":"","vorDoi":"10.1038/s41598-023-31169-8","vorDoiUrl":"https://doi.org/10.1038/s41598-023-31169-8","workflowStages":[]},"version":"v1","identity":"rs-1605700","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-1605700","identity":"rs-1605700","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0