Advancing the development of Deep Learning and Machine Learning models for oral drugs through diverse descriptor classes: A focus on Pharmacokinetic Parameters (Vdss and PPB) | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Advancing the development of Deep Learning and Machine Learning models for oral drugs through diverse descriptor classes: A focus on Pharmacokinetic Parameters (Vdss and PPB) Rakesh Bantu, Samiron Phukan, Simon Haydar This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5233119/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 11 Jun, 2025 Read the published version in Molecular Diversity → Version 1 posted 9 You are reading this latest preprint version Abstract In the present study, we report a predictive deep learning (DL) and machine learning (ML) model for pharmacokinetics (PK) parameters such as volume of distribution (Vdss) and plasma protein Binding (PPB). Using DL & ML algorithms our study provides a deeper and novel insights into the role of molecular descriptors in determining the PK parameters such as Vdss and PPB. FDA approved drugs with oral route of administration and having reported PK parameters were taken as the dataset. This was used for establishment of the foundational datasets followed by computation of different molecular descriptor classes. Feature engineering by Boruta algorithm exhibited significant increase in accuracy of the models. Features identified by Boruta algorithm, were trained for different models separately- for both Vdss and PPB. The highest predictive scores amongst the models were achieved in Gradient Boosting (GB) and Stacking Classifier with 80% and 78% for Vdss. In the case of PPB, Random Forest (RF) and Gradient Boosting (GB) algorithm predicted the highest scores of 73% and 71% respectively in comparison to all other algorithms. In summary we report here appropriate ML algorithms like Stacking Classifier - by utilizing an unreported feature engineering algorithm -to predict Vdss and PPB individually considering over 67 descriptors each with ≥80% accuracy and 73% accuracy respectively. Additionally, we developed models based on the shared descriptors between Vdss and PPB. Quantum chemical descriptors like MLFERs (MLFER_BH, MLFER_BO & MLFER_E) and Topological descriptors like piPC5, piPC6, piPC9 & TpiPC identified as the common drivers of the functional activity of Vdss and PPB together. Drug Metabolism and Pharmacokinetics (DMPK) Volume of distribution plasma protein binding Molecular descriptors cheminformatics Machine learning (ML) Deep learning (DL). Full Text Additional Declarations No competing interests reported. Supplementary Files Suplementaryfiles.zip Cite Share Download PDF Status: Published Journal Publication published 11 Jun, 2025 Read the published version in Molecular Diversity → Version 1 posted Editorial decision: Revision requested 06 Feb, 2025 Reviews received at journal 28 Oct, 2024 Reviews received at journal 24 Oct, 2024 Reviewers agreed at journal 20 Oct, 2024 Reviewers agreed at journal 16 Oct, 2024 Reviewers invited by journal 15 Oct, 2024 Editor assigned by journal 10 Oct, 2024 Submission checks completed at journal 10 Oct, 2024 First submitted to journal 09 Oct, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5233119","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":371013764,"identity":"640a4fe9-be0b-4407-984f-95c06078aba7","order_by":0,"name":"Rakesh Bantu","email":"","orcid":"","institution":"Aragen Lifesciences Ltd","correspondingAuthor":false,"prefix":"","firstName":"Rakesh","middleName":"","lastName":"Bantu","suffix":""},{"id":371013765,"identity":"e81da525-938e-4e44-8732-b4123f424f2e","order_by":1,"name":"Samiron Phukan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+ElEQVRIiWNgGAWjYJCCA2BSgoEZSNqQriWNFLsgWg4TVijffjrxcEXFHQZz6R5jg587zsubtx+/+LiAwU5OtwG7FoMzuRsOnjnzjMFyzhnjxN4ztw3nnMkpNp7BkGxsdgCHFgaglsa2wwwGN3KMD/C23WacwZCTJs3DcCBxGw4t8v1vEVoO/m07Zz+D/036b3xaGG4g2ZLM23YgcYZE+jFmfFoMbgBtaThzmMdyRlqxsWxbcvIMiTfM0jwGuP0i35+7+WNDxWE5c4nkzZJv2+xsZ/CnP/zMU2Enh0sLDPAYoLINcCtFuBDBZH9AhPpRMApGwSgYQQAA1z1he7UfapYAAAAASUVORK5CYII=","orcid":"","institution":"Aragen Lifesciences Ltd","correspondingAuthor":true,"prefix":"","firstName":"Samiron","middleName":"","lastName":"Phukan","suffix":""},{"id":371013766,"identity":"ec104376-0039-4ec1-b647-03e46b830ff7","order_by":2,"name":"Simon Haydar","email":"","orcid":"","institution":"Aragen Lifesciences Ltd","correspondingAuthor":false,"prefix":"","firstName":"Simon","middleName":"","lastName":"Haydar","suffix":""}],"badges":[],"createdAt":"2024-10-09 13:53:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5233119/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5233119/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s11030-025-11235-1","type":"published","date":"2025-06-11T15:57:27+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":84726491,"identity":"bff92665-ba6c-49d9-9a29-dfa90164c99f","added_by":"auto","created_at":"2025-06-16 16:05:50","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1024507,"visible":true,"origin":"","legend":"","description":"","filename":"ManuscriptMoleculardiversity20240910.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5233119/v1_covered_16169831-5d48-4b24-b9c5-20bca1cf5eae.pdf"},{"id":72752311,"identity":"3ab6cf89-9d5e-490c-b8e9-92ecfff1a7c5","added_by":"auto","created_at":"2025-01-01 15:37:49","extension":"zip","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":10177067,"visible":true,"origin":"","legend":"","description":"","filename":"Suplementaryfiles.zip","url":"https://assets-eu.researchsquare.com/files/rs-5233119/v1/9398072d6b260a8bd9886941.zip"}],"financialInterests":"No competing interests reported.","formattedTitle":"Advancing the development of Deep Learning and Machine Learning models for oral drugs through diverse descriptor classes: A focus on Pharmacokinetic Parameters (Vdss and PPB)","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"molecular-diversity","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"modi","sideBox":"Learn more about [Molecular Diversity](http://link.springer.com/journal/11030)","snPcode":"11030","submissionUrl":"https://submission.nature.com/new-submission/11030/3","title":"Molecular Diversity","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Drug Metabolism and Pharmacokinetics (DMPK), Volume of distribution, plasma protein binding, Molecular descriptors, cheminformatics, Machine learning (ML), Deep learning (DL).","lastPublishedDoi":"10.21203/rs.3.rs-5233119/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5233119/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eIn the present study, we report a predictive deep learning (DL) and machine learning (ML) model for pharmacokinetics (PK) parameters such as volume of distribution (Vdss) and plasma protein Binding (PPB). Using DL \u0026amp; ML algorithms our study provides a deeper and novel insights into the role of molecular descriptors in determining the PK parameters such as Vdss and PPB. FDA approved drugs with oral route of administration and having reported PK parameters were taken as the dataset. This was used for establishment of the foundational datasets followed by computation of different molecular descriptor classes.\u003c/p\u003e\n\u003cp\u003eFeature engineering by Boruta algorithm exhibited significant increase in accuracy of the models. Features identified by Boruta algorithm, were trained for different models separately- for both Vdss and PPB. The highest predictive scores amongst the models were achieved in Gradient Boosting (GB) and Stacking Classifier with 80% and 78% for Vdss. In the case of PPB, Random Forest (RF) and Gradient Boosting (GB) algorithm predicted the highest scores of 73% and 71% respectively in comparison to all other algorithms.\u003c/p\u003e\n\u003cp\u003eIn summary we report here appropriate ML algorithms like Stacking Classifier - by utilizing an unreported feature engineering algorithm -to predict Vdss and PPB individually considering over 67 descriptors each with ≥80% accuracy and 73% accuracy respectively.\u003c/p\u003e\n\u003cp\u003eAdditionally, we developed models based on the shared descriptors between Vdss and PPB. Quantum chemical descriptors like MLFERs (MLFER_BH, MLFER_BO \u0026amp; MLFER_E) and Topological descriptors like piPC5, piPC6, piPC9 \u0026amp; TpiPC identified as the common drivers of the functional activity of Vdss and PPB together.\u003c/p\u003e","manuscriptTitle":"Advancing the development of Deep Learning and Machine Learning models for oral drugs through diverse descriptor classes: A focus on Pharmacokinetic Parameters (Vdss and PPB)","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-01-01 15:37:44","doi":"10.21203/rs.3.rs-5233119/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-02-06T17:18:05+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-10-28T04:08:26+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-10-24T14:45:38+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"131105782833756519206274133044937215333","date":"2024-10-20T14:08:37+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"74257106303980075415692699722964139304","date":"2024-10-16T07:50:10+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-10-15T10:40:16+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-10-10T17:19:01+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-10-10T09:32:17+00:00","index":"","fulltext":""},{"type":"submitted","content":"Molecular Diversity","date":"2024-10-09T13:40:43+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"molecular-diversity","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"modi","sideBox":"Learn more about [Molecular Diversity](http://link.springer.com/journal/11030)","snPcode":"11030","submissionUrl":"https://submission.nature.com/new-submission/11030/3","title":"Molecular Diversity","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"4de24be8-32b5-4414-832b-0bc8b0d2b2ad","owner":[],"postedDate":"January 1st, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2025-06-16T16:00:36+00:00","versionOfRecord":{"articleIdentity":"rs-5233119","link":"https://doi.org/10.1007/s11030-025-11235-1","journal":{"identity":"molecular-diversity","isVorOnly":false,"title":"Molecular Diversity"},"publishedOn":"2025-06-11 15:57:27","publishedOnDateReadable":"June 11th, 2025"},"versionCreatedAt":"2025-01-01 15:37:44","video":"","vorDoi":"10.1007/s11030-025-11235-1","vorDoiUrl":"https://doi.org/10.1007/s11030-025-11235-1","workflowStages":[]},"version":"v1","identity":"rs-5233119","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5233119","identity":"rs-5233119","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.