Machine Learning Techniques with Fairness for Prediction of Completion of Drug and Alcohol Rehabilitation | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine Learning Techniques with Fairness for Prediction of Completion of Drug and Alcohol Rehabilitation Karen Roberts-Licklider, Theodore Trafalis This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4208301/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 7 You are reading this latest preprint version Abstract The aim of this study is to look at predicting whether a person will complete a drug and alcohol rehabilitation program and the number of times a person attends. The study is based on demographic data obtained from Substance Abuse and Mental Health Services Administration (SAMHSA) from both admissions and discharge data from drug and alcohol rehabilitation centers in Oklahoma. Demographic data is highly categorical which led to binary encoding being used and various fairness measures being utilized to mitigate bias of nine demographic variables. Kernel methods such as linear, polynomial, sigmoid, and radial basis functions were compared using support vector machines at various parameter ranges to find the optimal values. These were then compared to methods such as decision trees, random forests, and neural networks. Synthetic Minority Oversampling Technique Nominal (SMOTEN) for categorical data was used to balance the data with imputation for missing data. The nine bias variables were then intersectionalized to mitigate bias and the dual and triple interactions were integrated to use the probabilities to look at worst case ratio fairness mitigation. Disparate Impact, Statistical Parity difference, Conditional Statistical Parity Ratio, Demographic Parity, Demographic Parity Ratio, Equalized Odds, Equalized Odds Ratio, Equal Opportunity, and Equalized Opportunity Ratio were all explored at both the binary and multiclass scenarios. Support Vector Machines Kernel Methods Fairness Measures SMOTEN Decision Trees Random Forests Neural Networks Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 28 May, 2024 Reviewers agreed at journal 12 May, 2024 Reviewers agreed at journal 29 Apr, 2024 Reviewers invited by journal 27 Apr, 2024 Editor assigned by journal 20 Apr, 2024 Submission checks completed at journal 03 Apr, 2024 First submitted to journal 02 Apr, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4208301","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":286918476,"identity":"c981d26c-e644-4f96-86ab-e23453430426","order_by":0,"name":"Karen Roberts-Licklider","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABRElEQVRIie3Qv0sDMRQH8HcEcsurXa+U6r+QI3AiHO2/knIQl4JOpS41ky7qXME/QhCcI4G61PNfaCl0cildKoI1udbhji5uDvkO+fX4wHsB8PH5h2G62L6ho7A4CQo00HYFIKQoBqpM4u3dWOuIKAhsSUjFPsKhSsARcKSObB9JwrfZ/NyR19uX+cV6KA7CsdDTfnoGBJdzhLT1qEskxVPOR26WmzyLJ8IIilJrkcsTRWpPHEHyKgFJmwhraEe9pKGEnSIKle5eGdYxtWdbMt0KSeqL8AthAjzqHX8qMfwlG2YbW1iyqRIeSUoQxhCPekmgBLGEaku0I64BXSXxaEGayAyySZ41lCxmse3lmSWUxw8s4/dlwt5lsMKBOXQ/tlTp8PLoehzP1v02g7qZTT8G7dZdmewg4J7XXcnHx8fH5+/5AV6pdLfs5tcsAAAAAElFTkSuQmCC","orcid":"","institution":"University of Oklahoma","correspondingAuthor":true,"prefix":"","firstName":"Karen","middleName":"","lastName":"Roberts-Licklider","suffix":""},{"id":286918477,"identity":"da8bcf88-9ae2-4140-ba99-dbb1aa168095","order_by":1,"name":"Theodore Trafalis","email":"","orcid":"","institution":"University of Oklahoma","correspondingAuthor":false,"prefix":"","firstName":"Theodore","middleName":"","lastName":"Trafalis","suffix":""}],"badges":[],"createdAt":"2024-04-02 17:51:08","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4208301/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4208301/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":54264762,"identity":"fd8366ae-cb21-4565-9611-47c5f76c566c","added_by":"auto","created_at":"2024-04-08 04:33:23","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1271112,"visible":true,"origin":"","legend":"","description":"","filename":"SVMFairnessTreatmentDataSpringerJournal4.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4208301/v1_covered_d5cab41c-0223-43f2-bd5c-1acdf4cb8c9c.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Machine Learning Techniques with Fairness for Prediction of Completion of Drug and Alcohol Rehabilitation","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"international-journal-of-data-science-and-analytics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jdsa","sideBox":"Learn more about [International Journal of Data Science and Analytics](http://link.springer.com/journal/41060)","snPcode":"41060","submissionUrl":"https://submission.nature.com/new-submission/41060/3","title":"International Journal of Data Science and Analytics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Support Vector Machines, Kernel Methods, Fairness Measures, SMOTEN, Decision Trees, Random Forests, Neural Networks","lastPublishedDoi":"10.21203/rs.3.rs-4208301/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4208301/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The aim of this study is to look at predicting whether a person will complete a drug and alcohol rehabilitation program and the number of times a person attends. The study is based on demographic data obtained from Substance Abuse and Mental Health Services Administration (SAMHSA) from both admissions and discharge data from drug and alcohol rehabilitation centers in Oklahoma. Demographic data is highly categorical which led to binary encoding being used and various fairness measures being utilized to mitigate bias of nine demographic variables. Kernel methods such as linear, polynomial, sigmoid, and radial basis functions were compared using support vector machines at various parameter ranges to find the optimal values. These were then compared to methods such as decision trees, random forests, and neural networks. Synthetic Minority Oversampling Technique Nominal (SMOTEN) for categorical data was used to balance the data with imputation for missing data. The nine bias variables were then intersectionalized to mitigate bias and the dual and triple interactions were integrated to use the probabilities to look at worst case ratio fairness mitigation. Disparate Impact, Statistical Parity difference, Conditional Statistical Parity Ratio, Demographic Parity, Demographic Parity Ratio, Equalized Odds, Equalized Odds Ratio, Equal Opportunity, and Equalized Opportunity Ratio were all explored at both the binary and multiclass scenarios. ","manuscriptTitle":"Machine Learning Techniques with Fairness for Prediction of Completion of Drug and Alcohol Rehabilitation","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-04-08 04:25:14","doi":"10.21203/rs.3.rs-4208301/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2024-05-28T11:48:24+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"16212868340057788666254051389548993262","date":"2024-05-12T16:45:23+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"220131838100985424359245151535606122796","date":"2024-04-29T23:50:03+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-04-27T11:30:19+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-04-20T18:38:48+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-04-03T05:33:35+00:00","index":"","fulltext":""},{"type":"submitted","content":"International Journal of Data Science and Analytics","date":"2024-04-02T17:49:54+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"international-journal-of-data-science-and-analytics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"jdsa","sideBox":"Learn more about [International Journal of Data Science and Analytics](http://link.springer.com/journal/41060)","snPcode":"41060","submissionUrl":"https://submission.nature.com/new-submission/41060/3","title":"International Journal of Data Science and Analytics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"2390e886-ae32-40bd-ac44-c9eebdf52e77","owner":[],"postedDate":"April 8th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2024-07-24T23:14:11+00:00","versionOfRecord":[],"versionCreatedAt":"2024-04-08 04:25:14","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4208301","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4208301","identity":"rs-4208301","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.