Performance Analysis of Machine Learning Models with Optimized Feature Selection Techniques for Darknet Traffic Classification | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Method Article Performance Analysis of Machine Learning Models with Optimized Feature Selection Techniques for Darknet Traffic Classification Felix Etyang, Pramod Pavithran, Gideon Mwendwa, Sheenamol Yoosaf, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5677244/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 8 You are reading this latest preprint version Abstract Darknet traffic classification is a crucial area of cybersecurity, targeting the anonymity of network activities within an anonymized network. This study evaluates the efficacy of three feature selection methods, Boruta, Recursive Feature Elimination (RFE), and Lasso, across eight machine learning models, including Random Forest, Gradient Boosting, Neural Networks, and SVM. The research enhances computational efficiency and model accuracy through rigorous data preprocessing, such as encoding, normalization, and downsampling. Boruta emerged as the most effective method, with Random Forest achieving 98.7% accuracy and 99.66% ROC-AUC, showcasing its potential for high-accuracy applications. In contrast, Lasso excelled with simpler models like Naive Bayes, optimizing it significantly, while RFE was noted for the fastest training times, which is ideal for resource-constrained environments. This comprehensive performance analysis underscores the critical role of appropriate feature selection in darknet traffic classification. It sets the stage for future research to explore deeper, more complex machine learning frameworks. The findings illustrate significant trade-offs between computational efficiency and model performance, guiding the selection of feature selection techniques based on specific application need. Feature Selection Lasso RFE Darknet Traffic Classification Machine Learning Boruta Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 03 Apr, 2025 Reviews received at journal 30 Mar, 2025 Reviewers agreed at journal 26 Mar, 2025 Reviews received at journal 25 Mar, 2025 Reviewers agreed at journal 24 Mar, 2025 Reviewers invited by journal 24 Mar, 2025 Submission checks completed at journal 22 Mar, 2025 First submitted to journal 21 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5677244","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Method Article","associatedPublications":[],"authors":[{"id":433050206,"identity":"8c7b4d90-9241-43b3-9efe-167425fae0f7","order_by":0,"name":"Felix Etyang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABDElEQVRIiWNgGAWjYBAC9gbmhgMw9geJCiDFzNyAVwvPAUaQFgMQm43B4gxICyNhLQxwLZVtIJqQFvbGxsM8FX8S+2e3X3twc15tNH87UMuPim24tfAcbDg444xB4ow7Z8oNZ247njvjMGMDY8+Z2zi12EskNhz42GaQ2HAjJ0Factux3AagFmbGNtxaeOQfNhxI/GeQOB+k5e+cY7nzCWqRAIbYxwaDxA030o9JSDbU5G4gqIUnEeiXY8bGG2/kMBtIHDuQuxGo5SA+v/CwHz78madGTnbejfSHDyRq6nLnnT988MGPCtxaYMCxgYEHFDmHwbwDBNUDgT0wvTwA0nXEKB4Fo2AUjIIRBgCHNWdJVm0ZMAAAAABJRU5ErkJggg==","orcid":"","institution":"Cochin University of Science and Technology","correspondingAuthor":true,"prefix":"","firstName":"Felix","middleName":"","lastName":"Etyang","suffix":""},{"id":433050207,"identity":"89dd7203-17f5-42d1-a6f3-fdb1b309e034","order_by":1,"name":"Pramod Pavithran","email":"","orcid":"","institution":"Cochin University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Pramod","middleName":"","lastName":"Pavithran","suffix":""},{"id":433050208,"identity":"f10016e7-a54e-4800-883a-f1796b6bdf50","order_by":2,"name":"Gideon Mwendwa","email":"","orcid":"","institution":"National Forensic Sciences University","correspondingAuthor":false,"prefix":"","firstName":"Gideon","middleName":"","lastName":"Mwendwa","suffix":""},{"id":433050209,"identity":"f037ce28-8b11-4b19-b921-5cbf991183fa","order_by":3,"name":"Sheenamol Yoosaf","email":"","orcid":"","institution":"Cochin University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Sheenamol","middleName":"","lastName":"Yoosaf","suffix":""},{"id":433050210,"identity":"fd57b266-138b-467e-bbb3-4401288d1183","order_by":4,"name":"Ngaira Mandela","email":"","orcid":"","institution":"National Forensic Sciences University","correspondingAuthor":false,"prefix":"","firstName":"Ngaira","middleName":"","lastName":"Mandela","suffix":""}],"badges":[],"createdAt":"2024-12-19 13:38:31","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5677244/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5677244/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":79171110,"identity":"af2b873f-da7a-407a-b5dc-839cefec6a8e","added_by":"auto","created_at":"2025-03-25 09:24:28","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":766035,"visible":true,"origin":"","legend":"","description":"","filename":"Comparativeanalysis.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5677244/v1_covered_b6297409-a6dc-48a0-a27b-50473908cb30.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Performance Analysis of Machine Learning Models with Optimized Feature Selection Techniques for Darknet Traffic Classification","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"discover-computing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Discover Computing](https://link.springer.com/journal/10791)","snPcode":"10791","submissionUrl":"https://submission.springernature.com/new-submission/10791/3","title":"Discover Computing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Feature Selection, Lasso, RFE, Darknet Traffic Classification, Machine Learning, Boruta","lastPublishedDoi":"10.21203/rs.3.rs-5677244/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5677244/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDarknet traffic classification is a crucial area of cybersecurity, targeting the anonymity of network activities within an anonymized network. This study evaluates the efficacy of three feature selection methods, Boruta, Recursive Feature Elimination (RFE), and Lasso, across eight machine learning models, including Random Forest, Gradient Boosting, Neural Networks, and SVM. The research enhances computational efficiency and model accuracy through rigorous data preprocessing, such as encoding, normalization, and downsampling. Boruta emerged as the most effective method, with Random Forest achieving 98.7% accuracy and 99.66% ROC-AUC, showcasing its potential for high-accuracy applications. In contrast, Lasso excelled with simpler models like Naive Bayes, optimizing it significantly, while RFE was noted for the fastest training times, which is ideal for resource-constrained environments. This comprehensive performance analysis underscores the critical role of appropriate feature selection in darknet traffic classification. It sets the stage for future research to explore deeper, more complex machine learning frameworks. The findings illustrate significant trade-offs between computational efficiency and model performance, guiding the selection of feature selection techniques based on specific application need.\u003c/p\u003e","manuscriptTitle":"Performance Analysis of Machine Learning Models with Optimized Feature Selection Techniques for Darknet Traffic Classification","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-03-25 08:52:18","doi":"10.21203/rs.3.rs-5677244/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-04-03T11:41:25+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-03-30T13:46:01+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"317573256806288356300007815959753168261","date":"2025-03-26T04:59:18+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-03-25T20:58:08+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"197736592628077481684094723937919804259","date":"2025-03-24T09:15:42+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-03-24T09:09:13+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-03-22T08:00:41+00:00","index":"","fulltext":""},{"type":"submitted","content":"Discover Computing","date":"2025-03-21T19:38:18+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"discover-computing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Discover Computing](https://link.springer.com/journal/10791)","snPcode":"10791","submissionUrl":"https://submission.springernature.com/new-submission/10791/3","title":"Discover Computing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"dd7e1974-5014-48e7-a50c-13ecea71b4bc","owner":[],"postedDate":"March 25th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2025-06-13T07:23:29+00:00","versionOfRecord":[],"versionCreatedAt":"2025-03-25 08:52:18","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5677244","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5677244","identity":"rs-5677244","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.