Performance Analysis of Machine Learning Models with Optimized Feature Selection Techniques for Darknet Traffic Classification

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-16

This study analyzed Boruta, RFE, and Lasso feature selection methods with eight machine learning models for darknet traffic classification, finding Boruta and Random Forest achieved the highest accuracy.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-16 · read from full text

This preprint evaluates three feature selection methods (Boruta, recursive feature elimination, and Lasso) applied to eight machine learning models (including Random Forest, gradient boosting, neural networks, and SVM) for darknet traffic classification, using data preprocessing steps such as encoding, normalization, and downsampling. Across the comparisons, Boruta is reported as the most effective feature selection approach, with Random Forest reaching 98.7% accuracy and 99.66% ROC-AUC. The study also reports trade-offs: Lasso performs well with simpler models like Naive Bayes, while RFE is described as yielding the fastest training times for resource-constrained settings. A major caveat explicitly noted is that the work is a preprint that has not been peer reviewed by a journal and is under review. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Darknet traffic classification is a crucial area of cybersecurity, targeting the anonymity of network activities within an anonymized network. This study evaluates the efficacy of three feature selection methods, Boruta, Recursive Feature Elimination (RFE), and Lasso, across eight machine learning models, including Random Forest, Gradient Boosting, Neural Networks, and SVM. The research enhances computational efficiency and model accuracy through rigorous data preprocessing, such as encoding, normalization, and downsampling. Boruta emerged as the most effective method, with Random Forest achieving 98.7% accuracy and 99.66% ROC-AUC, showcasing its potential for high-accuracy applications. In contrast, Lasso excelled with simpler models like Naive Bayes, optimizing it significantly, while RFE was noted for the fastest training times, which is ideal for resource-constrained environments. This comprehensive performance analysis underscores the critical role of appropriate feature selection in darknet traffic classification. It sets the stage for future research to explore deeper, more complex machine learning frameworks. The findings illustrate significant trade-offs between computational efficiency and model performance, guiding the selection of feature selection techniques based on specific application need.
Full text 12,730 characters · extracted from preprint-html · click to expand
Performance Analysis of Machine Learning Models with Optimized Feature Selection Techniques for Darknet Traffic Classification | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Method Article Performance Analysis of Machine Learning Models with Optimized Feature Selection Techniques for Darknet Traffic Classification Felix Etyang, Pramod Pavithran, Gideon Mwendwa, Sheenamol Yoosaf, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5677244/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 8 You are reading this latest preprint version Abstract Darknet traffic classification is a crucial area of cybersecurity, targeting the anonymity of network activities within an anonymized network. This study evaluates the efficacy of three feature selection methods, Boruta, Recursive Feature Elimination (RFE), and Lasso, across eight machine learning models, including Random Forest, Gradient Boosting, Neural Networks, and SVM. The research enhances computational efficiency and model accuracy through rigorous data preprocessing, such as encoding, normalization, and downsampling. Boruta emerged as the most effective method, with Random Forest achieving 98.7% accuracy and 99.66% ROC-AUC, showcasing its potential for high-accuracy applications. In contrast, Lasso excelled with simpler models like Naive Bayes, optimizing it significantly, while RFE was noted for the fastest training times, which is ideal for resource-constrained environments. This comprehensive performance analysis underscores the critical role of appropriate feature selection in darknet traffic classification. It sets the stage for future research to explore deeper, more complex machine learning frameworks. The findings illustrate significant trade-offs between computational efficiency and model performance, guiding the selection of feature selection techniques based on specific application need. Feature Selection Lasso RFE Darknet Traffic Classification Machine Learning Boruta Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 03 Apr, 2025 Reviews received at journal 30 Mar, 2025 Reviewers agreed at journal 26 Mar, 2025 Reviews received at journal 25 Mar, 2025 Reviewers agreed at journal 24 Mar, 2025 Reviewers invited by journal 24 Mar, 2025 Submission checks completed at journal 22 Mar, 2025 First submitted to journal 21 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5677244","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Method Article","associatedPublications":[],"authors":[{"id":433050206,"identity":"8c7b4d90-9241-43b3-9efe-167425fae0f7","order_by":0,"name":"Felix Etyang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABDElEQVRIiWNgGAWjYBAC9gbmhgMw9geJCiDFzNyAVwvPAUaQFgMQm43B4gxICyNhLQxwLZVtIJqQFvbGxsM8FX8S+2e3X3twc15tNH87UMuPim24tfAcbDg444xB4ow7Z8oNZ247njvjMGMDY8+Z2zi12EskNhz42GaQ2HAjJ0Factux3AagFmbGNtxaeOQfNhxI/GeQOB+k5e+cY7nzCWqRAIbYxwaDxA030o9JSDbU5G4gqIUnEeiXY8bGG2/kMBtIHDuQuxGo5SA+v/CwHz78madGTnbejfSHDyRq6nLnnT988MGPCtxaYMCxgYEHFDmHwbwDBNUDgT0wvTwA0nXEKB4Fo2AUjIIRBgCHNWdJVm0ZMAAAAABJRU5ErkJggg==","orcid":"","institution":"Cochin University of Science and Technology","correspondingAuthor":true,"prefix":"","firstName":"Felix","middleName":"","lastName":"Etyang","suffix":""},{"id":433050207,"identity":"89dd7203-17f5-42d1-a6f3-fdb1b309e034","order_by":1,"name":"Pramod Pavithran","email":"","orcid":"","institution":"Cochin University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Pramod","middleName":"","lastName":"Pavithran","suffix":""},{"id":433050208,"identity":"f10016e7-a54e-4800-883a-f1796b6bdf50","order_by":2,"name":"Gideon Mwendwa","email":"","orcid":"","institution":"National Forensic Sciences University","correspondingAuthor":false,"prefix":"","firstName":"Gideon","middleName":"","lastName":"Mwendwa","suffix":""},{"id":433050209,"identity":"f037ce28-8b11-4b19-b921-5cbf991183fa","order_by":3,"name":"Sheenamol Yoosaf","email":"","orcid":"","institution":"Cochin University of Science and Technology","correspondingAuthor":false,"prefix":"","firstName":"Sheenamol","middleName":"","lastName":"Yoosaf","suffix":""},{"id":433050210,"identity":"fd57b266-138b-467e-bbb3-4401288d1183","order_by":4,"name":"Ngaira Mandela","email":"","orcid":"","institution":"National Forensic Sciences University","correspondingAuthor":false,"prefix":"","firstName":"Ngaira","middleName":"","lastName":"Mandela","suffix":""}],"badges":[],"createdAt":"2024-12-19 13:38:31","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5677244/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5677244/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":79171110,"identity":"af2b873f-da7a-407a-b5dc-839cefec6a8e","added_by":"auto","created_at":"2025-03-25 09:24:28","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":766035,"visible":true,"origin":"","legend":"","description":"","filename":"Comparativeanalysis.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5677244/v1_covered_b6297409-a6dc-48a0-a27b-50473908cb30.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Performance Analysis of Machine Learning Models with Optimized Feature Selection Techniques for Darknet Traffic Classification","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"discover-computing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Discover Computing](https://link.springer.com/journal/10791)","snPcode":"10791","submissionUrl":"https://submission.springernature.com/new-submission/10791/3","title":"Discover Computing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Feature Selection, Lasso, RFE, Darknet Traffic Classification, Machine Learning, Boruta","lastPublishedDoi":"10.21203/rs.3.rs-5677244/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5677244/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDarknet traffic classification is a crucial area of cybersecurity, targeting the anonymity of network activities within an anonymized network. This study evaluates the efficacy of three feature selection methods, Boruta, Recursive Feature Elimination (RFE), and Lasso, across eight machine learning models, including Random Forest, Gradient Boosting, Neural Networks, and SVM. The research enhances computational efficiency and model accuracy through rigorous data preprocessing, such as encoding, normalization, and downsampling. Boruta emerged as the most effective method, with Random Forest achieving 98.7% accuracy and 99.66% ROC-AUC, showcasing its potential for high-accuracy applications. In contrast, Lasso excelled with simpler models like Naive Bayes, optimizing it significantly, while RFE was noted for the fastest training times, which is ideal for resource-constrained environments. This comprehensive performance analysis underscores the critical role of appropriate feature selection in darknet traffic classification. It sets the stage for future research to explore deeper, more complex machine learning frameworks. The findings illustrate significant trade-offs between computational efficiency and model performance, guiding the selection of feature selection techniques based on specific application need.\u003c/p\u003e","manuscriptTitle":"Performance Analysis of Machine Learning Models with Optimized Feature Selection Techniques for Darknet Traffic Classification","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-03-25 08:52:18","doi":"10.21203/rs.3.rs-5677244/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-04-03T11:41:25+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-03-30T13:46:01+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"317573256806288356300007815959753168261","date":"2025-03-26T04:59:18+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-03-25T20:58:08+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"197736592628077481684094723937919804259","date":"2025-03-24T09:15:42+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-03-24T09:09:13+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-03-22T08:00:41+00:00","index":"","fulltext":""},{"type":"submitted","content":"Discover Computing","date":"2025-03-21T19:38:18+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"discover-computing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Discover Computing](https://link.springer.com/journal/10791)","snPcode":"10791","submissionUrl":"https://submission.springernature.com/new-submission/10791/3","title":"Discover Computing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Discover Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"dd7e1974-5014-48e7-a50c-13ecea71b4bc","owner":[],"postedDate":"March 25th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2025-06-13T07:23:29+00:00","versionOfRecord":[],"versionCreatedAt":"2025-03-25 08:52:18","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5677244","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5677244","identity":"rs-5677244","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-22T02:00:06.705733+00:00
License: CC-BY-4.0