Behavioral Feature Filtering for Source-Bias Mitigation in Hybrid Real-Simulated IoT Intrusion Detection Datasets

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract The proliferation of Internet of Things (IoT) devices has created an urgent need for effective intrusion detection, yet the lack of real attack traffic forces researchers to combine real benign data with simulated attacks. Such hybrid datasets suffer from a critical but understudied flaw: machine learning models learn to distinguish real from simulated traffic rather than benign from malicious behavior. This paper introduces a reproducible three-phase methodology to systematically diagnose, quantify, and mitigate this source bias. First, a controlled label-flipping experiment reveals that classifiers achieve 100% accuracy by learning data origin alone, formally characterizing source bias. Second, a novel behavioral feature filtering framework combining Kolmogorov–Smirnov testing (D KS < 0.3), Cohen’s d effect size (|d| 0.01 bits) reduces the feature set from 62 to 36 source-agnostic features, dropping source-prediction accuracy from 100% to 67.3% (only 1.8 percentage points above the 65.5% majority-class baseline). Third, we contribute a new validated hybrid dataset of 115,707 samples: 39,808 real benign flows from a 47-device, 119-day testbed plus 75,899 NS-3 simulated attacks across 12 categories. A gradient boosting ensemble (LightGBM, XGBoost, Cat-Boost) achieves binary weighted F1 = 99.41%, AUROC = 99.97%, and MCC = 0.9870. Fourth, zero-shot evaluation on four independent benchmarks containing real (non-NS-3) attack traffic yields an average weighted F1 of 0.933 (generalization gap = 6.1%), substantially improving over prior gaps of 13–40% without any domain adaptation; ablation confirms that behavioral filtering accounts for approximately 86% of this gain. The proposed methodology generalizes beyond intrusion detection to any hybrid real-simulated dataset where source origin correlates with class labels, and all code, NS-3 scripts, and the filtered dataset are released for reproducibility.
Full text 12,900 characters · extracted from preprint-html · click to expand
Behavioral Feature Filtering for Source-Bias Mitigation in Hybrid Real-Simulated IoT Intrusion Detection Datasets | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Behavioral Feature Filtering for Source-Bias Mitigation in Hybrid Real-Simulated IoT Intrusion Detection Datasets Hidsi Elmahdi, Mouncif Hicham This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9395215/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 5 You are reading this latest preprint version Abstract The proliferation of Internet of Things (IoT) devices has created an urgent need for effective intrusion detection, yet the lack of real attack traffic forces researchers to combine real benign data with simulated attacks. Such hybrid datasets suffer from a critical but understudied flaw: machine learning models learn to distinguish real from simulated traffic rather than benign from malicious behavior. This paper introduces a reproducible three-phase methodology to systematically diagnose, quantify, and mitigate this source bias. First, a controlled label-flipping experiment reveals that classifiers achieve 100% accuracy by learning data origin alone, formally characterizing source bias. Second, a novel behavioral feature filtering framework combining Kolmogorov–Smirnov testing (D KS < 0.3), Cohen’s d effect size (|d| 0.01 bits) reduces the feature set from 62 to 36 source-agnostic features, dropping source-prediction accuracy from 100% to 67.3% (only 1.8 percentage points above the 65.5% majority-class baseline). Third, we contribute a new validated hybrid dataset of 115,707 samples: 39,808 real benign flows from a 47-device, 119-day testbed plus 75,899 NS-3 simulated attacks across 12 categories. A gradient boosting ensemble (LightGBM, XGBoost, Cat-Boost) achieves binary weighted F1 = 99.41%, AUROC = 99.97%, and MCC = 0.9870. Fourth, zero-shot evaluation on four independent benchmarks containing real (non-NS-3) attack traffic yields an average weighted F1 of 0.933 (generalization gap = 6.1%), substantially improving over prior gaps of 13–40% without any domain adaptation; ablation confirms that behavioral filtering accounts for approximately 86% of this gain. The proposed methodology generalizes beyond intrusion detection to any hybrid real-simulated dataset where source origin correlates with class labels, and all code, NS-3 scripts, and the filtered dataset are released for reproducibility. Intrusion Detection Systems Internet of Things Hybrid Datasets Feature Engineering Ensemble Learning Cross-dataset Generalization Full Text Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Reviewers agreed at journal 20 May, 2026 Reviewers invited by journal 14 May, 2026 Editor assigned by journal 14 May, 2026 Submission checks completed at journal 17 Apr, 2026 First submitted to journal 12 Apr, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9395215","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":640719110,"identity":"2824ff33-0caa-4d4b-bdac-23c723f9f508","order_by":0,"name":"Hidsi Elmahdi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABA0lEQVRIie3PMWsCMRTA8XcEzqEJrpGD9hMUIsK1InJfJaFD106HgxRBOBela4fSzyAEMh886FR0PcjUb3Cly4FLIyjYIfVGh/zJkCE/Xh5AKHSJdQAIxBwEAQFy0oaQP+SzPQH33p2oaCG65Er9POX3cNchuv56x+z2pYzqBsyNj/TmTCevGw7DeZxzZVCZSpLeEmx/5iEC2TphhdsFaQqOyLQCSABs5CMZMr07kEGt3jBLtyXZOZJ5pxBmjlMEVzOMTCnj/RTlIxyZGdENpwLdLvLj0e2iiuFS2Acf6S5W2tL8+VpsUX8305H7GGLVTOzYR47Rk/t+cXHmfSgUCoX+7ReEA1JEOjNN8wAAAABJRU5ErkJggg==","orcid":"","institution":"Sultan Moulay Slimane University","correspondingAuthor":true,"prefix":"","firstName":"Hidsi","middleName":"","lastName":"Elmahdi","suffix":""},{"id":640719111,"identity":"a97682a7-2906-49b6-ade5-503f0aef584b","order_by":1,"name":"Mouncif Hicham","email":"","orcid":"","institution":"Sultan Moulay Slimane University","correspondingAuthor":false,"prefix":"","firstName":"Mouncif","middleName":"","lastName":"Hicham","suffix":""}],"badges":[],"createdAt":"2026-04-12 15:08:22","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9395215/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9395215/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":109433727,"identity":"e08236ec-58a2-4a77-928f-27bf526ac321","added_by":"auto","created_at":"2026-05-18 05:40:48","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1686402,"visible":true,"origin":"","legend":"","description":"","filename":"HybridRealSimulatedIoTIDSDataset.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9395215/v1_covered_23708d71-ef9c-4705-95ec-67f616626ebc.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Behavioral Feature Filtering for Source-Bias Mitigation in Hybrid Real-Simulated IoT Intrusion Detection Datasets","fulltext":[],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":true,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"cluster-computing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Cluster Computing](https://www.springer.com/journal/10586)","snPcode":"10586","submissionUrl":"https://submission.nature.com/new-submission/10586/3","title":"Cluster Computing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Intrusion Detection Systems, Internet of Things, Hybrid Datasets, Feature Engineering, Ensemble Learning, Cross-dataset Generalization","lastPublishedDoi":"10.21203/rs.3.rs-9395215/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9395215/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"The proliferation of Internet of Things (IoT) devices has created an urgent need for effective intrusion detection, yet the lack of real attack traffic forces researchers to combine real benign data with simulated attacks. Such hybrid datasets suffer from a critical but understudied flaw: machine learning models learn to distinguish real from simulated traffic rather than benign from malicious behavior. This paper introduces a reproducible three-phase methodology to systematically diagnose, quantify, and mitigate this source bias. First, a controlled label-flipping experiment reveals that classifiers achieve 100% accuracy by learning data origin alone, formally characterizing source bias. Second, a novel behavioral feature filtering framework combining Kolmogorov–Smirnov testing (D KS \u003c 0.3), Cohen’s d effect size (|d| \u003c 0.5), and mutual information (I \u003e 0.01 bits) reduces the feature set from 62 to 36 source-agnostic features, dropping source-prediction accuracy from 100% to 67.3% (only 1.8 percentage points above the 65.5% majority-class baseline). Third, we contribute a new validated hybrid dataset of 115,707 samples: 39,808 real benign flows from a 47-device, 119-day testbed plus 75,899 NS-3 simulated attacks across 12 categories. A gradient boosting ensemble (LightGBM, XGBoost, Cat-Boost) achieves binary weighted F1 = 99.41%, AUROC = 99.97%, and MCC = 0.9870. Fourth, zero-shot evaluation on four independent benchmarks containing real (non-NS-3) attack traffic yields an average weighted F1 of 0.933 (generalization gap = 6.1%), substantially improving over prior gaps of 13–40% without any domain adaptation; ablation confirms that behavioral filtering accounts for approximately 86% of this gain. The proposed methodology generalizes beyond intrusion detection to any hybrid real-simulated dataset where source origin correlates with class labels, and all code, NS-3 scripts, and the filtered dataset are released for reproducibility.","manuscriptTitle":"Behavioral Feature Filtering for Source-Bias Mitigation in Hybrid Real-Simulated IoT Intrusion Detection Datasets","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-18 05:39:52","doi":"10.21203/rs.3.rs-9395215/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"292231589350198533403577973316658060087","date":"2026-05-20T04:51:19+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-05-14T11:49:38+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-05-14T11:44:39+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-04-17T13:53:26+00:00","index":"","fulltext":""},{"type":"submitted","content":"Cluster Computing","date":"2026-04-12T15:02:58+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"cluster-computing","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"Learn more about [Cluster Computing](https://www.springer.com/journal/10586)","snPcode":"10586","submissionUrl":"https://submission.nature.com/new-submission/10586/3","title":"Cluster Computing","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"a32bb6e0-01f9-4dfe-b6a4-12ff9a95e67c","owner":[],"postedDate":"May 18th, 2026","published":true,"recentEditorialEvents":[{"type":"reviewerAgreed","content":"292231589350198533403577973316658060087","date":"2026-05-20T04:51:19+00:00","index":20,"fulltext":""},{"type":"reviewersInvited","content":"14","date":"2026-05-14T11:49:38+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-05-14T11:44:39+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-18T05:39:52+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-18 05:39:52","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9395215","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9395215","identity":"rs-9395215","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00