A comprehensive review of machine learning applications in cybersecurity: identifying gaps and advocating for cybersecurity auditing | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A comprehensive review of machine learning applications in cybersecurity: identifying gaps and advocating for cybersecurity auditing Ndaedzo Rananga, H. S. Venter This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4791216/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 11 You are reading this latest preprint version Abstract Cybersecurity threats present significant challenges in the ever-evolving landscape of information and communication technology (ICT). As a practical approach to counter these evolving threats, corporations invest in various measures, including adopting cybersecurity standards, enhancing controls, and leveraging modern cybersecurity tools. Exponential development is established using machine learning and artificial intelligence within the computing domain. Cybersecurity tools also capitalize on these advancements, employing machine learning to direct complex and sophisticated cyberthreats. While incorporating machine learning into cybersecurity is still in its preliminary stages, continuous state-of-the-art analysis is necessary to assess its feasibility and applicability in combating modern cyberthreats. The challenge remains in the relative immaturity of implementing machine learning in cybersecurity, necessitating further research, as emphasized in this study. This study used the preferred reporting items for systematic reviews and meta-analysis (PRISMA) methodology as a scientific approach to reviewing recent literature on the applicability and feasibility of machine learning implementation in cybersecurity. This study presents the inadequacies of the research field. Finally, the directions for machine learning implementation in cybersecurity are depicted owing to the present study’s systematic review. This study functions as a foundational baseline from which rigorous machine-learning models and frameworks for cybersecurity can be constructed or improved. cybersecurity auditing cybersecurity threats artificial intelligence cybersecurity tools modern and sophisticated cyberthreats machine learning state-of-the-art Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1 Introduction Observation confirms that a significant portion of personal, government, and corporate communication heavily depends on digital interactions within cyberspace. According to Aljabri et al. [1], the primary revenue generation source in the business and market relies profoundly on the appreciation of modern technology and the use of advanced business models. Modern advancements and their evolution should be applauded because they enhance the quality of daily life while improving corporate productivity. Inevitably, like other technological innovations or advancements, cyberspace is prone to a wide range of challenges. One of the frontrunners of the difficulties associated with cyberspace is the exponential development of more complex, advanced persistent threats (APTs) and sophisticated cyberthreats. Cremer et al. [2] emphasize the effect of cyberthreats, which can escalate to billions of dollars in total costs to the global economy. Considerable work has been conducted to combat the ever-growing cyberthreats associated with modern technologies. Corporations are implementing several resources to respond to evolving cyberthreats substantially. These resources include adopting security best practices and cybersecurity standards and using modern technologies and tools. Generally, cybersecurity can be defined as a collective effort of procedures, people’s actions, and technologies to counter cyberthreats and, therefore, protect valuable assets [3]. With artificial intelligence (AI), such as machine learning (ML), gaining ground in the current and future of modern technologies, cybersecurity tools are also exploiting ML to rejuvenate or revamp cyber capabilities against complex and sophisticated cyberthreats. Apruzzese et al.[4] emphasize that ML deployment in cybersecurity remains in its infancy, with gaps between the theories and the actual applicability of ML in cybersecurity practices. A need exists to perform a state-of-the-art review on ML applicability in modern cybersecurity capabilities using input from recent literature. Conversely, the problem is that due to the immaturity of implementing ML in the cybersecurity space, more research is needed to adopt ML applications within this space, therefore the need for this study. The study leverages the wide range of literature to perform such a state-of-the-art analysis or a systematic review to provide future directions on using ML in cybersecurity. As emphasized by Macs et al. [5], a wide range of learning, challenges, and opportunities persist for ML regarding cybersecurity. The study systematically reviewed the developments of ML in the cybersecurity space. The shortcomings of the literature and the future indicators of using ML in the cybersecurity space are presented. The remainder of this study is organized as follows: Section 2 is a background section summarizing the fundamental theoretical concepts used throughout this paper. This includes the ML concept and the background of cybersecurity. The methodology used to select reviewed studies from the literature is presented in Section 3 . Section 4 presents a detailed description of the reviewed literature. This section expands the current research topic by reviewing further studies at a prominent level. This includes the results and the discussion of the reviewed studies. In Section 5, gaps identified from the reviewed work are presented. Due to the nature of the study, only a specific scope is included, and other areas are left for potential future work, as presented in Section 6 . Before the study is concluded, a critical evaluation of the study is presented in Section 7. Section 8 concludes the study. After briefly introducing the present study, the subsequent section presents the fundamental background of terms used throughout this research. 2 Background In recent years, a subtle transformation has occurred in how users access, store, and communicate information. ML and cybersecurity are prominent in computing research and development. Before exploring the study details, the subsections concisely introduce ML and cybersecurity. 2.1 Machine learning The literature provides multiple definitions of ML, reflecting the preferences of diverse authors. These definitions do not directly contradict one another. This study succinctly defines ML as a computational approach that leverages data to train computers for tasks, minimizing or eliminating human intervention. As described by Dasgupta et al. [ 6 ], ML methods and algorithms rely on automated approaches to imitate human learning and improve efficiency while improving accuracy. The main objective of ML is to develop intelligence applications and systems able to learn how to perform tasks without continuous explicit programming and human intervention [ 7 ]. ML can be classified into three main categories: supervised learning, unsupervised learning, and reinforcement learning [ 3 ], described subsequently. 2.1.1 Supervised learning Supervised machine learning revolves around using a known dataset to train algorithms for predicting potential outcomes. In this approach, predefined functions leverage input data to estimate the corresponding output values. [ 8 ]. Supervised learning directs classifications and regression-type problems [ 7 ]. ML is problem-oriented; therefore, no universal rules exist regarding selecting ML techniques. According to Liu and Lang [ 9 ], supervised ML algorithms generally perform better than unsupervised ML algorithms. Their argument [ 9 ] emphasizes that unsupervised learning relies on labeling data before training a model; however, this process can be resource-demanding and complex for unstructured data, such as extracting data from traffic flow to perform the ML algorithm. In cybersecurity, such as in Anthi et al. [ 10 ], supervised classification algorithms can detect cyberthreats using a malicious data set as the input sample data set. Appreciating and incorporating ML in modern cybersecurity tools, approaches, and techniques are influenced by the need to rejuvenate cybersecurity capabilities to combat complex and sophisticated cyberattacks [ 11 ]. ML demonstrated remarkably accurate results [ 12 ], which can ultimately lead to increased efficiency and cost savings within the cybersecurity community. By reducing the time spent pursuing false-positive alerts, ML contributes significantly to improving effectiveness. 2.1.2 Unsupervised learning Unlike supervised learning, unsupervised ML uses algorithms to analyze unlabeled datasets to predict outputs. Unsupervised ML can identify and characterize an unlabeled dataset during model training [ 7 ]. Unlike supervised learning, unsupervised machine learning operates without human feedback or labeled data. It focuses on identifying patterns and structures within unlabeled datasets [ 3 ]. 2.1.3 Reinforcement learning As aforementioned, ML approaches are techniques trained to imitate human behavior and perform future tasks without human interventions, specifically in unsupervised learning. Reinforcement learning is inspired by psychological theories that reward anticipated behaviors and penalize unexpected actions. In this paradigm, AI-based systems learn through trial and error, adjusting their actions based on the outcomes they receive —rewards or punishments [ 3 ]. After a brief presentation of diverse categories of ML, owing to the approach by Sjarif et al. [ 8 ], Table 1 presents a summary of the categories. Table 1 Machine learning categories description summary Category Description Supervised learning These ML algorithms are trained to predict output using the provided input. For example, a model can be trained to label and classify malicious executables using an extract of the dataset from endpoint detection and response (EDR) or antivirus databases. Unsupervised learning These algorithms can discover patterns from unlabeled datasets without explicit human interventions. For example, a model can be trained to detect real-time network activity anomalies [ 13 ] based on behavioral attributes. This approach to cybersecurity can also be useful in combating zero-day cyberattacks. Reinforcement learning This ML model is based on reward and punishment, where applications are trained to make decisions that will reward them as the output. For example, an EDR ML-based application can be trained to isolate hosts from the network because of the detection of potential malicious behavior. The wide range of research to combat cyberthreats must be acknowledged; however, continual advancements are needed to enhance real-time cyber detection and prevention [ 12 ]. With the constant challenges of more sophisticated and complex cyberthreats, ML is a promising technique or approach for adequate resources to combat cyberthreats through more automated detections [ 14 ]. The subsection further elaborates on this concept. 2.2 Machine learning application in cybersecurity Sahin et al. [ 15 ] suggest that security vulnerabilities could exist in all technology, including modern advancements; therefore, the need to reinforce cybersecurity is becoming a primary concern. Cybersecurity is a concept of defined processes, technologies, tools, and people to counter the ever-evolving cyberthreats. Cybersecurity should protect assets, such as computer networks, users, programs, and data, from malicious actors and unauthorized access [ 5 ]. The rapid adoption of ML algorithms, such as support vector machines (SVM), deep neural networks (DNN), and reinforcement learning (RL), led to a transition to modern approaches empowered by ML from traditional cybersecurity mechanisms to modern ML-empowered approaches [ 16 ]. Comprehensive ML applications have been confirmed in cyberthreats. Simmons et al. [ 17 ] proposed a cyberattack taxonomy, such as AVOIDIT, helpful in classifying and categorizing diverse attacks based on their characteristics, methods, and effects, as presented in Fig. 1 . Figure 1 presents a cyberattack taxonomy, grouping common attacks based on their operational effect. These operational effects, as demonstrated in the taxonomy tree, include misuse of information, user and web compromise, installed malware, and service denial. Incorporating ML in these operational effects is evident in the modern cybersecurity realm. Computational intelligence is widely used in developing rigorous cyber prevention approaches toward this operational effect [ 18 ]. The subsections present some shared areas that appreciate using ML to compact these operational consequences. 2.2.1 Machine learning in misuse of (information) resources Technology simplifies tasks for end users, enabling them to easily share information resources while keeping them informed with the latest updates; however, like any technological advancement, the significant transition from traditional resources of sharing information using studies to using digital technology to share information is associated with challenges. One of the visible challenges within the modern information resources era is news disinformation . De Wet and Marivate [ 19 ] suggest that discerning legitimate shared information resources from fake ones is increasingly difficult for humans. In a quest to combat this challenge, ML adoption is vital. According to Huang [ 20 ], ML is adopted as a feasible countermeasure to information misuse, such as fake news. Misuse of information resources should be treated as a serious matter because it can range to the corporate level. For example, through social media, individuals can share fake news regarding supply chain tender opportunities and the operation of a company, which could have devastating effects on an organization. Akhtar et al. [ 21 ] realized the importance of incorporating ML to counter fake news at the corporate level. They proposed an ML model to detect fake news and disinformation to avoid supply chain disruptions. 2.2.2 Machine learning in user compromise Technology is becoming more accessible to most end users. According to Rananga and Venter [ 21 ], adopting technologies, such as mobile computing, is growing exponentially. In the modern days of digitization and cyberspace, cyberthreats are still increasing [ 8 ]; however, cybersecurity tools are becoming more aware and can remediate some common cyberattacks. As such, attackers consistently target users, considered the most vulnerable asset within the computing realm. Cyberattacks targeting human weaknesses, such as social engineering, are the easiest to initiate and have a high success probability rate [ 22 ]. Analogously, a need exists for more advanced techniques to protect users from social engineering-related cyberthreats. Several research projects explore automating social engineering attacks using ML methods [ 22 ]. 2.2.3 Machine learning in web compromise Web applications and services are increasingly accessible, attributable to technological advancements; however, these developments remain vulnerable to sophisticated and targeted cyberattacks, including ATP attacks. Web attacks include web defacement, where unauthorized users access the website, edit, delete, or replace the original content without the web owner’s consent. Although some attacks, such as defacement, do not have direct damage nor permanent damage on the web servers, as indicated by Maras [ 22 ], countering such attacks is necessary as they affect the integrity of website content. Besides defacement, the web is associated with several other attacks, such as SQL injection and DDoS. According to Kiruba et al. [ 23 ], most web-based attacks are initiated through the top 10 open worldwide application security project (OWASP) web vulnerabilities, such as injection, broken authentication, sensitive data disclosure, cross-site scripting (XSS), external entity (XSE), broken access control, insecure deserialization, insufficient logging and monitoring, and security misconfiguration. A wide range of ML adoption exists to improve cyber controls in web applications and services. Kaur et al. [ 24 ] proposed an ML approach to detect the XSS attacks; the results of the ML model were satisfactory. 2.2.4 Machine learning in installed malware The definition, type, and classification of malware are associated with the harm they cause [ 7 ]. Malware refers to malicious executable code that disrupts the normal functioning of a system or application. Common types of malware include viruses, worms, trojan horses, ransomware, adware, and spyware [ 19 ]. The distinct description of the malware does not form part of the present study; however, as aforementioned, the malware names conform to the harm they cause. For example, ransomware is a type of malware that, when successfully executed on the target machine or host, encrypts the data until a ransom payment is made. Malware detection techniques, such as static, dynamic, and hybrid analysis, have been proven to be effortlessly defeated by evasion techniques, such as obfuscation [ 5 ]; therefore, unlike in the past, when malware detection relied primarily on static and dynamic analysis of malware architecture to understand behavioral patterns, modern tools now leverage ML. These advanced ML techniques enable more efficient and accurate malware detection [ 20 ]. ML offers numerous opportunities for malware detection, including identifying malware, targeting smartphones, particularly those running on Android operating systems (OSs) [ 7 ]. 2.2.5 Machine learning in denial of services According to Wanjau et al. [ 25 ], one of the most predominant cyberattacks encountered in network security is brute force. Brute-force attacks apply the trial-and-error method to guess login credentials, web browser’s directory, and encryption keys. The most common brute-force attack may be an attempt by an intruder or an attacker to acquire unauthorized access to the system using guessing techniques to bypass the authentication controls [ 5 ]. In addition to brute-force attacks, DDoS is a common attack within cybersecurity. It is a cyberattack that disrupts the availability of services and applications to intended users. DDoS attacks typically involve consistently and repeatedly transmitting numerous requests to a target host. The goal is to overload or even exhaust the processing resources of the target host [ 26 ], causing a severely disrupted service. One of the critical factors influencing technology adoption and appreciation is its ease of use and availability [ 27 ]; therefore, countering cyberattacks, such as DDoS, that threaten the availability of services to end users is vital. Given the necessity of incorporating ML in cybersecurity, it is unsurprising that researchers are developing modern ML algorithms to combat the emergence of brute force and DDoS attacks. According to He et al. [ 28 ], ML can leverage detecting DDoS by preventing network packets from being distributed to external networks. Figure 1 presents the cyberattack taxonomy, including the remediation and mitigation mechanisms. According to Jayalaxmi et al. [ 29 ], intrusion detection systems (IDSs) and intrusion prevention systems (IPSs) are essential when developing robust cybersecurity controls. Conversely, EDR is also vital in endpoint cybersecurity controls [ 30 ]. After emphasizing some of the applications of ML in directing various cyberthreats, it is crucial also to discuss how ML is being utilized in intrusion detection systems (IDSs), intrusion prevention systems (IPSs), and endpoint detection and response (EDR) solutions. This section describes these common countermeasures to cyberthreats and explains how ML has been integrated to enhance their effectiveness. 2.2.6 Machine learning in network intrusion detection systems (IDSs) This study emphasizes that the rapid development in adopting modern technologies not only leverages daily operations but also exposes considerable security issues [ 31 ] while inviting more sophisticated and complex cyberthreats. One of the most feasible responses to combat the ever-evolving cyberthreats is using IDSs. One of the advantages of Intrusion Detection Systems (IDSs) is their ability to manage suspicious activities within the network without disrupting other operations within the network ecosystem [ 29 ]. With ML integration in modern technologies, it is unsurprising that more IDSs are being established using ML algorithms. These algorithms enhance efficiency by analyzing vast amounts of data and making detections faster than humans while also improving detection accuracy [ 31 ]. Liu and Lang [ 9 ] remark that ML approaches are also more vital to equipping the IDSs to detect unknown attacks. Rose et al. [ 32 ] illustrate how ML can be used to detect IoT-related cyberthreats. The authors emphasize that ML models can also process false-positive alarms, which may help cybersecurity defenses mitigate cyberthreats more effectively. 2.2.7 Machine learning in network intrusion prevention systems (IPSs) In addition to the extensive use of IDSs for detecting anomalies within networks, IPSs provide an additional layer of cybersecurity capabilities. IPSs are tools or techniques used for continuous network monitoring to detect suspicious activities and act upon detection. IPSs are considered a proactive approach to effective security threat prevention by blocking traffic from the offending host [ 29 ]. As specified by Chandre et al. [ 33 ], the primary distinction between an IDS and an IPS is that an IPS can take proactive measures to stop potential attacks. Conversely, a traditional IDS is limited to detecting attacks and alerting them to their presence without actively attempting to prevent them. Modern IPSs incorporate machine learning methods for real-time detection and prevention of cyberattacks with high accuracy [ 12 ]. An IPS can commonly act after detecting suspicious activity, including isolating the host from the rest of the network ecosystem. 2.2.8 Machine learning in endpoint detection and response (EDR) As more sophisticated and complex threats arise—usually real-time-based attacks—endpoints are no exception in these modern attacks. Endpoints are physical or virtual devices that exchange information and communicate with each other, employing computer networking. These devices, such as desktops, laptops, and mobile devices [ 8 ], are meant for end users’ applications. As specified by Xuan and Huong [ 34 ], one of the front runners of modern APTs is initiated by spreading malware through endpoints. Owning to the complexity of modern APT targeting endpoints, a need exists to rejuvenate EDR for adequate resources to protect endpoints. Adopting ML is acquiring popularity in the EDR realm of cybersecurity [ 8 ], like in Kaur and Tiwari [ 11 ], where the authors demonstrate how ML behavior and patterns can be used to characterize incidents targeting endpoints. The authors believe that rapidly incorporating ML to reinforce the endpoints’ security will continue to acquire ground in modern computing. The fundamental background of applying ML in the cybersecurity area is presented. As emphasized by Roest et al. [ 35 ], it is essential to perform an exhaustive systematic search to identify recent literature related to the study of interest. In the subsequent section, the state-of-the-art literature review of ML in cybersecurity is presented; however, before such a review is provided, a methodology adopted in the present study is presented first. 3 Review methodology Adopting online and digital library databases has resulted in a vast array of literature. Consequently, it is essential to define the selection criteria for choosing studies from the extensive range of available literature while excluding others. The following subsections present the review criteria adopted in the current research, along with a review included in this study. This study used a wide range of literature relevant to the application of machine learning in cybersecurity. Building upon existing literature without disregarding prior work, this study outlines the eligibility criteria for filtering literature. 3.1 Eligibility criteria The online libraries are chosen based on the PRISMA [ 35 ]. The PRISMA flow is a guideline on effective resources for selecting literature for a systematic review. In the present study, the included libraries are used owing to their popularity, the vast number of articles, and ease of access. The libraries used in the present study include IEEE Xplore, ScienceDirect, Google Scholar, Springer Link, and Wiley Online Library. Publications from these online libraries that were not older than five years at the time of this study were considered. 3.2 Search strategy and sources As adopted by Sibiya et al. [ 36 ], search indexing using key phrases can filter results from diverse digital libraries. In our case, however, we adopted a similar strategy, considering the availability of databases as presented earlier in the eligibility criteria section. The keywords used to search for studies relevant to the current research include ((“machine learning” OR “Advanced Threat Persistence” OR “cybersecurity” OR “modern cybersecurity EDR” OR “modern cybersecurity IDS/IPS” OR “Prisma statement in machine learning reviews”) AND “machine learning application in cybersecurity”). 3.2.1 Screening process and study selection After presenting the selection criteria adopted in the present study, the PRISMA flow diagram is presented in Fig. 1 . Traditionally, the PRISMA statement was intended to evaluate studies under health interventions; however, as suggested by Page et al. [ 37 ], the PRISMA approach applies to several other systematic review studies, which is why it is adopted in the current study. The PRISMA flow diagram can be divided into: Identifying the literature Screening of the studies The indication of the included literature reviewed and presented in the systematic review [ 35 ]. According to Fig. 1 , the PRISMA approach indicates how the authors selected the studies for the systematic review. As demonstrated by Hedley et al. [ 38 ], the abstract and title of the studies are used during the screening process, while only the studies the authors believe relate to the current research are incorporated. This study narrowed the scope of applying ML in cybersecurity, opportunities in cybersecurity, challenges of using ML in cybersecurity, and the future of ML in cybersecurity. As presented in Fig. 1 , after an exhaustive search of the digital libraries, 123 studies were identified. Twenty-three studies were omitted because they were duplicates. After the removal of duplicates, 100 were screened using the title and abstract. After the screening number, 81 articles were removed with confirmed reasons. Finally, 19 articles were reviewed and included in the present systematic review. Additional studies were excluded from the systematic review; however, they are still used in the literature and were considered applicable. The previous section provided the reader with the selection and filtering process to arrive at the selected studies reviewed. The subsequent section presents the evaluation criteria applied during each study’s comparison and critical analysis. 3.2.2 Evaluation criteria As demonstrated by Gupta et al. [ 3 ], the present study adopted some evaluation criteria in addition to the eligibility criteria, including the search strategy and sources. Several factors can be used to compare dissimilar ML techniques in cybersecurity. In Gupta et al. [ 3 ] in comparing various studies, the authors use these factors: “ Year, Authors, Algorithm used, Source of data, and Advantages and Limitations ” to perform a systematic review. Conversely, as presented in the previous section, Simmons et al. [ 17 ] proposed a cyberattack taxonomy that entails classification by: Attack vector Operational effect Defense Informational effect Attack target Emanating from those studies, the current research adopted similar factors presented by Gupta et al. [ 5 ] and Simmons et al. [ 17 ]; however, adapting them to be more relevant, as presented in the subsequent section. According to Macas et al. [ 5 ], a wide range of research explores ML applicability and AI-related techniques to combat the rapid development of cyberthreats. The subsections present available and related frameworks and models for this systematic review. 4 Systematic review Gupta et al. [ 3 ] remark that various ML approaches can secure ICT against cybersecurity threats. The background entailed throughout this study and the selection criteria are presented. The subsection presents the results from the reviewed studies. 4.1 Results As emphasized in the previous section, the choice of ML algorithm is problem-oriented; therefore, various problem scenarios can be solicited using diverse ML algorithms. As was evident in Shahrivari et al. [ 40 ], alternative ML classification methods were used in the processes of training data, and the accuracy results differed. The reviewed work indicates no specific ML algorithm that could solve the cybersecurity challenges. The challenges of consistently finding the best tools and techniques to combat particular challenges were also emphasized by L’Heureux et al. [ 39 ]. Figure 3 presents various ML algorithms that can combat advanced, and complex cyberattacks. These algorithms include random forest (RF), SVM, K-nearest neighbor (KNN), decision tree (DT), left-right parsing (LR), and nearest neighbor (NN). From Fig. 3 , the most used ML algorithm in cybersecurity, according to the literature, is RF. Nerurkar et al. [ 50 ] confirmed that RF is the most used ML algorithm in the literature. The dominant use of RF in ML is not a surprise because cybersecurity is a complex environment that changes more rapidly. RF can manage imbalanced data and non-linear relationships, as is consistently the case in cybersecurity datasets. For example, attacks consistently use more sophisticated resources to bypass security controls; therefore, using an ML algorithm that can combine multiple decisions to enable thorough detection is vital. The second most used ML algorithm in cybersecurity in the literature is SVM. SVM is believed to be more effective in instruction detection and malware classification [ 40 ]. The other aspect that makes the SVM the most commonly used algorithm in cybersecurity is that it can handle high-dimensional and complex data [ 41 ]. For example, in the case of Xia et al. [ 16 ], a case scenario application can be vital to supporting a study’s outcome. To further understand the reviewed studies from the literature, the authors analyzed how the research aligns with the category from the cyberattacks taxonomy presented in the previous section. This includes applying ML in web compromise, DoS, user compromise, misuse of information, installed malware, NIPS, network instruction detection systems (NIDS), and EDR, as presented in Fig. 4 . As depicted in Fig. 4 , ML applications in NIDS constitute the highest percentage, with an overall of 31.6%. ML applications in NIDS include using ML algorithms to detect patterns of unusual activities in the networking environment. Mohamed [ 42 ] confirms that using ML in NIDS is the most common ML application in the quest to combat cyberthreats. Pérez et al. [ 43 ] emphasize that NIDSs are one of the tools in cybersecurity; therefore, more ML algorithms are being developed in this area. Like in the case of Haddadpajouh et al. [ 44 ], one of the customary factors that can be used to assess the ML algorithm is performance and accuracy. From the reviewed studies, the algorithm’s accuracy was satisfactory, ranging from at least 80–99.9%. One of the most challenging aspects of ML is acquiring the correct data from the data sources. Usually, researchers rely on historical data to develop and assess models. Figure 5 presents some of the common data sources for ML applications in cybersecurity. From the reviewed studies, the most common data source for cybersecurity ML applications is the CICIDS2017 dataset developed by the Canadian Institute for Cybersecurity (CIC). Yang and Shami [ 45 ] suggest that the CICIDS2017 dataset is commonly used because it is vital in developing ML models to combat various types of cyberattacks, such as DoS attacks, port-scan attacks, brute-force attacks, web-based attacks, botnet attacks, and infiltration attacks. After presenting details of the reviewed studies and the most prominent data sources, the authors present the summary of the reviewed studies in a tabular format (Table 2 ). The shortcomings of each reviewed research study are also provided in the table. To expand on the idea behind the presented table, the shortcomings identified from the reviewed studies are categorized into subcategories (Table 2 ). If the deficiencies apply to a specific study, a tick (🗸) under the applicable category is used; otherwise, a blank space indicates that the shortcomings do not apply to a particular study. For example, the study entitled “Adversarial Attacks on Machine Learning Cybersecurity Defenses in Industrial Control Systems” has other shortcomings according to the applicable categories (🗸). The only one that does not apply is the algorithm performance category, which is left blank. Table 2 A summary of reviewed studies and their shortcomings Shortcomings category #Ref Title of the paper Scope limitation Poor algorithm performance Lack of security information and event management (SIEM) integration Lack of IPS inclusion Lack of APT consideration Absence of Cybersecurity Auditing Consideration [ 10 ] Adversarial attacks on machine learning cybersecurity defences in Industrial Control Systems 🗸 🗸 🗸 🗸 🗸 [ 39 ] Malware Family Classification Using Active Learning by Learning 🗸 🗸 🗸 🗸 🗸 [ 43 ] Phishing Detection Using Machine Learning Techniques 🗸 🗸 🗸 🗸 🗸 [ 44 ] Spam Detection Approach for Secure Mobile Message Communication Using Machine Learning Algorithms 🗸 🗸 🗸 🗸 [ 25 ] SSH-Brute Force Attack Detection Model Based on Deep Learning 🗸 🗸 🗸 🗸 🗸 [ 40 ] Clustering-based semi-supervised machine learning for DDoS attack classification 🗸 🗸 🗸 🗸 🗸 [ 12 ] Real-Time Network Intrusion Prevention System Based on Hybrid Machine Learning 🗸 🗸 🗸 [ 45 ] Deep H2O: Cyber attacks detection in water distribution systems using deep learning 🗸 🗸 🗸 🗸 [ 46 ] A novel machine learning inspired algorithm to predict real-time network intrusions 🗸 🗸 🗸 🗸 [ 10 ] An adversarial reinforcement learning-based system for cyber security 🗸 🗸 🗸 🗸 🗸 [ 14 ] Fake News Detection Using Machine Learning Models Malak 🗸 🗸 🗸 🗸 [ 34 ] A new approach for APT malware detection based on deep graph network for endpoint systems 🗸 🗸 🗸 🗸 [ 32 ] IDERES: Intrusion detection and response system using machine learning and attack graphs 🗸 🗸 🗸 🗸 🗸 [ 47 ] Detecting Port Scan Attempts with Comparative Analysis of Deep Learning and Support Vector Machine Algorithms 🗸 🗸 🗸 🗸 [ 48 ] Incorporating Machine Learning Algorithms to Detect Phishing Websites 🗸 🗸 🗸 🗸 [ 49 ] IoT Real-Time Attacks Classification Framework Using Machine Learning 🗸 🗸 🗸 [ 50 ] Machine Learning for Detecting Anomalies and Intrusions in Communication Networks 🗸 🗸 🗸 🗸 [ 51 ] Deep learning-based cyber bullying early detection using distributed denial of service flow 🗸 🗸 🗸 🗸 [ 1 ] Click fraud detection for online advertising using machine learning Malak 🗸 🗸 🗸 🗸 The review confirms a wide range of research regarding ML applications in cybersecurity. The authors believe the studies truly reflect the state-of-the-art over the past five years. The subsequent section includes a detailed presentation of gaps identified while presenting future work. 5 Identified gaps As specified by Gupta et al. [ 3 ], technological advancement and the rapid development of more complex and sophisticated cyberattacks have a direct proportional relationship. As modern technologies advance, the risk associated with such progress becomes more persistent and complex. The reviewed literature confirmed gaps within ML application in cybersecurity, which future studies should improve on. As mentioned by L’heureux et al. [ 39 ]To date, no specific ML algorithm has attended to these challenges. After thoroughly reviewing additional research, the authors realized that the deficiencies of the literature can be categorized. It is crucial to categorize these inadequacies to enable future researchers to thoroughly investigate the finer details of each, aiming to enhance ML applications in cybersecurity. The gaps identified from the literature are divided into several categories: Scope limitations Suboptimal algorithm performance A lack of SIEM integration Insufficient IPS consideration Inadequate APT consideration The absence of cybersecurity auditing considerations (Fig. 6 ) The explanation of each gap identified is also presented, following the figure. The summary of identified gaps in the figure emanates from Table 2 , presented in the previous section. The percentage differences between the lack of IPS consideration, APT consideration, SIEM integration, and cybersecurity auditing consideration are insignificant. Notably, the lack of cybersecurity auditing consideration and SIEM integration are the most common gaps identified in the literature. Unsurprisingly, cybersecurity auditing is still gaining prominence in auditing. Consequently, inevitable gaps require attention. The relatively new concept of integrating ML into SIEM systems also necessitates regulating specific gaps. One of the aspects that can be deduced is that most studies focus on the detection part of using ML in cybersecurity. The prevention aspect remains a severe gap in applying ML in cybersecurity. Among several other factors that might influence the lack of ML-based IPS, directing false positives remains a daunting exercise. In an ideal world, it is essential that the three cores of cybersecurity—confidentiality, integrity, and availability (CIA), are equally directed. False positives threaten the availability of services to the intended users. For example, suppose an adopted ML intrusion detection system experiences a high rate of false positives. In that case, the automated response feature might block specific authentic requests from the user, resulting in an unavailable service to the intended users. If an adopted service or system experiences a high rate of false positives, this might also result in corporations being reluctant to adopt ML-based IPS; therefore, these false-positive gaps in cybersecurity and ML must be dealt with. Advanced persistent threats (APTs) constitute significant challenges because of their complexity and sophistication. The rapid evolution of APTs complicates bridging gaps in machine learning applications for APTs; Therefore, continuous improvement in this area is essential. As aforementioned, the identified gaps from the literature are common, as depicted in Fig. 6 ; however, categorizing these gaps helps when narrowing the scope to be directed. These specifically directed gaps are discussed in more detail in the subsequent sections. 5.1 Scope limitation Machine learning algorithms are problem-oriented; however, from the reviewed literature, some proposed ML algorithms should have been tested using alternative case scenarios. As demonstrated by Shahin et al. [ 54 ], using a wide range of case scenarios is vital to developing a versatile ML algorithm that can be used to detect diverse cyberattacks. Expanding on ML applications in cybersecurity can result in such an algorithm being more relevant and applicable in real-life scenarios; therefore, it is vital to include a wide range of test scenarios when proposing ML algorithms to combat modern and advanced cybersecurity threats. 5.2 Poor algorithm performance From the reviewed literature, the performance of the proposed ML for cybersecurity was satisfactory; however, ML performance can consistently be improved. Emanating from the reviewed studies, the performance of ML ranged from the lowest accuracy of 80% to an accuracy of 99.9%. In an ideal cybersecurity landscape, a continual need exists to enhance ML performance to reduce false positives and increase the trustworthiness of cybersecurity controls. The manual process of verifying false positives can be a daunting exercise; therefore, ML algorithms must achieve high performance and minimize false positives. The validation of false positives can also be incorporated into ML algorithms, as emphasized by Otsuki et al. [ 55 ]. 5.3 A lack of security information and event management (SIEM) integration Adopting SIEM software is prevalent in the corporate world, as organizations use it to gain real-time visibility of network activities; however, the reviewed literature does not explicitly indicate how the results of proposed ML algorithms in cybersecurity can be integrated into SIEM software to enhance its capabilities. Modern cybersecurity controls should collect suspicious activity data and notify security administrators through a centralized SIEM system [ 43 ]. 5.4 Lack of intrusion prevention systems consideration Automation is crucial for effective and efficient cyberattack prevention. The results indicate that most proposed ML solutions focus on detecting malicious activities within networks; however, insufficient work exists to prevent these detected anomalies. The authors strongly believe that a need exists to develop and assess more machine learning algorithms for intrusion prevention, particularly those with a low rate of false positives. 5.5 Lack of advanced persistent threats considerations Although some studies regard machine learning to combat APTs, APTs are inherently complex and sophisticated malware; therefore, continuous improvement in this area is vital. Xuan and DT Huong [ 34 ] confirm a need to develop ML algorithms that can effectively and efficiently establish accurate malware profile behaviors for APTs. 5.6 Absence of cybersecurity auditing consideration Unlike in the past, when auditing primarily focused on financial statements and ITGC testing, modern auditing processes now incorporate cybersecurity. Recently, the concepts of CCM and CCA acquired attention in cybersecurity auditing. As the name suggests, CCM continuously monitors key cybersecurity controls within an organization [ 56 ]. Likewise, CCA is a concept that integrates auditing key controls continuously [ 57 ]. ML is crucial in these concepts; however, the reviewed literature provides insufficient guidance on how CCM and CCA can be leveraged using the ML approach. This study aimed to identify gaps in machine learning applications in cybersecurity. Although this field is still in its initial stages, indications of potential areas for future research exist. To further explore the identified gaps, the subsequent section outlines potential future research directions. 6 Future work Although hackers have not yet publicized the concept of hacking-as-a-service (HaaS), it is foreseeable that HaaS will soon become a major topic. HaaS will complicate cybersecurity reinforcement for individuals and corporations with limited budgets. HaaS involves commercializing hacking skills, which can be used to genuinely improve the security posture of systems or for malicious purposes by competitors and other intruders. The authors believe that ML is the most feasible approach to combat HaaS. Future studies should explore HaaS and the challenges it may present. The shortcomings confirm the need to expand the ML application on IPS further since more studies focus on detecting ML applications. SIEM integration is vital for real-time threat detection. Future researchers should investigate how SIEM can be improved using ML. The other prominent aspect of ML in cybersecurity is for CCM and CCA, a subset of cybersecurity auditing; therefore, a need exists to explore how ML can be used to leverage CCM and CCA in the cybersecurity landscape within ML and cybersecurity. The gaps and the potential future work identified have been presented. The main output of this study is the identified gaps, as explained in the previous section. The subsequent section presents the main contribution and the current study’s shortcomings. 7 Critical evaluation As illustrated in the previous section, a wide range of work has been conducted to combat the ever-evolving cybersecurity challenges. Among several other cybersecurity countermeasures, ML can be considered the most feasible approach to reinforcing cybersecurity controls. Before concluding this study, the critical evaluation of the study concerning its benefits and shortcomings is presented next. 7.1 Research benefits This study explored the current work on applying ML in cybersecurity. After that, the state-of-the-art ML application in cybersecurity is presented systematically. This study emphasizes some of the most common and prominent data sources researchers can rely on to extract data to develop and test more ML algorithms. Throughout the reviews, one of the common challenges in developing and improving ML algorithms is acquiring the right data sources; therefore, we needed to present some common data sources to provide future researchers with a reference guide. The challenges of data sources for ML algorithms are serious and have also been emphasized by L’heureux et al. [ 39 ]. The other main contribution of the current study is the shortcomings of applying ML in cybersecurity. Identifying these challenges is vital for the maturity of ML applications in the cybersecurity realm. Researchers can use these challenges as a baseline where advancing ML applications in cybersecurity can be built or improved. From the studies, the authors presented the categories of challenges associated with applying ML in cybersecurity. These challenges are categorized; therefore, a taxonomy can be built based on the presented categories for future cybersecurity challenges. 7.2 Research shortcomings Although cybersecurity is an established field, it still encounters several challenges; therefore, continuous studies need to develop more reinforced approaches to counter the ever-evolving cyberattacks. No tool or technique can be applied as a general approach to direct the ML challenges in cybersecurity. Although this study presents the shortcomings that need to be addressed urgently, it did not cover all aspects of ML and cybersecurity shortcomings. The present study only reviewed studies not older than five years; therefore, there are possibilities of exclusion of other studies that might point to other valuable insights. Attributable to the fast-evolving field of cybersecurity, however, observing the last five years of activity provides a sufficient corpus to base our findings. Gaps are identified, and a thorough exploration of ML applications’ future in cybersecurity is presented. The subsequent section concludes the study through a high-level study objective and its output. 8 Conclusion Cybersecurity is a complex and ever-evolving environment. The rapid increase of APT and other modern cyberattacks is a critical concern; therefore, developing more robust and advanced cybersecurity capabilities is consistently needed. From the literature, among several other advantages of implementing ML in cybersecurity, the most visible benefit was using ML to effectively and efficiently detect malicious activities in network communication. Conversely, the most used algorithms in ML applications in cybersecurity are RF and SVM. This study further explored the literature and deduced some limitations that need to be solicited soon. In addition to the identified gaps within ML and cybersecurity, this study presents potential future directions. Declarations Conflict of interest. All authors included in this paper have participated in the draft and review of this study. We further confirm that this manuscript has not been submitted to, nor is it under review at, another journal or other publishing venue. Ethical approval Not applicable. Author Contribution Ndaedzo Rananga wrote the main manuscript of the study and H.S. Venter reviewed and guided the manuscript throughout. Acknowledgement Elizabeth Marx References M. Aljabri, R. Mustafa, and A. Mohammad, “Click fraud detection for online advertising using machine learning,” Egypt. Informatics J. , vol. 24, no. 2, pp. 341–350, 2023, doi: 10.1016/j.eij.2023.05.006. Y. Li and Q. Liu, “A comprehensive review study of cyber-attacks and cyber security; Emerging trends and recent developments,” Energy Reports , vol. 7, pp. 8176–8186, 2021, doi: 10.1016/j.egyr.2021.08.126. C. Gupta, I. Johri, K. Srinivasan, Y. Hu, and S. M. Qaisar, “A Systematic Review on Machine Learning and Deep Learning,” Prog. Biophys. Mol. Biol. , no. June, 2022, [Online]. Available: https://doi.org/10.1016/j.pbiomolbio.2022.07.004 G. Apruzzese et al . , “The Role of Machine Learning in Cybersecurity,” Digit. Threat. Res. Pract. , vol. 4, no. 1, pp. 1–38, 2023, doi: 10.1145/3545574. M. Macas, C. Wu, and W. Fuertes, “A survey on deep learning for cybersecurity: Progress, challenges, and opportunities,” Comput. Networks , vol. 212, no. April, p. 109032, 2022, doi: 10.1016/j.comnet.2022.109032. D. Dasgupta, Z. Akhtar, and S. Sen, “Machine learning in cybersecurity: a comprehensive survey,” J. Def. Model. Simul. , vol. 19, no. 1, pp. 57–106, 2022, doi: 10.1177/1548512920951275. A. S. Review, “Android Mobile Malware Detection Using Machine Learning :,” pp. 1–34, 2021. N. N. A. Sjarif et al . , “Endpoint Detection and Response: Why Use Machine Learning?,” ICTC 2019 - 10th Int. Conf. ICT Converg. ICT Converg. Lead. Auton. Futur. , pp. 283–288, 2019, doi: 10.1109/ICTC46691.2019.8939836. H. Liu and B. Lang, “Machine learning and deep learning methods for intrusion detection systems: A survey,” Appl. Sci. , vol. 9, no. 20, 2019, doi: 10.3390/app9204396. E. Anthi, L. Williams, M. Rhode, P. Burnap, and A. Wedgbury, “Adversarial attacks on machine learning cybersecurity defences in Industrial Control Systems,” J. Inf. Secur. Appl. , vol. 58, no. February, p. 102717, 2021, doi: 10.1016/j.jisa.2020.102717. H. Kaur and R. Tiwari, “Endpoint detection and response using machine learning,” J. Phys. Conf. Ser. , vol. 2062, no. 1, 2021, doi: 10.1088/1742-6596/2062/1/012013. W. Seo and W. Pak, “Real-Time Network Intrusion Prevention System Based on Hybrid Machine Learning,” IEEE Access , vol. 9, pp. 46386–46397, 2021, doi: 10.1109/ACCESS.2021.3066620. A. Tuor, S. Kaplan, B. Hutchinson, N. Nichols, and S. Robinson, “Deep learning for unsupervised insider threat detection in structured cybersecurity data streams,” AAAI Work. - Tech. Rep. , vol. WS-17-01-, no. 2012, pp. 224–234, 2017. M. Aljabri, D. M. Alomari, and M. Aboulnour, “Fake News Detection Using Machine Learning Models,” Proc. - 2022 14th IEEE Int. Conf. Comput. Intell. Commun. Networks, CICN 2022 , pp. 473–477, 2022, doi: 10.1109/CICN56167.2022.10008340. M. E. Sahin, L. Tawalbeh, and F. Muheidat, “The Security Concerns on Cyber-Physical Systems and Potential Risks Analysis Using Machine Learning,” Procedia Comput. Sci. , vol. 201, no. C, pp. 527–534, 2022, doi: 10.1016/j.procs.2022.03.068. S. Xia, M. Qiu, and H. Jiang, “An adversarial reinforcement learning based system for cyber security,” Proc. - 4th IEEE Int. Conf. Smart Cloud, SmartCloud 2019 3rd Int. Symp. Reinf. Learn. ISRL 2019 , pp. 227–230, 2019, doi: 10.1109/SmartCloud.2019.00046. C. Simmons, C. Ellis, S. Shiva, D. Dasgupta, and Q. Wu, “AVOIDIT: A cyber attack taxonomy, tech. Report,” Univ. Memphis, USA , pp. 1–9, 2009. S. Das and M. J. Nene, “A survey on types of machine learning techniques in intrusion prevention systems,” Proc. 2017 Int. Conf. Wirel. Commun. Signal Process. Networking, WiSPNET 2017 , vol. 2018-Janua, pp. 2296–2299, 2018, doi: 10.1109/WiSPNET.2017.8300169. H. De Wet and V. Marivate, “Is it Fake? news disinformation detection on south african news websites,” IEEE AFRICON Conf. , vol. 2021-Septe, 2021, doi: 10.1109/AFRICON51333.2021.9570905. J. Huang, “Detecting fake news with machine learning,” J. Phys. Conf. Ser. , vol. 1693, no. 1, 2020, doi: 10.1088/1742-6596/1693/1/012158. P. Akhtar, A. Mujahid, G. Haseeb, and U. Rehman, “Detecting fake news and disinformation using artificial.pdf,” pp. 633–657, 2023. M. Maras, Computer forensics: Cybercriminals, laws, and evidence 2nd edition . Jones & Bartlett Learning, 2015. B. Kiruba, V. Saravanan, T. Vasanth, and B. K. Yogeshwar, “OWASP Attack Prevention,” no. Icesc, pp. 1671–1675, 2022, doi: 10.1109/icesc54411.2022.9885691. J. Kaur, U. Garg, and G. Bathla, Detection of cross-site scripting (XSS) attacks using machine learning techniques: a review , no. 0123456789. Springer Netherlands, 2023. doi: 10.1007/s10462-023-10433-3. S. K. Wanjau, G. M. Wambugu, and G. N. Kamau, “SSH-Brute Force Attack Detection Model based on Deep Learning,” Int. J. Comput. Appl. Technol. Res. , vol. 10, no. 01, pp. 42–50, 2021, doi: 10.7753/ijcatr1001.1008. F. Musumeci, V. Ionata, F. Paolucci, F. Cugini, and M. Tornatore, “Machine-learning-assisted DDoS attack detection with P4 language,” IEEE Int. Conf. Commun. , vol. 2020-June, 2020, doi: 10.1109/ICC40277.2020.9149043. N. Rananga and H. S. Venter, “Mobile Cloud Computing Adoption Model as a Feasible Response to Countries’ Lockdown as a Result of the COVID-19 Outbreak and beyond,” 2020 IEEE Conf. e-Learning, e-Management e-Services, IC3e 2020 , no. Mcc, pp. 61–66, 2020, doi: 10.1109/IC3e50159.2020.9288402. Z. He, T. Zhang, and R. B. Lee, “Machine Learning Based DDoS Attack Detection from Source Side in Cloud,” Proc. - 4th IEEE Int. Conf. Cyber Secur. Cloud Comput. CSCloud 2017 3rd IEEE Int. Conf. Scalable Smart Cloud, SSC 2017 , pp. 114–120, 2017, doi: 10.1109/CSCloud.2017.58. P. L. S. Jayalaxmi, R. Saha, G. Kumar, M. Conti, and T. H. Kim, “Machine and Deep Learning Solutions for Intrusion Detection and Prevention in IoTs: A Survey,” IEEE Access , vol. 10, no. November, pp. 121173–121192, 2022, doi: 10.1109/ACCESS.2022.3220622. A. Wolsey, “The State-of-the-Art in AI-Based Malware Detection Techniques: A Review,” arXiv Prepr. arXiv2210.11239 , pp. 1–18, 2022, [Online]. Available: https://arxiv.org/abs/2210.11239%0Ahttps://arxiv.org/pdf/2210.11239 T. Saranya, S. Sridevi, C. Deisy, T. D. Chung, and M. K. A. A. Khan, “Performance Analysis of Machine Learning Algorithms in Intrusion Detection System: A Review,” Procedia Comput. Sci. , vol. 171, no. 2019, pp. 1251–1260, 2020, doi: 10.1016/j.procs.2020.04.133. J. R. Rose et al . , “IDERES: Intrusion detection and response system using machine learning and attack graphs,” J. Syst. Archit. , vol. 131, no. July, p. 102722, 2022, doi: 10.1016/j.sysarc.2022.102722. P. R. Chandre, P. N. Mahalle, and G. R. Shinde, “Machine Learning Based Novel Approach for Intrusion Detection and Prevention System: A Tool Based Verification,” Proc. - 2018 IEEE Glob. Conf. Wirel. Comput. Networking, GCWCN 2018 , pp. 135–140, 2019, doi: 10.1109/GCWCN.2018.8668618. C. Do Xuan and D. Huong, “A new approach for APT malware detection based on deep graph network for endpoint systems,” Appl. Intell. , pp. 14005–14024, 2022, doi: 10.1007/s10489-021-03138-z. C. Roest, S. J. Fransen, T. C. Kwee, and D. Yakar, “Comparative Performance of Deep Learning and Radiologists for the Diagnosis and Localization of Clinically Significant Prostate Cancer at MRI: A Systematic Review,” Life , vol. 12, no. 10, 2022, doi: 10.3390/life12101490. G. Sibiya, H. S. Venter, and T. Fogwill, “Digital forensics in the Cloud: The state of the art,” 2015 IST-Africa Conf. IST-Africa 2015 , pp. 1–9, 2015, doi: 10.1109/ISTAFRICA.2015.7190540. M. J. Page et al . , “The PRISMA 2020 statement: An updated guideline for reporting systematic reviews,” Int. J. Surg. , vol. 88, no. March, pp. 2020–2021, 2021, doi: 10.1016/j.ijsu.2021.105906. P. L. Hedley, C. M. Hagen, C. Wilstrup, and M. Christiansen, “The use of artificial intelligence and machine learning methods in first trimester pre-eclampsia screening: a systematic review protocol,” medRxiv , p. 2022.07.20.22277873, 2022, doi: 10.1371/journal.pone.0272465. A. L’Heureux, K. Grolinger, H. F. Elyamany, and M. A. M. Capretz, “Machine Learning with Big Data: Challenges and Approaches,” IEEE Access , vol. 5, pp. 7776–7797, 2017, doi: 10.1109/ACCESS.2017.2696365. C. W. Chen, C. H. Su, K. W. Lee, and P. H. Bair, “Malware Family Classification using Active Learning by Learning,” Int. Conf. Adv. Commun. Technol. ICACT , vol. 2020, pp. 590–595, 2020, doi: 10.23919/ICACT48636.2020.9061419. M. Aamir and S. M. Ali Zaidi, “Clustering based semi-supervised machine learning for DDoS attack classification,” J. King Saud Univ. - Comput. Inf. Sci. , vol. 33, no. 4, pp. 436–446, 2021, doi: 10.1016/j.jksuci.2019.02.003. N. Mohamed, “Current trends in AI and ML for cybersecurity: A state-of-the-art survey,” Cogent Eng. , vol. 10, no. 2, 2023, doi: 10.1080/23311916.2023.2272358. S. Iglesias Pérez, S. Moral-Rubio, and R. Criado, “A new approach to combine multiplex networks and time series attributes: Building intrusion detection systems (IDS) in cybersecurity,” Chaos, Solitons and Fractals , vol. 150, 2021, doi: 10.1016/j.chaos.2021.111143. H. Haddadpajouh, A. Azmoodeh, A. Dehghantanha, and R. M. Parizi, “MVFCC: A Multi-View Fuzzy Consensus Clustering Model for Malware Threat Attribution,” IEEE Access , vol. 8, pp. 139188–139198, 2020, doi: 10.1109/ACCESS.2020.3012907. L. Yang and A. Shami, “IDS-ML: An open source code for Intrusion Detection System development using Machine Learning[Formula presented],” Softw. Impacts , vol. 14, no. November, p. 100446, 2022, doi: 10.1016/j.simpa.2022.100446. J. Rashid, T. Mahmood, M. W. Nisar, and T. Nazir, “Phishing Detection Using Machine Learning Technique,” Proc. - 2020 1st Int. Conf. Smart Syst. Emerg. Technol. SMART-TECH 2020 , pp. 43–46, 2020, doi: 10.1109/SMART-TECH49988.2020.00026. L. Guangjun, S. Nazir, H. U. Khan, and A. U. Haq, “Spam Detection Approach for Secure Mobile Message Communication Using Machine Learning Algorithms,” Secur. Commun. Networks , vol. 2020, 2020, doi: 10.1155/2020/8873639. M. N. K. Sikder, M. B. T. Nguyen, E. D. Elliott, and F. A. Batarseh, “Deep H2O: Cyber attacks detection in water distribution systems using deep learning,” J. Water Process Eng. , vol. 52, no. October 2022, 2023, doi: 10.1016/j.jwpe.2023.103568. D. Aksu and M. A. Aydin, “Detecting Port Scan Attempts with Comparative Analysis of Deep Learning and Support Vector Machine Algorithms,” Int. Congr. Big Data, Deep Learn. Fight. Cyber Terror. IBIGDELFT 2018 - Proc. , pp. 77–80, 2019, doi: 10.1109/IBIGDELFT.2018.8625370. N. J. Sinthiya, T. A. Chowdhury, and A. B. Haque, “Incorporating Machine Learning Algorithms to Detect Phishing Websites,” 9th Int. Conf. ICT Smart Soc. Recover Together, Recover Stronger Smarter Smartization, Gov. Collab. ICISS 2022 - Proceeding , pp. 1–5, 2022, doi: 10.1109/ICISS55894.2022.9915211. N. Karmous, M. O. E. Aoueileyine, M. Abdelkader, and N. Youssef, “IoT Real-Time Attacks Classification Framework Using Machine Learning,” 2022 9th Int. Conf. Commun. Networking, ComNet 2022 - Proc. , pp. 1–5, 2022, doi: 10.1109/ComNet55492.2022.9998441. Z. Li, A. L. G. Rios, and L. Trajkovic, “Machine Learning for Detecting Anomalies and Intrusions in Communication Networks,” IEEE J. Sel. Areas Commun. , vol. 39, no. 7, pp. 2254–2264, 2021, doi: 10.1109/JSAC.2021.3078497. M. H. Zaib, F. Bashir, K. N. Qureshi, S. Kausar, M. Rizwan, and G. Jeon, “Deep learning based cyber bullying early detection using distributed denial of service flow,” Multimed. Syst. , vol. 28, no. 6, pp. 1905–1924, 2022, doi: 10.1007/s00530-021-00771-z. M. Shahin, F. F. Chen, A. Hosseinzadeh, H. Bouzary, and R. Rashidifar, “A deep hybrid learning model for detection of cyber attacks in industrial IoT devices,” Int. J. Adv. Manuf. Technol. , vol. 123, no. 5–6, pp. 1973–1983, 2022, doi: 10.1007/s00170-022-10329-6. Y. Kurogome et al . , “Eiger: Automated IOC generation for accurate and interpretable endpoint malware detection,” ACM Int. Conf. Proceeding Ser. , pp. 687–701, 2019, doi: 10.1145/3359789.3359808. K. Singh and P. Best, “Auditing during a pandemic – can continuous controls monitoring (CCM) address challenges facing internal audit departments?,” Pacific Account. Rev. , vol. 35, no. 5, pp. 727–745, 2023, doi: 10.1108/PAR-07-2022-0103. R. Van Hillo and H. Weigand, “Continuous Auditing & Continuous Monitoring: Continuous value?,” Proc. - Int. Conf. Res. Challenges Inf. Sci. , vol. 2016-August, no. Cm, pp. 1–11, 2016, doi: 10.1109/RCIS.2016.7549279. Additional Declarations No competing interests reported. Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 30 Sep, 2024 Reviews received at journal 20 Sep, 2024 Reviews received at journal 18 Sep, 2024 Reviewers agreed at journal 11 Sep, 2024 Reviewers agreed at journal 11 Sep, 2024 Reviewers agreed at journal 11 Sep, 2024 Reviewers agreed at journal 09 Sep, 2024 Reviewers invited by journal 09 Sep, 2024 Editor assigned by journal 26 Jul, 2024 Submission checks completed at journal 26 Jul, 2024 First submitted to journal 23 Jul, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4791216","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":342454208,"identity":"fe420cb0-3200-4116-896c-a81d5c167884","order_by":0,"name":"Ndaedzo Rananga","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAv0lEQVRIiWNgGAWjYBACgwOMD4CUDRAzNh4gSovFAWYDIJUG0tJAnBYbiJbDYA6RWm4fZpMuqDhvt7b9MNCWGptoglrMziWzSc84czt525lEoJZjabkNBLWc4T8mzdt2O9nsAFALY8NhwlqMzzCzSfP+O5dsdv4hkVoMwVoaDtiZ3SDWFqAWZmueY8kJZjeAtiQQ4xeDM8yMt3lq7OzNzqc/fPChxoawFhhIBKtMIFY5CNiTongUjIJRMApGGAAAgQ1F3mBJ65cAAAAASUVORK5CYII=","orcid":"","institution":"University of Pretoria","correspondingAuthor":true,"prefix":"","firstName":"Ndaedzo","middleName":"","lastName":"Rananga","suffix":""},{"id":342454209,"identity":"41e3d3a7-90e7-4cfc-b15d-5790d6630a5d","order_by":1,"name":"H. S. Venter","email":"","orcid":"","institution":"University of Pretoria","correspondingAuthor":false,"prefix":"","firstName":"H.","middleName":"S.","lastName":"Venter","suffix":""}],"badges":[],"createdAt":"2024-07-23 21:23:16","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4791216/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4791216/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":63085138,"identity":"26457ff0-3a90-4623-8b90-41f8f321be14","added_by":"auto","created_at":"2024-08-23 02:41:12","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":158715,"visible":true,"origin":"","legend":"\u003cp\u003eCyber Attack taxonomy [17]\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-4791216/v1/ced22805ec69589895d7b201.png"},{"id":63085133,"identity":"71f2831e-aead-4c31-96d5-cedbda2a7e8c","added_by":"auto","created_at":"2024-08-23 02:41:12","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":98325,"visible":true,"origin":"","legend":"\u003cp\u003ePRISMA Flow diagram\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-4791216/v1/123cb55e8f166a93d5aee27d.png"},{"id":63085499,"identity":"e1886131-38f9-4cd8-ab2c-8fb48f85f855","added_by":"auto","created_at":"2024-08-23 02:49:12","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":71269,"visible":true,"origin":"","legend":"\u003cp\u003eMachine learning algorithms most used in applying cybersecurity\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-4791216/v1/23bc5fe1838f0988dfef4f3c.png"},{"id":63085134,"identity":"4fa66307-e37c-4448-ab9a-a65f37982d89","added_by":"auto","created_at":"2024-08-23 02:41:12","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":146978,"visible":true,"origin":"","legend":"\u003cp\u003eMachine learning application in diverse cybersecurity scenarios\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-4791216/v1/5f51ecf1a68f673c560ee935.png"},{"id":63085136,"identity":"011be95b-ef30-4ab5-ba38-aeed57e5c0d8","added_by":"auto","created_at":"2024-08-23 02:41:12","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":94027,"visible":true,"origin":"","legend":"\u003cp\u003eCommon data sources for machine learning applications in cybersecurity\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-4791216/v1/d130e13f8a5df6e1afe302a4.png"},{"id":63085137,"identity":"f4d135f9-6ea0-42c0-8634-4de3f0e301ff","added_by":"auto","created_at":"2024-08-23 02:41:12","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":77019,"visible":true,"origin":"","legend":"\u003cp\u003eSummary of identified gaps\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-4791216/v1/c7896c360bd899c46858b063.png"},{"id":63085910,"identity":"35f65f3c-c564-4106-89c7-7f6a466804de","added_by":"auto","created_at":"2024-08-23 02:57:13","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1620944,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4791216/v1/763ce2fe-38cd-4db3-9d56-2205c62e47a0.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"A comprehensive review of machine learning applications in cybersecurity: identifying gaps and advocating for cybersecurity auditing","fulltext":[{"header":"1 Introduction","content":"\u003cp\u003eObservation confirms that a significant portion of personal, government, and corporate communication heavily depends on digital interactions within cyberspace. According to Aljabri et al.\u0026nbsp;[1], the primary revenue generation source in the business and market relies profoundly on the appreciation of modern technology and the use of advanced business models. Modern advancements and their evolution should be applauded because they enhance the quality of daily life while improving corporate productivity. Inevitably, like other technological innovations or advancements, cyberspace is prone to a wide range of challenges. One of the frontrunners of the difficulties associated with cyberspace is the exponential development of more complex, advanced persistent threats (APTs) and sophisticated cyberthreats. Cremer et al.\u0026nbsp;[2]\u0026nbsp;emphasize the effect of cyberthreats, which can escalate to billions of dollars in total costs to the global economy.\u003c/p\u003e\n\u003cp\u003eConsiderable work has been conducted to combat the ever-growing cyberthreats associated with modern technologies. Corporations are implementing several resources to respond to evolving cyberthreats substantially. These resources include adopting security best practices and cybersecurity standards and using modern technologies and tools. Generally, cybersecurity can be defined as a collective effort of procedures, people\u0026rsquo;s actions, and technologies to counter cyberthreats and, therefore, protect valuable assets\u0026nbsp;[3]. With artificial intelligence (AI), such as machine learning (ML), gaining ground in the current and future of modern technologies, cybersecurity tools are also exploiting ML to rejuvenate or revamp cyber capabilities against complex and sophisticated cyberthreats.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eApruzzese et al.[4]\u0026nbsp;emphasize that ML deployment in cybersecurity remains in its infancy, with gaps between the theories and the actual applicability of ML in cybersecurity practices. A need exists to perform a state-of-the-art review on ML applicability in modern cybersecurity capabilities using input from recent literature. \u003cem\u003eConversely, the problem is that due to the immaturity of implementing ML in the cybersecurity space, more research is needed to adopt ML applications within this space, therefore the need for this study.\u003c/em\u003e The study leverages the wide range of literature to perform such a state-of-the-art analysis or a systematic review to provide future directions on using ML in cybersecurity. As emphasized by Macs et al.\u0026nbsp;[5], a wide range of learning, challenges, and opportunities persist for ML regarding cybersecurity.\u003c/p\u003e\n\u003cp\u003eThe study systematically reviewed the developments of ML in the cybersecurity space. The shortcomings of the literature and the future indicators of using ML in the cybersecurity space are presented. The remainder of this study is organized as follows:\u0026nbsp;\u003c/p\u003e\n\u003cul\u003e\n \u003cli\u003e\u003cem\u003eSection\u0026nbsp;\u003c/em\u003e2 is a background section summarizing the fundamental theoretical concepts used throughout this paper. This includes the ML concept and the background of cybersecurity.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eThe methodology used to select reviewed studies from the literature is presented in \u003cem\u003eSection 3\u003c/em\u003e.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cem\u003eSection\u003c/em\u003e 4 presents a detailed description of the reviewed literature. This section expands the current research topic by reviewing further studies at a prominent level. This includes the results and the discussion of the reviewed studies.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eIn\u003cem\u003e\u0026nbsp;Section 5,\u0026nbsp;\u003c/em\u003egaps identified from the reviewed work are presented.\u003cem\u003e\u0026nbsp;\u003c/em\u003e\u003c/li\u003e\n \u003cli\u003eDue to the nature of the study, only a specific scope is included, and other areas are left for potential future work, as presented in \u003cem\u003eSection 6\u003c/em\u003e.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003eBefore the study is concluded, a critical evaluation of the study is presented in \u003cem\u003eSection\u003c/em\u003e 7.\u0026nbsp;\u003c/li\u003e\n \u003cli\u003e\u003cem\u003eSection\u0026nbsp;\u003c/em\u003e8\u003cem\u003e\u0026nbsp;\u003c/em\u003econcludes the study.\u003c/li\u003e\n\u003c/ul\u003e\n\u003cp\u003eAfter briefly introducing the present study, the subsequent section presents the fundamental background of terms used throughout this research.\u003c/p\u003e"},{"header":"2 Background","content":"\u003cp\u003eIn recent years, a subtle transformation has occurred in how users access, store, and communicate information. ML and cybersecurity are prominent in computing research and development. Before exploring the study details, the subsections concisely introduce ML and cybersecurity.\u003c/p\u003e \u003cdiv id=\"Sec2\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Machine learning\u003c/h2\u003e \u003cp\u003eThe literature provides multiple definitions of ML, reflecting the preferences of diverse authors. These definitions do not directly contradict one another. This study succinctly defines ML as a computational approach that leverages data to train computers for tasks, minimizing or eliminating human intervention. As described by Dasgupta et al. [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], ML methods and algorithms rely on automated approaches to imitate human learning and improve efficiency while improving accuracy. The main objective of ML is to develop intelligence applications and systems able to learn how to perform tasks without continuous explicit programming and human intervention [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eML can be classified into three main categories: supervised learning, unsupervised learning, and reinforcement learning [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], described subsequently.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section3\"\u003e \u003ch2\u003e2.1.1 Supervised learning\u003c/h2\u003e \u003cp\u003eSupervised machine learning revolves around using a known dataset to train algorithms for predicting potential outcomes. In this approach, predefined functions leverage input data to estimate the corresponding output values. [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Supervised learning directs classifications and regression-type problems [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. ML is problem-oriented; therefore, no universal rules exist regarding selecting ML techniques.\u003c/p\u003e \u003cp\u003eAccording to Liu and Lang [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], supervised ML algorithms generally perform better than unsupervised ML algorithms. Their argument [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] emphasizes that unsupervised learning relies on labeling data before training a model; however, this process can be resource-demanding and complex for unstructured data, such as extracting data from traffic flow to perform the ML algorithm. In cybersecurity, such as in Anthi et al. [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], supervised classification algorithms can detect cyberthreats using a malicious data set as the input sample data set. Appreciating and incorporating ML in modern cybersecurity tools, approaches, and techniques are influenced by the need to rejuvenate cybersecurity capabilities to combat complex and sophisticated cyberattacks [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. ML demonstrated remarkably accurate results [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e], which can ultimately lead to increased efficiency and cost savings within the cybersecurity community. By reducing the time spent pursuing false-positive alerts, ML contributes significantly to improving effectiveness.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section3\"\u003e \u003ch2\u003e2.1.2 Unsupervised learning\u003c/h2\u003e \u003cp\u003eUnlike supervised learning, unsupervised ML uses algorithms to analyze unlabeled datasets to predict outputs. Unsupervised ML can identify and characterize an unlabeled dataset during model training [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Unlike supervised learning, unsupervised machine learning operates without human feedback or labeled data. It focuses on identifying patterns and structures within unlabeled datasets [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section3\"\u003e \u003ch2\u003e2.1.3 Reinforcement learning\u003c/h2\u003e \u003cp\u003eAs aforementioned, ML approaches are techniques trained to imitate human behavior and perform future tasks without human interventions, specifically in unsupervised learning. Reinforcement learning is inspired by psychological theories that reward anticipated behaviors and penalize unexpected actions. In this paradigm, AI-based systems learn through trial and error, adjusting their actions based on the outcomes they receive \u0026mdash;rewards or punishments [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAfter a brief presentation of diverse categories of ML, owing to the approach by Sjarif et al. [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e presents a summary of the categories.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMachine learning categories description summary\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"2\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCategory\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSupervised learning\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eThese ML algorithms are trained to predict output using the provided input. For example, a model can be trained to label and classify malicious executables using an extract of the dataset from endpoint detection and response (EDR) or antivirus databases.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUnsupervised learning\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eThese algorithms can discover patterns from unlabeled datasets without explicit human interventions. For example, a model can be trained to detect real-time network activity anomalies [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] based on behavioral attributes. This approach to cybersecurity can also be useful in combating zero-day cyberattacks.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eReinforcement learning\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eThis ML model is based on reward and punishment, where applications are trained to make decisions that will reward them as the output. For example, an EDR ML-based application can be trained to isolate hosts from the network because of the detection of potential malicious behavior.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe wide range of research to combat cyberthreats must be acknowledged; however, continual advancements are needed to enhance real-time cyber detection and prevention [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. With the constant challenges of more sophisticated and complex cyberthreats, ML is a promising technique or approach for adequate resources to combat cyberthreats through more automated detections [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. The subsection further elaborates on this concept.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Machine learning application in cybersecurity\u003c/h2\u003e \u003cp\u003eSahin et al. [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] suggest that security vulnerabilities could exist in all technology, including modern advancements; therefore, the need to reinforce cybersecurity is becoming a primary concern. Cybersecurity is a concept of defined processes, technologies, tools, and people to counter the ever-evolving cyberthreats. Cybersecurity should protect assets, such as computer networks, users, programs, and data, from malicious actors and unauthorized access [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. The rapid adoption of ML algorithms, such as support vector machines (SVM), deep neural networks (DNN), and reinforcement learning (RL), led to a transition to modern approaches empowered by ML from traditional cybersecurity mechanisms to modern ML-empowered approaches [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. Comprehensive ML applications have been confirmed in cyberthreats. Simmons et al. [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] proposed a cyberattack taxonomy, such as AVOIDIT, helpful in classifying and categorizing diverse attacks based on their characteristics, methods, and effects, as presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e presents a cyberattack taxonomy, grouping common attacks based on their operational effect. These operational effects, as demonstrated in the taxonomy tree, include misuse of information, user and web compromise, installed malware, and service denial. Incorporating ML in these operational effects is evident in the modern cybersecurity realm. Computational intelligence is widely used in developing rigorous cyber prevention approaches toward this operational effect [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. The subsections present some shared areas that appreciate using ML to compact these operational consequences.\u003c/p\u003e \u003cdiv id=\"Sec7\" class=\"Section3\"\u003e \u003ch2\u003e2.2.1 Machine learning in misuse of (information) resources\u003c/h2\u003e \u003cp\u003eTechnology simplifies tasks for end users, enabling them to easily share information resources while keeping them informed with the latest updates; however, like any technological advancement, the significant transition from traditional resources of sharing information using studies to using digital technology to share information is associated with challenges. One of the visible challenges within the modern information resources era is \u003cem\u003enews disinformation\u003c/em\u003e. De Wet and Marivate [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e] suggest that discerning legitimate shared information resources from fake ones is increasingly difficult for humans. In a quest to combat this challenge, ML adoption is vital.\u003c/p\u003e \u003cp\u003eAccording to Huang [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], ML is adopted as a feasible countermeasure to information misuse, such as fake news. Misuse of information resources should be treated as a serious matter because it can range to the corporate level. For example, through social media, individuals can share fake news regarding supply chain tender opportunities and the operation of a company, which could have devastating effects on an organization. Akhtar et al. [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] realized the importance of incorporating ML to counter fake news at the corporate level. They proposed an ML model to detect fake news and disinformation to avoid supply chain disruptions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section3\"\u003e \u003ch2\u003e2.2.2 Machine learning in user compromise\u003c/h2\u003e \u003cp\u003eTechnology is becoming more accessible to most end users. According to Rananga and Venter [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e], adopting technologies, such as mobile computing, is growing exponentially. In the modern days of digitization and cyberspace, cyberthreats are still increasing [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]; however, cybersecurity tools are becoming more aware and can remediate some common cyberattacks. As such, attackers consistently target users, considered the most vulnerable asset within the computing realm. Cyberattacks targeting human weaknesses, such as social engineering, are the easiest to initiate and have a high success probability rate [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Analogously, a need exists for more advanced techniques to protect users from social engineering-related cyberthreats. Several research projects explore automating social engineering attacks using ML methods [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section3\"\u003e \u003ch2\u003e2.2.3 Machine learning in web compromise\u003c/h2\u003e \u003cp\u003eWeb applications and services are increasingly accessible, attributable to technological advancements; however, these developments remain vulnerable to sophisticated and targeted cyberattacks, including ATP attacks. Web attacks include web defacement, where unauthorized users access the website, edit, delete, or replace the original content without the web owner\u0026rsquo;s consent. Although some attacks, such as defacement, do not have direct damage nor permanent damage on the web servers, as indicated by Maras [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e], countering such attacks is necessary as they affect the integrity of website content. Besides defacement, the web is associated with several other attacks, such as SQL injection and DDoS.\u003c/p\u003e \u003cp\u003eAccording to Kiruba et al. [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e], most web-based attacks are initiated through the top 10 open worldwide application security project (OWASP) web vulnerabilities, such as injection, broken authentication, sensitive data disclosure, cross-site scripting (XSS), external entity (XSE), broken access control, insecure deserialization, insufficient logging and monitoring, and security misconfiguration. A wide range of ML adoption exists to improve cyber controls in web applications and services. Kaur et al. [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e] proposed an ML approach to detect the XSS attacks; the results of the ML model were satisfactory.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section3\"\u003e \u003ch2\u003e2.2.4 Machine learning in installed malware\u003c/h2\u003e \u003cp\u003eThe definition, type, and classification of malware are associated with the harm they cause [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Malware refers to malicious executable code that disrupts the normal functioning of a system or application. Common types of malware include viruses, worms, trojan horses, ransomware, adware, and spyware [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. The distinct description of the malware does not form part of the present study; however, as aforementioned, the malware names conform to the harm they cause. For example, ransomware is a type of malware that, when successfully executed on the target machine or host, encrypts the data until a ransom payment is made.\u003c/p\u003e \u003cp\u003eMalware detection techniques, such as static, dynamic, and hybrid analysis, have been proven to be effortlessly defeated by evasion techniques, such as obfuscation [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]; therefore, unlike in the past, when malware detection relied primarily on static and dynamic analysis of malware architecture to understand behavioral patterns, modern tools now leverage ML. These advanced ML techniques enable more efficient and accurate malware detection [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. ML offers numerous opportunities for malware detection, including identifying malware, targeting smartphones, particularly those running on Android operating systems (OSs) [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section3\"\u003e \u003ch2\u003e2.2.5 Machine learning in denial of services\u003c/h2\u003e \u003cp\u003eAccording to Wanjau et al. [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e], one of the most predominant cyberattacks encountered in network security is brute force. Brute-force attacks apply the trial-and-error method to guess login credentials, web browser\u0026rsquo;s directory, and encryption keys. The most common brute-force attack may be an attempt by an intruder or an attacker to acquire unauthorized access to the system using guessing techniques to bypass the authentication controls [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. In addition to brute-force attacks, DDoS is a common attack within cybersecurity. It is a cyberattack that disrupts the availability of services and applications to intended users. DDoS attacks typically involve consistently and repeatedly transmitting numerous requests to a target host.\u003c/p\u003e \u003cp\u003eThe goal is to overload or even exhaust the processing resources of the target host [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e], causing a severely disrupted service. One of the critical factors influencing technology adoption and appreciation is its ease of use and availability [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]; therefore, countering cyberattacks, such as DDoS, that threaten the availability of services to end users is vital. Given the necessity of incorporating ML in cybersecurity, it is unsurprising that researchers are developing modern ML algorithms to combat the emergence of brute force and DDoS attacks. According to He et al. [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], ML can leverage detecting DDoS by preventing network packets from being distributed to external networks.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e presents the cyberattack taxonomy, including the remediation and mitigation mechanisms. According to Jayalaxmi et al. [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e], intrusion detection systems (IDSs) and intrusion prevention systems (IPSs) are essential when developing robust cybersecurity controls. Conversely, EDR is also vital in endpoint cybersecurity controls [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. After emphasizing some of the applications of ML in directing various cyberthreats, it is crucial also to discuss how ML is being utilized in intrusion detection systems (IDSs), intrusion prevention systems (IPSs), and endpoint detection and response (EDR) solutions. This section describes these common countermeasures to cyberthreats and explains how ML has been integrated to enhance their effectiveness.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section3\"\u003e \u003ch2\u003e2.2.6 Machine learning in network intrusion detection systems (IDSs)\u003c/h2\u003e \u003cp\u003eThis study emphasizes that the rapid development in adopting modern technologies not only leverages daily operations but also exposes considerable security issues [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e] while inviting more sophisticated and complex cyberthreats. One of the most feasible responses to combat the ever-evolving cyberthreats is using IDSs. One of the advantages of Intrusion Detection Systems (IDSs) is their ability to manage suspicious activities within the network without disrupting other operations within the network ecosystem [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. With ML integration in modern technologies, it is unsurprising that more IDSs are being established using ML algorithms. These algorithms enhance efficiency by analyzing vast amounts of data and making detections faster than humans while also improving detection accuracy [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. Liu and Lang [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e] remark that ML approaches are also more vital to equipping the IDSs to detect unknown attacks. Rose et al. [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e] illustrate how ML can be used to detect IoT-related cyberthreats. The authors emphasize that ML models can also process false-positive alarms, which may help cybersecurity defenses mitigate cyberthreats more effectively.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section3\"\u003e \u003ch2\u003e2.2.7 Machine learning in network intrusion prevention systems (IPSs)\u003c/h2\u003e \u003cp\u003eIn addition to the extensive use of IDSs for detecting anomalies within networks, IPSs provide an additional layer of cybersecurity capabilities. IPSs are tools or techniques used for continuous network monitoring to detect suspicious activities and act upon detection. IPSs are considered a proactive approach to effective security threat prevention by blocking traffic from the offending host [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. As specified by Chandre et al. [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e], the primary distinction between an IDS and an IPS is that an IPS can take proactive measures to stop potential attacks. Conversely, a traditional IDS is limited to detecting attacks and alerting them to their presence without actively attempting to prevent them. Modern IPSs incorporate machine learning methods for real-time detection and prevention of cyberattacks with high accuracy [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. An IPS can commonly act after detecting suspicious activity, including isolating the host from the rest of the network ecosystem.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section3\"\u003e \u003ch2\u003e2.2.8 Machine learning in endpoint detection and response (EDR)\u003c/h2\u003e \u003cp\u003eAs more sophisticated and complex threats arise\u0026mdash;usually real-time-based attacks\u0026mdash;endpoints are no exception in these modern attacks. Endpoints are physical or virtual devices that exchange information and communicate with each other, employing computer networking. These devices, such as desktops, laptops, and mobile devices [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], are meant for end users\u0026rsquo; applications. As specified by Xuan and Huong [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e], one of the front runners of modern APTs is initiated by spreading malware through endpoints. Owning to the complexity of modern APT targeting endpoints, a need exists to rejuvenate EDR for adequate resources to protect endpoints.\u003c/p\u003e \u003cp\u003eAdopting ML is acquiring popularity in the EDR realm of cybersecurity [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], like in Kaur and Tiwari [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e], where the authors demonstrate how ML behavior and patterns can be used to characterize incidents targeting endpoints. The authors believe that rapidly incorporating ML to reinforce the endpoints\u0026rsquo; security will continue to acquire ground in modern computing.\u003c/p\u003e \u003cp\u003eThe fundamental background of applying ML in the cybersecurity area is presented. As emphasized by Roest et al. [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e], it is essential to perform an exhaustive systematic search to identify recent literature related to the study of interest. In the subsequent section, the state-of-the-art literature review of ML in cybersecurity is presented; however, before such a review is provided, a methodology adopted in the present study is presented first.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"3 Review methodology","content":"\u003cp\u003eAdopting online and digital library databases has resulted in a vast array of literature. Consequently, it is essential to define the selection criteria for choosing studies from the extensive range of available literature while excluding others. The following subsections present the review criteria adopted in the current research, along with a review included in this study. This study used a wide range of literature relevant to the application of machine learning in cybersecurity. Building upon existing literature without disregarding prior work, this study outlines the eligibility criteria for filtering literature.\u003c/p\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Eligibility criteria\u003c/h2\u003e \u003cp\u003eThe online libraries are chosen based on the PRISMA [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. The PRISMA flow is a guideline on effective resources for selecting literature for a systematic review. In the present study, the included libraries are used owing to their popularity, the vast number of articles, and ease of access. The libraries used in the present study include IEEE Xplore, ScienceDirect, Google Scholar, Springer Link, and Wiley Online Library. Publications from these online libraries that were not older than five years at the time of this study were considered.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Search strategy and sources\u003c/h2\u003e \u003cp\u003eAs adopted by Sibiya et al. [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e], search indexing using key phrases can filter results from diverse digital libraries. In our case, however, we adopted a similar strategy, considering the availability of databases as presented earlier in the \u003cspan refid=\"Sec16\" class=\"InternalRef\"\u003eeligibility criteria\u003c/span\u003e section. The keywords used to search for studies relevant to the current research include ((\u0026ldquo;machine learning\u0026rdquo; OR \u0026ldquo;Advanced Threat Persistence\u0026rdquo; OR \u0026ldquo;cybersecurity\u0026rdquo; OR \u0026ldquo;modern cybersecurity EDR\u0026rdquo; OR \u0026ldquo;modern cybersecurity IDS/IPS\u0026rdquo; OR \u0026ldquo;Prisma statement in machine learning reviews\u0026rdquo;) AND \u0026ldquo;machine learning application in cybersecurity\u0026rdquo;).\u003c/p\u003e \u003cdiv id=\"Sec18\" class=\"Section3\"\u003e \u003ch2\u003e3.2.1 Screening process and study selection\u003c/h2\u003e \u003cp\u003eAfter presenting the selection criteria adopted in the present study, the PRISMA flow diagram is presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. Traditionally, the PRISMA statement was intended to evaluate studies under health interventions; however, as suggested by Page et al. [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e], the PRISMA approach applies to several other systematic review studies, which is why it is adopted in the current study. The PRISMA flow diagram can be divided into:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eIdentifying the literature\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eScreening of the studies\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eThe indication of the included literature reviewed and presented in the systematic review [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e].\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eAccording to Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, the PRISMA approach indicates how the authors selected the studies for the systematic review. As demonstrated by Hedley et al. [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e], the abstract and title of the studies are used during the screening process, while only the studies the authors believe relate to the current research are incorporated.\u003c/p\u003e \u003cp\u003eThis study narrowed the scope of applying ML in cybersecurity, opportunities in cybersecurity, challenges of using ML in cybersecurity, and the future of ML in cybersecurity. As presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, after an exhaustive search of the digital libraries, 123 studies were identified. Twenty-three studies were omitted because they were duplicates. After the removal of duplicates, 100 were screened using the title and abstract. After the screening number, 81 articles were removed with confirmed reasons. Finally, 19 articles were reviewed and included in the present systematic review. Additional studies were excluded from the systematic review; however, they are still used in the literature and were considered applicable.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe previous section provided the reader with the selection and filtering process to arrive at the selected studies reviewed. The subsequent section presents the evaluation criteria applied during each study\u0026rsquo;s comparison and critical analysis.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section3\"\u003e \u003ch2\u003e3.2.2 Evaluation criteria\u003c/h2\u003e \u003cp\u003eAs demonstrated by Gupta et al. [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], the present study adopted some evaluation criteria in addition to the eligibility criteria, including the search strategy and sources. Several factors can be used to compare dissimilar ML techniques in cybersecurity. In Gupta et al. [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e] in comparing various studies, the authors use these factors: \u0026ldquo;\u003cem\u003eYear, Authors, Algorithm used, Source of data, and Advantages and Limitations\u003c/em\u003e\u0026rdquo; to perform a systematic review. Conversely, as presented in the previous section, Simmons et al. [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] proposed a cyberattack taxonomy that entails classification by:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eAttack vector\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eOperational effect\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eDefense\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eInformational effect\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eAttack target\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eEmanating from those studies, the current research adopted similar factors presented by Gupta et al. [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e] and Simmons et al. [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]; however, adapting them to be more relevant, as presented in the subsequent section.\u003c/p\u003e \u003cp\u003eAccording to Macas et al. [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], a wide range of research explores ML applicability and AI-related techniques to combat the rapid development of cyberthreats. The subsections present available and related frameworks and models for this systematic review.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"4 Systematic review","content":"\u003cp\u003eGupta et al. [\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e] remark that various ML approaches can secure ICT against cybersecurity threats. The background entailed throughout this study and the selection criteria are presented. The subsection presents the results from the reviewed studies.\u003c/p\u003e\n\u003cdiv id=\"Sec21\" class=\"Section2\"\u003e\n \u003ch2\u003e4.1 Results\u003c/h2\u003e\n \u003cp\u003eAs emphasized in the previous section, the choice of ML algorithm is problem-oriented; therefore, various problem scenarios can be solicited using diverse ML algorithms. As was evident in Shahrivari et al. [\u003cspan class=\"CitationRef\"\u003e40\u003c/span\u003e], alternative ML classification methods were used in the processes of training data, and the accuracy results differed. The reviewed work indicates no specific ML algorithm that could solve the cybersecurity challenges. The challenges of consistently finding the best tools and techniques to combat particular challenges were also emphasized by L\u0026rsquo;Heureux et al. [\u003cspan class=\"CitationRef\"\u003e39\u003c/span\u003e]. Figure \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e presents various ML algorithms that can combat advanced, and complex cyberattacks. These algorithms include random forest (RF), SVM, K-nearest neighbor (KNN), decision tree (DT), left-right parsing (LR), and nearest neighbor (NN).\u003c/p\u003e\n \u003cp\u003eFrom Fig. \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e, the most used ML algorithm in cybersecurity, according to the literature, is RF. Nerurkar et al. [\u003cspan class=\"CitationRef\"\u003e50\u003c/span\u003e] confirmed that RF is the most used ML algorithm in the literature. The dominant use of RF in ML is not a surprise because cybersecurity is a complex environment that changes more rapidly. RF can manage imbalanced data and non-linear relationships, as is consistently the case in cybersecurity datasets. For example, attacks consistently use more sophisticated resources to bypass security controls; therefore, using an ML algorithm that can combine multiple decisions to enable thorough detection is vital. The second most used ML algorithm in cybersecurity in the literature is SVM. SVM is believed to be more effective in instruction detection and malware classification [\u003cspan class=\"CitationRef\"\u003e40\u003c/span\u003e]. The other aspect that makes the SVM the most commonly used algorithm in cybersecurity is that it can handle high-dimensional and complex data [\u003cspan class=\"CitationRef\"\u003e41\u003c/span\u003e].\u003c/p\u003e\n \u003cp\u003eFor example, in the case of Xia et al. [\u003cspan class=\"CitationRef\"\u003e16\u003c/span\u003e], a case scenario application can be vital to supporting a study\u0026rsquo;s outcome. To further understand the reviewed studies from the literature, the authors analyzed how the research aligns with the category from the cyberattacks taxonomy presented in the previous section. This includes applying ML in web compromise, DoS, user compromise, misuse of information, installed malware, NIPS, network instruction detection systems (NIDS), and EDR, as presented in Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e\n \u003cp\u003eAs depicted in Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e, ML applications in NIDS constitute the highest percentage, with an overall of 31.6%. ML applications in NIDS include using ML algorithms to detect patterns of unusual activities in the networking environment. Mohamed [\u003cspan class=\"CitationRef\"\u003e42\u003c/span\u003e] confirms that using ML in NIDS is the most common ML application in the quest to combat cyberthreats. P\u0026eacute;rez et al. [\u003cspan class=\"CitationRef\"\u003e43\u003c/span\u003e] emphasize that NIDSs are one of the tools in cybersecurity; therefore, more ML algorithms are being developed in this area.\u003c/p\u003e\n \u003cp\u003eLike in the case of Haddadpajouh et al. [\u003cspan class=\"CitationRef\"\u003e44\u003c/span\u003e], one of the customary factors that can be used to assess the ML algorithm is performance and accuracy. From the reviewed studies, the algorithm\u0026rsquo;s accuracy was satisfactory, ranging from at least 80\u0026ndash;99.9%.\u003c/p\u003e\n \u003cp\u003eOne of the most challenging aspects of ML is acquiring the correct data from the data sources. Usually, researchers rely on historical data to develop and assess models. Figure\u0026nbsp;5 presents some of the common data sources for ML applications in cybersecurity. From the reviewed studies, the most common data source for cybersecurity ML applications is the CICIDS2017 dataset developed by the Canadian Institute for Cybersecurity (CIC). Yang and Shami [\u003cspan class=\"CitationRef\"\u003e45\u003c/span\u003e] suggest that the CICIDS2017 dataset is commonly used because it is vital in developing ML models to combat various types of cyberattacks, such as DoS attacks, port-scan attacks, brute-force attacks, web-based attacks, botnet attacks, and infiltration attacks.\u003c/p\u003e\n \u003cp\u003eAfter presenting details of the reviewed studies and the most prominent data sources, the authors present the summary of the reviewed studies in a tabular format (Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e). The shortcomings of each reviewed research study are also provided in the table. To expand on the idea behind the presented table, the shortcomings identified from the reviewed studies are categorized into subcategories (Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e). If the deficiencies apply to a specific study, a tick (🗸) under the applicable category is used; otherwise, a blank space indicates that the shortcomings do not apply to a particular study.\u003c/p\u003e\n \u003cp\u003eFor example, the study entitled \u0026ldquo;Adversarial Attacks on Machine Learning Cybersecurity Defenses in Industrial Control Systems\u0026rdquo; has other shortcomings according to the applicable categories (🗸). The only one that does not apply is the algorithm performance category, which is left blank.\u003c/p\u003e\n \u003cp\u003e\u003c/p\u003e\u0026nbsp;\u003ctable id=\"Tab2\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eA summary of reviewed studies and their shortcomings\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\" colspan=\"6\"\u003e\n \u003cp\u003eShortcomings category\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\" colspan=\"2\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e#Ref\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eTitle of the paper\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eScope limitation\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003ePoor algorithm performance\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eLack of security information and event management (SIEM) integration\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eLack of IPS inclusion\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eLack of\u003c/p\u003e\n \u003cp\u003eAPT consideration\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eAbsence of Cybersecurity Auditing Consideration\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e10\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eAdversarial attacks on machine learning cybersecurity defences in Industrial Control Systems\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e39\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eMalware Family Classification Using Active Learning by Learning\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e43\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003ePhishing Detection Using Machine Learning Techniques\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e44\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eSpam Detection Approach for Secure Mobile Message Communication Using Machine Learning Algorithms\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e25\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eSSH-Brute Force Attack Detection Model Based on Deep Learning\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e40\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eClustering-based semi-supervised machine learning for DDoS attack classification\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e12\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eReal-Time Network Intrusion Prevention System Based on Hybrid Machine Learning\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e45\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eDeep H2O: Cyber attacks detection in water distribution systems using deep learning\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e46\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eA novel machine learning inspired algorithm to predict real-time network intrusions\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e10\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eAn adversarial reinforcement learning-based system for cyber security\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e14\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eFake News Detection Using Machine Learning Models\u003c/p\u003e\n \u003cp\u003eMalak\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e34\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eA new approach for APT malware detection based on deep graph network for endpoint systems\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e32\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eIDERES: Intrusion detection and response system using machine learning and attack graphs\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e47\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eDetecting Port Scan Attempts with Comparative Analysis of Deep Learning and Support Vector Machine Algorithms\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e48\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eIncorporating Machine Learning Algorithms to Detect Phishing Websites\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e49\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eIoT Real-Time Attacks Classification Framework Using Machine Learning\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e50\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eMachine Learning for Detecting Anomalies and Intrusions in Communication Networks\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e51\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eDeep learning-based cyber bullying early detection using distributed denial of service flow\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u003cspan class=\"CitationRef\"\u003e1\u003c/span\u003e]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cem\u003eClick fraud detection for online advertising using machine learning\u003c/em\u003e\u003c/p\u003e\n \u003cp\u003e\u003cem\u003eMalak\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003e🗸\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cp\u003e\u003c/p\u003e\n \u003cp\u003eThe review confirms a wide range of research regarding ML applications in cybersecurity. The authors believe the studies truly reflect the state-of-the-art over the past five years. The subsequent section includes a detailed presentation of gaps identified while presenting future work.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"5 Identified gaps","content":"\u003cp\u003eAs specified by Gupta et al. [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], technological advancement and the rapid development of more complex and sophisticated cyberattacks have a direct proportional relationship. As modern technologies advance, the risk associated with such progress becomes more persistent and complex. The reviewed literature confirmed gaps within ML application in cybersecurity, which future studies should improve on. As mentioned by L\u0026rsquo;heureux et al. [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]To date, no specific ML algorithm has attended to these challenges.\u003c/p\u003e \u003cp\u003eAfter thoroughly reviewing additional research, the authors realized that the deficiencies of the literature can be categorized. It is crucial to categorize these inadequacies to enable future researchers to thoroughly investigate the finer details of each, aiming to enhance ML applications in cybersecurity. The gaps identified from the literature are divided into several categories:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eScope limitations\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSuboptimal algorithm performance\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eA lack of SIEM integration\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eInsufficient IPS consideration\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eInadequate APT consideration\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eThe absence of cybersecurity auditing considerations (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e6\u003c/span\u003e)\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eThe explanation of each gap identified is also presented, following the figure. The summary of identified gaps in the figure emanates from Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, presented in the previous section.\u003c/p\u003e \u003cp\u003eThe percentage differences between the lack of IPS consideration, APT consideration, SIEM integration, and cybersecurity auditing consideration are insignificant. Notably, the lack of cybersecurity auditing consideration and SIEM integration are the most common gaps identified in the literature. Unsurprisingly, cybersecurity auditing is still gaining prominence in auditing. Consequently, inevitable gaps require attention. The relatively new concept of integrating ML into SIEM systems also necessitates regulating specific gaps.\u003c/p\u003e \u003cp\u003eOne of the aspects that can be deduced is that most studies focus on the detection part of using ML in cybersecurity. The prevention aspect remains a severe gap in applying ML in cybersecurity. Among several other factors that might influence the lack of ML-based IPS, directing false positives remains a daunting exercise. In an ideal world, it is essential that the three cores of cybersecurity\u0026mdash;confidentiality, integrity, and availability (CIA), are equally directed. False positives threaten the availability of services to the intended users.\u003c/p\u003e \u003cp\u003eFor example, suppose an adopted ML intrusion detection system experiences a high rate of false positives. In that case, the automated response feature might block specific authentic requests from the user, resulting in an unavailable service to the intended users. If an adopted service or system experiences a high rate of false positives, this might also result in corporations being reluctant to adopt ML-based IPS; therefore, these false-positive gaps in cybersecurity and ML must be dealt with.\u003c/p\u003e \u003cp\u003eAdvanced persistent threats (APTs) constitute significant challenges because of their complexity and sophistication. The rapid evolution of APTs complicates bridging gaps in machine learning applications for APTs; Therefore, continuous improvement in this area is essential.\u003c/p\u003e \u003cp\u003eAs aforementioned, the identified gaps from the literature are common, as depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e6\u003c/span\u003e; however, categorizing these gaps helps when narrowing the scope to be directed. These specifically directed gaps are discussed in more detail in the subsequent sections.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cdiv id=\"Sec23\" class=\"Section2\"\u003e \u003ch2\u003e5.1 Scope limitation\u003c/h2\u003e \u003cp\u003eMachine learning algorithms are problem-oriented; however, from the reviewed literature, some proposed ML algorithms should have been tested using alternative case scenarios. As demonstrated by Shahin et al. [\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e], using a wide range of case scenarios is vital to developing a versatile ML algorithm that can be used to detect diverse cyberattacks. Expanding on ML applications in cybersecurity can result in such an algorithm being more relevant and applicable in real-life scenarios; therefore, it is vital to include a wide range of test scenarios when proposing ML algorithms to combat modern and advanced cybersecurity threats.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003e5.2 Poor algorithm performance\u003c/h2\u003e \u003cp\u003eFrom the reviewed literature, the performance of the proposed ML for cybersecurity was satisfactory; however, ML performance can consistently be improved. Emanating from the reviewed studies, the performance of ML ranged from the lowest accuracy of 80% to an accuracy of 99.9%. In an ideal cybersecurity landscape, a continual need exists to enhance ML performance to reduce false positives and increase the trustworthiness of cybersecurity controls. The manual process of verifying false positives can be a daunting exercise; therefore, ML algorithms must achieve high performance and minimize false positives. The validation of false positives can also be incorporated into ML algorithms, as emphasized by Otsuki et al. [\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec25\" class=\"Section2\"\u003e \u003ch2\u003e5.3 A lack of security information and event management (SIEM) integration\u003c/h2\u003e \u003cp\u003eAdopting SIEM software is prevalent in the corporate world, as organizations use it to gain real-time visibility of network activities; however, the reviewed literature does not explicitly indicate how the results of proposed ML algorithms in cybersecurity can be integrated into SIEM software to enhance its capabilities. Modern cybersecurity controls should collect suspicious activity data and notify security administrators through a centralized SIEM system [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec26\" class=\"Section2\"\u003e \u003ch2\u003e5.4 Lack of intrusion prevention systems consideration\u003c/h2\u003e \u003cp\u003eAutomation is crucial for effective and efficient cyberattack prevention. The results indicate that most proposed ML solutions focus on detecting malicious activities within networks; however, insufficient work exists to prevent these detected anomalies. The authors strongly believe that a need exists to develop and assess more machine learning algorithms for intrusion prevention, particularly those with a low rate of false positives.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec27\" class=\"Section2\"\u003e \u003ch2\u003e5.5 Lack of advanced persistent threats considerations\u003c/h2\u003e \u003cp\u003eAlthough some studies regard machine learning to combat APTs, APTs are inherently complex and sophisticated malware; therefore, continuous improvement in this area is vital. Xuan and DT Huong [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e] confirm a need to develop ML algorithms that can effectively and efficiently establish accurate malware profile behaviors for APTs.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec28\" class=\"Section2\"\u003e \u003ch2\u003e5.6 Absence of cybersecurity auditing consideration\u003c/h2\u003e \u003cp\u003eUnlike in the past, when auditing primarily focused on financial statements and ITGC testing, modern auditing processes now incorporate cybersecurity. Recently, the concepts of CCM and CCA acquired attention in cybersecurity auditing. As the name suggests, CCM continuously monitors key cybersecurity controls within an organization [\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e]. Likewise, CCA is a concept that integrates auditing key controls continuously [\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e]. ML is crucial in these concepts; however, the reviewed literature provides insufficient guidance on how CCM and CCA can be leveraged using the ML approach.\u003c/p\u003e \u003cp\u003eThis study aimed to identify gaps in machine learning applications in cybersecurity. Although this field is still in its initial stages, indications of potential areas for future research exist. To further explore the identified gaps, the subsequent section outlines potential future research directions.\u003c/p\u003e \u003c/div\u003e"},{"header":"6 Future work","content":"\u003cp\u003eAlthough hackers have not yet publicized the concept of hacking-as-a-service (HaaS), it is foreseeable that HaaS will soon become a major topic. HaaS will complicate cybersecurity reinforcement for individuals and corporations with limited budgets. HaaS involves commercializing hacking skills, which can be used to genuinely improve the security posture of systems or for malicious purposes by competitors and other intruders. The authors believe that ML is the most feasible approach to combat HaaS. Future studies should explore HaaS and the challenges it may present.\u003c/p\u003e \u003cp\u003eThe shortcomings confirm the need to expand the ML application on IPS further since more studies focus on detecting ML applications. SIEM integration is vital for real-time threat detection. Future researchers should investigate how SIEM can be improved using ML. The other prominent aspect of ML in cybersecurity is for CCM and CCA, a subset of cybersecurity auditing; therefore, a need exists to explore how ML can be used to leverage CCM and CCA in the cybersecurity landscape within ML and cybersecurity.\u003c/p\u003e \u003cp\u003eThe gaps and the potential future work identified have been presented. The main output of this study is the identified gaps, as explained in the previous section. The subsequent section presents the main contribution and the current study\u0026rsquo;s shortcomings.\u003c/p\u003e"},{"header":"7 Critical evaluation","content":"\u003cp\u003eAs illustrated in the previous section, a wide range of work has been conducted to combat the ever-evolving cybersecurity challenges. Among several other cybersecurity countermeasures, ML can be considered the most feasible approach to reinforcing cybersecurity controls. Before concluding this study, the critical evaluation of the study concerning its benefits and shortcomings is presented next.\u003c/p\u003e \u003cdiv id=\"Sec31\" class=\"Section2\"\u003e \u003ch2\u003e7.1 Research benefits\u003c/h2\u003e \u003cp\u003eThis study explored the current work on applying ML in cybersecurity. After that, the state-of-the-art ML application in cybersecurity is presented systematically. This study emphasizes some of the most common and prominent data sources researchers can rely on to extract data to develop and test more ML algorithms. Throughout the reviews, one of the common challenges in developing and improving ML algorithms is acquiring the right data sources; therefore, we needed to present some common data sources to provide future researchers with a reference guide. The challenges of data sources for ML algorithms are serious and have also been emphasized by L\u0026rsquo;heureux et al. [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe other main contribution of the current study is the shortcomings of applying ML in cybersecurity. Identifying these challenges is vital for the maturity of ML applications in the cybersecurity realm. Researchers can use these challenges as a baseline where advancing ML applications in cybersecurity can be built or improved. From the studies, the authors presented the categories of challenges associated with applying ML in cybersecurity. These challenges are categorized; therefore, a taxonomy can be built based on the presented categories for future cybersecurity challenges.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec32\" class=\"Section2\"\u003e \u003ch2\u003e7.2 Research shortcomings\u003c/h2\u003e \u003cp\u003eAlthough cybersecurity is an established field, it still encounters several challenges; therefore, continuous studies need to develop more reinforced approaches to counter the ever-evolving cyberattacks. No tool or technique can be applied as a general approach to direct the ML challenges in cybersecurity. Although this study presents the shortcomings that need to be addressed urgently, it did not cover all aspects of ML and cybersecurity shortcomings. The present study only reviewed studies not older than five years; therefore, there are possibilities of exclusion of other studies that might point to other valuable insights. Attributable to the fast-evolving field of cybersecurity, however, observing the last five years of activity provides a sufficient corpus to base our findings.\u003c/p\u003e \u003cp\u003eGaps are identified, and a thorough exploration of ML applications\u0026rsquo; future in cybersecurity is presented. The subsequent section concludes the study through a high-level study objective and its output.\u003c/p\u003e \u003c/div\u003e"},{"header":"8 Conclusion","content":"\u003cp\u003eCybersecurity is a complex and ever-evolving environment. The rapid increase of APT and other modern cyberattacks is a critical concern; therefore, developing more robust and advanced cybersecurity capabilities is consistently needed. From the literature, among several other advantages of implementing ML in cybersecurity, the most visible benefit was using ML to effectively and efficiently detect malicious activities in network communication. Conversely, the most used algorithms in ML applications in cybersecurity are RF and SVM. This study further explored the literature and deduced some limitations that need to be solicited soon. In addition to the identified gaps within ML and cybersecurity, this study presents potential future directions.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eConflict of interest.\u003c/h2\u003e\n\u003cp\u003eAll authors included in this paper have participated in the draft and review of this study. We further confirm that this manuscript has not been submitted to, nor is it under review at, another journal or other publishing venue.\u003c/p\u003e\n\u003ch2\u003eEthical approval\u003c/h2\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\n\u003cp\u003eNdaedzo Rananga wrote the main manuscript of the study and H.S. Venter reviewed and guided the manuscript throughout.\u003c/p\u003e\n\u003ch2\u003eAcknowledgement\u003c/h2\u003e\n\u003cp\u003eElizabeth Marx\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eM. Aljabri, R. Mustafa, and A. Mohammad, \u0026ldquo;Click fraud detection for online advertising using machine learning,\u0026rdquo; \u003cem\u003eEgypt. Informatics J.\u003c/em\u003e, vol. 24, no. 2, pp. 341\u0026ndash;350, 2023, doi: 10.1016/j.eij.2023.05.006.\u003c/li\u003e\n \u003cli\u003eY. Li and Q. Liu, \u0026ldquo;A comprehensive review study of cyber-attacks and cyber security; Emerging trends and recent developments,\u0026rdquo; \u003cem\u003eEnergy Reports\u003c/em\u003e, vol. 7, pp. 8176\u0026ndash;8186, 2021, doi: 10.1016/j.egyr.2021.08.126.\u003c/li\u003e\n \u003cli\u003eC. Gupta, I. Johri, K. Srinivasan, Y. Hu, and S. M. Qaisar, \u0026ldquo;A Systematic Review on Machine Learning and Deep Learning,\u0026rdquo; \u003cem\u003eProg. Biophys. Mol. Biol.\u003c/em\u003e, no. June, 2022, [Online]. Available: https://doi.org/10.1016/j.pbiomolbio.2022.07.004\u003c/li\u003e\n \u003cli\u003eG. Apruzzese et al\u003cem\u003e.\u003c/em\u003e, \u0026ldquo;The Role of Machine Learning in Cybersecurity,\u0026rdquo; \u003cem\u003eDigit. Threat. Res. Pract.\u003c/em\u003e, vol. 4, no. 1, pp. 1\u0026ndash;38, 2023, doi: 10.1145/3545574.\u003c/li\u003e\n \u003cli\u003eM. Macas, C. Wu, and W. Fuertes, \u0026ldquo;A survey on deep learning for cybersecurity: Progress, challenges, and opportunities,\u0026rdquo; \u003cem\u003eComput. Networks\u003c/em\u003e, vol. 212, no. April, p. 109032, 2022, doi: 10.1016/j.comnet.2022.109032.\u003c/li\u003e\n \u003cli\u003eD. Dasgupta, Z. Akhtar, and S. Sen, \u0026ldquo;Machine learning in cybersecurity: a comprehensive survey,\u0026rdquo; \u003cem\u003eJ. Def. Model. Simul.\u003c/em\u003e, vol. 19, no. 1, pp. 57\u0026ndash;106, 2022, doi: 10.1177/1548512920951275.\u003c/li\u003e\n \u003cli\u003eA. S. Review, \u0026ldquo;Android Mobile Malware Detection Using Machine Learning :,\u0026rdquo; pp. 1\u0026ndash;34, 2021.\u003c/li\u003e\n \u003cli\u003eN. N. A. Sjarif et al\u003cem\u003e.\u003c/em\u003e, \u0026ldquo;Endpoint Detection and Response: Why Use Machine Learning?,\u0026rdquo; \u003cem\u003eICTC 2019 - 10th Int. Conf. ICT Converg. ICT Converg. Lead. Auton. Futur.\u003c/em\u003e, pp. 283\u0026ndash;288, 2019, doi: 10.1109/ICTC46691.2019.8939836.\u003c/li\u003e\n \u003cli\u003eH. Liu and B. Lang, \u0026ldquo;Machine learning and deep learning methods for intrusion detection systems: A survey,\u0026rdquo; \u003cem\u003eAppl. Sci.\u003c/em\u003e, vol. 9, no. 20, 2019, doi: 10.3390/app9204396.\u003c/li\u003e\n \u003cli\u003eE. Anthi, L. Williams, M. Rhode, P. Burnap, and A. Wedgbury, \u0026ldquo;Adversarial attacks on machine learning cybersecurity defences in Industrial Control Systems,\u0026rdquo; \u003cem\u003eJ. Inf. Secur. Appl.\u003c/em\u003e, vol. 58, no. February, p. 102717, 2021, doi: 10.1016/j.jisa.2020.102717.\u003c/li\u003e\n \u003cli\u003eH. Kaur and R. Tiwari, \u0026ldquo;Endpoint detection and response using machine learning,\u0026rdquo; \u003cem\u003eJ. Phys. Conf. Ser.\u003c/em\u003e, vol. 2062, no. 1, 2021, doi: 10.1088/1742-6596/2062/1/012013.\u003c/li\u003e\n \u003cli\u003eW. Seo and W. Pak, \u0026ldquo;Real-Time Network Intrusion Prevention System Based on Hybrid Machine Learning,\u0026rdquo; \u003cem\u003eIEEE Access\u003c/em\u003e, vol. 9, pp. 46386\u0026ndash;46397, 2021, doi: 10.1109/ACCESS.2021.3066620.\u003c/li\u003e\n \u003cli\u003eA. Tuor, S. Kaplan, B. Hutchinson, N. Nichols, and S. Robinson, \u0026ldquo;Deep learning for unsupervised insider threat detection in structured cybersecurity data streams,\u0026rdquo; \u003cem\u003eAAAI Work. - Tech. Rep.\u003c/em\u003e, vol. WS-17-01-, no. 2012, pp. 224\u0026ndash;234, 2017.\u003c/li\u003e\n \u003cli\u003eM. Aljabri, D. M. Alomari, and M. Aboulnour, \u0026ldquo;Fake News Detection Using Machine Learning Models,\u0026rdquo; \u003cem\u003eProc. - 2022 14th IEEE Int. Conf. Comput. Intell. Commun. Networks, CICN 2022\u003c/em\u003e, pp. 473\u0026ndash;477, 2022, doi: 10.1109/CICN56167.2022.10008340.\u003c/li\u003e\n \u003cli\u003eM. E. Sahin, L. Tawalbeh, and F. Muheidat, \u0026ldquo;The Security Concerns on Cyber-Physical Systems and Potential Risks Analysis Using Machine Learning,\u0026rdquo; \u003cem\u003eProcedia Comput. Sci.\u003c/em\u003e, vol. 201, no. C, pp. 527\u0026ndash;534, 2022, doi: 10.1016/j.procs.2022.03.068.\u003c/li\u003e\n \u003cli\u003eS. Xia, M. Qiu, and H. Jiang, \u0026ldquo;An adversarial reinforcement learning based system for cyber security,\u0026rdquo; \u003cem\u003eProc. - 4th IEEE Int. Conf. Smart Cloud, SmartCloud 2019 3rd Int. Symp. Reinf. Learn. ISRL 2019\u003c/em\u003e, pp. 227\u0026ndash;230, 2019, doi: 10.1109/SmartCloud.2019.00046.\u003c/li\u003e\n \u003cli\u003eC. Simmons, C. Ellis, S. Shiva, D. Dasgupta, and Q. Wu, \u0026ldquo;AVOIDIT: A cyber attack taxonomy, tech. Report,\u0026rdquo; \u003cem\u003eUniv. Memphis, USA\u003c/em\u003e, pp. 1\u0026ndash;9, 2009.\u003c/li\u003e\n \u003cli\u003eS. Das and M. J. Nene, \u0026ldquo;A survey on types of machine learning techniques in intrusion prevention systems,\u0026rdquo; \u003cem\u003eProc. 2017 Int. Conf. Wirel. Commun. Signal Process. Networking, WiSPNET 2017\u003c/em\u003e, vol. 2018-Janua, pp. 2296\u0026ndash;2299, 2018, doi: 10.1109/WiSPNET.2017.8300169.\u003c/li\u003e\n \u003cli\u003eH. De Wet and V. Marivate, \u0026ldquo;Is it Fake? news disinformation detection on south african news websites,\u0026rdquo; \u003cem\u003eIEEE AFRICON Conf.\u003c/em\u003e, vol. 2021-Septe, 2021, doi: 10.1109/AFRICON51333.2021.9570905.\u003c/li\u003e\n \u003cli\u003eJ. Huang, \u0026ldquo;Detecting fake news with machine learning,\u0026rdquo; \u003cem\u003eJ. Phys. Conf. Ser.\u003c/em\u003e, vol. 1693, no. 1, 2020, doi: 10.1088/1742-6596/1693/1/012158.\u003c/li\u003e\n \u003cli\u003eP. Akhtar, A. Mujahid, G. Haseeb, and U. Rehman, \u0026ldquo;Detecting fake news and disinformation using artificial.pdf,\u0026rdquo; pp. 633\u0026ndash;657, 2023.\u003c/li\u003e\n \u003cli\u003eM. Maras, \u003cem\u003eComputer forensics: Cybercriminals, laws, and evidence 2nd edition\u003c/em\u003e. Jones \u0026amp; Bartlett Learning, 2015.\u003c/li\u003e\n \u003cli\u003eB. Kiruba, V. Saravanan, T. Vasanth, and B. K. Yogeshwar, \u0026ldquo;OWASP Attack Prevention,\u0026rdquo; no. Icesc, pp. 1671\u0026ndash;1675, 2022, doi: 10.1109/icesc54411.2022.9885691.\u003c/li\u003e\n \u003cli\u003eJ. Kaur, U. Garg, and G. Bathla, \u003cem\u003eDetection of cross-site scripting (XSS) attacks using machine learning techniques: a review\u003c/em\u003e, no. 0123456789. Springer Netherlands, 2023. doi: 10.1007/s10462-023-10433-3.\u003c/li\u003e\n \u003cli\u003eS. K. Wanjau, G. M. Wambugu, and G. N. Kamau, \u0026ldquo;SSH-Brute Force Attack Detection Model based on Deep Learning,\u0026rdquo; \u003cem\u003eInt. J. Comput. Appl. Technol. Res.\u003c/em\u003e, vol. 10, no. 01, pp. 42\u0026ndash;50, 2021, doi: 10.7753/ijcatr1001.1008.\u003c/li\u003e\n \u003cli\u003eF. Musumeci, V. Ionata, F. Paolucci, F. Cugini, and M. Tornatore, \u0026ldquo;Machine-learning-assisted DDoS attack detection with P4 language,\u0026rdquo; \u003cem\u003eIEEE Int. Conf. Commun.\u003c/em\u003e, vol. 2020-June, 2020, doi: 10.1109/ICC40277.2020.9149043.\u003c/li\u003e\n \u003cli\u003eN. Rananga and H. S. Venter, \u0026ldquo;Mobile Cloud Computing Adoption Model as a Feasible Response to Countries\u0026rsquo; Lockdown as a Result of the COVID-19 Outbreak and beyond,\u0026rdquo; \u003cem\u003e2020 IEEE Conf. e-Learning, e-Management e-Services, IC3e 2020\u003c/em\u003e, no. Mcc, pp. 61\u0026ndash;66, 2020, doi: 10.1109/IC3e50159.2020.9288402.\u003c/li\u003e\n \u003cli\u003eZ. He, T. Zhang, and R. B. Lee, \u0026ldquo;Machine Learning Based DDoS Attack Detection from Source Side in Cloud,\u0026rdquo; \u003cem\u003eProc. - 4th IEEE Int. Conf. Cyber Secur. Cloud Comput. CSCloud 2017 3rd IEEE Int. Conf. Scalable Smart Cloud, SSC 2017\u003c/em\u003e, pp. 114\u0026ndash;120, 2017, doi: 10.1109/CSCloud.2017.58.\u003c/li\u003e\n \u003cli\u003eP. L. S. Jayalaxmi, R. Saha, G. Kumar, M. Conti, and T. H. Kim, \u0026ldquo;Machine and Deep Learning Solutions for Intrusion Detection and Prevention in IoTs: A Survey,\u0026rdquo; \u003cem\u003eIEEE Access\u003c/em\u003e, vol. 10, no. November, pp. 121173\u0026ndash;121192, 2022, doi: 10.1109/ACCESS.2022.3220622.\u003c/li\u003e\n \u003cli\u003eA. Wolsey, \u0026ldquo;The State-of-the-Art in AI-Based Malware Detection Techniques: A Review,\u0026rdquo; \u003cem\u003earXiv Prepr. arXiv2210.11239\u003c/em\u003e, pp. 1\u0026ndash;18, 2022, [Online]. Available: https://arxiv.org/abs/2210.11239%0Ahttps://arxiv.org/pdf/2210.11239\u003c/li\u003e\n \u003cli\u003eT. Saranya, S. Sridevi, C. Deisy, T. D. Chung, and M. K. A. A. Khan, \u0026ldquo;Performance Analysis of Machine Learning Algorithms in Intrusion Detection System: A Review,\u0026rdquo; \u003cem\u003eProcedia Comput. Sci.\u003c/em\u003e, vol. 171, no. 2019, pp. 1251\u0026ndash;1260, 2020, doi: 10.1016/j.procs.2020.04.133.\u003c/li\u003e\n \u003cli\u003eJ. R. Rose et al\u003cem\u003e.\u003c/em\u003e, \u0026ldquo;IDERES: Intrusion detection and response system using machine learning and attack graphs,\u0026rdquo; \u003cem\u003eJ. Syst. Archit.\u003c/em\u003e, vol. 131, no. July, p. 102722, 2022, doi: 10.1016/j.sysarc.2022.102722.\u003c/li\u003e\n \u003cli\u003eP. R. Chandre, P. N. Mahalle, and G. R. Shinde, \u0026ldquo;Machine Learning Based Novel Approach for Intrusion Detection and Prevention System: A Tool Based Verification,\u0026rdquo; \u003cem\u003eProc. - 2018 IEEE Glob. Conf. Wirel. Comput. Networking, GCWCN 2018\u003c/em\u003e, pp. 135\u0026ndash;140, 2019, doi: 10.1109/GCWCN.2018.8668618.\u003c/li\u003e\n \u003cli\u003eC. Do Xuan and D. Huong, \u0026ldquo;A new approach for APT malware detection based on deep graph network for endpoint systems,\u0026rdquo; \u003cem\u003eAppl. Intell.\u003c/em\u003e, pp. 14005\u0026ndash;14024, 2022, doi: 10.1007/s10489-021-03138-z.\u003c/li\u003e\n \u003cli\u003eC. Roest, S. J. Fransen, T. C. Kwee, and D. Yakar, \u0026ldquo;Comparative Performance of Deep Learning and Radiologists for the Diagnosis and Localization of Clinically Significant Prostate Cancer at MRI: A Systematic Review,\u0026rdquo; \u003cem\u003eLife\u003c/em\u003e, vol. 12, no. 10, 2022, doi: 10.3390/life12101490.\u003c/li\u003e\n \u003cli\u003eG. Sibiya, H. S. Venter, and T. Fogwill, \u0026ldquo;Digital forensics in the Cloud: The state of the art,\u0026rdquo; \u003cem\u003e2015 IST-Africa Conf. IST-Africa 2015\u003c/em\u003e, pp. 1\u0026ndash;9, 2015, doi: 10.1109/ISTAFRICA.2015.7190540.\u003c/li\u003e\n \u003cli\u003eM. J. Page et al\u003cem\u003e.\u003c/em\u003e, \u0026ldquo;The PRISMA 2020 statement: An updated guideline for reporting systematic reviews,\u0026rdquo; \u003cem\u003eInt. J. Surg.\u003c/em\u003e, vol. 88, no. March, pp. 2020\u0026ndash;2021, 2021, doi: 10.1016/j.ijsu.2021.105906.\u003c/li\u003e\n \u003cli\u003eP. L. Hedley, C. M. Hagen, C. Wilstrup, and M. Christiansen, \u0026ldquo;The use of artificial intelligence and machine learning methods in first trimester pre-eclampsia screening: a systematic review protocol,\u0026rdquo; \u003cem\u003emedRxiv\u003c/em\u003e, p. 2022.07.20.22277873, 2022, doi: 10.1371/journal.pone.0272465.\u003c/li\u003e\n \u003cli\u003eA. L\u0026rsquo;Heureux, K. Grolinger, H. F. Elyamany, and M. A. M. Capretz, \u0026ldquo;Machine Learning with Big Data: Challenges and Approaches,\u0026rdquo; \u003cem\u003eIEEE Access\u003c/em\u003e, vol. 5, pp. 7776\u0026ndash;7797, 2017, doi: 10.1109/ACCESS.2017.2696365.\u003c/li\u003e\n \u003cli\u003eC. W. Chen, C. H. Su, K. W. Lee, and P. H. Bair, \u0026ldquo;Malware Family Classification using Active Learning by Learning,\u0026rdquo; \u003cem\u003eInt. Conf. Adv. Commun. Technol. ICACT\u003c/em\u003e, vol. 2020, pp. 590\u0026ndash;595, 2020, doi: 10.23919/ICACT48636.2020.9061419.\u003c/li\u003e\n \u003cli\u003eM. Aamir and S. M. Ali Zaidi, \u0026ldquo;Clustering based semi-supervised machine learning for DDoS attack classification,\u0026rdquo; \u003cem\u003eJ. King Saud Univ. - Comput. Inf. Sci.\u003c/em\u003e, vol. 33, no. 4, pp. 436\u0026ndash;446, 2021, doi: 10.1016/j.jksuci.2019.02.003.\u003c/li\u003e\n \u003cli\u003eN. Mohamed, \u0026ldquo;Current trends in AI and ML for cybersecurity: A state-of-the-art survey,\u0026rdquo; \u003cem\u003eCogent Eng.\u003c/em\u003e, vol. 10, no. 2, 2023, doi: 10.1080/23311916.2023.2272358.\u003c/li\u003e\n \u003cli\u003eS. Iglesias P\u0026eacute;rez, S. Moral-Rubio, and R. Criado, \u0026ldquo;A new approach to combine multiplex networks and time series attributes: Building intrusion detection systems (IDS) in cybersecurity,\u0026rdquo; \u003cem\u003eChaos, Solitons and Fractals\u003c/em\u003e, vol. 150, 2021, doi: 10.1016/j.chaos.2021.111143.\u003c/li\u003e\n \u003cli\u003eH. Haddadpajouh, A. Azmoodeh, A. Dehghantanha, and R. M. Parizi, \u0026ldquo;MVFCC: A Multi-View Fuzzy Consensus Clustering Model for Malware Threat Attribution,\u0026rdquo; \u003cem\u003eIEEE Access\u003c/em\u003e, vol. 8, pp. 139188\u0026ndash;139198, 2020, doi: 10.1109/ACCESS.2020.3012907.\u003c/li\u003e\n \u003cli\u003eL. Yang and A. Shami, \u0026ldquo;IDS-ML: An open source code for Intrusion Detection System development using Machine Learning[Formula presented],\u0026rdquo; \u003cem\u003eSoftw. Impacts\u003c/em\u003e, vol. 14, no. November, p. 100446, 2022, doi: 10.1016/j.simpa.2022.100446.\u003c/li\u003e\n \u003cli\u003eJ. Rashid, T. Mahmood, M. W. Nisar, and T. Nazir, \u0026ldquo;Phishing Detection Using Machine Learning Technique,\u0026rdquo; \u003cem\u003eProc. - 2020 1st Int. Conf. Smart Syst. Emerg. Technol. SMART-TECH 2020\u003c/em\u003e, pp. 43\u0026ndash;46, 2020, doi: 10.1109/SMART-TECH49988.2020.00026.\u003c/li\u003e\n \u003cli\u003eL. Guangjun, S. Nazir, H. U. Khan, and A. U. Haq, \u0026ldquo;Spam Detection Approach for Secure Mobile Message Communication Using Machine Learning Algorithms,\u0026rdquo; \u003cem\u003eSecur. Commun. Networks\u003c/em\u003e, vol. 2020, 2020, doi: 10.1155/2020/8873639.\u003c/li\u003e\n \u003cli\u003eM. N. K. Sikder, M. B. T. Nguyen, E. D. Elliott, and F. A. Batarseh, \u0026ldquo;Deep H2O: Cyber attacks detection in water distribution systems using deep learning,\u0026rdquo; \u003cem\u003eJ. Water Process Eng.\u003c/em\u003e, vol. 52, no. October 2022, 2023, doi: 10.1016/j.jwpe.2023.103568.\u003c/li\u003e\n \u003cli\u003eD. Aksu and M. A. Aydin, \u0026ldquo;Detecting Port Scan Attempts with Comparative Analysis of Deep Learning and Support Vector Machine Algorithms,\u0026rdquo; \u003cem\u003eInt. Congr. Big Data, Deep Learn. Fight. Cyber Terror. IBIGDELFT 2018 - Proc.\u003c/em\u003e, pp. 77\u0026ndash;80, 2019, doi: 10.1109/IBIGDELFT.2018.8625370.\u003c/li\u003e\n \u003cli\u003eN. J. Sinthiya, T. A. Chowdhury, and A. B. Haque, \u0026ldquo;Incorporating Machine Learning Algorithms to Detect Phishing Websites,\u0026rdquo; \u003cem\u003e9th Int. Conf. ICT Smart Soc. Recover Together, Recover Stronger Smarter Smartization, Gov. Collab. ICISS 2022 - Proceeding\u003c/em\u003e, pp. 1\u0026ndash;5, 2022, doi: 10.1109/ICISS55894.2022.9915211.\u003c/li\u003e\n \u003cli\u003eN. Karmous, M. O. E. Aoueileyine, M. Abdelkader, and N. Youssef, \u0026ldquo;IoT Real-Time Attacks Classification Framework Using Machine Learning,\u0026rdquo; \u003cem\u003e2022 9th Int. Conf. Commun. Networking, ComNet 2022 - Proc.\u003c/em\u003e, pp. 1\u0026ndash;5, 2022, doi: 10.1109/ComNet55492.2022.9998441.\u003c/li\u003e\n \u003cli\u003eZ. Li, A. L. G. Rios, and L. Trajkovic, \u0026ldquo;Machine Learning for Detecting Anomalies and Intrusions in Communication Networks,\u0026rdquo; \u003cem\u003eIEEE J. Sel. Areas Commun.\u003c/em\u003e, vol. 39, no. 7, pp. 2254\u0026ndash;2264, 2021, doi: 10.1109/JSAC.2021.3078497.\u003c/li\u003e\n \u003cli\u003eM. H. Zaib, F. Bashir, K. N. Qureshi, S. Kausar, M. Rizwan, and G. Jeon, \u0026ldquo;Deep learning based cyber bullying early detection using distributed denial of service flow,\u0026rdquo; \u003cem\u003eMultimed. Syst.\u003c/em\u003e, vol. 28, no. 6, pp. 1905\u0026ndash;1924, 2022, doi: 10.1007/s00530-021-00771-z.\u003c/li\u003e\n \u003cli\u003eM. Shahin, F. F. Chen, A. Hosseinzadeh, H. Bouzary, and R. Rashidifar, \u0026ldquo;A deep hybrid learning model for detection of cyber attacks in industrial IoT devices,\u0026rdquo; \u003cem\u003eInt. J. Adv. Manuf. Technol.\u003c/em\u003e, vol. 123, no. 5\u0026ndash;6, pp. 1973\u0026ndash;1983, 2022, doi: 10.1007/s00170-022-10329-6.\u003c/li\u003e\n \u003cli\u003eY. Kurogome et al\u003cem\u003e.\u003c/em\u003e, \u0026ldquo;Eiger: Automated IOC generation for accurate and interpretable endpoint malware detection,\u0026rdquo; \u003cem\u003eACM Int. Conf. Proceeding Ser.\u003c/em\u003e, pp. 687\u0026ndash;701, 2019, doi: 10.1145/3359789.3359808.\u003c/li\u003e\n \u003cli\u003eK. Singh and P. Best, \u0026ldquo;Auditing during a pandemic \u0026ndash; can continuous controls monitoring (CCM) address challenges facing internal audit departments?,\u0026rdquo; \u003cem\u003ePacific Account. Rev.\u003c/em\u003e, vol. 35, no. 5, pp. 727\u0026ndash;745, 2023, doi: 10.1108/PAR-07-2022-0103.\u003c/li\u003e\n \u003cli\u003eR. Van Hillo and H. Weigand, \u0026ldquo;Continuous Auditing \u0026amp; Continuous Monitoring: Continuous value?,\u0026rdquo; \u003cem\u003eProc. - Int. Conf. Res. Challenges Inf. Sci.\u003c/em\u003e, vol. 2016-August, no. Cm, pp. 1\u0026ndash;11, 2016, doi: 10.1109/RCIS.2016.7549279.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"international-journal-of-information-security","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ijis","sideBox":"Learn more about [International Journal of Information Security](http://link.springer.com/journal/10207)","snPcode":"10207","submissionUrl":"https://submission.nature.com/new-submission/10207/3","title":"International Journal of Information Security","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"cybersecurity auditing, cybersecurity threats, artificial intelligence, cybersecurity tools, modern and sophisticated cyberthreats, machine learning, state-of-the-art","lastPublishedDoi":"10.21203/rs.3.rs-4791216/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4791216/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eCybersecurity threats present significant challenges in the ever-evolving landscape of information and communication technology (ICT). As a practical approach to counter these evolving threats, corporations invest in various measures, including adopting cybersecurity standards, enhancing controls, and leveraging modern cybersecurity tools. Exponential development is established using machine learning and artificial intelligence within the computing domain. Cybersecurity tools also capitalize on these advancements, employing machine learning to direct complex and sophisticated cyberthreats. While incorporating machine learning into cybersecurity is still in its preliminary stages, continuous state-of-the-art analysis is necessary to assess its feasibility and applicability in combating modern cyberthreats. The challenge remains in the relative immaturity of implementing machine learning in cybersecurity, necessitating further research, as emphasized in this study. This study used the preferred reporting items for systematic reviews and meta-analysis (PRISMA) methodology as a scientific approach to reviewing recent literature on the applicability and feasibility of machine learning implementation in cybersecurity. This study presents the inadequacies of the research field. Finally, the directions for machine learning implementation in cybersecurity are depicted owing to the present study\u0026rsquo;s systematic review. This study functions as a foundational baseline from which rigorous machine-learning models and frameworks for cybersecurity can be constructed or improved.\u003c/p\u003e","manuscriptTitle":"A comprehensive review of machine learning applications in cybersecurity: identifying gaps and advocating for cybersecurity auditing","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-08-23 02:41:08","doi":"10.21203/rs.3.rs-4791216/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-09-30T07:50:25+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-09-20T06:10:26+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-09-18T08:01:19+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"279179209399273386675534657412517752912","date":"2024-09-11T16:25:50+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"50359553747408733891234423664748416706","date":"2024-09-11T14:11:08+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"65627608121358745431310270948182008946","date":"2024-09-11T10:47:37+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"296993822437349062104578280316300222800","date":"2024-09-09T11:00:42+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-09-09T10:44:09+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-07-27T02:58:36+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-07-27T02:58:18+00:00","index":"","fulltext":""},{"type":"submitted","content":"International Journal of Information Security","date":"2024-07-23T21:15:12+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"international-journal-of-information-security","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ijis","sideBox":"Learn more about [International Journal of Information Security](http://link.springer.com/journal/10207)","snPcode":"10207","submissionUrl":"https://submission.nature.com/new-submission/10207/3","title":"International Journal of Information Security","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"b0077967-3566-4f2b-98ea-3ceeeaec80eb","owner":[],"postedDate":"August 23rd, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-04-21T09:42:04+00:00","versionOfRecord":[],"versionCreatedAt":"2024-08-23 02:41:08","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4791216","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4791216","identity":"rs-4791216","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.