<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">31641</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2023.031641</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>BotSward: Centrality Measures for Graph-Based Bot Detection Using Machine Learning</article-title>
<alt-title alt-title-type="left-running-head">BotSward: Centrality Measures for Graph-Based Bot Detection Using Machine Learning</alt-title>
<alt-title alt-title-type="right-running-head">BotSward: Centrality Measures for Graph-Based Bot Detection Using Machine Learning</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Shinan</surname><given-names>Khlood</given-names></name><xref ref-type="aff" rid="aff-1">1</xref>
<xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Alsubhi</surname><given-names>Khalid</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Usman Ashraf</surname><given-names>M.</given-names></name><xref ref-type="aff" rid="aff-3">3</xref><email>usman.ashraf@gcwus.edu.pk</email></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Science, College Computer Science in Al-Leith, Umm Al-Qura University</institution>, <addr-line>Mecca 21421</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Computer Science, Faculty of Computing and Information Technology, King Abdulaziz University</institution>, <addr-line>Jeddah 21589</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Computer Science, GC Women University Sialkot</institution>, <country>Pakistan</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: M. Usman Ashraf. Email: <email>usman.ashraf@gcwus.edu.pk</email></corresp>
</author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2022-08-16"><day>16</day>
<month>08</month>
<year>2022</year></pub-date>
<volume>74</volume>
<issue>1</issue>
<fpage>693</fpage>
<lpage>714</lpage>
<history>
<date date-type="received"><day>23</day><month>4</month><year>2022</year></date>
<date date-type="accepted"><day>08</day><month>6</month><year>2022</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Shinan et al.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Shinan et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_31641.pdf"></self-uri>
<abstract>
<p>The number of botnet malware attacks on Internet devices has grown at an equivalent rate to the number of Internet devices that are connected to the Internet. Bot detection using machine learning (ML) with flow-based features has been extensively studied in the literature. Existing flow-based detection methods involve significant computational overhead that does not completely capture network communication patterns that might reveal other features of malicious hosts. Recently, Graph-Based Bot Detection methods using ML have gained attention to overcome these limitations, as graphs provide a real representation of network communications. The purpose of this study is to build a botnet malware detection system utilizing centrality measures for graph-based botnet detection and ML. We propose BotSward, a graph-based bot detection system that is based on ML. We apply the efficient centrality measures, which are Closeness Centrality (CC), Degree Centrality (CC), and PageRank (PR), and compare them with others used in the state-of-the-art. The efficiency of the proposed method is verified on the available Czech Technical University 13 dataset (CTU-13). The CTU-13 dataset contains 13 real botnet traffic scenarios that are connected to a command-and-control (C&#x0026;C) channel and that cause malicious actions such as phishing, distributed denial-of-service (DDoS) attacks, spam attacks, etc. BotSward is robust to zero-day attacks, suitable for large-scale datasets, and is intended to produce better accuracy than state-of-the-art techniques. The proposed BotSward solution achieved 99&#x0025; accuracy in botnet attack detection with a false positive rate as low as 0.0001&#x0025;.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Network security</kwd>
<kwd>botnet detection</kwd>
<kwd>graph-based features</kwd>
<kwd>machine learning</kwd>
<kwd>measure centrality</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>The explosion of network applications and connected devices on the Internet has resulted in a high level of network complexity in terms of management and increased concerns over data security. Many institutions are continually exposed to a variety of security threats, which can result in financial and reputational harm. Malware such as viruses, spyware, trojans, worms, key-loggers, and botnets damages various systems and online businesses. A botnet is an overlay network formed by a large number of hosts (bots or zombies) that have been infected with bots and are controlled remotely by an attacker (botmaster). Botnets are one of the most dangerous types of serious cybersecurity threats, as they are a primary vector for large-scale attack campaigns including email spam, click fraud, financial theft, and DDoS attacks [<xref ref-type="bibr" rid="ref-1">1</xref>].</p>
<p>Botnet attacks will get worse because they are not typically created to infect just a single computer; they are designed to infect millions of them. Botherders frequently use a Trojan horse virus to install botnets on personal computers (PCs). Users often infect their own computers by downloading email attachments, clicking on malicious pop-up ads, or installing malicious software from a website. After infecting devices, botnets are then free to attack other computers, access and modify personal information, and commit other crimes. Botnets that are more advanced can even self-propagate, discovering and infecting devices on their own. Botnets take a long time to develop. Many will remain inactive on devices until the botmaster calls them for a DDoS attack or spam dissemination [<xref ref-type="bibr" rid="ref-2">2</xref>].</p>
<p>The first botnet attack to receive widespread attention, known as &#x201C;EarthLink Spammer&#x201D;, was an email spammer created by Khan K. Smith in 2000. He made at least 3 million dollars from the botnet, which sent 1.25 million emails with phishing schemes [<xref ref-type="bibr" rid="ref-3">3</xref>]. Botnets have grown over the previous 20 years to become increasingly complex and destructive [<xref ref-type="bibr" rid="ref-4">4</xref>]. In 2016, Methbot, one of the first large and sophisticated ad fraud and takedown operations, appeared. It was the world&#x2019;s largest botnet for defrauding the advertising industry, using sophisticated bots that pretended to be premium publishers to watch 300 million video advertisements each day on spoofed websites. Over 6,000 premium domains have been spoofed [<xref ref-type="bibr" rid="ref-5">5</xref>]. The year 2019 also produced one of the major ransomware attacks, when the French cyber police freed over 850,000 Windows computers from a botnet named &#x201C;Read up.&#x201D; When antivirus company Avast discovered a flaw in Retadup&#x2019;s C&#x0026;C communications protocol, it alerted the French National Gendarmerie, who in turn seized the servers [<xref ref-type="bibr" rid="ref-6">6</xref>].</p>
<p>Therefore, detecting botnet propagation activities early and determining the expected size of the attack is critical. However, bots can avoid detection by mimicking normal traffic, modifying packet structures, and hiding their payload characteristics through encryption. Attackers began replacing internet relay chat (IRC) with hypertext transfer protocol (HTTP) in order to blend into normal HTTP traffic and avoid detection. With port-based filtering, botnets can easily get around intrusion detection systems (IDSs) and firewalls [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>Botnet detection has received a lot of attention in wide-ranging studies, and various techniques have been proposed for this. Signature-based techniques [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>] generally rely on identifying pre-computed hashes of existing malware binaries and a database of known threats that must be frequently updated. However, these approaches are unable to detect unknown botnets and variations because of polymorphous attacks, zero-day attacks, and related techniques. Anomaly-based detection algorithms [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-11">11</xref>] are based on the idea that botnets have communication patterns that are distinct from benign hosts in networks. These are commonly used for botnet detection since they can identify unknown botnets based on the number of anomalies in network traffic, large traffic volumes, traffic on unusual ports, high network latency, and other unusual system behavior. The main drawbacks of anomaly-based detection approaches are the high rate of false alarms and the restrictions of their training data. Other detection approaches are based on honeypot technology [<xref ref-type="bibr" rid="ref-12">12</xref>], which is a network set up with intentional vulnerabilities used as a trap without exposing the real network, but which only detects existing bots and performs poorly in real-time. Community-based anomaly detection algorithms [<xref ref-type="bibr" rid="ref-13">13</xref>] can be defined as techniques that detect a group of highly correlated and highly interactive nodes that interact unusually frequently. These approaches cannot accurately determine botnets when entire communication graphs are unavailable. Detection methods that are based on certain structures and protocols [<xref ref-type="bibr" rid="ref-14">14</xref>] can&#x2019;t find botnets that have different structures or protocols.</p>
<p>Botnet detection techniques based on ML have been considered the most effective. Feature extraction is an important step in ML techniques that involves about 80&#x0025; of the time before processing data, which helps to minimize data dimensionality and improve the accuracy of ML models [<xref ref-type="bibr" rid="ref-15">15</xref>]. Flow-based features are the features most often focused on in bot detection research. However, this approach suffers from limitations such as not completely capturing the communication patterns that can expose additional aspects of malicious hosts, involves significant computational cost, and can potentially be evaded by tweaking behavioral characteristics, such as changing packet structures [<xref ref-type="bibr" rid="ref-16">16</xref>]. In order to overcome these limitations in flow-based features, we used graph-based features derived from flow-level information to reflect the true behavior of hosts. Furthermore, these properties enable the models to combine inputs from several networks and enhance spatial stability [<xref ref-type="bibr" rid="ref-1">1</xref>]. We think that adding graph-based features to ML makes it more resistant to attacks from unknown sources and communication patterns that are hard to understand [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p>The contributions of this work are as follows:
<list list-type="simple">
<list-item><label>1)</label><p>By proposing BotSward, an effective graph-based bot detection system that converts NetFlow traffic features to graph features, we demonstrate that our system is resistant to zero-day attacks and suitable for large datasets.</p></list-item>
<list-item><label>2)</label><p>Using Centrality Measures to extract valuable graph-based features that achieve high precision in bot detection while consuming less time.</p></list-item>
<list-item><label>3)</label><p>Using the new graph-based features noted above, we designed an ML-based detection system that can detect botnets with an accuracy of 99&#x0025;.</p></list-item>
<list-item><label>4)</label><p>We train and validate our graph-based botnet detection approach on real CTU-13 botnet datasets.</p></list-item>
<list-item><label>5)</label><p>We test and compare the efficacy of our proposed graph-based features with flow-based features from other studies using the same datasets and ML algorithms.</p></list-item>
<list-item><label>6)</label><p>We test and compare the efficacy of our proposed ML algorithms with different ML algorithms from different studies using the same datasets and graph-based features.</p></list-item>
</list></p>
<p>Organization. The remainder of this work is organized in the following manner. In Section 2, we review the background and related work, while the methodology and system design of this paper, including the graph algorithmic definitions and ML analysis techniques, are outlined in Section 3. In Section 4, we introduce the evaluation metrics, including the datasets, tools and instruments, and performance metrics. In Section 5, we present the results of our proposed botnet detection for CTU-13 datasets, followed by a summary discussion. In Section 6, we discuss the comparison evaluation approach and state-of-the-art studies. In Section 7, we give our conclusions and final thoughts. Before moving further, <xref ref-type="table" rid="table-1">Tab. 1</xref> lists down the acronyms used in this study for a better understanding.</p>
<table-wrap id="table-1"><label>Table 1</label><caption><title>List of abbreviations used in this study</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Acronyms</th>
<th align="left">Used for</th>
<th align="left">Acronyms</th>
<th align="left">Used for</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">ML</td>
<td align="left">Machine learning</td>
<td align="left">GPUs</td>
<td align="left">Graphics processing units</td>
</tr>
<tr>
<td align="left">CTU-13</td>
<td align="left">Czech technical university13 dataset</td>
<td align="left">DNN</td>
<td align="left">Deep neural network</td>
</tr>
<tr>
<td align="left">BC</td>
<td align="left">Betweenness centrality</td>
<td align="left">SVM</td>
<td align="left">Support vector machine</td>
</tr>
<tr>
<td align="left">CC</td>
<td align="left">Closeness centrality</td>
<td align="left">DT</td>
<td align="left">Decision tree</td>
</tr>
<tr>
<td align="left">PR</td>
<td align="left">PageRank</td>
<td align="left">KNN</td>
<td align="left">K-nearest neighbor</td>
</tr>
<tr>
<td align="left">ID</td>
<td align="left">In degree</td>
<td align="left">RF</td>
<td align="left">Random forest</td>
</tr>
<tr>
<td align="left">OD</td>
<td align="left">Out degree</td>
<td align="left">SGD</td>
<td align="left">Stochastic gradient descent</td>
</tr>
<tr>
<td align="left">AC</td>
<td align="left">Alpha centrality</td>
<td align="left">ACC</td>
<td align="left">Accuracy</td>
</tr>
<tr>
<td align="left">EC</td>
<td align="left">Eigenvector centrality</td>
<td align="left">F</td>
<td align="left">F-measure</td>
</tr>
<tr>
<td align="left">LCC</td>
<td align="left">Local clustering coefficient centrality</td>
<td align="left">P</td>
<td align="left">Precision</td>
</tr>
<tr>
<td align="left">C&#x0026;C</td>
<td align="left">Command-and-control</td>
<td align="left">R</td>
<td align="left">Recall</td>
</tr>
<tr>
<td align="left">DDoS</td>
<td align="left">Distributed denial-of-service</td>
<td align="left">SMOTE</td>
<td align="left">Synthetic minority oversampling technique</td>
</tr>
<tr>
<td align="left">IDS</td>
<td align="left">Intrusion detection systems</td>
<td align="left">T-links</td>
<td align="left">Tomek links</td>
</tr>
<tr>
<td align="left">TPUs</td>
<td align="left">Tensor processing units</td>
<td align="left">HTTP</td>
<td align="left">Hypertext transfer protocol</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2"><label>2</label><title>Background and Related Work</title>
<p>In general, botnet detection techniques rely on deep packet inspection (DPI) and flow-based and graph-based botnet detection methods.</p>
<sec id="s2_1"><label>2.1</label><title>Deep Packet Inspection</title>
<p>DPI is an advanced technique for examining the full content of data packets as they pass through a monitored network checkpoint. It is highly effective in preventing overflow attacks, denial of service (DoS) attacks, buffer overflow attacks, and even some forms of malware. However, the existing botnet detection systems, which depend on DPI, suffer from high computational costs and can be exploited by crooks to facilitate similar attacks, as well as being inefficient at recognizing unknown payload signatures. Moreover, most new malware uses evasion techniques to avoid detection, such as protocol encapsulation, obfuscation, and payload encryption. As well, inspecting every packet on a high-speed network is a costly task because of the continuously increasing network speed and the daily increases in the amount of data transferred on the network [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
<p>Gadelrab&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-19">19</xref>] presented BotCap, a botnet detection model based on DPI and ML techniques. The authors inspected network traffic packets in-depth in order to extract statistical features and then trained J48 decision tree and Support Vector Machine (SVM) algorithms to distinguish between benign traffic and malicious botnet traffic. The study&#x2019;s results showed an accuracy level of 80&#x0025; for HTTP botnets and 95&#x0025; for IRC botnets.</p>
</sec>
<sec id="s2_2"><label>2.2</label><title>Flow-Based Detection</title>
<p>Flow-based detection is a network protocol that collects IP network traffic as it enters or exits an interface. Flow-based features are described as a collection of features based on information that exists in the headers of packets with similar parameters, such as source and destination IP, protocol type, source, and destination port, etc. [<xref ref-type="bibr" rid="ref-1">1</xref>]. Flow-based features are used to identify anomalies like botnets in large-volume data and high-speed networks. Extensive studies have made major contributions in the area of botnet detection using flow-based features; however, there are serious drawbacks to these approaches. Firstly, the approaches only capture the characteristics of the effects of bots on individual links, rather than using the topological structure of the communication graph as a whole, which can expose additional aspects of malicious hosts. Secondly, flow-based detection methods suffer from high computational overhead, since instead of holistically monitoring a given network&#x2019;s behaviors to identify malicious traffic, it requires the comparison of each particular flow of traffic to all the other flows. Finally, flow-based detection techniques can be evaded by attackers through data volume changes, using encrypted commands, or tweaking behavioral characteristics by changing the packet structure [<xref ref-type="bibr" rid="ref-20">20</xref>].</p>
<p>In 2014, Beigi&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-21">21</xref>] discussed the effectiveness of flow-based features in combination with a C4.5 decision tree-supervised learning algorithm for botnet detection. The authors achieved a moderate detection rate, at 75&#x0025;, and a low false-positive rate of only 2.3&#x0025;. Gahelot&#x00A0;et&#x00A0;al.&#x00A0;[<xref ref-type="bibr" rid="ref-22">22</xref>] proposed a J48 decision tree and na&#x00EF;ve Bayes classifiers to classify flow-based P2P botnet traffic using a combined network flow PeerRush dataset and a botnet (2014) dataset. When using a J48 decision tree, their botnets detection accuracy was 99.94&#x0025;, with a very low rate of false positives, and when using na&#x00EF;ve Bayes classifiers, the accuracy was 99.46&#x0025;. The main disadvantages of their approach, however, were the complexity of the model and the long processing runtime.</p>
<p>To overcome these limitations, using graph-based features to detect botnets has been the focus of another research path. Using graph-based features is more computationally efficient compared to using flow-based methods because it avoids the requirement to compare each particular traffic flow to all the other flows in the dataset [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
</sec>
<sec id="s2_3"><label>2.3</label><title>Graph-Based Detection</title>
<p>Graph-based features are derived from flow-based features and reflect the real structure of host behavior, interactions, and communications, offering an option that overcomes flow-based limitations. This approach is robust and suitable for all botnet datasets because it attempts to create a graph from the captured traffic, which discards all features except IP address features to create nodes and packet features to represent directed edges. By using centrality-based graph measurements, data from packets is ignored in favor of focusing on the topological communications structure between hosts. Then, from each node, we can calculate centrality-based graph measures and extract features like DC, CC, and PR &#x2026; etc.</p>
<p>The authors of [<xref ref-type="bibr" rid="ref-16">16</xref>] proposed BotChase, a hybrid two-phase ML approach that leverages both unsupervised and supervised ML and graph-based bot detection systems. They modeled network communications as graphs, where hosts are vertices and communications between hosts are edges. However, the model incurs a high computational overhead when computing the features of a large communication graph, such as the betweenness centrality (BC) measures used by shortest-path algorithms computed for centrality measures.</p>
<p>By combining flow-based features with graph-based network traffic in [<xref ref-type="bibr" rid="ref-23">23</xref>], the study detects bots based on the hybrid analysis of these two types of features. This technique achieved a good performance, with a 96.62&#x0025; F-score.</p>
<p>BotGM [<xref ref-type="bibr" rid="ref-24">24</xref>] suggested an unsupervised graph mining tool for detecting bots through abnormal communication patterns. For each pair of source and destination IP addresses, the authors first create a graph sequence of ports, then compare each graph using the Graph-Edit Distance (GED). They achieve a very good level of accuracy, ranging from 78&#x0025; to 95&#x0025;. However, the GED is computed only once for each pair of graphs, and its computation is known to be NP-complete. BotHunter [<xref ref-type="bibr" rid="ref-25">25</xref>] is designed to detect the infection and coordinate communication that takes place after a successful malware infection. CAMNEP [<xref ref-type="bibr" rid="ref-26">26</xref>] combines various state-of-the-art anomaly detection methods to improve accuracy. However, some communication patterns between hosts that are unique to a botnet are missed by these approaches. Furthermore, since the complexity of computing per-flow features is significantly high, BClus clusters the traffic sent by each IP address based on their behavioral similarity.</p>
</sec>
<sec id="s2_4"><label>2.4</label><title>Research Gap</title>
<p>Botnets have evolved into one of the most significant cyber security concerns for networks because they impact numerous areas such as law enforcement, finance, cyber security, health care, and more. The majority of current botnet detection research relies on flow-based traffic analysis and mining communication patterns. However, these methods may not be capable of detecting bot activities in an efficient and effective manner. Existing flow-based methods have significant computational overhead and do not capture all network communication patterns, which might reveal additional aspects of malicious hosts. Moreover, flow-based features are often characteristic of specific protocol-based botnets and are not generalizable to newer botnet types. A graph-based approach is rather intuitive for overcoming flow-based approaches&#x2019; limitations, as graphs are true representations of network communications. The drawback of graph-based detection systems compared to flow-based detection systems is the time required to extract features such as BC computation, which can sometimes be considerable. Although a big improvement was made over the initial BC computation algorithm, many researchers argued that the BC algorithm is still too costly for large graphs.</p>
<p>To solve this problem, some graph-based feature algorithms such as BC and CC are easily parallelizable, as using more cores will result in speed improvements. Also, we obtain superior results in terms of time reduction when we exclude BC and replace it with CC. The development of a quick and non-rule-based technique for detecting botnets is a significant step toward building a new graph-based detection methodology. Simultaneously, the approach must be validated on a real-world dataset with various botnet types. This detection scheme must also be sufficiently robust to detect any type of botnet present in the dataset. We show that incorporating graph-based features into ML yields robustness against complex communication patterns and unknown attacks. Furthermore, cross-network ML model training and inference are possible.</p>
</sec>
</sec>
<sec id="s3"><label>3</label><title>System Design</title>
<p>We illustrate our proposed system, BotSward, as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. Botsward is a classification model for botnet detection, where the current botnet anomaly detection approach based on NetFlow features can potentially be evaded by attackers through data volume changes, using encrypted commands, or tweaking behavioral characteristics by changing the packet structure. This problem is known as misleading. It works in four steps: traffic generation, data preprocessing, model building, and bot identification. In the first step, we collect a very large amount of network traffic flow from the real-world public CTU-13 dataset discussed in Section 4. In the second step, we balance the distribution of the dataset classes and Select 3 features from 42 features, and then clean the data to remove duplicate or irrelevant data. In the third step, the system creates a graph G(V; E) from the features extracted, where V is a set of nodes and E is a set of directed edges. After that, we extract seven graph-based features from network traffic to characterize the behavior of the botnets and save them in a file to use in the model training phase. In the model training, we compared the results of seven different ML models and a Deep Neural Network (DNN), and in the last step, we identified bot traffic and normal traffic and then calculated the performance of the different models.</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>BotSward architecture</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_31641-fig-1.png"/></fig>
<p>In this work, the software implementation is primarily based on Python. We used Jupyter Notebook, which is an open-source application that facilitates data visualization, data preprocessing, ML, statistical modeling, and much more, as well as Google Colab Notebook, which is hosted on Google cloud servers and provides access to the Graphics Processing Units (GPUs) and Tensor Processing Units (TPUs) for tasks that can be done in a Jupyter notebook. Python 3.6 was used to implement the base classifiers and was used as the scripting language for writing most of the code. For this study, we used different important libraries such as Networkx, Scikit-learn, Keras, Tensorflow, Pandas, NumPy, Matplotlib, and Seaborn [<xref ref-type="bibr" rid="ref-27">27</xref>]. The main library in our work to compute the centrality of the graph features is NetworkX, which is a Python library for studying graphs and networks, which are mathematical structures used to model pairwise relations between objects. The experiment was carried out on a Windows 10 Pro workstation with an Intel(R) Core (TM) i9-8950HK CPU @ 2.90&#x2005;GHz, a 64-bit operating system, and 32.0&#x2005;GB of RAM.</p>
<sec id="s3_1"><label>3.1</label><title>Data Preprocessing</title>
<p>Preprocessing is crucial for developing a reliable and effective ML-based detection model since it enhances classification accuracy. This phase has three steps, which are: balanced dataset, feature selection, and data cleaning.</p>
<sec id="s3_1_1"><label>3.1.1</label><title>Balanced Dataset</title>
<p>The term &#x201C;balanced dataset&#x201D; refers to balancing the distribution of dataset classes, as in an unbalanced dataset class model is biased towards the majority class; in such a case, the accuracy percentage obtained may be illusory [<xref ref-type="bibr" rid="ref-28">28</xref>], Oversampling, undersampling, and hybrid approaches are the most common dataset class resampling techniques. The oversampling approach is used when the quantity of data is insufficient [<xref ref-type="bibr" rid="ref-29">29</xref>]. It mainly involves replicating the rare samples in some classes in order to equalize the distribution of classes in an unbalanced dataset. The most commonly utilized form of oversampling is the synthetic minority oversampling technique (SMOTE) [<xref ref-type="bibr" rid="ref-30">30</xref>]. The undersampling approach [<xref ref-type="bibr" rid="ref-31">31</xref>], in contrast to the oversampling approach, is used to balance dataset classes when the quantity of data is sufficient. The aim of this technique is to reduce the size of the abundant class by randomly selecting an equal number of samples from that class in order to equalize the number of samples in the rare class. For further modeling, a new, balanced dataset is retrieved. Tomek links (T-links) are a common undersampling approach for balancing dataset classes. Hybrid methods use a combination of oversampling and undersampling methods to eliminate the lack of training data produced by undersampling and to prevent overfitting problems caused by oversampling. In this work, the T-links approach was utilized to balance the datasets in both training and validation since the quantity of data was sufficient.</p>
</sec>
<sec id="s3_1_2"><label>3.1.2</label><title>Feature Selection</title>
<p>The process of automatically or manually selecting the features that contribute the most to the prediction variable or output in which you are interested is known as feature selection. Irrelevant features in a dataset may reduce model accuracy and lead the model to train on irrelevant features [<xref ref-type="bibr" rid="ref-32">32</xref>]. There are several advantages to selecting features before modeling data [<xref ref-type="bibr" rid="ref-33">33</xref>], including:
<list list-type="simple">
<list-item><p>&#x22C5;Reducing training time: fewer data processing allows algorithms to train faster and reduces algorithm complexity.</p></list-item>
<list-item><p>&#x22C5;Reducing overfitting: fewer data processing means less opportunity to make decisions based on noise.</p></list-item>
<list-item><p>&#x22C5;Improving accuracy: fewer misleading data processing means that modeling accuracy improves. To reduce the dimensionality of the data, we selected three features, which are Source IP, Destination IP, and Total Packets.</p></list-item>
</list></p>
</sec>
<sec id="s3_1_3"><label>3.1.3</label><title>Data Preprocessing</title>
<p>The first and most crucial step in creating a ML model is preprocessing the data [<xref ref-type="bibr" rid="ref-34">34</xref>]. This step is important for improving data quality for ML module training and more accurate decision-making. Bad data might lead to inaccurate results, and so to get the appropriate results from the beginning, the dataset is examined and formatting is done. In this study, data cleansing, normalization, and transformation are utilized to create a reliable dataset [<xref ref-type="bibr" rid="ref-35">35</xref>].
<list list-type="simple">
<list-item><label>1)</label><p>Data Cleansing: Data cleansing is the process of identifying incorrect, irrelevant, inaccurate, or incomplete parts of the data and then deleting, replacing, or modifying the dirty or coarse data. In our work, we drop the rows in the CTU-13 dataset that have null values by using the dropna () function of Pandas.</p></list-item>
<list-item><label>2)</label><p>Normalization: Normalization is a scaling technique that provides a consistent scale in which values are shifted and re-scaled so that they end up ranging between 0 and 1. The following is the mathematical equation for min-max feature scaling:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /><mml:mn>0</mml:mn><mml:mspace width="thickmathspace" /><mml:mo>&#x2264;</mml:mo><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2264;</mml:mo><mml:mn>1</mml:mn><mml:mspace width="thickmathspace" /></mml:math></disp-formula></p></list-item>
</list></p>
<p>We get <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mo>{</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> by using <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mo>.</mml:mo><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mspace width="thickmathspace" /><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mo>.</mml:mo><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mspace width="thickmathspace" /><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> functions of pandas.
<list list-type="simple">
<list-item><label>3)</label><p>Transformation: Transformation is the process of converting data from one format to another. Many categorical features in the CTU-13 dataset include non-numeric data that needed to be translated to numeric format for the ML algorithms&#x2019; analyses to deal with the algebraic format. Label Encoder is a technique that many data analysts use to do label encoding. We use the SciKit-learn library to import it.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s3_2"><label>3.2</label><title>Build Model</title>
<p>In the third step, we build the model, including a graph-based detector and an ML model.</p>
<sec id="s3_2_1"><label>3.2.1</label><title>Graph-Based Detector</title>
<p>The traffic flows that forward to our system BotSward are bidirectional network flows. These flows are transformed into a set <italic>T</italic> that includes 4-tuple flows <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mrow><mml:mtext mathvariant="italic">ScrAddr</mml:mtext></mml:mrow><mml:mo>;</mml:mo><mml:mrow><mml:mtext mathvariant="italic">DstAddr</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext mathvariant="italic">TotPkts</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Label</mml:mtext></mml:mrow><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. Where <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mrow><mml:mtext mathvariant="italic">ScrAddr</mml:mtext></mml:mrow></mml:math></inline-formula> is the source host <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>I</mml:mi><mml:mi>P</mml:mi></mml:math></inline-formula> address that uniquely identifies a source host known as <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mrow><mml:mtext mathvariant="italic">DstAddr</mml:mtext></mml:mrow></mml:math></inline-formula> is the destination host <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>I</mml:mi><mml:mi>P</mml:mi></mml:math></inline-formula> address that uniquely identifies a destination host known as <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mrow><mml:mtext mathvariant="italic">TotPkts</mml:mtext></mml:mrow></mml:math></inline-formula> quantifies the number of data packets sent by source host <italic>i</italic> to destination host <italic>j</italic>, <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mrow><mml:mtext mathvariant="italic">Label</mml:mtext></mml:mrow></mml:math></inline-formula> is the output class &#x201C;Bot&#x201D; or &#x201C;Normal&#x201D;. The system creates a graph <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>G</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>V</mml:mi><mml:mo>;</mml:mo><mml:mi>E</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <italic>V</italic> is the set of nodes <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, whereas E is the set of directed edges <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, for all i; j such that <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>V</mml:mi></mml:math></inline-formula>. In this study, each node denotes a unique IP address and each edge denotes the connection between one IP address to another. Set <italic>A</italic> is a set of tuples that have exclusive source and destination hosts. Where <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>A</mml:mi></mml:math></inline-formula>, such that <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mo>;</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext mathvariant="italic">TotPkts</mml:mtext></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>. The set of nodes <italic>V</italic> is a union of source and destination hosts from set <italic>A</italic> such that
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x222A;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:mrow></mml:munder><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mo>&#x222A;</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>For every <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> in A, There exist directed edges <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mrow><mml:mtext>e</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow><mml:mo>;</mml:mo><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mrow><mml:mtext>e</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow><mml:mo>;</mml:mo><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> from <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> to <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> to <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, respectively, such that <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mrow><mml:mtext>sip</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>i</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mrow><mml:mtext>dip</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>j</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. Therefore,
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mrow><mml:mi>E</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x222A;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:mrow></mml:munder><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x222A;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mi>s</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:msub><mml:mrow><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula></p>
<sec id="s3_2_1_1"><title>Graph Algorithmic Properties</title>
<p>Understanding networks requires the use of centrality measurements, also known as graphs. These algorithms use graph theory to calculate the importance of any given node in a network. It can capture all network communication patterns and disregard packet payloads, focusing on the topological communications structure between hosts. Every graph feature has a degree of importance, so you need to understand how they work to find the best one for your graph visualization applications. We present and discuss the following set of important Centrality Measures for Graph-Based features that are widely utilized which are DC, BC, CC, Alpha Centrality (AC), local clustering coefficient (LCC), and Eigenvector centrality (EC) Where BotSward replaced BC With CC.</p>
<p><bold>Definition 1 (Indegree and outdegree):</bold></p>
<p>In Degree (ID) can be described as the total number of head ends into a particular node coming from adjacent nodes in a directed graph (arrows pointing towards the node). A high ID value of node implies that the node is the crucial and pivotal node which could be C&#x0026;C servers or bots, where the adjacent nodes&#x2019; tendency to create additional connections, whereas a low value implies the opposite. In contrast, The out-degree (OD) is the total number of the tail of the edge ends going out of a particular node to adjacent nodes in a directed graph (arrows pointing outwards from the node). A high OD number for a node indicates that it makes more connections with other adjacent nodes, whereas a low value indicates the opposite. Bots tend to create additional connections with other possible victim machines in order to expand the botnet&#x2019;s reach or communicate with the C&#x0026;C domain for transferring information. As a result, on a graph, OD can be a good indicator of botnet activity. In degree weight (IDW) indicates the total amount of data packets received by a particular node from its adjacent linked nodes. Whereas the total amount of data packets sent by a particular node to its adjacent linked nodes is known as the out-degree weight (ODW) [<xref ref-type="bibr" rid="ref-16">16</xref>]. There are two types of graphs as directed and undirected graphs, Directed graphs have edges with direction whereas Undirected graphs have edges that do not have a direction. Let <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>G</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>E</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thickmathspace" /></mml:math></inline-formula> be a graph where <italic>n</italic> is total number of nodes in the graph and <italic>v</italic> in <italic>V</italic>, The mathematical expression for node degree is:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:mi>deg</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mn>2</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>In Direct graph</mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>In undirect graph</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>In direct graph G&#x2009;&#x003D;&#x2009;(V, A) where V is set of vertices and A is a set of ordered pairs of vertices,</p>
<p><inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:msup><mml:mi>g</mml:mi><mml:mrow><mml:mspace width="thickmathspace" /><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> denotes to ID and <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:msup><mml:mi>g</mml:mi><mml:mrow><mml:mspace width="thickmathspace" /><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> denotes OD, G is balanced if and only if the in-and out-edges of each vertex have the same weight. Thus the graph is called a balanced directed graph when:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:munderover><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:msup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi mathvariant="bold">V</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:munderover><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:mrow><mml:msup><mml:mrow><mml:mi>g</mml:mi></mml:mrow><mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><xref ref-type="fig" rid="fig-2">Fig. 2</xref> shows the example of ID and OD, ID of vertex <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:math></inline-formula> is number of edges termination from vertex <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>v</mml:mi><mml:mspace width="thickmathspace" /><mml:mn>0</mml:mn></mml:math></inline-formula>, whereas OD of vertex <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:math></inline-formula> is number of edges starting from vertex <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>v</mml:mi><mml:mn>1</mml:mn><mml:mspace width="thickmathspace" /></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<fig id="fig-2"><label>Figure 2</label><caption><title>In-degree and OD example</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_31641-fig-2.png"/></fig>
<p><bold>Definition 5 (Betweenness centrality):</bold></p>
<p>BC of a node <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>V</mml:mi></mml:math></inline-formula> is a measure of centrality in graph theory based on the number of shortest paths from all nodes to all others that pass through that node. The mathematical expression for node BC is:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mi>f</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:msub><mml:mi></mml:mi><mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo>&#x2260;</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>&#x2260;</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /><mml:mspace width="thickmathspace" /></mml:mrow></mml:munderover><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x03C3;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:msub><mml:mrow><mml:mi>&#x03C3;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the total number of shortest paths from node pairs <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>V</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the number of shortest paths that pass-through node <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /></mml:math></inline-formula> this feature has a high computational overhead with time complexity [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>].
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:mi mathvariant="bold">O</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="bold">V</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>.</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="bold">E</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>.</mml:mo><mml:mrow><mml:mi>log</mml:mi></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="bold">V</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Node BC can be a valuable feature for detecting botnets, particularly in P2P botnets when bots are more interconnected and there is no central C&#x0026;C structure. However, it has the potential to alienate bots when they make their first connections, when the bots&#x2019; IDW and ODW are low. As a result, it would be more favorable for the network&#x2019;s shortest routes to pass via the host. There is an inverse relationship between node&#x2019;s DC and node&#x2019;s BC, when the IDW and ODW increase, the BC of a node decreases immensely, as it is less favored for being included in shortest paths.</p>
<p><bold>Definition 6 (Local Clustering Coefficient):</bold></p>
<p>LCC of a vertex (node) in a graph indicates how close its neighbors are to being a complete graph connected (clique). It has a lower computational overhead. The mathematical expression of the LCC for node a can be given by:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:mi>L</mml:mi><mml:mi>C</mml:mi></mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>v</mml:mi><mml:mspace width="thickmathspace" /></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2.</mml:mn><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>where node known as <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi>v</mml:mi><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> indicates the number of node&#x2019;s neighbors <italic>v</italic> and <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> indicates the number of link connected pairs between all neighbors of node <italic>v</italic> where the output always between <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mn>0</mml:mn><mml:mo>&#x226A;</mml:mo><mml:mi>L</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x226A;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>L</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> feature can play a significant impact to distinguish malicious hosts behavior especially in P2P botnets detection. As bots, successfully infected hosts have a greater <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mi>L</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p><bold>Definition 7 (PageRank algorithm):</bold></p>
<p>PR is a Google Search algorithm that ranks web pages in search engine results and way of measuring the importance of website pages. According to Google, PR calculates the importance of a website by measuring the quantity and quality of links that point to this page. Intuitively, A node that is often connected frequently is important, and nodes connected to important nodes are also considered important. This rule corresponds to bots and C&#x0026;C servers in botnet detection. The definition for PR score of node <italic>u</italic> as
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mrow><mml:mi>P</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:munderover><mml:mfrac><mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>where <italic>d</italic> is the damping factor, <italic>N</italic> is the total number of nodes in the network, <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi>u</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>v</mml:mi></mml:math></inline-formula> are web pages. <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>n</mml:mi><mml:mi>b</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>u</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents all the neighbor nodes of node <italic>u</italic> [<xref ref-type="bibr" rid="ref-23">23</xref>].</p>
<p><bold>Definition 8 (Eigenvector centrality):</bold></p>
<p>EC is a measurement criterion of the influence of a node in a graph. The weight of a node in a graph is effective and important. The score of a node is influenced more by links to high-scoring nodes than by connections to low-scoring nodes, therefore each node is assigned a relative value. The EC of a node is determined by the total of the EC of all nodes connected to it. In other words, a node with a high EC is linked to other nodes with a high EC. Let, <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> be the adjacency matrix where
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if node</mml:mtext></mml:mrow><mml:mspace width="thickmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thickmathspace" /><mml:mrow><mml:mtext>is linked to node</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mo>,</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if node</mml:mtext></mml:mrow><mml:mspace width="thickmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mspace width="thickmathspace" /><mml:mrow><mml:mtext>is not linked with node</mml:mtext></mml:mrow><mml:mspace width="thickmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Then the centrality score can be given as:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>&#x03BB;</mml:mi></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /></mml:mrow></mml:munderover><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>w</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>M</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the set of neighbors of node <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>v</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>&#x03BB;</mml:mi></mml:math></inline-formula> is a constant. Now <xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref> can be</p>
<p>rewritten as
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mrow><mml:mi>A</mml:mi><mml:mi>x</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi>&#x03BB;</mml:mi></mml:mrow><mml:mrow><mml:mspace width="thickmathspace" /><mml:mi>x</mml:mi></mml:mrow></mml:math></disp-formula></p>
<p>Using the power technique and the Perron&#x2013;Frobenius theorem, a positive solution &#x03BB; with the last eigenvector exists. &#x03BB; is also the largest eigenvalue related with the adjacency matrix&#x2019;s eigenvector. EC is an extension of DC, where DC provides equal scores for each node connected [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p><bold>Definition 9: (Alpha centrality):</bold></p>
<p>The concept of AC was inspired by social network research, EC refers to a node&#x2019;s relative weight in the network, with connections to high-scoring nodes contributing more to the <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> score [<xref ref-type="bibr" rid="ref-16">16</xref>]. Hence, AC is given as
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow><mml:msubsup><mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>e</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula> is influence factor that controls focus between external sources to internal influence, <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the adjacency matrix, and <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the external influence of node <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p><bold>Definition 10: (Closeness centrality)</bold></p>
<p>CC is a method of detecting nodes that eligible efficiently transmit information throughout a graph. The concept of CC is based on the average farness (inverse distance) of a node to all other nodes. Nodes with a high closeness score have the shortest distances to all other nodes to spread information quickly, these nodes in the graph have important influence of the network. The mean distance between a vertex and other vertices is measured by CC [<xref ref-type="bibr" rid="ref-36">36</xref>].</p>
<p>For node <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, the closeness is calculated as the average shortest path between that node and all other nodes in the graph <italic>G</italic>. This is, let <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> be the shortest path between <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, the closeness is calculated as
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msub><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>n</mml:mi></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:munderover><mml:mrow><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi>u</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>u</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the shortest-path distance between <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mspace width="thickmathspace" /><mml:mi>v</mml:mi></mml:math></inline-formula> and <italic>u</italic>, and n is the number of nodes that can reach <italic>u</italic>.</p>
</sec>
<sec id="s3_2_1_2"><title>Graph Features Importance</title>
<p>Feature importance refers to a class of techniques for assigning scores to input features to a predictive model that indicates the relative importance of each feature when making a prediction. In BotSward, we applied all features discussed above and we found CC is the highest relative importance by 40&#x0025;, followed by PR, which has a relative importance of nearly 30&#x0025;. Whereas BC and LCC are the lowest features importance. Thus, BotSward discarded the BC and the LCC and applied the following set of important features relative which are DC, CC, and PR. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> presents all the relative importance of each graph feature.</p>
<fig id="fig-3"><label>Figure 3</label><caption><title>Graph features importance</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_31641-fig-3.png"/></fig>
</sec>
</sec>
<sec id="s3_2_2"><label>3.2.2</label><title>Machine Learning Analysis Techniques</title>
<p>The features generated previously are crucial and are used by ML and DL models for the classification of botnets. To train and evaluate these models, we used the CTU-13 dataset, large capture of real botnet traffic mixed with normal traffic and background traffic. The models that were trained in this study are Random Forest (RF), Gradient Boosting Classifier (GBC), Logistic Regression (LR), Stochastic Gradient descent (SGD), Decision Tree (DT), K-Nearest Neighbor (KNN), SVM, and DNN [<xref ref-type="bibr" rid="ref-37">37</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>].</p>

</sec>
</sec>
</sec>
<sec id="s4"><label>4</label><title>Performance Evaluation</title>
<sec id="s4_1"><label>4.1</label><title>Dataset</title>
<p>We used the CTU-13 dataset, which is one of the largest publicly available labeled datasets and includes botnet traffic mixed with normal traffic and background traffic [<xref ref-type="bibr" rid="ref-39">39</xref>]. It was developed in 2011 at the CTU and distinguishes 13 scenarios for different botnet attacks. In CTU-13, there are various types of botnet samples, such as IRC, Port Scan, Click Fraud Spam traffic, Fast Flux, and DDoS attacks [<xref ref-type="bibr" rid="ref-1">1</xref>]. The distinctive characteristics of the CTU13 dataset are that there are real botnet attacks, real-world traffic, and multiple types of botnets. Firstly, to have a clear idea about our analysis, we concatenate all 13 scenarios&#x2019; datasets that include benign and bot traffic. For these categories, 60&#x0025; of the cases are chosen as training data, while the rest of the dataset, 40&#x0025;, is used as validation data. However, the dataset is high-dimensional and largely class-imbalanced. The botnet traffic distribution in the CTU-13 dataset is just 1.5&#x0025; of the entire network traffic. In this proposed work, the Tomeklinks under-sampling approach (T-links) is used for handling the data imbalance. <xref ref-type="table" rid="table-2">Tab. 2</xref> shows the distribution of normal and bot flows in each scenario and the botnet used.</p>
<table-wrap id="table-2"><label>Table 2</label><caption><title>Scenarios from the CTU-13 dataset used for running experiments</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Scenario</th>
<th align="left">Total flows</th>
<th align="left">Normal flows</th>
<th align="left">Bot flows</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">2,824,636</td>
<td align="left">39,933 (1.41&#x0025;)</td>
<td align="left">2,784,703 (98.58&#x0025;)</td>
</tr>
<tr>
<td align="left">1</td>
<td align="left">1,808,122</td>
<td align="left">1,787,181 (98.84&#x0025;)</td>
<td align="left">20,941 (1.16&#x0025;)</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">4,710,638</td>
<td align="left">4,683,816 (99.43&#x0025;)</td>
<td align="left">26,822 (0.57&#x0025;)</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">1,121,076</td>
<td align="left">1,118,496 (99.77&#x0025;)</td>
<td align="left">2,580 (0.23&#x0025;)</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">129,832</td>
<td align="left">128,931 (99.31&#x0025;)</td>
<td align="left">901 (0.69&#x0025;)</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">558,919</td>
<td align="left">554,289 (99.17&#x0025;)</td>
<td align="left">4,630 (0.83&#x0025;)</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">114,077</td>
<td align="left">114,014 (99.94&#x0025;)</td>
<td align="left">63 (0.06&#x0025;)</td>
</tr>
<tr>
<td align="left">8</td>
<td align="left">2,954,230</td>
<td align="left">2,948,103 (99.79&#x0025;)</td>
<td align="left">6,127 (0.21&#x0025;)</td>
</tr>
<tr>
<td align="left">9</td>
<td align="left">2,087,508</td>
<td align="left">1,902,521 (91.14&#x0025;)</td>
<td align="left">184,987 (8.86&#x0025;)</td>
</tr>
<tr>
<td align="left">10</td>
<td align="left">1,309,791</td>
<td align="left">1,203,439 (91.88&#x0025;)</td>
<td align="left">106,352 (8.12&#x0025;)</td>
</tr>
<tr>
<td align="left">11</td>
<td align="left">107,251</td>
<td align="left">99,087 (92.39&#x0025;)</td>
<td align="left">8,164 (7.61&#x0025;)</td>
</tr>
<tr>
<td align="left">12</td>
<td align="left">325,471</td>
<td align="left">323,303 (99.33&#x0025;)</td>
<td align="left">2,168 (0.67&#x0025;)</td>
</tr>
<tr>
<td align="left">13</td>
<td align="left">1,925,149</td>
<td align="left">1,885,146 (97.92&#x0025;)</td>
<td align="left">40,003 (2.08&#x0025;)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2"><label>4.2</label><title>Performance Metrics</title>
<p>A confusion matrix is a table depicting the performance measurement technique for the ML classification model by presetting and relating the botnet detection and mitigation scheme. A set of confusion matrix performance measures was used to evaluate the botnet detection models&#x2019; robustness and effectiveness in classifying normal and attack traffic, which requires a high detection rate, high accuracy, and low false alarm rate. It contains four primary values that are utilized to generate the performance indicators listed below with the following short explanation for this technique:
<list list-type="bullet">
<list-item><p>True Positive (TP): Total actual number of attack records correctly classified.</p></list-item>
<list-item><p>False Negative (FN): Total actual number of attack records incorrectly classified.</p></list-item>
<list-item><p>False-positive (FP): Total actual number of normal records incorrectly classified.</p></list-item>
<list-item><p>True Negative (TN): Total actual number of normal records correctly classified.</p></list-item>
</list></p>
<p>Based on the defined confusion matrix values, the most widely adopted four metrics namely, the most widely adopted four metrics namely, Accuracy (ACC), Precision (P), Recall (R), F-measure (F), were used as in most previous botnet detection literature [<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-41">41</xref>]. These metrics are defined as follows,
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mrow><mml:mi>A</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mrow><mml:mtext>sensitivity</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mi>F</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mn>2.</mml:mn><mml:mfrac><mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mo>.</mml:mo><mml:mrow><mml:mi>R</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>P</mml:mi></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mi>R</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
</sec>
<sec id="s4_3"><label>4.3</label><title>Results</title>
<p>This section summarizes the findings and contributions made, and presents a performance evaluation of the results obtained from the experiment. The purpose of this evaluation was to see how well the centrality measurement algorithms and ML-based classifications performed. The results demonstrate the effectiveness of the proposed BotSward in mitigating botnet attacks.</p>
<sec id="s4_3_1"><label>4.3.1</label><title>Performance of Preprocessing Experiment on CTU-13 Datasets</title>
<p>One concern about the findings for the graph features was about the execution time of the centrality measures. Thus, we did two experiments to calculate the execution time of these measures. The first experiment yielded that the compute all graph features discussed in Section 3 for determining the most time-consuming. We found that this method can deliver good results with high accuracy (reaching 99&#x0025;), but its calculation of BC was quite time-consuming and expensive. The second experiment yielded that the compute all graph features discard BC. We found that this method can deliver good results with high accuracy (reaching 99&#x0025;) and with low time consumption. <xref ref-type="table" rid="table-3">Tab. 3</xref> shows the comparison between the results of the two experiments. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows the Experimenting time of BotSward when using BC and after excluded BC.</p>
<table-wrap id="table-3"><label>Table 3</label><caption><title>Comparing of experimenting time with BC and without</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">DS</th>
<th align="left">Nodes</th>
<th align="left">All features with BC (Seconds)</th>
<th align="left">Result</th>
<th align="left">All features without BC (F)</th>
<th align="left">Result</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1</td>
<td align="left">8283</td>
<td align="left">85.8262</td>
<td align="left">99&#x0025;</td>
<td align="left">23.226</td>
<td align="left">100&#x0025;</td>
</tr>
<tr>
<td align="left">2</td>
<td align="left">3628</td>
<td align="left">22.5601</td>
<td align="left">98&#x0025;</td>
<td align="left">11.6547</td>
<td align="left">98&#x0025;</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">6345</td>
<td align="left">77.1845</td>
<td align="left">83&#x0025;</td>
<td align="left">34.8127</td>
<td align="left">83&#x0025;</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">2235</td>
<td align="left">8.0149</td>
<td align="left">94&#x0025;</td>
<td align="left">5.5411</td>
<td align="left">93&#x0025;</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">357</td>
<td align="left">0.1292</td>
<td align="left">93&#x0025;</td>
<td align="left">0.0602</td>
<td align="left">90&#x0025;</td>
</tr>
<tr>
<td align="left">6</td>
<td align="left">1471</td>
<td align="left">1.1609</td>
<td align="left">99&#x0025;</td>
<td align="left">0.3392</td>
<td align="left">99&#x0025;</td>
</tr>
<tr>
<td align="left">7</td>
<td align="left">280</td>
<td align="left">0.0939</td>
<td align="left">33&#x0025;</td>
<td align="left">0.055</td>
<td align="left">33&#x0025;</td>
</tr>
<tr>
<td align="left">8</td>
<td align="left">6385</td>
<td align="left">79.3971</td>
<td align="left">100&#x0025;</td>
<td align="left">27.6409</td>
<td align="left">100&#x0025;</td>
</tr>
<tr>
<td align="left">9</td>
<td align="left">5950</td>
<td align="left">96.5348</td>
<td align="left">98&#x0025;</td>
<td align="left">60.2349</td>
<td align="left">99&#x0025;</td>
</tr>
<tr>
<td align="left">10</td>
<td align="left">2342</td>
<td align="left">39.3802</td>
<td align="left">75&#x0025;</td>
<td align="left">33.6711</td>
<td align="left">75&#x0025;</td>
</tr>
<tr>
<td align="left">12</td>
<td align="left">735</td>
<td align="left">0.2778</td>
<td align="left">100&#x0025;</td>
<td align="left">0.1117</td>
<td align="left">89&#x0025;</td>
</tr>
<tr>
<td align="left">13</td>
<td align="left">4728</td>
<td align="left">61.5901</td>
<td align="left">99&#x0025;</td>
<td align="left">18.749</td>
<td align="left">99&#x0025;</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-4"><label>Figure 4</label><caption><title>Experimenting time</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_31641-fig-4.png"/></fig>
</sec>
<sec id="s4_3_2"><label>4.3.2</label><title>Performance of Classification Algorithms Experiment on CTU-13 Datasets</title>
<p>In order to ensure the best results from the classifier, grid searching can be applied across ML to discover the optimal hyperparameters for a model that provide the most &#x201C;accurate&#x201D; predictions. Grid-search will create a model based on each possible parameter combination. It iterates through every parameter combination and stores a model for each one. Several state-of-the-art binary classification methods are experimentally assessed in the context of botnet detection. Eight widely used classification algorithms are discussed, where the predictive powers of RF, GBC, LR, SGD, SVM, DT, and KNN are contrasted. For most of these algorithms, we report on several models using alternative hyperparameter optimization techniques to obtain the best parameters for a given model. <xref ref-type="table" rid="table-4">Tab. 4</xref> shows the results of the grid search along with the range of parameters used, and the best parameter values are highlighted [<xref ref-type="bibr" rid="ref-42">42</xref>].</p>
<table-wrap id="table-4"><label>Table 4</label><caption><title>Hyperparameter search comparison (Grid <italic>vs.</italic> Random)</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Classifier</th>
<th align="left" colspan="2">Random</th>
<th align="left" colspan="2">Grid search</th>
</tr>
<tr>
<th/>
<th align="left">Parameter</th>
<th align="left">ACC</th>
<th align="left">Optimal</th>
<th align="left">ACC</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">RF</td>
<td align="left">criterion&#x2009;&#x003D;&#x2009;&#x2018;gini, n_estimators&#x2009;&#x003D;&#x2009;700, max_depth&#x2009;&#x003D;&#x2009;6</td>
<td align="left">97&#x0025;</td>
<td align="left">Criterion&#x2009;&#x003D;&#x2009;&#x2018;entropy&#x2019; n_estimators&#x2009;&#x003D;&#x2009;300, max_depth&#x2009;&#x003D;&#x2009;8</td>
<td align="left">99&#x0025;</td>
</tr>
<tr>
<td align="left">GBC</td>
<td align="left">n_estimators&#x2009;&#x003D;&#x2009;100, max_depth&#x2009;&#x003D;&#x2009;3</td>
<td align="left">98&#x0025;</td>
<td align="left">n_estimators&#x2009;&#x003D;&#x2009;300 max_depth&#x2009;&#x003D;&#x2009;4</td>
<td align="left">99&#x0025;</td>
</tr>
<tr>
<td align="left">LR</td>
<td align="left">penalty&#x2009;&#x003D;&#x2009;l2, solver&#x2009;&#x003D;&#x2009;&#x2018;sag, C&#x2009;&#x003D;&#x2009;1.0, random_state&#x2009;&#x003D;&#x2009;33</td>
<td align="left">81&#x0025;</td>
<td align="left">Penalty&#x2009;&#x003D;&#x2009;&#x2018;none&#x2019;, Solver&#x2009;&#x003D;&#x2009;&#x2018;newton-cg&#x2019;, C&#x2009;&#x003D;&#x2009;0.01</td>
<td align="left">82&#x0025;</td>
</tr>
<tr>
<td align="left">SGD</td>
<td align="left">penalty&#x2009;&#x003D;&#x2009;l2, loss&#x2009;&#x003D;&#x2009;&#x2018;squaredloss, learning_rate&#x2009;&#x003D;&#x2009;&#x2018;optimal&#x2019; random_state&#x2009;&#x003D;&#x2009;33</td>
<td align="left">23&#x0025;</td>
<td align="left">penalty&#x2009;&#x003D;&#x2009;l2, eta0&#x2009;&#x003D;&#x2009;0.1, learning_rate&#x2009;&#x003D;&#x2009;&#x2018;constant&#x2019;</td>
<td align="left">50&#x0025;</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="left">kernel&#x2009;&#x003D;&#x2009;&#x2018;rbf&#x2019;, max_iter&#x2009;&#x003D;&#x2009;100, C&#x2009;&#x003D;&#x2009;1.0, Gamma&#x2009;&#x003D;&#x2009;&#x2018;auto&#x2019;</td>
<td align="left">72&#x0025;</td>
<td align="left">max_iter&#x2009;&#x003D;&#x2009;300, C&#x2009;&#x003D;&#x2009;0.1, Gamma&#x2009;&#x003D;&#x2009;&#x2018;auto&#x2019;</td>
<td align="left">76&#x0025;</td>
</tr>
<tr>
<td align="left">DT</td>
<td align="left">criterion&#x2009;&#x003D;&#x2009;&#x2018;gini&#x2019;, max_depth&#x2009;&#x003D;&#x2009;3, random_state&#x2009;&#x003D;&#x2009;33</td>
<td align="left">97&#x0025;</td>
<td align="left">Criterion&#x2009;&#x003D;&#x2009;&#x2018;gini&#x2019;, max_depth&#x2009;&#x003D;&#x2009;8</td>
<td align="left">99&#x0025;</td>
</tr>
<tr>
<td align="left">KNN</td>
<td align="left">n_neighbors&#x2009;&#x003D;&#x2009;5, weights&#x2009;&#x003D;&#x2009;&#x2018;uniform&#x2019;, algorithm&#x2009;&#x003D;&#x2009;&#x2018;auto</td>
<td align="left">92&#x0025;</td>
<td align="left">criterion&#x2009;&#x003D;&#x2009;&#x2018;gini&#x2019;, max_depth&#x2009;&#x003D;&#x2009;8</td>
<td align="left">94&#x0025;</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3_3"><label>4.3.3</label><title>Discussion</title>
<p>The results demonstrated in this study match state-of-the-art methods and go beyond previous reports, showing a better result than in other studies. Graph-based features in previous phases extract full communication patterns to enhance traffic classification for the learning algorithms. Our findings in preprocessing suggest that excluding the BC feature and calculating PR, EC, and CC saves around 69.5&#x0025; of time while maintaining the same high accuracy 99&#x0025; in the classification models. RF, GBC, LR, SGD, SVC, DT, KNN, and DNN classification success is attributed to their operational procedures and ability to select the optimal parameters using the grid-search technique. As a result, the SGD algorithm had the lowest values in all performance metrics, providing an accuracy of 50&#x0025; when compared to other algorithms due to SGD evaluating the error for each training example within the dataset. This means that the parameters for each training example are updated one by one. Thus, the frequency of updates causes noise in the gradients, which has an impact on convergence and affects the classification accuracy. The SVM algorithm had moderate results in all performance metrics, providing an accuracy of 76&#x0025; when compared to other algorithms due to the classifiers not being efficient with very large datasets. KNN scored higher than the LR by approximately 14.64&#x0025; in classification accuracy, and even outperformed SVM by 23.69&#x0025; for the CTU-13 dataset. The results in <xref ref-type="table" rid="table-5">Tab. 5</xref> prove that KNN performs better than LR SVM for classification precision in complex situations [<xref ref-type="bibr" rid="ref-43">43</xref>].</p>
<table-wrap id="table-5"><label>Table 5</label><caption><title>Performance comparison of ML algorithms</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Classifier</th>
<th align="left">Precision (P)</th>
<th align="left">Recall (R)</th>
<th align="left">F-score (F)</th>
<th align="left">Accuracy (ACC)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">RF</td>
<td align="left">99&#x0025;</td>
<td align="left">99&#x0025;</td>
<td align="left">99&#x0025;</td>
<td align="left">99&#x0025;</td>
</tr>
<tr>
<td align="left">GBC</td>
<td align="left">99&#x0025;</td>
<td align="left">99&#x0025;</td>
<td align="left">99&#x0025;</td>
<td align="left">99&#x0025;</td>
</tr>
<tr>
<td align="left">LR</td>
<td align="left">86&#x0025;</td>
<td align="left">82&#x0025;</td>
<td align="left">81&#x0025;</td>
<td align="left">82&#x0025;</td>
</tr>
<tr>
<td align="left">SGD</td>
<td align="left">75&#x0025;</td>
<td align="left">50&#x0025;</td>
<td align="left">34&#x0025;</td>
<td align="left">50&#x0025;</td>
</tr>
<tr>
<td align="left">SVM</td>
<td align="left">79&#x0025;</td>
<td align="left">76&#x0025;</td>
<td align="left">75&#x0025;</td>
<td align="left">76&#x0025;</td>
</tr>
<tr>
<td align="left">DT</td>
<td align="left">98&#x0025;</td>
<td align="left">98&#x0025;</td>
<td align="left">98&#x0025;</td>
<td align="left">98&#x0025;</td>
</tr>
<tr>
<td align="left">KNN</td>
<td align="left">94&#x0025;</td>
<td align="left">94&#x0025;</td>
<td align="left">94&#x0025;</td>
<td align="left">94&#x0025;</td>
</tr>
<tr>
<td align="left">DNN</td>
<td align="left">95&#x0025;</td>
<td align="left">95&#x0025;</td>
<td align="left">95&#x0025;</td>
<td align="left">95&#x0025;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The RF and DT models scored almost the same results, up to 99&#x0025; and 98&#x0025;, respectively, since they execute the classification based on building decision trees to make quick data-driven decisions, as a DT is easier to interpret and faster to train on large datasets, particularly linear ones. The RF model needs rigorous training and consumes more training time, up to eight times more than the DT training time recorded in our experiment. Due to the additive learning technique, which uses gradient descent to identify challenges in the learners&#x2019; predictions and improve weak learners&#x2019; predictions, GBC attained a superior performance of up to 99&#x0025; as well as RF compared to other algorithms. Also, GBC overcomes the overfitting problem and attains full computational resource utilization.</p>
<p>Our approach is robust even when the dataset includes missing data because we are just collecting source IP, destination IP, and total packets, which are usually available in each flow of traffic, and then from these features, we extract 7 graph features to feed the ML model for bot detection.</p>
</sec>
</sec>
</sec>
<sec id="s5"><label>5</label><title>Evaluation Approach and Discussion</title>
<p>The evaluation of our proposed classification model is based on two aspects: performance of preprocessing to state-of-the-art and performance of the classification model to state-of-the-art.</p>
<sec id="s5_1"><label>5.1</label><title>Performance Comparison of Preprocessing to State-of-the-Art</title>
<p>To validate the effectiveness of our preprocessing steps, we evaluated the performance of the same ML algorithms and same datasets described earlier with different methods for extracting the features. Here we compare the results of the proposed method with BotSward with those of the traditional methods in [<xref ref-type="bibr" rid="ref-44">44</xref>] and [<xref ref-type="bibr" rid="ref-45">45</xref>]. Our proposed work on BotSward and these studies all used the same CTU-13 datasets and ML classifiers, which were RF and DT, but the methodologies for extracting the features were different. [<xref ref-type="bibr" rid="ref-44">44</xref>] and [<xref ref-type="bibr" rid="ref-45">45</xref>] used conversation-based features, while we used graph-based features. The framework in [<xref ref-type="bibr" rid="ref-44">44</xref>] employed the RF model to extract conversation-based features, and in their study, the RF had the greatest detection rate of all the classification methods, reaching up to 93.6&#x0025;, with only a 0.3&#x0025; false alarm rate, which is 10 times lower than detection based on traffic flow features in [<xref ref-type="bibr" rid="ref-46">46</xref>]. Also, [<xref ref-type="bibr" rid="ref-45">45</xref>] proposed an effective two-stage traffic classification method to detect botnet-based on DT on conversation features. The DT algorithm&#x2019;s success rate was as high as 94.4&#x0025;. From these results, it is clear that BotSward obtained the most robust results, with 99&#x0025; with both RF and DT and only a 0.001&#x0025; false alarm rate. The comparison of these results is summarized in <xref ref-type="table" rid="table-6">Tabs. 6</xref> and <xref ref-type="table" rid="table-7">7</xref>.</p>
<table-wrap id="table-6"><label>Table 6</label><caption><title>Performance of [<xref ref-type="bibr" rid="ref-44">44</xref>] compared to BotSward</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Algorithm</th>
<th align="left">Features</th>
<th align="left">Dataset</th>
<th align="left">Classifier</th>
<th align="left">Recall (R)</th>
<th align="left">F-score (F)</th>
<th align="left">Accuracy (ACC)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
<td align="left">Conversation-based</td>
<td align="left">CTU-13</td>
<td align="left">RF</td>
<td align="left">93.6&#x0025;</td>
<td align="left">93.6&#x0025;</td>
<td align="left">93.6&#x0025;</td>
</tr>
<tr>
<td align="left">BotSward</td>
<td align="left">Graph-based</td>
<td align="left">CTU-13</td>
<td align="left">RF</td>
<td align="left">99&#x0025;</td>
<td align="left">99&#x0025;</td>
<td align="left">99&#x0025;</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-7"><label>Table 7</label><caption><title>Performance of [<xref ref-type="bibr" rid="ref-45">45</xref>] compared to BotSward</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Algorithm</th>
<th align="left">Features</th>
<th align="left">Dataset</th>
<th align="left">Classifier</th>
<th align="left">Recall (R)</th>
<th align="left">F-score (F)</th>
<th align="left">Accuracy (ACC)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">[<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td align="left">Conversation-based</td>
<td align="left">CTU-13</td>
<td align="left">DT</td>
<td align="left">94.4&#x0025;</td>
<td align="left">94.4&#x0025;</td>
<td align="left">94.4&#x0025;</td>
</tr>
<tr>
<td align="left">BotSward</td>
<td align="left">Graph-based</td>
<td align="left">CTU-13</td>
<td align="left">DT</td>
<td align="left">98&#x0025;</td>
<td align="left">98&#x0025;</td>
<td align="left">99&#x0025;</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_2"><label>5.2</label><title>Performance Comparison of Classification Algorithm to State-of-the-Art</title>
<p>To validate the effectiveness of our classification model in BotSward, we evaluate the performance of the different ML classifier algorithms using the same feature extraction methods and dataset. We used the CTU-13 dataset to compare our model to the state-of-the-art in graph-based botnet detection methods, namely BClus, CAMNEP, BotHunter, BotGM, and Botchase, which are described in Section 2. <xref ref-type="table" rid="table-8">Tab. 8</xref> reports the results for each solution and all scenarios from the test set, which includes Scenarios 1, 2, 6, 8, and 9, as recommended in [<xref ref-type="bibr" rid="ref-26">26</xref>]. Our results are very competitive as we reached an accuracy of between 97&#x0025; and 100&#x0025;, while other studies with BClus, CAMNEP, and BotHunter only provide an accuracy of between 30&#x0025; and 50&#x0025;. Only Botchase achieved up to 99&#x0025; accuracy for Scenario 9, but it achieved moderate accuracy in the remaining scenarios. To compare BotSward, Botchase, and BotGM for graph-based botnet detection, we used the performance metrics for Scenarios 1, 2, 8, 6, and 9 to measure how well they worked.</p>
<table-wrap id="table-8">
<label>Table 8</label><caption><title>Accuracy of different algorithms evaluated in [<xref ref-type="bibr" rid="ref-26">26</xref>] and compared to BotSward</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Classifier</th>
<th align="left">1</th>
<th align="left">2</th>
<th align="left">6</th>
<th align="left">8</th>
<th align="left">9</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">BClus</td>
<td align="left">50&#x0025;</td>
<td align="left">50&#x0025;</td>
<td align="left">40&#x0025;</td>
<td align="left">30&#x0025;</td>
<td align="left">40&#x0025;</td>
</tr>
<tr>
<td align="left">CAMNEP</td>
<td align="left">50&#x0025;</td>
<td align="left">40&#x0025;</td>
<td align="left">40&#x0025;</td>
<td align="left">50&#x0025;</td>
<td align="left">50&#x0025;</td>
</tr>
<tr>
<td align="left">BotHunter</td>
<td align="left">40&#x0025;</td>
<td align="left">30&#x0025;</td>
<td align="left">38&#x0025;</td>
<td align="left">42&#x0025;</td>
<td align="left">40&#x0025;</td>
</tr>
<tr>
<td align="left">BotGM</td>
<td align="left">91&#x0025;</td>
<td align="left">78&#x0025;</td>
<td align="left">95&#x0025;</td>
<td align="left">89&#x0025;</td>
<td align="left">83&#x0025;</td>
</tr>
<tr>
<td align="left">BotChase</td>
<td align="left">98&#x0025;</td>
<td align="left">97&#x0025;</td>
<td align="left">94&#x0025;</td>
<td align="left">84&#x0025;</td>
<td align="left">99&#x0025;</td>
</tr>
<tr>
<td align="left">BotSward</td>
<td align="left">99&#x0025;</td>
<td align="left">97&#x0025;</td>
<td align="left">99&#x0025;</td>
<td align="left">100&#x0025;</td>
<td align="left">99&#x0025;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>BotChase <italic>vs.</italic> BotSward, according to the BotChase experiment, they applied ID, OD, ODW, BC, LCC, and AC graph-based algorithms, and therefore they employed a two-layer detection approach based on supervised and unsupervised learning. However, the BC feature used in Botchase has a high computational overhead, with a time complexity of (O&#x007C;V&#x007C;.&#x007C;E&#x007C;&#x2009;&#x002B;&#x2009;&#x007C;V&#x007C;2. log &#x007C;V&#x007C;). In contrast, BotSward focuses on the measures of centrality used in Botchase and replaces BC with PR, EC, and CC, followed by a one-layer detection technique based on GBC ensemble ML algorithms. In Scenario 8, Botchase achieved 84&#x0025;, whereas BotSward achieved outstanding results with 100&#x0025;. Overall, the accuracy in BotSward overcame that obtained in Botchase by the ratio of 2&#x0025; to 16&#x0025; and reduced the time complexity by the ratio of 69.55&#x0025;, as shown in <xref ref-type="table" rid="table-8">Tab. 8</xref>.</p>
<p>Reference [<xref ref-type="bibr" rid="ref-23">23</xref>] <italic>vs.</italic> BotSward, [<xref ref-type="bibr" rid="ref-23">23</xref>] proposed an anomaly-based botnet detection method by hybrid analysis of flow-based features and graph-based features of network traffic. They extracted 7 graph features from network traffic, which are ID, OD, IDW, ODW, LCC, BC, and PR. Their results achieved a good performance with a 96.62&#x0025; F-score. In contrast, BotSward focuses on the bot classification by measuring the centrality used in [<xref ref-type="bibr" rid="ref-23">23</xref>] and replacing BC with CC [<xref ref-type="bibr" rid="ref-23">23</xref>] achieved a moderate performance, with a 96.62&#x0025; F-score compared to a high performance of more than 99&#x0025; in BotSward.</p>
<p>BotGM <italic>vs.</italic> BotSward, while both of these systems use graph methodologies in attempting to identify behavior that is malicious or outlying, the BotGM model only achieved accuracy of up to 95&#x0025; in Scenario 6 of the test datasets, in contrast to 99&#x0025; for BotSward. In addition, the lower accuracy of the BotGM model was shown in Scenario 2 with 78&#x0025;, and in contrast, BotSward achieved 97&#x0025;, as shown in <xref ref-type="table" rid="table-8">Tab. 8</xref>. However, BotGM generates multiple graphs for every single host, which entails a high overhead. According to our results, it can be concluded that BotSward outperformed BotGm in terms of lower overhead and higher accuracy, as shown in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<fig id="fig-5"><label>Figure 5</label><caption><title>Accuracy of different algorithms evaluated in [<xref ref-type="bibr" rid="ref-26">26</xref>] and compared to BotSward</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_31641-fig-5.png"/></fig>
<p>In summary, using graph-based features in the ML approach provides robustness in defending against unknown attacks and complicated communication patterns, and provides enhanced resource utilization, a low error rate, a high detection rate, and detection and mitigation of botnet attacks. Put simply, BotSward is suitable to work with a variety of datasets because it only requires the extraction of source IP, destination IP, and total packets from the dataset, and then we can calculate measures of centrality for graph-based botnets, which include DC, CC, EC, and AC, and then use the ML classification to detect unknown botnets by using these features.</p>
</sec>
</sec>
<sec id="s6"><label>6</label><title>Conclusion</title>
<p>This paper proposes BotSward, a botnet attack defense approach based on centrality measures for graph-based and ML algorithms. BotSward utilizes centrality measures for graph-based, which employs a set of important graph algorithm features that are widely utilized. The graph-based feature extraction in all the previous studies suffered from high computational overhead because of the time complexity of the BC, and so we suggest excluding BC and LCC, and replacing them with CC, as we found the same high accuracy with a time reduction of up to 69.5&#x0025;. Our proposed solution then employs a detection model based on the ML algorithm to distinguish different botnet attacks from normal traffic. We compared various types of ML such as RF, GBC, LR, SGD, SVM, DT, KNN, and DNN and evaluated them using the CTU-13 dataset, and GBC showed the top accuracy, up to 99&#x0025; in all dataset scenarios and 100&#x0025; in some scenarios compared with other ML algorithms with a false positive rate as low as 0.0001&#x0025;. BotSward is also able to detect bots that rely on different protocols and proves robust against unknown attacks. In future work, we intend to deploy our BotSward approach in the Software-Defined Network (SDN) environment, which is an emerging network architecture that provides centralized administration through a single controller that decouples the control and data planes, addressing the limitations of traditional networks. However, this evolution has made the controller a critical target for malicious users to launch attacks such as DDoS attacks, where botnets provide platforms for DDOS. Several ways to stop botnet attacks in SDNs have been discussed, but the problems still exist.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="other"><p><bold>Funding Statement:</bold> The authors received no funding for this study.</p></fn>
<fn fn-type="conflict"><p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p></fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Shinan</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Alsubhi</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Alzahrani</surname></string-name> and <string-name><given-names>M. U.</given-names> <surname>Ashraf</surname></string-name></person-group>, &#x201C;<article-title>Machine learning-based botnet detection in software-defined network: A systematic review</article-title>,&#x201D; <source>Symmetry</source>, vol. <volume>13</volume>, no. <issue>5</issue>, pp. <fpage>866</fpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Heron</surname></string-name></person-group>, &#x201C;<article-title>Working the botnet: How dynamic DNS is revitalising the zombie army</article-title>,&#x201D; <source>Network Security</source>, vol. 15, no. <issue>1</issue>, pp. <fpage>9</fpage>&#x2013;<lpage>11</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. A.</given-names> <surname>Kamal</surname></string-name>, <string-name><given-names>L. M.</given-names> <surname>Ibrahim</surname></string-name> and <string-name><given-names>A. A.</given-names> <surname>Al-alusi</surname></string-name></person-group>,&#x201C;<article-title>Dolphin and elephant herding optimization swarm intelligence algorithms used to detect neris botnet</article-title>,&#x201D; <source>Journal of Engineering Science and Technology</source>, vol. <volume>15</volume>, no. <issue>5</issue>, pp. <fpage>2906</fpage>&#x2013;<lpage>2923</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="thesis"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Izzillo</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Pellegrini</surname></string-name></person-group>, &#x201C;<article-title>Graph and flow-based distributed detection and mitigation of botnet attacks</article-title>,&#x201D; M.S.thesis, <publisher-name>Dept. Engineering in Computer Science,University of Rome</publisher-name>, <publisher-loc>Roma, Italy</publisher-loc>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Lange</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Kettani</surname></string-name></person-group>, &#x201C;<article-title>On security threats of botnets to cyber systems</article-title>,&#x201D; in <conf-name>Proc. IEEE 6th Int. Conf. on Signal Processing and Integrated Networks (SPIN)</conf-name>, <conf-loc>Noida, India</conf-loc>, pp. <fpage>176</fpage>&#x2013;<lpage>183</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Blaise</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Bouet</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Conan</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Secci</surname></string-name></person-group>, &#x201C;<article-title>Botfp: Fingerprints clustering for bot detection</article-title>,&#x201D; in <conf-name>IEEE/IFIP Network Operations and Management Symp. (NOMS)</conf-name>, Budapest, Hungary, pp. <fpage>1</fpage>&#x2013;<lpage>7</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W. N. H.</given-names> <surname>Ibrahim</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Anuar</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Selamat</surname></string-name>, <string-name><given-names>O.</given-names> <surname>Krejcar</surname></string-name>, <string-name><given-names>R. G.</given-names> <surname>Crespo</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Multilayer framework for botnet detection using machine learning algorithms</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>9</volume>, pp. <fpage>48753</fpage>&#x2013;<lpage>48768</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Ghafir</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Svoboda</surname></string-name> and <string-name><given-names>V.</given-names> <surname>Prenosil</surname></string-name></person-group>, &#x201C;<article-title>A survey on botnet command and control traffic detection</article-title>,&#x201D; <source>Int. J. Adv. Comput. Netw. Secur.</source>, vol. <volume>5</volume>, no. <issue>2</issue>, pp. <fpage>7580</fpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Abualkibash</surname></string-name></person-group>, &#x201C;<article-title>Machine learning in network security using knime analytics</article-title>,&#x201D; <source>International Journal of Network Security &#x0026; Its Applications (IJNSA)</source>, vol. <volume>11</volume>, no. <issue>5</issue>, pp. <fpage>564</fpage>&#x2013;<lpage>578</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H. R.</given-names> <surname>Zeidanloo</surname></string-name>, <string-name><given-names>M. J. Z.</given-names> <surname>Shooshtari</surname></string-name>, <string-name><given-names>P. V.</given-names> <surname>Amoli</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Safari</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Zamani</surname></string-name></person-group>, &#x201C;<article-title>A taxonomy of botnet detection techniques</article-title>,&#x201D; in <conf-name>Proc. IEEE 3rd Int. Conf. on Computer Science and Information Technology (ICCSIT)</conf-name>, <conf-loc>Chengdu, China</conf-loc>, pp. <fpage>158</fpage>&#x2013;<lpage>162</lpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Vania</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Meniya</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Jethva</surname></string-name></person-group>, &#x201C;<article-title>A review on botnet and detection technique</article-title>,&#x201D; <source>International Journal of Computer Trends and Technology</source>, vol. <volume>4</volume>, no. <issue>1</issue>, pp. <fpage>23</fpage>&#x2013;<lpage>29</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Limarunothai</surname></string-name> and <string-name><given-names>M. A.</given-names> <surname>Munlin</surname></string-name></person-group>, &#x201C;<article-title>Trends and challenges of botnet architectures and detection techniques</article-title>,&#x201D; <source>Journal of Information Science and Technology</source>, vol. <volume>5</volume>, no. <issue>1</issue>, pp. <fpage>51</fpage>&#x2013;<lpage>57</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>I. C.</given-names> <surname>Paschalidis</surname></string-name></person-group>, &#x201C;<article-title>Botnet detection based on anomaly and community detection</article-title>,&#x201D; <source>IEEE Transactions on Control of Network Systems</source>, vol. <volume>4</volume>, no. <issue>2</issue>, pp. <fpage>392</fpage>&#x2013;<lpage>404</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Karim</surname></string-name>, <string-name><given-names>R. B.</given-names> <surname>Salleh</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Shiraz</surname></string-name>, <string-name><given-names>S. A. A.</given-names> <surname>Shah</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Awan</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Botnet detection techniques: Review, future trends, and issues</article-title>,&#x201D; <source>Journal of Zhejiang University Science C</source>, vol. <volume>15</volume>, no. <issue>11</issue>, pp. <fpage>943</fpage>&#x2013;<lpage>983</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R. C.</given-names> <surname>Fernandez</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Deng</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Mansour</surname></string-name>, <string-name><given-names>A. A.</given-names> <surname>Qahtan</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Tao</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>A demo of the data civilizer system</article-title>,&#x201D; in <conf-name>Proc. ACM Int. Conf. on Management of Data (MOD)</conf-name>, <conf-loc>Chicago, USA</conf-loc>, pp. <fpage>1639</fpage>&#x2013;<lpage>1642</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Abou Daya</surname></string-name>, <string-name><given-names>M. A.</given-names> <surname>Salahuddin</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Limam</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Boutaba</surname></string-name></person-group>, &#x201C;<article-title>Botchase: Graph-based bot detection using machine learning</article-title>,&#x201D; <source>IEEE Transactions on Network and Service Management</source>, vol. <volume>17</volume>, no. <issue>1</issue>, pp. <fpage>15</fpage>&#x2013;<lpage>29</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Venkatesh</surname></string-name>, <string-name><given-names>S. H.</given-names> <surname>Choudhury</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Nagaraja</surname></string-name> and <string-name><given-names>N.</given-names> <surname>Balakrishnan</surname></string-name></person-group>, &#x201C;<article-title>Botspot: Fast graph based identification of structured p2p bots</article-title>,&#x201D; <source>Journal of Computer Virology and Hacking Techniques</source>, vol. <volume>11</volume>, no. <issue>4</issue>, pp. <fpage>247</fpage>&#x2013;<lpage>261</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Biswas</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Roy</surname></string-name></person-group>, &#x201C;<article-title>Botnet traffic identification using neural networks</article-title>,&#x201D; <source>Multimedia Tools and Applications</source>, vol. <volume>80</volume>, no. <issue>16</issue>, pp. <fpage>24147</fpage>&#x2013;<lpage>24171</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. S.</given-names> <surname>Gadelrab</surname></string-name>, <string-name><given-names>M.</given-names> <surname>ElSheikh</surname></string-name>, <string-name><given-names>M. A.</given-names> <surname>Ghoneim</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Rashwan</surname></string-name></person-group>, &#x201C;<article-title>Botcap: Machine learning approach for botnet detection based on statistical features</article-title>,&#x201D;<source>International Journal of Communication Networks and Information Security (IJCNIS)</source>, vol. <volume>10</volume>, no. <issue>3</issue>, pp. <fpage>563</fpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Miller</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Busby-Earle</surname></string-name></person-group>, &#x201C;<article-title>The role of machine learning in botnet detection</article-title>,&#x201D; in <conf-name>Proc. IEEE 11th Int. Conf. for Internet Technology and Secured Transactions (ICITST)</conf-name>, <conf-loc>Barcelona, Spain</conf-loc>, pp. <fpage>359</fpage>&#x2013;<lpage>364</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>E. B.</given-names> <surname>Beigi</surname></string-name>, <string-name><given-names>H. H.</given-names> <surname>Jazi</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Stakhanova</surname></string-name> and <string-name><given-names>A. A.</given-names> <surname>Ghorbani</surname></string-name></person-group>, &#x201C;<article-title>Towards effective feature selection in machine learning-based botnet detection approaches</article-title>,&#x201D; in <conf-name>Proc. IEEE Conf. on Communications and Network Security (CNS)</conf-name>, <conf-loc>San Francisco, CA, USA</conf-loc>, pp. <fpage>247</fpage>&#x2013;<lpage>255</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Gahelot</surname></string-name> and <string-name><given-names>N.</given-names> <surname>Dayal</surname></string-name></person-group>, &#x201C;<article-title>Flow based botnet traffic detection using machine learning</article-title>,&#x201D; in <conf-name>Proc. ICETIT</conf-name>, <conf-loc>Cham, Switzerland</conf-loc>, <publisher-name>Springer</publisher-name>, pp. <fpage>418</fpage>&#x2013;<lpage>426</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Shang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Botnet detection with hybrid analysis on flow based and graph based features of network traffic</article-title>,&#x201D; in <conf-name>Proc. ICCCS</conf-name>, <conf-loc>Switzerland</conf-loc>, <publisher-name>Springer</publisher-name>, pp. <fpage>612</fpage>&#x2013;<lpage>621</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Lagraa</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Fran&#x00E7;ois</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Lahmadi</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Miner</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Hammerschmidt</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Botgm: Unsupervised graph mining to detect botnets in traffic flows</article-title>,&#x201D; in <conf-name>Proc. 1st Cyber Security in Networking Conf. (CSNet)</conf-name>, <conf-loc>Rio de Janeiro, Brazil</conf-loc>, <publisher-name>IEEE</publisher-name>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Gu</surname></string-name>, <string-name><given-names>P. A.</given-names> <surname>Porras</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Yegneswaran</surname></string-name>, <string-name><given-names>M. W.</given-names> <surname>Fong</surname></string-name> and <string-name><given-names>W.</given-names> <surname>Lee</surname></string-name></person-group>, &#x201C;<article-title>Bothunter: Detecting malware infection through ids-driven dialog correlation</article-title>,&#x201D; in <conf-name>USENIX Security Symp.</conf-name>, Vancouver, BC, Canada, vol. <volume>7</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>16</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Garcia</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Grill</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Stiborek</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Zunino</surname></string-name></person-group>, &#x201C;<article-title>An empirical comparison of botnet detection methods</article-title>,&#x201D; <source>Computers &#x0026; Security</source>, vol. <volume>45</volume>, pp. <fpage>100</fpage>&#x2013;<lpage>123</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="thesis"><person-group person-group-type="author"><string-name><given-names>A. R.</given-names> <surname>Vishwakarma</surname></string-name></person-group>, &#x201C;<article-title>Network traffic based botnet detection using machine learning</article-title>,&#x201D; M.S. thesis, <publisher-name>Dept. Computer Science, San Jose State University</publisher-name>, <publisher-loc>Washington, United States</publisher-loc>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Gonzalez-Cuautle</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Hernandez-Suarez</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Sanchez-Perez</surname></string-name>, <string-name><given-names>L. K.</given-names> <surname>Toscano-Medina</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Portillo-Portillo</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Synthetic minority oversampling technique for optimizing classification tasks in botnet and intrusiondetection-system datasets</article-title>,&#x201D; <source>Applied Sciences</source>, vol. <volume>10</volume>, no. <issue>3</issue>, pp. <fpage>794</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>A. K.</given-names> <surname>Sangaiah</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Small object detection via precise region-based fully convolutional networks</article-title>,&#x201D; <source>Computers, Materials and Continua</source>, vol. <volume>69</volume>, no. <issue>2</issue>, pp. <fpage>1503</fpage>&#x2013;<lpage>1517</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Pokhrel</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Abbas</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Aryal</surname></string-name></person-group>, &#x201C;<article-title>Iot security: Botnet detection in iot using machine learning</article-title>,&#x201D; arXiv preprint arXiv:2104.02231, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. M.</given-names> <surname>Rahman</surname></string-name> and <string-name><given-names>D. N.</given-names> <surname>Davis</surname></string-name></person-group>, &#x201C;<article-title>Addressing the class imbalance problem in medical datasets</article-title>,&#x201D; <source>International Journal of Machine Learning and Computing</source>, vol. <volume>3</volume>, no. <issue>2</issue>, pp. <fpage>224</fpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z. M.</given-names> <surname>Algelal</surname></string-name>, <string-name><given-names>E. A. G.</given-names> <surname>Aldhaher</surname></string-name>, <string-name><given-names>D. N.</given-names> <surname>Abdul-Wadood</surname></string-name> and <string-name><given-names>R. H. A.</given-names> <surname>AlSagheer</surname></string-name></person-group>, &#x201C;<article-title>Botnet detection using ensemble classifiers of network flow</article-title>,&#x201D; <source>International Journal of Electrical and Computer Engineering</source>, vol. <volume>10</volume>, no. <issue>3</issue>, pp. <fpage>2543</fpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Hung</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<article-title>A botnet detection system based on machine-learning using flow-based features</article-title>,&#x201D; in <conf-name>Proc. of the SECURWARE 2018: The Twelfth Int. Conf. on Emerging Security Information</conf-name>, <conf-loc>Italy</conf-loc>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Mishra</surname></string-name> and <string-name><given-names>S. K.</given-names> <surname>Jha</surname></string-name></person-group>, &#x201C;<chapter-title>Survey on botnet detection techniques</chapter-title>,&#x201D; in <source>Internet of Things and Its Applications, Lecture Notes in Electrical Engineering</source>, vol. <volume>825</volume>. <publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>, pp. <fpage>441</fpage>&#x2013;<lpage>449</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Dagon</surname></string-name>, <string-name><given-names>C. C.</given-names> <surname>Zou</surname></string-name>, and <string-name><given-names>W.</given-names> <surname>Lee</surname></string-name></person-group>, &#x201C;<article-title>Modeling botnet propagation using time zones</article-title>.,&#x201D; <source>NDSS</source>, vol. <volume>6</volume>, pp. <fpage>2</fpage>&#x2013;<lpage>13</lpage>, <year>2006</year>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Sanatinia</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Noubir</surname></string-name></person-group>, &#x201C;<article-title>Onionbots: Subverting privacy infrastructure for cyber attacks</article-title>,&#x201D; in <conf-name>Proc. IEEE 45th Annual IEEE/IFIP Int. Conf. on Dependable Systems and Networks (DSN)</conf-name>, <conf-loc>Janeiro, Brazil</conf-loc>, pp. <fpage>69</fpage>&#x2013;<lpage>80</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zou</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Lei</surname></string-name>, <string-name><given-names>R. S.</given-names> <surname>Sherratt</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Research on recurrent neural network based crack opening prediction of concrete dam</article-title>,&#x201D; <source>Journal of Internet Technology</source>, vol. <volume>21</volume>, no. <issue>4</issue>, pp. <fpage>1161</fpage>&#x2013;<lpage>1169</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>He</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Liao</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Li</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Parameters compressing in deep learning</article-title>,&#x201D; <source>Computers Materials &#x0026; Continua</source>, vol. <volume>62</volume>, no. <issue>1</issue>, pp. <fpage>321</fpage>&#x2013;<lpage>336</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Ryu</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Yang</surname></string-name></person-group>, &#x201C;<article-title>A comparative study of machine learning algorithms and their ensembles for botnet detection</article-title>,&#x201D; <source>Journal of Computer and Communications</source>, vol. <volume>6</volume>, no. <issue>5</issue>, pp. <fpage>119</fpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Nie</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Using CFW-net deep learning models for X-ray images to detect COVID-19 patients</article-title>,&#x201D; <source>International Journal of Computational Intelligence Systems</source>, vol. <volume>14</volume>, no. <issue>1</issue>, pp. <fpage>199</fpage>&#x2013;<lpage>207</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Luo</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Woodland labeling in chenzhou, China, via deep learning approach</article-title>,&#x201D;<source>International Journal of Computational Intelligence Systems</source>, vol. <volume>13</volume>, no. <issue>1</issue>, pp. <fpage>1393</fpage>&#x2013;<lpage>1403</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Comaneci</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Dobre</surname></string-name></person-group>, &#x201C;<article-title>Securing networks using sdn and machine learning</article-title>,&#x201D; in <conf-name>Proc. IEEE Int. Conf. on Computational Science and Engineering (CSE)</conf-name>, <conf-loc>Bucharest, Romania</conf-loc>, pp. <fpage>194</fpage>&#x2013;<lpage>200</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Yan</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Shao</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Multi-attributed heterogeneous graph convolutional network for bot detection</article-title>,&#x201D; <source>Information Sciences</source>, vol. <volume>537</volume>, pp. <fpage>380</fpage>&#x2013;<lpage>393</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Niu</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhuo</surname></string-name>, and <string-name><given-names>F.</given-names> <surname>Lv</surname></string-name></person-group>,&#x201C;<article-title>An effective conversation-based botnet detection method</article-title>,&#x201D; <source>Mathematical Problems in Engineering</source>, vol. <volume>2017</volume>, pp. <fpage>334</fpage>&#x2013;<lpage>344</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R. U.</given-names> <surname>Khan</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Alazab</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>A hybrid technique to detect botnets, based on P2P traffic similarity</article-title>,&#x201D; in <conf-name>Proc. Cybersecurity and Cyberforensics Conf. (CCC)</conf-name>, <conf-loc>Melbourne, Australia</conf-loc>, pp. <fpage>136</fpage>&#x2013;<lpage>142</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Kirubavathi</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Anitha</surname></string-name></person-group>, &#x201C;<article-title>Botnet detection via mining of traffic flow characteristics</article-title>,&#x201D; <source>Computers and Electrical Engineering</source>, vol. <volume>50</volume>, pp. <fpage>91</fpage>&#x2013;<lpage>101</lpage>, <year>2016</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>