<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Evaluation of Machine Learning Methods to Detect DoS / DDoS  Attacks on IoT</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Leonid Galchynsky</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mykola Graivoronskyi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleh Dmytrenko</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute for Information Recording of National Academy of Sciences of Ukraine</institution>
          ,
          <addr-line>2, Mykoly Shpaka Street, Kyiv, 03113</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Technical University «Igor Sikorsky Kyiv Polytechnic Institute»</institution>
          ,
          <addr-line>37, Prospekt Peremohy, Kyiv, 03056</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>225</fpage>
      <lpage>236</lpage>
      <abstract>
        <p>   This paper was aimed at the increase of validity of the chosen method and the accuracy of detection of DDoS attacks on IoT using machine learning. A study of the effectiveness of detecting DoS/DDoS attacks in IoT networks was conducted. The machine learning model was built, the procedure for selecting the most significant 10 features from the dataset data was analyzed, relevant machine learning algorithms were selected to identify traffic anomalies in the IoT network, criteria for evaluating the effectiveness of machine learning algorithms were defined, and performance estimates according to four criteria (time, recall, precision, and f-measure) were obtained. The ranking of the selected and tested machine learning algorithms was obtained by applying the ELECTRE III procedure. As a result, the Random Forest method received the highest rank. The obtained results provide the background for prospective application of the suggested method in real systems for detection of DoS/DDoS attacks in IoT networks.</p>
      </abstract>
      <kwd-group>
        <kwd>  Internet of Things</kwd>
        <kwd>DoS/DDoS attacks</kwd>
        <kwd>traffic anomalies</kwd>
        <kwd>machine learning</kwd>
        <kwd>ELECTRE method </kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>including sufficient speed, accuracy of assessment and a number of other indicators. In our opinion,
solutions with machine learning (ML) are one of the most promising for this problem. Although ML
information security solutions have gained some popularity, the issue of detecting attacks on IoT
networks is still poorly understood.</p>
    </sec>
    <sec id="sec-2">
      <title>2  Formulation of the problem </title>
      <p>
        Ensuring cybersecurity of IoT networks remains a challenge for professionals because of their
features, mainly due to a combination of different technologies, some of which are outdated. These
technologies have traditional flaws in data privacy and security, and they need to be addressed to the
specifics of IoT. Many researchers are working to address various security issues in the IoT, but the
level of security of IoT devices is still not fully in line with user needs. IoT networks are vulnerable to
known cyberattacks, including denial of service (DoS), distributed denial of service, replay attacks, man
in the middle, routing and eavesdropping attacks [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        The cybersecurity system should provide network protection at all layers of the Open System
Interconnection Reference Model (OSI), providing connection control, analysis of the structure and
content of network packets, traffic monitoring, assessment of system elements functioning states [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        However, the current practice shows that the analysis of security threats and methods of their
detection are limited to one layer of the network interaction while for IoT networks, real attack can be
a combination of threats from multiple network layers. Standard protection techniques developed for
certain layers of IoT and sensor networks may be ineffective against attacks from other layers. It follows
that a reliable assessment requires an analysis of the network as a whole. However, solving this systemic
problem in general is too complex to be completed in one study. Therefore, we will consider only one
of the biggest cyber threats to IoT networks – DoS/DDoS attacks. The main danger of the vast majority
of DoS/DDoS attacks – in their absolute clarity and almost "normal" behavior. Software errors, when
discovered, are immediately fixed and promptly corrected. But in this case the only apparent indication
of the attack is the full consumption of resources, what is almost normal behavior for modern
information systems. Standard methods of statistical analysis do not allow us to detect previously
unknown attacks, and hence machine learning algorithms become the tool to solve this problem which
is actively studied and used [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. As the analysis shows, the most effective means to detect and prevent
DoS/DDoS attacks are those that use intelligent methods (IDMS) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In recent years, many studies
have been conducted on the use of machine learning as a tool to detect DDoS attacks. However, they
are often not sufficiently substantiated, there are few details of the results obtained, and they also rarely
specialize in IoT networks, and evaluate the results only by accuracy, or do not clearly describe the
approach by which the solution was chosen.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3  Analysis of recent research and publications </title>
      <p>
        The objective of detecting DDoS attacks consists in making the most reliable estimate of whether
devices are under attack in a given period of time or not by analyzing data from devices. The speed of
the chosen method is the most desirable, because the slightest delay can lead to significant losses.
Standard methods of statistical analysis do not allow to detect previously unknown attacks, therefore
the appropriate tool to solve this problem is to use machine learning algorithms [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Such a strategy for
detecting DoS/DDoS attacks is implemented through intrusion detection systems (IDS). Network
Intrusion Detection Systems (NIDS) are usually located on the node that connects the internal network
to the Internet, in order to scan incoming and outgoing traffic for known attack patterns [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. As known
malicious patterns are detected, messages will be displayed and examination algorithms may be run,
which reduces response time and potentially increases the accuracy of the investigation. This is why
IDS should include ML in its anomaly detection process with classification, clustering, and other
methods used to detect anomalous traffic.
      </p>
      <p>To create a prototype of a real ML-based intrusion detection system, it was necessary to develop a
software application and to test it on the appropriate data set. For the experiments, we chose the BotIoT
data set, created in the laboratory of the Australian Center for Cyber Security (ACCS) specifically for
researchers to study the possibilities of machine learning in detecting attacks on IoT networks.</p>
      <p>
        Due to the limited power consumption of IoT devices, the key management system may be
compromised. Simplício et al have shown that simplified key authentication systems are insufficiently
reliable in high-scale implementations [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Network protocols such as IPv6 and IPv4 are targets for
attempting to establish remote access. Czyz et al demonstrated that IoT devices could be accessed
remotely via the command line interface using the Telnet protocol [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. IoT networks often process big
data continuously. Loss of data or denial of service can lead to huge amounts of network traffic and loss
of control over the system. According to Angrishi [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], malicious software can remain dormant in IoT
devices, and after launch, it turns the IoT device into a bot and uses it to carry out a DDoS attack.
      </p>
      <p>
        In 2013, the Austrian and German power grids began to fail due to a DDoS attack and filled the
central command center with traffic [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. A DDoS attack can temporarily disable the communication
network and IoT device protocol. The automated smart grid system relies heavily on the IoT monitoring
network. The sensor network collects real-time information data to monitor the condition of equipment
and control systems. As a result, in the event of an attack, the station will lose the ability to monitor the
transmission and distribution network.
      </p>
      <p>The aim of this paper is to increase the validity of the chosen method and the accuracy of detection
of DDoS attacks on IoT using machine learning.</p>
    </sec>
    <sec id="sec-4">
      <title>4  Results and discussion  4.1  The formal statement of the problem </title>
      <p>The formal statement of the problem of detecting DDoS attacks is as follows:
let a certain set of input data X and a set of responses Y exist, where xi ⊂ X is a set of characteristics
of traffic data, yi ⊂ Y determines the affiliation of a particular xi ⊂ X to the attack.</p>
      <p>We assume there is a training sample { x1, … , xn } ⊂ X and a corresponding sample of correct
responses { y1, … , yn } ⊂ Y.</p>
      <p>The problem of detecting an attack in this case may be formulated as the definition of such a rule
(algorithm or model), which will give the value of Y closest to the correct values on the whole set X:
a: Х → Y
The characteristics xi ⊂ X are divided into types:
 Quantitative (values from the set of real numbers);
 Boolean {0, 1};
 Nominal (values from a finite subset of N);
 Ordinal, representing nominal characteristics, with a linear order.</p>
      <p>A Boolean characteristic can be a column of data that indicates the presence or absence of an attack.
The protocol to which this package belongs can be a nominal characteristic. The number of flags in the
package headers can be an ordinal characteristic. An example of quantitative characteristics is the
package length.</p>
      <p>The training sample is created from a specific data set in the form of a feature matrix, in which the
rows represent examples of traffic, as well as the corresponding known states.</p>
      <p>Y can also consist of features of different types: Boolean, nominal, intersecting nominal, and
example data can belong to several types of attacks simultaneously. The peculiarity of the last two
options is that first it should be indicated whether the example is an attack, and only then may its type
be determined.
4.2 </p>
    </sec>
    <sec id="sec-5">
      <title>Two approaches to solve the problem </title>
      <p>Two approaches are proposed to solve this problem:
 detection of attacks based on signatures;
 detection of attacks based on anomalies.</p>
      <p>Signature-based identification is commonly used to identify known types of attacks. No description
of typical actions is required, but a database with known attack signatures is required to detect these
types of attacks. This is the main disadvantage of this approach. However, even for known signatures,
there may be another vulnerability: if there are too many signatures, searching the database can take too
long to detect intrusions in a timely manner.</p>
      <p>Anomaly-based intrusion detection methods recognize unusual activity in network traffic. For
example: abuse of system rules (hiding the IP address interval, performing a standard transaction on a
hidden port); unusual traffic patterns (more UDP packets than TCP packets); suspicious examples in
the data section of the program, etc.</p>
      <p>The most difficult problem for anomaly-based detection methods is the detection of atypical system
behavior, as well as the choice of boundaries to prevent false alarms. Statistical methods of analysis do
not allow to detect previously unknown attacks. The usual way to create a model that has the ability to
avoid a lot of false positives is extremely difficult. Therefore, in fact, the only way out is to use machine
learning tools to detect various threats [12].
4.3 </p>
    </sec>
    <sec id="sec-6">
      <title>Construction of a machine learning model </title>
      <p>Constructing a model of machine learning involves the following steps:
1. Preparation and collection of data;
2. Selection of characteristics (parameters);
3. Selection of the set of algorithms;
4. Choice of criteria and factors;
5. Training;
6. Evaluation of the effectiveness of the results.</p>
      <p>The scheme of step-by-step construction of machine learning model is shown in Fig.1.</p>
      <sec id="sec-6-1">
        <title>Figure 1: Stages of construction of the ML model </title>
        <p>The first stage involves filtering and normalizing the data to provide input data for algorithms and
characteristics. Those that are important for the learning process are selected. Data may contain noise,
may not have a clear structure, so they must be pre-processed. Good data preparation allows for
effective analysis, limits errors and inaccuracies that may occur with the data during processing, and
makes all processed data easy to use.</p>
        <p>At the second stage, among all the existing characteristics, those that are relevant to the learning
process are selected. The number of characteristics in the data sets can reach hundreds. And not all of
them are really important (i.e. have a relationship with the target variable), so they actually are
redundant (or noise). Deleting such parameters helps the researcher to better understand the data, as
well as reduces the time to set up the model and the accuracy of its forecasts.</p>
        <p>
          At the third stage, relevant algorithms are selected to solve a specific problem. Analysis of
publications [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], [13] allowed to choose five algorithms that are most suitable for detecting anomalies
in traffic. These include:
 The method of k-nearest neighbors (KNN);
 Naive Bayesian Classifier (NB);
 Random Forest;
 Logistic Regression;
 Decision Tree.
        </p>
        <p>The fourth stage involves the selection and justification of criteria by which one can assess the
quality of the algorithm in terms of the problem, as well as the optimal parameters of algorithms to
solve the problem are selected.</p>
        <p>At the fifth stage, the model is taught using part of the data set as training data. Part of the data is
taken from the whole set and given to the algorithm to detect patterns; afterwards it should be able to
make assumptions based on the data obtained in this stage.</p>
        <p>Finally, at the sixth stage, the model is tested on data other than the training ones in order to find
out the quality of the forecast.
4.4 </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Creation of a working prototype </title>
      <p>To create a prototype of a real ML-based intrusion detection system, it was necessary to develop a
software application and test the appropriate data set. For the experiments, we chose the BotIoT data
set, created in the laboratory of the Australian Center for Cyber Security (ACCS) [14] specifically for
researchers to study the possibilities of machine learning in detecting attacks on IoT networks.</p>
      <p>The reasons for this choice are as follows: it is open, it contains traffic captured from the IoT
network, it has samples of a wide variety of attacks, including a large number of examples of DoS/DDoS
attacks using different protocols, it includes real traffic, and provides the possibility to generate new
characteristics from the raw data set. The published recorded .pcap files are 69.3 GB in size and contain
over 72,000,000 entries. However, so far we have used only 5% of the original data set retrieved using
MySQL queries. This extracted 5% consists of four files with a total size of about 1.07 GB and includes
about 3 million records. This set is known as Bot-IoT. The data set includes real and simulated IoT
network traffic along with different types of attacks [15]. Attacks in it are divided into three types:
probing attacks, DoS and information theft. However, it also needs careful handling. Figure 2 provides
a diagram of the distribution of the number of packages by category. In this diagram, we can see a
significant uneven amount of data generated during DoS/DDoS attacks compared to other attacks and
normal data. On the one hand, this indicates an imbalance in the data in the dataset, but on the other
hand, it reflects the real situation.</p>
      <p>In the data set, the instances of the attack are marked "1", while the usual traffic "0" for learning and
testing machine learning models using binary code. In addition, attributes of the attack category and
subcategory were introduced that could be used to train and test multiclass classification models.</p>
      <p>Then a two-stage procedure of selection of significant characteristics was carried out. This process
was designed primarily to improve the predictability of classifiers. First, the selection from raw network
traffic data was made using the CICFlowMeter tool [16]. CICFlowMeter reads a pcap file and creates
a visual representation of the removed functions, as well as generates csv files for the data set.</p>
      <p>The selection of features in this software product is based on the Pearson correlation matrix, and in
addition to the already specified characteristics, there is a procedure for generating aggregate
characteristics. 14 characteristics were obtained in the first stage of 33 characteristics available in the
data set. And then the procedure of selection of characteristics on value of Ginny's impurity was carried
out to make further selection from these 14 characteristics. In the end, 10 significant characteristics
were finally selected, according to the combined assessment presented in Table 1.</p>
      <p>To ensure correct assessment of the quality of classifiers and to avoid the problem of overfitting, the
data were divided into training and test parts.</p>
      <p>Models learn only from the data of the training data set. Even when using cross-validation to
determine the optimal parameters of the model, cross-validation is performed only from training data.</p>
      <p>There are no elements from the training part in the final data set. Thus, the model we derived by
performing learning from training data does not know the exact answer to any of the elements of the
final sample and will make predictions based on the rules and patterns developed during training.</p>
      <p>In our study, we divided the original data set at 60% for the training sample and 40% for the test
sample. This gives a fairly even distribution, so the large size of the training sample reduces the
likelihood of underfitting when the model lacks data for training. To better divide each class into
training and test samples, data shuffling was used when sampling the random_state parameter to ten.</p>
      <p>After data analysis in BoT-IoT the following columns were coded: saddr, daddr, proto, target.
Missing values were replaced by NaN, we also left only columns with the code:
 DDoS_HTTP – 0,
 DDoS_TCP – 1,
 DDoS_UDP – 2,
 DoS_HTTP – 3,
 DoS_TCP – 4,
 DoS_UDP – 5,
 Normal_Normal – 6.</p>
      <p>Next, we conducted a series of tests to evaluate the effectiveness of anomaly-based detection of
DoS/DDoS attacks. The machine learning algorithms defined above were tested using the described
data set. The description of the selected algorithms is detailed in the literature. The experiments
presented in this work were performed on a Dell Inspiron 15 3000, with Windows 10 operating system
with AMD A6-6310 processor at 1.80 GHz, and 4 GB memory. For the software implementation,
scikitlearn, Mat-plotlib, Pandas, and NumPy machine learning libraries written for the Python 3.9
programming language, which are open source and widely used in modern machine learning research,
were used.</p>
      <p>Before testing ML algorithms, it is important to determine which indicators to evaluate. The most
important performance indicators are as follows: prediction time, precision, recall, accuracy, and
fmeasure as shown in the equations below [14]:
(1)
(2)
(3)
(4)
where TP is the number of true negative results, FN is false negative, TN is true negative, and FP is
false positive. In particular, the f-measure is indicative, because it reduces two other fundamental
metrics, accuracy and completeness, to one number.</p>
      <p>Precision is interpreted as part of the objects called positive by the classifier which are really
positive, and recall shows what part of positive class objects of all positive class objects was found by
the algorithm. The very introduction of precision does not allow us to write all objects into one class,
because in this case we get an increase in the level of False Positive. Thus, recall demonstrates the
ability of the algorithm to detect the positive class in general, and precision – the ability to distinguish
this class from other classes. Accuracy metrics are not very useful in problems with unequal object
classes [14]. Precision and recall do not depend, in contrast to accuracy, on the class ratio and are
therefore applicable in unbalanced samples.</p>
      <p>To implement the learning phase, a program was written using built-in functions in the sklearn
library. To prepare the algorithms and select the optimal parameters, the implementation of algorithms
in the sklearn library was used.</p>
      <p>For training, we selected random samples of data that were more balanced. Figure 3 shows a
screenshot of the data distribution chart by category.</p>
      <p>Next, experimental tests of the ability to determine anomalies for each trained model were performed
100 times on the test set of the data set. To average the metrics, this procedure was repeated 10 times.</p>
      <p>All data collected for testing algorithms on the BoT-IoT data set are presented in the table. As you
can see from the Table 2, all algorithms showed very high results in detecting instances of DoS attacks
via HTTP protocol. Deviations in the results are insignificant, so to detect this attack in real applications
using the fastest option could be recommended. However, such conclusion could be made at first glance
only. In fact, the purpose of this study is to obtain a sounder solution for the different types of DoS and
DDoS attacks. For example, based on the data obtained in the case of logistic regression, naive Bayesian
classifier and decision tree give a significantly lower percentage of detection of normal traffic, which
means an increased risk of false positives.</p>
      <p>By analyzing the time spent on obtaining the forecast, one can find a clear "outsider" – the KNN
method, which is ten times slower than the random forest, which in turn is slower than the rest. For
Logistic Regression, the decision tree, and the naive Bayesian classifier, the time spent on 100
predictions is comparable and is within 1.5 seconds. In this case, without an integrated assessment of
the criteria, it is difficult to determine which of the algorithms best coped with the task, and which ratio
of time and accuracy can be considered optimal.</p>
      <p>4.5 </p>
    </sec>
    <sec id="sec-8">
      <title>Choice of the best algorithm </title>
      <p>In order to obtain sound estimates of the quality of machine learning algorithms in the detection of
DoS/DDoS attacks, it is necessary to set the task of objectively choosing the best option from several,
that is to set and solve the problem of decision making. The results of evaluating the selected algorithms
using the selected dataset on these five metrics did not show a clear advantage of any of them. For a
reasonable choice of the best, it is necessary to apply a method that takes into account the multi-criteria
nature of such a choice. In a situation of multi-criteria choice, it is not clear which solution is better, so
you need to find a compromise solution that takes into account the importance of each criterion. This
leads to the concept of effective (Pareto optimal) solution. The property of efficiency (in extreme cases,
poor efficiency) must have any solution that claims to be called the best.</p>
      <p>Since there is no general solution to the Pareto optimal choice problem, the decision support
methodology offers many methods, from the set of which one should choose the one that would
correspond to the task. Among the many methods of multi-criteria selection, two classes can be
distinguished – methods that focus on the utility function, and methods that focus on pairwise
comparison of criteria with the subsequent matching procedure. In our case, it is advisable to rely on
the latter approach, because the generally accepted utility function for this task does not yet exist.</p>
      <p>One of the leaders of this approach is the method ELimination Et Choix Traduisant la REalite
(ELECTRE) [17]. This method is characterized by four aspects of advantage modeling: modeling of
fluctuations (partial and comprehensive), modeling of comprehensive incomparabilities, the concept of
consistency and relative importance of criteria, and the concept of inconsistency and veto thresholds. A
number of factors influenced the specific choice of the ELECTRE III method for the algorithm rating
problem. Second, ELECTRE was designed for the fuzzy (inaccurate and uncertain) nature of
decisionmaking, using thresholds of indifference and preference.
Table 2 </p>
      <sec id="sec-8-1">
        <title>The results of the evaluation of algorithms </title>
      </sec>
      <sec id="sec-8-2">
        <title>Criterion </title>
      </sec>
      <sec id="sec-8-3">
        <title>Time for 100 predictions </title>
      </sec>
      <sec id="sec-8-4">
        <title>Random  </title>
      </sec>
      <sec id="sec-8-5">
        <title>Forest </title>
        <p>2,87 </p>
        <p>Another feature of ELECTRE, which distinguishes it from many criteria of solution methods, is that
it is fundamentally non-compensatory. This means, in particular, that a very poor score on a criterion
cannot be offset by a good score on other criteria. The original feature is that the ELECTRE models
allow for incomparability. Incomparability, not to be confused with indifference, arises between any of
0,817 </p>
        <p>NB 
1,7 
0,5 
0,96 
0,34 
0,77 
0,87 
0,65 
0,53 
0,85 
0,92 
0,84 
0,99 
0,999 
0,999 
0,999 
0,364 
0,27 
0,56 
0,618 
0,46 
0,95 
the alternatives a and b, when there is no clear evidence in favor of which of the alternatives to prefer.
This is the situation we are faced with, choosing the best: speed with more errors, or vice versa.</p>
        <p>The purpose of applying ELECTRE methods is to narrow the Pareto set of alternative solutions. This
is done as follows: for each of the criteria (it is assumed that they are numerical) the "weight" is
determined by the results of the survey – a number that characterizes the importance of the criterion. In
all modifications of the ELECTRE method, an attempt is made to obtain qualitative information on the
relative importance of the criteria (statements such as "criteria 3 and 4 are of equal importance and
considered together are more important than criterion 1") and convert it into quantitative, numerical. In
all modifications of the ELECTRE method, in the first stage the weights of the criteria are determined
with the help of a decision-maker – positive real numbers, which are the greater, the more the
corresponding criterion is important for the decision-maker. This approach, of course, has a drawback
– the estimated determination of weights. However, it is impossible to completely avoid subjective
assessments in the decision-making process, it is only necessary to approach the determination of
weight with great care. In addition, these assessments should be provided by experts, not in general, but
in individual aspects.</p>
        <p>The ELECTRE III method involves the following steps – calculating the concordance matrix;
calculating the divergence matrix; calculation of the reliability matrix; calculation of the matrix of
preferences. In the last matrix is the resulting rating for each of the alternatives.</p>
        <p>Evaluations from five experts provided information on the relative importance of the criteria. All
experts determined that the detection of all subcategories of DoS/DDoS attacks is equally important,
and the correct work in detecting normal traffic is more important than individual categories of attacks.
Averaging the estimates, we obtain the following weights: time – 0.8, detection of normal traffic – 2,
detection of denial-of-service attacks – 6 (1 for each criterion). In addition, a survey was conducted to
determine three thresholds, preference (p), indifference (q), and veto (v).</p>
        <p>To calculate the results, the XLSSTAT package was used, which provides a user interface for
applying the ELECTRE II and ELECTRE III methods. To evaluate the performance of algorithms, the
f-measure, as an integral indicator of recall and precision, as well as the time for which the algorithm
will perform estimates, were chosen as criteria. The f-measure value was calculated, and the time was
converted into a dimensionless value in the interval (0; 1], where 1 is the best value available (the fastest
result among the five selected algorithms).</p>
        <p>As a result of calculations, omitting intermediate values, we obtain a matrix of preferences in the
terminology of ELECTRE III (Table 3).</p>
        <p>The letter P (preference) in the corresponding cell means the superiority of the algorithm over
another algorithm, and NP (non-preference) – vice versa. This means that the ELECTRE III method
allowed to integrally evaluate the effectiveness of the tested machine learning algorithms for detecting
DoS/DDoS attacks. Random Forest is a better choice than 4 other algorithms, Decision Tree is better
than other alternatives, Logistic Regression is better than two algorithms, KNN is better than one
algorithm. NB was last in the rankings.</p>
        <p>A study of the effectiveness of detecting DoS/DDoS attacks in IoT networks was conducted. To
accomplish this task, a machine learning model was built, in which a set of data was selected
(BoTIoT dataset), the procedure for selecting the most significant 10 features from the dataset data was
analyzed, relevant machine learning algorithms were analyzed and selected (Random Forest, Decision
Tree, Logistic Regression, KNN, NB) to identify traffic anomalies in the IoT network, criteria for
evaluating the effectiveness of machine learning algorithms were defined, training on a selected set of
datasets was conducted, efficiency on a test sample of the dataset was evaluated, and performance
estimates on four criteria (time, recall, precision, and f-measure) were obtained. The ranking of the
selected and tested machine learning algorithms was obtained by applying the ELECTRE III
procedure, designed for multi-criteria selection, which focuses on a pairwise comparison of criteria
with the subsequent matching procedure. As a result, the Random Forest method received the highest
rank. The obtained results give perspective for application in real systems for detection of DoS/DDoS
attacks in IoT networks. This study cannot be considered complete and calls for further work.</p>
        <p>The prospect for further research is to refine the presented model by using a larger set of data
for training and testing, expanding the range of potentially suitable machine learning algorithms,
obtaining estimates of the effectiveness of algorithms by alternative methods.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>References </title>
      <p>[12] D. Parwani, A. Dutta, P. Kumar Shukla, et al, Various techniques of DDoS attacks detection and
prevention at cloud: a survey, Orient. J. Comp. Sci. and Technol. 8 (2015), no. 2, URL:
http://www.computerscijournal.org/?p=1983
[13] Jadel Alsemiri &amp; Khalid Alsubhi, Internet of Things Cyber Attacks Detection using ML,</p>
      <p>International Journal of Advanced Comp. Science and Applications (2019).
[14] Nour Moustafa, Designing an online and reliable statistical anomaly detection framework for
dealing with large high-speed network traffic, PhD thesis, University of New South Wales,
Canberra, Australia.
[15] K. Nickolaos, N. Moustafa, E. Sitnikova, and B. Turnbull, "Towards the development of realistic
botnet dataset in the IoT for network forensic analytics: Bot-iot dataset", Future Generation Comp.</p>
      <p>Systems 100 (2019).
[16] Mehryar Mohri, Afshin Rostamizadeh, &amp; Ameet Talwalkar, Foundations of Machine Learning,
2nd. ed., the MIT Press, 2018.
[17] Jose Figueira, Salvatore Greco, Bernard Roy, Roman Slowinski, Electre Methods: Main Features
and Recent Developments, 2010. hal-00876980</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Moustafa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Turnbull</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.-K.R. Choo</surname>
          </string-name>
          ,
          <article-title>Towards automation of vulnerability and exploitation identification in iiot networks</article-title>
          ,
          <source>in: 2018 IEEE International Conference on Industrial Internet. DOI:10</source>
          .1109/ICII.
          <year>2018</year>
          .00023 Corpus ID:
          <fpage>53948805</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Moustafa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Turnbull</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.-K.R. Choo</surname>
          </string-name>
          ,
          <article-title>An ensemble intrusion detection technique based on proposed statistical flow features for protecting network traffic of IoT</article-title>
          , IEEE Internet Things J. (
          <year>2018</year>
          ). DOI:
          <volume>10</volume>
          .1109/ICII.
          <year>2018</year>
          .00023
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kasinathan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pastrone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.A.</given-names>
            <surname>Spirito</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Vinkovits</surname>
          </string-name>
          ,
          <article-title>Denial of service detection in 6LoWPAN based internet of things</article-title>
          ,
          <source>in: IEEE 9th International Conference on Wireless and Mobile Computing, Networking and Communications (WiMob)</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>600</fpage>
          -
          <lpage>607</lpage>
          . DOI:
          <volume>10</volume>
          .1109/WiMOB.
          <year>2013</year>
          .6673419
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Salnyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Storchak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mykytyuk</surname>
          </string-name>
          ,
          <source>Information Technology and Security</source>
          <volume>7</volume>
          (
          <year>2019</year>
          ),
          <source>Iss. 1</source>
          , pp.
          <fpage>25</fpage>
          -
          <lpage>34</lpage>
          . URL: http://nbuv.gov.ua/UJRN/inftech_2019_7_
          <issue>1</issue>
          _
          <fpage>5</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tayyab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Belaton</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Anbar</surname>
          </string-name>
          , “
          <article-title>ICMPv6-based DoS and DDoS attacks detection using machine learning techniques, open challenges, and blockchain applicability: a review,” IEEE Access 8 (</article-title>
          <year>2020</year>
          ), pp.
          <fpage>170529</fpage>
          -
          <lpage>170547</lpage>
          . DOI:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2020</year>
          .3022963
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <string-name>
            <surname>Sheikh</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          <article-title>Study of ML Algorithms for DDoS Detection</article-title>
          ,
          <source>International Journal for Research in Applied Science and Engineering Technology</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Alshboul</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Streff</surname>
          </string-name>
          , “
          <article-title>Analyzing Information Security Model for Small-Medium Sized Businesses”</article-title>
          ,
          <source>in Proc. 21st Americas Conference on Information Systems, Puerto Rico</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>[8] DoS and DDoS vulnerability of IoT: A review</article-title>
          , URL: https:// www.researchgate.net/publication/39862422_DoS_
          <article-title>and_DDoS_vulnerability_of_IoT_A_review</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Simplício</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. V. M.</given-names>
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. C. A.</given-names>
            <surname>Alves</surname>
          </string-name>
          , and
          <string-name>
            <surname>T. K. C. Shibata</surname>
          </string-name>
          , “
          <article-title>Lightweight and escrowless authenticated key agreement for the internet of things”, Comput</article-title>
          . Commun.,
          <volume>98</volume>
          (
          <year>2017</year>
          ), pp.
          <fpage>43</fpage>
          -
          <lpage>51</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Czyz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Luckie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Allman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Bailey</surname>
          </string-name>
          , “
          <article-title>Don't forget to lock the back door! a characterization of ipv6 network security policy”</article-title>
          ,
          <string-name>
            <surname>in</surname>
            <given-names>NDSS</given-names>
          </string-name>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>K.</given-names>
            <surname>Angrishi</surname>
          </string-name>
          , “
          <article-title>Turning internet of things (iot) into internet of vulnerabilities (iov): Iot botnets”</article-title>
          ,
          <source>CoRR</source>
          , vol.
          <source>abs/1702.03681</source>
          ,
          <year>2017</year>
          . arXiv:
          <volume>1702</volume>
          .03681. URL: http://arxiv.org/abs/1702.03681.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>