<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Comparative study of machine learning applications in malware forensics⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bayan Abduraimova</string-name>
          <email>abduraimovabk@mail.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sergiy Gnatyuk</string-name>
          <email>sergio.gnatyuk@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Albina Nurmukhanbetova</string-name>
          <email>nuralbina9898@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CPITS-II 2024: Workshop on Cybersecurity Providing in Information and Telecommunication Systems II</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>L. N. Gumilyov Eurasian National University</institution>
          ,
          <addr-line>2 Satbayev str., 010000 Astana</addr-line>
          ,
          <country country="KZ">Kazakhstan</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>National Aviation University</institution>
          ,
          <addr-line>1 Liubomyra Huzara ave., 03058 Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>139</fpage>
      <lpage>152</lpage>
      <abstract>
        <p>In the digital world, one of the most significant threats is malicious software, known as malware, developed by cyber attackers to intentionally cause damage to computer systems or gain access to them. The form and behavior of malware have developed year by year at the same time methods of detection from malicious software have evolved to ensure security. Earlier for disclosure of malicious software some classical methods, such as signature-based, heuristic, and so on. Traditional methods for detecting malware have failed to defeat new generations of malware and their sophisticated obfuscation tactics. However, at present, the use of detection methods based on machine learning has been recognized as one of the most modern and prominent methods. Methods based on machine learning provide fast malware prediction with excellent detection and analysis rates for various types of malware. Therefore, this research work represents a comparison of various machine learning approaches, including neural networks, decision trees, the support vector machine, and ensemble methods applied to the analysis of behavioral and static characteristics of malware. The research paper analyzes the advantages and limitations of each method, their effectiveness in various types of malicious attacks, and the possibilities of their adaptation to changing threats. The current trends and directions of development in the use of machine learning in malware forensics, including the use of deep learning and big data analysis technologies, are also considered. The objective of this research is to conduct a comparative analysis of various machine learning techniques applied in malware forensics, focusing on their effectiveness in detecting, classifying, and analyzing malicious software. The study aims to identify the strengths and weaknesses of these techniques in real-world digital forensic scenarios, providing insights into their applicability and performance in enhancing malware investigation processes. The results of the study emphasize the prospects of using machine learning to improve malware detection and control, as well as the importance of further research to develop new models and methods of data analysis to protect information systems from modern cyber threats.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;cyber security</kwd>
        <kwd>malicious software</kwd>
        <kwd>malware forensics</kwd>
        <kwd>machine learning algorithms</kwd>
        <kwd>SVM</kwd>
        <kwd>decision tree</kwd>
        <kwd>random forest1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Malicious software, known as malware, is developed to
harm digital devices on purpose. Nowadays, there are many
types of malware, such as trojan horses, worms, viruses,
bots and botnets, ransomware, adware, and spyware.
Attackers targeted individual computers and networks,
which led to an increase of holes in the system of security.
The most significant issue in this situation is confidentiality
and leakage of personal information. Moreover, financial
damage caused by malware software all over the world
increased proportionally. According to the report of McAfee
about COVID-19 Threats and Malware Surges, the new
malware samples averaged 648 new threats per minute [1].
Also, recent estimates of the AV-Test show that every day
over 450,000 new malicious programs and potentially
unwanted applications [2]. To ensure an answer to cyber
threats, which are complex and unsafe, most of the farms
developed different instruments for cyber security.
Researchers used machine learning and deep learning
algorithms to create effective models to solve these
problems by conducting detailed research. Malware
protection methods are also becoming more sophisticated as
the variety of malware increases.</p>
      <p>Earlier, malware was written by using simple codes and
they were easily detected. Currently, creators of malicious
software complicate the code, as a result even advanced
methods cannot detect them. The next generation of
malicious programs is hard to identify by comparison with
traditional malicious programs designed to run in the
kernel. These types of malware programs slip away easily
0000-0003-3913-1895 (B. Abduraimova);
0000-0003-4992-0564 (S. Gnatyuk);
0009-0009-7123-9522 (A. Nurmukhanbetova)
© 2024 Copyright for this paper by its authors. Use permitted under
Creative Commons License Attribution 4.0 International (CC BY 4.0).
from the security system, such as firewalls and antiviruses.
There are different approaches to malware detection, based
on different functions, such as signature, heuristic,
behavior, model verification, cloud, mobile devices, Internet
of things, machine learning, and deep learning. At first, the
detection of malware programs used signature-based
detection. Despite its speed, it cannot identify complex
malware. Traditional signature-based methods, such as
pattern matching, do not meet the requirements for
malware detection. So, in this case, the use of machine
learning algorithms is suitable for the detection of malicious
software now [3].</p>
      <p>
        This paper represents malware detection using machine
learning algorithms and recommends each of them. In
Sections 1 and 2 we review traditional and machine learning
methods of detection of malicious software. Section 3
contains a detailed comparison table of machine learning
algorithms and shows the practical part in Python. Finally,
section 4 keeps the conclusion and final notes [
        <xref ref-type="bibr" rid="ref3 ref4">4, 5</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Traditional method of detection of malicious software</title>
      <p>
        Signature-based detection: In this approach, malware is
detected based on certain signatures or file patterns. This
traditional method allows the detection fastly known
malicious software compared to other methods. Most of the
antivirus programs are realized by using signature-based
methods. Signature of malware generated and stored in the
database of detection to verify the signature of an unknown
file. If the signature of the file matches, then it is declared a
malicious file, otherwise, it is a benign file. The main issue
of this method is that the file signature changes even when
one byte of the file is changed. Thus, for every modified and
new malware must be generated a new signature. After that,
the detector of malicious software can find a new signature.
Also, the main disadvantage of this detection is that it takes
a lot of time and effort to detect malware and extract
different signatures for different types of malware [
        <xref ref-type="bibr" rid="ref5">6</xref>
        ].
      </p>
      <p>
        Behavior-based malware detection: This method is
based on program behavior monitoring and decision
making whether it is malware or not. This method detected
malicious software by its suspicious behavior difference
from other typical programs. The behavior of malware
determines its importance in this type, and sometimes
different malware programs are found under the same
signature [
        <xref ref-type="bibr" rid="ref6">7</xref>
        ]. The benefit compared to the traditional
signature-based approach is that this method enables the
detection of new malware without requiring human
analysis and extraction of their signatures, as the behavior
of new malware may align with known patterns of
malicious behavior. However, it faces challenges related to
time consumption and false positive rates [
        <xref ref-type="bibr" rid="ref7">8</xref>
        ].
      </p>
      <p>
        Heuristic-based malware detection: This method studies
the behavior of malware by using machine learning
algorithms and intellectual data analysis. It solves many
problems, which are existing in signature-based and
behavior-based detection methods [
        <xref ref-type="bibr" rid="ref8">9</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Machine learning methods and algorithms</title>
      <p>
        Machine learning is modern technology, which is supposed
to be learning machines based on experience. Machine
learning can detect hidden patterns, which users can detect
easily themselves [
        <xref ref-type="bibr" rid="ref9">10</xref>
        ]. The process begins by providing the
machine learning algorithm with an input dataset. Then
constructed a model, which can do accurate predictions for
new datasets. Machine learning algorithms are divided into
three main categories. Firstly, supervised learning, which
includes machine learning by using a set of examples with
their answers. The algorithm uses these examples to
respond to any new input based on what it has learned.
      </p>
      <p>Secondly, an unsupervised model uses data sets without
answers. Here, the algorithm classified input data based on
the similarity between these values.</p>
      <p>
        Analysis of malicious software includes the process of
analyzing and understanding the malicious software for
determination of its functionality, behavior, and potential
impacts. Analysis of malware provides a clear view of
malware in the code and byte string values, which is used
by an attacker to steal information or modify the source
code. In the process of analysis, various key features of the
code are extracted, which provides information and
functionality of malware activity in the system or network
system [
        <xref ref-type="bibr" rid="ref10">11</xref>
        ].
There are three types of the most frequently used
methods [
        <xref ref-type="bibr" rid="ref11">12</xref>
        ], as shown in Fig. 2.
      </p>
      <p>Static analysis: Static analysis of malicious software
includes the study of malware or files without execution.</p>
      <p>
        This method of analysis allows users and analytics of
security to understand structure, behavior, and potential
threats, coming from malware, without the risk of
infecting the system.
Static analysis can present valuable information about
behavior, opportunities, and potential influence of
malware on systems. This analysis is useful for the fast
detection of known malware signatures and the
extraction of important information from the code or
files [
        <xref ref-type="bibr" rid="ref12">13</xref>
        ]. The functionality of malware can be analyzed
by checking the internal malware code. It provides
information about the identity of the malware, library,
URL addresses, and programming languages as shown
in Fig. 3. The process of static analysis executes faster
and provides deeper knowledge about the malware
execution path. However, it has drawbacks: it does not
detect new variants of malware families or polymorphic
malware specifically designed to evade static analysis
[
        <xref ref-type="bibr" rid="ref13">14</xref>
        ].
      </p>
      <p>
        Dynamic analysis: At the same time, malicious code
is executed in a controlled environment to observe its
behavior and understand its capabilities. Difference
from the static analysis that researches the code without
its execution [
        <xref ref-type="bibr" rid="ref14">15</xref>
        ]. Dynamic analysis provides real-time
monitoring of how malware adapts and behaves in
response to specific environmental conditions.
      </p>
      <p>
        Observing malicious software in action, analysts can
gain insight into its purpose, distribution methods, and
the potential damage it can cause [
        <xref ref-type="bibr" rid="ref15">16</xref>
        ]. However,
dynamic analysis carries a certain risk, since malware is
actively working and there is a possibility of unforeseen
consequences. This is why, it is significant to carry out
dynamic analysis in a controlled and isolated
environment to minimize the potential impact on the
host system and the network as a whole [
        <xref ref-type="bibr" rid="ref16">17</xref>
        ]. The
process of dynamic analysis is shown in Fig. 4.
      </p>
      <p>
        Hybrid malware analysis: Hybrid malware analysis,
also known as combined analysis or integrated analysis,
is an approach combining several methods, such as static
analysis, dynamic analysis, and behavior analysis, to get
a complete understanding of malware. Using the strong
sides of different analysis methods, hybrid analysis is
aimed at overcoming the limitations of individual
methods and providing a more reliable and accurate
assessment of the behavior, capabilities, and potential
impact of malware [
        <xref ref-type="bibr" rid="ref17">18</xref>
        ].
      </p>
      <p>Machine learning is the subset of AI, used for tasks
of detection and classification in different areas.</p>
      <p>Algorithms of machine learning, such as RF, NB, KNN,
SVM, and DT, are widely used for the detection and
classification of malicious programs.</p>
      <p>Naive Bayes (NB) is an algorithm of classification
based on supervised learning. It works in the base of
probability functions, in which each attribute belongs to
a specific class. A strong assumption is required that the
attributes are conditionally independent. An
assumption, which is used in algorithm NB simplifies
the probability calculation. It calculates whether a data
point can fall into a certain class or not. It can accurately
predict test data sets for binary and many class datasets.
It is effective and scales well, therefore, it works well for
small irrelevant data sets. It used the NB theorem for the
decomposition of conditional probability. However, NB
gives poor performance because attributes are highly
correlated with each other, and it is believed that
attributes should be independent.</p>
      <p>K-Nearest Neighbors is a nonparametric supervised
algorithm, which means that it doesn't make any
assumptions about the underlying distribution of the
data. It works based on the proximity of the similarity of
objects or searching for Nearest Neighbors in the
specified set and uses the majority of votes to classify a
new data point. KNN defines the similar characteristics
of new points on the base of previously saved data
points, using the Euclidean distance between two data
points. This algorithm is known as lazy learning because
it does not require parameter adjustment and works
without training the model. Instead of training the
model, it takes all the data points at the time of
forecasting. However, it has some disadvantages, such
as high cost, low speed, and less scalability for large data
sets.</p>
      <p>Decision Tree is a classification of supervised
learning, used for tasks of classification and regression.</p>
      <p>
        It can also work with numbers and category data. It
follows the tree model approach, in which the data
points are divided into two branches and a conclusion
about existing features in each node of the tree [
        <xref ref-type="bibr" rid="ref19">20</xref>
        ].
      </p>
      <p>Pros Cons
Fast and safe. High accuracy. It is impossible to analyze confusing and encrypting
Low memory consumption. malware. Unknown malware cannot be detected.
Can analyze confusing and encrypting Consumes a high number of resources and is unsafe
malware. Also, it can detect known and and slow.
unknown malware.</p>
      <p>The result is more accurate than static Higher complexity and more time are used.
and dynamic analysis.</p>
      <p>
        A decision tree is designed for the creation of a learning
flowchart structure, which can be used for the
classification of classes or values targeted variables on
the base of decision-making rules, which are extracted
from previous data. It can work well with huge and
noisy data sets compared to models of KNN and SVM
[
        <xref ref-type="bibr" rid="ref20">21</xref>
        ].
Support Vector Machine is a powerful supervised model,
which is used as a task of classification so it is for
regression. The goal of SVM is to create the best
hyperplane, which exactly divides data sets of one class
from another class of data sets of learning and testing.
      </p>
      <p>
        In the data sets of SVM a set of data points separated by
a line called a hyperplane is specified, which is used for
classification of data sets classes. This is an effective
classifier for effectively processing a large set of data. It
is used for problem decisions of linear functions in
multidimensional feature spaces. However, this does not
work well in the case of big data sets with noisy data sets
and requires more time for learning [
        <xref ref-type="bibr" rid="ref20">21</xref>
        ].
      </p>
      <p>
        Random Forest is a classifier based on ensemble
learning, which consists of several decision trees, that
work parallel, a forest that is constructed and is called
decision tree ensemble of the decision tree. Each tree in
the ensemble model contains data samples taken from
the training data with replacement, known as a
bootstrap sample, which is a set of learning models to
improve the overall results of the models [
        <xref ref-type="bibr" rid="ref19">20</xref>
        ]. The
decision about targeted classes was adopted by a
majority of votes in Random Forest. This is an easy,
flexible, and simple algorithm, which gives the best
result most of the time, even without setting the
hyperparameters. This reduces over-training, which
leads to improving the accuracy of the decision tree [
        <xref ref-type="bibr" rid="ref21">22</xref>
        ].
      </p>
      <p>
        It works for categorical and continuous data. However,
the speed of Random Forest is low, because it requires
more time on learning a set of decision trees for
constructing strong classificatory [
        <xref ref-type="bibr" rid="ref20">21</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Machine learning malware detection and its application</title>
      <p>Confusion matrix is a matrix that is constructed for
every classification model prediction and it shows the
number of test cases correctly and incorrectly classified.
It looks like this, which is shown in Table 3 (considering
1 (positive) and 0 (negative) are the target classes):</p>
      <p>+ 
+ 
+ 
+ 
It can implement many general issues, but it is not useful
when we are talking about unbalanced data sets. For
example, if we detect fraud in bank datasets, the ratio of
fraud cases to non-fraud cases can be 1:99. In this case,
if we use the accuracy, the model can prove to be 99%
accurate, predicting all test cases as non-fraud. This way,
accuracy is a false indicator of the model's performance,
and in this case, the metric is required that can focus on
fraud data points.</p>
      <p>First of all, for doing research, necessary libraries,
such as scipy, seaborn, and Tensorflow to analyze the
data, build graphs, and study neural networks to find
and classify images, which is shown in Fig.5.</p>
      <p>After that, this code which is shown below imports
pandas, numpy, tensorflow, keras, and matplotlib
libraries for working with data, creating and training
machine learning models, and visualization. Fixed
random number generator seeds are set to ensure
reproducibility of results. Some tools for data
preprocessing and model evaluation are commented out
and are not currently used.</p>
      <p>This code loads data from a CSV file using pandas and
stores it in the variable data. It then prints the total
number of missing values in this dataset, using the isna()
method to determine the missing values and the sum()
function to count them. Finally, the code prints the
DataFrame data itself.
This code parses the unique values and their counts in
the “class” column, performs label encoding using
“LabelEncoder”, replaces special characters with “NaN”,
removes rows with missing values, converts all columns
to numeric data types, and outputs the total number of
features in the DataFrame after all transformations.
The code, which is shown in Fig.8 prints the total
number of features in the DataFrame (excluding the
class column), and then creates and displays a bar chart
of the class distribution, showing the number of instanes
for each class.
oversamples the minority class to equalize the number
of instances with the majority class, and then merges
and shuffles the balanced dataset.</p>
      <p>The code below balances the classes in a dataset by
counting the number of instances of each class and
outputting that data. It then splits the data into training
and testing sets, where 80% is used to train the model
and 20% is used for testing, and outputs the sizes of both
sets.</p>
      <p>Below is the code that builds a heat map of
correlations between features in a dataset. It uses the
seaborn and matplotlib libraries for visualization,
displaying correlation coefficients as a color map. The
heat map helps to identify dependencies and
relationships between different features.</p>
      <p>
        The code below trains a decision tree classifier on
the training data, makes predictions on the test data, and
computes the accuracy of the model.
The code, which is shown in Fig. 12 creates a graph
showing the accuracy of different classification
algorithms. It plots a bar chart using a color map to show
the accuracy and adds labels to each bar. It also displays
a color scale indicating the accuracy range.
This paper represented malware detection methods
using machine learning along with basic concepts such
as malware detection programs, and machine learning.
Also, there were several classification algorithms and
the accuracy of methods varied depending on the
method used, the number of attributes, data sets,
preprocessing methods, as well as tools implemented in
the model. It also depends on the analysis method and
the function used. As seen in Table 4, the best possible
malware detection accuracy is a Sequential Neural
Network. This algorithm achieves about 100% [
        <xref ref-type="bibr" rid="ref22 ref23">23, 24</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>Sequential Neural Network refers to one of the neural
network’s architecture types, where data is processed
sequentially, step by step. This is useful for tasks,
including time series data, such as text or speech. Based
on work, this algorithm reaches the best results.
However, it should be recognized that this conclusion
does not allow us to claim that the Sequential NN is the
best algorithm for detecting malware. There are always
other measures to take care of, such as the dataset,
preprocessing methods, extracted features, and feature
selection methods used.</p>
      <p>Further research directions include several key
aspects. First, it is worth exploring the possibilities of
developing hybrid models that combine different
machine learning algorithms to improve overall
efficiency and robustness to changing data conditions.
Second, it is worth focusing on improving the
interpretability of complex models such as neural
networks to make them more suitable for forensic
practice. Third, it is necessary to explore adaptive
learning methods that will allow models to quickly
respond to the emergence of new types of malware. It is
also worth conducting a comparative analysis of
algorithms in different digital forensics scenarios to
identify specific requirements and optimal solutions for
each task. Finally, it is important to investigate the
impact of different tools and frameworks for
implementing machine learning on the analysis results,
which can significantly improve the practical
application of these technologies in forensic science.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>R.</given-names>
            <surname>Samani</surname>
          </string-name>
          ,
          <source>McAfee Labs Report Reveals Latest COVID-19 Threats and Malware Surges. McAfee Blog</source>
          (
          <year>2024</year>
          ). URL: https://www.mcafee.com/blogs/ other-blogs/
          <article-title>mcafee-labs/mcafee-labs-report-revealslatest-covid-19-threats-and-malware-surges/ AV-TEST | Antivirus &amp; Security Software &amp; Antimalware Reviews</article-title>
          . URL: https://av-test.org/ en/statistics/malware.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>O.</given-names>
            <surname>Mykhaylova</surname>
          </string-name>
          , et al.,
          <article-title>Person-of-Interest Detection on Mobile Forensics Data-AI-Driven Roadmap</article-title>
          ,
          <source>in: Workshop on Cybersecurity Providing in Information and Telecommunication Systems, CPITS</source>
          , vol.
          <volume>3654</volume>
          (
          <year>2024</year>
          )
          <fpage>239</fpage>
          -
          <lpage>251</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>V.</given-names>
            <surname>Buhas</surname>
          </string-name>
          , et al.,
          <article-title>Using Machine Learning Techniques to Increase the Effectiveness of Cybersecurity</article-title>
          ,
          <source>in: Cybersecurity Providing in Information and Telecommunication Systems</source>
          , vol.
          <volume>3188</volume>
          , no.
          <issue>2</issue>
          (
          <year>2021</year>
          )
          <fpage>273</fpage>
          -
          <lpage>281</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>V.</given-names>
            <surname>Zhebka</surname>
          </string-name>
          , et al.,
          <article-title>Methodology for Predicting Failures in a Smart Home based on Machine Learning Methods</article-title>
          ,
          <source>in: Workshop on Cybersecurity Providing in Information and Telecommunication Systems, CPITS</source>
          , vol.
          <volume>3654</volume>
          (
          <year>2024</year>
          )
          <fpage>322</fpage>
          -
          <lpage>332</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Scott</surname>
          </string-name>
          ,
          <article-title>Signature based Malware Detection is Dead, Institute for Critical Infrastructure Technology (</article-title>
          <year>2017</year>
          ). https://informationsecurity. report/Resources/Whitepapers/920fbb41-8dc9
          <string-name>
            <surname>-</surname>
          </string-name>
          4053
          <string-name>
            <surname>-</surname>
          </string-name>
          bd01-72f961db24d9
          <string-name>
            <surname>_ICIT-AnalysisSignature-Based-Malware-</surname>
          </string-name>
          Detection-is-Dead.pdf
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Shabtai</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>Anomaly</surname>
            :
            <given-names>A Behavioral</given-names>
          </string-name>
          <string-name>
            <surname>Malware</surname>
          </string-name>
          <article-title>Detection Framework for Android Devices</article-title>
          ,
          <source>Journal of Intelligent Information Systems</source>
          , vol.
          <volume>38</volume>
          , no.
          <issue>1</issue>
          (
          <year>2012</year>
          )
          <fpage>161</fpage>
          -
          <lpage>190</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>I.</given-names>
            <surname>Firdausi</surname>
          </string-name>
          , et al.,
          <source>Analysis of Machine Learning Techniques used in Behavior-based Malware Detection, in: Proceedings of the 2010 Second International Conference on Advances in Computing, Control, and Telecommunication Technologies</source>
          . Washington, DC, USA: IEEE Computer Society (
          <year>2010</year>
          )
          <fpage>201</fpage>
          -
          <lpage>203</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Bazrafshan</surname>
          </string-name>
          , et al.,
          <source>A Survey on Heuristic Malware Detection Techniques, in: 5th Conference on Information and Knowledge Technology</source>
          (
          <year>2013</year>
          )
          <fpage>113</fpage>
          -
          <lpage>120</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zakeri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Daneshgar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Abbaspour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Static</given-names>
            <surname>Heuristic</surname>
          </string-name>
          <article-title>Approach to Detecting Malware Targets, Security and Communication Networks</article-title>
          , vol.
          <volume>8</volume>
          , no.
          <volume>17</volume>
          (
          <year>2015</year>
          )
          <fpage>3015</fpage>
          -
          <lpage>3027</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ahmadi</surname>
          </string-name>
          , et al.,
          <source>Malware Detection by Behavioural Sequential Patterns, Comput. Fraud Secur.</source>
          , vol.
          <volume>8</volume>
          (
          <year>2013</year>
          )
          <fpage>11</fpage>
          -
          <lpage>19</lpage>
          . doi:
          <volume>10</volume>
          .1016/S1361-
          <volume>3723</volume>
          (
          <issue>13</issue>
          )
          <fpage>70072</fpage>
          -
          <lpage>1</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Akour</surname>
          </string-name>
          , I. Alsmadi,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alazab</surname>
          </string-name>
          ,
          <source>The Malware Detection Challenge of Accuracy, in: 2nd Int. Conf. Open Source Softw. Comput</source>
          . (
          <year>2017</year>
          )
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          , doi: 10.1109/osscom.
          <year>2016</year>
          .
          <volume>7863676</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Alazab</surname>
          </string-name>
          , et al.,
          <article-title>A Hybrid Wrapper-Filter Approach for Malware Detection</article-title>
          ,
          <source>Journal of Networks</source>
          , vol.
          <volume>9</volume>
          , no.
          <volume>11</volume>
          (
          <year>2014</year>
          )
          <fpage>2878</fpage>
          -
          <lpage>2891</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vasilescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Gheorghe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Tapus</surname>
          </string-name>
          ,
          <source>Practical Malware Analysis based on Sandboxing, Proc. RoEduNet</source>
          IEEE International Conference (
          <year>2014</year>
          ). doi:
          <volume>10</volume>
          .1109/RoEduNet-RENAM.
          <year>2014</year>
          .
          <volume>6955304</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Alazab</surname>
          </string-name>
          , et al.,
          <string-name>
            <given-names>A Hybrid</given-names>
            <surname>Wrapper-Filter Approach</surname>
          </string-name>
          for Malware Detection,
          <source>J. Networks</source>
          , vol.
          <volume>9</volume>
          , no.
          <volume>11</volume>
          (
          <year>1969</year>
          )
          <fpage>2878</fpage>
          -
          <lpage>2891</lpage>
          . doi:
          <volume>10</volume>
          .4304/jnw.9.11.
          <fpage>2878</fpage>
          -
          <lpage>2891</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kumar</surname>
          </string-name>
          , et al.,
          <article-title>Malicious Data Classification using Structural Information and Behavioral Specifications in Executables</article-title>
          , in: Recent Adv.
          <source>Eng. Comput. Sci. RAECS</source>
          (
          <year>2014</year>
          ) pp.
          <fpage>6</fpage>
          -
          <lpage>8</lpage>
          , doi: 10.1109/RAECS.
          <year>2014</year>
          .
          <volume>6799525</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [17]
          <string-name>
            <surname>M.A.M. Ali</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          <string-name>
            <surname>Maarof</surname>
          </string-name>
          ,
          <article-title>Dynamic Innate Immune System Model for Malware Detection</article-title>
          , in: Int. Conference IT Converg.
          <source>Security ICITCS</source>
          (
          <year>2013</year>
          )
          <fpage>3</fpage>
          -
          <lpage>6</lpage>
          , doi: 10.1109/ICITCS.
          <year>2013</year>
          .
          <volume>6717828</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>I.</given-names>
            <surname>You</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yim</surname>
          </string-name>
          ,
          <article-title>Malware Obfuscation Techniques: A Brief Survey</article-title>
          , in: Proceedings of International Conference Broadband, Wirel,
          <source>Comput. Commun. Appl</source>
          . (
          <year>2010</year>
          ),
          <fpage>297</fpage>
          -
          <lpage>300</lpage>
          , doi: 10.1109/BWCCA.
          <year>2010</year>
          .
          <volume>85</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <surname>A Survey</surname>
          </string-name>
          <article-title>on Machine Learningbased Malware Detection in Executable Files</article-title>
          ,
          <string-name>
            <surname>J. Syst. Archit.</surname>
          </string-name>
          , vol.
          <volume>112</volume>
          (
          <year>2021</year>
          )
          <article-title>101861</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.sysarc.
          <year>2020</year>
          .
          <volume>101861</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>W.</given-names>
            <surname>Kanyongo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. E.</given-names>
            <surname>Ezugwu</surname>
          </string-name>
          ,
          <source>Feature Selection and Importance of Predictors of NonCommunicable Diseases Medication Adherence from Machine Learning Research Perspectives, Informatics in Medicine Unlocked</source>
          , vol.
          <volume>38</volume>
          (
          <year>2023</year>
          )
          <article-title>art</article-title>
          .
          <fpage>101232</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>G.</given-names>
            <surname>Iashvili</surname>
          </string-name>
          , et al.,
          <source>Content-based Machine Learning Approach for Hardware Vulnerabilities Identification System, Lecture Notes on Data Engineering and Communications Technologies</source>
          , vol.
          <volume>83</volume>
          (
          <year>2021</year>
          )
          <fpage>117</fpage>
          -
          <lpage>126</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Guo</surname>
          </string-name>
          , et al.,
          <article-title>A Novel Deep Learning Model Integrating CNN and GRU to Predict Particulate Matter Concentrations</article-title>
          ,
          <source>Process Safety and Environmental Protection</source>
          , vol.
          <volume>173</volume>
          (
          <year>2023</year>
          )
          <fpage>604</fpage>
          -
          <lpage>613</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.psep.
          <year>2023</year>
          .
          <volume>03</volume>
          .052.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Imanbayev</surname>
          </string-name>
          , et al.,
          <source>Research of Machine Learning Algorithms for the Development of Intrusion Detection Systems in 5G Mobile Networks and Beyond, Sensors</source>
          , vol.
          <volume>22</volume>
          , iss.
          <volume>24</volume>
          (
          <year>2022</year>
          )
          <fpage>9957</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>N.</given-names>
            <surname>Baisholan</surname>
          </string-name>
          , et al.,
          <article-title>Implementation of Machine Learning Techniques to Detect Fraudulent Credit Card Transactions on a Designed Dataset</article-title>
          ,
          <source>Journal of Theoretical and Applied Information Technology</source>
          , vol.
          <volume>101</volume>
          (
          <issue>13</issue>
          ) (
          <year>2023</year>
          )
          <fpage>5279</fpage>
          -
          <lpage>5287</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>