<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Malware detection through low-level features and stacked denoising autoencoders</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessandra De Paola</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Salvatore Favaloro</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Salvatore Gaglio</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Lo Re</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Morana</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>firstname.lastname@community.unipa.it</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>In recent years, the di usion of malicious software through various channels has gained the request for intelligent techniques capable of timely detecting new malware spread. In this work, we focus on the application of Deep Learning methods for malware detection, by evaluating their e ectiveness when malware are represented by high-level, and lowlevel features respectively. Experimental results show that, when using high-level features, deep neural networks do not signi cantly improve the overall detection accuracy. On the other hand, when low-level features, i.e., small pieces of information extracted through a light processing, are chosen, they allow to increase the capability of correctly classifying malware.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Malware detection is one of the most critical issues faced by computer security. Nowadays, the
accidental execution of malicious software coming from di erent channels makes IT systems
constantly exposed to risks. In order to e ectively detect threats hidden behind heterogeneous
software, which is susceptible to unpredictable variations, intelligent techniques are required.</p>
      <p>
        The most common approach implemented by popular malware-detection tools is to perform
a static analysis of a short sequence of bytes, i.e., a signature [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Nevertheless, the availability
of softwares allowing the automatic generation of several variants of a certain malware greatly
reduces the e ectiveness of this technique [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
      </p>
      <p>One of the most promising direction followed to guarantee a high detection rate even with
the constantly increase of threats, is the adoption of cloud-based approaches where the malware
detection is performed remotely through machine learning techniques capable of analyzing a
huge amount of malware les. The e ectiveness of this solution often relies on the adoption
of high-level features, extracted from executable les, designed to be more invariant to code
obfuscation and polymorphism than classic signature-based methods.</p>
      <p>Nevertheless, it is worth noticing that designing \good" set of features to realize e cient
and e ective systems is not trivial, due to the underlying and not explicit dependences among
di erent parts of a le and its behavior. A promising way to face such a problem is represented
by deep learning, which is expected to enable the automatic extraction of relevant features,
directly from raw information.</p>
      <p>Although the potentiality of applying deep learning on simple raw data is con rmed by some
recent results, many works adopt deep learning to perform malware detection or classi cation
on high-level features obtained from executable les.</p>
      <p>
        The rst goal of the proposed work is to evaluate the real e ectiveness of malware detection
through deep learning, while processing high-level features. To this aim, we compare the system
proposed in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] with other classi ers of various complexity.
      </p>
      <p>
        Moreover, we want to verify if deep networks can be e ectively adopted to perform
malware detection by exploiting a reduced set of low-level features extracted from the executable
les. The solution we propose here is based on a deep network implemented through stacked
denoising autoencoders, which analyzes raw information contained in the portable executable
(PE) packaging of Windows executable les. Our aim is to compare the performance obtained
by such lighter malware detection system with the most complex approach proposed in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], in
order to establish whether deep learning is able to achieve good performance even with fewer
and less pre-processed data.
      </p>
      <p>
        The following of the paper is structured as follow. Sect. 2 presents an overview of the related
work. Sect. 3 describes the high-level features and the system proposed in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], besides the
simpler classi ers we use for the comparative analysis. Sect. 4 provides a description of the
proposed malware detection system based on low-level features. Finally, Sect. 5 reports the
experimental evaluation, and Sect. 6 presents our conclusions and possible future work.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Several works which adopt deep learning to perform malware detection have been presented in
the recent literature, based both on dynamic and static analysis [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The dynamic analysis
exploits a protected environment, i.e., a \sandbox", in which the malware can be executed, and
its behavior observed, without threaten a real working system, while the static analysis focuses
only on information included in the target le, and do not require that the potential malware is
run. As a consequence, static methods are generally faster and less greedy for resources, even
if they can exploit less information and are vulnerable to malicious code obfuscation.
      </p>
      <p>
        The system proposed in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] performs a dynamic analysis, by adopting deep learning to
automatically generate the signature which represents such behavior. The authors use a deep
belief network implemented through stacked denoising autoencoders, which processes the text
le containing the transcription of all the events occurred during a le run, adequately converted
in a binary form. The automatically-generated signature is then processed by a SVM to perform
the e ective classi cation. In [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], the sequence of system calls recorded while the executable
le is running is exploited by a Deep Neural Network (DNN) which combines convolutional
layers and recurrent layers, using Long Short Term Memory (LSTM) cells to increase malware
detection capabilities.
      </p>
      <p>
        In order to overcome the limitation of dealing with less information than those obtained
by dynamic analysis, some works based on static analysis performs deep learning on high-level
features extracted according to a well-designed process. The authors of [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] propose to process a
le to build a large vector of features (179 thousands of sparse binary features), which is reduced
through a random projection process. The resulting vector, composed of 4,000 elements, is then
analyzed by a deep classi er which is pre-trained through a Restricted Bolzmann Machines
(RBM). The reduction of the feature vectors through random projection is also adopted in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ],
where JavaScript sources are processed through a 5{layer deep neural network implemented
with stacked denoising autoencoders. In [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] four di erent sets of static features, converted in
a 1024-length binary vectors and classi ed by a classic DNN, are used. A more complex set of
features is propose in [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], where data is obtained by merging information coming from static
and dynamic analysis. Such set of features includes a sub-set automatically obtained by DNNs.
The resulting advanced set of features is classi ed through Multiple Kernel Learning.
      </p>
      <p>
        Even though malware detection algorithms which combine deep learning and high-level
features usually provide better accuracy values, some authors emphasize the convenience to adopt
features as simple as possible so as to design light and e cient malware-detection systems [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        The potentiality of applying deep learning on simple raw data is con rmed by some works
recently presented in the literature, such as [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], where an android malware detection system
which applies a deep convolutional neural network to the raw sequences of opcode extracted
from disassembled programs is presented.
      </p>
      <p>Our work aims to con rm this idea, by proving that deep learning can be e ciently adopted
to perform malware detection when executable les are represented by simple features.
3</p>
    </sec>
    <sec id="sec-3">
      <title>High-level features for malware detection</title>
      <p>Over the years, several type of features have been adopted for malware detection through static
analysis. Some of them are obtained from a heavy pre-processing of the executable le to
summarize its global characteristics, or through a simple and light pre-processing phase. The
former can be considered high-level features, since they allow to represent the executable le at
a higher level of abstraction. The latter can be de ned low-level features, since they are closer
to the raw le representation.</p>
      <p>
        Authors of [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] proved the e ectiveness of several machine learning methods using high-level
features for distinguishing between packed and non-packed executables. This work adopts a
small set of features which summarize high-level aspects of the analyzed software, i.e., the
number of standard and non-standard sections in the portable executable (PE) packaging of a
Windows executable le, the number of executable sections, the number of
readable/writable/executable sections, the number of entries in the Import Address Table, and the entropy of PE
header, code section, data section, and of the whole executable le. The authors of [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] analyze
the frequencies of opcode sequences classifying them through Support Vector Machines. Some
works focus on the analysis of the frequency of Windows API calls, such as [
        <xref ref-type="bibr" rid="ref21 ref25 ref3">3, 25, 21</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ],
sequences of strings and bytes extracted from the le are considered together with the set of
API calls. Authors of [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] propose to use n-grams of strings, whereas in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] n-grams
based on n-length system calls sequences are used. Authors of [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] propose to merge a large set of
complex features by selecting the best subset of features through a forward stepwise selection.
The set of features includes n-grams of bytes, the byte-level entropy vector, some descriptors
of the image obtained by interpreting each byte as a grey level, the histogram of the length
of strings, the frequency of some speci c operation codes, the frequency of API calls, together
with many other features.
      </p>
      <p>
        One of the goal of our work is to verity whether, still using high-level features, deep learning
can further improve the detection accuracy. As target system for our analysis, we consider
the work proposed in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], which adopts a deep neural network to classify vectors of
highlevel features extracted from executable les. The feature vector is obtained by merging four
complementary histograms, which depend on di erent information extracted from the analyzed
le. Two mono-dimensional histograms capture information from the PE, while the other two
bi-dimensional histograms depend on the whole le. The concatenation of these four histograms,
after the row-by-row concatenation of the bi-dimensional components, produces a single
monodimensional vector composed by 1024 elements. The rst mono-dimensional part of the feature
vectors is the PE import histogram, obtained by reading the import address table header of PE
and generating the tuples of DLLs and functions listed in the header. Each tuple is mapped in
the corresponding hash value and the component is computed as the histogram of such hash
values. The second mono-dimensional part is the PE metadata histogram, which depends on
labels and values of numerical elds contained in the PE packaging. Each label-value pair is
mapped in the corresponding hash value and the histogram of the hash values is computed.
The bi-dimensional components of the feature vector are the byte/entropy histogram and the
string histogram. The former is created by sliding a 1024-byte-sized window over the le with
step of 256 bytes, and computing the base-2 entropy for each window. For each byte in the
window, a byte-entropy pair is computed; then, the set of pairs is processed to create a 16x16
byte/entropy matrix, i.e., the nal 2D histogram. The string histogram is obtained by extracting
the le strings long at least 6 characters. Each string is processed to obtain a pair containing
the hash value of the string itself and the logarithm of its length. All the resulting pairs are
considered to obtain a 16x16 hash/length histogram.
      </p>
      <p>
        The malware detection process is based on a deep neural network composed by four layers:
an input layer consisting of 1024 nodes, two hidden layers, each consisting of 1024 nodes,
and a nal layer with a single node. Nodes of the two hidden layers adopt the Parametric
Recti ed Linear Unit (PReLU) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] as activation functions, whereas activation of the node in
the nal layer depends on a sigmoid. PReLU is an activation function, recently proposed in the
literature, which autonomously modi es its form according to a speci c parameters that can be
changed to speed up the training phase. Training is performed by means of a back-propagation
algorithm and the Adam gradient-based optimizer [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], a recent stochastic optimizer that uses
the rst and second moments of the gradient to minimize the objective function, i.e., the binary
cross-entropy. Finally, to prevent over tting, dropout regularization [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] is performed on the
rst three layers. The dropout technique disables some units of the neural network during the
training so as to process a narrow sub-graph rather than the whole network. On every learning
round, di erent units are randomly sampled to be disabled, so guaranteeing that the training
is performed on di erent sub-networks.
      </p>
      <p>Sect. 5 presents the comparative analysis of such target system with other classi ers of
decreasing complexity, which adopt the same set of high-level features. Surprisingly, the results
we obtained prove that with this well-designed set of high-features, deep learning does not have
a relevant impact on detection accuracy as compared with other simpler classi ers.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Low-level features and DNNs for malware detection</title>
      <p>
        Rather than exploiting high-level features, we propose to adopt deep networks to process
lowlevel features extracted through static analysis, in order to automatically learn which
information is relevant to malware detection. The system we propose is based on a two-phase
training deep neural network, where a rst unsupervised pre-training with stacked denoising
autoencoders [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] is followed by a supervised ne-tuning based on back-propagation.
      </p>
      <p>Features extraction is performed by accessing DOS Header, File Header, Optional Header
and Section Table of the PE packaging. Each header contains di erent elds, each analyzed
by extracting the corresponding value and two o sets. Simple values are handled as unsigned
integers, whilst timestamps, arrays and strings, are processed via a hashing function. O set
values allow to preserve spatial information: the local o set speci es the position of a eld
within the header section, and the global o set represents the position in the le.</p>
      <p>
        Since a le can have a variabile number of sections and the classi er is designed to process
xed-size data, we limit the number of Section Tables to be processed. To this aim, some
experiments were performed on real data showing that a reasonable threshold on the number of
sections is 13. If a le contains less sections than the threshold then the elements of the feature
vector related to the missing sections are set to zero; otherwise, if the number of sections is
higher than the threshold, the extra sections will be ignored. Thus, 19 elds of the DOS header,
7 of the le header, 30 of the optional header, 12 of the section section header (for each of the
13 sections) were analysed, obtaining feature vectors of 636 elements in the range [0; 1]. These
vectors are about 38% smaller than those proposed in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], and described in Sect. 3. Moreover,
      </p>
      <p>Unlabeled  
dataset  </p>
      <p>Unsupervised  
pre-­‐training  
labeled  
dataset  </p>
      <p>Fine  tuning  
input  
corrupted  
input  
encoder  
decoder  
since data are not extracted from the entire le but only from its headers, the feature extraction
process is fast and independent of the le size.</p>
      <p>The deep network proposed here is consists of ve layers. The input layer has 636 elements,
the three hidden layers contain 256, 64, and 16 nodes respectively, whilst the output layer
consists of a single node. Each node use a sigmoid activation function. In this neural network,
the total amount of trainable hyper-parameters is of 180; 577 elements.</p>
      <p>The training of the deep network is performed in two steps (see Fig.1). The rst phase
performs an unsupervised pre-training by using unlabeled dataset to obtain a rst estimation of
weights and biases of the hidden layers. Such phase is implemented through stacked denoising
autoencoders. Each hidden layer is pre-trained individually by means of a support network
consisting of an input layer, a corrupted layer, an encoder, and a decoder layer (see Fig. 2),
where the encoder layer corresponds to the hidden layer to be pre-trained. On the contrary,
the input, the corrupted, and the decoder layers have the same number of nodes, which is
equal to the number of inputs of the hidden layer to be pre-trained. This support network is
trained through back-propagation and processes each input vector by i) corrupting it with some
kind of noise, ii) encoding the resulting noisy signal, and then iii) reconstructing the encoded
signal. The aim of the corruption layer, which adds isotropic Gaussian noise to the input
signals, is to force the encoding layer to learn the most useful information from input vectors,
by automatically neglecting noise from the corrupted input. The iteration of the pre-training
for all the hidden layers produces a deep neural network in which each layer is able to extract
and represent features at a higher level of abstraction than its predecessors. The pre-training
of the rst hidden layer is performed using the original dataset, while the other hidden layers
are pre-trained through an encoded version of the dataset, obtained by exploiting the current
trained denoising autoencoders to build a temporary network which extracts the output of the
encoding layer that precedes the layer to be trained.</p>
      <p>The second training step is the ne-tuning of the network, which is implemented through
a supervised back-propagation algorithm with the Adam stochastic optimization, applied by
considering the binary cross entropy as objective function. During this stage, weights and bias
of all hidden layers are initialized with the values produced by the pre-training.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Experimental Evaluation</title>
      <p>The dataset used for the experimental evaluation was obtained by merging 12.000 samples of
malware obtained from VirusShare 1 and 11.874 samples of certi ed software obtained from a
clean Windows 10 installation. The malware samples are not labeled according to the speci c
malware family they belong to, but this characteristic does not limit the experimental evaluation
presented here, since it aims only to assess the detection accuracy.</p>
      <p>
        The evaluation of each classi er is performed through a K-fold cross validation with a
strati ed sampling, in order to preserve the percentage of samples per class. Due to the limited
size of the dataset used here, we adopted K = 5 for each test so as to guarantee that each
validation set contains an adequate number of samples. Each model was trained until the loss
value dropped below 0:02, or the number of training epochs exceeded 200, as proposed in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>The performance of each classi er were evaluated by analyzing the trend of the ROC curve
(Receiver Operating Characteristic) with respect to the training epochs. Furthermore we
computed the nal loss, the nal accuracy, as well as several other metrics, i.e., TPR (True Positive
Rate), FPR (False Positive Rate), Precision, and AUC (Area Under Curve) of ROC curve. This
last metrics allows to evaluate the classi cation performance independently on the threshold
adopted for the last layer.
5.1</p>
      <sec id="sec-5-1">
        <title>Deep Learning performance with high-level features</title>
        <p>In order to verify the impact of deep learning methods when high-level features are used to
represent les, we performed a comparative evaluation between the benchmark described in
Sect. 3, and di erent classi ers based on neural networks, which adopt the same set of features.</p>
        <p>The rst neural network involved in the comparison, named shallow classi er, is obtained
by merging the two hidden layers of the benchmark network into a single double-sized layer.
The shallow network di ers from the benchmark only for its topology, since it has almost
the same numbers of parameters. The comparison between the benchmark and the shallow
networks allows to evaluate the e ect of the hierarchical strati cation on the performances of
the classi er. The second network is obtained by removing the second hidden layer from the
benchmark network, by replacing the PReLUs with sigmoids, and by removing the dropout
step, thus obtaining a classic neural network. Then, by varying the number of nodes contained
in the remaining hidden layer, we obtained three neural networks with 1024, 512 and 8 nodes in
the hidden layer. Such networks have less parameters than the benchmark, and their evaluation
allows to analyze how much deep learning techniques can improve the classi cation accuracy,
when high-level features are adopted. Table 1 summarizes the topology and the number of
parameters for each network.</p>
        <p>Results are summarized in Table 2. Surprisingly, even if the benchmark network allows to
achieve the best results for almost every metric, the results obtained by the other classi ers
are quite similar, even those obtained by the classic network with only 8 nodes on the hidden
layer. Moreover, by analyzing the trend of the average accuracy with respect to the training
epochs (Fig. 3a) and the ROC curves (Fig. 3b), no relevant di erence among the considered
neural networks arises. Such results con rm that, if a set of well-designed high-level features
is adopted to represent malware les, deep neural networks do not signi cantly impact on the
performances.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Deep Learning performance with low-level features</title>
        <p>Several tests were performed to evaluate both the e ectiveness of the proposed approach, based
on the adoption of deep learning with low-level features, and the impact of the design choices
we made.</p>
        <p>Firstly, we wanted to investigate whether there is a connection between the number of
training epochs and performance. For this purpose we compared the performance obtained by
varying the maximum training epoch threshold (i.e., 600, 800, and 1000) for the pre-training
phase. The second evaluation aimed to assess the e ectiveness of the pre-training phase, by
comparing our system with a deep network characterized by the same topology, but trained only
through the second training phase, that is without pre-training. Finally, we intended to verify
the impact of the hierarchical strati cation of our deep classi er on the overall performances
of the malware detection system. To this aim, we compared our system with a classic neural
network obtained by removing the last two hidden layers from our classi er, so obtaining three
layer of 636, 256 and 1 nodes respectively.</p>
        <p>In Table 3 the average values of the considered performance metrics are shown. We can
observe that higher accuracy values are obtained by using the pre-trained classi ers, and this
con rms the e ectiveness of the two-phase training. However, increasing the number of
pretraining epochs (e.g., from 600 to 800 or 1000) does not signi cantly improve the accuracy, nor
reduce the values of loss.</p>
        <p>Classi er
Deep model (600)
Deep model (800)
Deep model (1000)
Deep model ( ne tuning only)
Classic-623-256</p>
        <p>Loss</p>
        <p>The performances of the model that uses ne tuning only, i.e., trained only with the second
training phase, are worse than the pre-trained models on all metrics. Furthermore, the lowest
values are achieved by using the classic model, suggesting that the hierarchical strati cation is
e ective, but only when combined with the pre-training phase.</p>
        <p>A more in-depth analysis of the average accuracy achieved using di erent models is shown in
Fig. 4a, which uses results from the second training phase. Results con rm that the pre-trained
deep models exhibit a better trend along all the time line. Fig. 4b shows the average ROC
curves from which AUC score was calculated. Even in this case, pre-trained models outperform
the others, with no signi cative di erences when using 800 or 1000 pre-training epochs.</p>
        <p>When adopting low-level features, our deep model produces an increase of more than 2% of
precision and of 1% for accuracy with respect to a traditional neural network. On the contrary,
the adoption of a deep network in the benchmark system increase the accuracy and precision
values of only 0:06% and 0:11% respectively.</p>
        <p>It is worth noticing that, whatever is the classi er, the adoption of optimal high-level features
yields to better classi cation performances. Thus, even though using a smaller and simpler set
of features makes the malware detection process easier, the proposed deep learning solution is
not able to achieve the same level of accuracy.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions and Future Work</title>
      <p>In this work we addressed the scenario of malware detection through deep learning approaches.
Firstly, we wondered if deep learning can makes a relevant di erence when combined with
welldesigned high-level features. Experiments proved that, when using an optimal set of features,
the adoption of a deep neural network does not signi cantly in uence the performance.</p>
      <p>Since deep learning is often adopted to achieve a high classi cation accuracy even by
processing raw data, we also investigated whether the use of low-level features (computed only on
the rst 1 or 2 KB of a le) combined with a deep learning approach may lead to comparable
performances. We proposed a deep neural network whose parameters are learned through a
twophase training which combines a unsupervised technique, i.e., stacked denoising autoencoders,
with a supervised step to perform ne tuning.</p>
      <p>Results showed the capability of our system of extracting relevant knowledge from low-level
features obtained through a static analysis of the considered les. Moreover, the two-phase
training was proved to be very e ective regardless of the number of pre-training epochs.</p>
      <p>
        However, regardless of the classi er, the use of optimal high-level features leads to the
best classi cation performances. Nevertheless, it is worth noticing that with a total of 180,577
parameters, the 9% of those used by the benchmark system [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], our system exhibits a reduction
of only 0:57% and 1:07% in precision and accuracy respectively.
      </p>
      <p>Such results con rm the potentiality of our approach, and also highlight the necessity of a
further update of the method we propose in order to overcome the limitation introduced by
adopting a smaller and simpler set of features.</p>
      <p>
        As future work, we want to verify if other classi er can be adopted for the rst, unsupervised,
training phase. For example a noisy-based autoencoder, like the DA-IC (Denoising Autoencoder
with Interdependent Codes ) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], or a deterministic one, like the contractive autoencoder [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] could
lead to better performances.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Abou-Assaleh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Cercone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Keselj</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Sweidan</surname>
          </string-name>
          .
          <article-title>N-gram-based detection of new malicious code</article-title>
          .
          <source>In Proc. of the 28th Annual Computer Software and Applications Conf</source>
          .
          <source>(COMPSAC</source>
          <year>2004</year>
          ), volume
          <volume>2</volume>
          , pages
          <fpage>41</fpage>
          {
          <fpage>42</fpage>
          . IEEE,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ahmadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ulyanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Semenov</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Tro mov, and</article-title>
          <string-name>
            <given-names>G.</given-names>
            <surname>Giacinto</surname>
          </string-name>
          .
          <article-title>Novel feature extraction, selection and fusion for e ective malware family classi cation</article-title>
          .
          <source>In Proc. of the Sixth ACM Conf. on Data and Application Security and Privacy</source>
          , pages
          <volume>183</volume>
          {
          <fpage>194</fpage>
          . ACM,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Alazab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Venkatraman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Watters</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Alazab</surname>
          </string-name>
          .
          <article-title>Zero-day malware detection based on supervised learning algorithms of api call signatures</article-title>
          .
          <source>In Proc. of the Ninth Australasian Data Mining Conf</source>
          . - Volume
          <volume>121</volume>
          , AusDM '
          <volume>11</volume>
          , pages
          <fpage>171</fpage>
          {
          <fpage>182</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G. E.</given-names>
            <surname>Dahl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Stokes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Deng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Yu</surname>
          </string-name>
          .
          <article-title>Large-scale malware classi cation using random projections and neural networks</article-title>
          .
          <source>In Proc. of the 2013 IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , pages
          <fpage>3422</fpage>
          {
          <fpage>3426</fpage>
          . IEEE,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>O. E. David and N. S.</given-names>
            <surname>Netanyahu</surname>
          </string-name>
          . Deepsign:
          <article-title>Deep learning for automatic malware signature generation and classi cation</article-title>
          .
          <source>In Proc. of the 2015 Int. Joint Conf. on Neural Networks (IJCNN)</source>
          , pages
          <fpage>1</fpage>
          <article-title>{8</article-title>
          . IEEE,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <article-title>Delving deep into recti ers: Surpassing human-level performance on imagenet classi cation</article-title>
          .
          <source>In Proc. of the IEEE Int. Conf. on computer vision</source>
          , pages
          <volume>1026</volume>
          {
          <fpage>1034</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kingma</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          .
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>In Proc. of the 3rd Int. Conf. on Learning Representations (ICLR</source>
          <year>2015</year>
          ), pages
          <fpage>1</fpage>
          {
          <fpage>15</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kolosnjaji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zarras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. D.</given-names>
            <surname>Webster</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Eckert</surname>
          </string-name>
          .
          <article-title>Deep learning for classi cation of malware system call sequences</article-title>
          .
          <source>In Proc. of the Australasian Conf. on Arti cial Intelligence</source>
          , volume
          <volume>9992</volume>
          of Lecture Notes in Computer Science, pages
          <volume>137</volume>
          {
          <fpage>149</fpage>
          . Springer,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Larochelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Erhan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Vincent</surname>
          </string-name>
          .
          <article-title>Deep learning using robust interdependent codes</article-title>
          .
          <source>In AISTATS</source>
          , pages
          <volume>312</volume>
          {
          <fpage>319</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>K.</given-names>
            <surname>Mathur</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Hiranwal</surname>
          </string-name>
          .
          <article-title>A survey on techniques in detection and analyzing malware executables</article-title>
          .
          <source>Int. J. of Advanced Research in Computer Science and Software Engineering</source>
          ,
          <volume>3</volume>
          (
          <issue>4</issue>
          ),
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>McLaughlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Martinez del Rincon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yerima</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sezer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Safaei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Trickel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Doupe</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. Joon</given-names>
            <surname>Ahn</surname>
          </string-name>
          .
          <article-title>Deep android malware detection</article-title>
          .
          <source>In Proc. of the Seventh ACM on Conf. on Data and Application Security and Privacy</source>
          ,
          <source>CODASPY '17</source>
          , pages
          <fpage>301</fpage>
          {
          <fpage>308</fpage>
          . ACM,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S. B.</given-names>
            <surname>Mehdi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Tanwani</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Farooq</surname>
          </string-name>
          .
          <article-title>Imad: in-execution malware analysis and detection</article-title>
          .
          <source>In Proc. of the 11th Annual Conf. on Genetic and evolutionary computation</source>
          , pages
          <volume>1553</volume>
          {
          <fpage>1560</fpage>
          . ACM,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Roberto</surname>
            <given-names>Perdisci</given-names>
          </string-name>
          , Andrea Lanzi, and
          <string-name>
            <given-names>Wenke</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <article-title>Classi cation of packed executables for accurate computer virus detection</article-title>
          .
          <source>Pattern Recognition Letters</source>
          ,
          <volume>29</volume>
          (
          <issue>14</issue>
          ):
          <year>1941</year>
          {
          <year>1946</year>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rifai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vincent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Glorot</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <article-title>Contractive auto-encoders: Explicit invariance during feature extraction</article-title>
          .
          <source>In Proc. of the 28th Int. Conf. on machine learning (ICML11)</source>
          , pages
          <fpage>833</fpage>
          {
          <fpage>840</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>I.</given-names>
            <surname>Santos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Brezo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sanz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Laorden</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. G.</given-names>
            <surname>Bringas</surname>
          </string-name>
          .
          <article-title>Using opcode sequences in single-class learning to detect unknown malware</article-title>
          .
          <source>IET information security</source>
          ,
          <volume>5</volume>
          (
          <issue>4</issue>
          ):
          <volume>220</volume>
          {
          <fpage>227</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>J.</given-names>
            <surname>Saxe</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Berlin</surname>
          </string-name>
          .
          <article-title>Deep neural network based malware detection using two dimensional binary program features</article-title>
          .
          <source>In Proc. of the 2015 10th Int. Conf. on Malicious and Unwanted Software (MALWARE)</source>
          , pages
          <fpage>11</fpage>
          {
          <fpage>20</fpage>
          . IEEE,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Schultz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Eskin</surname>
          </string-name>
          , E. Zadok, and
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Stolfo</surname>
          </string-name>
          .
          <article-title>Data mining methods for detection of new malicious executables</article-title>
          .
          <source>In IEEE Symposium on Security and Privacy</source>
          , pages
          <volume>38</volume>
          {
          <fpage>49</fpage>
          . IEEE Computer Society,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>N.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. E.</given-names>
            <surname>Hinton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          .
          <article-title>Dropout: a simple way to prevent neural networks from over tting</article-title>
          .
          <source>J. of Machine Learning Research</source>
          ,
          <volume>15</volume>
          (
          <issue>1</issue>
          ):
          <year>1929</year>
          {
          <year>1958</year>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>P.</given-names>
            <surname>Vincent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Larochelle</surname>
          </string-name>
          , I. Lajoie,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.-A.</given-names>
            <surname>Manzagol</surname>
          </string-name>
          .
          <article-title>Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion</article-title>
          .
          <source>J. of Machine Learning Research</source>
          ,
          <volume>11</volume>
          (Dec):
          <volume>3371</volume>
          {
          <fpage>3408</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          , W.-D. Cai, and
          <string-name>
            <given-names>P..</given-names>
            <surname>Wei</surname>
          </string-name>
          .
          <article-title>A deep learning approach for detecting malicious javascript code</article-title>
          .
          <source>Security and Communication Networks</source>
          ,
          <volume>9</volume>
          (
          <issue>11</issue>
          ):
          <volume>1520</volume>
          {
          <fpage>1534</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>M.</given-names>
            <surname>Weber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schmid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schatz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Geyer</surname>
          </string-name>
          .
          <article-title>A toolkit for detecting and analyzing malicious software</article-title>
          .
          <source>In Proc. of the 18th Annual Computer Security Applications Conf</source>
          ., pages
          <volume>423</volume>
          {
          <fpage>431</fpage>
          . IEEE,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>L.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Jayasena</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Cavazos</surname>
          </string-name>
          . Hadm:
          <article-title>Hybrid analysis for detection of malware</article-title>
          .
          <source>In Proc. of the SAI Intelligent Systems Conf. (IntelliSys)</source>
          , pages
          <fpage>1037</fpage>
          {
          <fpage>1047</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Adjeroh</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Iyengar</surname>
          </string-name>
          .
          <article-title>A survey on malware detection using data mining techniques</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>50</volume>
          (
          <issue>3</issue>
          ):
          <fpage>41</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          , E. Tas,
          <string-name>
            <given-names>U.</given-names>
            <surname>Gupta</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Abdulhayoglu</surname>
          </string-name>
          .
          <article-title>Combining le content and le relations for cloud based malware detection</article-title>
          .
          <source>In Proc. of the 17th ACM SIGKDD Int. Conf. on Knowledge Discovery and Data Mining</source>
          , pages
          <volume>222</volume>
          {
          <fpage>230</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ye</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Q.</given-names>
            <surname>Jiang</surname>
          </string-name>
          .
          <article-title>An intelligent pe-malware detection system based on association mining</article-title>
          .
          <source>J. in computer virology</source>
          ,
          <volume>4</volume>
          (
          <issue>4</issue>
          ):
          <volume>323</volume>
          {
          <fpage>334</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>