<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Explaining Vulnerabilities of Deep Learning to Adversarial Malware Binaries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Luca Demetrio</string-name>
          <email>luca.demetrio@dibris.unige.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Battista Biggio</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Lagorio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Roli</string-name>
          <email>fabio.rolig@diee.unica.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Armando</string-name>
          <email>alessandro.armandog@unige.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CSecLab, University of Genova</institution>
          ,
          <addr-line>Genova</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>PRALab, Department of Electrical and Electronic Engineering, University of Cagliari</institution>
          ,
          <addr-line>Cagliari</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Pluribus One</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Recent work has shown that deep-learning algorithms for malware detection are also susceptible to adversarial examples, i.e., carefullycrafted perturbations to input malware that enable misleading classi cation. Although this has questioned their suitability for this task, it is not yet clear why such algorithms are easily fooled also in this particular application domain. In this work, we take a rst step to tackle this issue by leveraging explainable machine-learning algorithms developed to interpret the black-box decisions of deep neural networks. In particular, we use an explainable technique known as feature attribution to identify the most in uential input features contributing to each decision, and adapt it to provide meaningful explanations to the classi cation of malware binaries. In this case, we nd that a recently-proposed convolutional neural network does not learn any meaningful characteristic for malware detection from the data and text sections of executable les, but rather tends to learn to discriminate between benign and malware samples based on the characteristics found in the le header. Based on this nding, we propose a novel attack algorithm that generates adversarial malware binaries by only changing few tens of bytes in the le header. With respect to the other state-of-the-art attack algorithms, our attack does not require injecting any padding bytes at the end of the le, and it is much more e cient, as it requires manipulating much fewer bytes.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Despite their impressive performance in many di erent tasks, deep-learning
algorithms have been shown to be easily fooled by adversarial examples, i.e.,
carefully-crafted perturbations of the input data that cause misclassi cations [1{
6]. The application of deep learning to the cybersecurity domain does not
constitute an exception to that. Recent classi ers proposed for malware detection,
including the case of PDF, Android and malware binaries, have been indeed
shown to be easily fooled by well-crafted adversarial manipulations [1, 7{11].
Despite the sheer number of new malware specimen unleashed on the Internet
(more than 8 millions in 2017 according to GData4) demands for the application
of e ective automated techniques, the problem of adversarial examples has
signi cantly questioned the suitability of deep-learning algorithms for these tasks.
Nevertheless, it is not yet clear why such algorithms are easily fooled also in the
particular application domain of malware detection.</p>
      <p>
        In this work, we take a rst step towards understanding the behavior of
deeplearning algorithms for malware detection. To this end, we argue that
explainable machine-learning algorithms, originally developed to interpret the black-box
decisions of deep neural networks [12{15], can help unveiling the main
characteristics learned by such algorithms to discriminate between benign and malicious
les. In particular, we rely upon an explainable technique known as feature
attribution [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] to identify the most in uential input features contributing to each
decision. We focus on a case study related to the detection of Windows Portable
Executable (PE) malware les, using a recently-proposed convolutional neural
network named MalConv [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. This network is trained directly on the raw
input bytes to discriminate between malicious and benign PE les, reporting good
classi cation accuracy. Recently, concurrent work [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ] has shown that it can
be easily evaded by adding some carefully-optimized padding bytes at the end of
the le, i.e., by creating adversarial malware binaries. However, no explanation
has been clearly given behind the surprising vulnerability of this deep network.
To address this problem, in this paper we adopt the aforementioned
featureattribution technique to provide meaningful explanations of the classi cation of
malware binaries. Our underlying idea is to extend feature attribution to
aggregate information on the most relevant input features at a higher semantic level,
in order to highlight the most important instructions and sections present in
each classi ed le.
      </p>
      <p>Our empirical ndings show that MalConv learns to discriminate between
benign and malware samples mostly based on the characteristics of the le header,
i.e., almost ignoring the data and text sections, which instead are the ones where
the malicious content is typically hidden. This means that also depending on the
training data, MalConv may learn a spurious correlation between the class labels
and the way the le headers are formed for malware and benign les.</p>
      <p>
        To further demonstrate the risks associated to using deep learning \as is"
for malware classi cation, we propose a novel attack algorithm that generates
adversarial malware binaries by only changing few tens of bytes in the le header.
With respect to the other state-of-the-art attack algorithms in [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ], our attack
does not require injecting any padding bytes at the end of the le (it modi es
the value of some existing bytes in the header), and it is much more e cient, as
it requires manipulating much fewer bytes.
      </p>
      <p>The structure of this paper is the following: in Section 2 we introduce how
to solve the problem of malware detection using machine learning techniques,
and we present the architecture of MalConv; (ii) we introduce the integrated
gradient technique as an algorithm that extracts which features contributes most
! ∈ ℝ7×9 Convolutional</p>
      <p>layer
,...
,.
,.
*...
,/</p>
      <p>Embedding
! = #(%)
Padding bytes
(set to 256)
'(
...
')
')*(
...
'+</p>
      <p>Convolutional
layer
σ</p>
      <p>Temporal
max pooling</p>
      <p>Fully
connected
for the classi cation problem; (iii) we collect the result of the method mentioned
above applied on MalConv, highlighting its weaknesses, and (iv) we show how
an adversary may exploit this information to craft an attack against MalConv.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Deep Learning for Malware Detection in Binary Files</title>
      <p>
        Ra et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] propose MalConv, a deep learning model which discriminates
programs based on their byte representation, without extracting any feature.
The intuition of this approach is based on spatial properties of binary programs:
(i) code, and data may be mixed, and it is di cult to extract proper features; (ii)
there is correlation between di erent portions of the input program; (iii) binaries
may have di erent length, as they are strings of bytes. Moreover, there is no clear
metrics that can be imposed on bytes: each value represents an instruction or a
piece of data, hence it is di cult to set a distance between them.
      </p>
      <p>To tackle all of these issues, Ra et al. develop a deep neural network that
takes as input a whole program, whose architecture is represented in Fig. 1. First,
the input le size is bounded to d = 221 bytes, i.e., 2 MB. Accordingly, the input
le x is padded with the value 256 if its size is smaller than 2 MB; otherwise,
it is cropped and only the rst 2 MB are analyzed. Notice that the padding
value does not correspond to any valid byte, to properly represent the absence
of information. The rst layer of the network is an embedding layer, de ned
by a function : f0; 1; : : : ; 256g ! Rd 8 that maps each input byte to an
8dimensional vector in the embedded space. This representation is learned during
the training process and allows considering bytes as unordered categorical values.
The embedded sample Z is passed through two one-dimensional convolutional
layers and then shrunk by a temporal max-pooling layer, before being passed
through the nal fully-connected layer. The output f (x) is given by a softmax
function applied to the results of the dense layer, and classi es the input x as
malware if f (x) 0:5 (and as benign otherwise).</p>
      <p>We debate the robustness of this approach, as it is not clear the rationale
behind what has been learned by MalConv. Moreover, we observe that the deep
neural network focuses on the wrong sequences of bytes of an input binary.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Explaining Machine Learning</title>
      <p>
        We address the problem mentioned at the end of Section 2 by introducing
techniques aimed to explain the decisions of machine-learning models. In particular,
we will focus on techniques that explain the local predictions of such models,
by highlighting the most in uential features that contribute to each decision. In
this respect, linear models are easy to interpret: they rank each feature with a
weight proportional to their relevance. Accordingly, it is trivial to understand
which feature caused the classi er to assign a particular label to an input
sample. Deep-learning algorithms provide highly non-linear decision functions, where
each feature may be correlated at some point with many others, making the
interpretation of the result nearly impossible and leading the developers to naively
trust their output. Explaining predictions of deep-learning algorithms is still an
open issue. Many researchers have proposed di erent techniques to explain what
a model learns and which features mostly contribute to its decisions. Among the
proposed explanation methods for machine learning, we decided to focus on a
technique called integrated gradients [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], developed by Sundararajan et al., for
two main reasons. First, it does not use any learning algorithm to explain the
result of another machine learning model; and, second, it is more e cient w.r.t.
the computations that are required by the other methods in the state-of-the-art.
Integrated gradients. We introduce the concept of attribution methods :
algorithms that compute the contribution of each feature for deciding which label
needs to be assigned to a single point. Contributions are calculated w.r.t. a
baseline. A baseline is a point in the input space that corresponds to a null
signal: for most image classi ers it may be identi ed as a black image, it is an
empty sentence for text recognition algorithms, and so on. The baseline serves
as ground truth for the model: each perturbation to the baseline should increase
the contributions computed for the modi ed features. Hence, each contribution
is computed w.r.t. the output of the model on the baseline. The integrated
gradients technique is based on two axioms and upon the concept of baseline.
Axiom I: Sensitivity.
      </p>
      <p>The rst axiom is called sensitivity : an attribution method satis es sensitivity
if, for every input that di er in one feature from the baseline but they are
classi ed di erently, then the attribution of the di ering feature should be
nonzero.</p>
      <p>Moreover, if the learned function does not mathematically depend on a
particular feature, the attribution should be zero. On the contrary, if sensitivity
is not satis ed, the model is focusing on irrelevant features, as the attribution
method fails to weight the contribution of each variable. Authors state that
gradients violate the sensitivity axioms, hence using them during the training
phase by applying back-propagation implies attributing wrong importance to
the wrong features.</p>
      <p>Axiom II: Implementation Invariance. The second axiom is called
implementation invariance, and it is built on top of the notion of functional
equivalence: two networks are functionally equivalent if their outputs are equal on
all inputs, despite being implemented in di erent ways. Thus, an attribution
method satis es implementation invariance if it produces the same attributions
for two functionally equivalent networks. On the contrary, if this axiom is not
satis ed, the attribution method is sensitive to the presence of useless aspects
of the model. On top of these two axioms, Sundararajan et al. propose the
integrated gradient method that satis es both sensitivity and implementation
invariance. Hence, this algorithm should highlight the properties of the input
model and successfully attributing the correct weights to the feature believed
relevant by the model itself.</p>
      <p>Integrated Gradients. Given the input model f , a point x and baseline x0,
the attribution for the ith feature is computed as follows:</p>
      <p>IGi(x) = (xi
x0 )
i</p>
      <p>x0))
0
This is the integral of the gradient computed on all points that lie on the line
passing through x and x0. If x0 is a good baseline, each point in the line should
add a small contribution to the classi cation output. This method satis es also
the completeness axiom: the attributions add up to the di erence between the
output of the model at the input x and x0. Hence, the features that are important
for the classi cation of x should appear by moving on that line.</p>
      <p>Since we can only compute discrete quantities, the integral can be
approximated using a summation, adding a new degree of freedom to the algorithm,
that is the number of points to use in the process: Sundararajan et al. state that
the number of steps could be chosen between 20 and 300, as they are enough to
approximate the integral within the 5% of accuracy.
4</p>
    </sec>
    <sec id="sec-4">
      <title>What Does MalConv Learn?</title>
      <p>
        We applied the integrated gradient technique for trying to grasp the intuition of
what is going on under the hood of MalConv deep network. For our experiments
we used a simpli ed version of MalConv, with an input dimension shrunk to 220
instead of 221, that is 1 MB instead of 2 MB, trained by Anderson et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]
and publicly available on GitHub 5. To properly comment the result generated
by the attribution method, we need to introduce the layout of the executables
that run inside Windows operating system.
      </p>
      <p>Windows Portable Executable format. The Windows Portable Executable
format6 (PE) describes the structure of a Windows executable. Each program
begins with the DOS header, which is left for retro-compatibility issues and for
telling the user that the program can not be run in DOS. The only two useful
information contained into the DOS header are the DOS magic number MZ and
the value contained at o set 0x3c, that is an o set value that point to the
real PE header. If the rst one is modi ed, Windows throws an error and the</p>
      <sec id="sec-4-1">
        <title>5 https://github.com/endgameinc/ember/tree/master/malconv 6 https://docs.microsoft.com/en-us/windows/desktop/debug/pe-format</title>
        <p>program is not loaded, while if the second is perturbed, the operating system
can not nd the metadata of the program. The PE header consists of a Common
Object File Format (COFF) header, containing metadata that describe which
Windows Subsystem can run that le, and more. If the le is not a library, but
a standalone executable, it is provided with an Optional Header, which contains
information that are used by the loader for running the executable. As part of
the optional header, there are the section headers. The latter is composed by
one entry for each section in the program. Each entry contains meta-data used
by the operating system that describes how that particular section needs to be
loaded in memory. There are other components speci ed by the Windows PE
format, but for this work this information is enough to understand the output
of the integrated gradients algorithm applied to MalConv.</p>
        <p>Integrated gradients applied to malware programs. The integrated
gradients technique works well with image and text les, but we need to evaluate
it on binary programs. First of all, we have to set a baseline: since the choice of
the latter is crucial for the method to return accurate results, we need to pick a
sample that satis es the constraints7 highlighted by the authors of the method:
(i) the baseline should be an empty signal with null response from the network;
(ii) the entropy of the baseline should be very low. If not, that point could be
an adversarial point for the network and not suitable for this method.</p>
        <p>Hence, we have two possible choices: (i) the empty le, and (ii) a le lled
with zero bytes. For MalConv, an empty le is a vector lled by the special
padding number 256, as already mentioned in Section 2. While both these
baselines satisfy the second constraint, only the empty le has a null response, as
the zero vector is seen as malware with the 20% of con dence. Hence, the empty
le is more suitable for being considered the ground truth for the network. The
method returns a matrix that contains all the attributions, feature per feature,
V 2 Rd 8, where the second dimension is produced by the embedding layer. We
compute the mean for each row of V to obtain a signed scalar number for each
feature, resulting in a point v0 2 Rd and it may be visualized in a plot. Figure 2
highlights the importance attributed by MalConv to the header of an input
malware example using the empty le as baseline, marking with colours each byte's
contribution. We can safely state that a sample is malware if it is scored as such
by VirusTotal8 which is an online service for scanning suspicious les. The cells
colored in red symbolize that MalConv considered them for deciding whether
the input sample is malware. On the contrary, cells with blue background are
the ones that MalConv retains representative of benign programs. Regarding our
analysis, we may notice that MalConv is giving importance to portions of the
input program that are not related to any malicious intent: some bytes in the
DOS header are colored both in red and blue, and a similar situation is found
for the other sections. Notably, the DOS header is not even read by modern
operating systems, as the only important values are the magic number and the
value contained at o set 0x3c, as said in the previous paragraph. All the other
7 https://github.com/ankurtaly/Integrated-Gradients/blob/master/howto.md
8 https://www.virustotal.com
bytes are repeated in every Windows executable, and they are not peculiar
neither for goodware nor malware programs. We can thus say that this portion
of the program should not be relevant for discriminating benign from malicious
programs. Accordingly, one would expected MalConv to give no attribution to
these portions of bytes.</p>
        <p>
          We can observe the attribution assigned by integrated gradients to every
byte of the program, aggregated for each component of the binary, in Figure 3.
Each entry of the histogram is the sum of all the contributions given by each
byte in that region. The histogram is normalized using the l1 norm of its
components. The color scheme is the same as the one described for Figure 2. MalConv
puts higher weights at the beginning of the le, and this fact has been already
formalized by Kolosnjaji et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] in the discussion of the results of the attack
they have developed. It is clear that the .text section has a large impact on the
classi cation, as it is likely that the maliciousness of the input program lies in
that portion of the malware. However, the contribution given by the COFF and
optional headers outmatch the other sections. This further implies that the
locations that are learned as important for classi cation are somewhat misleading.
We can state that among all the correlations that are present inside a program,
MalConv surely learned something that is not properly relevant for classi cation
of malware and benign programs. By knowing this fact, an adversary may
perturb these portions of malware and sneaking behind MalConv without particular
e ort: all she or he has to do is to compute the gradient of the model w.r.t. to
the input. Hence, we describe how a malicious user can deliver an attack by
perturbing bytes contained into an input binary to evade the classi er.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Evading Malconv by Manipulating the File Header</title>
      <p>
        We showed that MalConv bases its classi cations upon unreliable and spurious
features, as also indirectly shown by state-of-the-art attacks against the same
network, which inject bytes in speci c regions of the malware sample without
even altering the embedded malicious code [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ]. In fact, these attacks can not
alter the semantics of the malware sample, as it needs to evade the classi er
while preserving its malicious functionality. Hence, not all the bytes in the input
binary can be modi ed, as it is easy to alter a single instruction and break the
behavior of the original program. As already mentioned in Section 1, the only
bytes that are allowed to be perturbed are those placed in unused zones of the
program, such as the space that is left by the compiler between sections, or at the
end of the sample. Even though these attacks have been proven to be e ective,
we believe that it is not necessary to deploy such techniques as the network
under analysis is not learning what the developers could have guessed during
training and test phases. Recalling the results shown in Section 4, we believe
that, by perturbing only the bytes inside the DOS header, malware programs
can evade MalConv. There are two constraints: the MZ magic number and the
value at o set 0x3c can not be modi ed, as said in Section 4. Thus, we ag as
editable all the bytes inside the DOS header that are contained in that range.
Of course, one may also manipulate more bytes, but we just want to highlight
the severe weaknesses anticipated in Section 3.
      </p>
      <p>
        Attack Algorithm. Our attack algorithm is given as Algorithm 1. It rst
computes the representation of each byte of the input le x in the embedded space,
as Z (x). The gradient G of the classi cation function f (x) is then
computed w.r.t. the embedding layer. Note that we denote with gi; zi 2 R8 the ith
row of the aforementioned matrices. For each byte indexed in I, i.e., in the
subset of header bytes that can be manipulated, our algorithm follows the strategy
implemented by Kolosnjaji et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This amounts to selecting the closest byte
in the embedded space which is expected to maximally increase the probability
of evasion. The algorithm stops either if evasion is achieved or if it reaches a
maximum number of allowed iterations. The latter is imposed since, depending
on which bytes are allowed to be modi ed, evasion is not guaranteed.
      </p>
      <p>The result of the attack applied using an example can be appreciated in
Figure 4: the plot shows the probability of being recognized as malware, which
is the curve coloured in blue. Each iteration of the algorithm is visualized as black
vertical dashed lines, showing that we need to interpret the gradient di erent
times to produce an adversarial example. This is only one example of the success
of the attack formalized above, as we tested it on 60 di erent malware inputs,
taken from both The Zoo9 and Das Malverik.10 Experimental results con rm
our doubts: 52 malware programs over 60 evade MalConv by only perturbing
the DOS header. Note that a human expert would not be fooled by a similar
attack, as the DOS header can be easily stripped and replaced with a normal
one without perturbed bytes. Accordingly, MalConv should have learned that
these bytes are not relevant to the classi cation task.</p>
      <p>Algorithm 1: Evasion algorithm
input : x input binary, I indexes
of bytes to perturb, f
classi er, T max iterations.
output: Perturbed code that should</p>
      <p>achieve evasion against f .
t 0;
while f (x) 0:5 ^ t &lt; T do</p>
      <p>Z (x) 2 Rd 8;
G r f (x) 2 Rd 8;
forall i 2 I do
gi gi=kgik2;
s 0 2 R256;
d 0 2 R256;
forall b 2 0; :::; 255 do
zb (b);
gi&gt; (zb zi);
kzb (zi + sb gi)k;
arg minj:sj&gt;0 dj ;
sb
db
end
xi
t + 1;
end
t
end</p>
      <sec id="sec-5-1">
        <title>9 https://github.com/ytisf/theZoo 10 http://dasmalwerk.eu/</title>
        <p>Fig. 4: Evading MalConv by
perturbing only few bytes in the DOS
header. The black dashed lines
represent the iterations of the
algorithm. In each iteration, the
algorithm manipulates 58 header bytes
(excluding those that can not be
changed). The blue curve shows the
probability of the perturbed sample
being classi ed as malware across
the di erent iterations.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Related Work</title>
      <p>
        We discuss here some approaches that leverage deep learning for malware
detection, which we believe may exhibit similar problems to those we have highlighted
for MalConv in this work. We continue with a brief overview of the already
proposed attacks that address vulnerabilities in the MalCon architecture. Then, we
discuss other explainable machine-learning techniques that may help us to gain
further insights on what such classi ers e ectively learn, and potentially whether
and to which extent we can make them more reliable, transparent and secure.
Deep Malware Classi ers. Saxe and Berlin [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] use a deep learning system
trained on top of features extracted from Windows executables, such as byte
entropy histograms, imported functions, and meta-data harvested from the header
of the input executable. Hardy et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] propose DL4MD, a deep neural network
for recognizing Windows malware programs. Each sample is characterized by its
API calls, hence they are passed to a stacked auto-encoders network. David et
al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] use a similar approach, by exploiting de-noising auto-encoders for
classifying Windows malware programs. The features used by the author are harvested
from dynamic analysis. Yhuan et al. [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] extract features from both static and
dynamic analysis of an Android application, like the requested permissions, API
calls, network usage, sent SMS and so on, and samples are given in input to
a deep neural network. McLaughlin et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] propose a convolutional neural
network which takes as input the sequence of the opcodes of an input program.
Hence, they do not extract any features from data, but they merely extract the
sequence of instructions to feed to the model.
      </p>
      <p>
        Evasion attacks against MalConv. Kolosnjaji et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] propose an attack
that does not alter bytes that correspond to code, but they append bytes at the
end of the sample, preserving its semantics. The padding bytes are calculated
using the gradient of the cost function w.r.t. the input byte in a speci c location
in the malware, achieving evasion with high probability. The technique is similar
to the one developed by Goodfellow et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], the so-called Fast Sign Gradient
Method (FSGM). Similarly, Kreuk et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] propose to append bytes at the
end of the malware and ll unused bytes between sections. Their algorithm is
di erent: they work in the embedding domain, and they translate back in byte
domain after applying the FSGM to the malware, while Kolosnjaji et al. modify
one padding byte at the time in the embedded domain, and then they translate
it back to the byte domain.
      </p>
      <p>
        Explainable Machine Learning. Riberio et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] propose a method called
LIME, which is an algorithm that tries to explain which features are important
w.r.t. the classi cation result. It takes as input a model, and it learns how it
behaves locally around an input point, by producing a set of arti cial samples.
The algorithm uses LASSO to select the top K features, where K is a free
parameter. As it may be easily noticed, the dimension of the problem matters, and it
may become computational expensive on high dimensional data. Guo et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
propose LEMNA, which is an algorithm specialized on explaining deep learning
results for security applications. As said in Section 2, many malware detectors
exploit machine learning techniques in their pipeline. Hence, being able to
interpret the results of the classi cation may be helpful to analysts for identifying
vulnerabilities. Guo et al. use a Fused LASSO or a mixture Regression Model
to learn the decision boundary around the input point and compute the
attribution for K features, where K is a free parameter. Another important work
that may help us explaining predictions (and misclassi cations) of deep-learning
algorithms for malware detection is that by Koh and Liang [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], which allows
identifying the most relevant training prototypes supporting classi cation of a
given test sample. With these technique, it may be possible to associate or
compare a given test sample to some known training malware samples and their
corresponding families (or, to some relevant benign le).
7
      </p>
    </sec>
    <sec id="sec-7">
      <title>Conclusions and Future Work</title>
      <p>We have shown that, despite the success of deep learning in many areas, there
are still uncertainties regarding the precision of the output of these techniques.
In particular, we use integrated gradients to explain the results achieved by
MalConv, highlight the fact that the deep neural network attributes wrong
nonzero weights to well known useless features, that in this case are locations inside
a Windows binary.</p>
      <p>To further explain these weaknesses, we devise an algorithm inspired by
recent papers in adversarial machine learning, applied to the malware detection.
We show that perturbing few bytes is enough for evading MalConv with high
probability, as we applied this technique on a concrete set of samples taken from
the internet, achieving evasion on most of all the inputs.</p>
      <p>Aware of this situation, we are willing to keep investigating the weaknesses
of these deep learning malware classi ers, as they are becoming more and more
important in the cybersecurity world. In particular, we want to devise attacks
that can not be easily recognized by a human expert: the perturbations in the
DOS header are easy to detect, as mentioned in Section 5 and it may be patched
without substantial e ort. Instead, we are studying how to hide these modi
cations in an unrecognizable way, like lling bytes between functions: compilers
often leave space between one function and the other to align the function entry
point to multiple of 32 or 64, depending on the underlying architecture. The
latter is an example we are currently investigating. Hence, all classi ers that
rely on raw byte features may be vulnerable to such attacks, and our research
may serve as a proposal for a preliminary penetration testing suite of attacks
that developers could use for establishing the security of these new technologies.</p>
      <p>In conclusion, we want to emphasize that these intelligent technologies are far
from being considered secure, as they can be easily attacked by aware adversaries.
We hope that our study may focus the attention on this problem and increase
the awareness of developers and researchers about the vulnerability of these
models.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Battista</given-names>
            <surname>Biggio</surname>
          </string-name>
          , Igino Corona, Davide Maiorca, Blaine Nelson, Nedim Srndic, Pavel Laskov, Giorgio Giacinto, and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Roli</surname>
          </string-name>
          .
          <article-title>Evasion attacks against machine learning at test time</article-title>
          .
          <source>In Joint European conference on machine learning and knowledge discovery in databases</source>
          , pages
          <volume>387</volume>
          {
          <fpage>402</fpage>
          . Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Christian</given-names>
            <surname>Szegedy</surname>
          </string-name>
          , Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and
          <string-name>
            <given-names>Rob</given-names>
            <surname>Fergus</surname>
          </string-name>
          .
          <article-title>Intriguing properties of neural networks</article-title>
          .
          <source>In International Conference on Learning Representations</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ian</surname>
            J Goodfellow, Jonathon Shlens, and
            <given-names>Christian</given-names>
          </string-name>
          <string-name>
            <surname>Szegedy</surname>
          </string-name>
          .
          <article-title>Explaining and harnessing adversarial examples (</article-title>
          <year>2014</year>
          ).
          <source>arXiv preprint arXiv:1412</source>
          .
          <fpage>6572</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Nicolas</given-names>
            <surname>Papernot</surname>
          </string-name>
          ,
          <string-name>
            <surname>Patrick</surname>
            <given-names>McDaniel</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Somesh</given-names>
            <surname>Jha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Matt</given-names>
            <surname>Fredrikson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z Berkay</given-names>
            <surname>Celik</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Ananthram</given-names>
            <surname>Swami</surname>
          </string-name>
          .
          <article-title>The limitations of deep learning in adversarial settings</article-title>
          .
          <source>In Security and Privacy (EuroS&amp;P)</source>
          ,
          <source>2016 IEEE European Symposium on</source>
          , pages
          <volume>372</volume>
          {
          <fpage>387</fpage>
          . IEEE,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Nicholas</given-names>
            <surname>Carlini</surname>
          </string-name>
          and
          <string-name>
            <given-names>David</given-names>
            <surname>Wagner</surname>
          </string-name>
          .
          <article-title>Towards evaluating the robustness of neural networks</article-title>
          .
          <source>In 2017 IEEE Symposium on Security and Privacy (SP)</source>
          , pages
          <fpage>39</fpage>
          {
          <fpage>57</fpage>
          . IEEE,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Battista</given-names>
            <surname>Biggio</surname>
          </string-name>
          and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Roli</surname>
          </string-name>
          .
          <article-title>Wild patterns: Ten years after the rise of adversarial machine learning</article-title>
          .
          <source>Pattern Recognition</source>
          ,
          <volume>84</volume>
          :
          <fpage>317</fpage>
          {
          <fpage>331</fpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Ambra</given-names>
            <surname>Demontis</surname>
          </string-name>
          , Marco Melis, Battista Biggio, Davide Maiorca, Daniel Arp, Konrad Rieck, Igino Corona, Giorgio Giacinto, and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Roli</surname>
          </string-name>
          .
          <article-title>Yes, machine learning can be more secure! a case study on android malware detection</article-title>
          .
          <source>IEEE Trans. Dependable and Secure Computing</source>
          , In press.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Kathrin</given-names>
            <surname>Grosse</surname>
          </string-name>
          , Nicolas Papernot, Praveen Manoharan,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Backes</surname>
          </string-name>
          , and
          <string-name>
            <surname>Patrick McDaniel</surname>
          </string-name>
          .
          <article-title>Adversarial perturbations against deep neural networks for malware classi cation</article-title>
          .
          <source>arXiv preprint arXiv:1606.04435</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Bojan</given-names>
            <surname>Kolosnjaji</surname>
          </string-name>
          , Ambra Demontis, Battista Biggio, Davide Maiorca, Giorgio Giacinto, Claudia Eckert, and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Roli</surname>
          </string-name>
          .
          <article-title>Adversarial malware binaries: Evading deep learning for malware detection in executables</article-title>
          . arXiv preprint arXiv:
          <year>1803</year>
          .04173,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Felix</surname>
            <given-names>Kreuk</given-names>
          </string-name>
          , Assi Barak, Shir Aviv-Reuven, Moran Baruch, Benny Pinkas, and
          <string-name>
            <given-names>Joseph</given-names>
            <surname>Keshet</surname>
          </string-name>
          .
          <article-title>Adversarial examples on discrete sequences for beating wholebinary malware detection</article-title>
          .
          <source>arXiv preprint arXiv:1802.04528</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>Nedim</given-names>
            <surname>Srndic</surname>
          </string-name>
          and
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Laskov</surname>
          </string-name>
          .
          <article-title>Practical evasion of a learning-based classi er: A case study</article-title>
          .
          <source>In Proc. 2014 IEEE Symp. Security and Privacy</source>
          ,
          <source>SP '14</source>
          , pages
          <fpage>197</fpage>
          {
          <fpage>211</fpage>
          , Washington, DC, USA,
          <year>2014</year>
          . IEEE CS.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Mukund</surname>
            <given-names>Sundararajan</given-names>
          </string-name>
          , Ankur Taly, and
          <string-name>
            <given-names>Qiqi</given-names>
            <surname>Yan</surname>
          </string-name>
          .
          <article-title>Axiomatic attribution for deep networks</article-title>
          .
          <source>arXiv preprint arXiv:1703.01365</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. Marco Tulio Ribeiro,
          <string-name>
            <given-names>Sameer</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Carlos</given-names>
            <surname>Guestrin</surname>
          </string-name>
          .
          <article-title>"why should I trust you?": Explaining the predictions of any classi er</article-title>
          .
          <source>In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          , San Francisco, CA, USA,
          <year>August</year>
          13-
          <issue>17</issue>
          ,
          <year>2016</year>
          , pages
          <fpage>1135</fpage>
          {
          <fpage>1144</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Wenbo</surname>
            <given-names>Guo</given-names>
          </string-name>
          , Dongliang Mu, Jun Xu, Purui Su,
          <string-name>
            <given-names>Gang</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Xinyu</given-names>
            <surname>Xing</surname>
          </string-name>
          . Lemna:
          <article-title>Explaining deep learning based security applications</article-title>
          .
          <source>In Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security</source>
          , pages
          <volume>364</volume>
          {
          <fpage>379</fpage>
          . ACM,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>P. W.</given-names>
            <surname>Koh</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          .
          <article-title>Understanding black-box predictions via in uence functions</article-title>
          .
          <source>In International Conference on Machine Learning (ICML)</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Edward</surname>
            <given-names>Ra</given-names>
          </string-name>
          , Jon Barker, Jared Sylvester, Robert Brandon, Bryan Catanzaro, and Charles Nicholas.
          <article-title>Malware detection by eating a whole exe</article-title>
          .
          <source>arXiv preprint arXiv:1710.09435</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>H. S. Anderson</surname>
            and
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Roth. EMBER</surname>
          </string-name>
          :
          <article-title>An Open Dataset for Training Static PE Malware Machine Learning Models</article-title>
          . ArXiv e-prints,
          <year>April 2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>Joshua</given-names>
            <surname>Saxe</surname>
          </string-name>
          and Konstantin Berlin.
          <article-title>Deep neural network based malware detection using two dimensional binary program features</article-title>
          .
          <source>In Malicious and Unwanted Software (MALWARE)</source>
          ,
          <year>2015</year>
          10th International Conference on, pages
          <volume>11</volume>
          {
          <fpage>20</fpage>
          . IEEE,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <given-names>William</given-names>
            <surname>Hardy</surname>
          </string-name>
          , Lingwei Chen, Shifu Hou, Yanfang Ye, and
          <string-name>
            <given-names>Xin</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>Dl4md: A deep learning framework for intelligent malware detection</article-title>
          .
          <source>In Proceedings of the International Conference on Data Mining (DMIN)</source>
          ,
          <source>page 61. The Steering Committee of The World Congress in Computer Science</source>
          , Computer Engineering and Applied Computing (WorldComp),
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Omid</surname>
            <given-names>E</given-names>
          </string-name>
          <string-name>
            <surname>David and Nathan S Netanyahu. Deepsign</surname>
          </string-name>
          :
          <article-title>Deep learning for automatic malware signature generation and classi cation</article-title>
          .
          <source>In Neural Networks (IJCNN)</source>
          , 2015 International Joint Conference on, pages
          <fpage>1</fpage>
          <article-title>{8</article-title>
          . IEEE,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Zhenlong</surname>
            <given-names>Yuan</given-names>
          </string-name>
          , Yongqiang Lu, and
          <string-name>
            <given-names>Yibo</given-names>
            <surname>Xue</surname>
          </string-name>
          .
          <article-title>Droiddetector: android malware characterization and detection using deep learning</article-title>
          .
          <source>Tsinghua Science and Technology</source>
          ,
          <volume>21</volume>
          (
          <issue>1</issue>
          ):
          <volume>114</volume>
          {
          <fpage>123</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Niall</surname>
            <given-names>McLaughlin</given-names>
          </string-name>
          ,
          <source>Jesus Martinez del Rincon</source>
          ,
          <source>BooJoong Kang</source>
          , Suleiman Yerima, Paul Miller, Sakir Sezer, Yeganeh Safaei, Erik Trickel,
          <string-name>
            <given-names>Ziming</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Adam</given-names>
            <surname>Doupe</surname>
          </string-name>
          , et al.
          <article-title>Deep android malware detection</article-title>
          .
          <source>In Proceedings of the Seventh ACM on Conference on Data and Application Security and Privacy</source>
          , pages
          <volume>301</volume>
          {
          <fpage>308</fpage>
          . ACM,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>