<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>June</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Sparse dictionaries for the explanation of classi cation systems?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>A. Apicella</string-name>
          <email>andrea.apicella@unina.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>F. Isgro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>R. Prevete</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>G. Tamburrini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Vietri</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dipartimento di Ingegneria Elettrica e delle Teconologie dell'Informazione Universita degli Studi di Napoli Federico II</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>4</volume>
      <issue>2019</issue>
      <abstract>
        <p>Providing algorithmic explanations for the decisions of machine learning systems to end users, data protection o cers, and other stakeholders in the design, production, commercialization and use of machine learning systems pipeline is an important and challenging research problem. Crucial motivations to address this research problem can be advanced on both ethical and legal grounds. Notably, explanations of the decisions of machine learning systems appear to be needed to protect the dignity, autonomy and legitimate interests of people who are subject to automatic decision-making. Much work in this area focuses on image classi cation, where the required explanations can be given in terms of images, therefore making explanations relatively easy to communicate to end users. In this paper we discuss how the representational power of sparse dictionaries can be used to identify local image properties as main ingredients for producing humanly understandable explanations for the decisions of a classi er developed on the basis of machine learning methods.</p>
      </abstract>
      <kwd-group>
        <kwd>XAI</kwd>
        <kwd>Explainable Arti cial Intelligence</kwd>
        <kwd>machine learning sparse coding</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Machine Learning (ML) techniques make possible to develop systems that learn
from observations. Many ML techniques (e.g., Support Vector Machines (SVM)
and Deep Neural Networks (DNN)) give rise to systems the behavior of which
is often hard to interpret [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. A crucial ML interpretability issue concerns the
generation of explanations for an ML system behavior that are understandable
to a human being. In general, this issue is addressed as a scienti c and
technological problem by so-called explainable arti cial intelligence (XAI). Providing
      </p>
      <p>XAI solutions to the ML explainability problem is important for many AI and
computer science research areas: to improve intelligent systems design, testing
and revision processes, to make the rationale of automatic decisions more
transparent to end users and systems managers, thereby leading to better forms of
HCI and HRI involving learning systems, to improve interactions between
learning agents in Distributed AI, and so on. Providing XAI solutions to the ML
explainability problem is quite important from ethical and legal viewpoints as
well. Learning systems are being increasingly used to make or to support
decisions that are most signi cant for the life of persons, including career, court,
medical diagnosis, insurance risk pro les and loan decisions. Thus, obtaining
explanations for classi cations and automatic decisions is arguably very important
on ethical grounds, in order to respect and protect the dignity, autonomy and
legitimate interests of people who are subject to automatic decision-making. On
more properly legal grounds, it is su cient to recall here that art. 22 of the
European Union GDPR establishes the right of a person to contest an automatic
decision and to address her complaint to the data protection o cer who is in
charge of the decision-making system. The data protection o cer would be in
a better position to evaluate these personal complaints, if he/she had an
understanding of the reasons, if any, underlying the contested automatic decision.
Moreover, in the case of repeated and undesired system behaviors, having good
explanations of learning systems decisions can be very helpful to identify the
sources of ethically unacceptable biases of learning systems, and to take those
corrective actions that are impelled by codes of professional ethics.</p>
      <p>
        Although some ML techniques come with reasonably interpretable
mechanisms and Input/Output (I/O) relationships (e.g., decision trees), this is not the
case for a wide variety of ML systems, whose processing and I/O relationships
are often di cult to understand [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. A ML system may have multiple sources
of opacity for human bounded rationality, notably including the large numbers
of features and ML parameters. As a consequence, the output of ML systems
may depend from inner data representations and processing which escape full
human understanding and interpretation. Indeed, in the ML system
representation space, small di erences or key features that cannot be easily made sense
of in the framework of human classi cation strategies may play a decisive role
for classi cation outcomes. Various senses of interpretability and
explainability for learning systems have been distinguished and analyzed [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and various
approaches to overcoming their opaqueness are now being pursued [
        <xref ref-type="bibr" rid="ref22 ref9">9, 22</xref>
        ]. For
example, in [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] a series of techniques for the interpretation of DNN are discussed,
and in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] a wide variety of motivations underlying interpretability needs are
examined, thereby re ning the notion of interpretability in ML systems. In the
context of this multifaceted interpretability problem [
        <xref ref-type="bibr" rid="ref27 ref28">27, 28</xref>
        ], we focus on the
issue of what it is to explain the behavior of ML perceptual classi cation systems
for which only I/O relationships are accessible, i.e., the learning system is seen
as a black-box. In literature, this type of approach is known as model agnostic
[
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].
      </p>
      <p>
        Various model agnostic approaches have been developed to give global
explanations exhibiting a class prototype which the input data can be associated
to [
        <xref ref-type="bibr" rid="ref19 ref22 ref27 ref9">9, 22, 27, 19</xref>
        ]. These explanations are given in response to explanation
requests that are usually expressed as why-questions: \Why were input data x
associated to class C?". Speci c why-questions which may arise in connection
with actual learning systems are : \Why was this loan application rejected?"
and \Why was this image classi ed as a fox?". However, prototypes often make
rather poor explanations available. For instance, if an image x is classi ed as
\fox", the explanation provided by means of a fox-prototype is nothing more
than a \because it looks like this" explanation: one would not be put in the
position to understand what features (parts) of the prototype are associated to
what characteristics (parts) of x. In order to go beyond this impoverished level
of understanding, instead of merely giving the user a global explanation, one
might attempt to provide a local explanation, which highlights salient parts of
the input [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]
      </p>
      <p>
        In this paper, we propose a model agnostic framework that returns local
explanations of classi cations based on dictionaries of local and humanly
interpretable elements of the input. This framework can be functionally described
in terms of a three-entity model, composed of an Oracle (an ML system, e.g. a
classi er), an Interrogator raising explanations requests about the Oracle's
responses, and a Mediator helping the Interrogator to understand the answer given
by the Oracle. The three-entity model is resumed in Figure 1. In this framework,
local explanations are provided by a system (the Mediator) which does not
coincide with the system which classi es inputs. A similar situation may occur in
the human brain where, for instance, the visual system provides classi cations
and recognition of objects present in a visual scene, but the reasons why a given
input is a recognised as a \cat" rather than, say, a '`dog", may involve other
areas of the brain, including those storing and processing semantic memories. In
this framework, the Mediator plays the crucial explanatory role, by advancing
hypotheses on what humanly interpretable elements are likely to have in uenced
the Oracle output. More speci cally, elements are computed which represent
humanly interpretable features of the input data, with the constraint that both
prototypes and input can be reconstructed as linear combinations of these
elements. Thus, one can establish meaningful associations between key features of
the prototype and key features of the input. To this end, we exploit the
representational power of sparse dictionaries learned from the data, where atoms of
the dictionary selectively play the role of humanly interpretable elements,
insofar as they a ord a local representation of the data. Indeed, these techniques
provide data representations that are often found to be accessible to human
interpretation [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. The dictionaries are obtained by a Non-negative Matrix
Factorization (NMF) method [
        <xref ref-type="bibr" rid="ref11 ref14 ref3">3, 14, 11</xref>
        ], and the explanation is determined using an
Activation-Maximization (AM) [
        <xref ref-type="bibr" rid="ref27 ref9">9, 27</xref>
        ] based technique, that we call Explanation
Maximization.
      </p>
      <p>The article is organized as follows: Section 2 brie y reviews related
approaches, in Section 3 we present the overall architecture; experiments and
results are discussed in Section 4, while Section 5 is devoted to concluding remarks
and future developments.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        In recent years, various attempts have been made to interpret and explain the
output of a classi cation system. Initial attempts concerned SVM classi ers (see
for example [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]) or rule-based systems [
        <xref ref-type="bibr" rid="ref5 ref6">6, 5</xref>
        ].
      </p>
      <p>
        In the neural network context, recent surveys on explainable AI are
proposed in [
        <xref ref-type="bibr" rid="ref1 ref10 ref24 ref33">33, 24, 10, 1</xref>
        ]. A signi cant attempt to explain in terms of images what
a computational neural unit computes is found in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] using the Activation
Maximization method. AM-like approaches applied to CNN were proposed in [
        <xref ref-type="bibr" rid="ref17 ref27">27,
17</xref>
        ]. Additional attempts to give interpretability to CNNs were proposed in [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]
and [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], where Deconvolutional Network (already presented by [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] as a way to
do unsupervised learning) and up-convolutional network are proposed, while [
        <xref ref-type="bibr" rid="ref21 ref22">22,
21</xref>
        ] uses an image generator network (similar to GANs) as priors for AM
algorithm to produce synthetic preferred images. In these approaches, explanations
are given in terms of prototypes or approximate input reconstructions. However,
one does not take into account the issue whether the given explanations are in
some manner interpretable by humans. Moreover, the proposed approaches seem
to be model-speci c for CNN, di erently from our model which is to be
considered as model-agnostic, and consequently applicable in principle to any classi er.
From another point of view, [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] studies the in uence on the output of hardly
perceptible perturbation on the input, empirically showing that it is possible to
arbitrarily change the network's prediction even when the input is left
apparently unchanged. Although this type of noise is extremely unlikely to occur in
realistic situations, the fact that such noise is imperceptible to an observer opens
interesting questions about the semantics of network components. However,
approaches of this kind are quite distant from our present concerns, insofar as they
focus on entities that are hardly meaningful to humans. Important works are
also made into [
        <xref ref-type="bibr" rid="ref2 ref20 ref4">2, 4, 20</xref>
        ] where Pixel-Wise Decomposition, Layer-Wise Relevance
propagation ad Deep Taylor Decomposition are presented. [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ] presents a work
based on prediction di erence analysis [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] where a features relevance vector
is built which estimates how much each feature is \important" for the classi er
to return the predicted class. In [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] , the model-agnostic explainer LIME is
proposed, which takes into account the model behavior in the proximity of the
instance being predicted. The LIME framework is more similar to our approach
than the other approaches mentioned in this section, and many other approaches
found in the literature. The LIME framework di ers from our own mainly in its
use of super-pixels instead of a learned dictionary constrained in order to have
a compact representation.
Given an oracle , an input x and an 's answer c^ (regardless of whether it is
correct or not), we want to give an explanation of the answer provided by the
model that is humanly interpretable.
      </p>
      <p>As we want to obtain humanly interpretable elements which, combined
together, can provide an acceptable explanation for the choice made by , we
search for an explanation having the following qualitative properties:
{ 1) the explanation must be expressed in terms of a dictionary V whose
elements (atoms) are easily understandable by an interrogator;
{ 2) the elements of the dictionary V have to represent \local properties" of
the input x;
{ 3) the explanation must be composed by few dictionary elements.</p>
      <p>We claim that considering as elements atoms of a sparse coding from a sparse
dictionary, and using sparse coding methods together with an AM-like algorithm
we obtain explanations satisfying the properties described above.
3.1</p>
      <sec id="sec-2-1">
        <title>Sparse Dictionary learning</title>
        <p>The rst step of the proposed approach consists in nding a \good" dictionary
V that can represent data in terms of humanly interpretable atoms.</p>
        <p>Let us assume that we have a set D = f(x(1); c(1)); (x(2); c(2)): : : : ; (x(n); c(n))g
where each x(i) 2 Rd is a column vector representing a data point, and c(i) 2 C
its class. We can learn a Dictionary V 2 Rd k of k atoms across multiple classes
and an encoding H 2 Rk n s.t. X = V H + where X = (x(1)jx(2)j : : : jx(n)) and
is the error introduced by the coding. Every column x(i) in X can be expressed
as x(i) = V hi with hi i th column of H. The dictionary forms the basis of our
explanation framework for an ML system.</p>
        <p>
          We selected as dictionary learning algorithm an NMF scheme [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] with the
additional sparseness constraint proposed by [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]; this choice is motivated by
the fact that it respects our requirements described above, giving a \local"
representation of data, and non-negativity, that ensures only additive operations in
data representations, giving a better human understanding with respect to other
techniques. The sparsity level can be set using two parameters 1 and 2 which
control the sparsity on the dictionary and the encoding, respectively. .
3.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Explanation Maximization</title>
        <p>
          Unlike traditional dictionary-based coding approaches, our main goal is not to
get an \accurate" representation of the input data, but to get a representation
that helps humans to understand the decision taken by a trained model. To this
aim, we modify the AM algorithm so that, instead of looking for the input that
just maximizes the answer of the model, it searches for the dictionary-based
encoding h that maximizes the answer and, at the same time, is sparse enough
but without being \too far" from the original input x. More formally, indicating
with Pr(c^jx) the probability given by a learned model that input x belongs to
class c^ 2 C, V the chosen dictionary, S( ) a sparsity measure, the objective
function that we optimise is
max log Pr c^jV h
h 0
1jjV h
xjj2 + 2S h
(1)
where 1; 2 are hyper-parameters regulating the input reconstruction and the
encoding sparsity level, respectively. The rst regularization term leads the
algorithm to choose dictionary atoms that, with an appropriate encoding, form a
good representation of the input, while the second regularization term ensures a
certain sparsity degree, i.e., that only few atoms are used. The h 0 constraint
ensures that one has a purely additive encoding. Thus, each hi; 8i:1 i d,
measures the \importance" of the i-th atom. Equation 1 is solved by using a
standard gradient ascent technique, together with a projection operator given
by [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] that ensures both sparsity and non-negativity. The complete procedure
is reported in Algorithm 1.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experimental Assessment</title>
      <p>
        To test our framework, we chose as Oracle a convolutional neural network
architecture, LeNet-5 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], generally used for digit recognition as MNIST. We have
trained the network from scratch using two di erent datasets: MNIST [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
obtaining an accuracy of 98:86% on the test set, and Fashion-MNIST [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], obtaining
Algorithm 1: Explanation Maximization procedure
5
6 end
7 return h ;
an accuracy of 91:43% on the test set. The training set is composed of 50000
images, while the test set is composed of 10000 images; the model is learned
using the Adam algorithm [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        NMF with sparseness constraints [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is used to determine the dictionaries.
We set the number of atoms to 200, relying on PCA analysis which showed that
the rst 100 principal components explain more than 95% of the data variance.
We construct di erent dictionaries with di erent sparsity values in the range
1; 2 2 [0:6; 0:8] [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], then we choose the dictionaries having the best trade-o
between sparsity level and reconstruction error. The dictionaries are determined
by looking for a a good trade-o between reconstruction error and sparsity level.
      </p>
      <p>The atoms forming our explanations are selected by taking those with larger
encoding values (i.e., those that are more \important" in the representation). In
gure 2 we show the atoms forming the explanation on three inputs for MNIST
dataset on which the Oracle gave the correct answer. The chosen atoms seem
to describe well the visual impact of the input numbers, by providing elements
that appear to be discriminative, such as the crossed line and the bottom part
for the \eight", a bottom straight line and a smoother upper part for the \two",
and the straight upper part together with the curved bottom part for the \ ve".
To probe empirically the impact of sparsity on this representation, we performed
the same experiment using a dictionary with a very low sparsity (0:1), obtaining
encondings without any preponderant value, thereby making it di cult to select
appropriate atoms for explanation. In gure 3 we show the more \important"
atoms obtained on three input images for the Fashion MNIST dataset, a boot, a
dress and trousers, all of them correctly classi ed by the Oracle. Selecting the
atoms with higher encoding values seems to give rise to representative parts of
the selected input, returning parts that can be easily interpreted by an human
interrogator, (e.g., the neck and the sole for the boot, the sleeves for the dress
and the separation between the legs for the trousers).</p>
      <p>As for MNIST, we performed the same experiment using a dictionary with
low sparsity, ending up with results that are di cult to interpret.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>
        We proposed a model-agnostic framework to explain the answers given by
classi cation systems. To achieve this objective, we started by de ning a general
explanation framework based on three entities: an Oracle (providing the answers to
explain), an Interrogator (posing explanation requests) and a Mediator (helping
Interrogator to interpret the Oracle's decisions). We propose a Mediator using
known and established techniques of sparse dictionary learning, together with
Interpretability ML techniques, to give a humanly interpretable explanation of
a classi cation system outcomes. We tried our proposed approach by using an
NMF-based scheme as sparse dictionary learning technique. However, we expect
that any other technique that meets the requirements outlined in Section 3 may
be successfully used to instantiate the proposed framework. The results of the
experiments that we carried out are encouraging, insofar as the explanations
provided seem to be qualitatively signi cant. Nevertheless, more experiments
are necessary to probe the general interest of our approach to explanation. We
plan to perform both a quantitative assessment, to evaluate explanations by
techniques such as those proposed in [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], and a subjective quality assessment
to test how do humans perceive and interpret explanations of this kind.
      </p>
      <p>The proposed approach does not take so far into account factors such as
the internal structure of the dictionary used. Accordingly, the present work can
be extended by considering, for example, whether there are atoms that are
sufciently \similar" to each other or whether the presence in the dictionary of
atoms which can be expressed as combinations of other atoms may a ect the
explanations that are arrived at. Another interesting direction of research
concerns contrastive explanations, which enable one to answer \why not?" negative
questions, by explaining why some given object was not given another classi
cation, di ering from the classi cation that the Oracle actually provided. One
should be careful to note that \why not?" questions are particularly relevant,
from an ethical and legal viewpoint, to address user complaints about purported
misclassi cations and corresponding user requests to be classi ed otherwise.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Adadi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berrada</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Peeking inside the black-box: A survey on explainable arti cial intelligence (xai)</article-title>
          .
          <source>IEEE Access 6</source>
          ,
          <issue>52138</issue>
          {
          <fpage>52160</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bach</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Binder</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montavon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klauschen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , Muller,
          <string-name>
            <given-names>K.R.</given-names>
            ,
            <surname>Samek</surname>
          </string-name>
          , W.:
          <article-title>On pixel-wise explanations for non-linear classi er decisions by layer-wise relevance propagation</article-title>
          .
          <source>PloS one 10(7)</source>
          ,
          <year>e0130140</year>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bao</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quan</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Dictionary learning for sparse coding: Algorithms and convergence analysis</article-title>
          .
          <source>IEEE transactions on pattern analysis and machine intelligence</source>
          <volume>38</volume>
          (
          <issue>7</issue>
          ),
          <volume>1356</volume>
          {
          <fpage>1369</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Binder</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montavon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lapuschkin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Muller,
          <string-name>
            <given-names>K.R.</given-names>
            ,
            <surname>Samek</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.</surname>
          </string-name>
          :
          <article-title>Layer-wise relevance propagation for neural networks with local renormalization layers</article-title>
          .
          <source>In: International Conference on Arti cial Neural Networks</source>
          . pp.
          <volume>63</volume>
          {
          <fpage>71</fpage>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Caruana</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lou</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gehrke</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koch</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sturm</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elhadad</surname>
          </string-name>
          , N.:
          <article-title>Intelligible models for healthcare: Predicting pneumonia risk and hospital 30-day readmission</article-title>
          .
          <source>In: Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          . pp.
          <volume>1721</volume>
          {
          <fpage>1730</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cooper</surname>
            ,
            <given-names>G.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aliferis</surname>
            ,
            <given-names>C.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ambrosino</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aronis</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buchanan</surname>
            ,
            <given-names>B.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caruana</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fine</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glymour</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gordon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanusa</surname>
            ,
            <given-names>B.H.</given-names>
          </string-name>
          , et al.:
          <article-title>An evaluation of machine-learning methods for predicting pneumonia mortality</article-title>
          .
          <source>Arti cial intelligence in medicine 9(2)</source>
          ,
          <volume>107</volume>
          {
          <fpage>138</fpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Doran</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Besold</surname>
            ,
            <given-names>T.R.</given-names>
          </string-name>
          :
          <article-title>What does explainable ai really mean? a new conceptualization of perspectives</article-title>
          .
          <source>arXiv preprint arXiv:1710.00794</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Dosovitskiy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brox</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Inverting visual representations with convolutional networks</article-title>
          .
          <source>In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          . pp.
          <volume>4829</volume>
          {
          <issue>4837</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Erhan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , Courville, ., Vincent,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>Visualizing higher-layer features of a deep network</article-title>
          .
          <source>University of Montreal 1341(3)</source>
          ,
          <volume>1</volume>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Guidotti</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Monreale</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruggieri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turini</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giannotti</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pedreschi</surname>
            ,
            <given-names>D.:</given-names>
          </string-name>
          <article-title>A survey of methods for explaining black box models</article-title>
          .
          <source>ACM computing surveys (CSUR) 51(5)</source>
          ,
          <volume>93</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Hoyer</surname>
            ,
            <given-names>P.O.</given-names>
          </string-name>
          :
          <article-title>Non-negative matrix factorization with sparseness constraints</article-title>
          .
          <source>Journal of machine learning research 5(Nov)</source>
          ,
          <volume>1457</volume>
          {
          <fpage>1469</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Kingma</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ba</surname>
          </string-name>
          , J.:
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>In: International Conference on Learning Representations (12</source>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Lecun</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bottou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ha</surname>
            <given-names>ner</given-names>
          </string-name>
          , P.:
          <article-title>Gradient-based learning applied to document recognition</article-title>
          .
          <source>Proceedings of the IEEE</source>
          <volume>86</volume>
          (
          <issue>11</issue>
          ),
          <volume>2278</volume>
          {
          <fpage>2324</fpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>D.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seung</surname>
            ,
            <given-names>H.S.</given-names>
          </string-name>
          :
          <article-title>Algorithms for non-negative matrix factorization</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>556</volume>
          {
          <issue>562</issue>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Letham</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rudin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCormick</surname>
            ,
            <given-names>T.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madigan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , et al.:
          <article-title>Interpretable classi ers using rules and bayesian analysis: Building a better stroke prediction model</article-title>
          .
          <source>The Annals of Applied Statistics</source>
          <volume>9</volume>
          (
          <issue>3</issue>
          ),
          <volume>1350</volume>
          {
          <fpage>1371</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Lipton</surname>
            ,
            <given-names>Z.C.</given-names>
          </string-name>
          :
          <article-title>The mythos of model interpretability</article-title>
          .
          <source>Queue</source>
          <volume>16</volume>
          (
          <issue>3</issue>
          ),
          <volume>30</volume>
          :
          <fpage>31</fpage>
          {
          <fpage>30</fpage>
          :57 (Jun
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Mahendran</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vedaldi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Understanding deep image representations by inverting them</article-title>
          .
          <source>In: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          . pp.
          <volume>5188</volume>
          {
          <issue>5196</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Mensch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mairal</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varoquaux</surname>
          </string-name>
          , G.:
          <article-title>Dictionary learning for massive matrix factorization</article-title>
          .
          <source>In: Proceedings of The 33rd International Conference on Machine Learning</source>
          . pp.
          <volume>1737</volume>
          {
          <issue>1746</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Montavon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samek</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Muller, K.:
          <article-title>Methods for interpreting and understanding deep neural networks</article-title>
          .
          <source>Digital Signal Processing 73</source>
          ,
          <issue>1</issue>
          {
          <fpage>15</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Montavon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lapuschkin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Binder</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samek</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Muller, K.R.:
          <article-title>Explaining nonlinear classi cation decisions with deep taylor decomposition</article-title>
          .
          <source>Pattern Recognition</source>
          <volume>65</volume>
          ,
          <issue>211</issue>
          {
          <fpage>222</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clune</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dosovitskiy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yosinski</surname>
          </string-name>
          , J.:
          <article-title>Plug &amp; play generative networks: Conditional iterative generation of images in latent space</article-title>
          .
          <source>In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          . pp.
          <volume>4467</volume>
          {
          <issue>4477</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dosovitskiy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yosinski</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brox</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clune</surname>
          </string-name>
          , J.:
          <article-title>Synthesizing the preferred inputs for neurons in neural networks via deep generator networks</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          . pp.
          <volume>3387</volume>
          {
          <issue>3395</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23. Nun~ez, H.,
          <string-name>
            <surname>Angulo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Catala</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Rule extraction from support vector machines</article-title>
          .
          <source>In: Esann</source>
          . pp.
          <volume>107</volume>
          {
          <issue>112</issue>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Qin</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>How convolutional neural network see the worlda survey of convolutional neural network visualization methods</article-title>
          . arXiv preprint arXiv:
          <year>1804</year>
          .
          <volume>11191</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Ribeiro</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guestrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Why should i trust you?: Explaining the predictions of any classi er</article-title>
          .
          <source>In: Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining</source>
          . pp.
          <volume>1135</volume>
          {
          <fpage>1144</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Robnik-Sikonja</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Kononenko, I.:
          <article-title>Explaining classi cations for individual instances</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>20</volume>
          (
          <issue>5</issue>
          ),
          <volume>589</volume>
          {
          <fpage>600</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Simonyan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vedaldi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Deep inside convolutional networks: Visualising image classi cation models and saliency maps</article-title>
          .
          <source>arXiv preprint arXiv:1312.6034</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Sturm</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lapuschkin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samek</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Muller, K.:
          <article-title>Interpretable deep neural networks for single-trial eeg classi cation</article-title>
          .
          <source>Journal of neuroscience methods 274</source>
          , 141{
          <fpage>145</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Szegedy</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaremba</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bruna</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erhan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goodfellow</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fergus</surname>
          </string-name>
          , R.:
          <article-title>Intriguing properties of neural networks</article-title>
          .
          <source>arXiv preprint arXiv:1312.6199</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Xiao</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rasul</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vollgraf</surname>
          </string-name>
          , R.:
          <article-title>Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms</article-title>
          .
          <source>CoRR abs/1708</source>
          .07747 (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Zeiler</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fergus</surname>
          </string-name>
          , R.:
          <article-title>Visualizing and understanding convolutional networks</article-title>
          .
          <source>In: European conference on computer vision</source>
          . pp.
          <volume>818</volume>
          {
          <fpage>833</fpage>
          . Springer (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Zeiler</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , G.W.,
          <string-name>
            <surname>Fergus</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Adaptive deconvolutional networks for mid and high level feature learning</article-title>
          .
          <source>In: Computer Vision</source>
          (ICCV),
          <year>2011</year>
          IEEE International Conference on. pp.
          <year>2018</year>
          {
          <year>2025</year>
          . IEEE (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Visual interpretability for deep learning: a survey</article-title>
          .
          <source>Frontiers of Information Technology &amp; Electronic Engineering</source>
          <volume>19</volume>
          (
          <issue>1</issue>
          ),
          <volume>27</volume>
          {
          <fpage>39</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Zintgraf</surname>
            ,
            <given-names>L.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>T.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adel</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Welling</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Visualizing deep neural network decisions: Prediction di erence analysis</article-title>
          .
          <source>arXiv preprint arXiv:1702.04595</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>