<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On the Readability of Deep Learning Models: the role of Kernel-based Deep Architectures</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Danilo Croce</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniele Rossini</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Basili</string-name>
          <email>basilig@info.uniroma2.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department Of Enterprise Engineering University of Roma</institution>
          ,
          <addr-line>Tor Vergata</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. Deep Neural Networks achieve state-of-the-art performances in several semantic NLP tasks but lack of explanation capabilities as for the limited interpretability of the underlying acquired models. In other words, tracing back causal connections between the linguistic properties of an input instance and the produced classification is not possible. In this paper, we propose to apply Layerwise Relevance Propagation over linguistically motivated neural architectures, namely Kernel-based Deep Architectures (KDA), to guide argumentations and explanation inferences. In this way, decisions provided by a KDA can be linked to the semantics of input examples, used to linguistically motivate the network output.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. Le Deep Neural Network
raggiungono oggi lo stato dell’arte in
molti processi di NLP, ma la scarsa
interpretabilita´ dei modelli risultanti
dall’addestramento limita la
comprensione delle loro inferenze. Non e´ possibile
cioe´ determinare connessioni causali tra
le proprieta´ linguistiche di un esempio
e la classificazione prodotta dalla rete.
In questo lavoro, l’applicazione della
Layerwise Relevance Propagation alle
Kernel-based Deep Architecture(KDA)
e´ usata per determinare connessioni tra
la semantica dell’input e la classe di
output che corrispondono a spiegazioni
linguistiche e trasparenti della decisione.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        Deep Neural Networks are usually criticized as
they are not epistemologically transparent devices,
i.e. their models cannot be used to provide
explanations of the resulting inferences. An
example can be neural question classification (QC) (e.g.
        <xref ref-type="bibr" rid="ref6">(Croce et al., 2017)</xref>
        ). In QC the correct category of
a question is detected to optimize the later stages
of a question answering system,
        <xref ref-type="bibr" rid="ref10">(Li and Roth,
2006)</xref>
        . An epistemologically transparent learning
system should trace back the causal connections
between the proposed question category and the
linguistic properties of the input question. For
example, the system could motivate the decision:
”What is the capital of Zimbabwe?” refers to a
Location, with a sentence such as: Since it is
similar to ”What is the capital of California?”
which also refers to a Location. Unfortunately,
neural models, as for example Multilayer
Perceptrons (MLP), Long Short-Term Memory Networks
(LSTM),
        <xref ref-type="bibr" rid="ref7">(Hochreiter and Schmidhuber, 1997)</xref>
        , or
even Attention-based Networks
        <xref ref-type="bibr" rid="ref9">(Larochelle and
Hinton, 2010)</xref>
        , correspond to parameters that have
no clear conceptual counterpart: it is thus difficult
to trace back the network components (e.g.
neurons or layers in the resulting topology)
responsible for the answer.
      </p>
      <p>
        In image classification, Layerwise Relevance
Propagation (LRP)
        <xref ref-type="bibr" rid="ref1">(Bach et al., 2015)</xref>
        has been
used to decompose backward across the MLP
layers the evidence about the contribution of
individual input fragments (i.e. pixels of the input
images) to the final decision. Evaluation against
the MNIST and ILSVRC benchmarks suggests
that LRP activates associations between input and
output fragments, thus tracing back meaningful
causal connections.
      </p>
      <p>
        In this paper, we propose the use of a
similar mechanism over a linguistically motivated
network architecture, the Kernel-based Deep
Architecture (KDA),
        <xref ref-type="bibr" rid="ref6">(Croce et al., 2017)</xref>
        . Tree
Kernels
        <xref ref-type="bibr" rid="ref3">(Collins and Duffy, 2001)</xref>
        are here used to
integrate syntactic/semantic information within a
MLP network. We will show how KDA input
nodes correspond to linguistic instances and by
applying the LRP method we are able to trace back
causal associations between the semantic
classification and such instances. Evaluation of the LRP
algorithm is based on the idea that explanations
improve the user expectations about the
correctness of an answer and shows its applicability in
human computer interfaces.
      </p>
      <p>In the rest of the paper, Section 2 describes the
KDA neural approach while section 3 illustrates
how LRP connects to KDAs. In section 4 early
results of the evaluation are reported.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Training Neural Networks in Kernel</title>
    </sec>
    <sec id="sec-4">
      <title>Spaces</title>
      <p>
        Given a training set o 2 D, a kernel K(oi; oj )
is a similarity function over D2 that corresponds
to a dot product in the implicit kernel space,
i.e., K(oi; oj ) = (oi) (oj ). Kernel functions
are used by learning algorithms, such as Support
Vector Machines
        <xref ref-type="bibr" rid="ref11">(Shawe-Taylor and Cristianini,
2004)</xref>
        , to efficiently operate on instances in the
kernel space: their advantage is that the
projection function (o) = ~x 2 Rn is never explicitly
computed. The Nystro¨m method is a factorization
method applied to derive a new low-dimensional
embedding x~ in a l-dimensional space, with l n
so that G G~ = X~ X~ &gt;, where G = XX&gt; is
the Gram matrix such that Gij = (oi) (oj ) =
K(oi; oj ). The approximation G~ is obtained using
a subset of l columns of the matrix, i.e., a
selection of a subset L D of the available
examples, called landmarks. Given l randomly
sampled columns of G, let C 2 RjDj l be the
matrix of these sampled columns. Then, we can
rearrange the columns and rows of G and define
X = [X1 X2] such that:
      </p>
      <p>G =</p>
      <p>W
X2&gt;X1</p>
      <p>X1&gt;X2
X2&gt;X2
=</p>
      <p>C
X2&gt;X1
where W = X1&gt;X1, i.e., the subset of G that
contains only landmarks. The Nystro¨m
approximation can be defined as:</p>
      <p>G</p>
      <p>G~ = CW yC&gt;
(1)
where W y denotes the Moore-Penrose inverse of
W . If we apply the Singular Value Decomposition
(SVD) to W , which is symmetric definite
positive, we get W = U SV &gt; = U SU &gt;. Then it
is straightforward to see that W y = U S 1U &gt; =
U S 12 S 21 U &gt; and that by substitution G G~ =
(CU S 2 )(CU S 21 )&gt; = X~ X~ &gt;. Given an
exam1
ple o 2 D, its new low-dimensional representation
~x~ is determined by considering the corresponding
item of C as
1
~x~ = ~cU S 2
(2)
where ~c is the vector whose dimensions contain
the evaluations of the kernel function between o
and each landmark oj 2 L. Therefore, the method
produces l-dimensional vectors.</p>
      <p>
        Given a labeled dataset, a Multi-Layer
Perceptron (MLP) architecture can be defined, with a
specific Nystro¨m layer based on the Nystro¨m
embeddings of Eq. 2,
        <xref ref-type="bibr" rid="ref6">(Croce et al., 2017)</xref>
        .
      </p>
      <p>Such Kernel-based Deep Architecture (KDA)
has an input layer, a Nystro¨m layer, a possibly
empty sequence of non-linear hidden layers and a
final classification layer, which produces the
output. In particular, the input layer corresponds to
the input vector ~c, i.e., the row of the C matrix
associated to an example o. It is then mapped to
the Nystro¨m layer, through the projection in
Equation 2. Notice that the embedding provides also
the proper weights, defined by U S 21 , so that the
mapping can be expressed through the Nystro¨m
1
matrix HNy = U S 2 : it corresponds to a
pretraining stage based on the SVD. Formally, the
low-dimensional embedding of an input example
1
o, ~x~ = ~c HNy = ~c U S 2 encodes the kernel
space. Any neural network can then be adopted:
in the rest of this paper, we assume that a
traditional Multi-Layer Perceptron (MLP) architecture
is stacked in order to solve the targeted
classification problems. The final layer of KDA is the
classification layer whose dimensionality depends on
the classification task: it computes a linear
classification function with a softmax operator.</p>
      <p>
        A KDA is stimulated by an input vector c which
corresponds to the kernel evaluations K(o; li)
between each example o and the landmarks li.
Linguistic kernels (such as Semantic Tree
Kernels
        <xref ref-type="bibr" rid="ref4">(Croce et al., 2011)</xref>
        ) depend on the
syntactic/semantic similarity between the x and the
subset of li used for the space reconstruction. We will
see hereafter how tracing back through relevance
propagation into a KDA architecture corresponds
to determine which semantic landmarks contribute
mostly to the final output decision.
3
      </p>
    </sec>
    <sec id="sec-5">
      <title>Layer-wise Relevance Propagation in</title>
    </sec>
    <sec id="sec-6">
      <title>Kernel-based Deep Architectures</title>
      <p>
        Layer-wise Relevance propagation (LRP,
presented in
        <xref ref-type="bibr" rid="ref1">(Bach et al., 2015)</xref>
        ) is a framework which
allows to decompose the prediction of a deep
neural network computed over a sample, e.g. an
image, down to relevance scores for the single input
dimensions, such as a subset of pixels.
      </p>
      <p>Formally, let f : Rd ! R+ be a positive
realvalued function taking a vector ~x 2 Rd as input: f
quantifies, for example, the probability of ~x
characterizing a certain class. The Layer-wise
Relevance Propagation assigns to each dimension, or
feature, xd, a relevance score Rd(1) such that:
f (x)</p>
      <p>P R(1)
d d
Features whose score R(1) &gt; 0 (or d R(1) &lt; 0)
d d
correspond to evidence in favor (or against) the
output classification. In other words, LRP allows
to identify fragments of the input playing key roles
in the decision, by propagating relevance
backwards. Let us suppose to know the relevance score
R(l+1) of a neuron j at network layer l + 1, then it
j
can be decomposed into messages Ri(l;l+j1) sent to
neurons i in layer l:</p>
      <p>R(l+1) = X R(l;l+1)
j i j</p>
      <p>i2(l)
Hence the relevance of a neuron i at layer l can be
defined as:</p>
      <p>R(l) =
i</p>
      <p>X
j2(l+1)</p>
      <p>
        R(l;l+1)
i j
Note that 4 and 5 are such that 3 holds. In this
work, we adopted the -rule defined in
        <xref ref-type="bibr" rid="ref1">(Bach et
al., 2015)</xref>
        to compute the messages R(l;l+1), i.e.
i j
R(l;l+1) =
i j
zj +
zij R(l+1)
sign(zj ) j
where zij = xiwij and &gt; 0 is a numerical
stabilizing term and must be small. Notice that weights
wij correspond to weighted activations of input
neurons. If we apply LRP to a KDA it
implicitly traces the relevance back to the input layer,
i.e. to the landmarks. It thus tracks back
syntactic, semantic and lexical relations between a
question and the landmark and it grants high relevance
to the relations the network selected as highly
discriminating for the class representations it learned;
note that this is different from similarity in terms
of kernel-function evaluation as the latter is task
independent whereas LRP scores are not. Notice
also that each landmark is uniquely associated to
an entry of the input vector ~c, as shown in Sec 2,
and, as a member of the training dataset, it also
corresponds to a known class.
(3)
(4)
(5)
      </p>
    </sec>
    <sec id="sec-7">
      <title>Explanatory Models</title>
      <p>LRP allows the automatic compilation of
justifications for the KDA classifications: explanations are
possible using landmarks f`g as examples. The
f`g that the LRP method produces as the most
active elements in layer 0 are semantic analogues of
input annotated examples. An Explanatory Model
is the function in charge of compiling the
linguistically fluent explanation of individual analogies
(or differences) with the input case. The
meaningfulness of such analogies makes a resulting
explanation clear and should increase the user
confidence on the system reliability. When a sentence
o is classified, LRP assigns activation scores r`s to
each individual landmark `: let L(+) (or L( ))
denote the set of landmarks with positive (or
negative) activation scores.</p>
      <p>Formally, an explanation is characterized by a
triple e = hs; C; i where s is the input sentence,
C is the predicted label and is the modality of the
explanation: = +1 for positive (i.e. acceptance)
statements while = 1 correspond to rejections
of the decision C. A landmark ` is positively
activated for a given sentence s if there are not more
than k 1 other active landmarks1 `0 whose
activation value is higher than the one for `, i.e.
jf`0 2 L(+) : `0 6= ` ^ r`0
s
r`s &gt; 0gj &lt; k</p>
      <p>A landmark is negatively activated when: jf`0 2
L( ) : `0 6= ` ^ r`s0 r`s &lt; 0gj &lt; k. Positively
(or negative) active landmarks in Lk are assigned
to an activation value a(`; s) = +1 ( 1). For all
other not activated landmarks: a(`; s) = 0.</p>
      <p>Given the explanation e = hs; C; i, a landmark
` whose (known) class is C` is consistent (or
inconsistent) with e according to the fact that the
following function:</p>
      <p>(C`; C) a(`; q)
is positive (or negative, respectively), where
(C0; C) = 2 kron(C0 = C) 1 and kron is the
Kronecker delta.</p>
      <p>The explanatory model is then a function
M(e; Lk) which maps an explanation e, a sub set
Lk of the active and consistent landmarks L for e
into a sentence in natural language. Of course
several definitions for M(e; Lk) and Lk are possible.</p>
      <p>1k is a parameter used to make explanation depending on
not more than k landmarks, denoted by Lk.</p>
      <p>
        A general explanatory model would be:
8 “ s is C since it is similar to ` ”
&gt;
&gt;&gt;&gt;&gt;&gt; 8` 2 Lk+ if &gt; 0
&gt;
&gt;&gt;&gt;&gt; “ s is not C since it is different
M(e; Lk) = &lt; from ` which is C ”
if &lt; 0
&gt;&gt;&gt; 8` 2 Lk
&gt;
&gt;
&gt;
&gt;
&gt; “ s is C but I don’t know why ”
&gt;
&gt;:&gt; if Lk = ;
where Lk+,Lk Lk are the partitions of landmarks
with positive (and negative) relevance scores in
Lk, respectively. Here we provide examples for
two explanatory models, used during the
experimental evaluation. A first possible model returns
the analogy only with the (unique) consistent
landmark with the highest positive score if = 1
and lowest negative when = 1. The
explanation of a rejected decision in the Argument
Classification of a Semantic Role Labeling task
        <xref ref-type="bibr" rid="ref12">(Vanzo et al., 2016)</xref>
        , described by the triple e1 =
h’vai in camera da letto’; SOURCEBRINGING; 1i,
is:
      </p>
      <p>I think ”in camera da letto” IS NOT [SOURCE] of
[BRINGING] in ”Vai in camera da letto” (LU:[vai]) since
it’s different from ”sul tavolino” which is [SOURCE] of
[BRINGING] in “Portami il mio catalogo sul tavolino”
(LU:[porta])</p>
      <p>The second model uses two active
landmarks: one consistent and one contradictory
with respect to the decision. For the triple
e1 = h’vai in camera da letto’; GOALMOTION; 1i
the second model produces:</p>
      <p>I think ”in camera da letto” IS [GOAL] of [MOTION] in
”Vai in camera da letto” (LU:[vai]) since it recalls ”al
telefono” which is [GOAL] of [MOTION] in ”Vai al telefono
e controlla se ci sono messaggi” (LU:[vai]) and it IS NOT
[SOURCE] of [BRINGING] since different from ”sul
tavolino” which is the [SOURCE] of [BRINGING] in
”Portami il mio catalogo sul tavolino” (LU:[portami])
4.1</p>
      <sec id="sec-7-1">
        <title>Evaluation methodology</title>
        <p>
          In order to evaluate the impact of the produced
explanations, we defined the following task: given a
classification decision, i.e. the input o is classified
as C, to measure the impact of the explanation e
on the belief that a user exhibits on the statement
“o 2 C is true”. This information can be
modeled through the estimates of the following
probabilities: P (o 2 C) that characterizes the amount
of confidence the user has in accepting the
statement, and its corresponding form P (o 2 Cje),
i.e. the same quantity in the case the user is
provided by the explanation e. The core idea is that
semantically coherent and exhaustive explanations
must indicate correct classifications whereas
incoherent or non-existent explanations must hint
towards wrong classifications. A quantitative
measure of such an increase (or decrease) in
confidence is the Information Gain (IG,
          <xref ref-type="bibr" rid="ref8">(Kononenko
and Bratko, 1991)</xref>
          ) of the decision o 2 C. Notice
that IG measures the increase of probability
corresponding to correct decisions, and the reduction of
the probability in case the decision is wrong. This
amount suitably addresses the shift in uncertainty
log2(P ( )) between two (subjective) estimates,
i.e., P (o 2 C) vs. P (o 2 Cje).
        </p>
        <p>Different explanatory models M can be also
compared. The relative Information Gain IM
is measured against a collection of explanations
e 2 TM generated by M and then normalized
throughout the collection’s entropy E as follows:
IM =
1</p>
        <p>1
E j TM j e2TM</p>
        <p>X I(e)
where I(e) is the IG of each explanation2.
5</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Experimental Evaluation</title>
      <p>
        The effectiveness of the proposed approach has
been measured against two different semantic
processing tasks, i.e. Question Classification (QC)
over the UIUC dataset
        <xref ref-type="bibr" rid="ref10">(Li and Roth, 2006)</xref>
        and
Argument Classification in Semantic Role Labeling
(SRL-AC) over the HuRIC dataset
        <xref ref-type="bibr" rid="ref12 ref2">(Bastianelli et
al., 2014; Vanzo et al., 2016)</xref>
        . The adopted
architecture consisted in a LRP-integrated KDA with 1
hidden layers and 500 landmarks for QC, 2
hidden layers and 100 landmarks for SRL-AC and a
stabilization-term = 10e 8.
      </p>
      <p>
        We defined five quality categories and
associated each with a value of P (o 2 Cje), as
shown in Table 1. Three annotators then
independently rated explanations generated from a
collection composed of an equal number of correct
and wrong classifications (for a total amount of
300 and 64 explanations, respectively, for QC and
SRL-AC). This perfect balancing makes the prior
probability P (o 2 C) being 0.5, i.e. maximal
entropy with a baseline IG = 0 in the [ 1; 1] range.
Notice that annotators had no information on the
2More details are in
        <xref ref-type="bibr" rid="ref8">(Kononenko and Bratko, 1991)</xref>
        QC
0.548
0.580
SRL-AC
0.669
0.784
system classification performance, but just
knowledge of the explanation dataset entropy.
      </p>
      <sec id="sec-8-1">
        <title>5.1 Question Classification</title>
        <p>Experimental evaluations3 showed that both the
models were able to gain more than half the bit
required to ascertain whether the network statement
is true or not (Table 2). Consider:
I think ”What year did Oklahoma become a state ?” refers
to a NUMBER since recalls me ”The film Jaws was made in
what year ?”
Here the model returned a coherent supporting
evidence, a somewhat easy case as for the available
discriminative pair, i.e. ”What year”. The
system is able to capture semantic similarities even in
poorer conditions, e.g.:</p>
        <p>I think ”Where is the Mall of the America ?” refers to a
LOCATION since recalls me ”What town was the setting for</p>
        <p>The Music Man ?” which refers to a LOCATION.
This high quality explanation is achieved even if
with such poor lexical overlap. It seems that richer
representations are here involved with
grammatical and semantic similarity used as the main
information involved in the decision at hand. Let us
consider:</p>
        <p>I think ”Mexican pesos are worth what in U.S. dollars ?”
refers to a DESCRIPTION since it recalls me ”What is the</p>
        <p>Bernoulli Principle ?”
Here the provided explanation is incoherent, as
expected since the classification is wrong. Now
consider:</p>
        <p>I think ”What is the sales tax in Minnesota ?” refers to a
NUMBER since it recalls me ”What is the population of
Mozambique ?” and does not refer to a ENTITY since
different from ”What is a fear of slime ?”.</p>
        <p>
          3For details on KDA performance against the task, see
          <xref ref-type="bibr" rid="ref6">(Croce et al., 2017)</xref>
          Although explanation seems fairly coherent, it is
actually misleading as ENTITY is the annotated
class. This shows how the system may lack of
contextual information, as humans do, against
inherently ambiguous questions.
        </p>
      </sec>
      <sec id="sec-8-2">
        <title>5.2 Argument Classification</title>
        <p>
          Evaluation also targeted a second task, that is
Argument classification in Semantic Role Labeling
(SRL-AC): KDA is here fed with vectors from
tree kernel evaluations as discussed in
          <xref ref-type="bibr" rid="ref4">(Croce et
al., 2011)</xref>
          . The evaluation is carried out over
the HuRIC dataset
          <xref ref-type="bibr" rid="ref12">(Vanzo et al., 2016)</xref>
          , including
about 240 domotic commands in Italian,
comprising of about 450 roles. The system has an accuracy
of 91.2% on about 90 examples, while the training
and development set have a size of, respectively,
270 and 90 examples. We considered 64
explanations for measuring the IG of the two explanation
models. Table 2 confirms that both explanatory
models performed even better than in QC. This is
due to the narrower linguistic domain (14 frames
are involved) and the clearer boundaries between
classes: annotators seem more sensitive to the
explanatory information to assess the network
decision. Examples of generated sentences are:
I think ”con me” is NOT the MANNER of COTHEME in
”Robot vieni con me nel soggiorno? (LU:[vieni])” since it
does NOT recall me ”lentamente” which is MANNER in
”Per favore segui quella persona lentamente (LU:[segui])”.
It is rather COTHEME of COTHEME since it recalls me
”mi” which is COTHEME in ”Seguimi nel bagno
(LU:[segui])”.
6
        </p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Conclusion and Future Works</title>
      <p>
        This paper describes an LRP application to a KDA
that makes use of analogies as explanations of a
neural network decision. A methodology to
measure the explanation quality has been also
proposed and the experimental evidence confirms the
effectiveness of the method in increasing the trust
of a user upon automatic classifications. Future
work will focus on the selection of subtrees as
meaningful evidences for the explanation, or on
the modeling of negative information for
disambiguation as well as on more in depth investigation
of the landmark selection policies. Moreover,
improved experimental scenarios involving users and
dialogues will be also designed, e.g. involving
further investigation within Semantic Role Labeling,
using the method proposed in
        <xref ref-type="bibr" rid="ref5">(Croce et al., 2012)</xref>
        .
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Bach</surname>
          </string-name>
          , Alexander Binder, Gregoire Montavon, Frederick Klauschen,
          <string-name>
            <surname>Klaus-Robert Mller</surname>
            , and
            <given-names>Wojciech</given-names>
          </string-name>
          <string-name>
            <surname>Samek</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation</article-title>
          .
          <source>PLOS ONE</source>
          ,
          <volume>10</volume>
          (
          <issue>7</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Emanuele</given-names>
            <surname>Bastianelli</surname>
          </string-name>
          , Giuseppe Castellucci, Danilo Croce, Luca Iocchi, Roberto Basili, and
          <string-name>
            <given-names>Daniele</given-names>
            <surname>Nardi</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Huric: a human robot interaction corpus</article-title>
          .
          <source>In LREC</source>
          , pages
          <fpage>4519</fpage>
          -
          <lpage>4526</lpage>
          .
          <string-name>
            <given-names>European</given-names>
            <surname>Language Resources Association (ELRA).</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Michael</given-names>
            <surname>Collins</surname>
          </string-name>
          and
          <string-name>
            <given-names>Nigel</given-names>
            <surname>Duffy</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>New ranking algorithms for parsing and tagging: Kernels over discrete structures, and the voted perceptron</article-title>
          .
          <source>In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics (ACL '02), July 7-12</source>
          ,
          <year>2002</year>
          , Philadelphia, PA, USA, pages
          <fpage>263</fpage>
          -
          <lpage>270</lpage>
          . Association for Computational Linguistics, Morristown, NJ, USA.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Danilo</given-names>
            <surname>Croce</surname>
          </string-name>
          , Alessandro Moschitti, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Basili</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Structured lexical similarity via convolution kernels on dependency trees</article-title>
          .
          <source>In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>1034</fpage>
          -
          <lpage>1046</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Danilo</given-names>
            <surname>Croce</surname>
          </string-name>
          , Alessandro Moschitti, Roberto Basili, and
          <string-name>
            <given-names>Martha</given-names>
            <surname>Palmer</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Verb classification using distributional similarity in syntactic and semantic structures</article-title>
          .
          <source>In ACL (1)</source>
          , pages
          <fpage>263</fpage>
          -
          <lpage>272</lpage>
          . The Association for Computer Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Danilo</given-names>
            <surname>Croce</surname>
          </string-name>
          , Simone Filice, Giuseppe Castellucci, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Basili</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Deep learning in semantic kernel spaces</article-title>
          .
          <source>In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          , pages
          <fpage>345</fpage>
          -
          <lpage>354</lpage>
          , Vancouver, Canada, July. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Sepp</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and Ju¨rgen Schmidhuber.
          <year>1997</year>
          .
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural Comput.</source>
          ,
          <volume>9</volume>
          (
          <issue>8</issue>
          ):
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          , November.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Igor</given-names>
            <surname>Kononenko</surname>
          </string-name>
          and Ivan Bratko.
          <year>1991</year>
          .
          <article-title>Informationbased evaluation criterion for classifier's performance</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>6</volume>
          (
          <issue>1</issue>
          ):
          <fpage>67</fpage>
          -
          <lpage>80</lpage>
          , Jan.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Hugo</given-names>
            <surname>Larochelle</surname>
          </string-name>
          and
          <string-name>
            <given-names>Geoffrey E.</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Learning to combine foveal glimpses with a thirdorder boltzmann machine</article-title>
          .
          <source>In Proceedings of Neural Information Processing Systems (NIPS)</source>
          , pages
          <fpage>1243</fpage>
          -
          <lpage>1251</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Xin</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dan</given-names>
            <surname>Roth</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Learning question classifiers: the role of semantic information</article-title>
          .
          <source>Natural Language Engineering</source>
          ,
          <volume>12</volume>
          (
          <issue>3</issue>
          ):
          <fpage>229</fpage>
          -
          <lpage>249</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>John</surname>
            Shawe-Taylor and
            <given-names>Nello</given-names>
          </string-name>
          <string-name>
            <surname>Cristianini</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Kernel Methods for Pattern Analysis</article-title>
          . Cambridge University Press, Cambridge, UK.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Vanzo</surname>
          </string-name>
          , Danilo Croce, Roberto Basili, and
          <string-name>
            <given-names>Daniele</given-names>
            <surname>Nardi</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Context-aware spoken language understanding for human robot interaction</article-title>
          .
          <source>In Proceedings of Third Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2016</year>
          ) &amp;
          <article-title>Fifth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          .
          <source>Final Workshop (EVALITA</source>
          <year>2016</year>
          ), Napoli, Italy, December 5-
          <issue>7</issue>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>