<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mitigating Bias in Deep Nets with Knowledge Bases : the Case of Natural Language Understanding for Robots</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Martino Mensio</string-name>
          <email>martino.mensio@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emanuele Bastianelli?</string-name>
          <email>emanuele.bastianelli@hw.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ilaria Tiddiy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Rizzoz</string-name>
          <email>giuseppe.rizzo@linksfoundation.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Copyright c 2020 held by the author(s). In A. Martin, K. Hinkelmann</institution>
          ,
          <addr-line>H.-G. Fill, A. Gerber, D. Lenat, R. Stolle, F. van Harmelen (Eds.)</addr-line>
          ,
          <institution>Proceedings of the AAAI 2020 Spring Symposium on Combining Machine Learning and Knowledge Engineering in Practice (AAAI-MAKE 2020). Stanford University</institution>
          ,
          <addr-line>Palo Alto, California</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Knowledge Media Institute, The Open University</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we tackle the problem of lack of understandability of deep learning systems by integrating heterogeneous knowledge sources, and in the specific we present how we used FrameNet to guarantee the correct learning for an LSTM-based semantic parser in the task of Spoken Language Understanding for robots. The problem of the explainability of Artificial Intelligence (AI) systems, i.e. their ability to explain decisions to both experts and end users, has attracted growing attention in the latest years, affecting their credibility and trustworthiness. Trusting these systems is fundamental in the context of AI-based robotic companions interacting in natural language, as the users' acceptance of the robot also relies on the ability to explain the reasons behind its actions. Following similar approaches, we first use the values of the neural attention layers employed in the semantic parser as a clue to analyze and interpret the model's behavior and reveal the intrinsic bias induced by the training data. We then show how the integration of knowledge from external resources such as FrameNet can help minimizing, or mitigating, such bias, and consequently guarantee the model to provide the correct interpretations. Our preliminary, but promising results suggest that (i) attention layers can improve the model understandability; (ii) the integration of different knowledge bases can help overcoming the limitations of machine learning models; and (iii) an approach combining the strengths of both knowledge engineering and machine learning can foster the development of more transparent, understandable intelligent systems.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>With the dramatic success of new machine learning
techniques relying on deep architectures, the number of
Artificial Intelligence (AI)-based systems has rapidly increased.</p>
      <p>
        EB and IT developed the theoretical framework and directed the
project; EB and GR designed the experiments; MM derived the
models, performed experiments and analysed the results; IT and
EB wrote the manuscript in consultation with GR and MM.
Events such as the Cambridge Analytica scandal and the
disruptions of the 2016 US elections have brought researchers
and practitioners to question the explainability of these
systems, i.e. their ability to explain decisions to both experts
and end-users, resulting in a number of initiatives to improve
their understandability and trustworthiness
        <xref ref-type="bibr" rid="ref25 ref40">(cfr. DARPA’s
eXplainable AI program1; the “right to explanation”
requested by the European General Data Protection
Regulation; and the “Ethics guidelines for Trustworthy AI”
published by the European Union in December 20182)</xref>
        . In the
context of robotic companions interacting in natural
language using AI techniques, where our research is placed,
trust and transparency are fundamental aspects, as the users’
acceptance of the robot assistants will be also based on their
ability to explain the reasons behind their actions, if and
when required.
      </p>
      <p>
        Let us take the example of a robot understanding
spoken commands given by a human, e.g. “take the book
from the table”, and where a corresponding robot action
such as take(book, table) has to be instantiated
correctly. Such instantiation is generally triggered by a trained
model, where noise, over-fitting, and mislabeling could
indeed bring to an undesired output, e.g. the robot placing
the book on the table. In the view of symbiotic autonomous
robots
        <xref ref-type="bibr" rid="ref12 ref29">(Rosenthal, Biswas, and Veloso 2010)</xref>
        that rely on
humans to overcome their limitations and correct their actions,
a transparent model could help identifying and explicit the
reason(s) behind the wrong behavior of the robot.
      </p>
      <p>
        Our motivation is the semantic processing of robotic
commands (also called semantic parsing) from spoken language
utterances, i.e. the process of mapping natural language
sentences to formal meaning representations. The formal
meaning representation theory we rely upon is Frame
Semantics
        <xref ref-type="bibr" rid="ref14">(Fillmore 1985)</xref>
        , describing actions and events
expressed in language through conceptual structures called
semantic frames. This theory also states that a frame is evoked
in a sentence through the occurrence of specific lexical units,
i.e. words (such as verbs and nouns) that linguistically
express the underlying situation. To identify such frames, we
      </p>
      <sec id="sec-1-1">
        <title>1https://www.darpa.mil/program/explainable-artificial</title>
        <p>intelligence</p>
        <p>
          2https://ec.europa.eu/futurium/en/ai-allianceconsultation/guidelines
built a semantic parser based on a multi-layer Long-Short
Term Memory (LSTM) neural network with attention
          <xref ref-type="bibr" rid="ref24">(Mensio et al. 2018)</xref>
          , and trained it over the Human-Robot
Interaction Corpus (HuRIC)
          <xref ref-type="bibr" rid="ref6">(Bastianelli et al. 2014)</xref>
          . LSTMs as
many similar deep nets-based models have an opaque
nature, i.e. they do not give clear clues on the way they
behave, which may complicate the understanding of undesired
behaviors as, in our case, an incorrect robot behavior.
Moreover, understanding the inner workings of such models tend
to be harder when trained on small, domain-specific datasets
(such as HuRIC), as they often lack of effective
representativeness of the problem domain. The questions we wish to
answer in this work are therefore:
how can we better understand our LSTM-based model?
how we can we identify undesired behaviors in the model?
is there a way to mitigate such undesired behaviors?
To answer the first two questions, we rely on the idea that
linguistic theories could be used in the context of our
semantic parser to obtain more understanding of the model,
i.e. they could be exploited to provide to explain the model’s
behavior. Recent trends in deep learning have shown that
visual explanations for the models’ behavior could be obtained
through the analysis of the values of the attention layers in a
number of tasks (Machine Translation
          <xref ref-type="bibr" rid="ref3 ref39">(Bahdanau, Cho, and
Bengio 2014)</xref>
          , Sentiment Analysis
          <xref ref-type="bibr" rid="ref22">(Lin et al. 2017)</xref>
          , Image
Captioning
          <xref ref-type="bibr" rid="ref37">(Xu et al. 2015)</xref>
          ) for their ability of correlating
inputs and outputs. Inspired by these works, our
hypothesis is that we can use attentions to achieve some degree of
explainability for the LSTM-based parser, and that Frame
Semantics can be the key to drive the interpretation process.
We therefore use attentions to capture the interpretation of
spoken commands and, more specifically, use the values that
the attention layer assign to each word of a given sentence to
detect which word is the lexical unit evoking (i.e. causing)
the identified frame. We show how this not only gives us a
hint on the model behavior, but that attentions help
unveiling the intrinsic bias induced by our training data. Here, we
exploit the linguistic knowledge encoded in an external
resource such as FrameNet
          <xref ref-type="bibr" rid="ref5">(Baker, Fillmore, and Lowe 1998)</xref>
          in a data augmentation strategy, with the goal of mitigating
the corpus bias, improve the explanations that the model
provides and, consequently, the overall model results.
        </p>
        <p>Although preliminary, our promising results suggest that
attention layers combined with Frame Semantics do provide
a clue to a more explainable model, and that the integration
of external knowledge bases can help overcoming the inner
limitations of machine learning models. More importantly,
our method suggests that the combination of knowledge
engineering and machine learning techniques can be beneficial
for the development of more transparent, understandable
intelligent systems.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Motivation and Background</title>
      <p>In this section, we present the theoretical and technical
background of our work. We first discuss Frame Semantics,
which we use as linguistic theory of reference, and then
describe the technical details of our neural network-based
semantic parser.</p>
      <sec id="sec-2-1">
        <title>Fundamentals of Frame Semantics</title>
        <p>Frame Semantics is a theory that formalizes how a sentence
is related to semantic frames. Each frame is a conceptual
structure representing an action or, more in general, an event
or situation (e.g. the action of Taking). Frames are further
specified by a set of frame elements (e.g. the THEME,
representing the object taken while performing the action
Taking), which enhance the meaning of the frame with
additional information. According to Frame Semantics, frames
are evoked in sentences by specific words, called lexical
units (LU). Lexical units are responsible to convey the
meaning of the frames, representing hooks between the textual
surface and the theory itself. In the example of Figure 1, the
frame Taking is evoked in the sentence “take the book to the
table” by the LU take, while the book and to the table
represent the THEME and the GOAL frame elements respectively:
[take]Taking
[the book]THEME
[from the table]ORIGIN</p>
        <p>The process of annotating Frame Semantics over natural
language involves three different tasks. First, all the frames
evoked in a sentence are identified looking at the potential
LUs contained it. This task is generally called Frame
Prediction or Frame Induction. Here, we refer to it as Action
Detection (AD), as we are dealing with the action expressed
by the person uttering the command to the robot. The second
task is called Argument Identification (AI, sometimes also
called Boundary Detection) and is responsible to find the
spans of text corresponding to possible frame elements. The
last task is called Argument Classification (AC) and consists
in assigning a label to the spans identified during the AI.
Note that the AI and AD tasks are often referred together as
the process of Semantic Role Labelling.</p>
        <p>If we take the example of “take the book from the table”,
the frame Taking would be predicted in the AD step by
identifying the LU take. In the AI step, the book and from the
table would be identified as 2 frame element spans, and
respectively classified as THEME and ORIGIN frame elements
in the following AC step.</p>
      </sec>
      <sec id="sec-2-2">
        <title>A multi-layer LSTM-based parser</title>
        <p>
          In our previous work
          <xref ref-type="bibr" rid="ref24">(Mensio et al. 2018)</xref>
          , we presented a
semantic parser for robotic commands, called 3LSTM-ATT,
based on a multi-layer LSTM network exploiting attention
mechanisms. The 3LSTM-ATT topology is shown in
Figure 2. The network was adapted from
          <xref ref-type="bibr" rid="ref23">(Liu and Lane 2016)</xref>
          so that each layer could carry one of the three semantic
parsing tasks presented above. We briefly describe the network
in the following, and refer the reader to the original paper
for more details.
        </p>
        <p>
          The input to the network is a tokenized sentence, where
each token is embedded using the GloVe word
embeddings
          <xref ref-type="bibr" rid="ref27 ref39">(Pennington, Socher, and Manning 2014)</xref>
          , pre-trained
ac
over the Common Crawl resource3. The sequence is firstly
encoded with a bidirectional LSTM (L1). For the AD task,
a single contextual representation for the whole sequence c
is computed through an attention layer
          <xref ref-type="bibr" rid="ref3 ref39">(Bahdanau, Cho, and
Bengio 2014)</xref>
          , which is in turn passed through a fully
connected layer with a final softmax activation to obtain
perframe probabilities. The sequence out of L1 is further
encoded with a LSTM (L2) with self-attention
          <xref ref-type="bibr" rid="ref11 ref23">(Cheng, Dong,
and Lapata 2016)</xref>
          . Single hidden representations of the
tokens are classified through a dense layer with softmax into
IOB labels, which denote whether a word is the Beginning,
the Inside or it is Outside of a frame element span. The
LSTM at L2 is modified so that, at each time step, the
output of the dense layer at t 1 is provided as additional input
to the LSTM cell at time t. The third and final encoding
layer (L3) takes as input the output of L2 and the output of
L1 through highway connections. The same type of encoder
used in L2 is applied in L3, with the difference that the dense
layer outputs frame element labels instead of IOB ones.
        </p>
        <p>
          The simple attention mechanism
          <xref ref-type="bibr" rid="ref3 ref39">(Bahdanau, Cho, and
Bengio 2014)</xref>
          used for the AD task is a layer that gives an
insight of the contribution that a certain input gives in the
production of a given output. The final contextual
representation of a sentence c is evaluated as the weighted sum:
c =
        </p>
        <p>X aihi;</p>
        <p>i
where hi represents the encoding of the i-th token and the
attention value (or score) ai is evaluated through a simple
feedforward network fatt(hi). Roughly speaking, this
attention layer evaluates a value ai for each encoded input token.
Since the AD classification layer operates over the
contextual representation c, each value ai indicates how much each
word in a sentence contributes to the final classification of
a frame. For this reason, it can intrinsically provide an
explanation for the model behavior, as it summarizes a much</p>
        <sec id="sec-2-2-1">
          <title>3http://commoncrawl.org/</title>
          <p>broader set of values that can be more difficult to interpret
(e.g. looking at all the values of the self-learned weights).
In fact, it enables to underline a restricted subset of features,
because not all the inputs have the same importance. The
self-attentions used in the two other layers (L2 and L3)
instead encode the relationship among all the input objects,
e.g. of much each token contributes to the representation of
all the other tokens for a given task. We point the reader to
the original paper for more details about the self-attention
layers.</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Hypotheses and Challenges</title>
        <p>Taking back our research questions, at this point we ask:
how can we better understand the LSTM-based model we
built?
how can we identify an undesired behavior in such model?
is there a way to mitigate any undesired behaviors?
Our first question can be answered looking at the attention
layer values to get hints on the model’s behavior. As
previously discussed, attentions give the chance to explore the
intermediate classification steps, enabling the interpretability
of how the system processes a given input – an aspect that
we can exploit for a better understanding of our process. As a
first attempt, this work aims at answering the previous
questions by taking into account the sole ability of the system
to detect the correct frame. For this reason, we will focus
on the analysis of the attention values for the sole AD task.
We leave the analysis of the other two tasks for forthcoming
work.</p>
        <p>We thus answer our second question by aligning
attentions and the linguistic theory. On the one hand, we have the
Frame Semantics theory that states that frames are evoked
in natural language by specific words called lexical units.
On the other, we have the attention values computed by the
network to balance the input words in the final contextual
representation used to classify the frame. Our assumption
therefore is that, by annotating data with Frame Semantics,
the algorithm learning from such data should encode
implicitly the theory itself, through an attempt of learning it (or a
good approximation of it). If the network is learning
correctly from the data, we should therefore observe an
alignment between the values produced by the attention of the
AD layer and what is stated by Frame Semantics, e.g. we
should notice relevant values attributed to words that could
possibly be lexical units for the classified frame. Should the
network not follow the underlying Frame Semantic theory,
this could mean not only that the model is only following
patterns statistically evident in the data (and not related to
the theory), but also that an incorrect explanation for its
behavior would be provided if requested. Our challenge is first
to verify whether the words receiving the highest attention
values are the correct lexical units of a classified frame (e.g.
given the sentence “take the book from the table”, the word
take should be given a high attention value).</p>
        <p>
          Finally, we need a mitigation strategy to overcome the
cases where the attention turns out to be focused on the
incorrect lexical element and consequently ensure that the
correct explanation for a decision can be provided. Given that
HuRIC’s annotations are based on Frame Semantics, we
propose to augment the dataset using additional examples from
the FrameNet corpus
          <xref ref-type="bibr" rid="ref5">(Baker, Fillmore, and Lowe 1998)</xref>
          .
Although FrameNet cover a different domain w.r.t HuRIC, i.e.
written vs. spoken language, we believe that, by using a data
augmentation strategy, the algorithm can be driven to rely
on patterns consistent with the theory, and thus to achieve
better generalization.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Approach</title>
      <p>In this section, we show the design of the overall approach,
namely (1) how we align the model to the Frame Semantics
theory; (2) how we use these alignments to identify
misbehavior by the model; and (3) the data augmentation strategy
we use to mitigate the bias in the model.</p>
      <sec id="sec-3-1">
        <title>Aligning Attentions and Linguistic Theory</title>
        <p>As previously explained, the attention values produced by
our 3LSTM-ATT parser during the AD stage can be used
to guess which words in a sentence are more relevant to
the classified frame. We can use these values to attempt an
alignment between words and the linguistic theory, namely
which words are lexical units or, other relevant words such
as prepositions.</p>
        <p>
          The parser has been trained over the previously mentioned
HuRIC dataset, which contains transcriptions of user
commands tagged with Frame Semantics. The annotated frames
generally correspond to actions like taking objects or
moving to a specific position. The dataset contains 585 frame
occurrences over 526 sentences on 16 different frame types for
an average of 36 sentences per frame. The results, obtained
over this dataset through a 5-fold cross validation stratified
on the frame types, are reported in Table 1. Compared to
results of
          <xref ref-type="bibr" rid="ref7">(Bastianelli et al. 2016)</xref>
          (BAS16 henceforth)4, our
parser obtains better results for both the AD and AI tasks.
Differently from BAS16, we can take advantage of the
attention layer in the AD step to understand our system’s behavior
when classifying a frame for a given sentence. As explained,
our assumption is that the word receiving the highest value
from the AD attention layer may be the LU for the classified
frame.
        </p>
        <p>In order to prove such hypothesis, we need to
quantitatively measure the alignment between the attention values
and the “gold” LU for a given frame. Let S = (w1; :::; wm)
be a sentence as a sequence of m words w. Gold LUs are
4Please note that the BAS16 makes use also of perceptual
features, while our parser relies only on linguistic inputs.
available in HuRIC, so let w^i be the gold LU for the i-th
sentence Si5. Let us consider the attention layer as a
(simplified) function fatt(w) that attributes an attention value to
a word w (for clarity, w is a shortcut for the hidden
representation h). The ALU (LU-alignment) measure can be then
calculated as follows:</p>
        <p>ALU =
1 N</p>
        <p>X I(arg max fatt(w) = w^i)
N i=1 w2Si
(1)
where I( ) is the indicator function.</p>
        <p>Although lexical units carry most of the meaning for a
frame, there are still many ambiguous cases, where a verb
alone may evoke different frames. Consider for example the
verb take, which may evoke the frame Bringing, e.g. in the
sentence take the book to the table, or Taking, e.g. take the
book from the table. The meaning in this case is not carried
only by the LU alone, but also by the co-occurrence with
other specific words or syntactic structures. The preposition
to in the first example clearly introduces an argument
representing the destination of a motion (i.e. the GOAL frame
element), helping in choosing the frame Bringing over
Taking for the word take. It is thus legit to think that, in these
cases, part of the attention values should also focus on such
discriminant words.</p>
        <p>We thus designed a second measure, that we call AD
(discriminant alignment), with the aim of taking into account
additional discriminant words in addition to the LU. To this
end, we annotated the discriminant words for each sentence
in the dataset. For each sentence S = (w1; :::; wm), we
created a vector of gold discriminant word indexes vg =
(gd1; :::; gdm) where each gdj 2 f0; 1g is set to 1 if its
position corresponds to a discriminant word in S. Given
the attentions values obtained from the AD layer, we
created a vector of classified discriminant word indexes vc =
(cd1; :::; cdm) where each cdj = I(fatt(wj ) 0:01)6.
Finally, we calculated Precision and Recall over these vectors
the following way:</p>
        <p>P =
R =</p>
        <p>Pm
j=1 I(cdj = gdj = 1)</p>
        <p>Pm</p>
        <p>j=1 cdj
Pm
j=1 I(cdj = gdj = 1)</p>
        <p>Pm
j=1 gdj
(2)
(3)
through which we obtained the F-Measure. The AD was
finally calculated as macro-average over the F-Measure of all
the sentences Si in the dataset.</p>
        <p>The HuRIC!HuRIC row of Table 2 shows the scores
for ALU and AD obtained when training and testing over
HuRIC. As we can see from the 11.17% on ALU and
20.53% on AD values, the model reaches good results on
the AD task (96.33% of F-Measure), but is quite misaligned
5Sentence splitting was applied in order to have 1 frame per
sentence, for the rare HuRIC cases containing more than one frame
per sentence.</p>
        <p>6This threshold was set to filter attention noises. The study how
to properly set this threshold is left for future work.
from the linguistic theory. Indeed, the error analysis we
carried on the attention values reported that the model is
folModel l oFrwamienggolldateFnratmpeaptretedrns, wgethich arethecompdlisehteesly unfrormelatedthteo thedining
Model t hFreamoeryg,oldrathFrearmethpraedn gengLeeUtralizingth-ethe lindisg-heusisticfrto-hmeory tah-se ex-din-ing
HuRICàHuRIC pecTatkeindg. In otThakeinrg words0,.L0U0t1he mo0d.9-e08l con0c.0e-1n8trate0.s0-4i1ts att0e.0-n21tion0.0-11
HuFRNIàCHàuHRuIRCIC onTTraaekkiicnnggurrentETwnatkeoirningrgds tha00t..09a09r14e not 00d..90i00s82crim0in.001a8tive w00..00i40t14h res0p.00e2c1t to0.0011
FNF+NHàuHàuHRuIRCIC theTTaarkkeiinnsggpectivETneatkeirfninrggame; y00..e9998t44, it wa0s.000a2ble t0o.000p1rodu00..c00e0142 the c0.o000r3rect 00
FN+HuàHuRIC claTsaskiinfigcationT.aking 0.984 0 0.001 0.012 0.003 0
Model
Model</p>
        <p>Frame gold Frame pred</p>
        <p>Frame gold Frame pred
HuRICàHuRIC
HuRICàHuRIC</p>
        <p>Taking
Taking</p>
        <p>Taking
Taking
take
taLkUe
0.L0U01
0.001</p>
        <p>An example of such behavior is reported in Figure 3:
while the Taking frame is indeed correctly identified, the
attention values reveal that the model attention falls on the two
words the and red, which do not convey any frame meaning
in this context, while the correct LU take receives only 0,1%
of attention. As an additional proof, a similar sentence with
a different frame, e.g. inspect the red shoes, is classified with
the same frame Taking (instead of Inspecting), with most of
the attention falling again on words the, red.</p>
        <p>A first consideration that can arise from the above
analysis is that linguistic phenomena are not equally represented
in HuRIC (i.e. some frames happen in correspondence of
more frequent, but not necessarily significant, grammatical
patterns), and this lack of representativeness might cause
intrinsic bias. This prevents the model to learn the underlying
linguistic theory, and to generalize from it.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Mitigating the Data Bias</title>
        <p>
          If we hypothesize that our model does not generalize
towards the linguistic theory as it should due to the lack of
representativeness of the dataset, a natural solution is to try
to increase the number of training examples to see if the
alignment measures improve without compromising the
performances. Since HuRIC is tagged with Frame Semantics
following the same scheme as the FrameNet corpus
          <xref ref-type="bibr" rid="ref5">(Baker,
Fillmore, and Lowe 1998)</xref>
          , the first solution at hand to
attenuate the bias with more examples consists in integrating
HuRIC with examples from FrameNet itself. For the purpose
of comparison, we selected only the FrameNet examples
annotated with frames also contained in HuRIC. This selection
resulted in a subset of 6,814 frame examples, for an average
of 425 examples per frame.
        </p>
        <p>Although sharing the same background linguistic theory,
however, the two datasets belong to two different domains,
namely written text vs. spoken commands. This indeed
may lead to a drop in terms of performances. Let us take
the example of FrameNet annotated-sentence for the frame
Taking:
roomIn the late 1870s, he defaulted on a loan from rancher
roo-mArchibald Stewart, so [Stewart]AGENT [took]Taking [the
0- Las Vegas Ranch]THEME [for his own]EXPLANATION .
00 Indeed, the label set of frame elements, and, in general,
th00e variability of the language in FrameNet is, in fact, much
h0igher than HuRIC. On the one hand, this can negatively
contribute to the overall performance, as the complexity of
the task increases. On the other, the network will access
more evidence in terms of theory-related patterns, e.g.
seeing more often the association of the frame Taking with
co-occurring verbs like take, than with other unrelated
words like shoes. Our aim is therefore to reach a good
trade-off between the model’s performance and its degree
of generalization that, in turn, reveals the degree of
understandability (explainability) of its behavior.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments and Results</title>
      <p>In order to support our hypotheses about the mitigation
strategy, we designed two additional experimental settings with
the goal of evaluating the changing in the model behavior:
FN!HuRIC: a model is trained over the full subset of
samples coming from FrameNet, and is tested on the
whole HuRIC dataset;
FN+Hu!HuRIC: the evaluation follows a 5-fold cross
validation. At each validation turn, the training set
consists in FrameNet + 80% of HuRIC, leaving the remaining
20% as test set. The distribution of frames is uniformly
stratified.</p>
      <p>Table 2 presents the results of the ALU and AD for both
configurations. The performances of the semantic parser in
terms of F-Measure for the AD, AI and AC tasks are
reported as well. Please note that HuRIC!HuRIC results
differ from the ones in Table 1, showing performances of the
single tasks in isolation (i.e. each task receives gold
information from the previous steps). Instead, we consider here
the full semantic parsing pipeline.</p>
      <p>It appears clear how the parser performances and the
alignment measure scores are reversed for the two
different settings. The models trained only on FrameNet do not
achieve high performances, reaching only approx. 68% for
the AD task. When the two datasets are combined, an
increase of 19% points is achieved for the same task. This
is still very low when compared to the 96.33% achieved
with HuRIC only. With that said, by looking at the
alignment measure scores, we notice that this drop of
performance comes at the advantage of the model’s explainability.
When trained only on FrameNet, in fact, the ALU and AD
scores reach 93.92% and 84.86% respectively. This confirms
that the AD attention layer is focusing on the relevant words,
hence giving us a hint that the model is correctly learning the
linguistic theory. The introduction of HuRIC to the training
sample helps in raising the parsing performances to
convincing levels, while not deteriorating completely the alignment.
The ALU and AD still drop by 40 and 34 points
respectively, but considering the performances reached by the</p>
      <p>In order to better demonstrate the trade-off between parser
performances and theory alignment, we also perform a
qualitative analysis on the test examples. In Figure 4, we show
the AD tagging and attention values produced by the
different training settings (i.e. HuRIC, FN, FN+Hu) for the
sentence get the dishes from the dining room. When trained
on HuRIC only, the network learns again unwanted patterns
and, although the frame Taking is correctly classified, the
attention mostly falls on the article the, also spreading with
minor values on the rest of the words. By using FrameNet
as training set, the attention falls back to the verb take that
corresponds to the current LU. However, the frame
classification fails, predicting Entering. The correct frame
classification (Taking) with attention values matching the correct
LU and, to a minor extent, the discriminant preposition from
is finally obtained when using a combination of the two
corpora, as in the last row.</p>
      <p>The same behavior can be observed if we consider also
other discriminant words in the sentence. Figure 5 shows
again frame parsing and attentions values over different
sentences. Discriminative words are here reported as well
(DISC). In all the four examples it appears clear that when
the system is trained only over the HuRIC resource, the
attention is unstable, i.e. either it distributes similarly among
more or less relevant words (5a), or more strongly
attending on non-discriminant words at all (5b). In other cases, the
attention indeed does attend on discriminant words, but
either the final frame classification is wrong for a lack of value
on the LU (5c), or, even if the frame is correct, we lose the
dependence of the classification outcome on the LU (5d).</p>
    </sec>
    <sec id="sec-5">
      <title>Related Work</title>
      <p>We divided the related work in three parts: (i) approaches to
enable more explainable deep learning-based applications,
with a particular focus on text classification and attention
methods, (ii) approaches to mitigate bias in data and (iii)
approaches for semantic parsing in the robotics domain.</p>
      <sec id="sec-5-1">
        <title>Explainability for Deep Learning</title>
        <p>
          Explainability for deep learning methods can be divided
in three families. A first family, including perturbation
experiments
          <xref ref-type="bibr" rid="ref39">(Zeiler and Fergus 2014)</xref>
          , saliency
mapbased methods
          <xref ref-type="bibr" rid="ref19 ref2 ref31">(Simonyan, Vedaldi, and Zisserman 2013)</xref>
          ,
LIME
          <xref ref-type="bibr" rid="ref23 ref28">(Ribeiro, Singh, and Guestrin 2016)</xref>
          and influence
functions
          <xref ref-type="bibr" rid="ref20">(Koh and Liang 2017)</xref>
          , relies on methods trying
to identify the relevant features treating the model as a black
box. An approximated model is built by observing
concurrent changes between the input and the output, so that it can
provide simple explanations.
        </p>
        <p>
          A second family of approaches focuses on inspecting the
internal representations and input processing. By observing
the inner parameters (weights of the neural network, or other
latent variables), these methods try to give a meaning to
layers and operations in a bottom-up way (Zhang and Z
          <xref ref-type="bibr" rid="ref18">hu
2018</xref>
          ). For this reason, their application is difficult to scale
for networks with lots of layers and parameters.
        </p>
        <p>A third family consists in the intrinsically explainable</p>
        <p>0.0002
FNàHHuuRRIICCàHuRIBCringing Searching</p>
        <p>Taking
FN+HuàFHNuàRHICuRICBringing SearchBinrignging</p>
        <p>Model Frame gold Frame pred bri-ng it- tLoU th-e side- oDfISC the- bath-tub
FigureH5u:RAICtàteHnuRtiIoCn anBarinlygisngis in reTalakitnigon to bLo0Uth the L-U0 and di0Ds.I0cS0rC0i4minan0.t0-1w92ords (0D.-0I3S4C3) for0-.t9h1e04thre0e.-03t5r2ainin g-0 settings.</p>
        <p>
          HuFRNIàCHàuHRuIRCIC BBrriinnggiinngg FBorlilnogwi ningg 0.00003 00..4060931 00.2.4871053 0.20205 0.01.7086506 0.12607 0.17096 0.001
models, whFicNFh+NHàauHràeuHRucIRCoImCpleBBxrriinnggeiinnnggoughBPtrlaioncginirngegach g o00od per0f.0o00r0-1 00.4.t16e70m753s, (Ado00 mavic0iu.5s309e4t al. 2000.81242)7 pro0p0ose to m00itigate the
bitmioanncleasy,eyrsetFeNpx+raoHcuvtàildyHiunpRgIrCogvoidoBedringhainingretslefvoaBrnrinicngeitnegmrperaest0au.t6ri8eo9nb.eAtwt0t.ee0e1n0n-7 0.3sa0ys0es2dte mcuasttioc0malegros’rirthatm0inagnsdafwtei0trhthaen cinlat0esrsaifictciavteio0uns,erb-oi nthtewrfaitche.a
the inputs and outputs, by learning a salience map between Several data augmentation methods for Generative
Adverthe two other nMeotwdeolrk laFyreamrse, gwolhdicFhramcaenprbede furrotbhoter vispuleaals-e tsaakerial Nettwheorks thamtuugse ima gtoe intenthseity nosrinmk alization,
roized using heat-maps independently from the d o-main con-- tLaUtion, re-s-caling, cr o-pping, flDIiSpCping, a-nd Gau-ssian noise
insidered. VisHuuaRlICaàttHeunRtIiCon (MBrinngiinhg et al.Ta2ki0ng14) has0 been use0d 0.j0e0c04tion w0.e0r1e92presen0.t0e3d43in the0.9c1o0n4tex0t.0o3f52medi0cal image
analin the automFaNtiàcHuIRmICage CBaripngtiinogningFotlaloswking(Xu e0t al. 200.14569;1 0.y47s0i3s (Droz0dzal et a0l..06200618; Hu0et al. 20018; Ro0th et al. 2015).
cYaoputioent aisl. g2FeN0n+1eH6rua)àteHwduh.ReITCrhee, gBaritivtnegeinnngtioann viBnariplnuguientsg pciacntubr0ee aobtseexrtvue0adl 0.b17e7tL3witetelen wk0onrokwhleadsgbee0ebnasdeosn0ef.8o2or2n7mhaocwh0itnoe elxepa0lroniitnaglisgynsmteemntss.
to highlight the area of the picture which most contributed The Knowledge Representation community has mostly
foto generate specific words in the caption. and which can cused on empirically analyzing the effects of data links,
be visualized using heat-maps independently from the do- i.e.
          <xref ref-type="bibr" rid="ref3 ref36 ref39">(Tiddi, d’Aquin, and Motta 2014)</xref>
          uses alignments to
main considered. Self-attentions
          <xref ref-type="bibr" rid="ref1 ref3 ref39">(Bahdanau, Cho, and Ben- quantify bias in datasets pairwise, without suggesting
mitgio 2014)</xref>
          have also been widely applied in many text pro- igation solutions; (Ding et al. 2010) discussed the confusion
cessing tasks, such as Sentiment Analysis
          <xref ref-type="bibr" rid="ref22">(Lin et al. 2017)</xref>
          of provenance and ground truth generated by owl:sameAs
and Question Answering
          <xref ref-type="bibr" rid="ref15">(Hermann et al. 2015)</xref>
          . Visual ex- in the context of bioinformatics datasets;
          <xref ref-type="bibr" rid="ref8">(Beek et al. 2018)</xref>
          planations were used in these cases to explain alignments gathers and fixed erroneous identity statements offering
between the words of the input and output sentences. them in a large-scale dataset.
        </p>
        <p>
          Knowledge bases integrated with deep nets have so far
Bias in Data been used to improve the embedding space at training time
In their work,
          <xref ref-type="bibr" rid="ref41">(Zhao et al. 2017)</xref>
          studies the problem of quan- or to explain the model’s outputs a posteriori (cfr.
          <xref ref-type="bibr" rid="ref16">(Hitzler
tifying gender bias in data and models for multi-label object et al. 2019)</xref>
          for a representative selection). To the best of our
classification and visual semantic role labeling, developing knowledge, our work is the first using an external knowledge
a calibration strategy that introduces frequency-constraints bases aligned to the training corpus to mitigate the bias in a
on the training corpus. In the context of recommender sys- training dataset in the context of deep nets.
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>Semantic Parsing for Robotic Applications</title>
        <p>
          A variety of approaches have been proposed in the last two
decades to create semantic parsers for commands of
virtual and real autonomous agents. With the breakthrough
of statistical models, many machine learning techniques
have been applied to semantically parse robot instructions,
from sequential labelling
          <xref ref-type="bibr" rid="ref21">(Kollar et al. 2010)</xref>
          , Statistical
Machine Translation
          <xref ref-type="bibr" rid="ref9">(Chen and Mooney 2011)</xref>
          ,
learningto-rank
          <xref ref-type="bibr" rid="ref19 ref2">(Kim and Mooney 2013)</xref>
          and probabilistic
graphical models
          <xref ref-type="bibr" rid="ref33">(Tellex et al. 2011)</xref>
          . Statistical methods have
also been applied to induce grammars to parse human
commands into suitable meaning representations as well
          <xref ref-type="bibr" rid="ref19 ref2 ref34">(Artzi
and Zettlemoyer 2013; Thomason et al. 2015)</xref>
          . These
approaches were implemented mostly in discretized
environments, relying on ad-hoc and formulaic representation
formalisms, and often dealing with constrained vocabularies.
Our work, on the contrary, builds upon the idea of relying
on linguistically sound theories of meaning representation,
e.g. Frame Semantics, to bridge between linguistic
knowledge and robot internal representations. We build upon
          <xref ref-type="bibr" rid="ref7">(Bastianelli et al. 2016)</xref>
          to design a parser to identify semantic
frames expressed in robot commands but rely on the
bidirectional LSTM network.
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>In this paper, we have presented an approach relying on
the integration of heterogeneous knowledge sources to
mitigate the biased results of a deep learning-based semantic
parser for Spoken Language Understanding for robots, and
improve the model’s understandability. We discussed how
current models do not necessarily learn the underlying
linguistic theory, but rather focus on unwanted, unexpected
patterns, because of an intrinsic bias induced by the size and
domain-specificity of the training dataset. We showed how
the values of the attention layers of the network can be used
as a clue to analyze and interpret the model’s behavior, as the
classification of frames in our case. Finally, we have provide
evidence that external resources such as FrameNet can help
to reduce the bias in the training data, also guaranteeing the
correct interpretations (or explanations) for the model’s
behavior. While being a preliminary attempt to measure a more
complex phenomenon, our work suggests that the strengths
of both knowledge engineering and machine learning can be
combined to foster the development of more transparent,
understandable intelligent systems.</p>
      <p>The future work will be focused in a first instance on
designing more thorough evaluation schemes to obtain better
quantitative understandings of the model’s behavior.
Secondly, we will focus on identifying the correct balance
between the domain-specific samples and the external ones,
also testing new pairs of datasets if possible. An analysis
carried by gradually combining the samples and showing
how the performances and the explainability measures
behave across several datasets and domain is indeed crucial.
Extending the use of more knowledge bases through their
links (e.g. WordNet, ConceptNet) is another route we wish
to follow. Finally, we will explore the idea of interactive,
symbiotic explanations, where the model can be corrected
through spoken dialogue with the user.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          2014.
          <article-title>De-biasing user preference ratings in recommender systems</article-title>
          . In RecSys 2014 Workshop on Interfaces and
          <article-title>Human Decision Making for Recommender Systems (IntRS</article-title>
          <year>2014</year>
          ),
          <fpage>2</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Artzi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Zettlemoyer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Weakly supervised learning of semantic parsers for mapping instructions to actions</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <fpage>49</fpage>
          -
          <lpage>62</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Bahdanau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ; and Bengio,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          <year>2014</year>
          .
          <article-title>Neural machine translation by jointly learning to align and translate</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>arXiv preprint arXiv:1409</source>
          .
          <fpage>0473</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>C. F.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Fillmore</surname>
            ,
            <given-names>C. J.;</given-names>
          </string-name>
          and
          <string-name>
            <surname>Lowe</surname>
            ,
            <given-names>J. B.</given-names>
          </string-name>
          <year>1998</year>
          .
          <article-title>The Berkeley FrameNet project</article-title>
          .
          <source>In Proceedings of ACL and COLING, Association for Computational Linguistics</source>
          ,
          <fpage>86</fpage>
          -
          <lpage>90</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Bastianelli</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Castellucci</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Croce</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Iocchi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Basili</surname>
          </string-name>
          , R.; and
          <string-name>
            <surname>Nardi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Huric: a human robot interaction corpus</article-title>
          .
          <source>In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14).</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Bastianelli</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Croce</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Vanzo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Basili</surname>
          </string-name>
          , R.; and
          <string-name>
            <surname>Nardi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>A discriminative approach to grounded spoken language understanding in interactive robotics</article-title>
          .
          <source>In Proceedings of the 2016 International Joint Conference on Artificial Intelligence (IJCAI).</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Beek</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ; Raad,
          <string-name>
            <given-names>J.</given-names>
            ;
            <surname>Wielemaker</surname>
          </string-name>
          , J.; and
          <string-name>
            <surname>Van Harmelen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>sameas. cc: The closure of 500m owl: sameas statements</article-title>
          .
          <source>In European Semantic Web Conference</source>
          ,
          <volume>65</volume>
          -
          <fpage>80</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>D. L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Mooney</surname>
            ,
            <given-names>R. J.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Learning to interpret natural language navigation instructions from observations</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>In Proceedings of the 25th AAAI Conference on AI</source>
          ,
          <volume>859</volume>
          -
          <fpage>865</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Cheng</surname>
          </string-name>
          , J.;
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ; and Lapata,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Long shortterm memory-networks for machine reading</article-title>
          .
          <source>In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <fpage>551</fpage>
          -
          <lpage>561</lpage>
          . Austin, Texas: Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          2010.
          <article-title>owl: sameas and linked data: An empirical study</article-title>
          .
          <source>In Proceedings of the Second Web Science Conference.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Drozdzal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Chartrand</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Vorontsov</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Shakeri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <given-names>Di</given-names>
            <surname>Jorio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ;
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ;
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ; and
            <surname>Kadoury</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Learning normalized inputs for iterative estimation in medical image segmentation</article-title>
          .
          <source>Medical image analysis</source>
          <volume>44</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Fillmore</surname>
            ,
            <given-names>C. J.</given-names>
          </string-name>
          <year>1985</year>
          .
          <article-title>Frames and the semantics of understanding</article-title>
          .
          <source>Quaderni di Semantica</source>
          <volume>6</volume>
          (
          <issue>2</issue>
          ):
          <fpage>222</fpage>
          -
          <lpage>254</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Hermann</surname>
            ,
            <given-names>K. M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kocisky</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Grefenstette</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Espeholt</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kay</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ; Suleyman,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ; and Blunsom,
          <string-name>
            <surname>P.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Teaching machines to read and comprehend</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          ,
          <volume>1693</volume>
          -
          <fpage>1701</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          2019.
          <article-title>Neural-symbolic integration and the semantic web</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Semantic</given-names>
            <surname>Web</surname>
          </string-name>
          (Preprint):
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Chung</surname>
            ,
            <given-names>A. G.</given-names>
          </string-name>
          ; Fieguth,
          <string-name>
            <given-names>P.</given-names>
            ;
            <surname>Khalvati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ;
            <surname>Haider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            ; and
            <surname>Wong</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>Prostategan: Mitigating data bias via prostate diffusion imaging synthesis with generative adversarial networks</article-title>
          .
          <source>arXiv preprint arXiv:1811</source>
          .05817.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Mooney</surname>
            ,
            <given-names>R. J.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Adapting discriminative reranking to grounded language learning</article-title>
          .
          <source>In ACL (1)</source>
          ,
          <fpage>218</fpage>
          -
          <lpage>227</lpage>
          . The Association for Computer Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Koh</surname>
            ,
            <given-names>P. W.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>Understanding blackbox predictions via influence functions</article-title>
          .
          <source>arXiv preprint arXiv:1703</source>
          .
          <fpage>04730</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Kollar</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Tellex</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Roy</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Roy</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <year>2010</year>
          .
          <article-title>Toward understanding natural language directions</article-title>
          .
          <source>In Proceedings of the 5th ACM/IEEE</source>
          , HRI '
          <volume>10</volume>
          ,
          <fpage>259</fpage>
          -
          <lpage>266</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Santos</surname>
            ,
            <given-names>C. N. d.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Xiang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ; and Bengio,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>A structured self-attentive sentence embedding</article-title>
          .
          <source>arXiv preprint arXiv:1703</source>
          .
          <fpage>03130</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Lane</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Attention-based recurrent neural network models for joint intent detection and slot filling</article-title>
          .
          <source>In INTERSPEECH</source>
          ,
          <fpage>685</fpage>
          -
          <lpage>689</lpage>
          . ISCA.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Mensio</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bastianelli</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Tiddi</surname>
            ,
            <given-names>I.;</given-names>
          </string-name>
          and Rizzo,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <article-title>A multi-layer lstm-based approach for robot command interaction modeling</article-title>
          .
          <source>Workshop on Language and Robotics</source>
          ,
          <string-name>
            <surname>IROS</surname>
          </string-name>
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Mnih</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Heess</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Graves</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; et al.
          <year>2014</year>
          .
          <article-title>Recurrent models of visual attention</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          ,
          <volume>2204</volume>
          -
          <fpage>2212</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Pennington</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Socher, R.; and
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Glove: Global vectors for word representation</article-title>
          .
          <source>In Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP)</source>
          ,
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Ribeiro</surname>
          </string-name>
          , M. T.;
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Guestrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Why should i trust you?: Explaining the predictions of any classifier</article-title>
          .
          <source>In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining</source>
          ,
          <fpage>1135</fpage>
          -
          <lpage>1144</lpage>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Rosenthal</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Biswas</surname>
            , J.; and Veloso,
            <given-names>M.</given-names>
          </string-name>
          <year>2010</year>
          .
          <article-title>An effective personal mobile robot agent through symbiotic humanrobot interaction</article-title>
          .
          <source>In Proceedings of the 9th International Conference on Autonomous Agents and Multiagent Systems: volume 1-</source>
          Volume
          <volume>1</volume>
          ,
          <fpage>915</fpage>
          -
          <lpage>922</lpage>
          . International Foundation for Autonomous Agents and
          <string-name>
            <given-names>Multiagent</given-names>
            <surname>Systems</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>H. R.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Seff</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Cherry</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ; and Summers,
          <string-name>
            <surname>R. M.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Improving computeraided detection using convolutional neural networks and random view aggregation</article-title>
          .
          <source>IEEE transactions on medical imaging 35</source>
          <volume>(5)</volume>
          :
          <fpage>1170</fpage>
          -
          <lpage>1181</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Simonyan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Vedaldi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <article-title>Deep inside convolutional networks: Visualising image classification models and saliency maps</article-title>
          .
          <source>arXiv preprint arXiv:1312</source>
          .
          <fpage>6034</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>Tellex</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Kollar,
          <string-name>
            <given-names>T.</given-names>
            ;
            <surname>Dickerson</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          ; Walter,
          <string-name>
            <given-names>M.</given-names>
            ;
            <surname>Banerjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Teller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ; and
            <surname>Roy</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          <year>2011</year>
          .
          <article-title>Approaching the symbol grounding problem with probabilistic graphical models</article-title>
          .
          <source>AI</source>
          Magazine
          <volume>32</volume>
          (
          <issue>4</issue>
          ):
          <fpage>64</fpage>
          -
          <lpage>76</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <surname>Thomason</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Zhang,
          <string-name>
            <surname>S.</surname>
          </string-name>
          ; Mooney, R.; and Stone,
          <string-name>
            <surname>P.</surname>
          </string-name>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <article-title>Learning to interpret natural language commands through human-robot dialog</article-title>
          .
          <source>In Proceedings of the 24th International Conference on Artificial Intelligence (IJCAI)</source>
          ,
          <source>IJCAI'15</source>
          ,
          <fpage>1923</fpage>
          -
          <lpage>1929</lpage>
          . AAAI Press.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <surname>Tiddi</surname>
          </string-name>
          , I.;
          <string-name>
            <surname>d'Aquin</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Quantifying the bias in data links</article-title>
          .
          <source>In International Conference on Knowledge Engineering and Knowledge Management</source>
          ,
          <fpage>531</fpage>
          -
          <lpage>546</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ba</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Kiros,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ; Cho,
          <string-name>
            <given-names>K.</given-names>
            ;
            <surname>Courville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ;
            <surname>Salakhudinov</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ; Zemel, R.; and Bengio,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Show, attend and tell: Neural image caption generation with visual attention</article-title>
          .
          <source>In International conference on machine learning</source>
          ,
          <fpage>2048</fpage>
          -
          <lpage>2057</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <surname>You</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Jin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Fang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Image captioning with semantic attention</article-title>
          .
          <source>In Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          ,
          <fpage>4651</fpage>
          -
          <lpage>4659</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <string-name>
            <surname>Zeiler</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. D.</surname>
          </string-name>
          , and
          <string-name>
            <surname>Fergus</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Visualizing and understanding convolutional networks</article-title>
          .
          <source>In European conference on computer vision</source>
          , 818-
          <fpage>833</fpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , Q.-s., and
          <string-name>
            <surname>Zhu</surname>
          </string-name>
          , S.-C.
          <year>2018</year>
          .
          <article-title>Visual interpretability for deep learning: a survey</article-title>
          .
          <source>Frontiers of Information Technology &amp; Electronic Engineering</source>
          <volume>19</volume>
          (
          <issue>1</issue>
          ):
          <fpage>27</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Yatskar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ordonez</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Chang</surname>
          </string-name>
          , K.-W.
          <year>2017</year>
          .
          <article-title>Men also like shopping: Reducing gender bias amplification using corpus-level constraints</article-title>
          .
          <source>arXiv preprint arXiv:1707</source>
          .
          <fpage>09457</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>