<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Criteria for Human-Compatible AI in Two-Player Vision-Language Tasks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cheolho Han</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sang-Woo Lee</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yujung Heo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wooyoung Kang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jaehyun Jun</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Byoung-Tak Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Interdisciplinary Program in Neuroscience, Seoul National University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Computer Science and Engineering, Seoul National University</institution>
        </aff>
      </contrib-group>
      <fpage>28</fpage>
      <lpage>33</lpage>
      <abstract>
        <p>We propose rule-based search systems that outperform not only the state-of-the-art but the human performance, measured in accuracy, in GuessWhat?!, a vision-language game where either of two players can be a human. Although those systems achieve the high accuracy, they do not meet other requirements to be considered as an AI system that communicates effectively with humans. To clarify what they lack, we suggest the use of three criteria to enable effective communication with humans in vision-language tasks. These criteria also apply to other two-player vision-language tasks that require communication with humans, e.g., ReferIt.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Recent advances in computer vision and natural language
processing have led researchers’ attentions to the
intersection of these two areas, vision-language tasks. An initiative
to this kind of task was image description [Kiros et al., 2014;
Vinyals et al., 2015; Xu et al., 2015; Johnson et al., 2016b;
Mao et al., 2016]. In the image description task, an image
is given, and the model is supposed to generate descriptions
or captions on the image. However, generated descriptions
have been difficult to evaluate, and results have not been
directly related to how well the model understand the image.
To show the comprehension of the model, visual question
answering (VQA) was introduced [Antol et al., 2015;
Johnson et al., 2016a; Agrawal et al., 2017; Fukui et al., 2016;
Kim et al., 2016]. In VQA, an image and a question about
the image is given, and the model is supposed to answer the
question. However, the communication occurs one-way, and
the model has the passive role only to answer questions.</p>
      <p>Some vision-language tasks require active bidirectional
communication between two (or possibly more) agents. To
interactively communicate over an image, visual dialogues
were introduced [Lazaridou et al., 2016; Das et al., 2016;
Mao et al., 2016; de Vries et al., 2017; Strub et al., 2017;
Das et al., 2017]. In visual dialogues, an image is given, and
y authors contributed equally.
two (or possibly more) agents communicate over the image.
Visual dialogues involving agents with specific roles or tasks
were mainly studied up to the present.</p>
      <p>ReferIt game [Kazemzadeh et al., 2014] is an example of
visual dialogue. It is a two player game referring to objects
in an image of natural scenes. One player is shown an image
with a target object and has to explain it to distinguish from
the other, where what it says is called the referring expression.
The other player is shown the same image and the referring
expression written by the other player and guesses the target
object.</p>
      <p>GuessWhat?! game is also an example of visual dialogue
(Fig. 1). It contains dialogue on question answering about the
given image over two players. Its goal is to locate an unknown
object in a rich image scene by asking a sequence of
questions. One player is randomly assigned an object in the image
and the other player is to locate the hidden object with a series
of Yes/No questions.</p>
      <p>On the two tasks mentioned above, each player also can
be human or agent. If one player is human and the other is
agent, the agent must generate meaningful dialog with
humans in natural and conversational language about a visual
image. This aspect is crucial to solve the task successfully.
Such tasks have been evaluated in terms of task-specific
performance metrics such as accuracy or success rate of the task.
However, they are not the only criteria to enable efficient
bidirectional communication with humans.</p>
      <p>In this paper, we pose the problem that we need more
criteria to measure and analyze the bidirectional communication
between human and agent other than metrics. To demonstrate
that, we first try to tackle GuessWhat?! game which we
mentioned above and show that our proposed rule-based search
systems outperformed not only state-of-the-art, but the
human performance measure, the success rate of the task. Then,
we suggest the use of some criteria to enable efficient
bidirectional communication between human and agent in
visionlanguage tasks such as GuessWhat?! game.</p>
      <p>The rest of the paper is organized as follows. First we
review related works to vision-language tasks on Section 2 and
we proposed rule-based search systems which outperformed
state-of-the-art performance on Section 3. Then, we suggest
the use of some criteria for measure and analyze the
bidirectional communication between human and agent on Section
4. Finally, we discuss conclusions and future work on Section
5.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <sec id="sec-2-1">
        <title>Image Description</title>
        <p>Automatic image description is a challenging problem that
involves analyzing an image, reasoning contextual
information between existing objects in the image and generating
textual descriptions. It has been first stage on research about
vision-language grounding. [Vinyals et al., 2015] proposed
neural image caption (NIC) generator inspired by advances
in machine translation. They replaced encoding step which
extracts abstract representations of source language using
RNN to using CNN fed into given image. Encouraged by
advances in employing attention in machine translation and
object recognition, attention mechanism is introduced by [Xu et
al., 2015]. The mechanism can attend to salient part of given
image while generating its caption so that demonstrated the
learned alignments correspond very well to human intuition.</p>
        <p>Many previous papers on image description have focused
on describing the entire image. On the other hand, [Johnson
et al., 2016b] address a new task, dense captioning, which
requires a model to predict a set of descriptions of regions of
given image. It is important to understand each object or part
not an entire image for high-level scene understanding. [Mao
et al., 2016] also focused on generating an unambiguous
description of a specific object or region in an image. They
considered both description generation and description
comprehension and jointly modeled both tasks combining CNN with
RNN.</p>
        <p>These models are just passive roles only to generate
description about given image on the task, not bidirectional
communication. And, generated descriptions have been
difficult to evaluate, and results have not been directly related to
how well the model understand the image. Therefore, the task
could be extended for bidirectional communication task such
as ReferIt or GuessWhat?! game and it needs further
consideration for evaluation and analysis.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>ReferIt</title>
        <p>ReferIt [Kazemzadeh et al., 2014; Lazaridou et al., 2016] is
a two player game referring an object in an image of
natural scenes. One player is shown an image with a target object
and has to explain it to distinguish from the other, where what
it says is called the referring expression. The other player is
shown the same image and the referring expression written by
the other player and guesses the target object. To accomplish
this game, both agents should cooperate and learn the
relation between vision and language. [Lazaridou et al., 2016]
designed referIt game between two agents, constitute the
referring expression by a binary vector between two agents,
formulated the game as classification, and solved the
classification problem by neural networks. On the task, these agents
develop their own artificial language from the need to
communicate in order to succeed at the game. It showed some
correlation with human language, but also showed some
mismatches. If one player is agent and the other is human, the
referring expression may not carry the exact meaning and
cause the confusion between the players. In terms of
meaninful vision-language integration between human and agent,
we argue that we need to analyze details of these referring
expressions other than metrics such as success rate of game
and accuracy.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>GuessWhat?!</title>
        <p>GuessWhat?! is a cooperative two-player guessing game
proposed by [de Vries et al., 2017]. The goal of the game is to
locate an unknown object in a rich image scene by asking a
sequence of questions. One player who called Oracle is
randomly assigned an object in the image and the other player
who called the Questioner does not know the object assigned
to the Oracle. The goal of the Questioner is to locate the
hidden object with a series of Yes/No questions which are
answered by the Oracle. If the Questioner selects a right object,
we consider the game successful.</p>
        <p>[de Vries et al., 2017] collect a large-scale
humanplayed GuessWhat?! game dataset consisting of 800K visual
question-answering pairs on 66K images and propose
baseline deep learning model. To solve the proposed task
successfully, [de Vries et al., 2017] suggested that an agent is
required higher-level image understanding, like spatial
reasoning, visual properties, object taxonomy, and interaction.
The authors also proposed that the agent should understand
the relationships between objects and how they are expressed
in natural language. The baseline model consists of 3 parts:
Oracle, Guesser and Question generator. Oracle is a
modelbased a simple neural network, which fed embedded inputs
and classifies answer among Yes/No or N/A. The role of
guesser is to predict one hidden object. It compares
dotproducts of embedded vectors of image, dialogue, and
information of candidate objects and classifies the most probable
object among the candidates. Question generator is to
generate questions reflecting the context of the previous question
answering pairs based on the Hierarchical Recurrent Encoder
Decoder(HRED) model.</p>
        <p>As a follow-up research, [Strub et al., 2017] present
endto-end reinforcement learning optimization for question
generation task to find the correct object efficiently. They define
GuessWhat?! game as a Markov Decision Process: A state xt
is the tokens generated on the dialogue until time t and an
action ut is to select a new word with zero-one reward
depending on the Questioner’s choice. They train the question
generator with policy gradient and obtain about 17% improvement
of accuracy as compared with the baseline model.</p>
        <p>GuessWhat?! game have been measured only by the
success rate of the game. However, as following Section 3, we
show rule-based search system which attains not only
stateof-the-art but also human performance. At the point, we
underline that it is not enough to measure the bidirectional
communications only by the metrics such as success rate of the
game or accuracy and propose criteria for meaningful
evaluation on Section 4.
3
3.1</p>
      </sec>
      <sec id="sec-2-4">
        <title>Methods</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Rule-Based Search Systems</title>
      <p>We constructed rule-based search systems using only the
spatial information of the target object for GuessWhat?!. We can
divide an image evenly into three parts by two vertical lines
and divide each part continuously by horizontal lines or
vertical lines in turns (Fig. 2). Then in a current region of interests
divided by two vertical lines, we say the left of the left-side
region, the right of the right-side region, and N/A of the
middle region. Similarly, we say the top, the bottom, and N/A
when a region is divided by horizontal lines. Through this
protocol or language, we can ask and answer for the location
of the center of the target object. Rule-based search systems
use this simple language to locate the target object. We may
also utilize statistics or the distribution of the spatial
information. For the first turn, we may divide an image not evenly.
We found that the range of the middle side of 0.18 covers 1/3
target objects, and the other sides of 0.82 covers 2/3 target
objects, which means the distribution of target objects was
denser in the evenly-divided middle side. Therefore, we set
the vertical lines to locate at 0.41 and 0.59.</p>
      <p>Given a segmentation model, we can further improve our
system. To explore this case, we stole a look at the
segmentation information of the candidates, which is supposed to
be known when the guesser selects a candidate after a series
of question-answer pairs. Then we can implement the binary
search based on the spatial information of the candidates. If a
segmentation model gives the segmentation similar with that
of candidates given in the dataset, we can get an algorithm
close to the binary search, which is optimal.</p>
      <p>Proposed rule-based search systems do not break the rule
of GuessWhat?! game. The spatial information is commonly
used by humans as appears in the dataset. Moreover, we can
substitute the spatial information with other features or
properties. We can choose a real-valued feature without the point
mass like the area and the color. Then the feature gives
another rule-based search algorithm. This search algorithm is
optimal (for the brightness of the center of the target
object) or near-optimal (given an optimal segmentation model
for the area.) We can also use real-valued features with the
point mass or unordered (nominal or categorical) features by
choosing a feature that evenly divide the candidates, which
gives a near-optimal algorithm.
3.2</p>
      <p>Results
[de Vries et al., 2017] and [Strub et al., 2017] constructed
the question generator by the hierarchical recurrent
encoderencoder (HRED) and recurrent neural network (RNN) with
reinforcement learning, respectively (Table 1). The oracle and
the guesser was trained before interacting with the question
generator. Therefore, the oracle and the guesser do not
benefit from interaction. In contrast, in proposed rule-based search
systems, the oracle and the questioner share a set of strict
rules. Proposed systems outperformed the state-of-the-art
accuracy in two turns and the human accuracy in four or five
turns. This result is remarkable improvement of accuracy
considering that [de Vries et al., 2017] and [Strub et al., 2017]
generated 5 and 8 questions respectively to get their
accuracy and that the minimum, mode, mean, and maximum of
the number of questions by humans are 1, 3, 5.2, and 24,
respectively. Furthermore, we improved the system by tuning
the first division of an image, utilizing statistics on spatial
information (denoted Fine-Tune). To explore the effect of
segmentation, we stole a look at the segmentation information of
the candidates (denoted Segment Info).
4</p>
    </sec>
    <sec id="sec-4">
      <title>More Criteria</title>
      <p>The search system exploits low-level features and strictly
follows a predefined set of rules without considering the
uncertainty or ambiguity. If either the oracle or the questioner is
a human, the search system may not successfully
communicate with the human. If the search system takes the part of
the questioner, then it may explain how it works at the
beginning. If the search system plays the role of the oracle, the
search system is unlikely to take the initiative, so it may not
have a chance to show how it works. To develop more
satisfying systems, we investigate characteristics of the effective
AI system for bidirectional communication with humans in
vision-language tasks. We will present some considerations
in designing such systems. Then we will review some criteria
in vision-language tasks and suggest the use of some criteria
to be adopted in GuessWhat?!.</p>
      <p>To develop effective AI systems for bidirectional
communication with humans in vision-language tasks, we need to
formulate a problem or task first. Many vision-language tasks
have been proposed for this purpose. A task naturally gives a
main set of criteria, called objective functions and constraints
in optimization. However, this main set may not be enough,
then we need to add more criteria. Several criteria in
visionlanguage tasks have been made. We may group these criteria
into subjective, task-specific, or similarity criteria.</p>
      <p>The subjective evaluation by humans is widely performed
in many AI researches as well as vision-language areas.
People observe or interact with systems and then evaluate how
well or alike humans systems behave. The subjective
evaluation is not objective so may not be considered scientific, but
the subjective evaluation is crucial because it tells how
humans actually feel about the system. However, the subjective
evaluation is costly in general.</p>
      <p>Task-specific criteria are essential to show how well the
system performs in each task. In vision-language tasks, some
cross-modal classification or retrieval metrics have been used
including accuracy, median rank (mRank), and precision /
recall at k (P/R@k) [Kent et al., 1955]. These criteria are
measured on data that was collected before constructing the
system, so they cost less than subjective criteria, which require
additional human efforts.</p>
      <p>Similarity criteria evaluate how human-like systems
behave. In language generations like machine translations
and text summarizations, some language similarity metrics
have been used including bilingual evaluation understudy
(BLEU) [Papineni et al., 2002], metric for evaluation of
translation with explicit ordering (METEOR) [Banerjee and
Lavie, 2005], recall-oriented understudy for gisting
evaluation (ROUGE) [Lin, 2004], and consensus-based image
description evaluation (CIDEr) [Vedantam et al., 2015]. These
similarity criteria are calculated systematically, so they cost
less than subjective criteria.</p>
      <p>Adversarial evaluation through neural networks [Bowman
et al., 2015; Kannan and Vinyals, 2017; Li et al., 2017] was
suggested as a similarity criterion. Unlike previous
similarity criteria, the adversarial evaluation is not fixed. Instead, it
changes when a neural network called a discriminator learns
whether the speaker is a human or not. After learning, the
discriminator tells the human from the other for the test data.
Since the discriminator is a neural network, it may catch
various complex patterns which could not be found by other
similarity metrics. However, training the discriminator is
necessary unlike other similarity metrics.</p>
      <p>Beyond the accuracy, we need to adopt more criteria in
GuessWhat?!. Basically, more criteria are considered the
better in evaluating vision-language tasks, which is partially
because we do not have a unique criterion that everyone agrees
with. However, we have limited resources, so we have to
choose some criteria among them. The accuracy is the
obvious task-specific criterion, but people would not be
satisfied only with the accuracy achieved by two AI players. We
need to solve the following constrained optimization
problem to develop systems that can communicate with humans in
GuessWhat?!. Given an objective functional f (the accuracy
in GuessWhat?!), human oracle Oh, and human questioner
Qh, the optimal oracle O and questioner Q are given by
solving
max f (O; Q)
O;Q
s.t. O is compatible with Qh</p>
      <sec id="sec-4-1">
        <title>Q is compatible with Oh</title>
        <p>where, for a threshold t &gt; 0,</p>
      </sec>
      <sec id="sec-4-2">
        <title>O is compatible with Qh</title>
        <p>if f (O; Qh) &gt; (1</p>
      </sec>
      <sec id="sec-4-3">
        <title>Q is compatible with Oh</title>
        <p>if f (Oh; Q) &gt; (1
t) max f (O; Qh)</p>
        <p>O
t) max f (Oh; Q)</p>
        <p>Q
This problem is a joint optimization problem and is difficult
to solve because it involves two optimization functions and
the interaction with human oracles as well as the observation
of human questioners. Instead, a two-phase greedy
optimization was commonly used in previous works. It involves the
observation of human questioners like the previous joint
optimization problem, but it involves only one optimization
function at each phase and the interaction with an oracle model
instead of human oracles.</p>
        <p>O^ = argmax f (O; Qh)</p>
        <p>O
Q^ = argmax f (O^; Q)</p>
        <p>Q
However, O^ and Q^ are marginally optimal, and Q^ may not be
compatible with Oh. We may employ human oracles to
determine whether Q^ is compatible with Oh even if it costs large.
Under the assumption that a questioner Q similar with the
human questioner Qh is compatible with Oh, similarity
criteria may complement or substitute with human evaluation.
Criteria that distinguish the AI system from humans are
necessary for this purpose. The adversarial metric is a promising
criterion among them because it involves the neural network,
which can learn complex patterns, as a discriminator or judge
who determines whether two players are human-like (Fig. 3).</p>
        <p>We reviewed some criteria in vision-language tasks and
suggested the use of some criteria to be adopted in
GuessWhat?!. We grouped the criteria into subjective, task-specific,
or similarity criteria and gave some examples. In
GuessWhat?!, we need more criteria other than the accuracy. If we
are affordable, then we can employ human evaluation.
Otherwise, we may choose similarity criteria. We recommended
the adversarial evaluation as a promising similarity criterion.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We proposed rule-based search systems which used only
the spatial information of the target object for GuessWhat?!
game. It can be regarded as the artificial language between
agents on the GuessWhat?! task. Our rule-based search
system outperformed the state-of-the-art accuracy in two turns
and the human accuracy in four or five turns. In the view of
(1)
(2)
(3)
(4)
the results, we argue that we need to measure the performance
of the system and analyze details of the results with more
concrete criteria not just task-specific metrics such as the success
rate of game and accuracy. We suggested the use of criteria
for bidirectional communication between humans and agents
in vision-language tasks. The adversarial evaluation can be
considered as a promising similarity criterion.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work was supported by the Institute for Information
&amp; Communications Technology Promotion
(2015-0-00310SW.StarLab) and Korea Evaluation Institute of Industrial
Technology (10044009-HRI.MESSI, 10060086-RISF)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Agrawal et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Aishwarya</given-names>
            <surname>Agrawal</surname>
          </string-name>
          , Aniruddha Kembhavi, Dhruv Batra, and
          <string-name>
            <given-names>Devi</given-names>
            <surname>Parikh</surname>
          </string-name>
          .
          <article-title>C-vqa: A compositional split of the visual question answering (vqa) v1. 0 dataset</article-title>
          .
          <source>arXiv preprint arXiv:1704.08243</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Antol et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Stanislaw</given-names>
            <surname>Antol</surname>
          </string-name>
          , Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra,
          <string-name>
            <given-names>C Lawrence</given-names>
            <surname>Zitnick</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Devi</given-names>
            <surname>Parikh</surname>
          </string-name>
          . Vqa:
          <article-title>Visual question answering</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Computer Vision</source>
          , pages
          <fpage>2425</fpage>
          -
          <lpage>2433</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <source>[Banerjee and Lavie</source>
          , 2005]
          <string-name>
            <given-names>Satanjeev</given-names>
            <surname>Banerjee</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alon</given-names>
            <surname>Lavie</surname>
          </string-name>
          .
          <article-title>Meteor: An automatic metric for mt evaluation with improved correlation with human judgments</article-title>
          .
          <source>In Proceedings of the acl workshop on</source>
          intrinsic and
          <article-title>extrinsic evaluation measures for machine translation and/or summarization</article-title>
          , volume
          <volume>29</volume>
          , pages
          <fpage>65</fpage>
          -
          <lpage>72</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Bowman et al.,
          <year>2015</year>
          ] Samuel R Bowman, Luke Vilnis, Oriol Vinyals, Andrew M Dai,
          <string-name>
            <given-names>Rafal</given-names>
            <surname>Jozefowicz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Samy</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <article-title>Generating sentences from a continuous space</article-title>
          .
          <source>arXiv preprint arXiv:1511.06349</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>[Das</surname>
          </string-name>
          et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Abhishek</given-names>
            <surname>Das</surname>
          </string-name>
          ,
          <string-name>
            <surname>Satwik Kottur</surname>
            , Khushi Gupta, Avi Singh,
            <given-names>Deshraj Yadav</given-names>
          </string-name>
          , Jose´ MF Moura, Devi Parikh, and
          <string-name>
            <given-names>Dhruv</given-names>
            <surname>Batra</surname>
          </string-name>
          .
          <article-title>Visual dialog</article-title>
          .
          <source>arXiv preprint arXiv:1611.08669</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>[Das</surname>
          </string-name>
          et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Abhishek</given-names>
            <surname>Das</surname>
          </string-name>
          ,
          <string-name>
            <surname>Satwik Kottur</surname>
            , Jose´ MF Moura,
            <given-names>Stefan</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>and Dhruv</given-names>
          </string-name>
          <string-name>
            <surname>Batra</surname>
          </string-name>
          .
          <article-title>Learning cooperative visual dialog agents with deep reinforcement learning</article-title>
          .
          <source>arXiv preprint arXiv:1703.06585</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>[de Vries</surname>
          </string-name>
          et al.,
          <year>2017</year>
          ] Harm de Vries, Florian Strub, Sarath Chandar, Olivier Pietquin, Hugo Larochelle, and
          <string-name>
            <given-names>Aaron</given-names>
            <surname>Courville</surname>
          </string-name>
          . Guesswhat?
          <article-title>! visual object discovery through multi-modal dialogue</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [Fukui et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Akira</given-names>
            <surname>Fukui</surname>
          </string-name>
          , Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and
          <string-name>
            <given-names>Marcus</given-names>
            <surname>Rohrbach</surname>
          </string-name>
          .
          <article-title>Multimodal compact bilinear pooling for visual question answering and visual grounding</article-title>
          .
          <source>arXiv preprint arXiv:1606</source>
          .
          <year>01847</year>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [Johnson et al., 2016a] Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li
          <string-name>
            <surname>Fei-Fei</surname>
            ,
            <given-names>C Lawrence</given-names>
          </string-name>
          <string-name>
            <surname>Zitnick</surname>
            , and
            <given-names>Ross</given-names>
          </string-name>
          <string-name>
            <surname>Girshick</surname>
          </string-name>
          .
          <article-title>Clevr: A diagnostic dataset for compositional language and elementary visual reasoning</article-title>
          .
          <source>arXiv preprint arXiv:1612.06890</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Johnson et al., 2016b] Justin Johnson, Andrej Karpathy, and
          <string-name>
            <surname>Li</surname>
          </string-name>
          Fei-Fei.
          <article-title>Densecap: Fully convolutional localization networks for dense captioning</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pages
          <fpage>4565</fpage>
          -
          <lpage>4574</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>[Kannan and Vinyals</source>
          , 2017]
          <string-name>
            <given-names>Anjuli</given-names>
            <surname>Kannan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Oriol</given-names>
            <surname>Vinyals</surname>
          </string-name>
          .
          <article-title>Adversarial evaluation of dialogue models</article-title>
          .
          <source>arXiv preprint arXiv:1701.08198</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Kazemzadeh et al.,
          <year>2014</year>
          ]
          <string-name>
            <given-names>Sahar</given-names>
            <surname>Kazemzadeh</surname>
          </string-name>
          , Vicente Ordonez,
          <string-name>
            <given-names>Mark</given-names>
            <surname>Matten</surname>
          </string-name>
          , and Tamara L Berg.
          <article-title>Referitgame: Referring to objects in photographs of natural scenes</article-title>
          .
          <source>In EMNLP</source>
          , pages
          <fpage>787</fpage>
          -
          <lpage>798</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [Kent et al.,
          <year>1955</year>
          ]
          <string-name>
            <given-names>Allen</given-names>
            <surname>Kent</surname>
          </string-name>
          , Madeline M Berry, Fred U Luehrs, and James W Perry.
          <article-title>Machine literature searching viii. operational criteria for designing information retrieval systems</article-title>
          .
          <source>Journal of the Association for Information Science and Technology</source>
          ,
          <volume>6</volume>
          (
          <issue>2</issue>
          ):
          <fpage>93</fpage>
          -
          <lpage>101</lpage>
          ,
          <year>1955</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [Kim et al.,
          <year>2016</year>
          ]
          <string-name>
            <surname>Jin-Hwa</surname>
            <given-names>Kim</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sang-Woo</surname>
            <given-names>Lee</given-names>
          </string-name>
          , Donghyun Kwak,
          <string-name>
            <surname>Min-Oh</surname>
            <given-names>Heo</given-names>
          </string-name>
          , Jeonghee Kim,
          <string-name>
            <surname>Jung-Woo Ha</surname>
          </string-name>
          , and
          <string-name>
            <surname>Byoung-Tak Zhang</surname>
          </string-name>
          .
          <article-title>Multimodal residual learning for visual qa</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          , pages
          <fpage>361</fpage>
          -
          <lpage>369</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Kiros et al.,
          <year>2014</year>
          ]
          <string-name>
            <given-names>Ryan</given-names>
            <surname>Kiros</surname>
          </string-name>
          , Ruslan Salakhutdinov, and Richard S Zemel.
          <article-title>Multimodal neural language models</article-title>
          .
          <source>In Icml</source>
          , volume
          <volume>14</volume>
          , pages
          <fpage>595</fpage>
          -
          <lpage>603</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [Lazaridou et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Angeliki</given-names>
            <surname>Lazaridou</surname>
          </string-name>
          , Nghia The Pham, and
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          .
          <article-title>Towards multi-agent communication-based language learning</article-title>
          .
          <source>arXiv preprint arXiv:1605.07133</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>[Li</surname>
          </string-name>
          et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Jiwei</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Will</given-names>
            <surname>Monroe</surname>
          </string-name>
          , Tianlin Shi,
          <string-name>
            <given-names>Alan</given-names>
            <surname>Ritter</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Dan</given-names>
            <surname>Jurafsky</surname>
          </string-name>
          .
          <article-title>Adversarial learning for neural dialogue generation</article-title>
          .
          <source>arXiv preprint arXiv:1701.06547</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <source>[Lin</source>
          ,
          <year>2004</year>
          ]
          <string-name>
            <surname>Chin-Yew Lin</surname>
          </string-name>
          .
          <article-title>Rouge: A package for automatic evaluation of summaries</article-title>
          .
          <source>In Text summarization branches out: Proceedings of the ACL-04 workshop</source>
          , volume
          <volume>8</volume>
          .
          <string-name>
            <surname>Barcelona</surname>
          </string-name>
          , Spain,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [Mao et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Junhua</given-names>
            <surname>Mao</surname>
          </string-name>
          , Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille,
          <article-title>and Kevin Murphy. Generation and comprehension of unambiguous object descriptions</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pages
          <fpage>11</fpage>
          -
          <lpage>20</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [Papineni et al.,
          <year>2002</year>
          ]
          <string-name>
            <given-names>Kishore</given-names>
            <surname>Papineni</surname>
          </string-name>
          , Salim Roukos, Todd Ward, and
          <string-name>
            <surname>Wei-Jing Zhu</surname>
          </string-name>
          .
          <article-title>Bleu: a method for automatic evaluation of machine translation</article-title>
          .
          <source>In Proceedings of the 40th annual meeting on association for computational linguistics</source>
          , pages
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          . Association for Computational Linguistics,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [Strub et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Florian</given-names>
            <surname>Strub</surname>
          </string-name>
          , Harm de Vries, Jeremie Mary, Bilal Piot, Aaron Courville, and Olivier Pietquin.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <article-title>End-to-end optimization of goal-driven and visually grounded dialogue systems</article-title>
          .
          <source>arXiv preprint arXiv:1703.05423</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [Vedantam et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Ramakrishna</given-names>
            <surname>Vedantam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C Lawrence</given-names>
            <surname>Zitnick</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Devi</given-names>
            <surname>Parikh</surname>
          </string-name>
          . Cider:
          <article-title>Consensusbased image description evaluation</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pages
          <fpage>4566</fpage>
          -
          <lpage>4575</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [Vinyals et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Oriol</given-names>
            <surname>Vinyals</surname>
          </string-name>
          , Alexander Toshev, Samy Bengio, and
          <string-name>
            <given-names>Dumitru</given-names>
            <surname>Erhan</surname>
          </string-name>
          .
          <article-title>Show and tell: A neural image caption generator</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          , pages
          <fpage>3156</fpage>
          -
          <lpage>3164</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [Xu et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Kelvin</given-names>
            <surname>Xu</surname>
          </string-name>
          , Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C Courville, Ruslan Salakhutdinov, Richard S Zemel, and
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <article-title>Show, attend and tell: Neural image caption generation with visual attention</article-title>
          .
          <source>In ICML</source>
          , volume
          <volume>14</volume>
          , pages
          <fpage>77</fpage>
          -
          <lpage>81</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>