<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Visually-Grounded Dialogue Models: Past, Present, and Future</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Raquel Fernandez</string-name>
          <email>raquel.fernandez@uva.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Amsterdam</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The past few years have seen an increasing interest in developing
neuralnetwork-based agents for visually-grounded dialogue, where the conversation
participants communicate about visual content. I will start by discussing how
visual grounding can be integrated with traditional task-oriented dialogue
system components. Most current work in the eld focuses on reporting
numeric results solely based on task success. I will argue that we can gain
more insight by (i) analysing the linguistic output of alternative systems
and (ii) probing the representations they learn. I will also introduce a new
dialogue dataset we have developed using a data-collection setup designed
to investigate linguistic common ground as it accumulates during
visuallygrounded interaction.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>